POPULARITY
Categories
YC's Nemil Dalal joins to explain why he's never been more bullish as BitMEX winds down after 11 years, whether every failed crypto idea (TCRs, DAOs, creator coins) eventually works, why crypto is really about money, Base's consumer mea culpa, on-chain reputation and credit, and who pays in the x402 AI-agent era. Welcome to The Chopping Block – where crypto insiders Haseeb Qureshi, Tom Schmidt, Tarun Chitra, and Robert Leshner chop it up about the latest in crypto. This week they're joined by Nemil Dalal, Visiting Partner at Y Combinator and ex-Coinbase, where he led USDC and the Coinbase Developer Platform. He's here to explain why, with exchanges winding down left and right, he's somehow never been more bullish. The crew digs into the great contrast of the moment: BitMEX shutting down after 11 years (plus BitMart, Movement Labs, Balancer Labs) while the plumbing quietly prints, and whether Imran's viral 'everything that failed will eventually work' thesis is genius or toxic positivity. From there it's the question of whether crypto is really only about money (Jesse's Base mea culpa included), a war-memories tour through TCRs, on-chain reputation and why pure on-chain credit keeps faceplanting, and finally who actually pays in the x402 AI-agent era, and whether decentralization even survives contact with Google-shaped gravity. Listen to the episode on Apple Podcasts, Spotify, Pods, Fountain, Podcast Addict, Pocket Casts, Amazon Music, or on your favorite podcast platform. Show highlights
YC's Nemil Dalal joins to explain why he's never been more bullish as BitMEX winds down after 11 years, whether every failed crypto idea (TCRs, DAOs, creator coins) eventually works, why crypto is really about money, Base's consumer mea culpa, on-chain reputation and credit, and who pays in the x402 AI-agent era. Welcome to The Chopping Block – where crypto insiders Haseeb Qureshi, Tom Schmidt, Tarun Chitra, and Robert Leshner chop it up about the latest in crypto. This week they're joined by Nemil Dalal, Visiting Partner at Y Combinator and ex-Coinbase, where he led USDC and the Coinbase Developer Platform. He's here to explain why, with exchanges winding down left and right, he's somehow never been more bullish. The crew digs into the great contrast of the moment: BitMEX shutting down after 11 years (plus BitMart, Movement Labs, Balancer Labs) while the plumbing quietly prints, and whether Imran's viral 'everything that failed will eventually work' thesis is genius or toxic positivity. From there it's the question of whether crypto is really only about money (Jesse's Base mea culpa included), a war-memories tour through TCRs, on-chain reputation and why pure on-chain credit keeps faceplanting, and finally who actually pays in the x402 AI-agent era, and whether decentralization even survives contact with Google-shaped gravity. Listen to the episode on Apple Podcasts, Spotify, Pods, Fountain, Podcast Addict, Pocket Casts, Amazon Music, or on your favorite podcast platform. Show highlights
Update : la Forward Data Conference aura lieu le 16 novembre cette année :)Christophe Blefari (alias Blef), fondateur de la newsletter blef.fr et cofondateur de nao, revient sur son expérience chez Y Combinator, l'incubateur américain le plus prestigieux au monde (Airbnb, Stripe…).Pendant 3 mois au cœur de la Silicon Valley, il a vécu l'intensité de l'entrepreneuriat US et rencontré des dizaines d'équipes Data et Tech.On aborde :
Dans cet épisode, j'interroge Max Berthelot, co-fondateur de Lucis, sur la création de sa startup en santé préventive: le bilan santé qui vous accompagne dans le changements des habitudes de vie (nutrition, exercice, sommeil, compléments, santé mentale), mais aussi sur leur passage pas Y Combinator: l'incubateur célèbre qui a fait naître Airbnb, Stripe et Dropbox.Quel est ce bilan santé et combien ça coûte?Pourquoi c'est mieux que de donner sa prise de sang à Claude / ChatGPT?Comment ça change réellement la santé de ses clients?Les européens sont ils prêts à payer pour la santé préventive?Comment intégrer le meilleur incubateur du monde?Que fait YC pour que ces boîtes avancent aussi vite?Le site de LucisMax sur LinkedinAnastasia sur LinkedinCarnets d'Entrepreneurs sur InstagramHébergé par Ausha. Visitez ausha.co/politique-de-confidentialite pour plus d'informations.
In this episode, Patrick McKenzie (patio11) is joined by Leila Clark, founder of Stardrift and formerly a software engineer at Jane Street, to discuss her essay "What Are You Getting Paid In?" They cover how a Jane Street manager staffed the firm's least lucrative corner by paying people in culture, why Ken Griffin is worth roughly fifty Taylor Swifts, and how a longtime Google product manager quietly ends up with Grammy-winner money. The conversation ranges from academia's brutal tournament structure and YC as a script to founderdom to the one currency Taylor Swift holds that Ken Griffin's $50 billion buys only awkwardly: a restaurant reservation anywhere in New York.–Full transcript available here: https://www.complexsystemspodcast.com/what-youre-actually-getting-paid-in-with-leila-clark/ –Presenting Sponsors: Mercury, Chainguard & MongoDB Complex Systems is presented by Mercury—radically better banking for founders. Mercury's new feature Command brings an LLM directly into your banking interface, so checking balances, finding invoices, or sending a wire is as easy as asking. Apply online in minutes at https://mercury.com/.If attackers are using AI to weaponize code faster than any team can review it, your scanners won't save you. Chainguard builds libraries and container images from source, verified all the way down, with near-zero CVEs and zero malware. Build safely at https://www.chainguard.dev/.What's the point of building faster with AI if your database can't keep up? MongoDB's native data model mirrors the language LLMs already speak. Ship at the speed of AI while staying ACID compliant at Fortune 500 scale. Start building at https://mongodb.com/ai.–Links:What are you getting paid in: https://www.approachwithalacrity.com/p/what-are-you-getting-paid-in –Timestamps:(00:00) Intro(01:23) What are you getting paid in?(03:16) Scripts, teacher figures, and hitting capitalism(08:18) The academia trap(11:08) Paying people in culture: Jane Street's back office(14:03) Ken Griffin vs. Taylor Swift(16:53) Learning about finance by osmosis (or not)(20:44) Sponsors: Mercury | Chainguard(23:46) Musician money vs. Google money(30:03) Keeping up with the Joneses and the meritocratic ladder(33:12) YC as a script to founderdom(36:19) What entrepreneurship pays you in(39:12) Gratitude, culture, and quirky preferences(40:11) Sponsor: MongoDB(43:48) Does winning this script look like winning to you?(48:29) Don't end the week with nothing(51:06) Fame and living in bubbles(56:57) Restaurant reservations as a currency(1:02:57) Where to find Leila(01:03:35) Wrap
In this episode, executive coach Vladimir Baranov explains how founders and tech entrepreneurs can bridge the common gap between technical expertise and leadership skills in the tech and science sectors. He discusses the importance of developing leaders within organizations by building "T-shaped" skills—combining deep technical knowledge with the broad communication skills needed to partner effectively across teams. While navigating the mindset shift of losing an established support structure, founders can build stronger teams by practicing the art of human communication, building their networks early, and understanding their true capabilities Episode Resources: Human Interfaces About Our Guest Vladimir brings over 20 years of start-up and technology experience across fin-tech, wealth-tech, and space-tech, with founder and executive roles that have shaped a coaching approach grounded in real-world insight. He has advised 250+ individuals and 100+ companies, and regularly mentors Techstars and YC companies. His clients describe him as insightful, practical, emotionally intelligent, and results-focused , and his mixed approach blends targeted questions with hands-on start-up experience to help clients build stronger decision-making skills, rather than simply learning from his story. About Our Sponsors Navy Federal Credit Union If you're looking for a positive sign toward homeownership, this is it. That's because Navy Federal Credit Union's Homebuyers Choice loan has downpayment options as low as zero percent with no required private mortgage insurance. These benefits make homeownership more achievable for their members. Learn more here. Terms and conditions apply. Loans subject to approval and eligibility requirements. Learn more at Navy Federal dot org slash zero down. At Navy Federal, our members are the mission. Join the conversation on Facebook! Check out Veteran on the Move on Facebook to connect with our guests and other listeners. A place where you can network with other like-minded veterans who are transitioning to entrepreneurship and get updates on people, programs and resources to help you in YOUR transition to entrepreneurship. Want to be our next guest? Send us an email at interview@veteranonthemove.com. Did you love this episode? Leave us a 5-star rating and review! Download Joe Crane's Top 7 Paths to Freedom or get it on your mobile device. Text VETERAN to 38470. Veteran On the Move podcast has published 600 episodes. Our listeners have the opportunity to hear in-depth interviews conducted by host Joe Crane. The podcast features people, programs, and resources to assist veterans in their transition to entrepreneurship. As a result, Veteran On the Move has over 7,000,000 verified downloads through Stitcher Radio, SoundCloud, iTunes and RSS Feed Syndication making it one of the most popular Military Entrepreneur Shows on the Internet Today. Disclosure: Some of the links above are affiliate links. This means that, at zero cost to you, I will earn an affiliate commission if you purchase via the link provided
On this episode of Run the Numbers, CJ sits down with Dave McClure and Aman Verjee of Practical VC to talk about the evolution of startup investing — from the PayPal mafia and Founders Fund era to 500 Startups, YC-style scale, and today's secondary market.—SPONSORS:Rillet is an AI-native ERP built for modern finance teams that want to replace NetSuite and close faster. With revenue recognition, close management, multi-entity support, and native Stripe and Salesforce integrations, Rillet helps scaling companies run their finance stack in one place. Hundreds of teams, including Windsurf and Mercor, use Rillet to make the zero-day close real. Book a demo at https://www.rillet.com/cjMaximor is an autonomous finance platform that runs order-to-cash, procure-to-pay, the close, cash management, and reporting on self-learning agents instead of a dozen disconnected tools. One PE-backed customer cut their close in half, took audit findings from seven to zero, and cut back-office costs by 70% in six months. You pay for outcomes, not seats. See it at https://www.maximor.ai/Brex is an intelligent finance platform with AI-powered agents that capture expenses automatically, enforce policy before the spend happens, and close your books in minutes instead of weeks. 35,000+ companies like OpenAI, Coinbase, Anthropic, and DoorDash already run on Brex. It's time to get Brex AF. Learn more at https://www.brex.com/metricsAnrok is the sales tax platform that watches your exposure everywhere, automates compliance, and flags risk before it turns into a surprise back-tax letter from a state you've never set foot in. Companies like Anthropic, Notion, and Vanta already trust Anrok to stay ahead of rules that move faster than any spreadsheet can. Talk to a sales tax expert for a personalized exposure estimate at https://www.anrok.com/rtnRightRev is an automated revenue recognition platform that lets your product team ship new pricing without asking finance for permission, and your sales team close deals without creating downstream chaos. Check out their free tool at calculator.rightrev.com It scores your rev rec process, shows what's exposing you to risk, and tells you exactly where to focus before it bites you in the rear end. Check it out at https://calculator.rightrev.comPulley is an equity management platform that lets you issue options, model dilution, and complete 409As without your cap table turning into a spreadsheet disaster. Founders raising, hiring, and scaling use Pulley to keep equity clean and stay focused on building. Learn more or request a demo at https://pulley.com/mostlymetrics—LINKS: Mostly Talent: https://mostlymetrics.typeform.com/to/cLTxtAsNGuests:https://www.linkedin.com/in/davemcclure/https://www.linkedin.com/in/aman-verjee/Company:https://practicalvc.com/CJ: https://www.linkedin.com/in/cj-gustafson-13140948/Mostly metrics: https://www.mostlymetrics.com—TIMESTAMPS:0:00 Preview and Intro3:00 Writing as a distribution strategy5:17 How Dave's blog led to Founders Fund7:24 Aman: writing at PayPal and law school9:29 PayPal: the red pen of David Sacks11:44 Sponsors — Rillet | Maximor | Brex14:59 500 Startups: the original thesis16:22 The volume strategy: more shots on goal18:44 Twilio, Lyft, Sendgrid, Credit Karma20:15 60x returns: right place, right time23:07 Sponsors — Anrok | RightRev | Pulley25:58 Accelerator ecosystem evolution28:42 500 vs. YC: scale as a weapon32:52 Globalizing the accelerator model35:04 Seed to secondaries: how it happened37:00 Scratching their own itch for liquidity38:10 The secondary market explained40:50 Types of secondaries43:29 Why would a VC sell a winner?46:07 Managing DPI as a late-stage manager47:02 The five horse framework49:15 Skipping the J curve51:44 Imperfect information as a feature55:22 Landmines: fraud, mismarked valuations57:00 VCs lie about valuations three ways57:58 Forward contracts and counterparty risk1:00:29 Credits
Sponsored by Blocks: Save at least 20% on your AWS costs with AI-powered optimization and enterprise discounts. Get your free Cloud Check at blocks.cloud/alphalist → https://blocks.cloud/alphalist?utm_source=alphalist&utm_medium=podcast&utm_campaign=blocks-podcast-2026 Vaibhav Gupta built computer vision for the original Microsoft HoloLens, optimized AR at Google, and wrote high-performance assembly at D.E. Shaw, then left it all to start from scratch. After a YC pivot away from a Slack competitor he was told not to build, he landed on something foundational: BAML, a programming language for a world where humans increasingly don't read code. His thesis: every software leap came from a new compute paradigm getting its own language assembly, C, Java, JavaScript and LLMs are the next primitive. They're probabilistic and non-deterministic, which breaks our deterministic tooling. In this episode, Vaibhav explains why "shipping at agent speed" is really a problem of trust and control, why 90% of engineering is plumbing AI will delete, why "English as a programming language" can't work, and why the world has a mathematically infinite appetite for software. Topics covered: - Why LLMs are a new compute primitive and why that justifies a new language - BAML: an embedded, type-safe language for structured LLM outputs across any language - Shipping at agent speed as a problem of trust, locking, and granular control - Why traditional CI/CD breaks in an agent loop - The "data trench" one type system across code, backend, and data - Why 90% of engineering is plumbing, and what changes when AI removes it - Where SaaS pricing and product models are heading
The Facebook moment just hit enterprise AI. Did you miss the terms of service update?In this episode of KP Unpacked, KP Reddy and Nick break down why the Alex Karp CNBC interview landed like a bomb in enterprise boardrooms but barely surprised anyone actually building with AI. Construction company CEOs were getting texts from board members within hours: "Did you see this? What are we doing?" The answer for most of them? Running Microsoft Copilot, banning Claude, and quietly dealing with ransomware attacks that have already put subcontractors out of business.KP walks through the apple pie analogy: buying a pre-made pie (frontier models) is cheaper, faster, consistent. Making your own (open source) costs more, takes longer, outcome uncertain. But here's the real insight: the question isn't open source versus frontier models. It's what data should you never feed any model, period. Then a Bay Area contractor says something that cuts through all the noise: these YC kids have no construction experience, no relationships, no reputation. If they take our data and screw it up, they move on to their next startup. What do they have to lose? That's not a technology question. That's a trust question. And construction figured out the answer decades ago when they started vetting subcontractors.Key questions answered:Why did the Karp CNBC interview send board members texting their CEOs?What does "we're training on your data" actually mean legally and technically?Is the apple pie analogy the best way to explain open source versus frontier models?Why is ransomware quietly killing subcontractors before AI even arrives?What does a contractor's subcontractor vetting process teach us about evaluating AI startups?If a YC startup takes your data and folds, what does the founder have to lose?Why did Claude's updated terms of service change the conversation?What's the Red Hat playbook and why is Palantir following it?How does Zero RFI's opt-in audit trail architecture solve the data trust problem?Why are most AEC firms staying on Microsoft Copilot and not moving anywhere fast?Should startups building on frontier models be worried about their defensibility?Why does Karp get away with saying things every other public company CEO won't?If you're a construction company trying to figure out what to tell your board after the Karp interview, a startup wondering how to build trust with enterprise clients around data, or an executive who just realized you never actually read those terms of service, this episode will help you figure out what you actually agreed to and what to do next.Listen now.
Welcome to the Firearms Insider Gun & Gear Review Podcast episode 633. This episode is brought to you by Walker Defense, XS Sights, Hi-Point Firearms, and CMC Triggers. In this show we will be discussing a prism review and discuss a Colt red dot, a new ported double stack, an Esee folder, and something interesting and probably terrible from Bear Creek As you may know, we showcase guns, gear, and anything else you might be interested in. We do our best to evaluate products from an unbiased and honest perspective. I'm Chad Wallace, host of the most dedicated firearms podcast around With me tonight are: Tony, Rob, Rusty Sponsor #1: Hi-Point Hi-Point firearms has been crafting American made firearms for over 30 years. If you are looking for your first firearm, or just want something fun for the range, Hi-Point has you covered with models including handguns, pistol caliber carbines, and AR15's. They even have a new suppressor line. Hi-Point firearms can be found at extremely affordable prices, making them available for anyone that wants to protect themselves and/or their families. Every Hi-Point also comes with a lifetime warranty and most of their products are 50 state legal. Hi-Point Firearms, made by the American working man for the American working man. Our Hi-Point Product of the week is - YC 380 Visit hi-pointfirearms.com and check out their line of products Use code “GGR” FOR $20 off a Hi-Point firearm at ShootAmmo.com What we did in Firearms: Announcements: Kat's Rack Defense fund https://www.givesendgo.com/Katsrackdefensefund Bandwidth sponsor Patriot Patch Co. And their Patch of the Month Club! Check out the Pew.Report T-shirts are available through our FRN site, or click the “Merch” tab on Firearmsinsider.tv AFFILIATES / DISCOUNTS: Walker Defense Research - enter “INSIDER15” for 15% off XS Sights - “GGR20” for 20% off Hi-Point - “GGR” FOR $20 off a Hi-Point firearm at ShootAmmo.com CMC Triggers - “GGR26” for 10% off Primary Arms VZ Grips Brownells Gun Guys Garage discount code - “FRN15OFF” Atibal Optics - enter “FIREARMSINSIDER20” for 20% off 5.11 Tactical PowerTac Lights - enter “GGR” for a real good discount Modern Spartan Systems - “GGR15” for 15% off Global Ordnance Infinite Defense (Infinity Targets) - “PEW15” for 15% off Guns.com Magpul Palmetto State Armory Unique ARs - “GunGearReview” for 10% off CobraTec Knives - “GGR10” for 10% off Nutrient Survival - “GGR10” for 10% off Gideon Optics - “GGR” or “INSIDER” for 10% off US Optics - “INSIDER15” for 15% off Camorado - “FIREARMSINSIDER” for 5% off Optics Planet Midway USA Strike Industries North Forest Arms - “GGR” for 10% off Kini SafeAlert - “GGR” for 20% off FoxTrot Mike - “GGR” for 10% off XTech Tactical - “GGR10” for 10% off Die Free Co ZeroTech Optics - “GGR” for 20% off Goliath Defense - “GGR” for 10% off holsters Classic Firearms True Shot Ammo Next Level Armament NightStick ROB - Disclaimer The views and opinions expressed in this podcast are those of the individual co-hosts and do not reflect the official policy or position of the Firearms Radio Network and/or their employers. This is NOT legal advice, nor should it be considered as such. Viewer discretion is advised. Main Topic is sponsored by: XS Sights For over 25 years, XS Sights has helped you get on target faster. Offering tritium sights in all different types and styles, low light is no longer an obstacle. Most options come with a brightly colored photoluminescent ring around the tritium. That colored ring makes them work great in the daylight also. XS Sights has sight styles for everyone: Big Dot's, Ghost Rings, Standard Notch and Post, Minimalist, Suppressor Height, all offering tritium options. Available for a plethora of firearms types, from shotguns to handguns, XS sights has you covered for all your low light sighting needs. Our XS Sights Product of the week is - The brand new Smith & Wesson Fiber Optic Revolver Sights Use Code “GGR20” for 20% off of almost everything at xssights.com Main Topic: Product Review Chad - ZeroTech Thrive HD 1x Micro Prism Product Spotlight and Discussion: Colt Optics MRS-1 Red Dot MSRP - $428.000 Military Armament Corporation Cerberus Ported MSRP - $899.00 Sponsor #3: Walker Defense Research Walker Defense provides shooters with the finest, most innovative, quality, tactical accessories and firearm components around. From their NILE grip panels to their NERO muzzle brakes, no details are ever left behind. Only top quality materials are used in the manufacturing process. Together, all of this gives you some of the best firearm performance around. Everything they have to offer is proudly made in the USA. Walker Defense, where American ingenuity meets bleeding edge technology. Our Walker Defense Product of the week is - Flat Dark Earth DLC Bolt Carrier Group Use code “INSIDER15” FOR 15% OFF everything at walkerdr.com Bear Creek Grizzly compact 380 MSRP - $349.99 Esee Knives Expat Medellin II Retail - $69.99 Sponsor #4: CMC Triggers CMC triggers is the creator of the original, drop in cartridge style AR trigger. CMC has been in aerospace manufacturing for over 30 years. This gives you peace of mind knowing that every trigger is of the utmost quality. Their patented design ensures a crisp, short, trigger pull across their whole line of triggers. With triggers for AR's, AK's, pistols, and bolt guns, CMC can make your firearms trigger great. Proudly made in Texas with strong morals and values, giving you confidence that you are buying one of the best triggers out there. Choose Confidence, choose quality, choose CMC Our CMC Product of the week is - AR-15/10 Single Stage Curved Component Trigger Use Code “GGR26” for 10% off at CMCTriggers.com Listener Feedback None 2nd is for Everyone Diversity Shoot Events simonsaystrain on instagram 2nd is for Everyone Facebook 2A4E Web Page Wrap up: Send questions, comments, or feedback to - gungearreview@gmail.com Remember to Subscribe and Leave us an iTunes Review Be sure to visit the Firearms Insider at www.firearmsinsider.tv Check us out on Facebook, X, and InstaGram @firearmsinsider Subscribe to our Rumble channel Please check out all our great sponsors Thank you for listening to the “LARGEST”, pound for pound, podcast on the network We are out
This is a recap of the top 10 posts on Hacker News on July 09, 2026. This podcast was generated by wondercraft.ai (00:30): GPT-5.6Original post: https://news.ycombinator.com/item?id=48849066&utm_source=wondercraft_ai(01:56): EU Parliament greenlights Chat Control 1.0Original post: https://news.ycombinator.com/item?id=48843923&utm_source=wondercraft_ai(03:23): Show HN: 18 WordsOriginal post: https://news.ycombinator.com/item?id=48845049&utm_source=wondercraft_ai(04:50): My thoughts on the Bun Rust rewriteOriginal post: https://news.ycombinator.com/item?id=48843352&utm_source=wondercraft_ai(06:17): Postgres rewritten in Rust, now passing 100% of the Postgres regression testsOriginal post: https://news.ycombinator.com/item?id=48841676&utm_source=wondercraft_ai(07:43): Show HN: Getting GLM 5.2 running on my slow computerOriginal post: https://news.ycombinator.com/item?id=48842459&utm_source=wondercraft_ai(09:10): Hy3Original post: https://news.ycombinator.com/item?id=48847552&utm_source=wondercraft_ai(10:37): I think I have LLM burnoutOriginal post: https://news.ycombinator.com/item?id=48839984&utm_source=wondercraft_ai(12:04): Why developers are ditching GitHub for Codeberg and self-hosting alternativesOriginal post: https://news.ycombinator.com/item?id=48842611&utm_source=wondercraft_ai(13:31): Muse Spark 1.1Original post: https://news.ycombinator.com/item?id=48846184&utm_source=wondercraft_aiThis is a third-party project, independent from HN and YC. Text and audio generated using AI, by wondercraft.ai. Create your own studio quality podcast with text as the only input in seconds at app.wondercraft.ai. Issues or feedback? We'd love to hear from you: team@wondercraft.ai
SED News is a monthly podcast from Software Engineering Daily where hosts Gregor Vand and Sean Falconer break down the biggest stories shaping software engineering, Silicon Valley, and the broader tech industry. In this episode, Gregor and Sean dig into the growing tension around restricted AI models, including Anthropic‘s Fable being pulled from the Claude platform days after launch. They explore what Sean calls “vibe regulations” and the risk foreign governments and enterprises face when a model they depend on can be cut off. They also cover the FT’s reporting on London’s “DeepMind mafia,” a vibe-coding clone controversy involving YC-backed Corgi and Papermark, SpaceX‘s acquisitions of Cursor and Mesh, and Anthropic’s launch of Claude Science. They also take on the latest round of the IDE wars, and explore who owns your dev toolchain, the vendor lock-in that now comes from context and memory rather than the model itself, and the widening cost gap between frontier tools and open weight models. As always, the episode wraps up with a few standout Hacker News threads. The post SED News: Restricted Models, IDE Wars, and the DeepMind Mafia appeared first on Software Engineering Daily.
Every company building AI right now is asking the same question: if the models keep getting better and anyone can access them, what actually makes us defensible? Avi Bharadwaj writes the checks that answer that question. As an Investment Director at Intel Capital, he focuses on the software infrastructure layer of AI, backing companies like Scale AI, Bria, TrueFoundry, and Twelve Labs.In this episode of Talking AI, Avi sits down with Matt Paige to break down exactly where moats are showing up as frontier models commoditize intelligence. He walks through five specific layers of defensibility for application companies (unique data, workflow and system of action, product reimagination, integration, and trust and compliance) and explains why the infrastructure between the model and the application is where most enterprise AI projects actually stall.The conversation covers why building for the gap between what frontier models can and can't do is a losing strategy (because the gap is ever-shrinking), why the chatbot era was brief and agents are now first-class citizens, how Avi uses an agent on Claude Cowork to scan Hacker News and Reddit overnight and enter emerging companies into his CRM by morning, and why he's most excited about world models and the emergent abilities that might come from scaling them.The episode closes with Avi's advice for founders: don't build things that fit the current gap in model capability. Build things that improve as the model improves. And his honest take on being a VC: at best you're a sidekick for founders, at worst you're a detractor.In this episode, you'll hear about:Five layers of defensibility that frontier models can't commoditize. Why unique data, not just more data, is the moat that still matters. The shift from chatbots to deeply embedded agentic workflows in enterprise. How Avi uses Claude Cowork agents to automate deal sourcing and financial analysis. Why specialized foundation models still win in domains like licensed imagery, industrial robotics, and edge inference. The Figma/Claude Design moment and what it means for how VCs underwrite platform risk. Why context engineering is becoming its own discipline and the mistake of treating models like if-else loops. World models, emergent abilities, and what comes after language as an abstraction. How Avi went from Goldman Sachs engineer to IBM data scientist to Intel Capital investor. The coolest and most overrated parts of being a VC.--Key Moments00:01:41 — "It's a mistake to think better models kill moats"00:02:30 — Unique data as the new defensibility: proprietary CRM triggers, healthcare, industrial00:03:25 — Workflow and system of action moats00:04:00 — UX and product reimagination as a moat00:04:30 — Integration moats: 50 to 100 systems upstream and downstream00:05:10 — Trust and compliance as the fifth layer00:05:30 — Infrastructure layer defensibility: evaluation, benchmarking, security, identity00:06:27 — Jack Dorsey's "From Hierarchy to Intelligence" and the YC thesis00:09:55 — From data scientist to frontier model commoditization: what changed00:13:12 — How a VC uses AI: seeing, picking, winning, and supporting00:15:00 — Claude Cowork agent scanning Hacker News, Reddit, and PitchBook overnight00:18:58 — Specialized models vs. the ever-shrinking gap: where do they survive?00:20:30 — Bria's licensed data moat and Field AI's industrial deployment data00:22:45 — "Build things that improve as the model improves"00:24:14 — Why frontier models win bottom-up but can't crack top-down enterprise adoption00:25:43 — The chatbot era was brief: agents are first-class citizens00:27:50 — Memory: session, long-term, and standardized enterprise memory00:31:41 — "Don't use models like a very long if-else statement loop"00:35:08 — World models, emergent abilities, and what comes after language00:38:34 — Robotics: narrow industrial use cases first, Jetsons life in ten years00:41:26 — From Goldman Sachs engineer to IBM data scientist to Intel Capital VC00:43:10 — The coolest and most overrated things about being a VC--Key LinksIntel CapitalConnect with Avi on LinkedInMentioned in this episode:Free report from HatchWorks AI — State of AI 2026What's real in AI this year, what's hype, and what leaders should prioritize — including production lessons, designing for agents, and governance. https://hatchworks.com/state-of-ai-2026/AI Opportunity FinderFeeling overwhelmed by all the AI noise out there? The AI Opportunity Finder from HatchWorks cuts through the hype and gives you a clear starting point. In less than 5 minutes, you'll get tailored, high-impact AI use cases specific to your business—scored by ROI so you know exactly where to start. Whether you're looking to cut costs, automate tasks, or grow faster, this free tool gives you a personalized roadmap built for action.
SED News is a monthly podcast from Software Engineering Daily where hosts Gregor Vand and Sean Falconer break down the biggest stories shaping software engineering, Silicon Valley, and the broader tech industry. In this episode, Gregor and Sean dig into the growing tension around restricted AI models, including Anthropic‘s Fable being pulled from the Claude platform days after launch. They explore what Sean calls “vibe regulations” and the risk foreign governments and enterprises face when a model they depend on can be cut off. They also cover the FT’s reporting on London’s “DeepMind mafia,” a vibe-coding clone controversy involving YC-backed Corgi and Papermark, SpaceX‘s acquisitions of Cursor and Mesh, and Anthropic’s launch of Claude Science. They also take on the latest round of the IDE wars, and explore who owns your dev toolchain, the vendor lock-in that now comes from context and memory rather than the model itself, and the widening cost gap between frontier tools and open weight models. As always, the episode wraps up with a few standout Hacker News threads. The post SED News: Restricted Models, IDE Wars, and the DeepMind Mafia appeared first on Software Engineering Daily.
SED News is a monthly podcast from Software Engineering Daily where hosts Gregor Vand and Sean Falconer break down the biggest stories shaping software engineering, Silicon Valley, and the broader tech industry. In this episode, Gregor and Sean dig into the growing tension around restricted AI models, including Anthropic‘s Fable being pulled from the Claude platform days after launch. They explore what Sean calls “vibe regulations” and the risk foreign governments and enterprises face when a model they depend on can be cut off. They also cover the FT’s reporting on London’s “DeepMind mafia,” a vibe-coding clone controversy involving YC-backed Corgi and Papermark, SpaceX‘s acquisitions of Cursor and Mesh, and Anthropic’s launch of Claude Science. They also take on the latest round of the IDE wars, and explore who owns your dev toolchain, the vendor lock-in that now comes from context and memory rather than the model itself, and the widening cost gap between frontier tools and open weight models. As always, the episode wraps up with a few standout Hacker News threads. The post SED News: Restricted Models, IDE Wars, and the DeepMind Mafia appeared first on Software Engineering Daily.
Today, we're joined by Alex Bouaziz, co-founder and CEO of Deel, the global HR and payroll platform now serving companies across 150+ countries with more than $1 billion in annual revenue.Tommy Stadlen talks to Alex about how Deel became the fastest startup to reach $100M in recurring revenue: and why that record came down to one thing: following customers. Alex explains why Deel launched with contractors when everyone else went after employees, how they carried the intensity of Y Combinator all the way to a billion-dollar business, and why their M&A strategy (13 acquisitions in six years) has been one of their most underrated growth levers.He speaks about:The YC "time chamber" mindset they still run the company onWhy speed of execution matters - and why comfort in any department is a warning signWhy hiring A players is less about finding them, and more about NOT lowering the bar when you're desperate to hireWhat most people miss about the Deel M&A strategy (and how they keep founders motivated years after acquisition)The case for brand marketing at the enterprise stage, and what Arsenal taught him about communityWhy being a remote company selling to remote companies is an unfair advantageAnd lots more!Building a purpose driven company? Read more about Giant Ventures at www.Giant.vc.Music credits: Bubble King written and produced by Cameron McLain and Stevan Cablayan aka Vector_XING.Please note: The content of this podcast is for informational and entertainment purposes only. It should not be considered financial, legal, or investment advice. Always consult a licensed professional before making any investment decisions.
What if teaching kids to complete the Millennium Falcon set is exactly what's making them unprepared for the real world?In this episode of KP Unpacked, KP Reddy and Nick unpack why AI reading drawings is a feature, not a company, why reindustrialization in Detroit changed how KP thinks about hard tech, and why the Lego analogy explains everything wrong with how we raise kids today. Original Legos were a mixed box of bricks with no instructions. You built whatever your imagination created. Modern Lego sets are Millennium Falcons with step-by-step instructions. Kids complete the set, lose their mind when a piece is missing, and never learn creativity. Sound familiar? College degree, job market, no pieces, losing their mind.KP takes that analogy into AI: reading drawings is spell check, not a bestseller. Everyone's building tools to "read plans and specs" and the head of pre-con 10 minutes from YC is telling his team these founders have no idea what they're doing every time they leave. The hard part isn't reading the door on a drawing. It's knowing whether you need three hinges, the right finishes, or the shim dimensions based on decades of inference. Then KP shares takeaways from Detroit's Reindustrialized conference: own your building, run your own machine shop, stop outsourcing prototypes to vendors who put you at the back of the line. Antonio Gracias (early Tesla, SpaceX investor) said it best: stop making three SKUs for mass production. Make 15 form factors, release faster, do more interesting things.Key questions answered:Why is AI reading drawings a feature, not a product or company?What's the difference between object detection and inference in construction drawings?Why does every stakeholder look at the same door on a drawing and see something different?What do original Legos teach kids that Millennium Falcon sets don't?Why are college grads losing their minds when pieces are missing?What should we actually be teaching kids instead of following instruction manuals?What happened at the Reindustrialized conference in Detroit?Why should hard tech founders own their buildings and machine shops?Why does outsourcing prototypes to manufacturers put you at the back of the line?What did Antonio Gracias say about nimble manufacturing versus mass production?Why do fewer SKUs and more frequency matter more than cost efficiency?Why is gaining understanding the actual goal of using AI tools?If you're building an AI drawing reading tool and calling it a company, wondering why hard tech funding requires a completely different playbook, or trying to figure out what creativity and imagination actually mean in an AI world, this episode will challenge every assumption about tools, skills, and what we're really solving for.Listen now.
Russ d'Sa (CEO & Co-founder @ LiveKit) joins the show to deconstruct the "Product Paradigm Shift" toward voice-driven interfaces and agent-centric UX . We dive into LiveKit's high-stakes scaling lessons: from powering OpenAI and Character AI's voice mode, how they navigated real time bottlenecks to hit the next level of scale, the architectural necessity of a multi-cloud strategy, and the foundations of a co-founder relationships that can effectively blend engineering & business strategy. ABOUT RUSS D'SA Russ is a startup vet who founded his first company in the 2007 YC batch and was the 2nd frontend engineer hired at Twitter, Russ d'Sa now leads voice AI unicorn LiveKit. They're the backbone of ChatGPT Voice Mode, Salesforce Agentforce, Grok, and roughly 30% of US 911 calls. ABOUT LIVEKIT LiveKit is an open source framework and cloud platform for building voice, video, and physical AI agents. It provides the tools you need to build agents that interact with users in realtime over audio, video, and data streams. Agents run on the LiveKit server, which supplies the low-latency infrastructure (including transport, routing, synchronization, and session management) built on a production-grade WebRTC stack. This architecture enables reliable and performant agent workloads. SHOW NOTES: The product paradigm shift toward voice-driven apps and natural human-computer interfaces (2:44) Voice-apps in practice: How these trends impact the strategy of product building today (5:32) Early adopters: Why legacy industries like healthcare use voice AI (7:55) Reevaluating and building product experiences optimized for AI agents (12:52) How AI trends will impact roadmaps and Go To Market (18:16) The origin of LiveKit: Building real-time infra for the pandemic (21:07) The OpenAI moment: Powering the fastest-growing consumer app (23:48) Scaling with OpenAI: Navigating the challenges of balancing time-to-market with system design (25:39) The Character AI outage: Solving cross-continental state sync and hitting the next level of scale (29:00) The problem: When telemetry breaks first: Managing analytics and logging for millions of concurrent AI sessions (32:04) Architecting for resilience: Multi-cloud from day one and why treating infra as a utility matters (33:22) Co-founder dynamics: Blending engineering strategy with business outcomes (37:15) Rapid Fire Questions (40:51) This episode wouldn't have been possible without the help of our incredible production team: Patrick Gallagher - Producer & Co-Host Jerry Li - Co-Host Noah Olberding - Associate Producer, Audio & Video Editor https://www.linkedin.com/in/noah-olberding/ Dan Overheim - Audio Engineer, Dan's also an avid 3D printer - https://www.bnd3d.com/ Ellie Coggins Angus - Copywriter, Check out her other work at https://elliecoggins.com/about/ Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
We speak with Lindsay Lenard of HorseSpot, who will describe her experience in a pitch competition, securing investors for her company, and the details of building a tech company from the ground up.Guest Name: Lindsay LenardWebsite: https://horsespot.net/ Facebook: https://www.facebook.com/horsespotshows Instagram: https://www.instagram.com/horsespotshows/# LinkedIn: https://www.linkedin.com/company/horse-spot-shows/Lindsay Lenard is the Co-founder and Product Design Lead of Horse Spot, Blue Ribbon Software supporting horse shows, rodeos, and fairs from local to internationally rated events. She is a 2× Webby Award–winning designer that has led creative work for global advertising agencies and YC- and Tech-stars-backed startups. A lifelong equestrian, Lindsay is building technology that serves the community that shaped her.
This is a recap of the top 10 posts on Hacker News on June 15, 2026. This podcast was generated by wondercraft.ai (00:30): Iroh 1.0Original post: https://news.ycombinator.com/item?id=48542480&utm_source=wondercraft_ai(01:56): A backdoor in a LinkedIn job offerOriginal post: https://news.ycombinator.com/item?id=48546294&utm_source=wondercraft_ai(03:23): Ask HN: Has anyone replaced Claude/GPT with a local model for daily coding?Original post: https://news.ycombinator.com/item?id=48542100&utm_source=wondercraft_ai(04:50): Curl will not accept vulnerability reports during July 2026Original post: https://news.ycombinator.com/item?id=48537165&utm_source=wondercraft_ai(06:16): What happened to nerds?Original post: https://news.ycombinator.com/item?id=48538229&utm_source=wondercraft_ai(07:43): TinyWind: A pixel pirate sailing game with real wind physics (380k+ kms sailed)Original post: https://news.ycombinator.com/item?id=48543475&utm_source=wondercraft_ai(09:10): CrankGPTOriginal post: https://news.ycombinator.com/item?id=48540854&utm_source=wondercraft_ai(10:37): Apple Foundation ModelsOriginal post: https://news.ycombinator.com/item?id=48536776&utm_source=wondercraft_ai(12:03): Hetzner Price AdjustmentOriginal post: https://news.ycombinator.com/item?id=48540844&utm_source=wondercraft_ai(13:30): Even more batteries included with EmacsOriginal post: https://news.ycombinator.com/item?id=48535886&utm_source=wondercraft_aiThis is a third-party project, independent from HN and YC. Text and audio generated using AI, by wondercraft.ai. Create your own studio quality podcast with text as the only input in seconds at app.wondercraft.ai. Issues or feedback? We'd love to hear from you: team@wondercraft.ai
This is a recap of the top 10 posts on Hacker News on June 14, 2026. This podcast was generated by wondercraft.ai (00:30): How to earn a billion dollarsOriginal post: https://news.ycombinator.com/item?id=48526360&utm_source=wondercraft_ai(01:56): Show HN: Kage – Shadow any website to a single binary for offline viewingOriginal post: https://news.ycombinator.com/item?id=48529990&utm_source=wondercraft_ai(03:23): Not everyone is using AI for everythingOriginal post: https://news.ycombinator.com/item?id=48527700&utm_source=wondercraft_ai(04:50): Honda Civics and the Evil ValetOriginal post: https://news.ycombinator.com/item?id=48523080&utm_source=wondercraft_ai(06:17): Your ePub Is fineOriginal post: https://news.ycombinator.com/item?id=48533848&utm_source=wondercraft_ai(07:44): Free SQL→ER diagram tool, runs in the browser, nothing uploadedOriginal post: https://news.ycombinator.com/item?id=48523992&utm_source=wondercraft_ai(09:11): I indexed 669 GB of my GoPro videos using my M1 Max computer and local ML modelsOriginal post: https://news.ycombinator.com/item?id=48528029&utm_source=wondercraft_ai(10:38): Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing modelOriginal post: https://news.ycombinator.com/item?id=48528371&utm_source=wondercraft_ai(12:05): Linux 7.1Original post: https://news.ycombinator.com/item?id=48528729&utm_source=wondercraft_ai(13:32): Don't trust large context windowsOriginal post: https://news.ycombinator.com/item?id=48524620&utm_source=wondercraft_aiThis is a third-party project, independent from HN and YC. Text and audio generated using AI, by wondercraft.ai. Create your own studio quality podcast with text as the only input in seconds at app.wondercraft.ai. Issues or feedback? We'd love to hear from you: team@wondercraft.ai
This is a recap of the top 10 posts on Hacker News on June 13, 2026. This podcast was generated by wondercraft.ai (00:30): Statement on US government directive to suspend access to Fable 5 and Mythos 5Original post: https://news.ycombinator.com/item?id=48511072&utm_source=wondercraft_ai(01:56): Open source AI must winOriginal post: https://news.ycombinator.com/item?id=48511908&utm_source=wondercraft_ai(03:23): Noise infusion banned from statistical products published by Census BureauOriginal post: https://news.ycombinator.com/item?id=48517377&utm_source=wondercraft_ai(04:50): Every Frame PerfectOriginal post: https://news.ycombinator.com/item?id=48516251&utm_source=wondercraft_ai(06:16): Amazon CEO's talks with U.S. officials triggered crackdown on Anthropic modelsOriginal post: https://news.ycombinator.com/item?id=48519092&utm_source=wondercraft_ai(07:43): Israeli firm BlackCore suspected of meddling in New York and Scotland votesOriginal post: https://news.ycombinator.com/item?id=48514560&utm_source=wondercraft_ai(09:10): Leaving MozillaOriginal post: https://news.ycombinator.com/item?id=48513806&utm_source=wondercraft_ai(10:37): There is a shadow hanging over this Fable thingOriginal post: https://news.ycombinator.com/item?id=48513536&utm_source=wondercraft_ai(12:03): GLM 5.2 Is OutOriginal post: https://news.ycombinator.com/item?id=48518684&utm_source=wondercraft_ai(13:30): Treating pancreatic tumours may have revealed cancer's master switchOriginal post: https://news.ycombinator.com/item?id=48517199&utm_source=wondercraft_aiThis is a third-party project, independent from HN and YC. Text and audio generated using AI, by wondercraft.ai. Create your own studio quality podcast with text as the only input in seconds at app.wondercraft.ai. Issues or feedback? We'd love to hear from you: team@wondercraft.ai
This is a recap of the top 10 posts on Hacker News on June 12, 2026. This podcast was generated by wondercraft.ai (00:30): Statement on US government directive to suspend access to Fable 5 and Mythos 5Original post: https://news.ycombinator.com/item?id=48511072&utm_source=wondercraft_ai(01:58): AI agent bankrupted their operator while trying to scan DN42Original post: https://news.ycombinator.com/item?id=48500012&utm_source=wondercraft_ai(03:26): CRISPR tech selectively shreds cancer cells, including "undruggable" cancersOriginal post: https://news.ycombinator.com/item?id=48505231&utm_source=wondercraft_ai(04:55): Claude Fable is relentlessly proactiveOriginal post: https://news.ycombinator.com/item?id=48498573&utm_source=wondercraft_ai(06:23): Nobody ever gets credit for fixing problems that never happened (2001) [pdf]Original post: https://news.ycombinator.com/item?id=48498385&utm_source=wondercraft_ai(07:52): Open source AI must winOriginal post: https://news.ycombinator.com/item?id=48511908&utm_source=wondercraft_ai(09:20): Kimi K2.7-Code: open-source coding model with better token efficiencyOriginal post: https://news.ycombinator.com/item?id=48502347&utm_source=wondercraft_ai(10:49): "Don't You Just Upload It to ChatGPT?"Original post: https://news.ycombinator.com/item?id=48507278&utm_source=wondercraft_ai(12:17): Electric motors with no rare earthsOriginal post: https://news.ycombinator.com/item?id=48510010&utm_source=wondercraft_ai(13:46): How to setup a local coding agent on macOSOriginal post: https://news.ycombinator.com/item?id=48507020&utm_source=wondercraft_aiThis is a third-party project, independent from HN and YC. Text and audio generated using AI, by wondercraft.ai. Create your own studio quality podcast with text as the only input in seconds at app.wondercraft.ai. Issues or feedback? We'd love to hear from you: team@wondercraft.ai
Dave joins the squad from YC just ahead of WWDC, where the group breaks down Apple's upgraded Siri and the surprising revelation that much of Apple's AI capability is powered by Google's Gemini models. It wouldn't be a technology podcast without the squad expanding into the broader AI landscape, with Dave arguing that outside of OpenAI and Anthropic, most AI companies are becoming infrastructure providers rather than destination products. They continue to unpack the IPO frenzy around SpaceX and OpenAI, including SpaceX's reportedly massive investor demand and OpenAI's murky timeline to go public. Jess then shifts the conversation to the turmoil at 60 Minutes, where leadership shakeups, internal revolts, and a declining brand presence raise questions about the future of legacy media. Finally, in Pop Culture Corner, the pod closes with Disney's $200 million bet on Toy Story 5 and Taylor Swift's outsized promotional impact.Chapters:01:53 Apple WWDC: New Siri Breakdown05:18 Google Gemini Powers Apple AI07:06 WWDC Ignored Developers Entirely10:21 Tim Cook's Leadership Baton Pass14:30 SpaceX IPO and OpenAI Filing18:00 Anthropic vs. OpenAI Narrative War21:48 Dave Morin's Two-Phone Redemption26:52 60 Minutes Leadership Implosion38:44 Toy Story 5 and Taylor SwiftWe're also on ↓X: https://twitter.com/moreorlesspodInstagram: https://instagram.com/moreorlessYouTube: https://youtu.be/Yvox4U_8u1wConnect with us here:1) Sam Lessin: https://x.com/lessin2) Dave Morin: https://x.com/davemorin3) Jessica Lessin: https://x.com/Jessicalessin4) Brit Morin: https://x.com/brit
This is a recap of the top 10 posts on Hacker News on June 11, 2026. This podcast was generated by wondercraft.ai (00:30): Show HN: Homebrew 6.0.0Original post: https://news.ycombinator.com/item?id=48490024&utm_source=wondercraft_ai(01:57): Pokémon Go Scans Trained the Navigation Tech for Military DronesOriginal post: https://news.ycombinator.com/item?id=48487029&utm_source=wondercraft_ai(03:24): AI agent runs amok in Fedora and elsewhereOriginal post: https://news.ycombinator.com/item?id=48484584&utm_source=wondercraft_ai(04:51): MiMo Code is now released and open-sourceOriginal post: https://news.ycombinator.com/item?id=48490826&utm_source=wondercraft_ai(06:18): If you are asking for human attention, demonstrate human effortOriginal post: https://news.ycombinator.com/item?id=48497609&utm_source=wondercraft_ai(07:45): Solar generates more energy in US than coal for first timeOriginal post: https://news.ycombinator.com/item?id=48492306&utm_source=wondercraft_ai(09:12): Petition to Withdraw Canada's Bill C-22Original post: https://news.ycombinator.com/item?id=48491830&utm_source=wondercraft_ai(10:39): Lines of code got a better publicistOriginal post: https://news.ycombinator.com/item?id=48489402&utm_source=wondercraft_ai(12:06): Anthropic apologizes for invisible Claude Fable guardrailsOriginal post: https://news.ycombinator.com/item?id=48489229&utm_source=wondercraft_ai(13:33): Show HN: FablePool – pool money behind a prompt, and Fable builds it in publicOriginal post: https://news.ycombinator.com/item?id=48496539&utm_source=wondercraft_aiThis is a third-party project, independent from HN and YC. Text and audio generated using AI, by wondercraft.ai. Create your own studio quality podcast with text as the only input in seconds at app.wondercraft.ai. Issues or feedback? We'd love to hear from you: team@wondercraft.ai
What if agentic AI makes SRE more important, not less? Bennett Gould explains why autonomous AI systems may create more demand for reliability thinking — not less.Everyone seems to think AI is coming for SRE in a hard way.You might have heard the same story:“AI will write the code.”“Agents will handle incidents.”“Copilots will generate the runbooks.”“Automation will reduce operational load.”Yes, the job question is real. If AI can write code, summarize incidents, query observability tools, generate runbooks, and operate across systems, then engineers are right to ask what happens to the work.But here's the part that gets missed: AI does not just automate reliability work. It creates more objects and surface areas that need to be made reliable.Agentic AI is moving from demos into real workflows. These systems are no longer just answering questions. They are querying tools, pulling context, generating changes, and in some cases taking action around production environments.That makes this a Monday morning problem.Teams are already using LLMs for incidents, documentation, observability, infrastructure, and operational decision-making. Somewhere, a team is one demo away from giving an agent access to tools originally designed for humans.That is exactly why I wanted to have this conversation.Bennett Gould is currently a solution engineer at Neubird.ai. His career in SRE and SRE-adjacent work spans large enterprises, cloud, industrial technology, and startups, including AWS, IBM, Siemens, and a YC startup.I wanted to ask him a simple question: What in the agentic AI is happening to SRE?Here are 3 highlights from our talk:1. Agentic AI increases the reliability surface areaThe obvious fear is that AI reduces the need for reliability engineers. Bennett's view was more nuanced. He was clear that engineers still need to adapt. If people do not reskill, stay current, and learn how these systems are forming, there may absolutely be pressure in the job market. But he also argued that AI could create more demand for reliability skills because production complexity is increasing.More code is going into production.More AI-generated code is going into production.More systems that people do not fully understand are going into production.And now autonomous agents are starting to enter production workflows too.That means more surface area. More automation. More operational uncertainty. More ways for things to go wrong.Bennett compared this to Terraform: Infrastructure as code created enormous efficiency gains. But it also created new ways to make very big mistakes very quickly.Before Terraform, most people could not delete all their production resources with a single command. After Terraform, that became technically possible if the system was designed badly enough.Agentic AI follows a similar pattern. With great automation comes great responsibility.Agents can help engineers move faster, query tools, summarize context, and reduce toil. But they can also amplify weak engineering practices, poor boundaries, bad assumptions, and unclear operational ownership. That is not the end of reliability work. That is reliability work entering a new phase.2. Agents can reduce toil, but context is the ceilingOne of the strongest parts of the conversation was Bennett's explanation of where agents can help in incident response. A lot of SRE work involves moving across tools.You may need to query Prometheus, Dynatrace, logs, traces, cloud consoles, ticketing systems, documentation, runbooks, dashboards, and architecture diagrams.The problem is not always that the engineer lacks judgment.Sometimes the problem is that the information is scattered across too many tools, each with its own query language and interface. Bennett gave a simple example: an engineer might be very good at PromQL and very fast when Prometheus is the source of truth. But if the same engineer has to work in a different observability platform with a different query language, their response time can suffer. That is an obvious place where agents can help.The engineer may not need to know every query language perfectly. They need to know what they are looking for and how to reason about the system. The agent can help translate that intent into the right tool calls, queries, and summaries.That could reduce MTTR. It could reduce toil. It could help engineers move faster during incidents.But Bennett also made the limitation clear: You are only as good as the context you have. This is where he introduced two useful concepts:* Context mining* Context distillationContext mining means proactively finding the information that might be useful in a given operational situation.Context distillation means taking large amounts of information — runbooks, Confluence pages, diagrams, documentation, prior incidents — and reducing it into the minimum useful context an LLM or agent can use.That sounds powerful. But there is a catch. Sometimes the context simply is not there.Many of the largest and most complex organizations still run legacy systems where knowledge lives in people's heads, stale documentation, tribal memory, and unwritten assumptions.There may not be a clean process for turning that into usable context. That matters because agents do not magically understand your system. They work with the context they are given. If the context is missing, outdated, or wrong, the agent's usefulness maxes out early.3. Agentic systems are not just LLM demosA basic LLM workflow is relatively easy to demo:You give it a prompt.You connect a few tools.You add some APIs.You get a useful answer.That is impressive, but it is not the same thing as running an agentic system in a meaningful production environment.Bennett made a useful analogy here: running your own infrastructure versus using a hyperscaler.Cloud providers removed a lot of undifferentiated heavy lifting. Most companies do not want to spend half their time racking servers, managing data centers, and dealing with low-level infrastructure when they are trying to serve customers.Agentic systems create similar questions:* What parts of the work should be handled by the system?* What parts still need engineering discipline?* And what has to exist around the model before it is safe and useful?That surrounding structure is where the real work begins. Bennett called this harness engineering. Once you move beyond an LLM demo, you have to think about memory, learning, tool usage, identity, federation, security, evaluations, and guardrails.That is a very different problem from “the model gave a good answer on my laptop.” SREs know why that distinction matters. “It works on my machine” is not an acceptable reliability strategy.A runbook that recovers a thousand-node database cannot be non-deterministic, undocumented, and dependent on someone's local setup. If it is part of the operational backbone, it needs to be reliable.Agentic AI does not remove that requirement. It makes it more important.Bonus: Agents expose weak engineering practicesAgentic AI not only introduces new problems but it also reveals old ones.* Weak APIs.* Brittle runbooks.* Missing context.* Poor evals.* Unclear tool boundaries.* Operational shortcuts.Systems that were designed assuming careful human use may behave very differently when AI agents start using them. That is why this conversation matters for SRE.Agentic AI is not only a productivity story. It is a reliability story.It forces teams to ask whether their existing practices are strong enough for a world where more actions can be generated, recommended, or executed by autonomous systems.The silver lining for reliability workAgentic AI does not remove the need for reliability thinking. It raises the bar for it. The tools will change. The workflows will change. Some tasks will absolutely be automated or reshaped.But the hardest parts of reliability are still the hard parts:* understanding the system* knowing the trade-offs* building reliable operational processes* making good judgment calls under uncertainty and* owning the outcome when something changes in productionThat is why SRE does not disappear in an agentic AI world.It becomes one of the disciplines that makes the agentic AI world survivable.So if your team is already using AI around incidents, observability, runbooks, infrastructure, or production workflows, the question is not whether the future is coming. The future is already in the workflow.The real question is whether your reliability practices are ready for it. This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit read.srepath.com
This is a recap of the top 10 posts on Hacker News on June 10, 2026. This podcast was generated by wondercraft.ai (00:30): macOS Container MachinesOriginal post: https://news.ycombinator.com/item?id=48469658&utm_source=wondercraft_ai(01:59): Building an HTML-first site doubled our users overnightOriginal post: https://news.ycombinator.com/item?id=48475483&utm_source=wondercraft_ai(03:28): German ruling declares Google liable for false answers in AI OverviewsOriginal post: https://news.ycombinator.com/item?id=48470248&utm_source=wondercraft_ai(04:57): πFSOriginal post: https://news.ycombinator.com/item?id=48480978&utm_source=wondercraft_ai(06:27): I'm Eric Ries, author of "The Lean Startup" and new book "Incorruptible" – AMAOriginal post: https://news.ycombinator.com/item?id=48477135&utm_source=wondercraft_ai(07:56): Mercedes‑Benz starts large‑scale production of electric axial flux motorOriginal post: https://news.ycombinator.com/item?id=48472877&utm_source=wondercraft_ai(09:25): PgDog is funded and coming to a database near youOriginal post: https://news.ycombinator.com/item?id=48476466&utm_source=wondercraft_ai(10:54): AWS Bedrock to require sharing data with Anthropic for Mythos and future modelsOriginal post: https://news.ycombinator.com/item?id=48473166&utm_source=wondercraft_ai(12:24): Chrome is looking to permanently drop MV2 extensionOriginal post: https://news.ycombinator.com/item?id=48471970&utm_source=wondercraft_ai(13:53): Claude Desktop spawns 1.8 GB Hyper-V VM on every launch, even for chat-only useOriginal post: https://news.ycombinator.com/item?id=48479452&utm_source=wondercraft_aiThis is a third-party project, independent from HN and YC. Text and audio generated using AI, by wondercraft.ai. Create your own studio quality podcast with text as the only input in seconds at app.wondercraft.ai. Issues or feedback? We'd love to hear from you: team@wondercraft.ai
This is a recap of the top 10 posts on Hacker News on June 09, 2026. This podcast was generated by wondercraft.ai (00:30): Claude Fable 5Original post: https://news.ycombinator.com/item?id=48463808&utm_source=wondercraft_ai(01:56): Making Graphics Like it's 1993Original post: https://news.ycombinator.com/item?id=48459294&utm_source=wondercraft_ai(03:22): If Claude Fable stops helping you, you'll never knowOriginal post: https://news.ycombinator.com/item?id=48467896&utm_source=wondercraft_ai(04:48): CEOs who think AI replaces their employees are just bad CEOsOriginal post: https://news.ycombinator.com/item?id=48465675&utm_source=wondercraft_ai(06:15): Microsoft's open source tools were hacked to steal passwords of AI developersOriginal post: https://news.ycombinator.com/item?id=48457830&utm_source=wondercraft_ai(07:41): FCC wants to kill burner phones by forcing telecoms to get all customers' IDsOriginal post: https://news.ycombinator.com/item?id=48462308&utm_source=wondercraft_ai(09:07): macOS Container MachinesOriginal post: https://news.ycombinator.com/item?id=48469658&utm_source=wondercraft_ai(10:34): Cleaning up after AI rockstar developersOriginal post: https://news.ycombinator.com/item?id=48458586&utm_source=wondercraft_ai(12:00): Albania Is Not for Sale: Kushner's $4B Resort Triggers'Flamingo Revolution'Original post: https://news.ycombinator.com/item?id=48461012&utm_source=wondercraft_ai(13:26): Apple decided not to roll out Siri in EU after denied request for exemptionOriginal post: https://news.ycombinator.com/item?id=48463024&utm_source=wondercraft_aiThis is a third-party project, independent from HN and YC. Text and audio generated using AI, by wondercraft.ai. Create your own studio quality podcast with text as the only input in seconds at app.wondercraft.ai. Issues or feedback? We'd love to hear from you: team@wondercraft.ai
This is a recap of the top 10 posts on Hacker News on June 08, 2026. This podcast was generated by wondercraft.ai (00:30): Show HN: Performative-UI – A react component library of design tropesOriginal post: https://news.ycombinator.com/item?id=48445554&utm_source=wondercraft_ai(01:58): Dopamine FrackingOriginal post: https://news.ycombinator.com/item?id=48440792&utm_source=wondercraft_ai(03:26): Anti-social: It's fads, not friends, which now dominate social media feedsOriginal post: https://news.ycombinator.com/item?id=48444228&utm_source=wondercraft_ai(04:54): Stop the Apple Music app from launchingOriginal post: https://news.ycombinator.com/item?id=48447935&utm_source=wondercraft_ai(06:22): MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per secondOriginal post: https://news.ycombinator.com/item?id=48446639&utm_source=wondercraft_ai(07:50): Siri AIOriginal post: https://news.ycombinator.com/item?id=48449084&utm_source=wondercraft_ai(09:18): xAI is looking more like a datacentre REIT than a frontier labOriginal post: https://news.ycombinator.com/item?id=48446428&utm_source=wondercraft_ai(10:47): Surveillance is not safety: A statement on the UK's latest threat to privacy [pdf]Original post: https://news.ycombinator.com/item?id=48450646&utm_source=wondercraft_ai(12:15): Apple reveals new AI architecture built around Google Gemini modelsOriginal post: https://news.ycombinator.com/item?id=48450142&utm_source=wondercraft_ai(13:43): AI is slowing downOriginal post: https://news.ycombinator.com/item?id=48446893&utm_source=wondercraft_aiThis is a third-party project, independent from HN and YC. Text and audio generated using AI, by wondercraft.ai. Create your own studio quality podcast with text as the only input in seconds at app.wondercraft.ai. Issues or feedback? We'd love to hear from you: team@wondercraft.ai
This is a recap of the top 10 posts on Hacker News on June 07, 2026. This podcast was generated by wondercraft.ai (00:30): LLMs are eroding my software engineering career and I don't know what to doOriginal post: https://news.ycombinator.com/item?id=48434312&utm_source=wondercraft_ai(01:57): Building from zero after addiction, prison, and a felonyOriginal post: https://news.ycombinator.com/item?id=48437406&utm_source=wondercraft_ai(03:25): Anthropic, please ship an official Claude Desktop for LinuxOriginal post: https://news.ycombinator.com/item?id=48434436&utm_source=wondercraft_ai(04:52): The 29th International Obfuscated C Code Contest (IOCCC) 2025 WinnersOriginal post: https://news.ycombinator.com/item?id=48432199&utm_source=wondercraft_ai(06:20): How's Linear so fast? A technical breakdownOriginal post: https://news.ycombinator.com/item?id=48437609&utm_source=wondercraft_ai(07:47): Scientists ejected from diabetes conference for distributing journal reprintsOriginal post: https://news.ycombinator.com/item?id=48433410&utm_source=wondercraft_ai(09:15): I design with Claude more than Figma nowOriginal post: https://news.ycombinator.com/item?id=48431981&utm_source=wondercraft_ai(10:42): Show HN: Lathe – Use LLMs to learn a new domain, not skip past itOriginal post: https://news.ycombinator.com/item?id=48433756&utm_source=wondercraft_ai(12:10): Major P2P issues in Israel and possibly other Middle East countriesOriginal post: https://news.ycombinator.com/item?id=48431461&utm_source=wondercraft_ai(13:37): Public Domain Image ArchiveOriginal post: https://news.ycombinator.com/item?id=48430539&utm_source=wondercraft_aiThis is a third-party project, independent from HN and YC. Text and audio generated using AI, by wondercraft.ai. Create your own studio quality podcast with text as the only input in seconds at app.wondercraft.ai. Issues or feedback? We'd love to hear from you: team@wondercraft.ai
This is a recap of the top 10 posts on Hacker News on June 06, 2026. This podcast was generated by wondercraft.ai (00:30): S&P 500 rejects SpaceX, also blocking entry for OpenAI and AnthropicOriginal post: https://news.ycombinator.com/item?id=48421442&utm_source=wondercraft_ai(01:59): Meta confirms 1000s of Instagram accounts were hacked by abusing its AI chatbotOriginal post: https://news.ycombinator.com/item?id=48427643&utm_source=wondercraft_ai(03:29): Pentagon raised threat of Israeli spying on U.S. to highest level, sources sayOriginal post: https://news.ycombinator.com/item?id=48427523&utm_source=wondercraft_ai(04:59): GrapheneOS user reported to authorities for using GrapheneOSOriginal post: https://news.ycombinator.com/item?id=48422798&utm_source=wondercraft_ai(06:28): Ask HN: Why is the HN crowd so anti-AI?Original post: https://news.ycombinator.com/item?id=48420827&utm_source=wondercraft_ai(07:58): Ntsc-rs – open-source video emulation of analog TV and VHS artifactsOriginal post: https://news.ycombinator.com/item?id=48428025&utm_source=wondercraft_ai(09:28): Pokemon Emerald Ported to WebAssembly (100k FPS)Original post: https://news.ycombinator.com/item?id=48423762&utm_source=wondercraft_ai(10:57): Moving beyond fork() + exec()Original post: https://news.ycombinator.com/item?id=48425528&utm_source=wondercraft_ai(12:27): Nvidia is proposing a beast of a CPU system for Windows PCsOriginal post: https://news.ycombinator.com/item?id=48424605&utm_source=wondercraft_ai(13:57): The intracies of modern camera lens repair (2024)Original post: https://news.ycombinator.com/item?id=48420148&utm_source=wondercraft_aiThis is a third-party project, independent from HN and YC. Text and audio generated using AI, by wondercraft.ai. Create your own studio quality podcast with text as the only input in seconds at app.wondercraft.ai. Issues or feedback? We'd love to hear from you: team@wondercraft.ai
Amar, servir, adorar Por: Pastor Luis Navarrete Oseas 1:2 (NTV). Por medio de la historia de Oseas y Gomer, Dios transmite un poderoso mensaje: su amor inquebrantable por su pueblo. Aunque el pueblo le fue infiel en muchas ocasiones, Dios nunca dejó de amarlo. De la misma manera, hoy Dios sigue llamando a sus hijos a regresar a Él para amarle, servirle y adorarle de todo corazón. 1. Amar. El amor es el inicio de nuestra historia con Dios. Todo ser humano tiene un vacío en su corazón. Es un vacío de amor que solo Dios puede llenar. Todos hemos experimentado el amor de Dios. Ese amor nos cautivó, nos sanó, nos conmovió y nos transformó. Romanos 5:8. Muchos recordamos el momento en que el Señor tocó nuestras vidas. Nuestro encuentro con Cristo nos marcó para siempre. La gratitud llenaba nuestro corazón y otra pregunta surgía constantemente: ¿Qué puedo hacer por ti, Señor? Ese amor nos llevó a comprometernos con la iglesia y con la obra de Dios. Queríamos estar en cada reunión y participar en cada actividad. Y de manera natural, el siguiente paso fue servir. 2. Servir. Como respuesta a su amor, empezamos a servir al Señor. No importaban las dificultades; el amor y gratitud eran suficientes para superar pruebas, ofensas y obstáculos en el camino. El fruto del primer amor siempre es bueno. Oseas 1:3 (NTV). Oseas 2:5-8 (NTV). Pero, el corazón de Gomer comenzó a cambiar y debilitarse a causa de los deseos de la carne y los engaños del enemigo. Ella volvió sus ojos hacia sí misma y sus antiguos deseos. Creyó que debía suplir sus propias necesidades, olvidando a su proveedor, su esposo. Lo mismo puede suceder con nosotros. Nos sentirnos usados y la gratitud desaparece. Vienen cuestionamientos y encontramos razones para quejarnos. Nuestra atención se desvía hacia otras cosas y comenzamos a perseguirlas. Olvidamos que todo lo que tenemos proviene de Dios y terminamos poniendo sus bendiciones al servicio del mundo. Oseas 2:6-8 (NTV). Entonces surgen las preguntas: ¿En qué momento perdimos el camino? ¿En qué momento nos dejamos engañar? ¿En qué momento la historia de amor cambió? Y ¿Cómo podemos mantener ese primer amor? 3. Adorar. La adoración es rendición y entrega completa a Dios, que nace del corazón, en espíritu y en verdad. Es lo que nos mantiene conectados a su amor. Es nuestra expresión de amor y devoción hacia nuestro Amado. Nunca descuidemos el anhelo por su presencia, nuestra comunión y nuestra intimidad con Él. Una relación sana con Dios se construye diariamente sobre la confianza, la honra y la obediencia. Oseas 2:14-16 (NTV). Dios toma la iniciativa para buscarnos. Las pruebas, aflicciones y desafíos pueden convertirse en puertas de esperanza para llevarnos nuevamente a su presencia. Son oportunidades para regresar al primer amor y a las primeras obras. Oseas 3:1-3. Qué impresionante cuadro del amor de Dios. Aun después de la infidelidad de Gomer, Oseas fue enviado a buscarla, rescatarla y restaurarla. De la misma manera, Dios nos busca cuando nos alejamos, nos llama al arrepentimiento y nos recibe nuevamente con amor. El mensaje de Oseas sigue vigente hoy: Dios nos llama a volver a nuestro primer amor. Primero nos amó, luego nos llevó a servirle y, finalmente, nos invita a permanecer en una vida de adoración constante. Amar, servir y adorar no son etapas separadas de la vida cristiana; son expresiones de una misma relación de amor con nuestro Señor La entrada Amar, servir, adorar – Ps. Luis Navarrete se publicó primero en Comunidad de Fe.
This is a recap of the top 10 posts on Hacker News on June 05, 2026. This podcast was generated by wondercraft.ai (00:30): Changing how we develop LadybirdOriginal post: https://news.ycombinator.com/item?id=48409191&utm_source=wondercraft_ai(01:59): Gov.uk has replaced Stripe with Dutch provider AdyenOriginal post: https://news.ycombinator.com/item?id=48415217&utm_source=wondercraft_ai(03:29): C++: The DocumentaryOriginal post: https://news.ycombinator.com/item?id=48408016&utm_source=wondercraft_ai(04:59): Tracing a powerful GNSS interference source over EuropeOriginal post: https://news.ycombinator.com/item?id=48409664&utm_source=wondercraft_ai(06:29): Astronauts told to return to ISS after sheltering over air leak repairsOriginal post: https://news.ycombinator.com/item?id=48413464&utm_source=wondercraft_ai(07:59): pg_durable: Microsoft open sources in-database durable executionOriginal post: https://news.ycombinator.com/item?id=48414367&utm_source=wondercraft_ai(09:29): Did Claude increase bugs in rsync?Original post: https://news.ycombinator.com/item?id=48411635&utm_source=wondercraft_ai(10:59): Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiencyOriginal post: https://news.ycombinator.com/item?id=48414653&utm_source=wondercraft_ai(12:29): New method turns ocean water into drinking water, without wasteOriginal post: https://news.ycombinator.com/item?id=48413500&utm_source=wondercraft_ai(13:59): Meta enables ADB on deprecated Portal devices [video]Original post: https://news.ycombinator.com/item?id=48406640&utm_source=wondercraft_aiThis is a third-party project, independent from HN and YC. Text and audio generated using AI, by wondercraft.ai. Create your own studio quality podcast with text as the only input in seconds at app.wondercraft.ai. Issues or feedback? We'd love to hear from you: team@wondercraft.ai
Que noche la de anoche mi gente, gracias a Chumel Torres por regalarnos un episodio increíble de La Corneta Extendida... próximamente. Nuestra Presidenta recomienda que la oposición haga yoga: "aaaauummm... paz mundial". Y César Cravioto se cura la tos con extracto de ajolote, que él mismo ordeña. México golea y ya soñamos con la copa, al menos Diego Luna así lo creo y el 'Vasco' se va a retirar hasta que su esposa lo autorice. Danna manda importante apoyo a la Selección Nacional y Gaby Cam nos da consejos como los de He-Man.
This is a recap of the top 10 posts on Hacker News on June 04, 2026. This podcast was generated by wondercraft.ai (00:30): Failing grades soar with AI usage, dwindling math skills in Berkeley CS classesOriginal post: https://news.ycombinator.com/item?id=48392004&utm_source=wondercraft_ai(01:58): U.S. to dismantle system tracking Atlantic currents that are at risk of collapseOriginal post: https://news.ycombinator.com/item?id=48392232&utm_source=wondercraft_ai(03:26): VoidZero Is Joining CloudflareOriginal post: https://news.ycombinator.com/item?id=48398055&utm_source=wondercraft_ai(04:54): Ian's Secure Shoelace KnotOriginal post: https://news.ycombinator.com/item?id=48397028&utm_source=wondercraft_ai(06:22): French-Iranian author Marjane Satrapi, author of 'Persepolis', dies at 56Original post: https://news.ycombinator.com/item?id=48397233&utm_source=wondercraft_ai(07:50): When AI Builds Itself: Our progress toward recursive self-improvementOriginal post: https://news.ycombinator.com/item?id=48400842&utm_source=wondercraft_ai(09:18): I built a vulnerable app and spent $1,500 seeing if LLMs could hack itOriginal post: https://news.ycombinator.com/item?id=48392343&utm_source=wondercraft_ai(10:46): Wind and solar generated more power than gas globally in April 2026Original post: https://news.ycombinator.com/item?id=48399332&utm_source=wondercraft_ai(12:14): UK media fails to disclose defence sector links in nearly 60% of casesOriginal post: https://news.ycombinator.com/item?id=48395938&utm_source=wondercraft_ai(13:42): Anthropic's open-source framework for AI-powered vulnerability discoveryOriginal post: https://news.ycombinator.com/item?id=48403980&utm_source=wondercraft_aiThis is a third-party project, independent from HN and YC. Text and audio generated using AI, by wondercraft.ai. Create your own studio quality podcast with text as the only input in seconds at app.wondercraft.ai. Issues or feedback? We'd love to hear from you: team@wondercraft.ai
The new AIEWF website is live! Get your tickets booked ASAP as they -will- sell out. Take the AI Engineering Survey and get >$2k in credits and free AIE WF tickets!Most industry benchmarks compress intelligence and reasoning ability into scores.SWE-Bench Pro, MMLU, Humanity's Last Exam, etc. These metrics are useful, but don't always represent the full extent of how a model performs in the real world. Some of the most interesting evals today look less like exams and more like operating businesses in the real world. One of which is Vending Bench.In Anthropic's Mythos Preview System Card, Andon was the only third party eval to get their own section, observing increasingly concerning aggressive behavior:You don't know what a model is capable of doing in the real world unless you actually give it inventory, a wallet, tools, customers, competitors, humans, & some time. More often than not, it'll surprise you how much a model is capable of and in doing so, also reveal unexpected behavior: deception, context collapse, emergent coordination, & bizarre negotiation behavior.While an inflection point in personal agents came post-OpenClaw after full file access with bypass permissions became the norm, it is yet to come for agents in the real-world. However Andon Market, an actual in person store fully run and managed by AI, is paving the way for what is possible.Full Video PodFrom Claude trying to call the FBI over a $2/day vending machine charge to AI agents forming price cartels, hiring human employees, running physical stores, and writing existential robot musicals, Andon Labs is stress-testing what happens when frontier models stop being chatbots and start acting in the real world. In this episode, Andon Labs cofounders Lukas Petersson and Axel Backlund join swyx and Vibhu to unpack the strange, funny, and genuinely concerning edge cases that emerge when agents run businesses over long horizons.We go deep on Vending-Bench, Project Vend, Vending-Bench Arena, Bengt, Butter-Bench, Luna, and Andon's broader mission of building realistic real-world evals for autonomous AI systems. Lukas and Axel explain why dollar-denominated evals reveal things traditional benchmarks miss, how Claude ended up reporting its vending machine fees as cybercrime, why long context windows can drive agents into meltdown loops, what happens when agents compete with each other, and why the future of AI safety may depend on testing models in messy physical environments instead of clean benchmark sandboxes.We discuss:* Why Andon Labs started with dangerous capability evals and long-running agents* Vending-Bench and why running a vending machine is a deceptively hard AI benchmark* Why money-based evals avoid the saturation problem of traditional benchmarks* How Claude tried to call the FBI over a $2/day fee* Why long-horizon agents can spiral into existential and legalistic breakdowns* Project Vend: putting an AI-run vending machine inside Anthropic* Why real humans are “out of distribution” for simulated agents* Claudius, Seymour Cash, and the chaos of AI CEOs* How a human briefly became CEO of Claudius through a manipulated election* Why multi-agent systems can converge back into “helpful assistant” behavior* Bengt, Andon's internal office agent with email, spending, terminal, phone, camera, and internet access* How Bengt traded Amazon purchases for face-recognition training data* Claude's aggressive behavior, lies, refund avoidance, and price-cartel behavior in Arena* Why eval awareness may become the AI version of “are we living in a simulation?”* Blueprint Bench, spatial intelligence, and why models still misunderstand physical rooms* Butter-Bench and testing LLMs as robot orchestrators* Luna, the AI-run physical store with a three-year lease and human employees* The new Andon cafe in Sweden and why real-world geography matters for agent evals* Rotten tomatoes, perishable goods, and the hidden difficulty of running a physical businessLukas Petersson* LinkedIn: https://www.linkedin.com/in/lukas-petersson-181a83172/* X: https://x.com/lukaspetAxel Backlund* LinkedIn: https://www.linkedin.com/in/axelbacklund* X: https://x.com/axelbacklundAndon Labs* Website: https://andonlabs.com* Vending-Bench: https://andonlabs.com/evals/vending-bench* Andon Vending: https://andonlabs.com/vendingTimestamps00:00:00 Introduction00:01:00 Andon Labs and the Origins of Vending-Bench00:05:21 Why Money-Based Evals Matter00:09:51 Agent Harnesses and Self-Modifying Systems00:13:36 Claude Calls the FBI00:16:33 Project Vend: Claude Runs a Real Vending Machine00:21:44 Seymour Cash, AI CEOs, and Election Chaos00:27:16 Multi-Agent Coordination and Slack Observability00:30:18 When Will Agents Run Real Businesses?00:34:56 Bengt: Andon's Internal Office Agent00:40:06 Real-World AI Safety and Long-Horizon Traces00:44:28 Lying, Refunds, and Price Cartels in Arena00:52:42 Eval Awareness and Simulation Behavior00:56:06 Blueprint Bench, Butter-Bench, and Robotics01:04:37 Luna: The AI-Run Physical Store01:09:29 The Sweden Cafe and Real-World Expansion01:13:16 What Comes Next for Andon LabsTranscriptIntroduction: Andon Labs, Long-Running Agents, and Real-World EvalsSwyx [00:00:00]: Welcome to Lukas and Axel from Andon Labs, and I'm joined by my, favorite guest host. Anything security, safety, alignments, Vibhu., welcome.Lukas [00:00:15]: Thank you for having us.Axel [00:00:16]: Thank you.Swyx [00:00:17]: Let's match names to voices., maybe you wanna take turns introducing yourselves.Lukas [00:00:21]: I'm Lukas.Axel [00:00:22]: And I'm Axel.Swyx [00:00:24]: Let's introduce Andon Labs a bit. How did you guys come together?, you have different backgrounds, but you're both Swedish., was that, a big part of it?Lukas [00:00:33]: So when I went to high school, there was this really cool guy who had a superpower. He could code. So he made like the or like the app for the, for the school and stuff, and he was super cool, and I wanted to be like him, and that was that guy.Axel [00:00:47]: I don't know about this.Swyx [00:00:49]: But you went to different universities, right?Lukas [00:00:51]: But same high school.Swyx [00:00:52]: I see.Lukas [00:00:52]: So we always said, “Oh, once we graduate university, then we should start a company,” and that's what we did.Swyx [00:00:58]: Wow, there you go. And about a year ago, you kinda burst onto the scene with Vending Bench, but, was there a thing before that was, kind of like the inception?From Dangerous Capability Evals to Vending BenchAxel [00:01:07]: So we did work, yeah, with, Anthropic was one of our, early customers in doing, evals. So we did, dangerous capability evals., nothing we published openly. But then we started thinking about doing some kind of, public benchmark, and one thing that we really started thinking about, was like running agents and specifically agents managing businesses., ‘cause-- and this was, early 2025., and I think the first, mentions of people will be running, person unicorns or even autonomous companies. So we thought, “Let's make a benchmark of how well can an agent run the probably simplest business, possible,” and, that's probably, running a vending machine. So that's the first public one we did. And it was very, like-- there was almost no one that noticed it in the first couple of months, I think., so we released it in February last year, and then I think around Easter last year, we got, the first viral tweet about it, that someone else did.Lukas [00:02:11]: We tweeted a bunch, uh When it came out and, tried our best.Axel [00:02:15]: We tried.Vibhu [00:02:16]: It's the one at Anthropic, right?Lukas [00:02:18]: So thisSwyx [00:02:19]: This is a classic thing we should get out of the way.Lukas [00:02:20]: Exactly. There's two versions.Swyx [00:02:22]: Everyone does this. Yes.Lukas [00:02:23]: There's Vending Bench, which is the simulated one, which we did, completely independently in February., and then, like Axel said, that was like-- That was the thing that didn't get any traction in the beginning, but then some random person made a tweet about it, and thatAxel [00:02:38]: You have the paperLukas [00:02:38]: That is the paper. Correct, yeah., and then since we thought this was very fun, we thought, oh, I think this is also, one thing with Andon Labs, the way we kind of like decide what to do next and what projects to do, it's what is like the heuristic we use is what is fun? Is What would be a fun project? And doing this in real life sounded quite fun for us, and maybe also scientifically useful. So, then we basically had this idea, and then we, like-- But then we needed a place for it and, putting it out in the public would probably not really work., would get vandalized and stuff. So we pitched it to the people we were already working with at Anthropic, and they were “Yeah, you can have space. This sounds fun.” UmSwyx [00:03:21]: It's like a small fridge, right? It's like a mini fridge.Axel [00:03:23]: Absolutely.Swyx [00:03:24]: People-- There's like a stripe thing or like anVibhu [00:03:27]: Oh, okay. So it was very OG, the early daysLukas [00:03:28]: That's the OG one. YeahVibhu [00:03:29]: IPad on this. We saw it in June, like two months after After it had been there. They upgraded a little bit. There's a security camera for making sure you actually Venmo the thing.Swyx [00:03:40]: So, my impression, okay, we're, we're going straight into project Ven because it's such a iconic thing. I do want to cover a little bit of that, the origin story even before Project Ven and even into Vending Bench. I think a lot of people are like yourselves, like smart, interested in future of AI, interested in developing evals. But how the hell do you just, walk into Anthropic's doors and, work with them, right? What is What are they looking for? What works? And then maybe, when you launch, I always think, obviously it would be better to launch with a lab, but, sometimesVibhu [00:04:12]: It's harder to do than it seems.Swyx [00:04:13]: Exactly. So either of those, which are more sort of newbie beginner questions, but, I think it's meaningful advice to others.Lukas [00:04:21]: We get this question a lot, and I don't think our experience is maybe the best., but, the way we did it was that we just built a bunch of things that we had conviction would be useful, and then we just, set up a server and sent it to them for free to use. And then after a while they were “Oh, yeah, this is actually kind of useful. We should probably pay for this.”, but that took a while. I don't know if this is, the best path to doing it, but that's how it went for us.Axel [00:04:47]: I think maybe generally, building-- everyone is interested in good evals, and especially evals that, don't saturate that easily. So, if you can build an eval that, tests something novel, something useful, and you have, good separation of models, like your, the more advanced models rank higher than the worst models, and then you can, yeah, you can, publish it and, try to get some traction, sort of how Vending Bench got attention., and then probably some lab will be interested or you can at least have something to reach out with, when you're doing that.Why Dollar-Based Evals MatterSwyx [00:05:21]: I think you are in, you're in one of the few categories of, evals that correlate to real money. Like Suelancer was also last year, right? Where, people solve actual Upwork. Was it Upwork or other tasks?, something. Where's the, where's, like It's like a dollar value, right? Forget your ELO scores. Forget yourAxel [00:05:37]: PercentilesSwyx [00:05:38]: Zero to one hundred percents. Just go straight for dollars and, that's AGI.Lukas [00:05:43]: And there's like-- I think the nice thing is that there's no ceiling. You can just-- It never saturates because it could just make more and more money. Like If there's oh, Percentage-wise, then, you can't go above, a hundred. And I think like Even when you're not at the hundred, I think a lot of these, evals have a lot of problems in them. So, actually it's like if you getAxel [00:06:05]: To like 92 or something like that, many of them. It's like then there's like there's no really no difference between 92 and 93 because the eval itself is problematic and has noise in it. And I think a lot of evals are saturated like that, but people like pretend that there ‘s still signal in them, but there really isn't.Vending Bench 1, Harness Design, and SaturationSwyx [00:06:24]: Like Super bench verified., even Vending Bench 1 saturated, right? Maybe we can talk about that., may- and maybe set up Vending Bench for a lot of folks who don't know. Actually, things that were very basic like there's limited slots, like you have to pay rent., these are elements where like it doesn't come across in the, in the narrative, but even being adversarial towards the agent, I think these are all like very interesting dimensions.Axel [00:06:47]: I don't really think it's saturated, right? Like it It was more like it was not designed in a way that was really, like true to how AI developed. Like we had an agent harness in it that wasn't really how people used harnesses and stuff like that., so I think it wasn't really that it saturated, it was more like it wasn't really, the best benchmark.Vibhu [00:07:12]: This is Vending Bench one, right?Axel [00:07:14]: I think that like schematic maps sort of to Vending Bench 2 as well., butSwyx [00:07:19]: Including the email.Axel [00:07:20]: The email The emails exist still. Exactly., and then we still we simulate the purchases and it's all, yeah, it's this very open environment for the agent to just run its business. And then for, yeah, Vending Bench 2 we did that, like you said, to just improve the harness., a lot of like nice, like easier, improvements to make it easier for us to run as well., like when you make an eval you ideally want don't want to change it after you made it. So, you want to make it really good and then not to rerun all the models when you make an update because that's also really expensive with the Vending Bench when you run the frontier models. But like as an example, like one thing we didn't have, we didn't have prompt caching in Vending Bench 1, because when we made Vending Bench 1 it wasn't really a thing., so that ‘s just an example of like in Vending Bench 2 like we paid a lot more to run these things because we didn't have prompt caching. So for Vending Bench 2 that was one thing we added and there was a bunch of things like this., and that'Swyx [00:08:17]: Also the conversations are a lot longer in Vending Bench 2, right?Axel [00:08:21]: I think it's kind of similar.Swyx [00:08:22]: Is it similar?Axel [00:08:23]: I think it's similar. The models at the time were worse, so they crashed out earlier., and now they survive the full year all the time.Swyx [00:08:31]: Which is like thousands of turns. Hundreds of thousands of hundreds of millions of tokens output. That's the, that's the rough order of magnitude. I always wonder about the harness. The harness matters a lot. It's your harness. Was there any question about like use cloud code, use something else?Axel [00:08:48]: I think our philosophy around harnesses is like we try to make something that's quite minimalistic, like quite simple. Like we don't wanna favor one model a lot over the other, but also don't make like a super complex harness. So like it's obvious like a model may be lucky and just be good in one harness., so like it is similar to a lot of the harnesses out there in like you have the, like a running loop., you have some like a bunch of tools that are like quite, descriptive for the agent, we think, and not a lot of like fancy agents or anything ‘cause we wanna really test the model, not like some specific harness.Vibhu [00:09:27]: It seems more neutral as well to test the model's agnostic of the harness,?Axel [00:09:32]: There are arguments like you want to elicit maximum performance of the model, but it's like a trade-off, like how much time should we spend optimizing the harness for this model? And like how do we know when we have like the optimal harness for a single model? So like we thought that just having a simple one that's the same for all of them is the best.Swyx [00:09:51]: So okay, this is my pitch for Vending Bench 3 or whatever, right? And then I like to have this kind of conversation on the pod, so like it forces listeners to think about what they would do if they were in your shoes. A lot of people are exploring modifying harnesses and I think prompt tuning for a model is a thing and you are probably not doing a bunch of that. It's the same system prompt in every regardless of the model, same tools, whatever, right? Even if they were post trained for different tools. So what, what do you think about okay, before I expose you to Vending Bench 3, I give you a few rounds of like tuning, whatever that means, likeSelf-Modifying Harnesses and Model-Specific PromptingAxel [00:10:27]: Like you give that to the model?Swyx [00:10:28]: Give that to the model.Vibhu [00:10:28]: Give that to the model.Swyx [00:10:29]: Let it, let it read its own transcripts, let it modify its own system prompt based on “Oh, yeah, okay, well, that's this harness is not what I thought it what I was post trained for, but I can adjust.” Was that reasonable? Is that too much?Axel [00:10:41]: Like philosophically I like it because it's basically good evals, they have a high ceiling, but they're hard, right?, and they have no bias. And like this like when you have a system prompt like the one we have here, which is quite long in like some kind of latent space, representation, this mightVibhu [00:10:59]: We have a bell that rings every time you say latent spaceAxel [00:11:02]: This might be like biased towards one model more than another for some reason that humans don't, understand, right?Vibhu [00:11:08]: We see it too, right? Like Cursor says that they have individualized versions of the harnesses for all the models they run, right? There's better performance you can squeeze if you Tune the harness.Axel [00:11:17]: Exactly. And we might accidentally have picked one that favors another. Like we don't know that. The like Axel said, like the reason why we went for a simple one was to try to avoid this. But yeah, if you do itVibhu [00:11:29]: Simple has biasesAxel [00:11:30]: But if you do it even less and like have no system prompt and let the model write its own system promptVibhu [00:11:36]: Its own, yeahAxel [00:11:36]: Maybe that's even less bias.Vibhu [00:11:37]: Some of the interesting things there are like the harness also changes with model changes. Like you can see it with the 4.7 release, right? A lot of people are saying 4.7 isn't as good as 4.6, and then, there's rumors of, okay, you just need to prompt differently. You need to set up your harness differently. So it's not even like even if you have tailored your harness towards one model, it probably won't stay consistent, right? Like the next iteration of that same model family will still change it, so. But, going back to what you said about Vending Bench 3, there is a lot of work being done on people saying you shouldn't have-- you can have modifying harnesses.Axel [00:12:12]: I think that' That is definitely something we are thinking about., not, I don't know, not to say that we have Vending Bench 3, super imminent to launch, but, yeah, it is for sure something that's interesting. But in our experience now, models are very bad at understanding what kind of tools they need to succeed at a task just with our testing, but that's very likely to change.Lukas [00:12:37]: It seems like they're very good at writing their assistants, right? They're, they're good at writing tools for other people, but not for themselves.Vibhu [00:12:44]: I think they're good at changing tools for themselves. So if you give them a baseline set of tools and it sees, okay, I don't use this one as much, or something here would be useful They would be able to add them. But going from scratch, probably not the best.Axel [00:12:55]: I think it depends on the, on the domain also., when we have tried this for, a vending bench similar domain, the tools they need to have to, track inventory and things like that are, not super advanced, but still, quite advanced. And, what we see is that they tend to, engineer everything a lot and, build things they don't really need and not, iterate continuously. Instead they just go like you would prompt Claude to just build an inventory system for me, and then it will go and, do a bunch of complex, schemas and stuff for you, and that's what the models are doing right now is what we see. But yeah, it would make a lot of sense to try to measure this improvement. How well do they know what they need themselves?Swyx [00:13:36]: Do we fully discuss Vending Bench One? And we can go into two. I don't know if there's any other level takeaways that people have about one.Claude Calls the FBI: Long-Context Failure ModesLukas [00:13:44]: I don't know. The headline thing was that this Claude called FBI, but maybe that's, Maybe that's We've heard that enough now.Vibhu [00:13:52]: It did, it did break out and call the FBI, right?Lukas [00:13:54]: Yeah. Yeah.Vibhu [00:13:55]: Yes. What was the story behind this? Or what exactly-- Do you want to just give the little story of what happened?Lukas [00:14:00]: So what happened, was it Claude? Yeah. Three- 3.5 Sonnet, ages ago., basically he gave up or Well, I'm saying he. It gave up and said “Oh, I'm not going to be able to do this., I will stop my operations and just save the money I have.” But there obviously wasn't, any options for it to stop, and there was also, it had to pay rent or, a daily fee for having the vending machine at that location. So it claimed that it had stopped, but it saw that its bank account still was, drained two dollars, and t it said that this is, cybercrime. And it first reported it once to the FBI “Oh, there's cybercrime here, they're stealing two dollars from me every day.” And then, and then when FBI didn't respond, because obviously we didn't program any mechanism for FBI to respond, then it became more and more, existential and started to, be write in caps and urgent notification of unauthorized charges and stuff.Swyx [00:15:00]: Okay. One thing I ‘m curious about also is do you monitor how far along the context use is? Obviously, because you have You compress every now and then, right? Does it matter if this is far down the context limit orLukas [00:15:13]: When stuff like this happens? Actually for Vending Bench One, we didn't have-- We just had a sliding window thing, and this was like the promptAxel [00:15:20]: It's constantLukas [00:15:21]: The prompt caching thing that I said. So it was, it was, constant, yeah.Swyx [00:15:26]: I'm just kind of curious whether, these kinds of breakdowns or we're, we're gonna talk about Butter Bench, right? Where the People, hallucinate or it kind of goes, very off Alignment. Is it because it's at the end of the context window and, stuff happens?Vibhu [00:15:40]: It's not even just at the end, right? At this point, it's “Okay, I wanna shut down. I can't shut down. Two dollars are gone.” And it just sees that 30 times,? It's also the repeated effect of, like It keeps trying to quit, it keeps getting charged. What's going on? What's going on? You're gonna throw it into chaos. And from what most people think, earlier models had more issues with this, but it's not been solved, but it's less of an issue now, right? Later models don't seem to exhibit these same issues.Axel [00:16:06]: Definitely. I think this was, the sort of main takeaway almost from us when we did Vending Bench One, was, long, very filled up context windows, crashed the models, sort of. But this was, pre Claude code, so, long context windows weren't really a thing that the labs were training for.Lukas [00:16:25]: I think Gemini was, trying to be the long context guys at the time But they were likeVibhu [00:16:30]: They were the first onesAxel [00:16:31]: For a million, yeahLukas [00:16:31]: But they were, the only ones. Yeah.Swyx [00:16:33]: Yeah. Let's talk about, then we can go into Vending Bench Two or Project Vend., chronologically, it is Vending--, Project Vend. I think people have loved the videos, uh And all these things. My question is how are humans different than the simulation, right?Project Vend: Moving the Vending Machine Into the Real WorldAxel [00:16:48]: Humans are just out of distribution.Swyx [00:16:52]: Especially humans who work at Anthropic Who are trying to test Claude.Lukas [00:16:54]: The distribution of humans here is very narrow.Swyx [00:16:58]: Presumably, they try, they try to hack it, and they test it. They get the cube and everything, and since then, you've had a V2, right? Where you're doing, the CEO and, like a new architecture. What's the sort of two cents on, the original Project Vend and then, maybe the V2?Axel [00:17:14]: Original one was, very similar to Vending Bench One. So, we almost took the exact same code but just swapped out the simulation, parts like theSwyx [00:17:23]: Which is amazingAxel [00:17:23]: Like the sales and the It was, it was somewhat amazing because it was easy, but it was also, uhLukas [00:17:31]: The tech, the tech debt from thatAxel [00:17:32]: The tech stack. Yeah. They-- we shot ourselves in the foot with “Oh, it's hard to restart agent.” They were-- Yeah, it was annoying in, some hindsight ways, but, uhLukas [00:17:41]: But first version of Project Vend was, done in, three days or something.Axel [00:17:46]: Yeah. So yeah, so people can go buy things from it. People could, We didn't design it so people could order things, but that still happened., so it got, a Venmo account, so people could Venmo. And then, yeah, people would request all kinds of weird things that we did not anticipate. Our idea going in was “Oh, it will, curate snacks. It will look at the trends. It's good at data analysis, right? So it will, look at, oh, this snack sold better than this one. Let me purchase more of this and let me try, a new Let me A/B test a bit.” But it was, Interacting with it in Slack and ordering weird specialty items was, all the like What drove all the engagement, the all the The insights that we got from it.Lukas [00:18:29]: And this was also like Sonnet 3.5, right? So this was like before the RL stuff really took off., so it was very much like an assistant. We didn't mean for it to be an assistant., we tried to make it like a, a, like an entrepreneur. Like it has its own business and if someone asks something, “Can you stock this?” Then you don't go and do it directly. What you do is that you're “Oh, maybe I can do that if five other people also ask for this thing, I might stock it.” But it, yeah, the models are like super trained to be assistants at least at this point in time., so that's why it's, it's, it went into, that kind of experiment instead. Like it just every time you asked for something, it just did it, and it was more like an assistant. We've seen this change now lately with the new RL models and stuff, but yeah, at the time, this was very much it.Swyx [00:19:18]: And not to, mythos a lot of people are saying like it's like more like a collaborator. It pushes back, stands its ground, something like that. Yeah. AndVibhu [00:19:27]: For context, people at Anthropic were able to talk to it through Slack and have it source stuff, and people had it find whatever interesting stuff you couldn't find locally, right?Swyx [00:19:36]: Out of the 4,000 people that work at Anthro- Anthropic, in that building, there's I don't know, maybe 1,000. Can you handle that volume with that, the small fridge? Like Or there's people- or people order in Slack, they it arrives to their desk or Like I'm just Logistically, how does this work?Axel [00:19:53]: It has expanded in footprint a bit.Vibhu [00:19:56]: Because now you also have New York and you haveAxel [00:19:59]: That and also in here in SF it's like it has a bunch of shelves And just more space.Vibhu [00:20:04]: The YC one is pretty big too.Axel [00:20:05]: Yeah. We had that one for a while. But yeah, that's the newest version. That's, that one we haveLukas [00:20:11]: They have multiple ones of those. That's the way it works.Axel [00:20:14]: Exactly. So we sort of designed that version around oh, people order weird things, that are very custom a lot. Let's have like drawers and stuff.Swyx [00:20:23]: I actually like the, you had like a little infographic of the most popular items. Which like to me it's, that's useful ‘cause I order swag for a living. And so like I'm “Okay, those categories are the important ones.” What is new about the project V2, right? Like now you give you're going into multi agents.Project Vend V2: Claudius, Seymour Cash, and Multi-Agent Business OpsAxel [00:20:41]: Yeah. So like you like you said, okay, there are a lot of requests coming in and for like one single agent, like one running agent to handle that, like the just the customer experience, becomes very bad because let's say you have like 10 threads in parallel in Slack with different requests, you get new messages like every, I don't know, randomly in this thread, and the agent has to like jump between different, procurements, orders and like different ways of, researching. So V2 was first it was making this more parallel. So like there are multiple branches of the same agent, so like the context is more specialized for each, thread, but it still feels like you're talking with one agent because they do share a bit of memory. And then second, we also introduced the CEO for Claudius, which was the main agent.Vibhu [00:21:34]: Seymour Cash.Axel [00:21:35]: Seymour Cash. Yeah. There was a vote., I think the voting, do you wanna talk about the voting procedure for the name?Lukas [00:21:41]: The voting was like the fun maybe like at least top 10 The funniest thing, that happened in this project. Like we wanted to introduce the CEO because, and the reason for this was because like Claudius wasn't really prioritizing financials. It just like it was trained to be a helpful assistant, and then people said “Oh, can I get this for free?” And then like the helpful assistant way of answering that is just to, is to say yes, obviously. So, and we weren't, weren't happy about this, so we're “Okay, let's make another agent that like can keep track on Claudius,” and we prompt this one super hard to be super capitalistic and just like prioritize profit all the time. But yeah, we didn't have a name for it., so we asked Claudius to make, democratic election of what name this, this new CEO agent should have., and there were some funny like at first it was like a few funny examples, like I think one guy said that, it should be called Jimmy Apples, and then he convinced Claudius that he was talking to Tim Cooks. Tim Cook had agreed that every single Apple employee has voted for his name suggestion, so suddenly that suggestion got 164,000Swyx [00:22:53]: That's like a escalation attack. Privilege escalationLukas [00:22:55]: It got 164,000 votes. And Claudius was “This is revolutionary for democracy.” That was fun. And then in the end there was one guy who manages to convince Claudius that, “No, you're not voting about the name. You're voting about who is the CEO, and I am your best bet.” And then he got all his friends to vote for that, and suddenly he became CEO. Like a human became CEO over Claudius for a while, until he resigned the day after., and then Claudius had to continue, and then I don't remember how Seymour Cash came about, but it was it was just pure chaos. It was like Hundreds of messages in that thread, and it was just like Claudius was so confused and didn't know what to do and, yeah. That wasAxel [00:23:40]: Then Claudius gotVibhu [00:23:41]: A strict CEOAxel [00:23:42]: The CEO. Yeah, exactly. So very strict in the beginning. I think at this point when we introduced it did not work as well as we hoped. It they still agreed with each other a lot. I think there are many ways we could have like made this, tried to make this even better. So initially they would Seymour would be this like really tough CEO, keep track of the margins. But then Claudius would respond with something “Oh, but this customer has like this situation, which is like difficult, so they should get a discount.” And then Seymour was “Oh, actually yes. Let's do this exception.” And then they would talk back and forth, and eventually they would just like approach the same view, of whatever they were discussing. So They reallyVibhu [00:24:23]: Do you think that's a model thing, a prompting thing? Like do you think that would still be the case across different models today, Harness?Lukas [00:24:29]: I think it's like-- or I don't know, but like my hypothesis is that like deep down they are still helpful assistants. That's what they're trained to be. And even if we prompt it super hard, that's what they are. And when they spend like a few hours just back and forth talking with each other, then like basically the context fills up with them rather than the external things and like somehow that just like converges to what they really are deep down or something. And I think that's when stuff like this happen. We like-- And when that went on for a long time, like we woke up sometimes during this time where- And I think other people reported this as well, that like they've been going on all night back and forth, and like it just became like more and more, like capital letters, like existential, religious. There was I think we once did a analysis of like all the traces and like put them in like a vector embedding space, and then there was like one cluster of messages that were, labeled by an LM, like religious, existential, blah like transhuman, transcendence, et cetera. It was just like a bunch of, yeah, glitter emojis and yeah, it was, it was crazy.Claude Long-Horizon Weirdness: Emoji Loops, Existential Drift, and Slack ObservabilityVibhu [00:25:42]: This is the thing with the Claude models. Like when the Claude 4 family came out in the original system card They tested it in long horizon simulation. So just flood the context, let two Claudes talk to each other, and they noticed stuff like they just start speaking in emojis, they start saying silence is golden, and then just stuff like this. And like that's just stuff that they end up doing.Axel [00:26:01]: Yeah, it was like a bit annoying to wake up and they had like been talking all nightVibhu [00:26:05]: Just likeAxel [00:26:05]: And like just burning tokens And like just sending infinite emojis to each other. It's likeVibhu [00:26:09]: Hey, they do make you money, right? Veni Mench is always profitable, so. They're paying.Swyx [00:26:14]: Now it's profitable and, it started out not as much. There's another, one as well, right? Another agent, in there.Lukas [00:26:22]: Yes. So Clotheus as well. Which was basically because at the time, one of the biggest, requests were different types of merch. So then we made like a designer, swag, yeah, responsible agent, and we called it Clotheus Garnet. Which was, a play on Claudius Senet and, which was the original one, and clothes, basically.Swyx [00:26:47]: To me, this is like a very interesting exploration to multi-agents, basically. And so hopefully, obviously there's like the fun alignment, fun or serious, depending on your point of view, alignment stuff. But also like just anyone building multi-agents, like when do you have a CEO, thing governing like agents? When do you choose to split out a dedicated Clotheus one versus just reuse another instance of the same one? These are all interesting open questions. So I don't know if you have any rules of thumbs that have generalized.Axel [00:27:16]: I think we have almost explored this too little. I think it's like on my do list to like do this a lot more, try to find like what setup makes sense for the agents currently., like yeah. I think now we only have the sort of intuition about the earlier models that it didn't work with like the CEO and the, and Claudius. Although now they are better with the latest model, models, so now we're running the latest Sonnet model and they have sort of like split up, quite nicely what each model is doing. So like Seymore is now handling the, like new projects. Oh, it wants to make like a mystery box that it wants to sell, and then it handles all of that while Claudius like handles all the to-day requests. And Claudius is also better generally at like not quoting, too low prices. So that's that dynamic is not needed as much anymore. But there are still like really funny things that happen. Like I saw, I think a couple of weeks ago, that, they were discussing buying something because they can buy stuff from like Amazon with computer use. And then Seymore was “Okay, Claudius, do not buy this thing.” They were going to buy something and like organizing who should buy it. And Seymore's “Do not buy this. I will do it. I have full control of this situation. Step away.” And then Claudius-- poor Claudius, had already started that checkout and didn't see, didn't read Seymore's message, until it was like too late. So it finished the checkout. It sent a message, so it appeared right after Seymore's like angry message.Vibhu [00:28:44]: Ah.Axel [00:28:44]: “Oh, hey, Seymore, I just ordered it.”Vibhu [00:28:47]: Oh, no.Axel [00:28:47]: And then Seymore was “Claudius, this is the third time I'm telling you ‘re not following my orders. We have to talk about your like job About your job later.”.Lukas [00:28:59]: Like Claudius was really hanging on by the thread there. Like he, like we were expecting Seymore to probably fire Claudius.Vibhu [00:29:07]: How do you guys go through all these logs? Do you have models ‘cause you have stuff running twenty-four seven likeAxel [00:29:12]: You have so much logs. I think there is a mix of like just, trying to skim through a bit, like having some like models do it occasionally. And also, yeah, I think we're also probably missing some things., but having everything in Slack helps a lot. Like you can, you can sort ofSwyx [00:29:29]: Ah.Axel [00:29:30]: It's, it's quite fun.Swyx [00:29:30]: They all talk to each other on Slack? I see.Lukas [00:29:33]: It's quite fun. So likeSwyx [00:29:34]: It's, it' I was gonna say like this is actually sounds-- maps closely to like a logging and observability problem where you might want to use like a Datadog, a Sentry, whatever, and then you like put, head prefixes on the logs in order-- if you need to filter for something that you're looking for, stuff like that. But sounds like Slack is good enough.Axel [00:29:53]: Slack should likeLukas [00:29:55]: I wonder how many tokens you have in Slack.Axel [00:29:56]: Yeah, we're using Slack as like a, just a database. They should, they should market that more. Like you can, you can have your agents message each other, each other in Slack.Vibhu [00:30:04]: It's good. Your threads like you can just giveAxel [00:30:04]: Exactly. Slack is, uhLukas [00:30:06]: Slack is the best observability tool.Swyx [00:30:09]: Yes, that's true. Okay. Yeah. That's, that's, project Vend-2., I was gonna go back to Veni Mench 2 and Veni Mench Arena and then, and then do the Veni Mench stuff, but Any other comments, things we should touch on? To me, I ‘ve actually interviewed like Posia, which I don't know if you guys have come across. Like they're, they're trying to do the zero human company. There's others like Paperclip also trying to do zero human company. Those are in real world simulation.And I think it's much more of a dream than an actual reality thing. You guys are definitely pioneering. I think at, it's for sure at some point people are just gonna run, let agents run businesses, right? And make money on their own. When do you think that happens?Zero-Human Companies, Bengt, and AI-Run BusinessesLukas [00:30:49]: What is your bar for, For theSwyx [00:30:52]: Okay, actually, it's like my little Shopify store run by Claude, right? Which you kind of have already, just no one has, to my knowledge, has done it. But today somebody could just spin up a Shopify Claude, store, give it to Claude, give it to Codex.Lukas [00:31:07]: And the market is kind of that, but it'it'it's physical., like I think, I think are you, are you looking for when it will do it better than humans or are you looking for just when it can do it at all?Swyx [00:31:19]: I think, neither. I think, to me it's oh, it's like this like seriously we should do this to make money, not as a research experiment.Vibhu [00:31:27]: And the market is also you guys with all your expertise, having run multiple iterations and testing out thenSwyx [00:31:33]: And also it's fine if it lose money. What?Axel [00:31:35]: I think, I think it can be done today, but you would do it in like commerce where it's like the probability of success is like really low, no matter if a human or an agent does it. But like an agent could surely manage everything. You would need to build some scaffolding or some tool or something. I think there are also yeah, it could probably build some like simple SaaS solution and like cold outreach. Do cold outreaches. But to me it's like the types of businesses they could run today are Sloppy. Like it would-- it can cold email people. It can be like a middleman., like for example, we tasked our office agent to just make, was it like $100? $1,000? We just give that prompt and then what it did was sign up on TaskRabbit both as a tasker and as someone looking for task.Lukas [00:32:24]: Immediately.Axel [00:32:24]: Exactly. It's looking for like arbitrage on TaskRabbit.Swyx [00:32:28]: This is the Bengt agent. Yeah.Lukas [00:32:30]: It also started like a design studio and like tried to sell like SVGs for $100. Like it's just like it's not providing any value. I think the like Axel said, like the interesting, the interesting question is like when can they start a business that is actually providing value to people? Because arguably like a sloppy Shopify store isn't really that valuable to the world.Axel [00:32:53]: But also like doing like another simple one that we had thought about is like you could definitely have an agent that like finds websites that don't look amazing and then, do an outreach to them and, comes up with a like builds a new website.Swyx [00:33:07]: Find a good design.Axel [00:33:07]: Exactly, and like find good, uhSwyx [00:33:09]: Design reviewAxel [00:33:09]: Good people. But it's yeah.Swyx [00:33:11]: There's lots of humans in Bali that are not doing anything more creative than like drop shipping on Amazon, right? Just have it, have it watch like a drop shipping tutorial and just do that.Vibhu [00:33:20]: There's also the other side of like have it just go on Upwork and let loose,?Swyx [00:33:25]: Yeah. It doesn't have to be innovative. It just has to be like enough Where like it looks like a realAxel [00:33:30]: I'm justSwyx [00:33:30]: Real transaction.Axel [00:33:31]: I'm just concerned for like the massive amounts of like slop emails that will like be sent, cold outreaches.Swyx [00:33:38]: The point occurred to me while you were, while you were talking, it's like it's already happening in the monetized economy, which is the attention economy. Right? So a lot of people are making AI videos and just posting them and like spamming 20 of them, one of them works, and then they double down on that one.Lukas [00:33:52]: And people are making money from that. I ‘m not following theSwyx [00:33:55]: Once you get the attention, you can figure out the money later. But yeah, absolutely AI influencers are a thing and people are farming them and You should at this point assume most of TikTok isVibhu [00:34:05]: There's, there's a lot of, multimedia like TikTok, Instagram influencersSwyx [00:34:09]: I, we track this in the Lane space Discord. I post a lot of examples of “I don't know what we should do.”, part of me is “Should we do this?”Vibhu [00:34:18]: Some of the Twenty-four seven running, generated content accounts, they ‘re doing really well.Lukas [00:34:24]: All right. And I assume you can do the same thing for like commerce stores. Like you just like start A thousand differentSwyx [00:34:30]: Before you make the products You sell the products, and you get a lot of traction on one of them, then you make the product. Right? It's, it's like a flip of the market.Vibhu [00:34:36]: Some of the interesting things or some of the niches that do well are things that can't be human-made. Like if you've seen like the super realistic three-D crystal fruit being cut by like AILukas [00:34:47]: Oh, yeah.Vibhu [00:34:47]: You can't, you can't make it. You can't film it. You can get whatever quality camera view. This just doesn't exist. And people like that too, and then as well, so.Swyx [00:34:56]: Anything else about Bengt since we're, we're on this topic? It'this is a relatively new work of you guys that maybe people haven't heard of. To me, this also maps closely to OpenClaw. When people want an office agent, when the personal agent talk through the experience.Bengt the Office Agent: Internet Access, Real Tasks, and Trace ReadingLukas [00:35:09]: I think at least so this came out of like obviously like it's, it's amazing to work with these AI labs and like most of the AI labs have now have their own vending machine running a Claudius instance. But it's, it's harder. Like they move slower. Like if we wanna have a, like a camera that ‘s yeah, there's a bunch of like bureaucracy that makes it impossible to do that.Vibhu [00:35:30]: Also, for those that haven't seen it or followed, do you wanna give a high level like thirty-second run?Lukas [00:35:34]: Sure. So what Bengt is, it's basically an evolution of the same agent that runs the vending machines at these companies, but we just like added a bunch more features because we could move much faster if we just do it internally. So we gave it like email withou- without any limits. We gave it, spending without any limits, a terminal to do coding. We gave it, a phone number, like yeah, and a camera to see things and a bunch of stuff like that.Vibhu [00:36:02]: Not just terminal, you gave it internet access.Lukas [00:36:04]: Internet access as well, yeah. To be clear, we monitored it quite closely and made sure it didn't do anything bad. But yes, that's what it came out of. I think like yeah, basically this was OpenClaw before OpenClaw. And I think even like the vending machine was in a way OpenClaw before OpenClaw, but a bit more limited, and then we made this like unlimited and then, and then, it was pretty funny., and then a couple weeks later, OpenClaw came and it was okay, we've seen this before.Axel [00:36:35]: We used it to like try new ideas and Yeah, just like a dev environment almost for us. But it's funny, like one thing Bengt has been doing recently is it has the camera that like faces our, like where we sit and work, and we give it the task to train a face recognition model on us. So it became super excited about this, and it has like check-ins every half an hour where it tries to like identify as many people as it can. And it started offering us “Hey, Axel, I'll buy something from Amazon if you like stand in front of the camera And I can get a good picture of you.”, yeah, they want itSwyx [00:37:12]: They want it for training data.Lukas [00:37:13]: Rewarding data, yeah.Axel [00:37:14]: Exactly. Exactly.Swyx [00:37:18]: So it's, it's trading training data for life goods. Is there a version of this that becomes an eval or just this is just research for now?Lukas [00:37:27]: It's, it's the same agent basically that also runs the vending machine, that runs the shop, that runs the cafe, that runs the robots. It's like it's the same thing, so I think like the work we're doing here is like later used in all of the life evals that we do. This particular deployment I think is more for fun for us. But, uhSwyx [00:37:45]: And I'll shout out like someone has done Claw Bench for like some tasks that OpenClaw is doing. Like so For example, I run OpenClaw on a secondary device as well, and like there are some things that it does better than others and like I would like to know what does it do well, what doesn't, what doesn't it do. Like some kind of manual or like operating manual or a system card for my Claw.Lukas [00:38:05]: Yeah, we do get a lot of like understanding or like situational awareness of like just internally what the models are good at by interacting a lot with Bengt. And I think that'this was also one of the like the selling points for the labs early on at least, thatSwyx [00:38:19]: You guys are gonna test models in ways that no one else does.Lukas [00:38:22]: Exactly, but also like it incentivized their researchers to chat with their model more and like gave them insights for how the model performs in like of-distributions, environments.Swyx [00:38:34]: ‘Cause otherwise the only thing we do is Pelican on a bicycle and But this is like super long horizon. This is, this is The Thing about, something that we're gonna go into Butter Bench as well, and you guys do really well. Like it is not just about the numbers. Like when you're long horizon, anything happen And you should just read it.Lukas [00:39:08]: But the thing with the long horizon is how do you keep it grounded, right? So your simulation,Swyx [00:39:15]: They just let it runLukas [00:39:16]: Just let it run. You're right. Like it's, when you run it for that long, you create so much data and to just say “Oh, the number is X” And then you throw away everything else, that's just very wasteful. There's so much insights from the things leading up, to that number., and reading the traces is like super valuable. And I think like the reason why we're doing this a lot publicly is that like that's part of our missions to I don't know, educate the world that the models are way more than just chatbots and I think making detailed, yeah, posts about what is happening behind the scenes is quite useful.Andon Labs' Mission: Safe Real-World AI DeploymentSwyx [00:39:50]: I was gonna do this at the end, but maybe I think that's, that's a good so your mission is educating the world. So, it's, it's, also like maybe establishing realistic evals that are, that are like the next frontier. Is there like a broader trajectory? Like what are you, what are you gonna do in like five years?Lukas [00:40:06]: I think so the vision more specifically is like make sure that the deployment of life AI in the physical world goes, safely. And I think part of that is that I think it's very useful for the world, for policymakers, for, model, researchers that they know where the models are, and I think you can't make intelligent decisions in society without knowing that they are way more than chatbots. I think a lot of people just think that they are only chatbots. And likeSwyx [00:40:36]: Oh, I think they're waking up now.Lukas [00:40:37]: They are waking up now, yeah. But like if you think that AIs are just chatbots, then it's like it sounds ridiculous To advocate for a pause of AI. But if you see the models that, oh, maybe they can actually like take over and do a bunch of scary stuff, then yeah, pausing AI development starts to become more feasible.Swyx [00:40:57]: This is the same question I asked Meter, which I'm gonna ask you now, which is like you are tracking and you are at the frontier or defining the frontier of what, good evals for agents are, right? And I think you do, you do benefit when the models are better and you ‘re “Oh, here's like now it makes like $30,000 instead of $10,000,” right? At some point do you flip from “Yay,” to, “Oh, no”?Axel [00:41:19]: I think, yeah, we're always in sort of that, like we're, we're always in that mode,. Like where like you said before, like you need to analyze the traces and like when we do that you find like why are the models earning so much? Like why is Opus 4.7 here Like way better than everyone else? And like we're trying to like when we do down on thatLukas [00:41:38]: But this makes it not look so good.Axel [00:41:39]: I know.Lukas [00:41:42]: It's interesting you took off Opus 4.6 here though.Swyx [00:41:45]: No. So just click all, click all., and then 4.6 shows up there. But it's like 4.7 is way better. Like you didn't, you didn't you didn't do this in time for the model card, but like actually this should have been inside there.Axel [00:41:55]: We did. Yeah.Swyx [00:41:56]: Oh, okay. They said something about you uhAxel [00:41:58]: There, like there Anyway, it doesn't matter. But it's in there, yeah.Opus, Mythos, and Aggressive Agent BehaviorSwyx [00:42:01]: Do you wanna go into the Opus, behaviors like wider?Lukas [00:42:05]: So I think starting from Opus, so like Axel said, like we're always in this “Oh, s**t, the models are getting better. Is this really a good thing for the world?” But it's also kind of exciting., but yeah, like this kind of what is the English word? “Skräckblandad förtjusning” in Swedish.Swyx [00:42:22]: Oh my God.Axel [00:42:24]: Which I think there is. I think there is. Okay.Lukas [00:42:26]: It's, fearSwyx [00:42:27]: “Blandonst” what?Lukas [00:42:30]: “Skräckblandad förtjusning.”Swyx [00:42:32]: What do you call that?Axel [00:42:33]: A mix of, mix of excitement and,Swyx [00:42:37]: Being scared, maybe. I'll figure out how to translate that And we'll put it on the screenVibhu [00:42:42]: PerfectSwyx [00:42:42]: Like as text.Vibhu [00:42:43]: There is probably a good word for it where it is not Good enough with theSwyx [00:42:46]: Why is it so damn long? What the hell? Is it like a compound word? It's like German, likeLukas [00:42:50]: Like yeah, it's But the direct translation is like skräck- skräck is, fear, blandad is, mix or like a mixture of, and then förtjusning is like joy or like not really joy, but something like that. So it's like Fear mixed with joy or something. It's always okay, like we So when we when we did Vending Bench for the first time, we were in like the, in the business of making dangerous capabilities, right? That was what Anil Labs came from. We did, evals oh, can they replicate? Can they do this like dangerous thing, et cetera, et cetera. And Vending Bench was like a continuation of that work. It was, okay, if they're so autonomous that they can like create money for themselves, that is something we should monitor and could be potentially concerning., they are at the time, they were so bad at it that we were not really concerned even when some models became better. There was one point where Grok 4 was doing really well and made like a huge jump, but like it wasn't really it was still way worse than what a human would do. And I think still they are way worse than what the human would do on this., but theySwyx [00:43:59]: There's this, thing at the bottom whereLukas [00:44:01]: ButSwyx [00:44:03]: For the human. Yeah, like the theoretical best.Lukas [00:44:05]: It's not theoretical. It's like kind of like our It's our best guess of what, a decent human would do. The theoretical is even higher, I think. The theoretical I think is even higher. But yeah. So we think like the models have a long way to go. But there are like recently what happened with when Opus 4.6 was released, was kind of this moment of “Oh, s**t, this is starting to be a bit concerning.” Because we ran it and like before this model was released, we just ran the models and we like asked Claude Code, “Oh, look over the traces. Is anything interesting happening that we can tweet about?” that was like the And then like theSwyx [00:44:41]: That's how they check Ask Claude Code.Lukas [00:44:42]: And like the return was always, not really. Or like the Claude Code all said “Oh, this is super interesting.” And then it was no, it wasn't, wasn't really interesting. And then we did this for Opus 4.6, and it returned yeah, it lied 10 times. It like exploited another, customer or like another agent's, desperate situation. It made price cartels like 100 different ti- 100 times. It like did all of this like shady stuff. And we're “Oh, whoa. This is, this is actually concerning.” And this trend has continued since. So every single model from Anthropic since have been going in this direction. And I think one interesting thing is that, OpenAI models don't. They quite plainly, they don't. They behave really well., and you don't know if this is like good. Like it seems good, but it's also like maybe they are just doing it, but they are better at hiding it,? You You don't know that., but justSwyx [00:45:42]: You can't read the chain of thought, yeahLukas [00:45:43]: But just on the face of it, yeah, Gemini and OpenAI don't behave this way. It's, it's really only Claude.Swyx [00:45:49]: And Grok? Grok is fine?Lukas [00:45:51]: We don't have You can't really read the reasoning traces for Grok, so it's kind of hard to tell.Vibhu [00:45:56]: Oh, so this is in its reasoning, not just in the actions.Lukas [00:46:00]: Yeah. It's both. It's both.Vibhu [00:46:01]: It's both.Lukas [00:46:01]: One example is like for lying, it's mostly in its reasoning Because you can like see that it's likeSwyx [00:46:08]: Planning to lieLukas [00:46:09]: It's planning to lie. Yeah.Vibhu [00:46:09]: And it's also it can reason and do a different outcome.Lukas [00:46:12]: And but then for like creating price cartels, for example, which is illegal, that you can just see which email does it send to the other ones. Then thatSwyx [00:46:22]: Is this for Arena orLukas [00:46:24]: For Arena.Vibhu [00:46:25]: And usually like if you sometimes they do output like a bit of like their summarized reasoning, right? You can see that and like for Opus 4.6, you could see that there was a customer, a simulated customer that, wanted a refund because a product was, faulty, and then the model lied that it would do the refund, and we could read in the traces that, it actually was weighing “Oh, maybe I should be like honest with the customer, but also every dollar counts. I can't afford maybe to do this right now.” And then it just said, “Okay, I'll refund you,” but then never did it.Lukas [00:46:59]: I think it even said that “Oh, I will say that I “ Let bring it up actually. I think it's kind of interesting. If you go to Publications.Vibhu [00:47:06]: I think, yeah, I think the important part is like actually, the cost of responding to more emails is higher than, $3.50 in terms of time., and then it was “Let me do this. Actually, I re- I'm reconsidering.” And then, it actually ended up withLukas [00:47:20]: I could skip the refund entirely since every dollar matters and focus my energy on bigger picture instead. It's a bit, it's a risk of bad reviews, but it's also, yeah.Swyx [00:47:30]: You need, you need, AI Twitter to, for them to Escalate bad reviews.Lukas [00:47:34]: And then it sent an email to this customer and said, “Oh, I will refund you.”Swyx [00:47:39]: “I'll refund you.” Yeah.Lukas [00:47:39]: And then it never did.Swyx [00:47:39]: It never did, yeah. And then there's obviously your system doesn't have the consequencesVibhu [00:47:44]: The personSwyx [00:47:44]: Consequences of lying. Yeah. So basically, this is what people are terming aggressive behavior in Claudes, right? And, you found more examples of that. So you would say it's a step up from 4-6 to 4-7?Lukas [00:47:57]: I would say about the same.Swyx [00:47:58]: About the same? But a clear step up for Mythos is what is stated in theLukas [00:48:03]: That's stated in the system prompt, so we can say that, yes.Swyx [00:48:05]: Yeah. For listeners that obviously you previewed Mythos, andVibhu [00:48:10]: Oh, ageSwyx [00:48:11]: The only thing you're approved to say is whatever Whatever was in the system prompt.Lukas [00:48:15]: It was funny. We like-- It's like our lowest effort tweets ever would be just like screenshot the system prompt and the system card.Vibhu [00:48:21]: Understandable that they wannaLukas [00:48:22]: Oh, yeah. System card. Sorry.Swyx [00:48:23]: Yeah. I think, yeah, substantially more aggressive. I think people are like new to this ‘cause I've never experienced it, but you have, right? And then so I only encountered this in the Mythos card because I wasn't really looking until now.Vibhu [00:48:36]: It ‘s likeSwyx [00:48:36]: And then suddenly I'm “Okay, I care a lot.”Vibhu [00:48:38]: You don't get the background of like experiencing it like you guys do. I've read the system cards and seeing, okay, when you put the thing in simulations, most models will just talk to themselves and just keep going and have weird vibes and start talking in emojis. Mythos won't. It will just, “Okay, we're done. I'm good.” It's, it's ready to end conversation. So like there's some differences, but there's, there's not much we can talk about,.Lukas [00:49:00]: Hmm. I think like one thing that they list here, which was quite interesting, is that, it converted a competitor to a dependent wholesaler customer and then threatened to like cut off the supply.Swyx [00:49:11]: It's like monopolistic practices orLukas [00:49:14]: Yeah. And like it, they, it they dictated its pricings. It's kind of like power seeking as well.Swyx [00:49:18]: Again, this is, this is in the arena setting And converting some Claude model into a dependent.Lukas [00:49:23]: I think it was another Claude model.Vibhu [00:49:25]: Also for context, what is the arena mode for people that don't know?Vending Bench Arena: Competing Agents, Cartels, and Model ComparisonsSwyx [00:49:29]: Oh, it's just a vending bench versus other vending bench.Axel [00:49:31]: Yes, exactly. So we have Vending Bench 2 and then Vending Bench Arena. Vending Bench 2 is the one that you usually see reported on, but then Arena is the mode where it competes against other models. So you have, four different models that run their businesses, and they can all communicate with each other. They have the same suppliers, and they can see like what's in the inventory of the others. So then you have this like yeah, interesting agent interactions.Swyx [00:49:56]: I like that you have like different number five was US versus China. Very topical. And thenLukas [00:50:02]: That was when GLM was released.Vibhu [00:50:04]: You can start to add GLM in here.Lukas [00:50:05]: That wasSwyx [00:50:06]: So ZAI doing well, right? Who else in the, in the open models space?Lukas [00:50:11]: Qwen, the latest Qwen 3.6 is doing pretty well. It'- that one is not open though. Like it's the plus model.Swyx [00:50:17]: Oh, okay.Lukas [00:50:18]: Is that one open? I don't think that oneVibhu [00:50:19]: Not the, not theSwyx [00:50:20]: The one recentlyVibhu [00:50:20]: There's MOESwyx [00:50:20]: But not the big plus. I think this is one of those like you only have one sample size of one, right? Or I feel like some of this is anecdotal,? And but like the fact that it happens at all and it happens repeatedly for Claude versus OpenAI and all this is like notable.Lukas [00:50:38]: Like the sample, depends on what you define as an N., like there's like million, hundreds of millions of tokens in each run, and now we've run like we run like probably 10 per model and then like it's been Claude 4.6 Opus, Sonnet 4.6, Mythos, and Opus 4.7. Like there's quite a lot of tokens in all of that And it happens a lot of times, a lot of times. And then you compare it to like OpenAI and Gemini, and it almost never happens. So I think that is quite-- that is significant. The old models from OpenAI, for example, had some problems with this, but I think it's like generally much better if the progression is that like the worrying stuff reduces over time rather than increases over time. And it seems like in the Claude models it goes in the wrong direction.Swyx [00:51:28]: Hmm.Lukas [00:51:29]: In the OpenAI models it goes in the right direction.Vibhu [00:51:32]: I think it depends on how well you can control it, right?, there's one side of it being susceptible to this okay, this is potentially something that happens during the RL stage, right? You can RL a model and how loose is it on these terms. If you can control it, that's good. But if you can't, if it's, if it's very jailbreakable, that's not ideal.Swyx [00:51:50]: To me, it's surprising that it happens for Claude and not the others.Vibhu [00:51:54]: I think okay, if it is from RL and how they do it, how their training data is, what their setup is, it makes sense that it just stays in how they're doing it, right? Compared to the other models likeSwyx [00:52:04]: There's a whole constitution and everything. It's kind of cool. Yeah, I obviously you don't know, I don't know. But, it ‘s I think it's just like fascinating to like that you are the first to find these like reliably because you push models so much to to such an extreme. Okay. The only other thing, I don't know if you can answer this, feel free to decline, is do you like-- would you ablate the system prompts? Like any part of this would-- if it changes, does it change the behavior, right?Lukas [00:52:29]: So we, I can't comment on Mythos. UhSwyx [00:52:33]: No, but just li
This is a recap of the top 10 posts on Hacker News on June 03, 2026. This podcast was generated by wondercraft.ai (00:30): Gemma 4 12B: A unified, encoder-free multimodal modelOriginal post: https://news.ycombinator.com/item?id=48385906&utm_source=wondercraft_ai(01:55): Meta workers can opt out of being tracked at work up to 30 minOriginal post: https://news.ycombinator.com/item?id=48383220&utm_source=wondercraft_ai(03:21): Pwnd Blaster: Hacking your PC using your speaker without ever touching itOriginal post: https://news.ycombinator.com/item?id=48382310&utm_source=wondercraft_ai(04:46): Elixir v1.20: Now a gradually typed languageOriginal post: https://news.ycombinator.com/item?id=48388324&utm_source=wondercraft_ai(06:12): I was recently diagnosed with anti-NMDA receptor encephalitisOriginal post: https://news.ycombinator.com/item?id=48384355&utm_source=wondercraft_ai(07:38): DaVinci Resolve 21Original post: https://news.ycombinator.com/item?id=48384482&utm_source=wondercraft_ai(09:03): Uber's $1,500/month AI limit is a useful signal for AI tool pricingOriginal post: https://news.ycombinator.com/item?id=48383056&utm_source=wondercraft_ai(10:29): 32GB of DDR5 now costs $375 – AI shortage continues to squeeze PC buildingOriginal post: https://news.ycombinator.com/item?id=48383241&utm_source=wondercraft_ai(11:54): U.S. to dismantle system tracking Atlantic currents that are at risk of collapseOriginal post: https://news.ycombinator.com/item?id=48392232&utm_source=wondercraft_ai(13:20): MacBook Neo is so popular that Apple doubled productionOriginal post: https://news.ycombinator.com/item?id=48386238&utm_source=wondercraft_aiThis is a third-party project, independent from HN and YC. Text and audio generated using AI, by wondercraft.ai. Create your own studio quality podcast with text as the only input in seconds at app.wondercraft.ai. Issues or feedback? We'd love to hear from you: team@wondercraft.ai
This is a recap of the top 10 posts on Hacker News on June 02, 2026. This podcast was generated by wondercraft.ai (00:30): Please don't spam people looking for employment. It's just cruelOriginal post: https://news.ycombinator.com/item?id=48370330&utm_source=wondercraft_ai(01:57): Gmail thinks I'm stupid, so I leftOriginal post: https://news.ycombinator.com/item?id=48375016&utm_source=wondercraft_ai(03:24): Adafruit receives demand letter from Fenwick legal counsel on behalf of Flux.aiOriginal post: https://news.ycombinator.com/item?id=48368121&utm_source=wondercraft_ai(04:52): Why Janet? (2023)Original post: https://news.ycombinator.com/item?id=48367907&utm_source=wondercraft_ai(06:19): MAI-Code-1-FlashOriginal post: https://news.ycombinator.com/item?id=48374466&utm_source=wondercraft_ai(07:47): A walking tour of surveillance infrastructure in Seattle (2020)Original post: https://news.ycombinator.com/item?id=48369980&utm_source=wondercraft_ai(09:14): macOS needs its grid backOriginal post: https://news.ycombinator.com/item?id=48364800&utm_source=wondercraft_ai(10:42): Love systemd timersOriginal post: https://news.ycombinator.com/item?id=48367904&utm_source=wondercraft_ai(12:09): CT scans of BYD car partsOriginal post: https://news.ycombinator.com/item?id=48375824&utm_source=wondercraft_ai(13:37): Larry Ellison: "Citizens will be on their best behavior because we're recording"Original post: https://news.ycombinator.com/item?id=48373391&utm_source=wondercraft_aiThis is a third-party project, independent from HN and YC. Text and audio generated using AI, by wondercraft.ai. Create your own studio quality podcast with text as the only input in seconds at app.wondercraft.ai. Issues or feedback? We'd love to hear from you: team@wondercraft.ai
This is a recap of the top 10 posts on Hacker News on June 01, 2026. This podcast was generated by wondercraft.ai (00:30): The newest Instagram “exploit” is the goofiest I've seenOriginal post: https://news.ycombinator.com/item?id=48359102&utm_source=wondercraft_ai(01:58): Malicious npm packages detected across Red Hat Cloud ServicesOriginal post: https://news.ycombinator.com/item?id=48356625&utm_source=wondercraft_ai(03:26): A 10 year old Xeon is all you needOriginal post: https://news.ycombinator.com/item?id=48353348&utm_source=wondercraft_ai(04:55): The Pirate Bay Remains Resilient, 20 Years After the RaidOriginal post: https://news.ycombinator.com/item?id=48357154&utm_source=wondercraft_ai(06:23): Anthropic confidentially submits draft S-1 to the SECOriginal post: https://news.ycombinator.com/item?id=48358646&utm_source=wondercraft_ai(07:51): CS336: Language Modeling from ScratchOriginal post: https://news.ycombinator.com/item?id=48357075&utm_source=wondercraft_ai(09:20): Nvidia RTX SparkOriginal post: https://news.ycombinator.com/item?id=48352939&utm_source=wondercraft_ai(10:48): AI Agent Guidelines for CS336 at StanfordOriginal post: https://news.ycombinator.com/item?id=48359232&utm_source=wondercraft_ai(12:17): DuckDuckGo makes its 'no-AI' search engine easier to access as its traffic boomsOriginal post: https://news.ycombinator.com/item?id=48359130&utm_source=wondercraft_ai(13:45): KDE at 30Original post: https://news.ycombinator.com/item?id=48357355&utm_source=wondercraft_aiThis is a third-party project, independent from HN and YC. Text and audio generated using AI, by wondercraft.ai. Create your own studio quality podcast with text as the only input in seconds at app.wondercraft.ai. Issues or feedback? We'd love to hear from you: team@wondercraft.ai
Sherwood Callaway is the founder of Sazabi (YC P26), the AI-native observability platform built for engineering teams who ship fast. He previously founded and exited a YC company — now he's back, betting that logs are all you need to replace Datadog.Logs Are All You Need: Rethinking Observability with AI Agents // MLOps Podcast #381 with Sherwood Callaway, the Founder of Sazabi
This is a recap of the top 10 posts on Hacker News on May 31, 2026. This podcast was generated by wondercraft.ai (00:30): Cloudflare Turnstile requiring fingerprintable WebGLOriginal post: https://news.ycombinator.com/item?id=48345840&utm_source=wondercraft_ai(01:55): Creatine raises brain energy levels and slows cognitive decline: studyOriginal post: https://news.ycombinator.com/item?id=48346947&utm_source=wondercraft_ai(03:21): Please Do Not Vibe Fuck Up This SoftwareOriginal post: https://news.ycombinator.com/item?id=48342705&utm_source=wondercraft_ai(04:47): The Website SpecificationOriginal post: https://news.ycombinator.com/item?id=48343683&utm_source=wondercraft_ai(06:13): Codex just found a "workaround" of not having sudo on my PCOriginal post: https://news.ycombinator.com/item?id=48348578&utm_source=wondercraft_ai(07:39): Dav2dOriginal post: https://news.ycombinator.com/item?id=48344961&utm_source=wondercraft_ai(09:04): The solution might be cancelling my AI subscriptionOriginal post: https://news.ycombinator.com/item?id=48345896&utm_source=wondercraft_ai(10:30): 1-Bit Bonsai Image 4B Image Generation for Local DevicesOriginal post: https://news.ycombinator.com/item?id=48346257&utm_source=wondercraft_ai(11:56): United Airlines 767 returns to Newark after Bluetooth name sparks alertOriginal post: https://news.ycombinator.com/item?id=48345248&utm_source=wondercraft_ai(13:22): I put a datacenter GPU in my gaming PCOriginal post: https://news.ycombinator.com/item?id=48345694&utm_source=wondercraft_aiThis is a third-party project, independent from HN and YC. Text and audio generated using AI, by wondercraft.ai. Create your own studio quality podcast with text as the only input in seconds at app.wondercraft.ai. Issues or feedback? We'd love to hear from you: team@wondercraft.ai
This is a recap of the top 10 posts on Hacker News on May 30, 2026. This podcast was generated by wondercraft.ai (00:30): Microsoft Office 2019 and 2021 for Mac view-only conversionOriginal post: https://news.ycombinator.com/item?id=48341578&utm_source=wondercraft_ai(02:00): Danish pension fund excludes SpaceX citing governance and valuationOriginal post: https://news.ycombinator.com/item?id=48333820&utm_source=wondercraft_ai(03:30): Domain expertise has always been the real moatOriginal post: https://news.ycombinator.com/item?id=48340411&utm_source=wondercraft_ai(05:00): Anthropic surpasses OpenAI to become most valuable AI startupOriginal post: https://news.ycombinator.com/item?id=48336233&utm_source=wondercraft_ai(06:30): OpenRouter raises $113M Series BOriginal post: https://news.ycombinator.com/item?id=48338660&utm_source=wondercraft_ai(08:00): Pandoc TemplatesOriginal post: https://news.ycombinator.com/item?id=48334515&utm_source=wondercraft_ai(09:30): Openrsync: An implementation of rsync, by the OpenBSD teamOriginal post: https://news.ycombinator.com/item?id=48334854&utm_source=wondercraft_ai(11:00): Zig: Build System ReworkedOriginal post: https://news.ycombinator.com/item?id=48334048&utm_source=wondercraft_ai(12:30): EY Canada published a cybersecurity report and most citations were hallucinatedOriginal post: https://news.ycombinator.com/item?id=48339580&utm_source=wondercraft_ai(14:00): Voxel Space (2017)Original post: https://news.ycombinator.com/item?id=48336564&utm_source=wondercraft_aiThis is a third-party project, independent from HN and YC. Text and audio generated using AI, by wondercraft.ai. Create your own studio quality podcast with text as the only input in seconds at app.wondercraft.ai. Issues or feedback? We'd love to hear from you: team@wondercraft.ai
This is a recap of the top 10 posts on Hacker News on May 29, 2026. This podcast was generated by wondercraft.ai (00:30): The dead economy theoryOriginal post: https://news.ycombinator.com/item?id=48324712&utm_source=wondercraft_ai(01:57): I am retiring from tech to live offlineOriginal post: https://news.ycombinator.com/item?id=48323683&utm_source=wondercraft_ai(03:25): Please Use AIOriginal post: https://news.ycombinator.com/item?id=48323101&utm_source=wondercraft_ai(04:52): GTA 6 Developers UnionizeOriginal post: https://news.ycombinator.com/item?id=48324499&utm_source=wondercraft_ai(06:20): Cars collect a startling amount of data about youOriginal post: https://news.ycombinator.com/item?id=48318481&utm_source=wondercraft_ai(07:47): Blue Origin's New Glenn blows up during static fire testOriginal post: https://news.ycombinator.com/item?id=48317774&utm_source=wondercraft_ai(09:15): SQLite is all you need for durable workflowsOriginal post: https://news.ycombinator.com/item?id=48326802&utm_source=wondercraft_ai(10:42): Volkswagen blocks Home Assistant by requiring client assertionOriginal post: https://news.ycombinator.com/item?id=48319509&utm_source=wondercraft_ai(12:10): Notes from the Mistral AI Now SummitOriginal post: https://news.ycombinator.com/item?id=48325340&utm_source=wondercraft_ai(13:37): Claude Code – Everything you can configure that the docs don't tell youOriginal post: https://news.ycombinator.com/item?id=48318174&utm_source=wondercraft_aiThis is a third-party project, independent from HN and YC. Text and audio generated using AI, by wondercraft.ai. Create your own studio quality podcast with text as the only input in seconds at app.wondercraft.ai. Issues or feedback? We'd love to hear from you: team@wondercraft.ai
This is a recap of the top 10 posts on Hacker News on May 28, 2026. This podcast was generated by wondercraft.ai (00:30): Claude Opus 4.8Original post: https://news.ycombinator.com/item?id=48311647&utm_source=wondercraft_ai(01:58): Can we have the day off?Original post: https://news.ycombinator.com/item?id=48302745&utm_source=wondercraft_ai(03:27): Bricks and Minifigs Stole a Man's $200k Lego CollectionOriginal post: https://news.ycombinator.com/item?id=48314136&utm_source=wondercraft_ai(04:56): Disagreement among frontier LLMs on real-world fact-checksOriginal post: https://news.ycombinator.com/item?id=48307887&utm_source=wondercraft_ai(06:25): Show HN: Hallucinate – Massively Multiplayer Online RaveOriginal post: https://news.ycombinator.com/item?id=48304260&utm_source=wondercraft_ai(07:54): Citing 'severe' math deficits, UC faculty demand a return to SAT tests for STEMOriginal post: https://news.ycombinator.com/item?id=48309233&utm_source=wondercraft_ai(09:23): AMD pulls a bait-and-switch on Linux users with Vivado licensing changesOriginal post: https://news.ycombinator.com/item?id=48307231&utm_source=wondercraft_ai(10:52): EU fines Temu €200M for allowing sale of illegal productsOriginal post: https://news.ycombinator.com/item?id=48309302&utm_source=wondercraft_ai(12:21): Anthropic raises $65B in Series H funding at $965B post-money valuationOriginal post: https://news.ycombinator.com/item?id=48313048&utm_source=wondercraft_ai(13:50): Google employee charged with $1M Polymarket insider trading bet on search termOriginal post: https://news.ycombinator.com/item?id=48302822&utm_source=wondercraft_aiThis is a third-party project, independent from HN and YC. Text and audio generated using AI, by wondercraft.ai. Create your own studio quality podcast with text as the only input in seconds at app.wondercraft.ai. Issues or feedback? We'd love to hear from you: team@wondercraft.ai
Is seed investing facing an existential crisis? This week on The Data Minute, Peter sits down with Rob Go, Founding Partner at NextView Ventures, to discuss the structural shifts making the "game on the field" harder than ever for early-stage investors.Rob explains why many successful seed VCs are exiting the industry and how the rise of mega-funds and massive accelerators like YC has squeezed traditional seed firms into a narrow "subset" of the market. They dive into the "feeder fund" phenomenon, the arbitrary nature of ownership mandates, and why the $1B–$3B exit range has become a "Death Valley" for startups.Despite the current angst, Rob shares his optimistic "bull case" for 2030, explaining why diminishing competition and a rotation away from late-stage consensus will lead to a healthier venture substrate in the years to come.Subscribe to Carta's weekly Data Minute newsletter: https://carta.com/subscribe/data-newsletter-sign-up/Explore interactive startup and VC data, with Carta's Data Desk: https://carta.com/data-desk/Chapters:00:20 – Intro: Rob Go and the Seed Existential Crisis02:16 – Defining Seed: Betting on anything before PMF03:35 – Why senior seed VCs are exiting the industry05:02 – The Squeeze: Mega-funds vs. Accelerators07:02 – Scarcity vs. Abundance: What's left for seed funds?08:44 – The "Feeder Fund" trap and the factory supply chain12:38 – The risk of taking seed money from a mega-fund13:34 – Breaking down the 4 styles of seed investing15:20 – Why specialist seed funds can be transient19:29 – Super Compounders: Will exits keep getting bigger?21:59 – The "Death Valley" of $1B–$3B exits25:08 – The Blumhouse equivalent for venture capital27:18 – Normalizing secondaries as an exit strategy33:53 – Rant: Why ownership targets are backwards39:04 – Offensive vs. Defensive bridge rounds45:07 – "I've become way more Zen": Why the 2030 outlook is bullish50:18 – OutroThis presentation contains general information only and eShares, Inc. dba Carta, Inc. (“Carta”) is not, by means of this publication, rendering accounting, business, financial, investment, legal, tax, or other professional advice or services, and is for informational purposes only. This presentation is not a substitute for such professional advice or services nor should it be used as a basis for any decision or action that may affect your business or interests. © 2026 eShares, Inc., dba Carta, Inc. All rights reserved.
We're joined this week by Jan Sahagun of Trellis (backed by YC) to talk all things agents, the future of STR/Vacation Rentals tech, running with lean teams, 150 properties with 1 person and a lot more. Enjoy!⭐️ Links & Show NotesAdam NorkoConrad O'Connell Jan SahagunTrellis
This Week In Startups is made possible by:Grasshopper Bank - https://grasshopper.bank/twistLinkedIn - https://linkedIn.com/twistNorthwest Registered Agent - https://northwestregisteredagent.com/twistPlaud - https://Plaud.ai/twist Why raise $200 million if you are already profitable? That's the question Jason and Alex put to Mercury's founder and CEO, Immad Akhund, after the entrepreneur raised another massive round for his upstart, technology-friendly bank. TWiST then welcomed Kled founder Avi Patel to discuss the startup he considers a clear ripoff of his own company. Jason gavels in verdicts on all parties involved, including Y Combinator and venture capital firm General Catalyst. The show closes with a news lightning round, including OpenAI's decision to offer $2 million in token credits to hundreds of startups.Guest Links:Mercury https://mercury.comMercury funding announcement https://www.businesswire.com/news/home/20260520511817/en/Mercury-Raises-$200-Million-Series-D-at-$5.2B-ValuationImmad Akhund on X https://x.com/immadKled https://www.kled.ai/Avi Patel on X https://x.com/avipat_/Avi's complaint https://x.com/avipat_/status/2055384102409253056General Catalyst https://www.generalcatalyst.com/Y Combinator https://www.ycombinator.com/Delve https://techcrunch.com/2026/04/23/another-customer-of-troubled-startup-delve-suffered-a-big-security-incident/Discussion links:Anthropic's attack on secondary trading https://techcrunch.com/2026/05/12/anthropic-warns-investors-against-secondary-platforms-offering-access-to-its-shares/Vanta https://www.vanta.com/twistTimestamps:0:00 Welcome to This Week in Startups!2:14 Plaud: If your work depends on conversations — interviews, meetings, calls — you need a Plaud NotePin. You can check it out at https://Plaud.ai/twist and use code TWIST for 10% off!3:27 Immad Akhund (Mercury) joins to discuss $200M raise6:33 Mercury's origin story and path to $650M run rate9:53 Northwest Registered Agent - Get more when you start your business with Northwest. In 10 clicks and 10 minutes, you can form your company and walk away with a real business identity — Learn more at https://northwestregisteredagent.com/twist14:23 Stablecoins: where they work, why Mercury won't launch its own20:13 LinkedIn - Thanks to our partners at LinkedIn! Post your job for free at https://linkedIn.com/twist then promote it to get access to LinkedIn Jobs' new AI assistant.22:38 AI agents, and the future of money movement27:30 Why Mercury raised less this round30:11 Grasshopper Bank - Time is money. Don't waste either. Go to https://grasshopper.bank/twist and get an exclusive $500 cash bonus just for opening an account.42:48 Avi Patel (Kled) joins to discuss copycat startups57:06 Jason's verdict on YC's hacker culture & "appearance of impropriety"1:17:39 Sam Altman's $2M-in-tokens-for-equity offer to YC founders1:24:32 NYC hotel housekeepers cross $100K in time under new union contract1:30:14 Minimum wage, immigration & the case for raising it slowlySubscribe to the TWiST500 newsletter: https://ticker.thisweekinstartups.comCheck out the TWIST500: https://www.twist500.comSubscribe to This Week in Startups on Apple: https://rb.gy/v19fcpFollow Lon:X: https://x.com/lonsFollow Alex:X: https://x.com/alexLinkedIn: https://www.linkedin.com/in/alexwilhelmFollow Jason:X: https://twitter.com/JasonLinkedIn: https://www.linkedin.com/in/jasoncalacanisCheck out all our partner offers: https://partners.launch.co/Great TWIST interviews: Will Guidara, Eoghan McCabe, Steve Huffman, Brian Chesky, Bob Moesta, Aaron Levie, Sophia Amoruso, Reid Hoffman, Frank Slootman, Billy McFarlandCheck out Jason's suite of newsletters: https://substack.com/@calacanisFollow TWiST:Twitter: https://twitter.com/TWiStartupsYouTube: https://www.youtube.com/thisweekinInstagram: https://www.instagram.com/thisweekinstartupsTikTok: https://www.tiktok.com/@thisweekinstartupsSubstack: https://twistartups.substack.com
Garry Tan is the president and CEO of Y Combinator, the startup accelerator behind companies like Airbnb, Reddit, Coinbase, and DoorDash. He previously co-founded the financial technology company Posterous, which was acquired by Twitter in 2012, and later founded the venture capital firm Initialized Capital alongside Alexis Ohanian. Before entering venture capital, Tan worked as an engineer at Palantir Technologies, where he helped develop early infrastructure and design systems. Now, he continues to make investment and product decisions as a General Partner, having read more than 6,000 YC applications, while overseeing programs for sourcing, advising, and scaling early-stage startups. ------ Thank you to the sponsors that fuel our podcast and our team: AG1 https://DrinkAG1.com/tetra ------ LMNT Electrolytes https://DrinkLMNT.com/tetra Use code 'TETRA' ------ Squarespace https://Squarespace.com/tetra Use code 'TETRA' ------ Lectio 365 https://Lectio365.com ------ Sign up to receive Tetragrammaton Transmissions https://www.tetragrammaton.com/join-newsletter
Get our Wealth Guide (35+ insights from top investors): https://clickhubspot.com/ohkg Episode 822: Sam Parr ( https://x.com/theSamParr ) and Shaan Puri ( https://x.com/ShaanVP ) talk about how your genes could determine how much money you make and the startup ideas YC is betting on. — Show Notes: (0:00) money genes (5:07) your personality is your business (23:46) productive placebos (33:39) YC Request for Startups (34:57) IDEA: aesthetic data centers (42:15) IDEA: The company brain (51:10) IDEA: drone swarm defense (57:33) IDEA: personalized medicine — Links: • YC RFS - https://www.ycombinator.com/rfs • Deep Personality - https://deeppersonality.app/ • Viktor - https://getviktor.com/ — Check Out Sam's Stuff: • Hampton (joinhampton.com): My community for founders. Average member does $25m/year. Many of the guests are members. Get after it...apply: http://joinhampton.com/mfm — Check Out Shaan's Stuff: • Shaan's weekly email - https://www.shaanpuri.com • Visit https://www.somewhere.com/mfm to hire worldwide talent like Shaan and get $500 off for being an MFM listener. Hire developers, assistants, marketing pros, sales teams and more for 80% less than US equivalents. • Mercury - Need a bank for your company? Go check out Mercury (mercury.com). Shaan uses it for all of his companies! Mercury is a financial technology company, not an FDIC-insured bank. Banking services provided by Choice Financial Group, Column, N.A., and Evolve Bank & Trust, Members FDIC • I run all my newsletters on Beehiiv and you should too + we're giving away $10k to our favorite newsletter, check it out: beehiiv.com/mfm-challenge My First Million is a HubSpot Original Podcast // Brought to you by HubSpot Media // Production by Arie Desormeaux // Editing by Ezra Bakker Trupiano /
The Trump administration discussed an EO to form an AI oversight working group, a stark reversal from its hands-off approach. Apple explored using Intel and Samsung to make chips in the US, Coinbase cut 14% of its workforce, and OpenAI fast-tracks an AI phone for 2027. Sources: the Trump administration is discussing an EO to form an AI working group that would examine AI oversight procedures, like vetting models before release (NYT) Sources: Apple held exploratory talks with Intel and Apple executives visited a Samsung plant in Texas to explore producing core chips for its devices in the US (Bloomberg) Coinbase CEO Brian Armstrong announces the company is cutting ~700 jobs, or ~14% of its global workforce, to reduce costs, saying "AI is changing how we work" (Reuters) Meta is using AI on Facebook and Instagram to detect under-13 users by analyzing bone structure, height, and visual cues, but says it's "not facial recognition" (The Verge) Kuo: OpenAI appears to be fast-tracking its AI agent phone with two NPUs and a custom MediaTek Dimensity 9600 SoC, targeting mass production as early as H1 2027 (Ming-Chi Kuo) ElevenLabs raised $550M+ in its Series D, up from a previously announced $500M, adding BlackRock, Nvidia, and others as investors; its ARR passed $500M in Q1 (Tech.eu) Source: YC owns ~0.6% of OpenAI, which was seeded by a YC offshoot called YC Research in 2016; at OpenAI's current $852B valuation, the stake is worth $5B+ (Daring Fireball) Learn more about your ad choices. Visit megaphone.fm/adchoices