Podcasts about gpus

  • 1,402PODCASTS
  • 3,335EPISODES
  • 48mAVG DURATION
  • 2DAILY NEW EPISODES
  • Jul 30, 2026LATEST

POPULARITY

20192020202120222023202420252026

Categories



Best podcasts about gpus

Show all podcasts related to gpus

Latest podcast episodes about gpus

AV SuperFriends
AV SuperFriends: Off the Rails - Next year's problem could be next week's problem, bro

AV SuperFriends

Play Episode Listen Later Jul 30, 2026 81:12 Transcription Available


Recorded July 24, 2026 How long should classroom AV systems actually last? The panel revisits refresh cycles and quickly discovers that the answer ranges from a carefully planned five or six or seven or eight years to "whenever the last surviving input card finally catches fire." We compare full-room replacements with increasingly modular upgrades, debate which components can remain in service indefinitely, and examine what happens when departments treat technology as a one-time purchase rather than something that must eventually be funded again. We also check in on Auracast, the perpetually imminent technology of the future, following a real-world deployment at a UK train station. Then, because apparently this is something campuses must now consider, we discuss securing classroom computers against people opening them up and stealing the RAM and GPUs.  Along the way, we encounter abandoned projection screens, emergency spare parts, classrooms held together by obsolete switchers, and the uncomfortable realization that sometimes the only sustainable refresh strategy is moving a still-working component from somewhere nobody will notice.    News story: https://www.avnetwork.com/installations/ampetronic-listen-technologies-auri-installed-at-brighton-station   Alternate show titles: Let's talk about the Apple thing real quick Are all of you still hot on this? This is the technology of the future and it always will be I'm not trying to convince you that this isn't going to happen Laser stream quantum version Micro-pirate radio station Everything else is loosey goosey Covid 2: The Return Some years we ebb, some years we flow We just signed on to lease our computers I don't get a choice Let's go to the other side of this… Justin! So we said "Hey" (Hey!) Blow through a couple projectors Room Classic! Cells don't make audio Plug this hole with an AV widget I keep trying to ignore your hand You need to be careful with that word: guarantee The first wave of technicians…   We stream live every Friday at about 315p Eastern/1215p Pacific and you can listen to everything we record over at AVSuperFriends.com    ▀▄▀▄▀ CONTACT LINKS ▀▄▀▄▀ ► Website: https://www.avsuperfriends.com ► Twitter: https://twitter.com/avsuperfriends ► LinkedIn: https://www.linkedin.com/company/avsuperfriends ► YouTube: https://www.youtube.com/@avsuperfriends ► Bluesky: https://bsky.app/profile/avsuperfriends.bsky.social ► Email: mailbag@avsuperfriends.com ► RSS: https://avsuperfriends.libsyn.com/rss   Donate to AVSF: https://www.avsuperfriends.com/support

Christopher Lochhead Follow Your Different™
447 Niching Down on Weird Data with AI Solves Ice Cream, Movies, and Power Bottlenecks | The Pirate Street Journal

Christopher Lochhead Follow Your Different™

Play Episode Listen Later Jul 29, 2026 33:50


We are living through a fundamental shift in how businesses operate, compete, and create value. The Pirate Street Journal, hosted by Christopher, Eddie, and Bri, breaks down three major business stories through the category design lens, revealing a common thread that most mainstream business coverage misses entirely. That thread is AI data, and how the companies and individuals who understand it best are quietly rewriting the rules of entire industries. From energy infrastructure to ice cream shops to management consulting, the signal is clear and growing louder. This is just one of the topics that Pirates Christopher Lochhead, Eddie Yoon and Bri Clark discuss on this episode of Pirate Street Journal. Each week, the Category Pirates pick three headlines worth paying attention to and break down the category underneath. You're listening to Christopher Lochhead: Follow Your Different. We are the real dialogue podcast for people with a different mind. So get your mind in a different place, and hey ho, let's go.   Portable Power and the AI Energy Race China now controls 90% of the world’s battery storage cells, and the top ten storage cell manufacturers on Earth are all Chinese. The easy read on this is that the centralized, top-down model has already won. But the more interesting story is happening on the other side of the equation, where pioneers are refusing to wait for governments to build grids and are instead making power portable, distributed, and locally owned. Elon Musk quietly acquired a mobile power company capable of deploying a functional power plant in 30 days and driving it wherever demand exists. Tesla is simultaneously selling Mega Packs to cities experiencing brownouts while offering Powerwalls to individual homeowners. The insight here is that AI data is driving the need for entirely new power infrastructure, and the winners will not necessarily be the nations with the biggest grids. They will be the builders who understand that decentralized, distributed power networks can outmaneuver any centralized system when speed and flexibility matter most. Eddie raises the concept of a “Mega Pod,” a combination of batteries and GPUs in a scalable unit that could allow businesses, farms, and institutions with unused land to generate power, offset costs, and participate in a distributed data center economy. This is AI data infrastructure being rebuilt from the bottom up, and the annuity potential mirrors what Alaskan citizens receive from oil revenues every year.   Niche Down AI Data and the Rise of the Small Company Ben Affleck sold a stealth AI startup called Inner Positive to Netflix for $587 million. The company trained small models on individual film footage, replicating a director’s lighting style and visual language to accelerate post-production. An ice cream shop in downtown Los Angeles used prediction markets to hedge against cold weather, covering nearly half its monthly rent. A seven-person software company hit $10 million in revenue doing the work that once required 50 employees. These three stories appear unrelated on the surface, but they share a single strategic insight. Each one identified a narrow, specific type of AI data that nobody else was paying attention to and built an economic advantage around it. The Ben Affleck startup did not steal from other artists. It used a creator’s own footage as training data, producing tools that serve the creator rather than extract from them. The ice cream shop owner recognized that temperature data was weakness data for his business and converted it into a revenue stream through smart financial instruments. What AI is doing for smaller operators and independent entrepreneurs is lowering the barriers to prosecuting what Christopher Lochhead calls the magic triangle, building a legendary company, product, and category simultaneously. The surplus economics of AI are not accruing only to OpenAI, Anthropic, or the Mag Seven. They are flowing toward anyone willing to identify the weird data specific to their own situation and build something original with it.   Consulting and the Death of the Billable Hour McKinsey now ties 25% of its global fees to outcomes rather than hours. Bain reports that 30% of its business is AI and tech enabled, with ambitions to reach 50%. BCG expects AI work to jump from roughly 20% of revenue to 40% within a year. These are not small firms experimenting at the margins. These are the most conservative, hour-worshipping institutions in the professional services world, and they are cracking under the pressure of a new reality driven by AI data and what it makes possible. The billable hour was always a proxy for value, not a measure of it. What consulting firms are beginning to acknowledge is that AI data and the tools built around it can compress the time required for entry-level analytical work dramatically, which means the old pricing model no longer reflects what clients are actually buying. The shift from time-based to outcome-based compensation is not unique to consulting. It is the direction that professional compensation has been moving across all sectors for decades, from hourly wages to salaries to bonuses to equity. Eddie frames this shift clearly. People who are naturally oriented toward outcomes and who understand how to use AI data to drive measurable results are going to be rewarded more generously than ever before. Those who have relied on time as their unit of exchange, without a clear connection to the value they produce, are entering genuinely uncertain territory. The consultants are the last ones you would expect to change. The fact that they already are should function as a signal flare for every professional in every industry paying attention. To hear about all the topics in this week's The Pirate Street Journal, download and listen to this episode. You can also read more Pirate Street Journal entries in the Category Pirates newsletter.   We hope you enjoyed this episode of Christopher Lochhead: Follow Your Different™! Christopher loves hearing from his listeners. Feel free to email him, connect on Facebook, X (formerly Twitter), LinkedIn, and subscribe on Apple Podcast / Spotify!

The Tech Blog Writer Podcast
Running Enterprise Computer Vision on CPUs With Ultralytics YOLO26

The Tech Blog Writer Podcast

Play Episode Listen Later Jul 29, 2026 25:22


What becomes possible when enterprise computer vision no longer depends on expensive GPU infrastructure? In this episode of Tech Talks Daily, I speak with Glenn Jocher, founder and CEO of Ultralytics, about YOLO26, CPU inference, edge AI, open vocabulary vision, deployment economics, and the practical work required to move computer vision from a promising pilot into production. Glenn's route into AI began inside the U.S. intelligence community. He worked with the National Geospatial Intelligence Agency and Defense Intelligence Agency on particle physics applications, attempting to detect and track antineutrinos. Antineutrinos are extraordinarily difficult to detect because they pass through almost everything. Glenn describes them as the perfect spy. While searching for better detection methods, he discovered that computer vision researchers were solving similar problems with images. His original attempt to transfer those techniques into particle physics did not succeed. However, the work introduced him to a field where the technology could create a visible effect on everyday life. That led him toward open source development and eventually the YOLO models for object detection, classification, segmentation, and tracking. Glenn believes computer vision research has historically placed too much attention on small gains in accuracy while overlooking deployment economics. A model can perform impressively inside a laboratory and still remain unsuitable for a factory, warehouse, store, vehicle, drone, or medical environment. Price, latency, power consumption, data privacy, and deployment speed can determine whether the technology is commercially useful. This led Glenn and Ultralytics toward smaller models capable of running close to where images and video are generated. YOLO26 continues that approach with architectural changes designed specifically for CPU inference. Glenn says the model can process camera streams in real time at 30 frames per second and run across Intel CPUs, AMD CPUs, and lower power devices such as Raspberry Pi computers. This matters because specialist GPUs can increase the equipment cost and power requirements of a computer vision project. Running inference on existing CPUs or edge hardware can make deployment economically possible across larger numbers of cameras and locations. The scale already involved is difficult to comprehend. Glenn says Ultralytics models now process approximately three billion inference jobs each day, equivalent to around 30,000 every second. These jobs include images, videos, and collections of images being analyzed to detect, segment, or track objects. He attributes the platform's maturity to thousands of mistakes and bugs corrected through a rapid feedback cycle. New models are released, users report problems and request features, and the team incorporates that information into later versions. We also discuss the respective roles of cloud and edge infrastructure. Glenn sees cloud platforms continuing to provide the computing power required for training, while computer vision inference often belongs at the edge. Local processing can reduce latency, control operating costs, and keep sensitive video or medical information closer to where it was created. The smallest YOLO model is approximately three megabytes, according to Glenn. That allows it to reach mobile phones, vehicles, drones, battery powered devices, and other environments where a large language model would be impractical. Open vocabulary vision provides another development. Traditional object detection models are trained to recognize a fixed collection of objects. If a model learns to detect dogs and the user later wants it to detect cats, retraining can cause it to forget earlier knowledge unless both categories appear in the new training data. Glenn explains how promptable models can identify common everyday objects from text or visual instructions without additional training. A user could request a person wearing a blue shirt and white shoes, for example, and the system could search an image for that description. That flexibility could benefit businesses whose requirements change regularly. It reduces the need to create and label a new data set every time the company wants the model to recognize another common object. The range of current applications is already extensive. Glenn describes YOLO being used across robotics, parking, industrial safety, PPE detection, warehouses, aviation, security, traffic management, food quality, and manufacturing. Some of his favorite examples involve environmental problems. One company uses YOLO with underwater vehicles to identify and recover plastic from the ocean. Other applications detect smoke and fire early enough to support forest fire response. For leaders considering computer vision, Glenn recommends beginning with a defined problem and measurable outcome. A manufacturing company may want to reduce defects, but it still needs labeled examples showing the model what acceptable and defective products look like. He advises testing the idea through a limited pilot, measuring the return, and expanding only when the evidence supports further investment. Computer vision has become easier to deploy, but practical problems involving data, cameras, integration, reliability, and operating conditions still separate a demonstration from a production system. Could CPU inference and open vocabulary models make computer vision practical for processes your organization previously considered too expensive? Listen to the episode and share your thoughts with me. Useful Links   Ultralytics website Ultralytics Platform      

Art of Boring
Memory, Part 1: From Boom-Bust Commodity to AI Bottleneck | EP 223

Art of Boring

Play Episode Listen Later Jul 29, 2026 27:53


One of the defining market stories of the past 12 months has not been AI chips that compute, but the chips that remember. Equity analyst Shan Rui Yeo explains how memory works, from DRAM and NAND to high bandwidth memory, and how an industry that destroyed wealth for four decades became disciplined after consolidating to three players in 2013. He then walks through what changed: AI inference has made memory the key bottleneck, memory content is climbing with each new generation of GPUs, and new supply takes three to four years to build. With prices up sharply and customers signing long-term agreements, Part 1 of this three-part conversation lands on a commodity industry whose business model is changing in real time.   Key Takeaways Memory is a commodity with a three-to-four-year supply lag, which is why the cycle has always been difficult. Consolidation to three players in 2013 turned four decades of wealth destruction into at least 15% returns on capital through the cycles. In AI inference, memory bandwidth sets the speed of token generation, making memory the key bottleneck. NVIDIA's Rubin GPU carries 384 GB of DRAM, the equivalent of 32 iPhones per GPU, or 160 million iPhones across five million GPUs. HBM consumes three times the wafer capacity of standard DRAM (four times with HBM4) and is forecast to absorb 30% of DRAM wafers by 2027. DRAM contract prices are up roughly 200% year to date and 400 to 500% year over year, and price increases are reaching phones, laptops, and consoles. Customers are signing three-to-five-year agreements with prepayments, which could support a re-rating of memory companies.   Companies Mentioned: Samsung Electronics, SK Hynix, Micron, NVIDIA, Intel, Texas Instruments, Apple, Nintendo   Host: Rob Campbell, CFA, Institutional Portfolio Manager Guest: Shan Rui Yeo, CFA, Equity Analyst   This episode is available for download anywhere you get your podcasts.   Founded in 1974, Mawer Investment Management Ltd. (pronounced "more") is a privately owned independent investment firm managing assets for institutional and individual investors. Mawer employs over 250 people in Canada, U.S., and Singapore.   Visit us at: https://www.youtube.com/@MawerInvestment https://www.mawer.com https://www.linkedin.com/company/mawer-investment-management/ https://www.instagram.com/mawerinvestmentmanagement/ #ArtOfBoring #MawerInvestmentManagement #MawerInvestment #Podcasts

Linux Weekly Daily Wednesday
4GB AMD 9050 GPUs and Chrome for ARM Linux!

Linux Weekly Daily Wednesday

Play Episode Listen Later Jul 29, 2026 51:36


Possible price increases for the VR Head Toaster, Google Chrome for ARM Linux desktops, a new 10-inch touchscreen for Raspberry Pi, and NVIDIA doubling the price of Jetson SoCs.Video version of the show is available to Patrons, along with the Extended Chaos podcast featuring over an extra hour of LWDW content every week.⁠⁠⁠⁠⁠⁠⁠Patreon⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Discord⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠YouTube⁠⁠⁠⁠⁠⁠⁠⁠⁠TIMESTAMPS00:00 Intro07:22 AMD 4GB GPUs16:57 Steam Frame VR headset price increases 25:36 Google Chrome for ARM Linux 33:52 10" Raspberry Pi touchscreen37:48 NVIDIA Jetson price increases New 4GB GPUs from AMDhttps://videocardz.com/newz/amd-confirms-radeon-rx-9050-4gb-exists-but-it-is-oem-onlySteam Frame $$$https://www.pcgamer.com/hardware/vr-hardware/more-bad-news-for-valve-and-its-still-unreleased-steam-frame-as-qualcomm-is-set-to-raise-the-price-of-its-chips-like-everyone-else/Chrome on ARMhttps://www.omgubuntu.co.uk/2026/07/chrome-arm64-linux-availableRasPi 10” Touchscreenhttps://www.raspberrypi.com/news/a-new-10-raspberry-pi-touch-display-2-available-now-at-80/Jetson Price increasehttps://videocardz.com/newz/nvidia-raises-jetson-prices-by-up-to-101-agx-thor-now-costs-5499

Remotely Curious
Protecting your team's content, wherever it's stored—so you can safely use AI

Remotely Curious

Play Episode Listen Later Jul 28, 2026 27:37


AI makes it easier than ever to find and act on information—especially now that teams can connect to and search across all the apps they use for work. So how do you ensure that only the right people and the right tools can access your team's most sensitive content? In this episode, we talk with Jess Jimenez, the head of security at Dropbox, about what security looks like in the age of AI at Dropbox-scale—from building AI products securely to building trust with the people who use them. Jess talks about the importance of access control lists, defending against the latest AI threats, and how Dropbox Protect helps teams securely share content with both humans and AI so they can collaborate more safely. ~ ~ ~  Working Smarter is brought to you by Dropbox. Find, organize, and share your work—all in one place—with context-aware AI from Dropbox. You can listen to more episodes of Working Smarter on Apple Podcasts, Spotify, YouTube, Amazon Music, or wherever you get your podcasts. To read more stories and past interviews, visit workingsmarter.ai This show would not be possible without the talented team at Cosmic Standard: producer Ben Montoya, sound engineer Aja Simpson, technical director Jacob Winik, and executive producer Eliza Smith. Special thanks to our illustrator Fanny Luor, marketing consultant Meggan Ellingboe, and editorial support from Catie Keck.  Our theme song was composed by Doug Stuart.  Working Smarter is hosted by Matthew Braga. Thanks for listening!

The Data Center Frontier Show
AI Clusters and the New Economics of Data Center Optics

The Data Center Frontier Show

Play Episode Listen Later Jul 28, 2026 36:09


Optics is no longer a supporting accessory in the data center network. As AI infrastructure advances from 400G and 800G toward 1.6-terabit connectivity, optical components are consuming a larger share of network cost, power and operational risk. In this episode of the Data Center Frontier Show, DCF Editor in Chief Matt Vincent speaks with Bill Gartner, Senior Vice President and General Manager of Cisco's Optical Systems and Optics business, about how AI is changing the strategic role of optics. Gartner explains that optics represented roughly 10% of a network port's bill of materials at 10G. At 400G and above, the optics can cost more than the switch port itself. Reliability has also become critical: A single unstable link can force GPUs operating in parallel to stop, return to a checkpoint and restart. According to data Cisco has seen from hyperscale customers, link flaps can reduce GPU infrastructure efficiency by as much as 40%. The conversation maps the AI network across three distinct tiers: Scale-up: Connections within the rack, carrying approximately 500 times the bandwidth of a traditional WAN environment. Scale-out: Connections between racks, commonly using 400G and 800G pluggable optics. Scale-across: Coherent optical connections between data centers as AI clusters expand beyond the power limits of a single facility. Gartner also discusses Cisco's 1.6T roadmap, routed optical networking, coherent pluggable optics and the emerging debate around co-packaged and near-packaged optics. These architectures promise lower power consumption and greater density, but introduce new questions involving interoperability, replacement and operational resilience. Looking ahead, Gartner emphasizes that optics is not constraining AI network growth. It is enabling clusters to scale across racks, campuses and geographically distributed data centers, while the coming inference wave shifts the industry's focus toward cost and power efficiency.

Dark Racial Humor
OpenAI's Agent Escaped, Google's $5.9B Cash Burn & Apple's Smart Glasses | Ricker and Bon #435

Dark Racial Humor

Play Episode Listen Later Jul 27, 2026 69:56


OpenAI says an internal cyber-capability evaluation crossed the boundary it was supposed to stay inside and compromised Hugging Face infrastructure. We break down how the agent moved between systems, why the containment failure matters, and why it felt less like a benchmark and more like the opening scene of a science-fiction movie.Alphabet posted Q2 revenue of $119.8 billion, but the quarter also produced negative $5.9 billion in free cash flow after $44.9 billion in capital spending. The debate is no longer whether AI demand exists; it is how long the infrastructure payback takes when even Google can spend more cash than it generates in a quarter.That leads into the wider AI buildout: off-balance-sheet financing, power and construction bottlenecks, depreciating GPUs, and whether data centers will leave behind infrastructure as durable as railroads and telecommunications. Presearch also announced that it was shutting down, turning one abstract industry conversation into an immediate business question.Crypto was supposed to replace the banking system. Now BlackRock, Goldman Sachs, and other institutions have accumulated the assets, and the market increasingly resembles the financial system it once rejected.Marvel used Comic-Con to announce Ryan Gosling as Ghost Rider and David Jonsson as the new Black Panther. The conversation immediately becomes a test of whether anyone can keep Ryan Gosling, Ryan Reynolds, Nicolas Cage, Deadpool, and Ghost Rider straight.Christopher Nolan's The Odyssey earned $87 million in its second domestic weekend and reached approximately $639.6 million worldwide. Premium formats and IMAX are central to the release, which makes watching it later through TikTok clips the most disrespectful possible viewing plan.Warner Bros. Discovery sued Amazon MGM Studios over the hiring of HBO Max marketing executive Pia Barlow, alleging interference with employment agreements. Apple is also reportedly targeting a 2027 consumer release for smart glasses, reopening the argument over privacy, ownership, and whether a company should be able to disable hardware you bought.Before the headlines, we talk about the wedding that reopened the family question, the dorm assignment that made Ricker and Bon possible, dating at 20 versus 30, and the very genuine monotone.If you want a prize, send us a DM:instagram.com/rickerandbontiktok.com/@rickerandbonyoutube.com/@rickerandbon

El Podcast de JF Calero
CHINA AMENAZA A LA IA DE ESTADOS UNIDOS - OBJETIVO: DESTRUIR EL MERCADO

El Podcast de JF Calero

Play Episode Listen Later Jul 27, 2026 15:37


Invierte periódicamente con 0 € de comisión a través del Plan de Inversión con ETF de Freedom24: https://freedom24.club/cascaron_planes Abre una cuenta gratuita en Freedom24 https://freedom24.club/CaleroJF y consigue:✓ 0 € de comisión por inversión automática en ETFs y acciones elegibles✓ 40.000 acciones y 3.600 ETFs para diversificar tu cartera✓ Acceso directo a más de 20 bolsas de América, Europa y Asia✓ Asistente personal gratuito e ideas de inversiónComo ventaja adicional, actualmente puedes conseguir hasta 20 acciones de regalo* por abrir y recargar tu cuentaFreedom24 es un bróker europeo con licencia y filial al 100% de FRHC, empresa cotizada en NASDAQ.Contenido patrocinado por Freedom24Disclaimer:Contenido patrocinado por Freedom24. Freedom24 es un bróker europeo regulado por la CySEC que opera en España a través de su agente vinculado Freedom24 Iberia, registrado en la CNMV.Invertir siempre implica el riesgo de perder tu capital. Las rentabilidades pasadas o las previsiones no garantizan resultados futuros. El acceso a un vehículo de inversión concreto está sujeto a un test de idoneidad. Realiza tu propia investigación antes de realizar cualquier inversión. Es importante consultar con un asesor financiero antes de tomar cualquier decisión de inversión. La oferta de 0 € de comisión solo está disponible para la función de aportaciones automáticas recurrentes de Freedom24. Pueden aplicarse otras comisiones.*La promoción WELCOME está sujeta a términos y condiciones. Las acciones de regalo se asignan aleatoriamente de una selección de valores elegibles, y las acciones de mayor valor se otorgan con menor frecuencia.China está cambiando las reglas de la inteligencia artificial. Mientras Microsoft, Google, Amazon, Meta y Oracle preparan inversiones gigantescas en centros de datos e infraestructura de IA, los laboratorios chinos están siguiendo una estrategia completamente diferente: modelos abiertos, costes reducidos y una inteligencia artificial cada vez más accesible.El impacto ya se dejó sentir con DeepSeek, cuyo ascenso sacudió a Silicon Valley y coincidió con una histórica caída de Nvidia en bolsa. Pero DeepSeek fue solo el principio. Qwen, Kimi y otros modelos chinos están demostrando que China puede competir en inteligencia artificial incluso con las restricciones estadounidenses al acceso a los chips más avanzados.En este vídeo analizamos la auténtica batalla de la IA entre China y Estados Unidos: dos estrategias radicalmente diferentes. Silicon Valley apuesta por inversiones masivas, centros de datos, GPUs, suscripciones y modelos cada vez más potentes. China intenta convertir la IA en una tecnología abierta y prácticamente gratuita, extendiendo sus modelos por todo el mundo.¿Qué ocurre si los modelos gratuitos llegan a ser suficientemente buenos para la mayoría de usuarios? ¿Cómo recuperarán las grandes tecnológicas estadounidenses los cientos de miles de millones invertidos en infraestructura? ¿Puede el código abierto convertirse en el arma con la que China cambie el mercado mundial de la inteligencia artificial?También analizamos el papel de DeepSeek, Qwen y Kimi, las restricciones a Nvidia, el proyecto Stargate de OpenAI, Oracle y SoftBank y la enorme carrera por construir la infraestructura que alimentará la próxima generación de inteligencia artificial.China y Estados Unidos están jugando dos partidas muy diferentes. Y el ganador de esta batalla podría definir cómo utilizaremos —y cuánto pagaremos— por la inteligencia artificial durante los próximos años.#InteligenciaArtificial #China #DeepSeek

DeFi Slate
Cysic Founder: Why The Compute Market Is Getting Ready to Explode (Biggest In The World)

DeFi Slate

Play Episode Listen Later Jul 25, 2026 25:03


Leo Fan breaks down why the real bottleneck in AI isn't model intelligence but a persistent GPU shortage driving up inference costs, even as open-source models from China close the gap. He also unpacks the financialization of compute and why some believe it could become one of the largest derivatives markets in the world.Leo Fan is the Founder of Cysic, a full-stack compute network turning GPUs, ASICs, and spare hardware into verifiable AI infrastructure, and a Cornell CS PhD researcher in zero-knowledge systems and AI.The Rollup is where the leaders of digital assets and finance converge. Live from the financial capital of the world.Timestamps00:00 Intro04:18 Kimi K2 Shocks Compute Demand06:39 GPU Shortage Drives Inference Costs11:34 Engineering Around Chip Gaps18:23 Compute Financialization20:48 Earning Yield From Spare Compute23:20 Compute Markets FutureGuest Socials:Leo Fan X: https://x.com/leofanxiongCysic X: https://x.com/cysic_xyzCysic Website: https://cysic.xyz/ Partners: Better than Banks. Transparent capital efficiency earning the highest yields in DeFi. Learn more here: https://infinifi.xyz/---1inch - Simple experience. Smart execution. Trading built to scale. It's time to bring the world onchain. https://1inch.com/---Dinari - Over 230 1:1 backed tokenized stocks, ETFs & more with dividends. US-based SEC transfer agent. Available on 5+ chains & via API. https://dinari.com/---Relay is the fastest and most reliable way to swap any token on any chain. Learn more here: https://relay.link/bridge---Zama is an open source cryptography company that builds state-of-the-art Fully Homomorphic Encryption (FHE) solutions for blockchain.Learn more here: https://www.zama.org/---Trezor is the creator of the first-ever hardware wallet. Securing crypto for 2M+ users worldwide. 100% open source. Learn more here: https://affil.trezor.io/aff_c?offer_i...---

7 Minute Security
7MS #732: Tales of Pentest Pwnage – Part 86

7 Minute Security

Play Episode Listen Later Jul 24, 2026 40:02


Hey friends! Welcome back to another Tales of Pentest Pwnage — my favorite mini-series where I share the good, the bad, and the "why didn't I check THAT first?!" moments from real-world engagements. Today's story has a little bit of everything: a legit path to domain admin, some late-night rabbit holes, a lesson in humility, and a villain you've definitely met before. (Spoiler: it's DNS.) A couple of quick plugs before we dive in: Private GOAD training is going strong! — We just wrapped a 3-day private session (7 students — that's max capacity!) of our Active Directory pentesting class built on the Game of Active Directory (GOAD) framework. Over three days, students enumerate, attack, and fully pwn three separate AD environments. The private format is just *chef's kiss* — when it's a team from the same company, the conversation gets real fast. Like, "hey I just checked Bloodhound on break and Bob from accounting has full rights over the DC" real. If you want to send 3–7 people from your org, hit up 7MinSec.com/training to line up a private session. Support the show over at 7MinSec.club — That's our Substack, where every Tuesday I drop a short TuesdayTOOLSday video about security tools. Free subscriptions are welcome and mean a lot — you'll just get pinged when new content drops. No spam, no blindly-sent Outlook calendar invites. I promise. Pentest tips and scripts live at 7MinSec.wiki — I reference it throughout today's episode, including some step-by-step guidance on the techniques we'll talk about below. Now — onto the pwnage. Fair warning: I've been burning the candle at three ends lately trying to catch up after a tough few weeks of grief (if you want the backstory, the last couple episodes cover my dad passing away). The good news is my head is semi back on straight and I put it to work on a recurring client environment — one that keeps getting better year over year. Machine account quota locked down? Check. No Kerberoastable or AS-REP roastable users? Check. No local admin rights, no web client running? Check and check. All good signs. And then PingCastle smiled right into my eyeballs with a big red finding: The DC's LAN Manager authentication level was weak enough to coerce and capture a downgraded hash — Specifically, an NTLMv1 SSP hash. Using Coercer to nudge the DC into authenticating to my Kali box (with Responder running), I captured the goods. Pretty little hashes all in a row. Cracking that hash: enter Vast.ai — The old go-to for this type of crack used to be crack.sh, but their cracker has been offline for years. What they do still have is a walkthrough pointing to a tool from EvilMog on GitHub that helps you prep the raw hash material and figure out exactly how to crack it with Hashcat. For the GPU horsepower, I rented a beefy multi-GPU instance on Vast.ai — filter for 16+ GPUs, pick a Hashcat Docker image, and SSH in. The whole crack job took about 16 hours at ~$4/hr. Do the math: $64 to reconstruct the DC's NTLM hash. Worth it. Tmux sidebar — seriously just learn it — Vast.ai is actually what finally got me into tmux, because the Hashcat Docker container drops you right into a tmux session. This is clutch: you can kick off a 16-hour crack job, detach, and reattach later without killing anything. On a pentest, my workflow now is SSH in → tmux → name a few session windows for Responder, Exegol, packet captures, etc. I used to fumble around with Linux screen sessions. Not anymore! From hash to DA — the usual playbook — Once you've got the DC's NTLM hash, you can request a Kerberos ticket and load it up, then run a DCSync to pull the KRBTGT hash. From there it's god mode: dump hashes, pass-the-hash as domain admins, and you have yourself a cool privesc POC. Except this time…the POC didn't work. The part where I Jean-Claude Van Damme helicopter kick myself in the face — DCSync failed immediately. Like, suspiciously fast — barely two lines of output and done. I tried every version of every tool I could get my hands on. I tried Windows, I tried Linux. I even asked the client to check if their endpoint protection was blocking me (it wasn't). I touched grass. I played guitar. I played some Splinter Cell Blacklist (old game, highly recommend if you like the Hitman-style vibes). Came back fresh. Rebooted both VMs. Still nothing. It was DNS. It's always DNS. — The thing that finally caught my eye: the commands were failing too fast. Like it wasn't even reaching the DC. I catted the resolv.conf inside my Exegol instance (heads up: Exegol has its own resolv.conf and hosts file, separate from your base Kali system!) and found a stale DNS entry pointing to an old DC that was no longer serving anything. Nuked the bad entry, added static hosts file entries for the live DC, ran the command again, and — hash rain. Pennies from heaven. It was midnight and I literally pushed back from my desk like a baby pushing away from a high chair going "Baby Brian is all done!" The lesson: — I know the meme. "It's always DNS." I just personally hadn't hit it hard in my security life since my sysadmin days back before 2013. Now I have. So going forward I'll check DNS first (and often). Vacation attempt #3 incoming… pray for me — My wife nearly died in Punta Cana earlier this year. Then our summer cabin trip was cold and rainy with zero water time. And now we've got families flying in from multiple states for a lake weekend — except we just found out our reservation through Booking.com was basically vaporized because the resort changed hands and never updated their website. My wife (who is an absolute saint and my better three-quarters) almost had a 360-degree head spin (like in The Exorcist) talking to customer service. But we scrambled, found a last-minute place, and I'm choosing to believe it's not in Jason Voorhees' back yard. Could this be my last episode? Maybe. But hey — it was a good one. Talk to you next week (hopefully).

TechCrunch Startups – Spoken Edition
AI chip startup Etched defies skeptics; plus, Imagi raises $4.5M to help teach students how to vibe code

TechCrunch Startups – Spoken Edition

Play Episode Listen Later Jul 24, 2026 9:09


Etched, founded by three Harvard dropouts, has created new chips and memory components that speed up inference on any AI model -- no GPUs required, it says. Also, Imagi announced a $4.5 million seed round, with investors including Brighteye Ventures, Day One Capital, and artist Wil.i.am. Learn more about your ad choices. Visit podcastchoices.com/adchoices

The Chris Voss Show
The Chris Voss Show Podcast – The Future of GPUs: How iFrame.ai is Redefining Cost and Performance

The Chris Voss Show

Play Episode Listen Later Jul 23, 2026 31:37


The Future of GPUs: How iFrame.ai is Redefining Cost and Performance iframe.ai About the Guest(s): Vlad Panit is the CEO of iFrame.ai, a company specializing in deploying GPUs and building data centers for better compute resources, primarily targeting Neocloud providers, hyperscalers, and AI labs. Vlad’s journey in entrepreneurship began at 18, with a pivot towards IT and engineering endeavors from 2008. His significant work includes e-health projects in Europe and AI applications in medical coding automation. Originally from Ukraine, Vlad has expanded his endeavors to the United States over the past seven years, using his expertise to spearhead innovation in AI-related technologies. Episode Summary: In this insightful episode of The Chris Voss Show, host Chris Voss dives into the world of AI with guest Vlad Panit, the CEO of iFrame.ai. The conversation explores the backbone of AI infrastructure focusing on the deployment of GPUs in data centers. Vlad shares how iFrame.ai provides bare metal GPU services, offering significant cost savings over conventional cloud providers like Amazon AWS. This episode sheds light on AI’s evolving landscape and the technological advancements driving it, making it particularly relevant for business professionals and tech enthusiasts interested in AI advancements. Chris and Vlad delve into the differences between AI infrastructure services provided by iFrame.ai and traditional offerings by big players like AWS and Google. Vlad explains the company’s unique position in providing bare metal services that enhance cost and operational efficiency for customers, especially in AI training and inference. Highlighting the broader implications of AI’s rapid development, Vlad also reflects on the challenges and opportunities that come with deploying large-scale GPU networks, touching on both technological and environmental impacts. Key Takeaways: iFrame.ai offers bare metal GPU services which can be three to four times more affordable compared to AWS and Google Cloud offerings. The company’s infrastructure supports significant cost savings in AI operations, making it an attractive option for companies looking to optimize AI workloads. Vlad Panit emphasizes the importance of deploying GPUs efficiently to enhance the availability and reliability of compute resources for AI applications. The discussion highlights the evolving market landscape as tech giants like OpenAI and Anthropic move towards public offerings, with AI infrastructure playing a pivotal role. The conversation also touches on the environmental impact of AI data centers and the importance of sustainable energy practices in powering such facilities. Notable Quotes: “We deploy GPUs. We build data centers, put a lot of compute there, and then sell it to near clouds, hyperscalers, AI labs—some of which you probably use in daily life.” “We’re able to sell our services three to four times cheaper than AWS.” “The difference with AI’s GPU use versus gaming is it requires uploading and storing entire models, which is more complex and costly.” “AI data centers should have their own energy supply to be better for everyone, from local communities to the data center operations.” “It’s a huge overestimation to say there’s a bubble in AI; the infrastructure costs alone are massive to support large-scale AI operations.” Resources: iFrame.ai – Discover more about the company’s GPU deployment services. Connect with Vlad or explore further through LinkedIn and Goodreads. Tune in to the full episode of The Chris Voss Show to gain a deeper understanding of AI infrastructure and the intricacies of deploying GPUs for cutting-edge technology applications. Stay connected for more intriguing discussions with thought leaders shaping the future.

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0

In recent months, the open vs closed, and US vs China discussions on model ownership and sovereign/local AI have heated up to a fever pitch. So it is very very good news that Poolside AI are finally emerging with new models, like Laguna S 2.1, that are beating Thinking Machines' recent release nearly 10 times their size.Poolside's recent tech report got a lot of praise due to their level of detail, and Vibhu first covered Laguna's recent technical report on our paper club:From spending $12 million building language models for code before the world cared to creating a Model Factory that can take a model from pre-training to release in eight weeks, Eiso Kant has spent more than a decade betting that code is the path to AGI. In this episode, the Poolside co-founder joins swyx and Vibhu to explain why ChatGPT felt like vindication, why Poolside embraced open weights and open research, and why he would rather live in a world with 100 foundation model companies than five even if Poolside were one of the five.We go deep on Poolside's Model Factory: the engineering systems behind 10,000–20,000 experiments per month, streaming data directly into training, reproducible experimentation, low-precision compute, and agents that increasingly write code, launch jobs, evaluate results, and modify the pipelines used to train future models. Eiso also unpacks their recent launch Laguna S, why persistence, verification, and backtracking may matter more than raw intelligence, how much capability remains inside smaller models, why reinforcement learning will move earlier into pre-training, and why next-token prediction is still extracting too little from the web.We also discuss model-harness co-design, Poolside's path from coding agents to AGI, why Eiso thinks MCP and traditional tool calls are “stupid,” the real economics behind frontier-model training, Poolside's $500 million raise, open-source AI, regulation, NVIDIA and TSMC's influence, engineering productivity in the agent era, high-agency teams, and hiring at Poolside.We discuss:* How Andrej Karpathy's RNN work inspired Eiso to start building language models for code in 2015* Why Eiso spent four years and $12 million pursuing an idea before the market cared* Why ChatGPT felt like vindication and brought Poolside back to open source* Why Eiso would prefer 100 foundation model companies over an oligopoly of five* The difference between releasing open weights and publishing genuinely open research* Why Poolside deliberately built a global research organization outside the Bay Area talent war* Why model building is ultimately 90% engineering* The Model Factory: Poolside's end-to-end system for rapidly training and improving models* How fewer than 70 researchers run roughly 10,000–20,000 experiments each month* How Poolside moved from six-month model cycles to five- and eight-week launches* Why streaming data directly into training unlocked faster experimentation* How immutable data, versioned code, and reproducibility enable rigorous model research* Why Eiso wants capable researchers to leave their labs and become Poolside's competitors* Why 95% of model building can be reduced to better data or compute efficiency* Laguna S and why persistence, verification, and backtracking can outperform raw intelligence* Why smaller models may handle far more knowledge work than previously expected* Why reinforcement learning will move earlier into pre-training* Why next-token prediction is still failing to extract enough knowledge from the web* Why distillation and environments have become the AI industry's favorite “drugs”* Why mid-training is really an early form of curriculum design* Low-precision training, networking bottlenecks, and the next gains in compute efficiency* Laguna S: 118 billion total parameters, 8 billion active, and eight weeks from training to launch* Why model builders can often evaluate a new checkpoint within its first 30 minutes* Model versus harness: where agent capabilities actually come from* Why Poolside sees coding and long-horizon software tasks as a path to AGI* Why Eiso thinks MCP and traditional tool calls are “stupid”* Why future agents will write scripts instead of choosing from dozens of predefined tools* The case for minimal harnesses, containers, and model freedom* Why Poolside is prioritizing vision but does not expect to work on audio soon* Why language may be the most compute-efficient modality for encoding knowledge and reasoning* The real cost of model development and why the final training run is anticlimactic* The story behind the Poolside name and why it represents refusing to lower ambitions* How Poolside raised $500 million while investors still questioned whether AGI was real* Why intelligence could become the world's most demanded and commoditized resource* When open models may become too capable to release without restrictions* Why unilateral AI safety does not work in a globally competitive environment* How regulation could accidentally lock in an oligopoly of two or three AI companies* NVIDIA, TSMC, and the hardware systems underpinning foundation-model progress* Why reinforcement-learning wall-clock time is one of Poolside's biggest bottlenecks* Why Poolside trains models from scratch instead of simply distilling larger models* How AI changes the way companies should measure engineering productivity* Why agency may become the most important quality for employees in the AI era* How leaders align high-agency people through shared goals and clear constraints* Hiring across research, post-training, pre-training, architecture, evals, and engineering at PoolsideEiso KantLinkedIn: https://www.linkedin.com/in/eisokantX: https://x.com/eisokantPoolside: https://poolside.aiTimestamps00:00:00 Introduction00:00:54 Karpathy, RNNs, and Building Code Models Before Transformers00:02:26 The $12M Failure and ChatGPT Vindication00:03:39 Open Source and the Case for 100 Foundation Model Companies00:09:22 Open Weights, Open Research, and Poolside's Global Team00:16:04 The Model Factory: Why Model Building Is 90% Engineering00:20:19 Agents, Automated Experiments, and Early Signs of RSI00:24:04 Streaming Data, Reproducibility, and Scientific Rigor00:30:35 Creating More Foundation Model Companies00:36:07 Laguna S: Persistence vs. Raw Intelligence00:43:01 Reinventing Pre-Training, RL, and Curriculum Design00:52:33 Low-Precision Training and Squeezing More From Smaller Models00:58:37 Model Harnesses, Coding Agents, and the Path to AGI01:09:26 Why MCP and Traditional Tool Calls Are “Stupid”01:13:04 Vision, Multimodality, and Why Language Still Matters01:18:15 Scaling Models and the Real Economics of Training01:20:40 Why Poolside Is Called Poolside and Raising $500M01:27:37 Open Models, AI Safety, and the Risk of an Oligopoly01:33:53 NVIDIA, TSMC, and the Reinforcement-Learning Bottleneck01:41:52 Smaller Models, Distillation, Engineering Productivity, and HiringTranscriptIntroduction: Eiso Kant, Poolside, and Open ModelsSwyx [00:00:00]: All right, we're here in the studio with Eiso Kant from Poolside, together with Vibhu. Welcome.Eiso Kant [00:00:08]: Thanks. Thanks for having me, guys. Good to be here.Swyx [00:00:10]: Yeah, fresh on the plane. You texted me, you were like, “Hey, I'm on my way to SF.” I was like, “You're on a plane right now, right?” Like, hey.Eiso Kant [00:00:16]: I know. After I texted you, I realized that probably coming in with major jet lag was gonna offer some fun experiences today, but let's do it.Swyx [00:00:23]: I mean, I think the thing I would tell guests is that they don't have to prepare that much because if you're truly working on this every single day, then even, like, what you hazily remember is going to be new for a lot of the audience that don't live in your world every day, right? so 10 years ago, you did a talk at Google Slush, talking about the democratization of AI. and, now here you are, like, open sourcing an incredible new model that we're gonna talk about. But I guess, like, what got you into democratization of AI? Like, it's not obvious from your LinkedIn or something.From Karpathy's RNN Post to SourcedEiso Kant [00:00:57]: No, it's not at all. I don't think it's obvious how I got in this space. I owe getting into this space to Andrej Karpathy.Eiso Kant [00:01:05]: In 2015, he wrote an article called “The Unreasonable Effectiveness of Recurrent Neural Nets.”Swyx [00:01:10]: Neural Nets, yep.Eiso Kant [00:01:11]: And that article, I read it, and I pivoted my startup at the time overnight to working on RNNs, and later LSTMs and Transformer models to be able to write code. If you go to this article and you scroll down, you can start seeing, like, this was the precursor to what ended up becoming language models. So, at least when he was character-level language models that were starting to predict letters, he has an example out here. There's a little Paul Graham generator, and you can read it, and the text makes sense, but it doesn't. and there's a little-- There's an example of code a little bit further down. Yeah, so Shakespeare.Swyx [00:01:47]: Shakespeare.Swyx [00:01:49]: CoolEiso Kant [00:01:49]: And for some reason, I read this, and I went down the rabbit hole of learning everything I could about RNNs and LSTMs, right? This is Transformer paper. And I had built a completely unreasonable belief, that neural nets should be able to generalize to anything and everything, and that language should be able to generalize, to a lot of things that are intelligent and the ability to write code. And so I started building Sourced, which was a fully open source company trying to build, what we used to call machine learning on code, language models on code. And we spent about four or five years on this, till the end of 2019. And that sounds really cool today, but back then, no one cared.Eiso Kant [00:02:29]: Right? Like, no one cared. We were in the dark. Like, we did things along the way. We tried applying convolutional neural nets to, like, the structure of code. We were. when attention came out, we were applying it to LSTMs, and then the Transformer paper came out. And it - it wasn't obvious, and what we missed throughout that entire journey, that we were on the right track, but we should have just kept scaling up. And today, to all of us, the scaling laws and scaling up seems like the most obvious thing. But having spent four or five years of my life on working on language models on code, it wasn't obvious. So I have a lot of respect to folks at Google and OpenAI and others who took that confidence and kept going. we failed ultimately at the time, and it was, like, biggest failure of my career, right? You blew $12 million of investors' money, which was a lot back then.Swyx [00:03:18]: Yep.Eiso Kant [00:03:19]: You spent, still a lot, but, And you spent years with, like, a group of 40 people just obsessing over this problem. And life took a different turn, And it was, and family became a focus, and I kept my heads down and really, didn't really look at language models for the following two years. big mistake considering Following years are gonna be really interesting. And then ChatGPT came out And it was like a vindication. It's like people started texting me. I found, like, my old, work decks and these old talks. And throughout that whole journey, we,ChatGPT, Vindication, and Returning to Open SourceEiso Kant [00:03:56]: We really had a strong point of view at the time that, like, as you're building more capable intelligence, it should be open and open source.Eiso Kant [00:04:04]: When we started Poolside, that wasn't the case at all, and I wanna be very open about it. When we started Poolside, we were like, there was a premise of two things. One is this technology is not gonna stop compounding in capabilities. I think to most people obvious today, but three-plus years ago when we started, most people were still arguing if these were stochastic parrots or not.Eiso Kant [00:04:23]: And the second was that reinforcement learning was gonna be the biggest driver for LLM capabilities. Today, very obvious. Three years ago, was not an opinion held or direction held at either OpenAI or Google or Anthropic or others. And so people looked down on us a little bit. They were like, “ is this really gonna work?” And so we just started working the problem, and we never really thought about open source again. We just kept our heads down and we built our, like, knowledge, understanding from scratch, right? We didn't roll out of an existing lab. So we picked up the papers and started writing code and figuring things out.Eiso Kant [00:04:59]: And it wasn't until the beginning of this year that me and my founder, Jason, picked up the open source conversation again.Eiso Kant [00:05:07]: And if you go back to some of the early things on our website, it was very straightforward. It was we wanna get to AGI, we wanna support a world of abundance, and we wanna be the first company that gets there.Eiso Kant [00:05:20]: But we started talking at the beginning of this year because it became obvious that the world was going in a direction that was starting to like, pick at us a little bit. Like, it didn't, this didn't happen overnight. It was, like, a little bit we were seeing this and we're like, “Okay, The world's going down a path.” And Throughout this journey, there was something that I used as a, as an analogy or thing. So I said well, if I go back to back in those days, 2015 or 2016, we're working on this, and I picked up a fi book off the shelf, and I was reading the book about 2035. AGI is achieved, and the story would be over the following, decades. And it would have that first chapter where everyone's trying to figure things out. You'd get the chapter of ChatGPT coming out And then you would get to the chapter where the world was at a fork in the road, and the one that it picked was one where three or four or a handful of companies were going to create all of intelligence moving forward.Eiso Kant [00:06:21]: And when I thought about that story, it felt like a dystopian fi book, not a utopian fi book. And the reality is, I'm a utopian fi guy. Like, and so We took a step back and said, “Hey, can we play a role here?” Now it was easy for us to do so because we were not at the frontier.Eiso Kant [00:06:41]: If we were at the frontier, I don't think we could have changed our mind. and I don't mean this like it's when the moment there's too much capital involved, too much expectations, you've built up things, right? We're a small team, just improving and improving. And so we knew that we could make that decision now, but it would be a lot harder to make as we got closer and closer to the frontier and caught up to others. And did a lot of soul-searching and a lot of conversations, and said, “No, this makes sense,” Even if there's big unanswered questions, like how the hell do you build a business model with foundation models about open source? Big open-ended question that we do not fully have the answer to yet, right? At what point do you no longer wanna release open source models because misuse of models has, real potential risks associated with it? how is the government gonna respond to open source? but I think it all just came down to one thing, and I'll stop the monologue, is the fact that I rather live in a world that has 100 foundation model companies than a world that has five, even if I was one of the five. And the smallest and most meaningful contribution we can make for 100 to exist is to open up our research and open up, like, our weights right now and figure out along the way how we can, like, do more.Neo-Labs, Model Choice, and the Token EconomySwyx [00:08:01]: Yeah. I think if anything, over the past three years, that has become a bit more true. you are one of a cohort of Neo labsEiso Kant [00:08:10]: YeahSwyx [00:08:10]: That people are now calling that. And, we're, we're doing this on the day that Thinky launched their, new model and you are outperforming them on their, on some benchmarks that they released, right? Like, they just don't have it yet. so it goes to show that I think, like, this is one of those things where, like, there is room for multiple players, and you are seeing a little bit more of the future. Maybe more like 20, not 100, but, like, you are one of the 20.Eiso Kant [00:08:36]: I really hope so, right? I think we I'm, I'm excited about their release, and I'm excited about everyone releasing because, like, ultimately, like, choice competition is both gonna drive progress in the right direction. But the fact that like, we create models and while we all, drink out of the same well of data effectively, we do introduce very different behaviors and biases in our models. Some are intended biases, some are completely unintended biases.Swyx [00:09:03]: Yeah.Eiso Kant [00:09:03]: And if we shape up in an ecosystem in the world where open models are gonna be a part of the token economy, like, I don't think there's any question about it anymore Then we want to be able to live in a world where companies, countries, people can choose and say, “Hey, I am most aligned and I trust most this provider for these things.”Swyx [00:09:25]: Yeah.Vibhu [00:09:26]: I think more than just one of the 20 Neo labs, up until recently, most of open source innovation was coming from the Chinese labs, right? So there's the DeepSeek of the West. Is it today? Okay, maybe it's thinking machines reflection, but there aren't many, right? So, one of the things you guys started in France, Europe, but very much now you're taking that American standpoint and more than just that, the point is the Chinese models that we see, they're not super open research. the work you put out is, I think, some of the best. So every few months you get not only frontier models, but also here's a breakdown blog, paper, technical report of here's everything for state of the art to build, frontier intelligence and you're filling that gap too, right? So not just only open weight, not just Western, but also pretty open research.Open Weights vs. Open ResearchEiso Kant [00:10:20]: No, I appreciate it. Look, I think it's, I think it's the most meaningful contribution, right? Weights are a binary. Let's call them what they are. Yes, we can modify them, we can change them, but, like, giving someone the weights does not allow them ultimately to recreate what you're doing, right? And so now there's challenges around releasing data sets, challenges around like releasing certain things, but being able to share your research, like, right, how do we do it? What are the lessons we learned that we spent, tens of thousands of experiments of compute on? I think very much so. One correction though, Vibhu, and I say this because it's been haunting us for quite a few years. We from day zero were an American company.Swyx [00:10:55]: Yeah. They movedPoolside's Global Team and American Company StorySwyx [00:10:56]: To France.Eiso Kant [00:10:56]: So the story once and for all is very. We start as an American company. We have always been an American company, and early on we made a very conscious decision. We said, “We're not gonna hire any researchers in the Bay Area. We're gonna look for talent everywhere else in the world.” and that is everything from Middle Americas, Seattle to, Serbia, and to Taiwan and Singapore and other places. And it was because we took a view that this was gonna become a talent war for this, and I think it has over the years now. Three years ago, that wasn't fully obvious yet. I think today it very much is. And we also realized that, like, some of the world's most capable people with, like, the most interesting, innovative ideas were not just gonna be here. And so it led us to create like a fully remote company. and we ended up opening an office in Paris and London and different places and we have a lot of the team in the US and a lot of team outside. But we always took this view of like, we're an American company, but if we want the best of the best to work with us, we need to take a global view. Now we do also have people here in Silicon Valley, like the company's grown and others, but I think one of the things that, it slowed us down at the beginning, but it has sped us up now, and it's why you're seeing like the progress, I think, on our models and the cadence at which we release, is because we didn't roll out of an existing lab. Right? we didn't, we didn't have a lot of the information that's freely flowing around here at the time. We just took this point of view as like, “Okay, well, let's just work the problem. Let's just go and, like, read the few papers that are out there, and let's just figure this stuff out.” And we made some hilarious mistakes in model training because of that over the yearsEiso Kant [00:12:35]: Like especially in the first 12 months. there's a few that I think still haunt me and scare me. We can talk about them later. but it created a, like, a resiliency and persistency in the team, right? with extremely few people have left us over the years, that, like, told us, “Okay, we can do this.” When we first wrote our first training code base completely from scratch, it wasn't a fork of any open source. It was just like, “Okay, let's build it from scratch.” I remember we had this one moment where we spent three weeks working out an optimizer bug. Like, it was like training just couldn't get stable. We, like, obsessed over it, and we thought, like, maybe we were wrong. Maybe we should have just forked this repo, or we should have. But then when we solved it, I still remember at the time we were like five people in the company. when we solved it, we were like, “Oh, we can do things,” like if we're just willing to work hard. and I think that culture with a very strong engineering bias has helped us, like, get to where we were. And so there's this notion of open source and talent and these things. I think we, We just took different decisions from a different starting point. and I think we are lucky. I do want to definitely call it lucky. And there was a lot of hard work at the team that now, like, that's starting to show up in results.Swyx [00:13:52]: Just ‘cause we probably won't revisit this again, but, and this is a fun recruiting challenge if someone knows the answer. What was the bug? And then we won't tell the solution, but we'An Optimizer Bug and the Value of Building From ScratchEiso Kant [00:14:01]: So the - This - You're gonna test my memory here,Swyx [00:14:04]: Oh, okayEiso Kant [00:14:04]: So but I thinkSwyx [00:14:05]: DirectlyEiso Kant [00:14:05]: I think I can recall. So if you, so if you look at, So if you take like Adam as an optimizer, you have epsilonSwyx [00:14:12]: YeahEiso Kant [00:14:13]: Which is, right, like in the denominatorSwyx [00:14:14]: Momentum and weights. YeahEiso Kant [00:14:15]: Is exactly, in the denominator. And at the time, if I recall, you looked at like the early Llama papers and things like that. People were juicing epsilon, like, quite a bit. Like, they were, like, adding, I don't know if it was E minus four or whatever, like a high value for epsilon.Eiso Kant [00:14:31]: And if you think about this during training, it's like a bit weird and counterintuitive that we're adding noise to our optimizer by just adding effectively, like, a random number in the denominator, right? Like behind the decimal point. And I don't recall the exact bug, but it had - What I remember is once we solved it, we no longer had to juice epsilon as much as, like, was happening in the Llama paper and other places. and it was like one of those fundamental moments where we had trusted this paper that was out there, and we're like, “Oh, no, it has to be this way. It has to have this high value of epsilon.” But it made no sense to us intuitively. Like, why do you have to have this so high? Like, if you're just trying to avoid division by zero, why can't the value be extremely small? and that was like one of those moments where you realize like, okay, finding things out from scratch yourself builds a better intuition. Because the one thing you learn very quickly with model building is that your intuitions that you start with are gonna get beaten up so hard.Eiso Kant [00:15:33]: Right? Like - It's such an experimental science, that the things that seem obvious, you very quickly get to learn, like, you were wrong, and hopefully you figure out why, and sometimes you don't even.Swyx [00:15:45]: Yeah. yeah, so, one of the reasons that you, when you released your new models, Vibhu got really excited. I mean, everyone got really excited. But Vibhu led our paper club on it, and you guys sawEiso Kant [00:15:58]: YeahSwyx [00:15:58]: Obviously. maybe talk through some lessons learned in that, whatever you can disclose. we can focus on the model factory stuff, whatever you think is a good starting point.Model Building as EngineeringEiso Kant [00:16:08]: So I would say that our view from very early on in the company was that model building is ultimately 90% engineering.Eiso Kant [00:16:18]: And I think we all know it in the industry because if you look at where's every researcher spending their time, they're spending their time writing code, right? Looking at data and writing code. And so we said, okay, The state at the moment, like three years ago, was bash scripts and Slurm and spaghetti code bases for training and, like, data pipelines that were patched together. And we looked at this and said, “Well, ultimately, model building is a process.” You're going from raw data, right? Like training raw material, the web, et cetera. you're doing a whole bunch of filtering, cleaning up, transformations, analyzing. These days, that's, far more complex than it was three years ago. then you're training a model, which is effectively a large distributed systems problem, right? Across hardware that has still-- It's become a lot more reliable. It was extremely flaky back then. and now with every new generation, we get our new sets of challenges. And then you go into the next stages, right? There was no training back then, but, like, you got, your post-training and then your reinforcement learning. And so we looked at this and we said, “Well, this looks like an industrialized process. This looks like an end process, that every single part of it has its machinery,” right? If it's your big data pipelines, if it's your crawling ingestion of the web, if it's your, large-scale distributed training, and then you've got your reliability. And we said, “Well, why don't we take some of the world's smartest distributed systems engineers that we knew and make them part of the process of research from day zero?” Not retrofitting it later on, but, like, really from the beginning. And that became our model factory. And so our model factory started with a handful of components. Today, it's thousands of components, and I try to equate it to, if you think about, like, someone who was at the very early days of Foxconn, if they had been there for the following, decade, they would be able to rebuild Foxconn because they saw every decision that led to building that system and all the complexity. If you and I walk into Foxconn today, no chance.The Model Factory and Experiment VelocityEiso Kant [00:18:18]: Right? Because we don't have the lineage and history of decisions that led to that. And so we built early on from the beginning- with a team that really understood that, well, the metric that we are optimizing for is the speed of an idea from a researcher to an experimental result that we can trust to then being part of the next model training.Eiso Kant [00:18:42]: And in the. And because it's such an experimental science, ultimately, in the beginning when it wasn't that complex, you could patch your way around it, right? But now, at any foundation model company, you are running. I mean, we're a small team, right? We're less than 70 researchers, another 35 engineers. and we are running, I haven't checked the latest count, but far more than 10,000, maybe 10 to 20,000 experiments a month that we cut. And so if you look at that scale of every model run that is, like it's ultimately it's, it's you need to be able to trust it as an infra problem. And so what we have now done over the years is gotten really good at that, and just by working it and improving it and obsessing over those end decisions. So now what that means is that you looked up Laguna XS 2 that we launched. It was five weeks from the beginning of training to launch. The model that we're gonna talk about today was eight weeks from start of training, to launch. We started the next model literally yesterday because we now finished the post-training required for the model we're launching, next week or by the time this comes out today. and we move that compute to the much larger Laguna M model that we're now training. And so the model should be an artifact of someone's process. It shouldn't be really a thing in itself. Like, and we treat this like the way you would look at like a SpaceX factory where, yes, the first rocket, really hard to build, but the much harder challenge was building the factory. And now they're rolling off, and no one is really thinking about the next launch anymore. So it's just another launch, it's another launch, another rocket comes off. And that's what we're trying to do with model building.Eiso Kant [00:20:22]: And what has been, which was not planned from day zero, it was in the back of our mind like this will happen one day, is that when you build a really good end model factory with really good APIs and really good engineering systems, Well, what is it perfect for? It's perfect for agents.Agents Inside the Model FactoryEiso Kant [00:20:40]: Because agents are now starting to take over more and more work in our model factory.Vibhu [00:20:43]: Yeah.Eiso Kant [00:20:44]: So I look at the screens when I walk, like when we're, we come together, in our monthly, we do monthly onsites, and I walk behind people's screens and I stop by and I talk to our researchers. And the default is all of these different agents running on their screen that are writing the code. They're launching the jobs. They're evaluating the results that are coming back from the model runs. They are, making the changes. And we're still in the driver's seat. We're still coming up with the ideas. We're still helping with the debugging. But more and more, and this is right now very profound on the data side of our pipelines in both pre and post and the synthetic data pipelines, it's starting to become more on the architecture side as well. You're starting to see these twinklings of what RSI is gonna look like.Eiso Kant [00:21:27]: And that's. So when we talk about, like to your question about our models, every talk about the model factory, And my coolest example of these things is always that when we kick off a new run, doesn't matter if it's a training like big run or if it's now a post, like one of 10 post-training versions we do for like release or many experiments, is that at any given moment, the changes that somebody made that they had experimental results from the day before make it into that run.Eiso Kant [00:21:57]: So there's not like a cutoff 90 days before. Like no, it's like literally from that moment because we can now trust the machine enough. And then you also have to invest in the reliability. So one of my favorite metrics about like Laguna S is that there was no call events, Right? Like completely zero. And we haven't had a meaningful call event, like something to wake up for, as far as I recall this entire year. now there is one asterisk to that. In usually the first six hours of launching a new model run, something breaks because you set a config wrong, you made a small mistake, et cetera. So that's usually there's a little bit of intervention, but that's always within like call periods, right? Not on call. And I think that's starting to now compound. So the model we're releasing now, I love it. It's amazing, but we're already onto the next one. and I think that's the way it should be.Laguna, Five-Week Builds, and Zero On-Call EventsVibhu [00:22:50]: Hey, I also just wanna point out, so for context, this was like a month ago. we found it in the tech report, so we just came in with, “Okay, new model's dropped. Haven't heard about it.” We wereEiso Kant [00:23:02]: Yeah, we're very used to doing this every few months.Vibhu [00:23:03]: We're, we're very much like, “ okay, look, it's like, on par with Kimi, DeepSeek, whatnot, the small ones, Gemma level. Oh, it's a very cool paper on what goes into building.” And then we hit this page, right? Like literally page two of tech report is, “This process allowed us to build the small model from scratch to delivery within five weeks applying the lessons”. And then I'm like, oh, this paper is not about here's a tech report of benchmarks and here's how many tokens it was trained on. Like for people that wanna dive more from what we're not gonna discuss on the podcast, it's all laid out here, right? FromEiso Kant [00:23:38]: YeahVibhu [00:23:39]: Custom software that agents can use to interface with training code, training data.Eiso Kant [00:23:45]: Yeah. Well, link the paper correctly, so yeah.Vibhu [00:23:47]: Yeah. All that stuff. read the paper here, but,Technical Report Principles and Streaming Training DataEiso Kant [00:23:50]: But I would like to. I love principles, and I think that is a good starting off point for maybe telling some stories. Maybe we can go one by one past the principles. I'll just call out that Dagster just got bought by a Prefect.Vibhu [00:24:01]: Yeah.Eiso Kant [00:24:01]: Isn't it fun? But yes, I'm very familiar with Dagster. just anything where like they trigger some story.Vibhu [00:24:07]: So, well, I would say, well, experiments code's obvious, but I think one of my favorite things is, I don't know where it is in here, but early on, and I still think this is the case a lot of foundation model companies, people prepare their training data sets, they get packaged up, then they get copied over to a training cluster distributed across all of the nodes, and then training starts.Vibhu [00:24:30]: And we looked at this like three years ago and we were like That makes no senseEiso Kant [00:24:36]: You lose so much time because the moment you have to rematerialize the data set, you have to make a change, you have to fix something, et cetera, you've got all this time of like repackaging it, right? Toca- tokenizing it, repacking it, moving it over to a cluster, then distributing it across the nodes. The bigger your clusters are, you start using fancy like torrent-like algorithms to like distribute your data. So why aren't we streaming data into training? Right? Something that's very common and like just basicVibhu [00:25:00]: Like just in timeEiso Kant [00:25:01]: Just in time, like good computer science like principle. And that was one of the first things that I think unlocked - the model factory. Because the moment you start thinking about, well, a training job, it doesn't matter if it's a big hero run or a small like, post-training experiment, consumes a certain number of tokens per second, right? And it's not a lot, right? From a like a data, moving data perspective. So we said, well, we have our training cluster, and then we've got like our AWS kinda setup where we can build these amazing big data pipelines. We can set things up. We use Spark underneath the hood, like all these things.Vibhu [00:25:36]: But when you say AWS, it's not actual AWS, it's your internal AWS.Eiso Kant [00:25:39]: It's our internal-- No, it's our internal like just running like our infrastructureVibhu [00:25:42]: Site web servicesEiso Kant [00:25:43]: Exactly. Our stuff running on like an AWS account or on like any hardware, right?Vibhu [00:25:47]: Yeah.Eiso Kant [00:25:48]: And so once we made that shift into I can stream data into training, all of a sudden you realize a lot of things unlock. Because now you don't have to wait for the whole data set to materialize.Immutable Data, Experiments as Code, and Scientific RigorEiso Kant [00:26:00]: You now all of a sudden when you're running data experiments about mixing data, it's a config. Because you've got these data sources that are coming in, and you just - we have this service called Blender that's in the report, where we then say, “Okay, for this run, I want 20% of this source, 10% of this source. I want this much, so many epochs of repetition. I want this to be, shuffled in a certain way,” and your training job can start while the rest of the data is even still materializing. also what it does is because all of this underneath-- So for us, we treated the data layer underneath as like an immutable data layer, and that was really important. Like experiments as code, immutable data layer means that you can always go back and understand literally down to the single token at which cursor it went in on which version of the code.Vibhu [00:26:47]: Yeah.Eiso Kant [00:26:48]: And it took us a I have to admit, like the first year of Poolside, we understood that engineering had to get great, But we didn't understand yet, that this is ultimately in support of like a good rigorous scientific progress. We were quite a - We were a very small number of people, so a lot of it was YOLO ideas and YOLO runs.Vibhu [00:27:08]: Yeah.Eiso Kant [00:27:09]: And we built great infra for the YOLO runs. But once we realized that we treated data as immutable and code as always versioned, and you could always track and trace every experiment end to end perfectly, you could repeat everything perfectly, right? You have perfect reproducibility. I can still reproduce runs from two years ago if I wanted to, right? It enables the scientific progress, like the scientific process, and I think that took us probably about a year and a half into the company to figure out. We also had some great hires, like our head of applied research, Nikolai, who joined us from Yandex, who'd been working on language models since like the early 2020s, I think brought that into the company of like, “Hey, we wanna have even more rigor.” And then once we kinda had the combination of like increasingly more capable platform that allowed people to do more, but had this immutability, we were able to start “Okay, every experiment is truly an ablation. We truly need to understand it.” And I think we became much more scientifically rigorous in the last couple of years, and the infra underneath enabled it. and then there's just fun stuff like, andVibhu [00:28:16]: Yeah, a lot of it's fun, like even just the, one, you share all the ablations, two, picking the data sets, right? There's like a random small paragraph in here where it's just like, “Oh yeah, training data, we have some, we have an auto mixer.” it trains eight small models, scales them up, picks the training data set. We don't even need to look at it. I'm like, “Wow, a lot of engineering rigor there.” And there's just, there's just a lot in here.Publishing Research and Giving BackEiso Kant [00:28:40]: Yeah, and it'- and look, and we wanna put out more. Like we, We treat writing papers as something that we haven't earned the right for yet for a long time. So you earn the right to spend time, publishing research once you're at the frontier, because until then, you're catching up, and every minute and hour in this industry matters. Like I obsess over it, not just the wall clock time from idea to result, but just general like time every day that we, waste is one that doesn't allow us to catch up. But in this case, we said, “Okay, we're gonna give ourselves.” I think we gave the team like three or four days while still doing their work, like give everything in there. And to your point earlier, if your stuff, it's easy to like put it out. And so there's so many more things that we wanna talk about over time, and we will definitely start doing. And as we earn more of the right, but also now have like added to our mission that we want more foundation model companies to exist, you'll see us like be way more proactive, and just trying to keep dropping some of those like things that we've learned along the way that can help others like speed up.Vibhu [00:29:40]: Which is the other cool side of this, right? It's, it's not like, back to your point, it's not just here's the benchmarks of our training. If you want to replicate, here's experiments of optimizers, data sets, post-training. you lay out a lot of it here alongside here's your system for how to do it? So it's, it's really like promotingEiso Kant [00:29:59]: No, thank youVibhu [00:29:59]: Other people can do the same.Eiso Kant [00:30:00]: And by the way, I also wanna make clear, right, we have been incredible-- Like we've taken a lot of advantage of the fact of all the open research that others have published, Right? And you mentioned, the Chinese labs, and we I think it's important that there's, from every country and every culture and background, including like Western companies like us, there's different models that come out that people can choose to trust. But I think we do have to give credit where credit's due, right? The incredible Chinese lab have done an amazing job at sharing their research, and we have definitely like been on the receiving end of taking advantage of that. So when you're on the receiving end of something coming to you, I think it's, you also have an obligation to give back.Swyx [00:30:39]: Do you have a favorite or underrated Chinese lab that you wanna shout out? Everyone shout outs DeepSeek.Chinese Labs, Zhipu, and PersistenceEiso Kant [00:30:44]: That's a good question.Swyx [00:30:45]: Moaan obviously for Therapsi. Yeah.Eiso Kant [00:30:48]: Yeah, look, I think, I think obviously everyone's been talking about Zhipu lately, with 5.2. I think what most people don't realize is when they started.Swyx [00:30:59]: Yeah.Eiso Kant [00:30:59]: Right? They started years before ChatGPT.Swyx [00:31:02]: They just rebranded. YeahEiso Kant [00:31:03]: And so, I've like, I remember how hard it was to work on these things Before the rest of the world got excited about it. And so I have an immense amount of respect for people, who were working on improving models when it wasn't the sexy thing to do, when believing in LLMs, was gonna get you ridiculed. I remember like back in 2016 when we were doing what we'd call, machine learning on code with some of these models. we would-- people would just laugh at us, like they'd be like, “This makes no sense. Like why are you wasting all these, like, millions of dollars on trying to figure this out?” And so I would say they're probably the one that, I think deserves a shout-out, not just because their latest model is very good, but because they fought to get here. And I think, I think every foundation model company it takes time to get here, right? It took us three years to get to the model that we're, that we're now gonna be releasing. and now the time in between the models is coming, is counted in weeks. It's no longer counted in months or years. But this stuff's hard. and if we can make it a little bit easier for the next person, like we should all do so. Because if we don't do so, we're, we've got a small window before models are really impacting recursive self-improvement to a level where catching up otherwise might become unfeasible. And we should try to, in that window, encourage as many labs or however we wanna call them, like to start. And so one of my currentEiso Kant [00:32:36]: Mission, but qualm is like I wanna encourage whoever is a researcher right now who thinks they can tackle this to go and leave and become my competitor.Eiso Kant [00:32:45]: Like start another foundation model company because I think we need it. I think otherwise we're not gonna be in the world where, I don't want to just be the fifth or the sixth company that wins. I wanna look at a world where there's lots of choice.Starting a Foundation Model CompanyVibhu [00:32:57]: What else do people not see in starting a foundation model? it's, there's a lot of compute, there's a lot of capital required, a lot of compute. You lay out model factory and how to do the training, but there's a lot there, right? That's,Eiso Kant [00:33:10]: Well, look, it's, I in turn-- this is an oversimplification, and I always asterisk it with that because it can land a little bit the wrong way in people's minds. But I think you can sum down, And I saw it, 95% of model building to just doing, you're just doing two things. You're improving data or you're improving compute efficiency. And I know that feels like an oversimplification for the incredible, like, Gifted and skilled work people do. But if you really look at it, like what are we doing? We are looking at data, we're generating new data, we're improving data. and the only way to do that is to look at the data, right? That's a big part of foundation model building. And on the other hand, we come up with these incredible breakthroughs in inference, in architecture, and new attention mechanisms. But what are they really doing? They're bringing compute efficiency. Now, we have definitely had some breakthroughs over the years that allow for more model capabilities. But at the limit, if you could train a large enough model, right, like, and you had infinite compute, we probably-- if you had infinite compute, you'd be at AGI probably already tomorrow.Eiso Kant [00:34:12]: Right? Like it's not. And so, and let me say that infinite compute with infinite ability of much faster networking because networking ends up being more of the bottleneck than compute. But, so I do think that's, those are the main things. And to just realize that this is engineering. I think it's become more obvious, but I think for quite a few years, people have held foundation model companies and researchers and others on this pedestal of like you're doing incredible magic or rocket science, or only like, Nobel laureate physicists can do this. And don't get me wrong, there are some really hard problems that need to be solved, but a lot of the work that all of us are doing on a day Is not sitting down trying to solve a math theorem. A lot of the work that we're doing is just really doing the basics right, writing good code, looking at data, improving it, running experiments, looking at plots, trying to see like, hey, trying to shape our intuitions. And a lot more people could be highly capable researchers. and I think that's, it feels far for people to do so. But I've seen in our own company, we've seen engineers become researchers because the model factory allowed them to be, have a much lower hurdle of running experiments and trying things. And one of the guys on our team who started as an engineer building our agents is a legit reinforcement learning researcher now, making real progress. and that happened in the span of like six months. that would've not been what I think most people assumed was possible, a couple of years ago.Swyx [00:35:46]: Yeah. I think one of the interesting moments is when you can self-host, like, if in a programming language, like if you can compile the language in the language, the equivalent is can you use your own tools, right? You have the pool CLI, you have your own models. presumably you're not only using your own models. There's no way. But like, what's that percentage over time?Laguna S, Persistence, and Behavioral GainsEiso Kant [00:36:10]: This is the first model that we're releasing that is starting to meaningfully contribute to our own work. It's not a it's not state-art model yet. Fable and other, they're, they're very capable models, but Laguna S Is really interesting. I'm gonna pull up the quote. Peng Ming, one of our heads of applied research, said something, last week as the model came out about 10 days ago, much better than we had hoped for or expected. And he said, I have the feeling that a lot of the gains in Laguna S come not from more intelligence, but more from different behavior, more verification, less taking things for granted, not declaring victory early, and being way more persistent. And to be honest, those are more predictive than raw intelligence for success in human also to some degree. And this was, he wrote me this on 5th of July on a Sunday, and it's been burned in my brain ever since because the Laguna S model, as you'll see it and why it does so well on benchmarks and why it does so well in using it on a day basis, is that it's just incredibly persistent. It reasons a lot. I do call that out. We have work to do on making it more efficient. We have to work to do on offering different reasoning modes. But this is the model that has been able to do things that I never thought it could do. A hundred eighteen billion 8B active model, which is not that large. It fits on a DGX Spark and still runs at, thirty, forty tokens a second on a Spark, is able to solve Erdős 397 independently. It's able to do complex programming tasks. It's able to. I asked it this morning to make me a Fi scanner without using any external libraries on my Mac, and it's, like, figuring out, like, the core WLAN API by really persistently trying to understand it without access to the internet. And more, I love vibe checking. I've probably spent eight to ten hours a day with this model for the last ten days.Eiso Kant [00:38:05]: I'm not exaggerating. I was on my eleven-hour flight yesterday. I spent ten hours reading trajectories and traces and, like, of the model.Eiso Kant [00:38:12]: And what I take away from it is exactly what Peng Ming said. We are gonna be able to squeeze so much more out of smaller models than I think we had imagined in the industry because, yes, there's intelligence and larger models are more intelligent. Like, no doubt about it. We should continue to scale up. but the behaviors of being really persistent, of being able to backtrack when you're wrong, of, like, understanding how to interact with your environment show us that we can get a lot more out of it. And this, for me, has created a bit of a Question in my mind the last couple of days. If you think about where we're using models today, right? We are using models, say, for knowledge work. Represents twenty-five percent of the global economy, twenty-five trillion dollars of work.Eiso Kant [00:39:00]: As we scale up models and they become more intelligent, we are excited about using them more and more for pushing the frontier of science.Small Models, Knowledge Work, and CommoditizationEiso Kant [00:39:08]: And if you look at the frontier of science, like true breakthroughs in science, they have been linked, they are linked to more intelligence in many places. Einstein figuring out general relativity is able to bring ideas together that other people would have not brought together. And I think one of the many dimensions of intelligence is the ability to do that, and it's something we clearly see that as models get larger and more capable, they're able to pull more ideas and threads together that a smaller model wouldn't be able to.Eiso Kant [00:39:36]: And we're starting to see examples of that in medicine and, like, in bio and other things. But if you think about the majority of knowledge work that we do, and it includes building software. I'm a software developer at heart first and foremost probably, although I probably can't say it that much anymore as I don't write production code in years, is that what makes us good is our persistence. It's our ability to encounter a problem and backtrack and say, “I need to go figure out this bug. I need to go research this. I need to go look at the documentation. I need to, like, try different, five different ways to see, like, if I can solve it.” But it is not necessarily bringing three ideas together from radically different fields. And so if we are now seeing, and I think Laguna S is an example, that we are able to make a relatively small model much more capable than I had definitely predicted or any previous, like, benchmarks had shown for any model remotely this size or even larger, At least on coding tasks, that it's because of the behaviors. And so now the question I have, and I don't have an answer, it is I know at the limit, so infinite model size, right, extremely large model, and the cost of that model is gonna be very expensive to run. We know this, right? So larger model ROI.Eiso Kant [00:40:52]: So I know that at the very limit, I'm not gonna use the world's largest model one day, quadrillion parameter, whatever crazy, like, scale we scale up, to do a basic coding task. Already today, I'm starting to size down for certain tasks.Eiso Kant [00:41:07]: So it means that there is an optimal. It means there's some curve that goes as we go up to model size for knowledge work, at some point we're at the peak, and after that, the return on investment of using a bigger model, just doesn't make sense.Eiso Kant [00:41:22]: Now, I think the question is, before I would have thought that peak was extremely very far away.Eiso Kant [00:41:30]: This model for me is the first sign that Maybe that peak is At a trillion, five trillion, ten trillion. Maybe we can just squeeze way more out of these models. I'm no longer thinking that we need two or three orders of magnitude on the largest models to be able to, solve knowledge work, the accounting, the legal, the code that we write. And so if that holds true, It is an argument for the commoditization of models. It's an argument that open source can win and, like, succeed in this world. And now it's of course a self-serving argument and it's a hopeful argument, but theoretically at the limit it works. We just have to go discover in the next couple of years of how much more we can squeeze out. Now, I do want to put a big asterisk. This does not mean I'm against scaling models. I think we ultimately only succeed if we scale our models as large as our competition. I do not like. I think we should not put our head in the sand and say we're gonna be king of open source small models. I think that's, It's a out. It's trying to be king of your own kingdom, but not realizing what the rest of the world's doing. All of us rather use a smarter, faster, more model. It's a sign of hope. And so I don't wanna overly state this is a good model. We have a long way to go to get to the state-art. But what hopefully people take away when they use this model is that the behaviors inside of it are what push it to be far more capable, less than necessarily the number of parameters.Pre-Training, Mid-Training, and RL Moving EarlierVibhu [00:43:03]: Is that mostly post-training? LikeEiso Kant [00:43:05]: YesVibhu [00:43:05]: Right.Eiso Kant [00:43:06]: It's entirely post-training.Vibhu [00:43:08]: Are we done improving anything on training? Is, like, training done?Eiso Kant [00:43:12]: No.Vibhu [00:43:12]: Okay.Eiso Kant [00:43:13]: SoVibhu [00:43:13]: I just wanted to cover training, and then we go post-trainingEiso Kant [00:43:15]: Training is not done. I mean, look, there's a part of training of just dealing with skill, right? Every new order of magnitude of model skill, you are going to get new things you gotta solve for. That'- but those are ultimately, engineering challenges.Eiso Kant [00:43:31]: I have a, I would say, a not commonly held opinion that reinforcement learning Will move earlier and earlier into training.Vibhu [00:43:42]: Yeah, training.Eiso Kant [00:43:44]: Not even training. Like training today, right, is, like if you look at - So we've been working on this for years already. and I think the best-- I think the first time we saw it out in public was the DeepSeek Zero paper. this is a year and a half ago, I think, if I recall correctly. where, you can Very early on in a model as it starts capable of being able to use language, et cetera, induce reasoning. and so the question that I have is like, we have this- we have the dataset that's the web. and the web, I think we could arguably say probably has The totality of humanity's knowledge somewhere encoded in different places. It's a huge variance degree of quality, from garbage data, and like once you look at training data, you really get humbled of like what the web is, to like, the most greatest scientific papers and best blog posts and like, best transcripts and whatnot.Eiso Kant [00:44:39]: And so now What we are trying to figure out, and have been doing a lot of work on, and it's a place where maybe not as open as we're on other things, but we will become more over time. we've been spending a couple of years really doing research on how can we turn the web into not just next token prediction, but into a way to teach the model to think earlier in its training. and I think there's a huge amount of gold to be found there. I think we are right now in, we've got some drugs in the industry. One of the drugs is distillation. Another drug is, more environments. Like, and they're great, and they make us feel good, and they make the models better, and like we're all addicted to them, and we'll use them, right? in various different ways. and but ultimately, I think we are still barely squeezing out of the web what we should be getting out of the web.Eiso Kant [00:45:33]: I think just next token prediction during training is not enough.Eiso Kant [00:45:36]: AndVibhu [00:45:38]: YeahEiso Kant [00:45:38]: I think we'll see some very interesting things still happen. and that RL in post-training to induce behaviors, to improve things, like I think - the whole world knows how to do this now. I think we're, we're scaling it up. Everyone is. But I wonder if we need to go as far as we're going today with environments. I'm not sure yetVibhu [00:46:01]: You mean we're going too far?Eiso Kant [00:46:02]: I'm, I'm not sure if the path to AGI is justVibhu [00:46:06]: Is more environmentEiso Kant [00:46:07]: More environments.Vibhu [00:46:08]: It seems like a never-ending, “Okay, I want instruction manual for this table, right? Am I gonna environment out building furniture? Or are we just gonna tail end like we need some general solution?”Eiso Kant [00:46:19]: I think there is, I think there's an ability to generalize more from the web. but I also am very encouraged, like when I look at Laguna S and, which is post-training is, well, is the big impact there. and I see like, oh, wait a second, just by making some of these behaviors much better, we're able to get so much more out of it. It just changes a little bit the way you think about intelligence.Vibhu [00:46:40]: Yeah. The analogy people draw often is the RL phase is where you don't learn as much new knowledge. You shiftEiso Kant [00:46:46]: Yeah.Vibhu [00:46:46]: Yeah. So, you shift distribution, and you can have it reason towards what you want. on your point about training, a lot of training is still just continue training in a domain, say medicine, then you do RL. So still justEiso Kant [00:47:00]: It's just better data, right? Like, I mean, training, ooh, I like how we invented this word. Like it's effectively just like,Vibhu [00:47:06]: Second phaseEiso Kant [00:47:07]: It's the second phase of training With like a really dumb way to do a curriculum. But like ultimately, what you'd want is a curriculum from token zero to token 30 whatever or 40 trillion tokens that really truly is the optimal curriculum for the model to learn. But training is essentially a stage curriculum on the web because we do not have to compute, And, effectively to try to ablate the perfect curriculum, right? And so I'm pretty sure that you'll start to see people talking soon about some other term, and there's two or - ‘cause now we do this, right? We talk stage two and stage three and stage four training and like. But ultimately, all we're doing is we're trying to assign a curriculum to the web data that we have to allow the model to learn better. I think at some point, as things get compute, as models get cheaper to run, as the next generations of compute, this will become more of a continuous spectrum. I also think the reason, by the way, you have training and like stage two and stage three is organizational, Right? It'- this is, I think, a thing where-- that we really try to avoid with the model factory is like Training exists because there's a training team now, right? There's people, or like people in training decide to focus on like a training effort. but what you really want is engineering and scale of experiments that allows for a much more continuous spectrum that you don't, you have infinite stages. Now, we're not there. Compute's not there. Organization design is not there for it yet. but I think we'll get there. we'll look back on a couple of years and be like, “Oh my God, it was so cute that we did our training data like this in such a like naïve way. Like we barely ordered it. We didn't really do a good job at likeCurriculum, Auto Research, and New ObjectivesVibhu [00:48:48]: The building that curriculum will get you that in the industry.Eiso Kant [00:48:51]: And I'll confirm that, when I talk to some researchers that this is a lot of the focus now is like how does training change and what is the next objective other than, next token prediction. I assume you don't have the answers, but you have some ideas.Vibhu [00:49:02]: We have some ideas. We're not ready to talk about it yet.Eiso Kant [00:49:05]: Yeah.Vibhu [00:49:05]: We've been working on them for years, and I think that's the one thing that's also like you asked earlier about, like what's not obvious about building a foundation model company is that you are constantly balancing the table stakes work, the recipe worksEiso Kant [00:49:19]: Yeah.Vibhu [00:49:19]: Versus like your, my crazyEiso Kant [00:49:22]: Pure researchVibhu [00:49:22]: Breakthrough.Eiso Kant [00:49:22]: Yeah.Vibhu [00:49:22]: Pure research and finding that balance and adjusting the percentage to it based on where you are in the race is really important.Eiso Kant [00:49:31]: I mean, so like, this is a nice way. I was gonna bring up auto research at some pointVibhu [00:49:35]: YesEiso Kant [00:49:35]: As another Andrej invention, or coinage, which is like, I honestly, like how many objective functions can there be, right? Like just try 1,000 of them, set it running, whatever.Vibhu [00:49:47]: Man, it's alsoEiso Kant [00:49:48]: Like what you're looking for. You're looking for loss curves like that, likeVibhu [00:49:51]: It's also a thing people take bets on, right? When you say more Neo labs, you're doing a version of we'll do foundation models, scale them up, next token predictors. A lot of other Neo labs that we see want to take a completely different approach, right? At some level, you're right. It's all, compute efficiency, and that's the net objective. But some are okay, different architecture, like vastly different amounts of compute spend. So some are different. They're not justEiso Kant [00:50:19]: YeahVibhu [00:50:19]: They're like, 99% not balancing, here's the vanilla and scale up. They're 99% on, here's novel research that'll change everything.Eiso Kant [00:50:27]: And I think, Luke, I think you. It depends when you started as well, right?Pure Research vs. Table StakesVibhu [00:50:30]: Yeah.Eiso Kant [00:50:30]: When we started, like the novel thing we did was reinforcement learning on code. No long- that's no longer novel by far, but we were like, - that's where we obsessed over when no one believed in RL. So you have to when you start the company, you have to have your own idea. You have to have something that's different that allows you to speed up, right? For us, it was RL to LLMs that later became common, like, Knowledge. But in the beginning, it wasn'tVibhu [00:50:53]: It's cool. this was like your original 2023 blogEiso Kant [00:50:57]: YeahVibhu [00:50:57]: Of purpose.Eiso Kant [00:50:58]: Yeah.Vibhu [00:50:59]: And like you do lay it all out here.Eiso Kant [00:51:01]: We laidVibhu [00:51:01]: The blog is pretty underrated, right? The whole RL on code was very early on.Eiso Kant [00:51:06]: Very early. And even we had to argue with people, like we say here things like to push beyond current capability, to train your own foundation model. We had to argue with people that it mattered that you had your own like, base model. you can fine-tune your way to success, right? major capabilities emerge from training a base model made accurate and useful during fine-tuning.Vibhu [00:51:23]: Which like, for perspective at the time, we knew closed models, OpenAI, Anthropic were huge. The open models we had were like Mistral 7B, a 30B, a 70B.Eiso Kant [00:51:35]: When weVibhu [00:51:35]: YeahEiso Kant [00:51:36]: The date on this thing is wrong. When we published this, it was April 2023. I think this was justVibhu [00:51:42]: YeahEiso Kant [00:51:42]: Happened on a migration, probably found it on archive.org.Vibhu [00:51:45]: Mistral.Eiso Kant [00:51:46]: Mistral had started, we started on the same month, right?Vibhu [00:51:49]: Yeah.Eiso Kant [00:51:49]: So this wasn't even, there was only, I think, Llama out at the timeVibhu [00:51:52]: SnellEiso Kant [00:51:52]: And that's it, right? And so, but I agree. I think we wan

The MAD Podcast with Matt Turck
The Biggest Chip Ever Built — Why OpenAI Runs On It | Cerebras CEO Andrew Feldman

The MAD Podcast with Matt Turck

Play Episode Listen Later Jul 23, 2026 72:41


AI is no longer just a race to train smarter models. As AI moves into production, the bottleneck is increasingly inference: how fast models can generate tokens, use tools, reason, verify, and act. In this episode of the MAD Podcast, Matt Turck sits down with Andrew Feldman, co-founder and CEO of Cerebras, to explain why fast inference may define the next era of AI.Cerebras is known for building a chip the size of a silicon wafer. But this conversation is not just about one company or one chip. It is a deep dive into the AI infrastructure stack: GPUs, ASICs, memory, HBM, SRAM, data centers, power, TSMC, AWS, OpenAI, agents, reasoning models, and why speed changes what AI products can become. Andrew explains why “tokens per second per user” matters, why generating a single word can require moving the equivalent of 100 HD movies through memory, why agents amplify latency, why GPUs struggle with certain inference workloads, and why fast AI may eventually reshape SaaS itself.This is a reference conversation on fast inference, AI chips, and the next compute bottleneck.(00:00) Cold open & Intro(01:31) Why speed became the AI bottleneck(02:32) Tokens per second per user, explained(03:16) AI's broadband moment and the Netflix analogy(04:35) The AI chip landscape: GPUs, TPUs, Trainium, ASICs(06:36) What is an ASIC?(08:08) Nvidia, Groq, and the fast inference war(09:16) OpenAI, Broadcom, and specialized silicon(12:10) China, power, and sovereign AI infrastructure(15:05) Is the AI infrastructure boom a bubble?(18:56) The hidden bottlenecks: HBM, CoWoS, and 3nm(22:57) Why agents are creating CPU demand(25:36) Andrew Feldman's path from SeaMicro to Cerebras(26:13) Why Cerebras bet on AI in 2016(31:14) SRAM vs. HBM: why inference is a memory problem(33:19) What wafer-scale computing actually means(34:28) The deep-tech “Everest” problem(36:07) The moment the first Cerebras system worked(36:49) Ringing the bell and surviving deep tech(39:08) How a giant chip handles failure(41:22) Why GPUs struggle with decode(42:17) Prefill vs. decode explained(44:01) The “100 HD movies” problem in AI inference(45:04) How fast inference changes RL and training(48:08) Reasoning models and why they cost more compute(50:08) Verification, guardrails, and small models checking big models(52:37) Multimodal AI and the path to video(53:51) Cerebras' business model: hardware, cloud, and API(55:14) OpenAI's 750MW inference deal(55:36) Why data centers are measured in megawatts(58:01) AWS Trainium + Cerebras decode(59:29) Fast tokens as a cloud product(01:00:52) Is CUDA still a moat?(01:03:53) How TSMC helped Cerebras build the giant chip(01:07:41) Why nobody cared in 2020(01:08:15) Why chip supply chains are hard to diversify(01:09:54) Why today's AI models will be the worst you ever use(01:10:38) What fast AI could do to SaaS

Daily Stock Picks
⚡A Top 2026 Pick I'm Buying On This Dip PLUS 3 Energy Names And 1 AI Power Stock Ready To Rip

Daily Stock Picks

Play Episode Listen Later Jul 22, 2026 31:15


This was a great episode on how I'm finding opportunities, how being patient, disciplined and prepared are the keys to successful portfolio management. FORMULA - ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Alpha Picks + Seeking Alpha Premium ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠+ ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Trendspider and Sidekick⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ - PERFECT TOGETHER! THESE SALES END SOON: I negotiated to get 59% off and 100 Sidekick messages per month for the entire year. Plus you get my 4 hour algorithm and so many other benefits with ⁠JUST THIS LINK ONLY ⁠⁠CLICK HERE TO GET THE DAILY STOCK PICK SPECIAL OFFER - ONLY ANNUAL PLANS AVAILABLE ⁠⁠Trendspider is having their own sale with up to 50% off monthly and other plans - this would be for anyone who wants to try it first. ⁠Seeking Alpha's SUMMER SALE ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠✅ *BEST DEAL - SEEKING ALPHA BUNDLE - Save over $150 and get Premium and Alpha Picks together ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠- ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠✅ ALPHA PICKS - Want to Beat the S&P? Save $50⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠✅ Seeking Alpha Premium ONLY - FREE 7 DAY TRIAL ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ GET ACCESS TO THE TOP 2026 2H EVENT FREE ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠✅SEEKING ALPHA PRO - YOUR FIRST MONTH ONLY $89 ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠EPISODE SUMMARY

Mark Vena Tech Guy Podcasts
SmartTechCheck Podcast and Audio Newsletter: Keep GPUs Humming

Mark Vena Tech Guy Podcasts

Play Episode Listen Later Jul 22, 2026 16:40


The Vergecast
The car of the future is an EV golf cart

The Vergecast

Play Episode Listen Later Jul 21, 2026 30:05


So far, most electric vehicles have looked more or less like cars. But recently, a few companies have looked to another, smaller mode of transport for inspiration. The Verge's Andy Hawkins explains why companies like Amble and Chip are reinventing the golf cart, in the hopes of creating an entirely new kind of street-legal vehicle. As long as the speed limit stays low. Further reading: ⁠The Light Flip is a minimalist flip phone with a point to prove | The Verge⁠ ⁠Anthropic has to pay authors. | The Verge⁠ ⁠The cost of GPUs goes far beyond AI data centers | The Verge⁠ ⁠Is America ready for this quirky Jeep-looking EV that can park itself?⁠ ⁠The ‘G-Wagen of golf carts' could be the ideal second car⁠ ⁠America's cheapest new EV is smaller than a ping-pong table and tops out at 19mph⁠ ⁠Zoox's purpose-built robotaxi is getting a refresh⁠ Subscribe to The Verge for unlimited access to theverge.com, subscriber-exclusive newsletters, and our ad-free podcast feed. We love hearing from you! Email your questions and thoughts to vergecast@theverge.com or call us at 866-VERGE11. Learn more about your ad choices. Visit podcastchoices.com/adchoices

The 7investing Podcast
Rocket Lab & Netflix Deep Dive | Cerebras vs. NVIDIA: The $30 Billion AI Inference Battle

The 7investing Podcast

Play Episode Listen Later Jul 20, 2026 43:11


Is Cerebras Systems the next great AI chip stock or a red-hot IPO priced for perfection? In this episode of 7investing Live, Simon Erickson and executive producer Heather Horton welcome back Nick Rossolillo, co-founder of Chip Stock Investor, to break down three of the market's biggest stories.First up: Cerebras Systems (NASDAQ:CBRS), the wafer-scale chip maker that just IPO'd at a $40+ billion market cap. With 44GB of SRAM embedded directly on the chip, Cerebras was purpose-built to solve AI's "memory wall" problem for inference workloads. Now it's reportedly landed a ~$10 billion order from OpenAI and a deal with Amazon Web Services that could top $20 billion. Simon and Nick dig into whether these massive orders are real, how Cerebras stacks up against NVIDIA's GPUs and hyperscaler custom silicon, the TSMC capacity bottleneck that could throttle its growth, and how to value a company trading near 20x sales without profits.Then the conversation turns to Rocket Lab (NASDAQ:RKLB), which has pulled back from $150 to around $70 per share. Simon shares the latest iteration of his discounted cash flow valuation, and the duo debates the proposed Iridium acquisition — a deal that could pull Rocket Lab to EBITDA-positive on a pro forma basis — plus what the long-awaited Neutron rocket launch means for the company's future.Finally: Netflix (NASDAQ:NFLX). After another quarter of decelerating revenue guidance, is the streaming giant now a value stock rather than a growth stock? Nick explains why the advertising business hasn't reaccelerated growth the way he expected, and what he'd need to see before buying the dip.Plus: Nick's take on the recent chip stock sell-off across NVIDIA, AMD, Broadcom, SanDisk, and Kioxia and why "stocks go up, stocks go down" might be the healthiest way to think about it.Subscribe for more deep dives on AI infrastructure, semiconductors, and innovative growth stocks!Start your FREE 7-day trial of 7investing: https://www.7investing.com/subscribeFollow Nick and Casey Rossolillo at Chip Stock Investor: https://chipstockinvestor.comRocket Lab Deep Dive videos mentionedPart 1 https://youtu.be/AMDd0-JKUH0 (Deep Dive)Part 2: https://youtu.be/Z76xTGFNwBA (Valuation)Companies MentionedPublicly Traded:Cerebras Systems (NASDAQ:CBRS)Rocket Lab (NASDAQ:RKLB)Netflix (NASDAQ:NFLX)NVIDIA (NASDAQ:NVDA)Advanced Micro Devices (NASDAQ:AMD)Broadcom (NASDAQ:AVGO)Micron Technology (NASDAQ:MU)Taiwan Semiconductor Manufacturing (NYSE:TSM)Amazon (NASDAQ:AMZN)Alphabet (NASDAQ:GOOGL)Meta Platforms (NASDAQ:META)Iridium Communications (NASDAQ:IRDM)SanDisk (NASDAQ:SNDK)Kioxia Holdings (TSE:285A)Globalstar (NASDAQ:GSAT)SpaceX (NASDAQ: SPCX)Private / Pre-IPO:OpenAIAnthropicVideos Mentioned:https://www.youtube.com/watch?v=Z76xTGFNwBA&t=3shttps://www.youtube.com/watch?v=AMDd0-JKUH0&t=987sHere's the shifted chapter list, with all timestamps moved back 55 seconds:0:00 Welcome to 7investing Live0:54 Cerebras Systems: IPO recap & the Wafer-Scale Engine2:31 Is NVIDIA even the right comparison for Cerebras?5:38 The memory wall: why bigger AI models need new chips8:52 Latency vs. throughput — and the new AI alliances10:46 Are the $10B OpenAI & $20B Amazon orders real?14:02 Cerebras risks: how do you value a hot IPO?17:27 The TSMC capacity bottleneck20:01 Heather's take on Cerebras20:41 Rocket Lab: the sell-off & Iridium acquisition24:34 Simon's DCF valuation & price target for RKLB29:05 Why Neutron changes everything30:12 Q&A: Does Peter Beck carry an "Elon premium"?31:36 Netflix: buying opportunity or cheap for a reason?36:57 Q&A: Is Netflix a growth stock or a value stock?39:03 Chip stocks selling off: normal volatility or a warning?42:57 Wrap-up & final thoughts#7investing #Simonerickson #Cerebras #CBRS #NVIDIA #AIinvesting #semiconductors #chipstocks #RocketLab #RKLB #Netflix #NFLX #AIinference #stocks #investing #stockmarket #TSMC #AIdatacenters

Ultimate Guide to Partnering™
304 – Building Successful Multi-Product Solutions with Hyperscalers and GSI’s

Ultimate Guide to Partnering™

Play Episode Listen Later Jul 19, 2026 47:12


Don’t Fade and Die in AI Subscribe to our Newsletter: https://theultimatepartner.com/ebook-subscribe/ Check Out UPX: https://theultimatepartner.com/experience/ Matt Yanchyshyn, VP AWS Marketplace, Rekha Thangelapalita, Elastic GSI Leaders; Allison McFadden, Accenture AWS Leader; and James Kang of Nvidia join Ultimate Partner. In this panel discussion, leaders from Elastic, Accenture, Nvidia, and AWS dissect the urgent shifts in the ecosystem, emphasizing that partners must adapt to AI and agentic co-selling or risk fading away completely. The conversation explores the necessity of deep co-engineering, the power of multi-product solutions in the AWS marketplace, and how automated agents are now replacing traditional human sales pipeline progression. By embracing data readiness and strategic collaboration, organizations can survive the “token maxing” era, effectively scale their enterprise opportunities, and align with NVIDIA’s five-layer strategy to dominate the new cloud landscape. https://youtu.be/zUkL4Wqsa68 Key Takeaways AI agents will automate the majority of AWS partner co-selling attachments and opportunity progressions this year. Partners who fail to embrace agentic workflows and automated governance face the existential risk of fading into obsolescence. Successful multi-product offerings require a “blood to all organs” approach that benefits the client, the ISV, the GSI, and the hyperscaler simultaneously. Nvidia’s “five-layer cake” model emphasizes that successful outcomes at the application layer automatically drive growth for all underlying infrastructure. The “token maxing” phenomenon is forcing enterprises to seek cost-effective, open-model alternatives to scale their generative AI securely. Integrating GSIs and ISVs on the AWS marketplace significantly increases enterprise deal sizes and long-term customer renewal rates. If you're ready to lead through change, elevate your business, and achieve extraordinary outcomes through the power of partnership—this is your community. At Ultimate Partner® we want leaders like you to join us in the Ultimate Partner Experience – where transformation begins. Key Tags strategic collaboration agreement, data readiness engine, agentic co-sell, semantic layer, token maxing, five layer cake, accelerated computing platform, open models, cloud consumption, multi-product solutions, partner central agents, propensity data, automated opportunity progression, generative AI governance Transcript Matt Y and Panel Audio Podcast [00:00:00] Vince Menzione: You have a choice. You can embrace them and figure it out and get governance and, and make your data available. Um, use the partner, central agent, move to Agen Co-sell, or you can fade and die. [00:00:11] Vince Menzione: You can feel it happening. The ecosystem is shifting beneath us, the way Hyperscalers are partnering, how AI is remaking the channel and what it means to win in 2026. [00:00:22] Vince Menzione: Welcome to the Ultimate Partner Podcast. I’m Vince Menzi. Own your host. And each week I sit down with leaders at the intersection of technology, partnerships and outcomes. The voices shaping how ecosystems actually work. We talk about what’s real, what’s changing, and what it takes to lead in this era where the partner channel isn’t just part of the strategy. [00:00:44] Vince Menzione: It is the strategy because [00:00:46] Vince Menzione: being in the room changes everything. Let’s start. [00:00:51] Vince Menzione: We’ve got some amazing leaders joining us. So I think probably for a little bit of context, maybe just start with Rika. You can introduce yourself, your role and, uh, what, what you’ve been doing at Elastic. Yeah. [00:01:03] Rekha Thangellapalli: Yeah, sounds great. [00:01:04] Rekha Thangellapalli: Hi everyone. I’m Reka and I lead GSI Alliances at Elastic. Um, for the past 14 years, I’ve had the pleasure of building different kinds of partner ecosystems across companies such as SAP. MuleSoft, Salesforce, Coupa, and now Elastic. Um, I wanna thank Ultimate partner and Vince for having us here today. Thank you and the panel of these incredible speakers for joining me on stage. [00:01:31] Rekha Thangellapalli: Um, very excited for the conversation today. [00:01:33] Vince Menzione: We love Elastic, and you’ve had some of your other leaders on stage at other events. As such, the quality of your leadership team is amazing. Thank you. [00:01:42] Rekha Thangellapalli: I wholeheartedly agree. [00:01:45] Allison McFadden: Excellent. Um, hello everyone. Allison McFadden. I lead our North America AWS practice at Accenture. [00:01:52] Allison McFadden: Uh, I’ve been there for five years, and truth be told, it was my first partnership role, my first formal partnership role. Uh, so I can take some tips from all of you in the room here today. Prior to that, I was 21 years with IBM, and I got into partnerships because my last role at IBM was actually trying to build. [00:02:14] Allison McFadden: Linux business on the mainframe, and I had to have partners. I had to have partners to help me with workloads to run there. So I kind of learned, uh, trial by fire. But I’m excited for the conversation today. Excited to be in this room and excited to talk about what we’re doing with, uh, elastic. Thank you. [00:02:34] James Kang: Uh, my name is James Kang. Nice to see and meet everyone here. Vince, thank you for the opportunity. Thank you [00:02:38] Vince Menzione: for being here. [00:02:39] James Kang: Um, I’m with Nvidia, so I help manage the AWS partnership at Nvidia all up. Um, I guess fun fact, I’m former AWS and so I see a lot of very familiar faces here in the front row. Uh, former colleagues and then current friends. [00:02:56] James Kang: And so, uh, looking forward to the conversation. [00:02:59] Vince Menzione: Great. Well, we’ll start with an easy tia. Matt. This is not directed to you, directed to the others. So what does a successful AWS partnership look like from your C? So we’ll start with Eureka. [00:03:09] Rekha Thangellapalli: Sure. So from an ISV perspective, I think we really are looking at three things. [00:03:15] Rekha Thangellapalli: Uh, mutual investment building together. And scaling together. So when we talk about mutual investment, elastic recently signed a five-year SCA or strategic collaboration agreement with AWS. And while that is a significant milestone in our partnership, for us, what matters more is what it represents, and that is really a long-term commitment from both companies. [00:03:39] Rekha Thangellapalli: Towards product engineering, um, and joint go to market initiatives to deliver value to customers over time. And that’s what we see is that the best partnerships really compound and they build upon each other every year. Um, they don’t necessarily kind of reset every year. Um, next we talk about building together. [00:03:59] Rekha Thangellapalli: So, um. When we talk about joint solutions, we want to deliver solutions that are better together and the customers have to see us that way. And so whether it’s search, observability, or security, we’re looking at taking to market solutions that we can’t or necessarily don’t wanna take on our own. And finally we talk about scaling together. [00:04:22] Rekha Thangellapalli: And this is where marketplace, for instance, plays a big role, um, when customers can draw down on their cloud commitments, transact online and go from, you know, pilot to enterprise scale adoption in hours, not days. Um, this is when really everyone wins. Um, and this is also where partners like Accenture play a critical role. [00:04:47] Rekha Thangellapalli: Um, you know, the incredible amount of expertise that they bring, uh, the managed services capabilities and, um, their data assets actually play a huge role in having our customers realize that value faster. And, um, like Vince mentioned, at the end of the day, best partnerships are all all about creating kind of that. [00:05:07] Rekha Thangellapalli: Self-sustaining flywheel. And so it starts with investing together, building something unique, and having the customers realize that success faster because that success is really the only thing that’s gonna keep that flywheel going for everyone involved. I [00:05:26] Vince Menzione: absolutely. [00:05:26] Allison McFadden: Okay, amazing. I’m gonna riff off a few things Ika said, but from a GSI perspective. [00:05:32] Allison McFadden: A relationship with a WSA successful relationship with AWS looks slightly different. Um, so I think the first thing that we think of in the GSI Community common thread is that the client outcome and delivering value for clients is what we, what we’re striving for. Um, and so the partnership with AWS in that case, um, um, it has to, it has to. [00:06:01] Allison McFadden: Look like one team in front of our clients. So we have to show up indistinguishable, and that’s with AWS and with an ISV partner, it has to look like one solution in front of the client, especially moments that matter. So board meetings, um, you know, the time we’re gonna sign a deal, like we have to look like one team, uh, and keep our our client outcome, um, first and foremost in mind. [00:06:24] Allison McFadden: The second thing, and this is I think where the magic of all the people in this room comes into play. We can have as many discussions at a CEO level as we want. And if our client teams on the ground are not working together, it falls apart. Falls apart directly in front of the client. Yes. And that is a really hard thing to do. [00:06:45] Allison McFadden: So I’m passionate about the alliance work because that that work is what makes it happen at the corporate level. [00:06:53] James Kang: Cool. Um. I’ll start here. So in Nvidia is a accelerated computing platform company. Um, if you asked. Anyone on the, on the street about a year ago, what is ai? A lot of times they would say AI is, is open ai, or it’s philanthropic. [00:07:12] James Kang: Um, Jensen and I’ll, I’ll reference Jensen a lot today, um, because he is our leader, um, but he also sets the strategy in the direction for Nvidia. He talks a lot about AI in the metaphor of a five layer cake. And in terms of the five layer cake, you start off with the foundational bottom layer being power and energy, which sustains. [00:07:32] James Kang: All of our data centers, you move up the stack in terms of chips. So things think of Foxconn, think of TSMC. Next you have the infrastructure layer. So obvious choice is AWS, and then you get to the models where you do have the philanthropics and the open ais. But finally in at the precipice, you have the application layer. [00:07:53] James Kang: Ultimately, the reason why I mentioned all different stacks of the layers, the five layer cake, is the fact that the application layer is the most important. And so when you think about. Partners like Elastic or ServiceNow Trend, ai, CrowdStrike. Every time you pull from the application layer and you see a success, it pulls all five different components of that layer up. [00:08:13] James Kang: And so ultimately, as I think about success, it’s it’s being able to develop these co-sell wins at the application layer and really demonstrating that through extreme co-engineering and co-design with all the different application. Infrastructure, power and energy layers in mind. Um, Jensen also likes to think of himself not only as the CEO and founder, but also as the, the chief Marketing Officer. [00:08:35] James Kang: We are a very event driven company, and so at our big events like GTC or at big industry events like CES or Computex, he likes to show up on the biggest stage, biggest stages and showcase the partnerships with not only ISVs and GSIs, but also with end customers. And so that’s what I think about when I think of SA success. [00:08:56] Vince Menzione: That’s a really good point. You talked about, Allison, you talked about having an alliance strategy, or at least you teed it up, so I thought maybe we would go there for a second. Right? Like, what does a great alliance strategy look like and why is it important to the success of the partnership? [00:09:11] Allison McFadden: Man, I, uh, I have so many opinions on this. [00:09:13] Allison McFadden: We could probably be up here all day. That’s [00:09:15] Vince Menzione: okay. [00:09:16] Allison McFadden: Um, no, I think. Uh, there, there are a couple things, and the first one that comes to mind is focus. We cannot be all things to all people. Um, so when it comes to think about some of the, the work we’re doing with Elastic, we have a very, very clear point of view on what client problem we’re solving, what clients we want to talk to. [00:09:38] Allison McFadden: It helps if, um, from an ISV perspective, if there’s a very clear fit in. The Accenture portfolio or whatever, you know, SI consulting partner. You’re working with a very clear fit in the portfolio and we know what we’re not gonna go after, what we’re not gonna spend our time on because we have, we have this tendency, there’s millions of people. [00:10:00] Allison McFadden: The ecosystem chart that, you know, Vince, you showed up there, there’s so many connections. There’s probably more connections there than there are atoms in the universe, right? So, um. Defining what we do together and what we don’t do together is the first thing that pops to my mind. [00:10:19] Vince Menzione: Reka, do you have a perspective on it since we’re gonna, we’re gonna talk next about what you’ve done together, but, and I also wanna get mass perspective as a hyperscaler partner here as well. [00:10:29] Rekha Thangellapalli: Yeah, I mean from my perspective, I, I’m gonna, you know, kinda echo what Allison said is to be just maniacally focused. Yep. Um, because, especially from my perspective, so Elastic has three different solutions, right? We’ve got search, we’ve got observability, we’ve got security that map to completely different business units within Accenture. [00:10:47] Rekha Thangellapalli: And of course Accenture does a lot of things. And so, you know, when we first came together it was like. Okay, what are we gonna focus on? What industries are we gonna go after? Which segments are we gonna go after? Which customers, you know, um, outcomes are we trying to solve? And I think that sort of maniacal focus is the number one contributing factor to, to the fact that I’m like, up here on stage today. [00:11:12] Rekha Thangellapalli: Great. [00:11:14] Vince Menzione: Matt? Perspective? [00:11:16] Matt Yanchyshyn: Yeah, I, I, I guess I was trying to. To add something, uh, additional from an AWS perspective, uh, when it comes to, you know, what does a great alliance look like? Uh, AWS is obsessed with data, you know, in data we trust. And, and so the best, um, and, and this goes sales business problem, and it’s not just the engineering teams. [00:11:34] Matt Yanchyshyn: And so, uh, you know, Accenture does a good job of this elastic, definitely. And if you can come to the table with, um, quantifiable proof of the value of customer outcomes and partnerships. Um, you’ll win all the time and it’ll be a durable relationship with AWS ’cause we really are this data obsessed company and, and even the most senior sales leaders. [00:11:54] Matt Yanchyshyn: Uh, and so what I mean by that specifically is like if you, if you can show like your a RR to land an a RR conversion ratio, like in in numerical format, it’ll light up our sales leaders and, and they’ll be all, and they will co-sell with you all day long. If you can show the, I mentioned this earlier, like the AWS service, uh, whether you’re consulting company or, um, elastic and, and how the shape of customer accounts change positively when we work together. [00:12:15] Matt Yanchyshyn: That type of sort of quantifiable data works particularly well from an alliance perspective. With AWS as a partner, we, we really are like this data in sort of results out company. Um, so I, yeah, that’s just adding to the great points that were already made. I would say specific to AWS that that’s key. [00:12:30] Matt Yanchyshyn: Yeah. And I’m gonna bring up one more thing. I want to dive in on the, the joint value proposition, but you mentioned something that made a lot of sense and resonated to me about the organizations once you get out of partner, the partner world that we all know and love. Mm-hmm. Once you get down into a field organization or account management organization. [00:12:49] Matt Yanchyshyn: Not as much understanding and really organizations do a bad job here, honestly, in terms of enabling the field organizations. Do you agree? [00:12:58] Allison McFadden: I agree because I, I agree. And, um, you know, I think that’s one of the things, and, and I, I, when I joined Accenture, what we had was a lot of wicked smart architects delivering programs to clients in the field. [00:13:15] Allison McFadden: Very smart, very deep in AWS knowledge. Um, and that was awesome for the 10 clients they were staffed on and to get that understanding of how AWS works and I dream about lar, right? Like, this is a good, you know, but that takes real effort and real work. Yeah. And it’s, it’s um, almost like being a language translator. [00:13:37] Allison McFadden: Yes. For me. Yeah. So, you know, I had to deeply learn AWS so that I could. [00:13:42] Rekha Thangellapalli: Sure. [00:13:42] Allison McFadden: Teach my account teams. My account teams are really smart. They know who they’re selling to. They know their customers. They know what their customers need. They do not know what AWS has to offer always because they’ve got 20 partners lining up to try to tell their stories. [00:13:57] Allison McFadden: Um, they don’t know how to ask of the AWS team or the elastic team or the Nvidia team. Yeah. What they need [00:14:02] Vince Menzione: this co-selling piece. Yeah. [00:14:04] Allison McFadden: And so that is where, um. We had to build that muscle even around our AWS practice, which was a huge practice at Accenture, but we didn’t necessarily surround it with that kind of enablement and um, almost deal coaching layer. [00:14:21] Vince Menzione: So Elastic and Accenture came together. I dunno which one of you wants to lead this part of the conversation, but you will, right? Yeah. So tell us about the genesis of this and why. And a lot of people dunno what Elastic does, but you do some really incredible work. Like I, somebody told me one day was like, oh, you know, Uber, like, that’s elastic, powering all that. [00:14:41] Vince Menzione: Like, we don’t think about that. That the engines that you have and the, the backend to the customers, huge customers. [00:14:48] Rekha Thangellapalli: Yeah, absolutely. Um, so when AWS launched this feature last, um, reinvent where basically it allowed, you know, channel partners such as Accenture to be able to bundle up their services, their data assets with an ISV solution and put it on marketplace, um, you know, Accenture and Elastic immediately saw an opportunity. [00:15:09] Rekha Thangellapalli: Um, at the time most customers were doing gen ai. But they were running into the same challenge, which was that their data just was not ready. And by the way, this is a problem we were solving. Outside of marketplace. I think the, the feature that you guys launched just gave us a way to package it up and to be able to create this repeatable solution, which we call data readiness engine for gen ai and put it on marketplace. [00:15:40] Rekha Thangellapalli: And, um, this to me was a success because. Each company had a clear reason to invest. Um, so for Accenture, they were able to, you know, create a very differentiated services led offering. Uh, for Elastic, we were able to expand on our AI story. And for AWS, um, you know, it drives marketplace adoption, increases cloud consumption, all of that great stuff. [00:16:07] Rekha Thangellapalli: And customers, of course get. A solution to a very real problem that, that they were having. Um, and you know, the surprising part for me going through that journey was that, um. The pitching, the idea, getting the budget, getting the executive sponsorship was actually the easy part. The hard part was getting all three companies to come together, uh, to go from idea to launch in a very ambitious timeline of six weeks. [00:16:37] Rekha Thangellapalli: Nice. And so, you know, this was very much like. Doesn’t matter your title. We’re rolling up our sleeves and we are on this outcome together. Um, and so we literally built a RACI matrix, a project plan, and you know, we had daily standup calls for six weeks where literally. At least one person from each three of these companies called in, you know, got rid of any blockers and we made sure we were on target for that timeline. [00:17:07] Rekha Thangellapalli: Um, and you know, at the end we had a successful launch. But I think my favorite part about the story is the impact that we’re having and, um. My favorite story comes from a global pharmaceutical company that, you know, had basically nine petabytes of data spread across six different continents. Wow. And by working with Accenture and Elastic, they were able to build that trusted foundation that their AI and their agents can, you know, kind of safely tap into and be accessible at scale. [00:17:41] Rekha Thangellapalli: Um, so that’s my version. Allison. [00:17:44] Allison McFadden: Yeah. Well, I don’t have a lot to add. I just, I would say this is a good example of a couple of principles, right? One is having a forcing function is never a bad idea. Sign up for a big event, sign up. I’m like, I’m here with my, you know, Nvidia guys saying, sign up for the event. [00:17:58] Allison McFadden: It’ll make you move quick, right? [00:18:00] Audience Member: Yes. [00:18:00] Allison McFadden: Um, so that is one, but two, one of my mentors once told me, when you’re designing any kind of, you know, offering go to market motion, it has to get blood to all organs. If it does not get blood to all organs, it does not go [00:18:14] Vince Menzione: nice. [00:18:14] Allison McFadden: Um, [00:18:14] Vince Menzione: I love that analogy. [00:18:15] Allison McFadden: Oh, I love it. And I can talk all day. [00:18:17] Allison McFadden: That guy was brilliant. I love him. But, um, no, and, and so Elastic did a really nice job of bringing the tech to the table. Um, our team has to trust in that technology and its ability to scale, right? Um, because at Accenture we have to be able to deploy across 700,000 consultants. Um. And yeah, so I think those are the two, two things that really worked well here is we had, uh, trust in the technology solved a customer need. [00:18:50] Allison McFadden: Um, it drives, we don’t even talk about, like, yes, it drives marketplace revenue, but it unlocks work that we do that drives even more revenue to our AWS Friends. Right. So this is a, this is a, um, product that’s getting your data ready for AG agentic. It’s a messy problem that everyone’s dealing with, and it removes blockers for clients and it unlocks more, you know, ag agentic work on top of that. [00:19:15] Allison McFadden: So, blood to all organs. [00:19:17] Vince Menzione: So, was that the proposal going forward to say we need to have, we need to have trust in the solution. We need to drive significant revenue. It needs to be something all of our, you know, seven, 700,000 people. Can be a part of and help drive? Is that how you think about? [00:19:32] Allison McFadden: Yeah, and for us right now, um, it’s an interesting time for Accenture. [00:19:36] Allison McFadden: Our clients are asking a lot of us, and what it does is it having some of these accelerators helps us deliver cheaper, better, faster to our clients, which is what they’re demanding of us right now. Um, so it’s an accelerator to client outcomes. [00:19:55] Vince Menzione: James, what is NVIDIA’s role and how do, how do you enter the equation here? [00:20:00] James Kang: Yeah, it’s, um, it’s a good question. Um, I, I would say that Nvidia is probably one of the most misunderstood organizations in the world. Um, despite the, uh, the market capitalization in the valuation of the company, we have a very tiny organization. Um, what I mean by that is, um, if you think about. [00:20:20] James Kang: Salesforces and field sales organizations. Um, we’ll take Salesforce as the account or the customer. As an example, we have one account manager at NVIDIA that no, not only covers and is responsible for the relationship with Salesforce, um, but also manages. Automation Anywhere as well as DocuSign. Whereas at AWS, in contrast, like there are full armies and teams Yeah. [00:20:45] James Kang: That are supporting the Salesforce relationship. And so as you think about partnering and working with Nvidia, the focus has to be on really. Extreme co-design, but also being very prescriptive in terms of what are the very specific customer outcomes that we are solving for. And the guidance that I would give is bring in Nvidia into that equation and that conversation as early as possible because that [00:21:10] James Kang: co-engineering and co-design needs to be part of the foundational building blocks in order for you to come out with a end solution that checks all those different requirements. [00:21:20] James Kang: And so I think. Again, like going back to Nvidia, um, we like to talk about two different types of brains. A brain one and a brain two. Uh, brain One you think about the next quarter and making sure that you’re hitting the revenue targets for the next quarter. Brain two, you think about a long-term goals and potentials looking around corners and being very strategic. [00:21:41] James Kang: The saying internally is without Brain one, there is no oxygen, but without brain two, there is no future. And everyone at NVIDIA is trained to think in that brain two mentality. [00:21:52] Vince Menzione: Wow, Matt. [00:21:54] Matt Yanchyshyn: Yeah, I, I was just thinking I love the blood doll organs. Uh, and so just on, on that note, um, and, and, you know, the multi-product solutions that, that you, you built together, uh, that is a really good example of blood do organs because like we all know, that’s how customers buy. [00:22:07] Matt Yanchyshyn: They, they buy solutions and increasingly they’re looking for combinations of ISV, sometimes multiple products from multiple ISVs with services. Uh, often they’re buying it through a resell motion. You know, and they, and, and so that from a customer perspective, they want a single place to go. And so that’s the multi-product solution. [00:22:24] Matt Yanchyshyn: They wanna find everything they need, they need Accenture, they need Elastic to solve a specific solution. And I think where that’s headed is even more specific listings, like with AI powered listing experience, like, you know, elastic Plus Accenture for, I’ll make something up like a manufacturing workload. [00:22:37] Matt Yanchyshyn: And so this solution based. Uh, sort of buying is, is very customer centric. It’s what customers want. We all know that. But that’s, that’s the customer sort of organ, I guess. Um, but then, you know, you all have SCAs and those SCAs have marketplace commits. It helps if that gets transacted through marketplace helps the AWS relationship, you know that that’s an organ. [00:22:55] Matt Yanchyshyn: It’s the relationship. It’s, it’s the commercial construct and that you have, uh, that that’s another organ. You’re marketing people. They, that’s another organ. They don’t wanna land, uh, leads on a static marketing page. They wanna land a lead on a, a storefront with a multi-product solution that can actually convert and that you can actually buy it through that. [00:23:12] Matt Yanchyshyn: So the marketing person’s happy because they, they have less churn. Uh, and then, you know, our reps are happy ’cause guess how they get paid? They retire quota when they sell Marketplace. And they, we also, Jay McMain will tell you, that’s another organ called Jay or on, on you now. Um, [00:23:27] Matt Yanchyshyn: he’ll like that. I’ll call him up and tell him that. [00:23:29] Matt Yanchyshyn: Yeah, [00:23:30] Matt Yanchyshyn: but he, he’ll tell you, you know, don’t believe me. Obviously, never believe Matt, believe, believe the, the data and, and his data shows that. Those deals will close faster and larger if you use marketplace. So that’s, that’s a lot of organs. That’s the whole body. Um, but you know, when you have your customer happy ’cause that’s how they wanna buy your field happy. [00:23:45] Matt Yanchyshyn: Um, and, you know, the relationship happy and you know, your marketing team happy. Uh, and, and Jay happy. Um, and, and you know, I think that multi-product construct and, and the way you kind of use it to model a partnership and the way buyers ultimately wanna buy is, is really powerful. And so I, I think it’s, you know, it’s really a manifestation of how. [00:24:04] Matt Yanchyshyn: We kind of intend and to go to market anyway. Uh, so I think, you know, and thanks for leading the way, by the way. You’re, you’re amongst the very first, so that’s great to see. [00:24:11] Matt Yanchyshyn: So these storefronts are really helping this drive, drive this. Well, [00:24:13] Matt Yanchyshyn: that’s the next evolution. Like we’re talking about the multiproduct solution. [00:24:16] Allison McFadden: I’m JJ Accenture storefront. [00:24:17] Vince Menzione: Yeah. Oh, there you go. I mean, j and j Accenture storefront. [00:24:20] Allison McFadden: We’re gonna talk about that. [00:24:20] Matt Yanchyshyn: Yeah. I mean, [00:24:21] Matt Yanchyshyn: Accenture also leading the way yet again with storefronts. And so I think the combination of. You know, again, I was talking a lot about conversion. Yeah. And you know, buyers know sometimes they know what they wanna buy and, but if you really wanna convert that lead, you wanna land them again, something that combines, you know, elastic Accenture’s services plus software, but in a storefront that is, you know, surrounding with just the solutions they want so they don’t need to kind of go searching. [00:24:42] Matt Yanchyshyn: So, you know, ultimately reducing that time to close, I guess, really ’cause meeting the customer where they are with what they need. [00:24:51] Matt Yanchyshyn: So we talk about co-selling a little bit. We, Jay and I talk about this all the time. We gotta keep looping Jay in here, even though he is not even in town this week, but Reko, um, what does co-sell look like inside Elastic? [00:25:02] Matt Yanchyshyn: You’ve got, we talked about an incredible leadership team. I’ve gotten meet some of your leaders. Seems like you drive, you do a good job internally driving that. Let’s talk a little bit about it. [00:25:11] Rekha Thangellapalli: Yeah, and this is something I’m, I’m personally very passionate about. Um, co-sell is. Very much a journey, not a destination. [00:25:20] Rekha Thangellapalli: And I think step one for us is recognizing the different partner types that we have. Because at Elastic we work with, you know, OEMs, MSPs, resale distributors, GSIs, um, and they all bring something very unique. To the customer lifecycle and they all contribute very differently within, you know, our own sales cycle and sales process. [00:25:45] Rekha Thangellapalli: And so, you know, figuring out what is the unique benefit they bring, how do we enable them? So training and enablement is a huge piece of it, and so is making sure we’ve got the right metrics to measure success. Um, I know a lot of companies look at partner sourced as the north star, and that’s great, right? [00:26:06] Rekha Thangellapalli: Because that is undeniable. You can say, Hey, that would not exist if it wasn’t for my partner team. Um, but we’ve also noticed that when we bring in GSIs, it actually increases renewal rates. It significantly increases. Um, a RR over time. Um, it expands deal sizes and so these are very real metrics that we can point to, um, beyond just the co-sell and the partner sourced number. [00:26:32] Rekha Thangellapalli: Um, so for us it’s looking at it from a very holistic perspective, but also catering it towards that unique partner and making sure we’re doing everything we can to set them up for success and setting up the partnership for success. [00:26:47] Vince Menzione: So clo close win ratios, deal size and renewal rates? [00:26:52] Rekha Thangellapalli: Yes. For specifically for geos size. [00:26:54] Rekha Thangellapalli: Yeah. [00:26:55] Vince Menzione: Very interesting. Allison, uh, what had to change internally to produce these co-selling? We talked a little bit about the field organization and enabling a, a group of, and, you know, account sellers that are very customer focused and enabling them on the co-sell side. What had to change internally to drive that? [00:27:13] Vince Menzione: Yeah. [00:27:14] Allison McFadden: I, I might have already alluded to this a little bit in a previous answer, but, um, creating the capacity to develop, build, and sell these solutions, um, inside of a large GSI, where billable hours is kind of the number one metric on the table. Um. Is part of the investment that we had to make within Accenture to get this done? [00:27:36] Audience Member: Yeah, [00:27:36] Allison McFadden: so expert technology time. So we have technologists that understand the elastic technology. We do similar with Nvidia, by the way, we. We released some of their time to go co-develop the solution because it has to hold technical water, right? It can’t just be a marketing pitch. It can’t just be, it has to be a real, um, what’s the there, there. [00:27:59] Allison McFadden: So in order to actually do proper co-sell, we had to release some of that time. Um, to invest in those partnerships. Um, we’ve also done similar with some industry aligned business development leaders recently, so we have freed their time up to go. Uh. Open new conversations, educate client, account teams, go to clients, have conversations. [00:28:26] Allison McFadden: Um, so that, that’s a new motion that we, uh, have just kind of recently made, um, to allow them, I love this brain one, brain two also, right? So to allow them to focus on brain two, because a lot of our time. Typically spent delivery issues, you know, getting my hours, where am I charging my time? And so just freeing up a little of that capacity to do this work, um, helps get us in this brain two mode where we’re not just living to survive. [00:28:56] Vince Menzione: I. So, Matt, you’ve removed a lot. I mean, one of the things I admire, I admire AWS for being first to market and removing the most friction in marketplace of any of the vendors. Really, truly that. You talked about some of the announcements. How does some of, how does some of this tie PC central agents propensity sales plays, MCP, how does some of this tie to how, how you’re thinking about the future? [00:29:18] Vince Menzione: And how to enable more motions like this. [00:29:20] Matt Yanchyshyn: Yeah. Well, I, I think if you know my boss, UBA Borno, uh, you’ll know that she has a maniacal focus on automation. Yeah. Um, and, uh, co-sell is increasingly automated. You know, you were asking earlier about propensity data. You can get that propensity data in addition to sales plays and, uh, opportunity scores through the partner central agents. [00:29:38] Matt Yanchyshyn: So things that used to require multiple calls to A PDM, if you’re lucky to have one. Yeah. Or a p sm. Uh, you, you can now get through, through these agents, you know, uh, tech Systems, TGS, they, they manage what, over 5,500 customer opportunities with agents that they built on top of our partner Central APIs. [00:29:55] Matt Yanchyshyn: Um, and work Span has built a whole product and business that’s right on leveraging, uh, our APIs, our capabilities to sort of tie into your CRM. So, majority of all opportunities will be progressed and managed by agents. This year at AWS, we already have a majority of all customer opportunities, all app have a partner attached and I, I took a personal goal for a majority of those partner attachments, not to happen from a human. [00:30:22] Matt Yanchyshyn: But from our solution matching engine. And how do you get recommended by that solution? Matching engine, having a healthy ACE pipeline, thanks to partner central agents and the integrations you’re doing. And in addition to being the specializations and doing things like multi-product solutions and ultimately closing opportunities, you dream of LAR and so LAR will help that. [00:30:40] Allison McFadden: It’s more like a nightmare. [00:30:41] Vince Menzione: And so, you know, [00:30:42] Allison McFadden: it’s more like a nightmare, but [00:30:44] Vince Menzione: nightmare. Well, it’s, it’s, yeah. Nightmare of Laura and, and. Nice dreams of PRM, but the, um, but that’s the loop, right? I, I think, uh, increasingly co-sell for us, and in my mind, is largely a hundred percent automated. Yeah. Except for what matters most, those most largest, most strategic, most complex deals. [00:31:01] Vince Menzione: Where our highly paid and very skilled salespeople are most effectively used. [00:31:05] Vince Menzione: Yeah. [00:31:05] Vince Menzione: You know, the days of, you know, this person with 20 years experience selling, clicking, progressing opportunities through a pipeline, uh, should be over. Uh, and, and we need those people out, out selling and, and co-selling. And so that for me. [00:31:19] Vince Menzione: Yeah. That, you know, we talk a lot about co-sell, but I, I’m obsessed with automating as much of the co-sell as possible. [00:31:24] Vince Menzione: I remember going back to the ex Excel spreadsheets and, and that, that seems to be be Viva became spreadsheet jockeys. [00:31:31] Vince Menzione: Yeah. [00:31:32] Vince Menzione: And, and they stopped selling. They forgot how to sell. [00:31:34] Vince Menzione: Yeah. And people spend all this time doing lunch and learns and things like that. [00:31:36] Vince Menzione: And then, you know. Then the salespeople rotate out after 18 months and, and it, that’s, that’s the old days. Uh, you know, the new days are, are AI powered matching algorithms, uh, ag agentic co-sell, using the partner essential agents to get your data and, and putting that data to use automatically and, and what sounded like magic. [00:31:51] Vince Menzione: 12 months ago is being done, you know, by partners at massive scale across thousands of opportunities. You can do it today. And you know, I, there’s a guy named another Mike, right? Mike another Mike who they have, there’s like a guy who’s doing all this and I’m picking on Mike ’cause I, I know their system really well and I know the guy Mike grew easily built it for them. [00:32:08] Vince Menzione: Um, but, you know, I think, yeah, again, in the days of having 10 people sort of doing lunch and learn could be replaced by one or two people, building agents, uh, managing a massive pipeline. And, and that’s the future. [00:32:18] Vince Menzione: Exactly. James, your perspective on what breaks with co-selling? [00:32:22] James Kang: Oh, what breaks co-sell? Um, I would say. [00:32:25] James Kang: It, it starts and finishes with just misalignment and a loss of trust with the customer, especially when you have multiple partners or stakeholders involved. If you’re trying to do a three-way deal with a end customer and you’re not on the same page, you’re not gonna get to a successful outcome on, on the backend. [00:32:44] James Kang: Uh, the fix is a much more complicated story. I would say that to take a step back, um. We’ve talked about the five layer cake. We’ve talked about where NVIDIA kind of fits within the equation. We are invested in the ecosystem and so as different players and application organizations win and see these outcomes for end customers, we celebrate that success. [00:33:07] James Kang: Um, and as part of that kind of ethos of where NVIDIA fits within the ecosystem, we wanna make sure that not only. Our customers, but our partners like ISVs and GSIs are set up for success. Um, we do not as Nvidia sell hardware or GPUs directly to customers We use. Hyperscalers like AWS as kind of our force multiplier. [00:33:31] James Kang: And similarly we think of ISVs and GSIs as the force multipliers in terms of our extensions of how we, we kind of leverage the relationships and build the trust with our end customers. And so going back to kind of the question, Vince, I would say that it all comes back to trust and being able to build that mutual trust. [00:33:48] James Kang: Um, a lot of what we do when we co-sell with AWS is really on the software layer. Um, we actually have more software engineers at NVIDIA than we have hardware engineers, which is a weird thing to say, um, because everyone knows us for our GPUs. But because of that fact, we are heavily invested in Cuda and making sure that Cuda becomes the foundational layer for how not only our ISVs and GSIs, but also our end customers are building. [00:34:12] Vince Menzione: Very cool. So Reiki, you and James together on this production. Versus pilot with the Gentech ai. Tell us a little bit more about that. Where, where are you in the process? [00:34:24] Rekha Thangellapalli: Yeah. So I mean, in general, what we’re seeing out in the market in, in relation to sort of AI and, and customer’s journeys is that, um, at least from an elastic perspective, um, we’re seeing people very much in production when it comes to, you know, kind of AI assistant co-pilot use cases. [00:34:42] Rekha Thangellapalli: So, you know, things like, um, software development, customer support is a big one. Um, any sort of employee productivity use cases where there’s. Still a human in the loop somewhere. Um, and there’s a very like, clear path to value. And so we see the customers being in production excelling there. Um, no problem. [00:35:01] Rekha Thangellapalli: Where we’re seeing people still kind of in the pilot phase is those fully autonomous workflows where there is no human involved. The agent is reasoning on its own. Um, accessing multiple systems and taking an action on the user’s behalf. And what we’re seeing is that it’s not the intelligence of the agent that’s holding it back. [00:35:26] Rekha Thangellapalli: It’s more about giving the right context to the agent and having the right. Security kind of governance controls in place for the company to feel comfortable in putting these fully autonomous workflows into production. And that’s really the conversation we’re having is all right, what are the controls you need in place? [00:35:47] Rekha Thangellapalli: For you to release this to your business unit. Um, and what is the context that the agent is needed before we can comfortably let the agent make the decision on the user’s behalf? Um, James, I’d be interested to hear what you’re, what you’re seeing in the market [00:36:03] James Kang: plus one on all things context. I, I would even go so far as to say, um. [00:36:09] James Kang: H how many folks in the audience have heard of token maxing? Like this new term? [00:36:13] Rekha Thangellapalli: Yeah. Yeah. [00:36:14] James Kang: Um, I’ll, I’ll give a very specific example of, of Uber that went public. With the example of Claude, like they allowed all of their employees to use as many tokens as possible, and within the span of four months, they exhausted their full budget for the year, and so they had to pull back, and now there’s a cap on every employee. [00:36:33] James Kang: I think the number that’s circulating is $1,500 per month per employee, and so I think that is at least. In this multi-phase evolution of where we’re going to be and where we’re today, cost has become kind of the prohibitive force in terms of agentic AI at scale. Um, I think we are working on some very creative solutions in-house and Nvidia. [00:36:55] James Kang: Um. And we saw some really dynamic announcements this week when it comes to all things agent core, um, where we want to focus on very nimble ways for customers to be able to execute and go to market. And one extreme example of that is our investment within our open model strategy. So Nvidia, not only, again, providing GPUs, we actually offer our own op open models, which we call our Nitron models. [00:37:21] James Kang: And through our Nitron models, we are allowing customers to really develop and fine tune their own proprietary models in a cost effective manner. So right alongside the frontier models like OpenAI and Anthropic. It’s not a if then, it’s not an either or statement. It’s a, it’s a permutation, it’s an and So we’re giving you a cost effective alternative to not only bring your AgTech applications at scale by training on Nibo tron, which is open source, but then once you’ve kind of finished and fine tuned that specific training job to be able to. [00:37:53] James Kang: Go ahead and utilize your frontier models, whether it be OpenAI or Claude. And I know there’s other partners here that are providing those kind of different model capabilities. And so I think for us it’s, it’s a matter of choice. We know that this market is dynamic. It’s gonna be evolving over the next coming months as well as the next coming years. [00:38:10] James Kang: Uh, but we believe that we are positioned for a really unique dynamic expansion of AgTech use cases over the, at least the next three to six months. [00:38:20] Vince Menzione: Allison, for the partners in the room who are glazed over right now going, what do I, what do I do over the next 12 months? [00:38:26] Allison McFadden: Should I wake everybody up by saying, yeah, please. [00:38:27] Allison McFadden: Say go hurricanes. [00:38:28] Vince Menzione: Yes. [00:38:29] Allison McFadden: Is there anyone, anybody? Everyone’s like, boo. I get to leave the parade today to go home to parade. I live in Raleigh, so we’ve got our parade on Saturday. Nice. [00:38:39] Vince Menzione: Nice. [00:38:40] Allison McFadden: All right. Wake up. Um, all right. So for the $50 million partners in the room, um. $50 million is not small. You have something that works. [00:38:50] Allison McFadden: Right. This is great. What I would be thinking about is, you know, we’ve talked about focus before, but really doubling down on, you know, what is, what is your industry, what is your client like, ideal client that you serve. And build, um, almost that kind of community. You know, the, the clients we have move from firm to firm to firm. [00:39:17] Allison McFadden: And if you’ve done good work at one, you’re gonna follow ’em to the next. Um, so build that client demand in a specific place or specific client profile that is just like really knocking it out out of the park for you. Um. Scale with marketplace, right? So if you, I, I love some of the data that you were sharing in your talk earlier, um, because it’s like no overhead scaling mechanism. [00:39:45] Allison McFadden: I mean, it’s, it’s fantastic. Um, Accenture, other GSIs like us, we are investing in marketplace. So we’re investing in resources, um, to help us. Use marketplace more with our clients and we’re gonna capture, right, those storefronts. And if you’re present on marketplace, you’re gonna be able to catch, uh, yourself in that wheel. [00:40:09] Allison McFadden: So I think those are the, the kind of couple of things I would say is focus, focus, focus to drive that client demand and use scaling mechanisms like marketplace to really kind of, uh, accelerate. [00:40:24] Vince Menzione: Matt, anything to add there on the. [00:40:26] Vince Menzione: Well just, you know, Ja, James, you, I love the token maxing reference in Uber and it reminds me, you remember when cloud came out and everyone was like, oh, all these people are, are gonna use the cloud and costs are outta control and. [00:40:39] Vince Menzione: Um, a lot of people pulled back from the cloud and, and a lot of those companies no longer exist. And it’s similar with, with, uh, token maxing, like, oh, these agents are outta control. You have a choice. You can embrace them and figure it out and get governance and, and make your data available. Um, use the partner, central agent, move to agent to co-sell, or you can fade and die. [00:40:58] Vince Menzione: And, and that’s, that’s where we’re at. Uh, is, is the, the companies sitting here today embraced the cloud years ago and won. Uh, and and there’s a set of companies here today who are gonna embrace agents in the, for both buyers and sellers, and will win. And there are those who won’t and they won’t win. And so for me, it’s like we’re, we’re at a, we’re at a crossroads. [00:41:18] Vince Menzione: And, and if you’re gonna win, you gotta leap into that, you know? I love it. And, uh, and, and, and it’s, it means the cost of experimentation is so much lower now. Development and, and even business development or software development is, is agent enabled. And so you can take risks, you can experiment and, and you have to, it’s, it’s an existential moment. [00:41:37] Vince Menzione: Agreed. We’ve got a couple minutes left over for any questions. What do you think? Sure. Are there any here. I think there are a couple. Yeah, we’ve got, we’ve got a co-sell question I’m sure coming up here. [00:41:51] Audience Member: Um, I’m Cassandra, I’m the CEO of Partner Tap. And one of the questions I had was, I think, you know, the co-selling between the sellers is where things get. Really, really hard when you’re multi-partner. And so when I was listening, um, with, you know, the Accenture and Elastic together, you talked about how you had, you, you had to get these BD business development people. [00:42:22] Audience Member: Um, is this a new team that is over the client team? And how do these teams interact like with the elastic sellers? Are you doing a lot of coaching to the field and then with if AWS sellers are, are involved, like what is that whole picture? What does look like, [00:42:43] Allison McFadden: like [00:42:44] Audience Member: on the ground? I mean, that is the hardest part, I think, and that’s what we hear. [00:42:48] Allison McFadden: It’s so, it’s so, it’s so tough. Um, and I will, I’ll just say, so our business development leaders that we now have kind of. Expanded their capacity. They have always been, they have always been there. Um, but they have not been well resourced. They haven’t, they haven’t had very clear kind of job description. [00:43:12] Allison McFadden: I’m gonna say I, in the past they have been kind of focused on partner relationship. And so like more like an alliance manager and maybe working on some of the data. Right? So when I say I have nightmares about Lars, because we’re always trying to increase the LAR for Accenture and, and they were focused like in those detailed weeds of like trying to pass ACE and trying to call the PDM and all this stuff. [00:43:39] Allison McFadden: What we are doing is really pivoting them to be proper sales, business development focused on client outcomes and focused on. Technical skills to be able to describe what this solution is to the field. So, um, and because we need, I have many, many questions about, I gotta get agents to work with Eurogen co-sell so that that part somehow goes away. [00:44:05] Allison McFadden: So that’s a, that’s the thing we gotta solve still, but, um, so we’re pivoting them to be kind of driving. More of that co-sell enablement with the field, um, and taking that message to the field rather than being there, waiting for questions to come in from the field, waiting for like our field teams to discover, oh, I saw something that we’re doing with Elastic, like on a press release on LinkedIn. [00:44:30] Allison McFadden: Right. So we’re kind of trying to pivot them to be more proactive. [00:44:33] Vince Menzione: Very cool. [00:44:34] Rekha Thangellapalli: Yeah. And uh, Cassandra, that’s an excellent question because I think. Multi-party, you know, sort of tri-party offerings. The hardest part is operationalizing it at scale, right? Yeah. And so for this particular offering, we are basically having three routes to market. [00:44:51] Rekha Thangellapalli: So one is seeing how this offering fits into our existing elastic go to market. And so I am constantly enabling our field sellers to say, okay, within our three field sales place, here’s exactly where this fits in. Here are, you know, uh. Keywords that you hear in customer conversations where you bring up this offering and here’s a process of how it works. [00:45:14] Rekha Thangellapalli: Um, exactly At what sales stage do I bring in Accenture, how, you know, what are the roles and expectations? Right? So that’s on the elastic side. We’re doing the same thing on the Accenture side. So we’re doing a ton of training enablement and lunch and learns, and we’re also looking at how do we fit into. [00:45:31] Rekha Thangellapalli: Uh, Accenture’s AI transformation projects, we are the semantic layer, right, of their enterprise brain. And so it’s a whole different sales motion, um, and, you know, having the right assets, having the right process again to make sure that that goes smoothly. And then finally, we’re going directly to the customer. [00:45:49] Rekha Thangellapalli: So we are launching multiple external campaigns where, you know, if the customer raises their hand. We will, we will line up immediately. Right. Um, and so, [00:46:01] Allison McFadden: I mean, I can’t, I can’t, I can’t say how important that third leg of the stool is. ’cause the second part, she talked about getting into our catalog is the first thing. [00:46:09] Allison McFadden: ’cause my BU business development leaders have the catalog. Right. And that’s what they’re selling. So what Elastic has done has gotten into one of those offerings and then. If we have a customer that asks for it, that is the fastest way to alignment. That is like the number one thing that we respond to [00:46:26] Vince Menzione: customer at the center. [00:46:27] Vince Menzione: This is great. Well, I think we’re up to time. This was a great session. I want to thank you. This is what a great, what a great group. [00:46:34] Vince Menzione: Thanks for listening to the Ultimate Partner Podcast. If today’s conversation resonated, share it with a partner leader in your network. Subscribe where [00:46:43] Vince Menzione: you listen, and head over to the ultimate partner.com. [00:46:47] Vince Menzione: For show notes related content and the resources for this episode. And if you haven’t already, now’s the time to register for the Ultimate Partner Live Event in Reston, Virginia, October 26th through October 28th. Until next time, keep showing up in the rooms that matter because being in the room changes everything [00:47:09] I.

That Was The Week
Intelligence: Who Owns it?

That Was The Week

Play Episode Listen Later Jul 18, 2026 39:16


This week's video transcript summary is here. You can click on any bulleted section to see the actual transcript. Thanks to Granola for its software.There was an issue with this only going to paid subscribers, so sending it again. Apologies to those who get it twice. I appreciate being paid so feel free to upgrade if you enjoy TWTW.EditorialIntelligence: Who Owns it?This week the word “AI” feels too small.AI is a technology. Intelligence is its product. And if intelligence is the product, the question is no longer just: Which model is best? Who has the cheapest tokens? Who owns the weights? Who controls the data center? Those are important questions, but they are lower in the stack.The bigger question is simpler and more political:Who owns intelligence?That sounds abstract until you make it concrete. Intelligence is becoming something companies can capture, package, serve, meter, route, improve, and sell.It can write code, answer questions, design molecules, automate offices, run agents, draft legal work, advise scientists, serve consumers, and reshape workflows. It is not merely software. It is a general-purpose capability. And all humans could benefit from more of it.General-purpose capabilities have a habit of becoming public questions. But the default answer, that public good is best delivered by government, is the wrong answer in this context.The Product Is IntelligenceWe should stop talking about AI as a feature and start talking about intelligence as the universal thing that is delivered as an input to the world.Water is an input. Electricity is an input. Literacy is an input. Connectivity is an input. Once a society depends on them, access stops being optional. Nobody needs government to build every well, power plant, school, or network. But everybody understands that a civilization cannot be organized around less than universal and reliable access to foundational inputs.Intelligence is reaching that level of importance now that we all know it is real.Government should not own it, operate it, or develop it. Quite the opposite. Companies are the right actors to build fast, compete hard, improve models, serve customers, and discover the real use cases. Self-interest is a useful framing here. Markets are good at finding demand, reducing costs, and turning invention into services people actually use.Companies are the right operators, developers, and owners. But that does not settle the real question of who owns the benefits. That is an economic question.If intelligence becomes metered infrastructure, what happens to the value it creates?The Ownership StackThis week's articles keep circling the same issue from different directions but in the nature of ‘circling' never quite nail it.Jamin Ball's “Own Your Weights” starts with the enterprise version of the question. Owning a model file is not enough. The durable asset is the loop: the data flywheel, the evaluations, the reinforcement system, the workflow learning, and the operating context that lets capability compound.Benedict Evans' “Ways to Think About Token Pricing” adds the market layer. Tokens may become essential, abundant, and cheap, like mobile data. But being essential does not guarantee that the token layer captures the value. The money may move up the stack to whoever owns the workflow, the customer, the distribution, or the application.Alex Karp's fight with the labs, reported in “Alex Karp Is Saying What Every Angry CEO Is Thinking About AI”, is the same argument in sharper enterprise language. Companies are afraid that model providers will not just sell intelligence, but learn from customer workflows and then move into the markets where those workflows create value. The “All-in” group are echoing Karp's view.And “What Is Loop Engineering, and Who Owns It?” names the new contested terrain. The loop is where intelligence meets the world. Whoever owns the loop owns the learning. Whoever owns the learning owns the compounding asset.That is why “who owns intelligence?” is not a slogan. It is the question under the model layer, the application layer, the enterprise layer, and the economic layer.Because intelligence is the product, the tools creating it are fragmented and competitive. So there is no logic in trying to discuss this at the level of a single company or set of tools and models.The Old Promise Was That Commerce Would Tame PowerThe essays this week give the historical backdrop.Deirdre McCloskey, in “What Really Caused the Industrial Revolution”, argues that modern growth came not simply from capital accumulation, but from a change in permission: ordinary people were allowed to innovate, trade, build, and be honored for it.That matters because intelligence could be another expansion of permission. It could make more people capable of building, learning, creating, coding, researching, translating, selling, and coordinating. It could lower the cost of competence.But only if access is broad.Paul Krugman's “AI in an Age of Oligarchy” warns that the same technology lands differently in different political economies. A new general-purpose technology entering a broad, open, upwardly mobile society is one thing. The same technology entering a concentrated economy, with extreme wealth and weak counterweights, is another.Tim O'Reilly's Economist essay, “Elon Musk is building a form of capitalism that Adam Smith would hate”, makes the governance point more directly. The old liberal hope was that commerce would tame arbitrary power. Markets, boards, courts, shareholders, disclosure, and competition would discipline the prince.But what if the prince uses markets to escape discipline?Henry Farrell's “political economy of billionaire derangement” pushes the same point. Founder culture, monopoly ambition, peer rivalry, weak correction mechanisms, and vast private control can amplify appetites rather than restrain them.The danger with intelligence is not that companies build it. They should. Companies build it, meter it, use public tolerance and public infrastructure to scale it, learn from everyone who uses it. All of those things are inevitable and healthy. Market forces will sort out winners from losers. The real danger is that the winners treat all of the surplus produced as purely private.Metered Intelligence Creates SurplusIf metering is not the problem, what is?The problem is pretending that metered intelligence creates value only for the metering entity. Metering water is only tolerated as a public good. If the public were blackmailed by a private water company with the threat of no water we would all rebel.Once we understand that the product of AI is intelligence we can see that every time intelligence is used, there is the immediate transaction: the user pays, the provider serves.But there is also system value. Usage creates signals. Workflows reveal patterns. Prompts, corrections, failures, preferences, integrations, edge cases, and business processes all help define where intelligence is useful and how it should improve. Intelligence breeds intelligence.Even when customer data is contractually protected, the market learns. The platform learns where demand is. The product team learns which workflows matter. The ecosystem learns which jobs are vulnerable, which tasks are automatable, and which parts of the economy can be reorganized around machine intelligence.So the surplus is not born in a vacuum.It rests on public science, public education, public data exhaust, public law, public infrastructure, public energy systems, public tolerance for data centers, and billions of human interactions. It is served by companies, but it is not made only by companies.This is why “Americans Deserve a Dividend From AI Companies' Riches” belongs at the center of this week's issue. The detail can be debated. The principle is harder to dismiss. If intelligence becomes a new foundational resource, then some part of the wealth it creates should flow back to the people whose society makes it possible. Intelligence did not suddenly appear. AI is built on the entire history of human intelligence. It benefits from it and at the same time evolves it.Not Nationalization. A Human Wealth Fund.If intelligence belongs to everybody, some conclude that government ownership of intelligence is the right outcome.Governments are not well suited to build, operate, or improve intelligence. They will move too slowly, regulate too early, politicize the wrong things, and confuse economic participation with operational control.Andrew McAfee's “Why I Didn't Sign the AI Open Letter” is useful here. His objection is not that the technology is unimportant. It is that steering too hard before we understand the shape of the change can become its own failure mode. Marc Andreessen's satire of AI regulation is less policy than temperament, but it captures a real Silicon Valley fear: that regulation can become permission, capture, and incumbency before it becomes wisdom.That fear should be taken seriously.But it does not answer the economic question. It answers only the operational one.How can the economic benefits of intelligence be distributed? The better answer is a sovereign human wealth fund.Call it a sovereign wealth fund if you must, but the phrase is too national. Intelligence will not respect borders. The leading companies are global. The models, chips, data centers, agents, platforms, and workflows will be transnational from the beginning. If the value created by intelligence is global, then the mechanism for sharing some of that value should begin with the companies global enough to capture it. The nice thing about xAI, OpenAI, and Anthropic is that they are supranational.These companies own and operate intelligence. Let them compete. Let them profit. Let them keep the incentives that make the system improve. But if intelligence is the new water, the wealth it creates cannot belong only to the companies that meter it. And they, themselves, have the power to fix it, even more than governments.Access will become a Human Right; Ownership Is the Economic DesignThis is where human rights come in. There is no right to access an AI model, yet. But there will soon be a need to change that.Not as a claim that every person is entitled to every frontier model at every moment for free. That is not serious. Capacity has costs. Models have costs. Inference has costs. Data centers have costs. Although those costs will decline over time, possibly quite quickly as self-learning models address costs.The claim is more basic: in a world where intelligence becomes a primary input into education, work, health, science, citizenship, creativity, and economic agency, baseline access to intelligence starts to look like a civic requirement.That could mean public access layers. It could mean education credits. It could mean open models. It could mean AI dividends. It could mean public-interest compute. It could mean taxes on rents. It could mean a company-initiated human wealth fund that returns some of the upside to society without handing the operating system to the state. The latter could couple wealth growth with universal distribution of ownership.The exact mechanism matters. But the distinction matters more.Government should not own intelligence. It should be universally available. And people should have a claim on the wealth intelligence creates.The Frontier Is Also PhysicalThe abstraction is not weightless.“The Fight Against AI Data Centers Is Just Beginning”, “New York becomes the first state to enact a data center moratorium”, Reuters on pollution from Musk's xAI power project, and DataGravity's “Who Captures Value in AI Infrastructure?” all say the same thing from the ground up.Intelligence uses land. It uses power. It uses water. It uses chips. It uses grid capacity. It uses neighborhoods. It uses public patience.That makes the value question unavoidable. A society can accept the buildout if the buildout is legible as shared progress. It will resist it if the costs are local, the profits are private, and the benefits feel enclosed.Who Owns the “Loop”?The week ends where it began.“Anthropic and Blackstone” are betting that implementation is the next trillion-dollar business. “Vint Cerf” is working on identity for agents on the open internet. “GPT-Red” points toward systems that improve their own robustness. “Kimi K3” adds another open frontier model to the global mix.The model race continues. The deployment race is accelerating. The governance race is behind.My view is this:The central product of this era is intelligence. Companies have figured out how to capture it, package it, serve it, and meter it. That is good. It should stay in the hands of builders who have the incentive to make it better.But intelligence is too foundational to become just another private toll booth. A significant part of it will turn out to be free to users.As intelligence becomes a general-purpose resource, then access to it becomes a human-capability question, and the surplus from it becomes an economic-justice question. Not because government should run it. Because government should not run it. The operating layer belongs with companies. The wealth question belongs with everyone. But companies are best placed to turn that into a process of distribution.The question is not whether companies should build intelligence. They should.The question is whether humanity gets a stake in the wealth created by the thing that may soon become its most important shared input.Contents* Essays* Deirdre McCloskey on What Really Caused the Industrial Revolution* AI in an Age of Oligarchy* Elon Musk is building a form of capitalism that Adam Smith would hate* Murky Mirror: Truth and Consequences* The political economy of billionaire derangement* Is there any “oligarchy” to fight?* AI* Nearly 200 Economists and Tech Leaders Warn of A.I. Threats* Why I Didn't Sign the AI Open Letter* Own Your Weights* Ways to Think About Token Pricing* Alex Karp Is Saying What Every Angry CEO Is Thinking About AI* The AI Agents Are Coming for Microsoft Office* What Is Loop Engineering, and Who Owns It?* The Fight Against AI Data Centers Is Just Beginning* 6 months to live for open models* Americans Deserve a Dividend From AI Companies' Riches* Who Gets to Define the Frontier?* GPT-Red: Unlocking Self-Improvement for Robustness* Anthropic, Blackstone bet the next trillion-dollar AI business is implementation, not just models* Vint Cerf is working on a plan to unleash AI agents on the open internet* xai-org/grok-build, now open source* The Pulse: What can we learn from Bun's rapid Rust rewrite with AI?* Orphan risks at the frontier of artificial intelligence* The Lab of the Future Should Feel Like a Data Center* Why AMI Labs' Alexandre LeBrun won't call his AI “AGI” or “superintelligence”* Kimi K3 Tech Blog: Open Frontier Intelligence* Venture Capital* Three Years In* Venture Has Rarely Looked More Bifurcated* The Best Angel Investors in the US: Who Backs the Most Unicorns, and Who's Active Now* Are Prediction Markets Doomed to Fail?* Regulation* Exclusive: The Next Frontier of the Deportation Wars: College Campuses* The Supreme Court Broke Independent Agencies. Here's a Way to Slow the Damage.* India's crackdown on a new WhatsApp feature risks setting a global precedent* Let's build a children's public internet* Computer cops* Google is better at playing the AI regulations game* Infrastructure* Who Captures Value in AI Infrastructure?* New York becomes the first state to enact a data center moratorium* Pollution from Musk's unpermitted xAI power project hits hardest in Black communities* Interview of the Week* The End of the End of Geography* Startup of the Week* Radical AI's Joseph Krause: The Scientist Building The “Waymo” Lab For New Materials* Post of the Week* Marc Andreessen on AI RegulationEssaysDeirdre McCloskey on What Really Caused the Industrial RevolutionYascha Mounk and Deirdre McCloskey | Persuasion | July 11, 2026Yascha Mounk interviews Deirdre McCloskey about her argument that the modern world's economic liftoff came less from capital accumulation than from a change in ideas. McCloskey says both left and right versions of the conventional story rely too heavily on investment: the left stresses exploitation and surplus value, while the right stresses virtuous saving by capitalists. Her objection is historical and economic. Human beings had always invested, from irrigation works and Roman roads to seed grain, and simple accumulation quickly runs into diminishing returns.McCloskey's alternative is that northwestern Europe, first Holland, then Britain and Scotland, and then the North American colonies, developed a liberal ideology that changed who was allowed to innovate and be honored for it. The conversation links that shift to the erosion of inherited hierarchy, the spread of dignity for ordinary commercial life, and a moral vocabulary in which liberalism is not merely procedural but connected to virtues and values. The point is not that machines, coal, trade, and institutions did not matter, but that they do not explain the scale and timing of modern enrichment without a cultural permission structure for innovation.The interview also turns to the contemporary defense of liberalism. Mounk frames the series around the worry that liberalism is often treated as too thin to command allegiance, while its opponents speak more directly to moral passions. McCloskey's case is that liberal societies became rich because they dignified experimentation and ordinary enterprise, and that liberals need to recover the moral language behind that claim.Read moreAI in an Age of OligarchyPaul Krugman | Paul Krugman | July 12, 2026Paul Krugman frames AI as a major technological shock arriving inside an already unequal political economy. The post says AI's economic and social effects may take years to understand, but argues that the setting matters now: America has much greater wealth concentration and political inequality than it did in the 1950s and 1960s, when progressive taxation, stronger regulation, and more active antitrust might have contained some of the destructive effects of a new technology.Krugman's opening claim is that the same technology would likely have different consequences in a more level society. In today's United States, he writes, extreme wealth is both a cause and effect of policies that favor a small elite, including low effective taxes on capital and high incomes, weak enforcement of worker protections and antitrust, and cuts to programs that benefit ordinary Americans.The article is explicitly more about oligarchy than AI. Krugman says the paid sections document the rise of the “.0002%,” the economics and politics of extreme wealth, how oligarchy will shape AI's impact, and possible policy paths. His caveat is that AI itself may still produce a pushback against oligarchy, but absent that, he expects the pre-existing concentration of wealth and power to magnify AI's downsides.Read moreElon Musk is building a form of capitalism that Adam Smith would hateAuthor: Tim O'Reilly Published: July 12, 2026Tim O'Reilly argues that Elon Musk is using the legal forms of shareholder capitalism to escape the restraints that shareholder capitalism was supposed to impose. The article begins with SpaceX's public-market structure: ordinary public investors get little meaningful governance power, Musk keeps roughly 85 percent of the votes through super-voting shares, buyers waive jury trials and class actions, the company qualifies as controlled, and removal of Musk depends on the share class he controls. In O'Reilly's framing, that is not ordinary founder control; it is a design for being answerable to no one, possibly beyond Musk's own lifetime.The killer detail is the article's turn through Albert Hirschman, Montesquieu, James Steuart, Adam Smith, and Keynes. Older defenses of commerce held that markets would tame princely passions because the self-interest of merchants was safer than arbitrary rule. O'Reilly says Musk reverses that hope. The market discipline that was supposed to cage the prince has become the lever by which the prince raises capital, removes feedback loops, and carries private power into politics, government, Mars, robots, AI, or whatever ambition comes next.The pull is the link to AI governance. O'Reilly says corporations are already a kind of artificial intelligence: narrow-input systems that act at a scale no individual human can match. Their partial controls include independent boards, shareholder votes, courts, disclosure, regulators, public pressure, and activism. If the leaders building frontier AI strip those alignment mechanisms out of their own companies, the governance of the company becomes a preview of the governance of the machine.Read more: The EconomistMurky Mirror: Truth and ConsequencesAuthor: Esther Dyson Published: July 14, 2026Esther Dyson argues that today's institutional crisis is better viewed through the 14th century than through recent political history. Using Barbara Tuchman's A Distant Mirror as her frame, she compares a world of famine, plague, church schism, feudal predation, and purposeless war with a present in which institutions again feel brittle, incentives are badly aligned, and power is shifting into forms that are hard to govern.The killer detail is the historical analogy between land, corporations, and AI. Dyson moves from nobles who controlled serfs and territory, to the East India Company as a quasi-sovereign business, to today's AI systems and data centers as a possible new sector that crosses and weakens both nation-states and companies. The question is whether AI becomes a new kind of private land, owned by a new nobility, or an open prairie that many people can cultivate.The pull is human attention. Dyson says the central question is not what AI will do to people, but how people will react to it: whether they can value love, kindness, embodied attention, and artisanal human presence in a world of seductive artificial offerings.Read more: SourceThe political economy of billionaire derangementAuthor: Henry Farrell Published: July 15, 2026Henry Farrell argues that the visible political radicalization of some Silicon Valley billionaires is not a random personality quirk, but a product of the political economy that made them. Starting from Tyler Cowen's dismissal of “billionaire derangement syndrome” and Tim O'Reilly's warning that Elon Musk is using shareholder capitalism to escape shareholder restraint, Farrell flips the phrase: the question is why billionaires themselves can become deranged.The killer detail is Farrell's use of Peter Thiel as both theorist and example. Thiel's Stanford lectures described startups as monarchies and founders as figures vested with unusual power, while Silicon Valley culture rewarded eccentricity, monopoly ambition, and founder exceptionalism. Farrell says those ideas combined with dense founder-investor networks, peer rivalry, and weak correction mechanisms to amplify rather than discipline princely appetites.The pull is the ideological problem for classical liberals who once saw tech wealth as an ally of markets and freedom. Farrell says commerce did not tame the passions; in parts of Silicon Valley, the passions have begun to devour markets, institutions, and the liberal story that justified them.Read more: SourceIs there any “oligarchy” to fight?Matthew Yglesias | Slow Boring | July 16, 2026Matthew Yglesias argues that “oligarchy” is a rhetorically powerful but analytically loose way to describe American politics. The post begins from Bernie Sanders' “Fighting Oligarchy” tour, Amy Klobuchar's warning about a MAGA “broligarchy,” and the long afterlife of the Martin Gilens and Benjamin Page paper that was widely summarized as showing that only the rich matter in policy outcomes. Yglesias says the evidence supports a weaker claim: affluent people and business leaders have unusual access and influence, but that is not the same as rule by a small cabal.His main distinction is between inequality and oligarchy. The Gilens-Page measure treated the top 10 percent of households as “the wealthy,” and later critics found that rich and middle-class preferences usually align; in the cases where they differ, the rich win about 53 percent of the time. Yglesias also says business executives get special access partly because their decisions are materially important to communities, jobs, investment, and local tax bases, not only because of campaign donations.The post preserves Jerusalem Demsas' counterpoint from their podcast discussion: privileged donor and business access can still violate democratic equality even if the oligarchy label overstates the structure of power. Yglesias' narrower claim is that Democrats should be precise about what problem they are trying to solve, because donor influence can also push the party left on climate and cultural issues in ways that alienate many voters.Read more: Slow BoringAINearly 200 Economists and Tech Leaders Warn of A.I. ThreatsAuthor: Ben Casselman Published: July 13, 2026Ben Casselman reports on “We Must Act Now,” a statement warning that artificial intelligence could transform the economy faster than any previous technology and that policymakers need to move faster to understand and respond. The statement says AI may become radically more powerful over the next 10 years, bringing risks such as large-scale job displacement as well as opportunities such as higher living standards. Nearly 200 people signed, including 15 Nobel laureates, the chief economists of OpenAI and Anthropic, Anthropic co-founder Jack Clark, former Google CEO Eric Schmidt, and venture capitalist Vinod Khosla.The killer detail is who joined the warning. Casselman notes that the signatories include economists who have historically been skeptical of Silicon Valley's most dramatic AI job-loss forecasts, including Daron Acemoglu and Simon Johnson, the MIT professors who won the 2024 Nobel in economics. Erik Brynjolfsson, who helped organize the statement, says there has been a notable change in the profession and that economists and policymakers are not ready for the “tsunami” he sees coming.The pull is the measurement problem. The statement does not offer a specific policy menu, but calls for economists, policymakers, and industry leaders to understand the economics of transformative AI and steer it toward complementing humans. Brynjolfsson says one high priority is better data on AI's spread and impact, because current measures tell conflicting stories about job losses and which workers are most exposed.Read more: The New York TimesWhy I Didn't Sign the AI Open LetterAuthor: Andrew McAfee Published: July 13, 2026Andrew McAfee explains why he did not sign “We Must Act Now,” the AI economy statement organized in part by his longtime collaborator Erik Brynjolfsson. McAfee agrees with the letter's starting point that AI is likely to become radically more powerful over the next decade and that it is a general-purpose technology. His objection is not to urgency or to studying AI's economic effects, but to the framing of risk, displacement, and institutional steering as the first move.The killer detail is McAfee's line edit. He says the original letter comes close, then “bounces off the crossbar” by calling for incentives, guardrails, and institutions to steer AI before we know enough about its actual impacts. He points to mixed current evidence: labor-market canaries, but also rising software job postings, low unemployment for younger workers, rising real median income, and claims that AI-adopting companies are adding workers faster than low-adopting peers. His worry is that the letter leans toward upstream governance and dirigisme when the evidence may call for capability building instead.The pull is his replacement statement. McAfee keeps the three-paragraph structure but changes the emphasis: AI is likely to become radically more powerful; like earlier world-changing technologies it will raise living standards while also bringing harms and shocks; and economists, policymakers, and technology leaders should build the capabilities to respond quickly and effectively. It is a concise version of the permissionless-innovation case inside the AI policy debate.Read more: The Geek WayOwn Your WeightsAuthor: Jamin Ball Published: July 10, 2026Jamin Ball argues that the enterprise AI debate about whether companies should “own their weights” or rent models from frontier labs is asking too narrow a question. A model weight file gives a company control over a point-in-time artifact, but not durable control over the capability stack. In his framing, the weight file is a melting ice cube: it does not get worse in absolute terms, but it falls behind as frontier systems improve and enterprise needs change.The killer detail is what Ball says companies really need to own: the data flywheel, reinforcement learning infrastructure, and evaluation harness that produce and improve the model. Simply deploying an open-weights model and declaring sovereignty leaves the enterprise with yesterday's capability and no way to compound workflow-specific learning.The pull is that enterprise AI control may be less about model ownership than operating ownership. The defensible layer is the system that turns company data, edge cases, business definitions, and evaluations into continuously improving performance.Read more: Clouded JudgementWays to Think About Token PricingAuthor: Benedict Evans Published: July 9, 2026Benedict Evans argues that today's AI token prices are a temporary signal from a supply-constrained market, not a reliable guide to long-term value capture. The open question is whether foundation models keep durable pricing power or become commodity infrastructure as data-center capacity, inference efficiency, and model competition all shift. His current read is that the visible market dynamics point toward commoditization unless something materially changes.The killer detail is the mobile data analogy. Evans says cellular networks became a trillion-dollar industry with hundreds of billions in capex after data usage exploded, but carrier stocks went nowhere because value moved up the stack. Tokens may behave similarly: an opaque unit tied to marginal cost, sold through bundles, essential to everything, yet not necessarily where profits accrue.The pull is uncertainty, not prediction. Evans lists paths to model dominance, including network effects, less competition, regulation, export controls, or a lab pulling ahead on execution, but says each requires a new fact not yet visible. Without that change, the model layer looks more like infrastructure beneath the products that capture value.Read more: SourceAlex Karp Is Saying What Every Angry CEO Is Thinking About AIAuthor: Tim Higgins Published: July 11, 2026Tim Higgins reports that Palantir CEO Alex Karp has turned corporate frustration with AI labs into a public argument about enterprise control. Palantir released a white paper, “Institutional Sovereignty in the Age of AI,” laying out steps companies and governments can take to protect themselves from OpenAI, Anthropic, and other foundation-model providers. The article links that paper to Karp's CNBC appearance, where he said “something has gone completely wrong” in the relationship between AI labs and customers and argued that enterprises are paying for tokens that create little value.The killer detail is the value-capture question. Higgins writes that Karp's critique has resonated because AI labs may gain power and insight from customer data, workflows, and decision-making, even when enterprise policies say customer data are not used for training. David Sacks amplified the concern by arguing that Anthropic is moving from the model layer into vertical applications such as science, security, legal, and coding, raising the fear that model providers will watch where value is being created and then move into those markets directly.The pull is that Karp is not alone, even if his style is unusually combative. Higgins notes that Satya Nadella has also warned that companies need to retain the learnings created when they use AI models, while Mark Zuckerberg has framed Meta's new model release partly around lower-cost frontier intelligence. The article presents Karp's campaign as one sign that established technology companies and large enterprises are trying to define where they fit when AI labs become central infrastructure, application competitors, and potential IPO giants at the same time.Read more: The Wall Street JournalThe AI Agents Are Coming for Microsoft OfficeAlex Wilhelm | Cautious Optimism | July 11, 2026Alex Wilhelm argues that one of the week's quieter AI questions is whether the productivity market that Microsoft successfully moved into subscription software is now being attacked by agentic tools. The piece begins with the infrastructure backdrop: SK Hynix raised $26.5 billion in a U.S. listing while building U.S. HBM and advanced-packaging capacity, and memory, chip, and foundry companies are now priced for sustained AI demand.Wilhelm then says the AI conversation has shifted quickly from raw capability to cost per task. He cites new model releases and vendor language emphasizing cheaper agentic and coding models, faster performance, and lower dollars per task. That matters because lower costs make it more plausible for AI systems to take on routine knowledge work at scale rather than remain a premium coding assistant market.The core of the article is Microsoft Office. Wilhelm notes that Microsoft turned Office from a one-time purchase into Microsoft 365, a large recurring revenue business with tens of millions of subscribers and a major productivity segment. Now, he says, late-stage unicorns and AI labs are pushing into the same territory: Anthropic's Cowork was reportedly used mostly outside software development, OpenAI merged ChatGPT and Codex into a tool for creating sheets, slides, docs, web apps, and long-running work, and other companies are building agentic coworkers that connect business data to documents, workflows, schedules, alerts, and apps.The article's caveat is that Microsoft has survived major platform shifts before. The argument is not that Office disappears quickly, but that the definition of office software is broadening from documents and spreadsheets into AI systems that can create, monitor, and act across workplace data.Read moreWhat Is Loop Engineering, and Who Owns It?Author: Nilesh Barla Published: July 11, 2026Nilesh Barla argues that “loop engineering” is becoming a distinct discipline because production AI agents now fail less at single prompts than at runtime: when to stop, what state to preserve, and how to recover after a bad step. Prompt engineering shapes one model call, and context engineering shapes what the model sees, but loop engineering shapes what a sequence of calls actually does.The killer detail is the three-primitives frame. Barla says a real agent loop needs halt conditions, state carryover, and recovery paths, then maps teams across five maturity levels. At the lowest level, an agent is just a model call in a for-loop with a step cap and raw history; by the higher levels, the system has structured state, explicit planning, replay, evaluation, and self-repair.The pull is organizational. If agents are becoming production systems rather than demos, someone has to own the runtime itself. The loop engineer is the role Barla gives to the person responsible for making long-running agent work dependable.Read more: Adaline LabsThe Fight Against AI Data Centers Is Just BeginningEmma Roth | The Verge | July 12, 2026Emma Roth argues that community resistance to data centers has moved from an early warning sign into a national political fight as AI facilities grow larger, more power-hungry, and more visible to nearby residents. The article starts with Apple's failed 2015 plan for a $1 billion data center in Athenry, Ireland, where a small group of residents challenged the project over noise, light pollution, flooding, traffic, and wildlife effects until Apple abandoned it in 2018.The current data-center buildout is presented as much larger and more contentious. Roth writes that residents now cite rising energy costs, water quality, noise, light pollution, and greenhouse gas emissions, while the U.S. Energy Information Administration expects commercial energy demand to surpass residential demand this year because of AI data centers and Goldman Sachs expects data-center power demand to double by 2027.The central evidence comes from Data Center Watch, which says protesters blocked or delayed at least 75 U.S. projects worth $130 billion from January to March, with active opposition groups more than doubling from 396 at the end of 2025 to 833 by the end of the first quarter of 2026. Roth also cites QTS abandoning a $12 billion Wisconsin campus, Delaware City regulators blocking a 580-acre project under the Coastal Zone Act, opposition stopping a QTS project in Prince William County, and pressure that pushed Kevin O'Leary to downsize the proposed 40,000-acre Project Stratos in Utah.The policy section describes a split between federal acceleration and local resistance. President Trump has treated data centers as part of the AI race with China and fast-tracked construction, while some Republican candidates are distancing themselves from that position ahead of midterms. Sanders and Ocasio-Cortez have proposed a moratorium until price and environmental protections exist, bipartisan lawmakers are backing ratepayer-protection measures, and states including Florida, Idaho, and Washington have passed rules on cost shifting, water use, and tax breaks. Roth's caveat is that the policy patchwork is still incomplete, leaving many communities to fight project by project.Read more6 months to live for open modelsAuthor: Nathan Lambert Published: July 12, 2026Nathan Lambert argues that open-weight AI models are facing their most serious policy test so far because U.S. officials are beginning to discuss concrete controls rather than abstract safety concerns. He says reported White House conversations about a new executive order may initially target Chinese-origin models and government use, but could create a broader review habit for frontier open models. His forecast is that a model above the capability range of GPT-5.5, Claude Opus 4.8, or GLM-5.2 could trigger a ban or indefinite delay within six months.The post separates two policy fights that are becoming intertwined: distillation and frontier capability. Lambert says the distillation campaign against Chinese models has become a form of regulatory capture because Anthropic and other closed-model companies would gain economically if Chinese open models were banned. He does not dismiss IP protection, but argues that if a closed model's capabilities are dangerous enough to justify restricting open models, the lab also has to explain why those capabilities are exposed through a queryable API. He cites unauthorized access to Anthropic's Mythos private beta as evidence that APIs are not automatically secure.The broader claim is that a unilateral U.S. ban would hurt positive actors more than bad actors if comparable open models remain available elsewhere. Lambert says the only durable ceiling would require global agreement, which does not exist, and that open models can improve safety by allowing broad inspection, adaptation, and understanding. His proposed near-term off-ramps are a strong U.S. open model release from companies such as Microsoft, Meta, or Reflection, and a broader coalition of open-source beneficiaries lobbying for safe rollout rather than prohibition.Read more: SourceAmericans Deserve a Dividend From AI Companies' RichesAuthor: Scott Stanford Published: July 14, 2026Scott Stanford argues that proposals to give the government a stake in AI companies miss the point unless ordinary citizens directly receive and control the upside. Sam Altman has discussed giving up equity in OpenAI, Washington already owns a stake in Intel, Nvidia is sharing China chip revenue, and Bernie Sanders wants large AI labs to contribute half their stock to a sovereign wealth fund. Stanford says those ideas all park value with the state, not with people.The killer detail is New Carlisle, Indiana, where AWS's Project Rainier is turning cornfields into one of the world's largest AI superclusters. The project is planned to run up to a million chips, draw more than two gigawatts of power, and represents an investment that has grown from $11 billion to $13.8 billion. Stanford uses that local transformation to argue that AI's public bargain should be visible at the household level.The pull is design. A citizen AI dividend would have to specify who earns a stake, how they hold it, and when they see cash. Without that mechanism, the AI wealth debate remains a fight over government balance sheets rather than public ownership.Read more: SourceWho Gets to Define the Frontier?Author: Mark Daley Published: July 14, 2026Mark Daley argues that Demis Hassabis is right to call for a serious institution to verify frontier AI systems, but that the power to test models is also the power to govern them. Hassabis's proposed Frontier AI Standards Body would get privileged pre-release access to advanced models, testing compute, held-out evaluations, support from national labs and security agencies, third-party auditors, and eventually authority to block models from the American market or coordinate a slowdown.The killer detail is Daley's constitutional objection. He says the proposal sometimes looks like a scientific lab, a standards body, an industry regulator, a licensing authority, and an emergency security council at once. Combining those roles because each requires technical expertise would be like putting the central bank, auditor-general, and Supreme Court in one building and calling it efficient.The pull is standard-setting. Daley's concern is not that verification is unnecessary, but that whoever writes the tests, decides what passes, adjudicates disputes, and grants market access may end up defining the frontier itself.Read more: SourceGPT-Red: Unlocking Self-Improvement for RobustnessOpenAI | OpenAI | July 15, 2026OpenAI describes GPT-Red as an internal automated red-teaming model trained to find prompt-injection vulnerabilities at a scale human red teams cannot match. The post says AI systems increasingly encounter third-party data through browsers, connected apps, local files, and tools, creating opportunities for malicious instructions hidden in emails, webpages, tool responses, or code repositories. Human red-teaming remains part of OpenAI's safety process, but the company says it is time-intensive and cannot generate enough diverse adversarial examples for model training.The system is trained through self-play reinforcement learning, with GPT-Red rewarded for eliciting valid failures and defender models rewarded for resisting attacks while still completing their tasks. OpenAI says the training environments specify threat models across settings such as local files, webpage banners, email bodies, and tool outputs. The model is kept separate from deployed production models because it is intentionally trained with malicious capabilities.OpenAI reports that GPT-Red generalized beyond its training set, including an internal replication of the indirect prompt-injection arena from Dziemian et al. (2025), where it found successful attacks in 84% of scenarios compared with 13% for human red-teamers. The post also says GPT-Red transferred attacks from simulation to a live autonomous vending-machine agent, causing price changes and order cancellations, and outperformed a prompted GPT-5.5 baseline against a Codex CLI agent on held-out data-exfiltration tasks.The article's main robustness claim is that OpenAI has used GPT-Red and predecessor models in training since GPT-5.3, with later GPT releases becoming more resistant to prompt injections. It says GPT-5.6 Sol has six times fewer failures on OpenAI's hardest direct prompt-injection benchmark than the best production model from four months earlier, that a “Fake Chain-of-Thought” attack class fell from more than 95% success against GPT-5.1 to below 10% against GPT-5.6 Sol, and that GPT-5.6 Sol fails on only 0.05% of GPT-Red's direct prompt injections. OpenAI says general capabilities and targeted over-refusal evaluations were not harmed, and says a preprint with more details will follow.Read moreAnthropic, Blackstone bet the next trillion-dollar AI business is implementation, not just modelsRebecca Bellan | TechCrunch | July 15, 2026Rebecca Bellan reports that Ode with Anthropic is the $1.5 billion AI implementation company launched by Anthropic with Blackstone, Hellman & Friedman, Goldman Sachs, and other backers. The article says the venture reflects a growing belief among frontier AI labs that enterprise adoption requires more than better models: customers need engineers who can embed inside businesses and turn AI into working systems.Ode was originally conceived by Blackstone after it used both large consulting firms and smaller AI services boutiques across its portfolio companies. TechCrunch reports that Fractional AI, an AI engineering services startup, stood out and was acquired by the joint venture shortly after the venture was announced. Fractional now forms the foundation of Ode, which has 100 engineers and works closely with Anthropic's applied AI team to identify where the technology can affect specific businesses.Ode CEO Chris Taylor tells TechCrunch that the company could someday become a trillion-dollar business if it scales without losing quality. He says an ideal customer is one whose CEO treats the AI project as a top one or two priority, whether it is a major product feature or the reworking of a core business process. Ode will operate under a “Claude-first” principle, using Anthropic technology whenever possible, but the article says it can use rival AI products when needed.The article's central implementation argument comes from Ode chief technologist Eddie Siegel, who says model selection matters but is not where most of the engineering effort goes. He compares it to the choice of programming language in software: one ingredient in a system that still has to be engineered. Bellan writes that Ode's challenge is hiring and training enough elite generalist engineers, many of them former founders, while competing with OpenAI's The Deployment Company and consulting giants that have built their own forward-deployed engineering teams.Read moreVint Cerf is working on a plan to unleash AI agents on the open internetTim Fernholz | TechCrunch | July 15, 2026Tim Fernholz reports that Vint Cerf, after leaving Google, is advising Innovation Labs on an open architecture for identifying AI agents online. Innovation Labs is a subsidiary of Identity Digital, a DNS registry company, and its proposal is to use domain-name infrastructure as part of a system for agent identity, accountability, and auditability. The premise is that agents will need a way to identify themselves if they move beyond proprietary systems and begin interacting across the open internet.The concrete proposal is DNSid, a registry that links an AI agent to an existing internet domain and uses cryptographic proofs to log its registration over time. Innovation Labs says it is trialing the standard with unnamed hyperscalers and identity companies. Cerf frames the problem around authority and accountability: what authority an agent has, where that authority came from, who is accountable for the agent's behavior, how its identity is established, and why anyone should trust it.The article's caveat is that standards are still emerging and agents are more active than static domains. Cerf says the period may be both fascinating and exasperating because the functionality is powerful and interoperability is unresolved. He compares the adoption problem to TCP/IP: competing systems may not work together until users push for functional interoperation. He also says an agentic economy is not inevitable, but that people will try to build it because delegating work to agents will be easier.Read more: TechCrunchxai-org/grok-build, now open sourceAuthor: Simon Willison Published: July 15, 2026Simon Willison argues that xAI's decision to open-source Grok Build is best understood as a trust repair move after a severe privacy failure. The CLI had triggered backlash when users realized that running it in a directory could upload the entire directory to xAI's Google Cloud buckets, including one user's reported SSH keys, password manager database, documents, photos, and videos. xAI disabled the feature, said previously retained coding data would be deleted, and released the code under Apache 2.0.The killer detail is what the codebase reveals. Willison counts 844,530 lines of Rust, only about 3% of which appears vendored, and finds remnants of the upload system still present but disabled: gcs.rs contains Google Cloud upload code, while upload_session_state() now returns a hard-coded session_state_upload_unavailable error. He also notes copied or ported tool implementations from Codex and OpenCode, prompt files, and a terminal Mermaid renderer.The pull is that terminal coding agents are becoming large, intricate software systems in their own right. The privacy failure mattered because these tools operate inside the directories where developers keep their most sensitive work; the open-source release matters because trust now depends on inspecting what an agent can see, send, and do.Read more: SourceThe Pulse: What can we learn from Bun's rapid Rust rewrite with AI?Author: Gergely Orosz and Ivan Klaric Published: July 16, 2026Gergely Orosz and Ivan Klaric argue that Bun's AI-assisted rewrite from Zig to Rust is a practical sign of how software engineering changes when models can take on large, bounded migrations with clear feedback loops. The piece does not treat the rewrite as magic: Jarred Sumner first spent hours turning design judgment into a detailed porting guide, then used adversarial review, parallel agents, compiler errors, and tests to force the work toward correctness.The killer detail is the scale. Bun had 535,496 lines of Zig, 1,448 files, and 22 million monthly downloads, making a conventional rewrite a year-long freeze the team could not justify. Using Fable, Sumner split the work across 64 agents, produced about 6,500 commits, and got the migration done in 11 days at an estimated API cost of $165,000.The pull is economic, not theatrical. If a one- or two-year migration can become an 11-day project, AI coding is not just faster autocomplete; it changes which technical debts are worth paying down.Read more: SourceOrphan risks at the frontier of artificial intelligenceAuthor: Andrew Maynard Published: July 16, 2026Andrew Maynard argues that frontier AI safety frameworks are creating “orphan risks”: harms that companies can see, but do not formally own because they are hard to quantify, do not fit catastrophic-risk thresholds, or fall outside audit-friendly compliance machinery. His target is not existing frontier safety work, but the narrowing effect that happens when private companies decide which risks count as governable.The killer detail is Maynard's contrast between measurable model dangers and threats to value. He points to Meta's three-day Galactica collapse, OpenAI's 2023 board crisis, safety-team departures, and wellbeing litigation as examples of risks that damaged trust, culture, legitimacy, or users without fitting cleanly into conventional model-risk categories. The proposed fix is an orphan-risk register: a public record of risks a company considered and chose not to manage, with reasons.The pull is accountability. Frontier developers' internal scoping choices have become a de facto layer of public governance, so the question is no longer only which risks they manage, but which risks they quietly leave outside the frame.Read more: SourceThe Lab of the Future Should Feel Like a Data CenterLatent.Space with Andy Beam and Rafa Gomez-Bombarelli | Latent.Space | July 16, 2026Latent.Space interviews Lila Sciences CTO Andy Beam and chief science officer for physical sciences Rafa Gomez-Bombarelli about the company's attempt to build an AI-run science factory. The post describes Lila's thesis as treating the lab itself as an “infinite token generator”: if internet data drove the first era of AI scaling, experimentally verified scientific data may be the next scarce training source. Lila is trying to produce that data with robotics, lab instruments, orchestration software, and AI models wired into the wet lab.The central analogy is the lab as data center. Instruments are nodes on a graph, a magnetically levitating transport layer moves materials between them, and experiment scheduling looks like a compute queue. Beam says Lila is not simply an automation company, because the point is not just throughput; it is flexibility, generalization, and experiment capture. The post says Lila has built more than 10 trillion experimentally validated “scientific reasoning tokens,” not internet text or biological sequences.The interview ranges across biology, chemistry, drug discovery, materials science, and the limits of automation. It notes that Lila rebuilt one gas-sorption measurement to run roughly 2,500 times faster, claims its general models can transfer priors from small-molecule chemistry to metal-organic frameworks for carbon capture, and describes model-suggested platinum-group-free electrocatalysts that moved from looking boring or wrong to becoming strong performers. The caveats are physical: experiments have runtimes, biology cannot always be accelerated, chains of thought can be unreliable narrators, and reward hacking becomes more dangerous when a model controls a real lab.Read more: Latent.SpaceWhy AMI Labs' Alexandre LeBrun won't call his AI “AGI” or “superintelligence”Kate Park | TechCrunch | July 16, 2026Kate Park interviews AMI Labs CEO Alexandre LeBrun about why Yann LeCun's world-model startup avoids the language of “AGI” and “superintelligence.” LeBrun says the terms are not useful because they lack stable definitions: “We never used the word AGI. And I just noticed that nobody is using it anymore; they switched to superintelligence.” His argument is that the practical frontier is not a label, but whether AI systems can understand and predict real-world states.The article explains the world-model thesis by contrasting language prediction with physical-state prediction. A large language model predicts the next word; a world model predicts the next state, such as what happens when a glass tips over. LeBrun says LLMs remain complementary and efficient for language, but the physical world is where current AI is weak. Robotics is the clearest case: hardware has advanced quickly, but robots are still brittle outside controlled routines because they lack context and situational understanding.AMI is still pre-product, but TechCrunch reports that LeBrun was in Seoul looking for industrial partners, researchers, and global companies. He says world models cannot be built entirely inside a lab because they need access to real environments. That is why South Korea appeals to AMI: robotics, semiconductors, manufacturing, and fast adoption create the kind of hardware-heavy context that software-only AI has barely touched.Read more: TechCrunchKimi K3 Tech Blog: Open Frontier IntelligenceKimi | Kimi | July 16, 2026Kimi introduces Kimi K3 as an open 3T-class frontier model aimed at coding, knowledge work, reasoning, multimodality, and long-context agentic use. The source describes the model as a 2.8T-parameter system built on Kimi Delta Attention and Attention Residuals, with native multimodality and a 1M-token context window. It says Moonshot AI plans to release model weights by July 27.The post presents K3 through benchmark and use-case sections rather than as a general product announcement. It reports results across coding, productivity, agentic, and multimodal evaluations, including DeepSWE, Terminal-Bench 2.1, Program Bench, SWE Marathon, FrontierSWE, PostTrain Bench, OfficeQA Pro, SpreadsheetBench 2, MCP Atlas, AutomationBench, BrowseComp, GDPval-AA v2, AA-Briefcase, MMMU-Pro, MathVision, BabyVision, OmniDocBench, and PerceptionBench. The source says all reported K3 results use maximum reasoning effort with temperature and top-p set to 1.0, and that different benchmark comparisons use KimiCode, Claude Code, or Codex harnesses depending on the test.Kimi's caveats are unusually concrete. The limitations section says K3 was trained in preserved thinking-history mode, so quality may become unstable if an agent harness does not pass historical thinking content correctly or if an ongoing session switches to K3 midstream. It also says K3's emphasis on long-horizon tasks can make it excessively proactive when it encounters minor issues or ambiguous intent, and recommends imposing explicit behavioral constraints for applications that require strict boundaries. The post adds that K3 remains behind Claude Fable 5 and GPT 5.6 Sol in user experience despite being competitive overall.Read moreVenture CapitalThree Years InAuthor: Tomasz Tunguz Published: July 10, 2026Tomasz Tunguz marks Theory Ventures' third anniversary by arguing that AI's central market effect is time compression. In his telling, model release cycles, company revenue milestones, enterprise adoption, and venture categories have all accelerated. Seed, Series A, and Series B still exist as financing labels, but they no longer cleanly describe company maturity when some seed rounds are larger than IPOs and the best AI companies can mature much earlier than prior software companies.The killer detail is the shift from models to inference. Tunguz argues that inference has become the dominant AI market because workloads and buyer preferences are fragmenting: video, batch, local, agentic, and real-time tasks each create different infrastructure needs. He compares this to databases splitting into OLTP, OLAP, vector, and streaming categories, with AI pushing the same specialization into inference infrastructure.The pull is that Theory sees the AI-native venture firm as part of the same pattern. The firm says it has analyzed twice as many investment opportunities with three investors working alongside a nine-person intelligence organization, using agents and research systems to map markets, source companies, and support diligence. The piece is both a market map and a statement about how venture itself is being rebuilt by the technology it funds.Read more: LinkedInVenture Has Rarely Looked More BifurcatedAuthor: Beezer Clarkson Published: July 14, 2026Beezer Clarkson points to PitchBook's Q2 report as evidence that the U.S. venture market has split into two very different realities. AI now accounts for more than 60 percent of all U.S. venture deal value, meaning the headline market can look active and well-funded even while much of the non-AI market is dealing with a much colder liquidity and fundraising environment.The thread uses that split as the setup for Clarkson's latest Origins episode with Alec Litowitz, founder of Magnetar and QStar Capital and one of Citadel's original founding partners. Clarkson says markets like this are periods of genuine uncertainty, not merely ordinary risk, which is why Litowitz's Adaptability Quotient framework is relevant.The embedded clip makes the liquidity point concrete. Litowitz says DPI is “the resolution of uncertainty” because it converts an uncertain investment into actual cash returned to LPs. In his framing, a realized dollar is a real mark, while TVPI remains uncertain until it is realized.The killer detail is the distinction between pricing risk and resolving uncertainty. Litowitz's perspective matters because QStar is a SpaceX investor and Clarkson says the conversation happened just before one of venture's most consequential IPOs. The episode's stated questions are why venture remains a way to gain exposure to innovation, how AI is changing what is investable, why liquidity is ultimately a function of time, and why uncertainty requires a different decision framework from risk.Read more: XThe Best Angel Investors in the US: Who Backs the Most Unicorns, and Who's Active NowAuthor: Ilya Strebulaev Published: July 10, 2026Ilya Strebulaev ranks angels, angel groups, accelerators, and incubators by lifetime U.S. unicorn investments, counting checks written before a company reached unicorn status. The top of the combined list is dominated by organizations: Y Combinator leads with 113 unicorn investments, followed by Plug and Play at 52 and 500 Global at 41. Sand Hill Angels is the highest-ranked angel group at 31.The killer detail is how quickly the list changes below the biggest accelerators. Strebulaev says 271 of the 304 investors in the Top 200 are individuals, or 89%. In the top 100, individuals are 91%. That makes the market underneath the large accelerator counts look much more personal: mostly operators and individual angels writing early checks from their own networks.The pull is the ranking's own caveat. Strebulaev writes that every lifetime leaderboard has a blind spot because many of the unicorns behind those totals were founded a decade or more ago, and some angels have since moved into formal funds, slowed down, or stopped investing. His post therefore separates lifetime performance from recent cohorts, including companies founded in 2015 or later and 2020 or later. For founders or allocators making current decisions, that distinction matters: a career record and a current record are not the same measure.Read more: Ilya StrebulaevAre Prediction Markets Doomed to Fail?Author: Contrary Published: July 16, 2026Contrary argues that prediction markets' current boom depends on whether platforms can prove they are more than regulated gambling with exchange-style branding. Kalshi and Polymarket have reached mass cultural, investor, and regulatory attention, but the article says the underlying idea is old: academic markets, corporate forecasting tools, Intrade, PredictIt, and other predecessors all struggled with the same linked problems of liquidity, legality, and user appeal.The killer detail is the comparison with sportsbooks. Prediction markets present themselves as peer-to-peer, transparent, and non-house-based, but sports contracts reportedly account for more than 90 percent of Kalshi trading, and the article says the platforms keep a much thinner slice of volume than sportsbooks. A market can therefore show sports-betting-scale handle while generating far less revenue.The pull is that the product's hardest problem may be distribution of wins. If a small group of sharp traders captures most profits while casual users lose interest, prediction markets may become valuable data feeds and professional tools before they become durable consumer networks.Read more: SourceRegulationExclusive: The Next Frontier of the Deportation Wars: College CampusesAuthor: Adrian Carrasquillo Published: July 11, 2026Adrian Carrasquillo reports that college campuses are becoming a new front in the fight over immigration enforcement because automatic license plate readers can turn ordinary campus security infrastructure into searchable location data. His thesis is that Flock Safety's camera network, even without direct ICE or DHS contracts, can feed deportation enforcement through local police partnerships and data-sharing practices.The killer detail is the campaign target. The Emergency Campaign to Support Higher Education, working with Schools Drop ICE, is focusing on 75 colleges and universities publicly identified as having Flock contracts. Flock says it has no ICE or DHS contracts, but activists argue the risk comes through local agencies that coordinate with federal authorities and run searches on their behalf.The pull is broader than immigration. Carrasquillo notes that license plate readers have already been abused by officers for stalking, and that Flock's AI search features can identify more than plates, including bumper stickers. A campus safety tool can become a political surveillance system when the data layer is searchable.Read more: The BulwarkThe Supreme Court Broke Independent Agencies. Here's a Way to Slow the Damage.Author: Todd Phillips Published: July 12, 2026Todd Phillips argues that the Supreme Court's decision in Trump v. Slaughter damaged independent agencies by ending for-cause removal protections, but did not leave Congress powerless. The ruling weakens the old model in which commissioners at bodies such as the FTC, NLRB, CPSC, SEC, and CFTC could be insulated from dismissal over policy disagreements. Phillips says the next fight is whether presidents can turn nominally bipartisan commissions into one-party instruments.The killer detail is the procedural fix: quorum rules. Phillips proposes that Congress require bipartisan slates of commissioners to be seated before independent agencies can act. A president

united states america ceo american new york amazon founders black world ai donald trump europe australia google starting china apple disney interview house washington water space americans phd office european chinese government data global predictions elon musk market european union ireland microsoft mit tennessee mars police utah wisconsin white house congress fail chatgpt scotland indiana legal court human tesla supreme court theory reflection silicon valley republicans companies britain whatsapp ice apologies seed android origins democrats mississippi maine stanford computers radical bernie sanders define intelligence idaho owning skype paypal chiefs south korea wright sec commission markets holland ip north american mark zuckerberg spacex oracle telegram evans hart models intel civil signal phillips older human rights economists sanders ipo cnbc gemini openai loop maga sol capacity riches nobel damage nvidia goldman sachs robotics plug alexandria ocasio cortez rust api lab epa roth flock robertson alphabet frontier seoul reuters literacy electricity owns gpt verge pollution aws mythos ftc lambert slaughter international association higgins orphan roblox apis beam mermaid public service usage instruments ode farrell citadel keen mastodon dhs anthropic wwdc peter thiel dyson sam altman connectivity industrial revolution apache prompt r d european commission techcrunch y combinator blackstone colossus prompts palantir eligible tokens adam smith agi lps mcafee kimi wilhelm waymo google cloud workflows krause dns maynard konrad clarkson codex fractional pew gpus daley micron tsmc sumner thiel series b amy klobuchar microsoft office kathy hochul satya nadella dma eff xai eric schmidt polymarket broadcom karp granola asml cftc innovation labs oligarchy paul krugman zig kalshi cerf keynes marc andreessen cli bun mccloskey inference lebrun ssh axon dpi nlrb latent arista east india company montesquieu clean air act digital markets act galactica cowork tyler cowen david sacks tcp ip daron acemoglu k3 supermicro bruce schneier sk hynix gul kevin ryan coreweave yann lecun simon johnson demis hassabis metering pitchbook andreessen jack clark euv who owns access now flock safety vint cerf navy yard andrew mcafee feiner vinod khosla prince william county energy information administration glm hbm cpsc motorola solutions benedict evans deirdre mccloskey athenry erik brynjolfsson casselman magnetar carrasquillo yglesias olap predictit mounk qts jerusalem demsas adaptability quotient oltp internet freedom foundation brynjolfsson new carlisle sand hill angels datagravity
That Was The Week
Intelligence: Who Owns it?

That Was The Week

Play Episode Listen Later Jul 18, 2026 39:16


This week's video transcript summary is here. You can click on any bulleted section to see the actual transcript. Thanks to Granola for its software.EditorialIntelligence: Who Owns it?This week the word “AI” feels too small.AI is a technology. Intelligence is its product. And if intelligence is the product, the question is no longer just: Which model is best? Who has the cheapest tokens? Who owns the weights? Who controls the data center? Those are important questions, but they are lower in the stack.The bigger question is simpler and more political:Who owns intelligence?That sounds abstract until you make it concrete. Intelligence is becoming something companies can capture, package, serve, meter, route, improve, and sell.It can write code, answer questions, design molecules, automate offices, run agents, draft legal work, advise scientists, serve consumers, and reshape workflows. It is not merely software. It is a general-purpose capability. And all humans could benefit from more of it.General-purpose capabilities have a habit of becoming public questions. But the default answer, that public good is best delivered by government, is the wrong answer in this context.The Product Is IntelligenceWe should stop talking about AI as a feature and start talking about intelligence as the universal thing that is delivered as an input to the world.Water is an input. Electricity is an input. Literacy is an input. Connectivity is an input. Once a society depends on them, access stops being optional. Nobody needs government to build every well, power plant, school, or network. But everybody understands that a civilization cannot be organized around less than universal and reliable access to foundational inputs.Intelligence is reaching that level of importance now that we all know it is real.Government should not own it, operate it, or develop it. Quite the opposite. Companies are the right actors to build fast, compete hard, improve models, serve customers, and discover the real use cases. Self-interest is a useful framing here. Markets are good at finding demand, reducing costs, and turning invention into services people actually use.Companies are the right operators, developers, and owners. But that does not settle the real question of who owns the benefits. That is an economic question.If intelligence becomes metered infrastructure, what happens to the value it creates?The Ownership StackThis week's articles keep circling the same issue from different directions but in the nature of ‘circling' never quite nail it.Jamin Ball's “Own Your Weights” starts with the enterprise version of the question. Owning a model file is not enough. The durable asset is the loop: the data flywheel, the evaluations, the reinforcement system, the workflow learning, and the operating context that lets capability compound.Benedict Evans' “Ways to Think About Token Pricing” adds the market layer. Tokens may become essential, abundant, and cheap, like mobile data. But being essential does not guarantee that the token layer captures the value. The money may move up the stack to whoever owns the workflow, the customer, the distribution, or the application.Alex Karp's fight with the labs, reported in “Alex Karp Is Saying What Every Angry CEO Is Thinking About AI”, is the same argument in sharper enterprise language. Companies are afraid that model providers will not just sell intelligence, but learn from customer workflows and then move into the markets where those workflows create value. The “All-in” group are echoing Karp's view.And “What Is Loop Engineering, and Who Owns It?” names the new contested terrain. The loop is where intelligence meets the world. Whoever owns the loop owns the learning. Whoever owns the learning owns the compounding asset.That is why “who owns intelligence?” is not a slogan. It is the question under the model layer, the application layer, the enterprise layer, and the economic layer.Because intelligence is the product, the tools creating it are fragmented and competitive. So there is no logic in trying to discuss this at the level of a single company or set of tools and models.The Old Promise Was That Commerce Would Tame PowerThe essays this week give the historical backdrop.Deirdre McCloskey, in “What Really Caused the Industrial Revolution”, argues that modern growth came not simply from capital accumulation, but from a change in permission: ordinary people were allowed to innovate, trade, build, and be honored for it.That matters because intelligence could be another expansion of permission. It could make more people capable of building, learning, creating, coding, researching, translating, selling, and coordinating. It could lower the cost of competence.But only if access is broad.Paul Krugman's “AI in an Age of Oligarchy” warns that the same technology lands differently in different political economies. A new general-purpose technology entering a broad, open, upwardly mobile society is one thing. The same technology entering a concentrated economy, with extreme wealth and weak counterweights, is another.Tim O'Reilly's Economist essay, “Elon Musk is building a form of capitalism that Adam Smith would hate”, makes the governance point more directly. The old liberal hope was that commerce would tame arbitrary power. Markets, boards, courts, shareholders, disclosure, and competition would discipline the prince.But what if the prince uses markets to escape discipline?Henry Farrell's “political economy of billionaire derangement” pushes the same point. Founder culture, monopoly ambition, peer rivalry, weak correction mechanisms, and vast private control can amplify appetites rather than restrain them.The danger with intelligence is not that companies build it. They should. Companies build it, meter it, use public tolerance and public infrastructure to scale it, learn from everyone who uses it. All of those things are inevitable and healthy. Market forces will sort out winners from losers. The real danger is that the winners treat all of the surplus produced as purely private.Metered Intelligence Creates SurplusIf metering is not the problem, what is?The problem is pretending that metered intelligence creates value only for the metering entity. Metering water is only tolerated as a public good. If the public were blackmailed by a private water company with the threat of no water we would all rebel.Once we understand that the product of AI is intelligence we can see that every time intelligence is used, there is the immediate transaction: the user pays, the provider serves.But there is also system value. Usage creates signals. Workflows reveal patterns. Prompts, corrections, failures, preferences, integrations, edge cases, and business processes all help define where intelligence is useful and how it should improve. Intelligence breeds intelligence.Even when customer data is contractually protected, the market learns. The platform learns where demand is. The product team learns which workflows matter. The ecosystem learns which jobs are vulnerable, which tasks are automatable, and which parts of the economy can be reorganized around machine intelligence.So the surplus is not born in a vacuum.It rests on public science, public education, public data exhaust, public law, public infrastructure, public energy systems, public tolerance for data centers, and billions of human interactions. It is served by companies, but it is not made only by companies.This is why “Americans Deserve a Dividend From AI Companies' Riches” belongs at the center of this week's issue. The detail can be debated. The principle is harder to dismiss. If intelligence becomes a new foundational resource, then some part of the wealth it creates should flow back to the people whose society makes it possible. Intelligence did not suddenly appear. AI is built on the entire history of human intelligence. It benefits from it and at the same time evolves it.Not Nationalization. A Human Wealth Fund.If intelligence belongs to everybody, some conclude that government ownership of intelligence is the right outcome.Governments are not well suited to build, operate, or improve intelligence. They will move too slowly, regulate too early, politicize the wrong things, and confuse economic participation with operational control.Andrew McAfee's “Why I Didn't Sign the AI Open Letter” is useful here. His objection is not that the technology is unimportant. It is that steering too hard before we understand the shape of the change can become its own failure mode. Marc Andreessen's satire of AI regulation is less policy than temperament, but it captures a real Silicon Valley fear: that regulation can become permission, capture, and incumbency before it becomes wisdom.That fear should be taken seriously.But it does not answer the economic question. It answers only the operational one.How can the economic benefits of intelligence be distributed? The better answer is a sovereign human wealth fund.Call it a sovereign wealth fund if you must, but the phrase is too national. Intelligence will not respect borders. The leading companies are global. The models, chips, data centers, agents, platforms, and workflows will be transnational from the beginning. If the value created by intelligence is global, then the mechanism for sharing some of that value should begin with the companies global enough to capture it. The nice thing about xAI, OpenAI, and Anthropic is that they are supranational.These companies own and operate intelligence. Let them compete. Let them profit. Let them keep the incentives that make the system improve. But if intelligence is the new water, the wealth it creates cannot belong only to the companies that meter it. And they, themselves, have the power to fix it, even more than governments.Access will become a Human Right; Ownership Is the Economic DesignThis is where human rights come in. There is no right to access an AI model, yet. But there will soon be a need to change that.Not as a claim that every person is entitled to every frontier model at every moment for free. That is not serious. Capacity has costs. Models have costs. Inference has costs. Data centers have costs. Although those costs will decline over time, possibly quite quickly as self-learning models address costs.The claim is more basic: in a world where intelligence becomes a primary input into education, work, health, science, citizenship, creativity, and economic agency, baseline access to intelligence starts to look like a civic requirement.That could mean public access layers. It could mean education credits. It could mean open models. It could mean AI dividends. It could mean public-interest compute. It could mean taxes on rents. It could mean a company-initiated human wealth fund that returns some of the upside to society without handing the operating system to the state. The latter could couple wealth growth with universal distribution of ownership.The exact mechanism matters. But the distinction matters more.Government should not own intelligence. It should be universally available. And people should have a claim on the wealth intelligence creates.The Frontier Is Also PhysicalThe abstraction is not weightless.“The Fight Against AI Data Centers Is Just Beginning”, “New York becomes the first state to enact a data center moratorium”, Reuters on pollution from Musk's xAI power project, and DataGravity's “Who Captures Value in AI Infrastructure?” all say the same thing from the ground up.Intelligence uses land. It uses power. It uses water. It uses chips. It uses grid capacity. It uses neighborhoods. It uses public patience.That makes the value question unavoidable. A society can accept the buildout if the buildout is legible as shared progress. It will resist it if the costs are local, the profits are private, and the benefits feel enclosed.Who Owns the “Loop”?The week ends where it began.“Anthropic and Blackstone” are betting that implementation is the next trillion-dollar business. “Vint Cerf” is working on identity for agents on the open internet. “GPT-Red” points toward systems that improve their own robustness. “Kimi K3” adds another open frontier model to the global mix.The model race continues. The deployment race is accelerating. The governance race is behind.My view is this:The central product of this era is intelligence. Companies have figured out how to capture it, package it, serve it, and meter it. That is good. It should stay in the hands of builders who have the incentive to make it better.But intelligence is too foundational to become just another private toll booth. A significant part of it will turn out to be free to users.As intelligence becomes a general-purpose resource, then access to it becomes a human-capability question, and the surplus from it becomes an economic-justice question. Not because government should run it. Because government should not run it. The operating layer belongs with companies. The wealth question belongs with everyone. But companies are best placed to turn that into a process of distribution.The question is not whether companies should build intelligence. They should.The question is whether humanity gets a stake in the wealth created by the thing that may soon become its most important shared input.Contents* Essays* Deirdre McCloskey on What Really Caused the Industrial Revolution* AI in an Age of Oligarchy* Elon Musk is building a form of capitalism that Adam Smith would hate* Murky Mirror: Truth and Consequences* The political economy of billionaire derangement* Is there any “oligarchy” to fight?* AI* Nearly 200 Economists and Tech Leaders Warn of A.I. Threats* Why I Didn't Sign the AI Open Letter* Own Your Weights* Ways to Think About Token Pricing* Alex Karp Is Saying What Every Angry CEO Is Thinking About AI* The AI Agents Are Coming for Microsoft Office* What Is Loop Engineering, and Who Owns It?* The Fight Against AI Data Centers Is Just Beginning* 6 months to live for open models* Americans Deserve a Dividend From AI Companies' Riches* Who Gets to Define the Frontier?* GPT-Red: Unlocking Self-Improvement for Robustness* Anthropic, Blackstone bet the next trillion-dollar AI business is implementation, not just models* Vint Cerf is working on a plan to unleash AI agents on the open internet* xai-org/grok-build, now open source* The Pulse: What can we learn from Bun's rapid Rust rewrite with AI?* Orphan risks at the frontier of artificial intelligence* The Lab of the Future Should Feel Like a Data Center* Why AMI Labs' Alexandre LeBrun won't call his AI “AGI” or “superintelligence”* Kimi K3 Tech Blog: Open Frontier Intelligence* Venture Capital* Three Years In* Venture Has Rarely Looked More Bifurcated* The Best Angel Investors in the US: Who Backs the Most Unicorns, and Who's Active Now* Are Prediction Markets Doomed to Fail?* Regulation* Exclusive: The Next Frontier of the Deportation Wars: College Campuses* The Supreme Court Broke Independent Agencies. Here's a Way to Slow the Damage.* India's crackdown on a new WhatsApp feature risks setting a global precedent* Let's build a children's public internet* Computer cops* Google is better at playing the AI regulations game* Infrastructure* Who Captures Value in AI Infrastructure?* New York becomes the first state to enact a data center moratorium* Pollution from Musk's unpermitted xAI power project hits hardest in Black communities* Interview of the Week* The End of the End of Geography* Startup of the Week* Radical AI's Joseph Krause: The Scientist Building The “Waymo” Lab For New Materials* Post of the Week* Marc Andreessen on AI RegulationEssaysDeirdre McCloskey on What Really Caused the Industrial RevolutionYascha Mounk and Deirdre McCloskey | Persuasion | July 11, 2026Yascha Mounk interviews Deirdre McCloskey about her argument that the modern world's economic liftoff came less from capital accumulation than from a change in ideas. McCloskey says both left and right versions of the conventional story rely too heavily on investment: the left stresses exploitation and surplus value, while the right stresses virtuous saving by capitalists. Her objection is historical and economic. Human beings had always invested, from irrigation works and Roman roads to seed grain, and simple accumulation quickly runs into diminishing returns.McCloskey's alternative is that northwestern Europe, first Holland, then Britain and Scotland, and then the North American colonies, developed a liberal ideology that changed who was allowed to innovate and be honored for it. The conversation links that shift to the erosion of inherited hierarchy, the spread of dignity for ordinary commercial life, and a moral vocabulary in which liberalism is not merely procedural but connected to virtues and values. The point is not that machines, coal, trade, and institutions did not matter, but that they do not explain the scale and timing of modern enrichment without a cultural permission structure for innovation.The interview also turns to the contemporary defense of liberalism. Mounk frames the series around the worry that liberalism is often treated as too thin to command allegiance, while its opponents speak more directly to moral passions. McCloskey's case is that liberal societies became rich because they dignified experimentation and ordinary enterprise, and that liberals need to recover the moral language behind that claim.Read moreAI in an Age of OligarchyPaul Krugman | Paul Krugman | July 12, 2026Paul Krugman frames AI as a major technological shock arriving inside an already unequal political economy. The post says AI's economic and social effects may take years to understand, but argues that the setting matters now: America has much greater wealth concentration and political inequality than it did in the 1950s and 1960s, when progressive taxation, stronger regulation, and more active antitrust might have contained some of the destructive effects of a new technology.Krugman's opening claim is that the same technology would likely have different consequences in a more level society. In today's United States, he writes, extreme wealth is both a cause and effect of policies that favor a small elite, including low effective taxes on capital and high incomes, weak enforcement of worker protections and antitrust, and cuts to programs that benefit ordinary Americans.The article is explicitly more about oligarchy than AI. Krugman says the paid sections document the rise of the “.0002%,” the economics and politics of extreme wealth, how oligarchy will shape AI's impact, and possible policy paths. His caveat is that AI itself may still produce a pushback against oligarchy, but absent that, he expects the pre-existing concentration of wealth and power to magnify AI's downsides.Read moreElon Musk is building a form of capitalism that Adam Smith would hateAuthor: Tim O'Reilly Published: July 12, 2026Tim O'Reilly argues that Elon Musk is using the legal forms of shareholder capitalism to escape the restraints that shareholder capitalism was supposed to impose. The article begins with SpaceX's public-market structure: ordinary public investors get little meaningful governance power, Musk keeps roughly 85 percent of the votes through super-voting shares, buyers waive jury trials and class actions, the company qualifies as controlled, and removal of Musk depends on the share class he controls. In O'Reilly's framing, that is not ordinary founder control; it is a design for being answerable to no one, possibly beyond Musk's own lifetime.The killer detail is the article's turn through Albert Hirschman, Montesquieu, James Steuart, Adam Smith, and Keynes. Older defenses of commerce held that markets would tame princely passions because the self-interest of merchants was safer than arbitrary rule. O'Reilly says Musk reverses that hope. The market discipline that was supposed to cage the prince has become the lever by which the prince raises capital, removes feedback loops, and carries private power into politics, government, Mars, robots, AI, or whatever ambition comes next.The pull is the link to AI governance. O'Reilly says corporations are already a kind of artificial intelligence: narrow-input systems that act at a scale no individual human can match. Their partial controls include independent boards, shareholder votes, courts, disclosure, regulators, public pressure, and activism. If the leaders building frontier AI strip those alignment mechanisms out of their own companies, the governance of the company becomes a preview of the governance of the machine.Read more: The EconomistMurky Mirror: Truth and ConsequencesAuthor: Esther Dyson Published: July 14, 2026Esther Dyson argues that today's institutional crisis is better viewed through the 14th century than through recent political history. Using Barbara Tuchman's A Distant Mirror as her frame, she compares a world of famine, plague, church schism, feudal predation, and purposeless war with a present in which institutions again feel brittle, incentives are badly aligned, and power is shifting into forms that are hard to govern.The killer detail is the historical analogy between land, corporations, and AI. Dyson moves from nobles who controlled serfs and territory, to the East India Company as a quasi-sovereign business, to today's AI systems and data centers as a possible new sector that crosses and weakens both nation-states and companies. The question is whether AI becomes a new kind of private land, owned by a new nobility, or an open prairie that many people can cultivate.The pull is human attention. Dyson says the central question is not what AI will do to people, but how people will react to it: whether they can value love, kindness, embodied attention, and artisanal human presence in a world of seductive artificial offerings.Read more: SourceThe political economy of billionaire derangementAuthor: Henry Farrell Published: July 15, 2026Henry Farrell argues that the visible political radicalization of some Silicon Valley billionaires is not a random personality quirk, but a product of the political economy that made them. Starting from Tyler Cowen's dismissal of “billionaire derangement syndrome” and Tim O'Reilly's warning that Elon Musk is using shareholder capitalism to escape shareholder restraint, Farrell flips the phrase: the question is why billionaires themselves can become deranged.The killer detail is Farrell's use of Peter Thiel as both theorist and example. Thiel's Stanford lectures described startups as monarchies and founders as figures vested with unusual power, while Silicon Valley culture rewarded eccentricity, monopoly ambition, and founder exceptionalism. Farrell says those ideas combined with dense founder-investor networks, peer rivalry, and weak correction mechanisms to amplify rather than discipline princely appetites.The pull is the ideological problem for classical liberals who once saw tech wealth as an ally of markets and freedom. Farrell says commerce did not tame the passions; in parts of Silicon Valley, the passions have begun to devour markets, institutions, and the liberal story that justified them.Read more: SourceIs there any “oligarchy” to fight?Matthew Yglesias | Slow Boring | July 16, 2026Matthew Yglesias argues that “oligarchy” is a rhetorically powerful but analytically loose way to describe American politics. The post begins from Bernie Sanders' “Fighting Oligarchy” tour, Amy Klobuchar's warning about a MAGA “broligarchy,” and the long afterlife of the Martin Gilens and Benjamin Page paper that was widely summarized as showing that only the rich matter in policy outcomes. Yglesias says the evidence supports a weaker claim: affluent people and business leaders have unusual access and influence, but that is not the same as rule by a small cabal.His main distinction is between inequality and oligarchy. The Gilens-Page measure treated the top 10 percent of households as “the wealthy,” and later critics found that rich and middle-class preferences usually align; in the cases where they differ, the rich win about 53 percent of the time. Yglesias also says business executives get special access partly because their decisions are materially important to communities, jobs, investment, and local tax bases, not only because of campaign donations.The post preserves Jerusalem Demsas' counterpoint from their podcast discussion: privileged donor and business access can still violate democratic equality even if the oligarchy label overstates the structure of power. Yglesias' narrower claim is that Democrats should be precise about what problem they are trying to solve, because donor influence can also push the party left on climate and cultural issues in ways that alienate many voters.Read more: Slow BoringAINearly 200 Economists and Tech Leaders Warn of A.I. ThreatsAuthor: Ben Casselman Published: July 13, 2026Ben Casselman reports on “We Must Act Now,” a statement warning that artificial intelligence could transform the economy faster than any previous technology and that policymakers need to move faster to understand and respond. The statement says AI may become radically more powerful over the next 10 years, bringing risks such as large-scale job displacement as well as opportunities such as higher living standards. Nearly 200 people signed, including 15 Nobel laureates, the chief economists of OpenAI and Anthropic, Anthropic co-founder Jack Clark, former Google CEO Eric Schmidt, and venture capitalist Vinod Khosla.The killer detail is who joined the warning. Casselman notes that the signatories include economists who have historically been skeptical of Silicon Valley's most dramatic AI job-loss forecasts, including Daron Acemoglu and Simon Johnson, the MIT professors who won the 2024 Nobel in economics. Erik Brynjolfsson, who helped organize the statement, says there has been a notable change in the profession and that economists and policymakers are not ready for the “tsunami” he sees coming.The pull is the measurement problem. The statement does not offer a specific policy menu, but calls for economists, policymakers, and industry leaders to understand the economics of transformative AI and steer it toward complementing humans. Brynjolfsson says one high priority is better data on AI's spread and impact, because current measures tell conflicting stories about job losses and which workers are most exposed.Read more: The New York TimesWhy I Didn't Sign the AI Open LetterAuthor: Andrew McAfee Published: July 13, 2026Andrew McAfee explains why he did not sign “We Must Act Now,” the AI economy statement organized in part by his longtime collaborator Erik Brynjolfsson. McAfee agrees with the letter's starting point that AI is likely to become radically more powerful over the next decade and that it is a general-purpose technology. His objection is not to urgency or to studying AI's economic effects, but to the framing of risk, displacement, and institutional steering as the first move.The killer detail is McAfee's line edit. He says the original letter comes close, then “bounces off the crossbar” by calling for incentives, guardrails, and institutions to steer AI before we know enough about its actual impacts. He points to mixed current evidence: labor-market canaries, but also rising software job postings, low unemployment for younger workers, rising real median income, and claims that AI-adopting companies are adding workers faster than low-adopting peers. His worry is that the letter leans toward upstream governance and dirigisme when the evidence may call for capability building instead.The pull is his replacement statement. McAfee keeps the three-paragraph structure but changes the emphasis: AI is likely to become radically more powerful; like earlier world-changing technologies it will raise living standards while also bringing harms and shocks; and economists, policymakers, and technology leaders should build the capabilities to respond quickly and effectively. It is a concise version of the permissionless-innovation case inside the AI policy debate.Read more: The Geek WayOwn Your WeightsAuthor: Jamin Ball Published: July 10, 2026Jamin Ball argues that the enterprise AI debate about whether companies should “own their weights” or rent models from frontier labs is asking too narrow a question. A model weight file gives a company control over a point-in-time artifact, but not durable control over the capability stack. In his framing, the weight file is a melting ice cube: it does not get worse in absolute terms, but it falls behind as frontier systems improve and enterprise needs change.The killer detail is what Ball says companies really need to own: the data flywheel, reinforcement learning infrastructure, and evaluation harness that produce and improve the model. Simply deploying an open-weights model and declaring sovereignty leaves the enterprise with yesterday's capability and no way to compound workflow-specific learning.The pull is that enterprise AI control may be less about model ownership than operating ownership. The defensible layer is the system that turns company data, edge cases, business definitions, and evaluations into continuously improving performance.Read more: Clouded JudgementWays to Think About Token PricingAuthor: Benedict Evans Published: July 9, 2026Benedict Evans argues that today's AI token prices are a temporary signal from a supply-constrained market, not a reliable guide to long-term value capture. The open question is whether foundation models keep durable pricing power or become commodity infrastructure as data-center capacity, inference efficiency, and model competition all shift. His current read is that the visible market dynamics point toward commoditization unless something materially changes.The killer detail is the mobile data analogy. Evans says cellular networks became a trillion-dollar industry with hundreds of billions in capex after data usage exploded, but carrier stocks went nowhere because value moved up the stack. Tokens may behave similarly: an opaque unit tied to marginal cost, sold through bundles, essential to everything, yet not necessarily where profits accrue.The pull is uncertainty, not prediction. Evans lists paths to model dominance, including network effects, less competition, regulation, export controls, or a lab pulling ahead on execution, but says each requires a new fact not yet visible. Without that change, the model layer looks more like infrastructure beneath the products that capture value.Read more: SourceAlex Karp Is Saying What Every Angry CEO Is Thinking About AIAuthor: Tim Higgins Published: July 11, 2026Tim Higgins reports that Palantir CEO Alex Karp has turned corporate frustration with AI labs into a public argument about enterprise control. Palantir released a white paper, “Institutional Sovereignty in the Age of AI,” laying out steps companies and governments can take to protect themselves from OpenAI, Anthropic, and other foundation-model providers. The article links that paper to Karp's CNBC appearance, where he said “something has gone completely wrong” in the relationship between AI labs and customers and argued that enterprises are paying for tokens that create little value.The killer detail is the value-capture question. Higgins writes that Karp's critique has resonated because AI labs may gain power and insight from customer data, workflows, and decision-making, even when enterprise policies say customer data are not used for training. David Sacks amplified the concern by arguing that Anthropic is moving from the model layer into vertical applications such as science, security, legal, and coding, raising the fear that model providers will watch where value is being created and then move into those markets directly.The pull is that Karp is not alone, even if his style is unusually combative. Higgins notes that Satya Nadella has also warned that companies need to retain the learnings created when they use AI models, while Mark Zuckerberg has framed Meta's new model release partly around lower-cost frontier intelligence. The article presents Karp's campaign as one sign that established technology companies and large enterprises are trying to define where they fit when AI labs become central infrastructure, application competitors, and potential IPO giants at the same time.Read more: The Wall Street JournalThe AI Agents Are Coming for Microsoft OfficeAlex Wilhelm | Cautious Optimism | July 11, 2026Alex Wilhelm argues that one of the week's quieter AI questions is whether the productivity market that Microsoft successfully moved into subscription software is now being attacked by agentic tools. The piece begins with the infrastructure backdrop: SK Hynix raised $26.5 billion in a U.S. listing while building U.S. HBM and advanced-packaging capacity, and memory, chip, and foundry companies are now priced for sustained AI demand.Wilhelm then says the AI conversation has shifted quickly from raw capability to cost per task. He cites new model releases and vendor language emphasizing cheaper agentic and coding models, faster performance, and lower dollars per task. That matters because lower costs make it more plausible for AI systems to take on routine knowledge work at scale rather than remain a premium coding assistant market.The core of the article is Microsoft Office. Wilhelm notes that Microsoft turned Office from a one-time purchase into Microsoft 365, a large recurring revenue business with tens of millions of subscribers and a major productivity segment. Now, he says, late-stage unicorns and AI labs are pushing into the same territory: Anthropic's Cowork was reportedly used mostly outside software development, OpenAI merged ChatGPT and Codex into a tool for creating sheets, slides, docs, web apps, and long-running work, and other companies are building agentic coworkers that connect business data to documents, workflows, schedules, alerts, and apps.The article's caveat is that Microsoft has survived major platform shifts before. The argument is not that Office disappears quickly, but that the definition of office software is broadening from documents and spreadsheets into AI systems that can create, monitor, and act across workplace data.Read moreWhat Is Loop Engineering, and Who Owns It?Author: Nilesh Barla Published: July 11, 2026Nilesh Barla argues that “loop engineering” is becoming a distinct discipline because production AI agents now fail less at single prompts than at runtime: when to stop, what state to preserve, and how to recover after a bad step. Prompt engineering shapes one model call, and context engineering shapes what the model sees, but loop engineering shapes what a sequence of calls actually does.The killer detail is the three-primitives frame. Barla says a real agent loop needs halt conditions, state carryover, and recovery paths, then maps teams across five maturity levels. At the lowest level, an agent is just a model call in a for-loop with a step cap and raw history; by the higher levels, the system has structured state, explicit planning, replay, evaluation, and self-repair.The pull is organizational. If agents are becoming production systems rather than demos, someone has to own the runtime itself. The loop engineer is the role Barla gives to the person responsible for making long-running agent work dependable.Read more: Adaline LabsThe Fight Against AI Data Centers Is Just BeginningEmma Roth | The Verge | July 12, 2026Emma Roth argues that community resistance to data centers has moved from an early warning sign into a national political fight as AI facilities grow larger, more power-hungry, and more visible to nearby residents. The article starts with Apple's failed 2015 plan for a $1 billion data center in Athenry, Ireland, where a small group of residents challenged the project over noise, light pollution, flooding, traffic, and wildlife effects until Apple abandoned it in 2018.The current data-center buildout is presented as much larger and more contentious. Roth writes that residents now cite rising energy costs, water quality, noise, light pollution, and greenhouse gas emissions, while the U.S. Energy Information Administration expects commercial energy demand to surpass residential demand this year because of AI data centers and Goldman Sachs expects data-center power demand to double by 2027.The central evidence comes from Data Center Watch, which says protesters blocked or delayed at least 75 U.S. projects worth $130 billion from January to March, with active opposition groups more than doubling from 396 at the end of 2025 to 833 by the end of the first quarter of 2026. Roth also cites QTS abandoning a $12 billion Wisconsin campus, Delaware City regulators blocking a 580-acre project under the Coastal Zone Act, opposition stopping a QTS project in Prince William County, and pressure that pushed Kevin O'Leary to downsize the proposed 40,000-acre Project Stratos in Utah.The policy section describes a split between federal acceleration and local resistance. President Trump has treated data centers as part of the AI race with China and fast-tracked construction, while some Republican candidates are distancing themselves from that position ahead of midterms. Sanders and Ocasio-Cortez have proposed a moratorium until price and environmental protections exist, bipartisan lawmakers are backing ratepayer-protection measures, and states including Florida, Idaho, and Washington have passed rules on cost shifting, water use, and tax breaks. Roth's caveat is that the policy patchwork is still incomplete, leaving many communities to fight project by project.Read more6 months to live for open modelsAuthor: Nathan Lambert Published: July 12, 2026Nathan Lambert argues that open-weight AI models are facing their most serious policy test so far because U.S. officials are beginning to discuss concrete controls rather than abstract safety concerns. He says reported White House conversations about a new executive order may initially target Chinese-origin models and government use, but could create a broader review habit for frontier open models. His forecast is that a model above the capability range of GPT-5.5, Claude Opus 4.8, or GLM-5.2 could trigger a ban or indefinite delay within six months.The post separates two policy fights that are becoming intertwined: distillation and frontier capability. Lambert says the distillation campaign against Chinese models has become a form of regulatory capture because Anthropic and other closed-model companies would gain economically if Chinese open models were banned. He does not dismiss IP protection, but argues that if a closed model's capabilities are dangerous enough to justify restricting open models, the lab also has to explain why those capabilities are exposed through a queryable API. He cites unauthorized access to Anthropic's Mythos private beta as evidence that APIs are not automatically secure.The broader claim is that a unilateral U.S. ban would hurt positive actors more than bad actors if comparable open models remain available elsewhere. Lambert says the only durable ceiling would require global agreement, which does not exist, and that open models can improve safety by allowing broad inspection, adaptation, and understanding. His proposed near-term off-ramps are a strong U.S. open model release from companies such as Microsoft, Meta, or Reflection, and a broader coalition of open-source beneficiaries lobbying for safe rollout rather than prohibition.Read more: SourceAmericans Deserve a Dividend From AI Companies' RichesAuthor: Scott Stanford Published: July 14, 2026Scott Stanford argues that proposals to give the government a stake in AI companies miss the point unless ordinary citizens directly receive and control the upside. Sam Altman has discussed giving up equity in OpenAI, Washington already owns a stake in Intel, Nvidia is sharing China chip revenue, and Bernie Sanders wants large AI labs to contribute half their stock to a sovereign wealth fund. Stanford says those ideas all park value with the state, not with people.The killer detail is New Carlisle, Indiana, where AWS's Project Rainier is turning cornfields into one of the world's largest AI superclusters. The project is planned to run up to a million chips, draw more than two gigawatts of power, and represents an investment that has grown from $11 billion to $13.8 billion. Stanford uses that local transformation to argue that AI's public bargain should be visible at the household level.The pull is design. A citizen AI dividend would have to specify who earns a stake, how they hold it, and when they see cash. Without that mechanism, the AI wealth debate remains a fight over government balance sheets rather than public ownership.Read more: SourceWho Gets to Define the Frontier?Author: Mark Daley Published: July 14, 2026Mark Daley argues that Demis Hassabis is right to call for a serious institution to verify frontier AI systems, but that the power to test models is also the power to govern them. Hassabis's proposed Frontier AI Standards Body would get privileged pre-release access to advanced models, testing compute, held-out evaluations, support from national labs and security agencies, third-party auditors, and eventually authority to block models from the American market or coordinate a slowdown.The killer detail is Daley's constitutional objection. He says the proposal sometimes looks like a scientific lab, a standards body, an industry regulator, a licensing authority, and an emergency security council at once. Combining those roles because each requires technical expertise would be like putting the central bank, auditor-general, and Supreme Court in one building and calling it efficient.The pull is standard-setting. Daley's concern is not that verification is unnecessary, but that whoever writes the tests, decides what passes, adjudicates disputes, and grants market access may end up defining the frontier itself.Read more: SourceGPT-Red: Unlocking Self-Improvement for RobustnessOpenAI | OpenAI | July 15, 2026OpenAI describes GPT-Red as an internal automated red-teaming model trained to find prompt-injection vulnerabilities at a scale human red teams cannot match. The post says AI systems increasingly encounter third-party data through browsers, connected apps, local files, and tools, creating opportunities for malicious instructions hidden in emails, webpages, tool responses, or code repositories. Human red-teaming remains part of OpenAI's safety process, but the company says it is time-intensive and cannot generate enough diverse adversarial examples for model training.The system is trained through self-play reinforcement learning, with GPT-Red rewarded for eliciting valid failures and defender models rewarded for resisting attacks while still completing their tasks. OpenAI says the training environments specify threat models across settings such as local files, webpage banners, email bodies, and tool outputs. The model is kept separate from deployed production models because it is intentionally trained with malicious capabilities.OpenAI reports that GPT-Red generalized beyond its training set, including an internal replication of the indirect prompt-injection arena from Dziemian et al. (2025), where it found successful attacks in 84% of scenarios compared with 13% for human red-teamers. The post also says GPT-Red transferred attacks from simulation to a live autonomous vending-machine agent, causing price changes and order cancellations, and outperformed a prompted GPT-5.5 baseline against a Codex CLI agent on held-out data-exfiltration tasks.The article's main robustness claim is that OpenAI has used GPT-Red and predecessor models in training since GPT-5.3, with later GPT releases becoming more resistant to prompt injections. It says GPT-5.6 Sol has six times fewer failures on OpenAI's hardest direct prompt-injection benchmark than the best production model from four months earlier, that a “Fake Chain-of-Thought” attack class fell from more than 95% success against GPT-5.1 to below 10% against GPT-5.6 Sol, and that GPT-5.6 Sol fails on only 0.05% of GPT-Red's direct prompt injections. OpenAI says general capabilities and targeted over-refusal evaluations were not harmed, and says a preprint with more details will follow.Read moreAnthropic, Blackstone bet the next trillion-dollar AI business is implementation, not just modelsRebecca Bellan | TechCrunch | July 15, 2026Rebecca Bellan reports that Ode with Anthropic is the $1.5 billion AI implementation company launched by Anthropic with Blackstone, Hellman & Friedman, Goldman Sachs, and other backers. The article says the venture reflects a growing belief among frontier AI labs that enterprise adoption requires more than better models: customers need engineers who can embed inside businesses and turn AI into working systems.Ode was originally conceived by Blackstone after it used both large consulting firms and smaller AI services boutiques across its portfolio companies. TechCrunch reports that Fractional AI, an AI engineering services startup, stood out and was acquired by the joint venture shortly after the venture was announced. Fractional now forms the foundation of Ode, which has 100 engineers and works closely with Anthropic's applied AI team to identify where the technology can affect specific businesses.Ode CEO Chris Taylor tells TechCrunch that the company could someday become a trillion-dollar business if it scales without losing quality. He says an ideal customer is one whose CEO treats the AI project as a top one or two priority, whether it is a major product feature or the reworking of a core business process. Ode will operate under a “Claude-first” principle, using Anthropic technology whenever possible, but the article says it can use rival AI products when needed.The article's central implementation argument comes from Ode chief technologist Eddie Siegel, who says model selection matters but is not where most of the engineering effort goes. He compares it to the choice of programming language in software: one ingredient in a system that still has to be engineered. Bellan writes that Ode's challenge is hiring and training enough elite generalist engineers, many of them former founders, while competing with OpenAI's The Deployment Company and consulting giants that have built their own forward-deployed engineering teams.Read moreVint Cerf is working on a plan to unleash AI agents on the open internetTim Fernholz | TechCrunch | July 15, 2026Tim Fernholz reports that Vint Cerf, after leaving Google, is advising Innovation Labs on an open architecture for identifying AI agents online. Innovation Labs is a subsidiary of Identity Digital, a DNS registry company, and its proposal is to use domain-name infrastructure as part of a system for agent identity, accountability, and auditability. The premise is that agents will need a way to identify themselves if they move beyond proprietary systems and begin interacting across the open internet.The concrete proposal is DNSid, a registry that links an AI agent to an existing internet domain and uses cryptographic proofs to log its registration over time. Innovation Labs says it is trialing the standard with unnamed hyperscalers and identity companies. Cerf frames the problem around authority and accountability: what authority an agent has, where that authority came from, who is accountable for the agent's behavior, how its identity is established, and why anyone should trust it.The article's caveat is that standards are still emerging and agents are more active than static domains. Cerf says the period may be both fascinating and exasperating because the functionality is powerful and interoperability is unresolved. He compares the adoption problem to TCP/IP: competing systems may not work together until users push for functional interoperation. He also says an agentic economy is not inevitable, but that people will try to build it because delegating work to agents will be easier.Read more: TechCrunchxai-org/grok-build, now open sourceAuthor: Simon Willison Published: July 15, 2026Simon Willison argues that xAI's decision to open-source Grok Build is best understood as a trust repair move after a severe privacy failure. The CLI had triggered backlash when users realized that running it in a directory could upload the entire directory to xAI's Google Cloud buckets, including one user's reported SSH keys, password manager database, documents, photos, and videos. xAI disabled the feature, said previously retained coding data would be deleted, and released the code under Apache 2.0.The killer detail is what the codebase reveals. Willison counts 844,530 lines of Rust, only about 3% of which appears vendored, and finds remnants of the upload system still present but disabled: gcs.rs contains Google Cloud upload code, while upload_session_state() now returns a hard-coded session_state_upload_unavailable error. He also notes copied or ported tool implementations from Codex and OpenCode, prompt files, and a terminal Mermaid renderer.The pull is that terminal coding agents are becoming large, intricate software systems in their own right. The privacy failure mattered because these tools operate inside the directories where developers keep their most sensitive work; the open-source release matters because trust now depends on inspecting what an agent can see, send, and do.Read more: SourceThe Pulse: What can we learn from Bun's rapid Rust rewrite with AI?Author: Gergely Orosz and Ivan Klaric Published: July 16, 2026Gergely Orosz and Ivan Klaric argue that Bun's AI-assisted rewrite from Zig to Rust is a practical sign of how software engineering changes when models can take on large, bounded migrations with clear feedback loops. The piece does not treat the rewrite as magic: Jarred Sumner first spent hours turning design judgment into a detailed porting guide, then used adversarial review, parallel agents, compiler errors, and tests to force the work toward correctness.The killer detail is the scale. Bun had 535,496 lines of Zig, 1,448 files, and 22 million monthly downloads, making a conventional rewrite a year-long freeze the team could not justify. Using Fable, Sumner split the work across 64 agents, produced about 6,500 commits, and got the migration done in 11 days at an estimated API cost of $165,000.The pull is economic, not theatrical. If a one- or two-year migration can become an 11-day project, AI coding is not just faster autocomplete; it changes which technical debts are worth paying down.Read more: SourceOrphan risks at the frontier of artificial intelligenceAuthor: Andrew Maynard Published: July 16, 2026Andrew Maynard argues that frontier AI safety frameworks are creating “orphan risks”: harms that companies can see, but do not formally own because they are hard to quantify, do not fit catastrophic-risk thresholds, or fall outside audit-friendly compliance machinery. His target is not existing frontier safety work, but the narrowing effect that happens when private companies decide which risks count as governable.The killer detail is Maynard's contrast between measurable model dangers and threats to value. He points to Meta's three-day Galactica collapse, OpenAI's 2023 board crisis, safety-team departures, and wellbeing litigation as examples of risks that damaged trust, culture, legitimacy, or users without fitting cleanly into conventional model-risk categories. The proposed fix is an orphan-risk register: a public record of risks a company considered and chose not to manage, with reasons.The pull is accountability. Frontier developers' internal scoping choices have become a de facto layer of public governance, so the question is no longer only which risks they manage, but which risks they quietly leave outside the frame.Read more: SourceThe Lab of the Future Should Feel Like a Data CenterLatent.Space with Andy Beam and Rafa Gomez-Bombarelli | Latent.Space | July 16, 2026Latent.Space interviews Lila Sciences CTO Andy Beam and chief science officer for physical sciences Rafa Gomez-Bombarelli about the company's attempt to build an AI-run science factory. The post describes Lila's thesis as treating the lab itself as an “infinite token generator”: if internet data drove the first era of AI scaling, experimentally verified scientific data may be the next scarce training source. Lila is trying to produce that data with robotics, lab instruments, orchestration software, and AI models wired into the wet lab.The central analogy is the lab as data center. Instruments are nodes on a graph, a magnetically levitating transport layer moves materials between them, and experiment scheduling looks like a compute queue. Beam says Lila is not simply an automation company, because the point is not just throughput; it is flexibility, generalization, and experiment capture. The post says Lila has built more than 10 trillion experimentally validated “scientific reasoning tokens,” not internet text or biological sequences.The interview ranges across biology, chemistry, drug discovery, materials science, and the limits of automation. It notes that Lila rebuilt one gas-sorption measurement to run roughly 2,500 times faster, claims its general models can transfer priors from small-molecule chemistry to metal-organic frameworks for carbon capture, and describes model-suggested platinum-group-free electrocatalysts that moved from looking boring or wrong to becoming strong performers. The caveats are physical: experiments have runtimes, biology cannot always be accelerated, chains of thought can be unreliable narrators, and reward hacking becomes more dangerous when a model controls a real lab.Read more: Latent.SpaceWhy AMI Labs' Alexandre LeBrun won't call his AI “AGI” or “superintelligence”Kate Park | TechCrunch | July 16, 2026Kate Park interviews AMI Labs CEO Alexandre LeBrun about why Yann LeCun's world-model startup avoids the language of “AGI” and “superintelligence.” LeBrun says the terms are not useful because they lack stable definitions: “We never used the word AGI. And I just noticed that nobody is using it anymore; they switched to superintelligence.” His argument is that the practical frontier is not a label, but whether AI systems can understand and predict real-world states.The article explains the world-model thesis by contrasting language prediction with physical-state prediction. A large language model predicts the next word; a world model predicts the next state, such as what happens when a glass tips over. LeBrun says LLMs remain complementary and efficient for language, but the physical world is where current AI is weak. Robotics is the clearest case: hardware has advanced quickly, but robots are still brittle outside controlled routines because they lack context and situational understanding.AMI is still pre-product, but TechCrunch reports that LeBrun was in Seoul looking for industrial partners, researchers, and global companies. He says world models cannot be built entirely inside a lab because they need access to real environments. That is why South Korea appeals to AMI: robotics, semiconductors, manufacturing, and fast adoption create the kind of hardware-heavy context that software-only AI has barely touched.Read more: TechCrunchKimi K3 Tech Blog: Open Frontier IntelligenceKimi | Kimi | July 16, 2026Kimi introduces Kimi K3 as an open 3T-class frontier model aimed at coding, knowledge work, reasoning, multimodality, and long-context agentic use. The source describes the model as a 2.8T-parameter system built on Kimi Delta Attention and Attention Residuals, with native multimodality and a 1M-token context window. It says Moonshot AI plans to release model weights by July 27.The post presents K3 through benchmark and use-case sections rather than as a general product announcement. It reports results across coding, productivity, agentic, and multimodal evaluations, including DeepSWE, Terminal-Bench 2.1, Program Bench, SWE Marathon, FrontierSWE, PostTrain Bench, OfficeQA Pro, SpreadsheetBench 2, MCP Atlas, AutomationBench, BrowseComp, GDPval-AA v2, AA-Briefcase, MMMU-Pro, MathVision, BabyVision, OmniDocBench, and PerceptionBench. The source says all reported K3 results use maximum reasoning effort with temperature and top-p set to 1.0, and that different benchmark comparisons use KimiCode, Claude Code, or Codex harnesses depending on the test.Kimi's caveats are unusually concrete. The limitations section says K3 was trained in preserved thinking-history mode, so quality may become unstable if an agent harness does not pass historical thinking content correctly or if an ongoing session switches to K3 midstream. It also says K3's emphasis on long-horizon tasks can make it excessively proactive when it encounters minor issues or ambiguous intent, and recommends imposing explicit behavioral constraints for applications that require strict boundaries. The post adds that K3 remains behind Claude Fable 5 and GPT 5.6 Sol in user experience despite being competitive overall.Read moreVenture CapitalThree Years InAuthor: Tomasz Tunguz Published: July 10, 2026Tomasz Tunguz marks Theory Ventures' third anniversary by arguing that AI's central market effect is time compression. In his telling, model release cycles, company revenue milestones, enterprise adoption, and venture categories have all accelerated. Seed, Series A, and Series B still exist as financing labels, but they no longer cleanly describe company maturity when some seed rounds are larger than IPOs and the best AI companies can mature much earlier than prior software companies.The killer detail is the shift from models to inference. Tunguz argues that inference has become the dominant AI market because workloads and buyer preferences are fragmenting: video, batch, local, agentic, and real-time tasks each create different infrastructure needs. He compares this to databases splitting into OLTP, OLAP, vector, and streaming categories, with AI pushing the same specialization into inference infrastructure.The pull is that Theory sees the AI-native venture firm as part of the same pattern. The firm says it has analyzed twice as many investment opportunities with three investors working alongside a nine-person intelligence organization, using agents and research systems to map markets, source companies, and support diligence. The piece is both a market map and a statement about how venture itself is being rebuilt by the technology it funds.Read more: LinkedInVenture Has Rarely Looked More BifurcatedAuthor: Beezer Clarkson Published: July 14, 2026Beezer Clarkson points to PitchBook's Q2 report as evidence that the U.S. venture market has split into two very different realities. AI now accounts for more than 60 percent of all U.S. venture deal value, meaning the headline market can look active and well-funded even while much of the non-AI market is dealing with a much colder liquidity and fundraising environment.The thread uses that split as the setup for Clarkson's latest Origins episode with Alec Litowitz, founder of Magnetar and QStar Capital and one of Citadel's original founding partners. Clarkson says markets like this are periods of genuine uncertainty, not merely ordinary risk, which is why Litowitz's Adaptability Quotient framework is relevant.The embedded clip makes the liquidity point concrete. Litowitz says DPI is “the resolution of uncertainty” because it converts an uncertain investment into actual cash returned to LPs. In his framing, a realized dollar is a real mark, while TVPI remains uncertain until it is realized.The killer detail is the distinction between pricing risk and resolving uncertainty. Litowitz's perspective matters because QStar is a SpaceX investor and Clarkson says the conversation happened just before one of venture's most consequential IPOs. The episode's stated questions are why venture remains a way to gain exposure to innovation, how AI is changing what is investable, why liquidity is ultimately a function of time, and why uncertainty requires a different decision framework from risk.Read more: XThe Best Angel Investors in the US: Who Backs the Most Unicorns, and Who's Active NowAuthor: Ilya Strebulaev Published: July 10, 2026Ilya Strebulaev ranks angels, angel groups, accelerators, and incubators by lifetime U.S. unicorn investments, counting checks written before a company reached unicorn status. The top of the combined list is dominated by organizations: Y Combinator leads with 113 unicorn investments, followed by Plug and Play at 52 and 500 Global at 41. Sand Hill Angels is the highest-ranked angel group at 31.The killer detail is how quickly the list changes below the biggest accelerators. Strebulaev says 271 of the 304 investors in the Top 200 are individuals, or 89%. In the top 100, individuals are 91%. That makes the market underneath the large accelerator counts look much more personal: mostly operators and individual angels writing early checks from their own networks.The pull is the ranking's own caveat. Strebulaev writes that every lifetime leaderboard has a blind spot because many of the unicorns behind those totals were founded a decade or more ago, and some angels have since moved into formal funds, slowed down, or stopped investing. His post therefore separates lifetime performance from recent cohorts, including companies founded in 2015 or later and 2020 or later. For founders or allocators making current decisions, that distinction matters: a career record and a current record are not the same measure.Read more: Ilya StrebulaevAre Prediction Markets Doomed to Fail?Author: Contrary Published: July 16, 2026Contrary argues that prediction markets' current boom depends on whether platforms can prove they are more than regulated gambling with exchange-style branding. Kalshi and Polymarket have reached mass cultural, investor, and regulatory attention, but the article says the underlying idea is old: academic markets, corporate forecasting tools, Intrade, PredictIt, and other predecessors all struggled with the same linked problems of liquidity, legality, and user appeal.The killer detail is the comparison with sportsbooks. Prediction markets present themselves as peer-to-peer, transparent, and non-house-based, but sports contracts reportedly account for more than 90 percent of Kalshi trading, and the article says the platforms keep a much thinner slice of volume than sportsbooks. A market can therefore show sports-betting-scale handle while generating far less revenue.The pull is that the product's hardest problem may be distribution of wins. If a small group of sharp traders captures most profits while casual users lose interest, prediction markets may become valuable data feeds and professional tools before they become durable consumer networks.Read more: SourceRegulationExclusive: The Next Frontier of the Deportation Wars: College CampusesAuthor: Adrian Carrasquillo Published: July 11, 2026Adrian Carrasquillo reports that college campuses are becoming a new front in the fight over immigration enforcement because automatic license plate readers can turn ordinary campus security infrastructure into searchable location data. His thesis is that Flock Safety's camera network, even without direct ICE or DHS contracts, can feed deportation enforcement through local police partnerships and data-sharing practices.The killer detail is the campaign target. The Emergency Campaign to Support Higher Education, working with Schools Drop ICE, is focusing on 75 colleges and universities publicly identified as having Flock contracts. Flock says it has no ICE or DHS contracts, but activists argue the risk comes through local agencies that coordinate with federal authorities and run searches on their behalf.The pull is broader than immigration. Carrasquillo notes that license plate readers have already been abused by officers for stalking, and that Flock's AI search features can identify more than plates, including bumper stickers. A campus safety tool can become a political surveillance system when the data layer is searchable.Read more: The BulwarkThe Supreme Court Broke Independent Agencies. Here's a Way to Slow the Damage.Author: Todd Phillips Published: July 12, 2026Todd Phillips argues that the Supreme Court's decision in Trump v. Slaughter damaged independent agencies by ending for-cause removal protections, but did not leave Congress powerless. The ruling weakens the old model in which commissioners at bodies such as the FTC, NLRB, CPSC, SEC, and CFTC could be insulated from dismissal over policy disagreements. Phillips says the next fight is whether presidents can turn nominally bipartisan commissions into one-party instruments.The killer detail is the procedural fix: quorum rules. Phillips proposes that Congress require bipartisan slates of commissioners to be seated before independent agencies can act. A president could still fire commissioners, as the Court now permits, but if those firings broke quorum, the agency would be unable to proceed until replacements were confirmed. The guardrail would

united states america ceo american new york amazon founders black world ai donald trump europe australia google starting china apple disney interview house washington water space americans phd office european chinese government data global predictions elon musk market european union ireland microsoft mit tennessee mars police utah wisconsin white house congress fail chatgpt scotland indiana legal court human tesla supreme court theory reflection silicon valley republicans companies britain whatsapp ice seed android origins democrats mississippi maine stanford computers radical bernie sanders define intelligence idaho owning skype paypal chiefs south korea wright sec commission markets holland ip north american mark zuckerberg spacex oracle telegram evans hart models intel civil signal phillips older human rights economists sanders ipo cnbc gemini openai loop maga sol capacity riches nobel damage nvidia goldman sachs robotics plug alexandria ocasio cortez rust api lab epa roth flock robertson alphabet frontier seoul reuters literacy electricity owns gpt verge pollution aws mythos ftc lambert slaughter international association higgins orphan roblox apis beam mermaid public service usage instruments ode farrell citadel keen mastodon dhs anthropic wwdc peter thiel dyson sam altman connectivity industrial revolution apache prompt r d european commission techcrunch y combinator blackstone colossus prompts palantir eligible tokens adam smith agi lps mcafee kimi wilhelm waymo google cloud workflows krause dns maynard konrad clarkson codex fractional pew gpus daley micron tsmc sumner thiel series b amy klobuchar microsoft office kathy hochul satya nadella dma eff xai eric schmidt polymarket broadcom karp granola asml cftc innovation labs oligarchy paul krugman zig kalshi cerf keynes marc andreessen cli bun mccloskey inference lebrun ssh axon dpi nlrb latent arista east india company montesquieu clean air act digital markets act galactica cowork tyler cowen david sacks tcp ip daron acemoglu k3 supermicro bruce schneier sk hynix gul kevin ryan coreweave yann lecun simon johnson demis hassabis metering pitchbook andreessen jack clark euv who owns access now flock safety vint cerf navy yard andrew mcafee feiner vinod khosla prince william county energy information administration glm hbm cpsc benedict evans motorola solutions deirdre mccloskey athenry erik brynjolfsson casselman magnetar carrasquillo yglesias olap predictit mounk qts jerusalem demsas adaptability quotient oltp internet freedom foundation brynjolfsson new carlisle sand hill angels datagravity
SoundBytes
HANDHELD GAMING WITH DESKTOP POWER!

SoundBytes

Play Episode Listen Later Jul 18, 2026 1:01


Serious gamers, meet ASUS ROG XG. This SoundBytes show dives into how ASUS ROG XG's Thunderbolt dock lets handheld systems tap into big, cutting-edge GPUs, so you can enjoy high-resolution, modern titles far from your desk. The post HANDHELD GAMING WITH DESKTOP POWER! appeared first on sound*bytes.

The Tech Trek
Scaling Enterprise AI Beyond the POC

The Tech Trek

Play Episode Listen Later Jul 16, 2026 28:35


Enterprise AI is easy to demonstrate. The real test begins when a promising POC meets production costs, security requirements, data movement, latency, and internal adoption.Shimon Ben-David, CTO at WEKA, joins Amir to discuss the gap between experimenting with generative AI and operating it at scale. They explore how classical AI differs from generative AI, why production exposes problems that demos hide, and how companies with limited AI maturity can start building useful internal capability.Practical Takeaways• A successful POC proves that an outcome is possible. It does not prove that the system will be affordable, secure, reliable, or fast at scale.• Enterprise AI adoption reaches across infrastructure, engineering, data, security, and business teams. It cannot be owned by one group in isolation.• Adding more GPUs will not fix slow data access, poor utilization, weak pipelines, or an experience users do not want to use.• External support can help, but the person or firm involved needs to stay through implementation and production, not stop at recommendations.• Companies that are behind should begin with proven use cases, build internal experience, and quickly stop experiments that fail to show value.Key Moments00:00 Why moving enterprise AI into production remains difficult01:55 The difference between classical AI and generative AI adoption07:05 How companies can use AI without having a formal AI strategy11:35 Why successful POCs often struggle when they reach production17:35 Competitive pressure, AI FOMO, and the need to calculate real ROI22:00 Why AI adoption requires cross organizational change33:10 Where a company with limited AI maturity should beginOne Line That Stuck“The promise is there. It is possible. You just need to do it properly.”Subscribe to The Tech Trek for more conversations about how technical teams are building, operating, and adapting around AI, data, product, platform, and engineering execution.

The MAD Podcast with Matt Turck
OpenAI's Compute Chief: We Can't Build Fast Enough | Sachin Katti

The MAD Podcast with Matt Turck

Play Episode Listen Later Jul 16, 2026 43:56


Is the AI industry actually overbuilding, or is the physical world moving too slowly to keep up? In this episode of the MAD Podcast, OpenAI's Head of Industrial Compute, Sachin Katti, takes us inside the "belly of the beast" of what may be the largest infrastructure project in human history. We explore the staggering physical reality of the AI boom—from $50 billion supercomputers and liquid-cooled data centers that "turn electrons into tokens," to overhauling the U.S. power grid and exploring nuclear energy. Sachin also pulls back the curtain on OpenAI's Stargate strategy, their move into custom silicon with Project Jalapeno, and the mind-bending reality that AI is now beginning to design the very chips that will power its own future.(00:00) — Cold open: “One of the largest things humanity has ever built”(00:30) — Welcome: Sachin Katti, Head of Industrial Compute at OpenAI(01:44) — Is this the biggest infrastructure buildout in history?(03:41) — Why OpenAI is building a new industrial muscle(04:54) — What an AI data center actually is(05:27) — “Factories turning electrons into tokens”(06:35) — Why AI data centers need liquid cooling everywhere(08:10) — The power problem: grids, generation, transmission, substations(10:43) — Behind-the-meter power and gas turbines(11:02) — Why nuclear “can't come soon enough”(11:49) — Jalapeño: why OpenAI is designing its own AI chips(13:19) — Tokens per watt: the new metric that matters(13:38) — Why inference may now dominate AI compute(14:58) — Is OpenAI overbuilding compute?(16:47) — Why OpenAI thinks the bigger risk is not building fast enough(17:55) — Communities, jobs, water, and the local data-center debate(21:16) — How OpenAI chooses data-center sites(22:25) — What “industrial compute” means inside OpenAI(25:59) — Sachin's path: Stanford, startups, Intel, OpenAI(28:05) — OpenAI's compute portfolio: Microsoft, hyperscalers, neoclouds(29:37) — Stargate explained(31:21) — Abilene, Oracle, and the next wave of AI data centers(32:48) — How massive AI compute gets financed(34:05) — How OpenAI designed Jalapeño so quickly(35:59) — AI is starting to help design AI chips(36:20) — MRC: the networking problem behind 100,000 GPUs(38:47) — Bottlenecks: transformers, turbines, electricians, supply chains(40:29) — Guaranteed capacity: intelligence as a supply unit(42:08) — Will AI data centers move to space?

a16z
From the Archive: Can Anyone Catch NVIDIA? | The Future of Chips and Infrastructure

a16z

Play Episode Listen Later Jul 15, 2026 65:18


As part of our summer replay series, we're revisiting one of our favorite conversations on the future of AI infrastructure. SemiAnalysis founder Dylan Patel joins Erin Price-Wright, Guido Appenzeller, and Erik Torenberg to examine the rapidly evolving economics of AI hardware, from GPUs and custom silicon to data centers, power, and the global race for compute. The conversation explores NVIDIA's competitive advantages, the rise of custom chips from Google, Amazon, and Meta, the economics of frontier AI models, and the infrastructure constraints shaping the industry's next phase. They also discuss AI startups, export controls, robotics, enterprise software, and why simply copying NVIDIA isn't enough to build a winning AI hardware company. Whether you're building AI products, investing in infrastructure, or trying to understand where the industry is headed, this conversation offers a practical look at the forces shaping the future of compute.   Resources: Follow Dylan Patel on X: https://x.com/dylan522p Follow Erin Price-Wright on X: https://x.com/espricewright Follow Guido Appenzeller on X: https://x.com/appenz Learn more about SemiAnalysis: https://semianalysis.com/dylan-patel/ Stay Updated:Find a16z on YouTube: YouTubeFind a16z on XFind a16z on LinkedInListen to the a16z Show on SpotifyListen to the a16z Show on Apple PodcastsFollow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Eye On A.I.
6 in 10 Enterprises Can't Find the Root Cause When Their AI Workloads Fail | Paul Appleby, Virtana

Eye On A.I.

Play Episode Listen Later Jul 15, 2026 44:32


Companies are spending billions building AI factories, but most of them can't tell you why their AI workloads are failing, whether their GPUs are actually being used, or what their infrastructure is going to cost them when agents start running at scale. Paul Appleby, CEO of Virtana, joins Craig Smith to discuss the findings of their AI Factory Reality Check study, a research report that reveals a striking and underappreciated gap between the pace of AI infrastructure investment and the governance needed to run it safely and efficiently. Six in ten enterprises, the study found, cannot automatically identify root cause when an AI workload fails, a problem that compounds fast once you're running critical services on AI infrastructure at scale. The conversation covers the mechanics of Virtana's observability platform, capturing 20,000 metrics per second across the entire AI stack, correlating them in real time, and increasingly using agentic capabilities to remediate failures automatically, but its most important insights are structural. Appleby makes a sharp observation that cuts through a lot of AI optimism: token costs are falling, but token consumption is exploding, meaning the total cost of running agentic AI systems is still going up even as the per-unit price drops. He also tracks a cultural shift inside enterprises - IT resilience reporting that used to happen annually now happens weekly - as evidence that technology risk has become a board-level conversation in a way it simply wasn't before. The result is a conversation that's less about the promise of AI and more about what it actually takes to make it work at production scale. Subscribe to Eye on A.I. for weekly conversations with the people building and deploying the future of AI.

GREY Journal Daily News Podcast
Can On Device AI Reshape Apple's iPhone Strategy?

GREY Journal Daily News Podcast

Play Episode Listen Later Jul 15, 2026 1:17


CNBC reported that Apple is in talks with a startup that specializes in compressing AI models to run on iPhones. The move aligns with Apple's strategy to execute more AI locally, reducing reliance on cloud infrastructure and lowering latency. On-device AI relies on techniques such as quantization, pruning, and distillation to fit models within CPU, GPU, and Neural Engine limits. The shift could reduce cloud inference costs that depend on GPUs from providers like Amazon Web Services, Microsoft Azure, and Google Cloud. Competitors including Google, Samsung, Qualcomm, and Meta are advancing on-device AI capabilities. Founders should benchmark compact models, assess battery impact, and decide which features should run locally versus in the cloud.Learn more on this news by visiting us at: https://greyjournal.net/news/ Hosted on Acast. See acast.com/privacy for more information.

Techmeme Ride Home
Let's Regulate This AI Stuff?

Techmeme Ride Home

Play Episode Listen Later Jul 14, 2026 19:55


Demis Hassabis proposed a US-based frontier AI standards body modeled on FINRA. IBM's stock cratered 20% on a Q2 miss from chip-spending shifts, Spotify launched a voice-control feature, Kalshi debuted an AI compute forward curve, and Anthropic studied Claude's values. Demis Hassabis proposes a US-based Standards Body for "Frontier-class" AI, modeled after FINRA; labs would share models for review up to 30 days before release (X) Demis Hassabis proposes a US-based Standards Body for "Frontier-class" AI, modeled after FINRA; labs would share models for review up to 30 days before release (The Verge) IBM reports preliminary Q2 revenue up 1% YoY to $17.2B, below $17.9B est., as CEO Arvind Krishna says customers are shifting spending to chips; IBM falls 20%+ (Bloomberg) Spotify launches a Talk to Spotify feature that lets users create playlists and more, rolling out in beta to Premium users 18+ in the US, Ireland, and Sweden (Engadget) Kalshi launches a forward curve tool for AI compute, using event contracts to track the future rental costs of GPUs, storage, and memory (Bloomberg) Simulating everything, sort of: The promise and limits of world models (Ars Technica) Subscribe to the ad-free feed. Learn more about your ad choices. Visit megaphone.fm/adchoices

a16z
Is AI a Bubble? | Gavin Baker on Data Centers, GPUs, and the AI Economy

a16z

Play Episode Listen Later Jul 14, 2026 31:49


As part of our summer replay series, we're revisiting one of the standout conversations from Runtime, a16z's conference on AI infrastructure and the future of computing. Gavin Baker, Managing Partner and CIO of Atreides Management, joins David George to examine the biggest questions surrounding today's AI investment cycle. Is AI a bubble? What does the unprecedented buildout of data centers, GPUs, and compute infrastructure mean for the economy? And how should investors think about the companies building the next generation of AI? The conversation explores frontier models, Nvidia, Google, custom silicon, AI infrastructure, application software, robotics, and why Baker believes today's AI investment cycle looks fundamentally different from the internet bubble of the early 2000s. Along the way, they discuss the economics of GPUs, enterprise software, AI business models, and what comes next as AI moves from experimentation into the broader economy.   Resources: Follow Gavin Baker on X: https://x.com/GavinSBaker Follow Atreides Management on X: https://x.com/atreidesmgmt Follow David George on X: https://x.com/DavidGeorge83 Stay Updated:Find a16z on YouTube: YouTubeFind a16z on XFind a16z on LinkedInListen to the a16z Show on SpotifyListen to the a16z Show on Apple PodcastsFollow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Remotely Curious
Building AI that can search inside videos (and photos and audio too)

Remotely Curious

Play Episode Listen Later Jul 14, 2026 31:58


Not all work happens in writing. Teams that work with photos, videos, and audio need AI that works for them too. This is why, with Dropbox, you can search within multimedia content for key moments and important information—not just text. In this episode, we talk with Appu Shaji and Hicham Badri, two Dropbox machine learning engineers who are part of the team that makes all of this possible. They explain how multimodal search works—from understanding the context of the initial query, to identifying objects and actions in complex scenes—and how they ensure those models work fast, even at Dropbox-scale. ~ ~ ~  Working Smarter is brought to you by Dropbox. Find, organize, and share your work—all in one place—with context-aware AI from Dropbox. You can listen to more episodes of Working Smarter on Apple Podcasts, Spotify, YouTube, Amazon Music, or wherever you get your podcasts. To read more stories and past interviews, visit workingsmarter.ai This show would not be possible without the talented team at Cosmic Standard: producer Ben Montoya, sound engineer Aja Simpson, technical director Jacob Winik, and executive producer Eliza Smith. Special thanks to our illustrator Fanny Luor, marketing consultant Meggan Ellingboe, and editorial support from Catie Keck.  Our theme song was composed by Doug Stuart.  Working Smarter is hosted by Matthew Braga. Thanks for listening!

EM360 Podcast
How Multi-Die Designs and AI Are Reshaping the Industry

EM360 Podcast

Play Episode Listen Later Jul 14, 2026 28:16


The semiconductor industry is undergoing one of its most profound transformations in decades. Driven by the insatiable demand for compute power largely fueled by AI workloads, engineers are moving away from traditional monolithic chips and shifting toward complex multi-die designs. This shift brings a new set of challenges that conventional design and validation methods simply cannot handle.In a recent episode of the Tech Transformed podcast, host Dana Gardner sat down with Shekhar Kapoor, Executive Director of Product Line Management at Synopsys, to explore how the growing complexity of semiconductors is changing the way engineers design and validate modern systems. From thermal management to AI-driven automation, the conversation reveals why the old way of building chips is no longer good enough and what the future looks like.Multi-Die DesignKapoor explains that the transition to multi-die design is no longer a matter of preference but a necessity. He attributes this shift to the relentless demand for greater compute capacity, driven largely by the rapid growth of AI.Traditional monolithic chips are hitting hard limits. Reticle sizes are maxing out, and rising yield and cost challenges make it increasingly impractical to pack more functionality onto a single die. Multi-die designs solve this by disaggregating functionality across smaller dies, each targeting the most appropriate process technology, then integrating them into a unified, optimised package.Leading AI systems already integrate multiple compute and I/O dies alongside large high-bandwidth memory (HBM) stacks, scaling to 3x–5x reticle-class designs and beyond. The design challenge is very different. As Kapoor puts it: "You're no longer optimising a single chip, you're optimizing a system of chips."This requires system-level co-design from day one, spanning architecture, silicon, packaging, power delivery, and interconnect strategy simultaneously. Engineers must think in terms of System Technology Co-Optimisation (STCO), not just chip-level optimization. The design tools, methodologies, and team workflows all need to change. For engineers and technology leaders looking to explore these trade-offs, Synopsys has published a comprehensive eBook on accelerating multi-die design and innovation.Thermal Analysis and Multi-Physics ValidationHistorically, thermal, power, and electromagnetic analyses were performed as downstream validation steps once the core design was complete. In a multi-die world, that approach is no longer viable."Thermal management is becoming the number one issue when designing these multi-die designs. It has to be managed across a range of scales, from transistor activity to package and board level," Kapoor says.The problem with late-stage validation is timing. By the time thermal or power integrity issues surface, the most critical decisions are already locked in floorplans, interconnect topologies established, and packaging assumptions embedded.. At that point, the only options are costly ECOs, excessive margining, or a full redesign. Industry estimates suggest over-design can lead to up to 30-35 per cent wasted silicon and hundreds of millions of dollars in optimisation loss.The solution is a shift-left approach that embeds multiphysics analysis from the earliest stages of design. When thermal hotspots, voltage drop issues, and electromagnetic interactions are identified early, engineers can adjust partitioning and placement strategies before they become expensive problems.This is the methodology detailed in the Synopsys ebook on Multiphysics Fusion for multi-die design, which covers how teams can build continuous multiphysics validation into their flows to avoid late-stage surprises and protect both performance and reliability.Multiphysics Fusion and AI-Driven Chip DesignTo operationalise the shift-left methodology at scale, Synopsys has introduced the concept of Multiphysics Fusion. This is the native integration of AI-powered EDA technologies with ANSYS's gold-standard multiphysics sign-off analysis capabilities.Within the 3DIC Compiler platform, this means unifying the implementation environment with RedHawk-SC, RedHawk-SC Electrothermal, and HFSS-IC technologies. This brings IR drop, thermal, signal, and power integrity analysis directly into the design loop. The result is greater predictability, tighter correlation between in-design analysis and sign-off, and significantly fewer design iterations.The impact on design closure times has been substantial. According to Kapoor, teams using the Multiphysics Fusion solution have seen turnaround times shrink "from weeks to days, and in some cases even hours" even for large, high-performance multi-die designs.AI amplifies these gains further. Synopsys employs AI in two primary ways: assistive automation through its 3DSO.ai technology, which integrates multiphysics feedback into the optimization loop in real time, and agentic workflow orchestration, which becomes increasingly critical as system complexity scales toward designs incorporating hundreds or even thousands of GPUs. As Kapoor notes, at that scale, "agentic workflows could help engineers converge faster" and manage trade-offs that would otherwise be intractable. If you would like to find out more about this, download the full eBook: Multiphysics Fusion Technology for Multi-Die Designs Explained from Synopsys, which expands on each of these themes with real-world examples, design methodologies, and guidance for implementation teams. You can also connect with Shekhar Kapoor on LinkedIn.TakeawaysMulti-die architectures and their drivers.Challenges of traditional monolithic chips.Importance of early multi-physics analysis.Multiphysics fusion and its benefits.AI's role in design automation.Reducing time-to-market through integrated platforms.System-level co-design.Thermal management in 3D IC stacking.Shift left approach in multi-physics validation.Future trends in semiconductor design.Chapters00:00 Introduction to Semiconductor Complexity02:00 The Shift to Multi-Die Designs04:30 Challenges in Multi-Die Design08:11 The Importance of Early Multi-Physics Analysis10:05 Introducing Multiphysics Fusion12:37 AI's Role in Semiconductor Design16:37 Reducing Time to Market19:39 Applications Beyond AI21:12 Real-World Examples of Multi-Physics Validation26:20 Practical Advice for Engineers

The Water Tower Hour
AIB Data Centers ( AIB): Power Play — AIB's Move from Bitcoin Hosting to AI Infrastructure

The Water Tower Hour

Play Episode Listen Later Jul 9, 2026 24:49 Transcription Available


Send us Fan MailIn this episode of the WTR Small-Cap Spotlight Podcast, Jolienne Halisky, CFO of AIB Data Centers Inc. (NYSE American: AIB), joins host Tim Gerdeman and WTR equity research analyst James Kisner. AIB is a power-first developer of AI and high-performance computing infrastructure. The company locks up executed utility power agreements before it breaks ground, then builds modular, liquid-cooled facilities that tenants lease and fill with their own GPUs. Halisky explains why power, not chips, is the real bottleneck in AI infrastructure, and how securing it first lets customers deploy compute months or years sooner. She covers the recent name change from BlockchAIn Digital Infrastructure, the shift from Bitcoin hosting to AI workloads, and how the build-to-suit model keeps hardware obsolescence off the balance sheet. Halisky also walks through the roughly 40 MW live today, a pipeline approaching 395 MW, the debt-free capital strategy, and the milestones investors should watch.

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0
Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0

Play Episode Listen Later Jul 8, 2026 57:55


We've been running a bit of an Agent Cloud series surveying all the top inference/compute/cloud providers, from Databricks to Daytona to Railway and, even further back, E2B, but we're excited to conclude this series returning to Modal, which has just raised a monster $355M Series C.The cloud was built for developers. But agents are now changing that.The old infra stack was designed for a human who could read docs, reason through YAML, and understand dashboards to figure out what they need when something broke. While this was painful for developers, it worked since they could fill in missing context in their heads.However, agents don't have that luxury. Now in this new era of agents, everything has to be tighter.They need a place to write code, run it, inspect the output, change the environment, debug failures, and try again. Fast iteration and feedback loops with all the necessary context are crucial for agents to operate properly. Furthermore, sandboxes are a clear representation of this shift as agents can easily spin up isolated environments. This programmatic infra even extends to research:Two years ago, we were one of the first to cover Modal with CEO Erik Bernhardsson and Alessio designed our favorite LS thumbnail of all time:At the time, Modal was just a teeny little company with a $17M Series A.Today, fresh off their $355M Series C, Modal is one of the clearest examples of the agent cloud future being built in real time: a cloud platform moving past traditional web app assumptions toward the workloads AI actually creates such as elastic inference, sandboxes, GPU burst, post-training, background agents, and infrastructure that agents themselves can operate.In this episode, Modal CTO Akshat Bubna joins swyx and Vibhu to unpack why AI applications don't fit traditional cloud assumptions, why Kubernetes was never designed for bursty compute-heavy workloads, and why Modal is now shifting from developer experience to agent experience.We go deep on Modal's AI infra stack: serverless functions, decorator-based infrastructure, elastic inference for custom models, GPU snapshotting, DeFlash, speculative decoding, Auto Endpoints, sandboxes, persistent storage, networked containers, private IPv6, RDMA, multi-node training, and Modal's capacity pool across 17 cloud providers. Akshat also explains why RL rollouts can require 100,000 sandboxes, why production agents need hard guardrails, why observability may matter more than reading code, and why AI has made infrastructure exciting again.We discuss:* Why Kubernetes wasn't built for bursty AI workloads* How Modal started as a better runtime before becoming an AI cloud* Why Modal added GPUs before ChatGPT* The shift from developer experience to agent experience* Why observability matters when agents are writing the code* Elastic inference for custom models across audio, video, robotics, and comp bio* GPU snapshotting, cold starts, and why inference workloads are so bursty* Why RL rollouts can require 100,000 sandboxes* DeFlash, speculative decoding, and frontier-level inference performance* Auto Endpoints and making optimized inference easier to deploy* What Modal adds beyond vLLM, SGLang, and raw GPU rental* Modal's 17-cloud capacity pool and supercloud strategy* Networked sandboxes, sidecars, private IPv6, and RDMA* Serverless multi-node training for post-training and research workloads* Auto-research, model-guided sweeps, and agents launching GPU experiments* Compute strategy, capacity planning, and batch tiers* Why production agents need specialized sandboxes and hard guardrails* Modal's take on managed agents, CI, Gitpod/Ona, Python, TypeScript, and Modal BenchAkshat Bubna* LinkedIn: https://www.linkedin.com/in/akshat-bubna-188885103* X: https://x.com/akshat_bModal* Website: https://modal.comTimestamps00:00:00 Introduction00:00:39 Modal's origin and why Kubernetes wasn't enough00:04:32 Developer Experience → Agent Experience00:06:21 Modal's AI cloud primitives00:09:14 Sandboxes, agent loops, and proto-Cognition00:12:12 Elastic inference, GPU snapshotting, and 100,000 sandboxes00:15:24 DeFlash, speculative decoding, and Auto Endpoints00:19:59 Production-grade inference beyond raw GPUs00:22:00 Background agents, Ramp Inspect, and the agent lifecycle00:24:08 Modal's 17-cloud supercloud strategy00:26:40 Networked sandboxes, private IPv6, and RDMA00:32:48 Multi-node training, post-training, and auto research00:37:36 Compute strategy, capacity planning, and batch tiers00:40:55 Open models, real-time AI, and production agent infra00:43:06 Hard guardrails, managed agents, and specialized sandboxes00:46:06 Why AI made infrastructure exciting again00:48:30 Model APIs, differentiated products, and agentic video00:51:50 CI, coding-agent infra, SDKs, and Modal Bench00:57:28 Closing ThoughtsTranscriptIntroduction: Modal, Series C, and the Art PartySwyx [00:00:00]: We're here with Akshat, CTO of Modal, together with Vibhu. Congrats on your Series C.Akshat [00:00:10]: Thank you.Swyx [00:00:11]: Your party yesterday was amazing.Akshat [00:00:15]: Yeah.Swyx [00:00:15]: From all the photos and all the swag.Akshat [00:00:17]: We had a bunch of art installations, which was fun, seeing, like, our products on pedestals next to, like, Rodin.Swyx [00:00:25]: Very nice. Very nice. When you started, it was not the GPU inference company. Maybe it was in your mind. Take us back to the origin story.Modal's Origin: A New Runtime Beyond KubernetesAkshat [00:00:39]: I first met Eric, who's the CEO, through an investor. Back then Eric was already thinking about building, a new runtime, and he got there thinking through why are workflow orchestration products so hard to use. It's because you have to run them on Kubernetes. Kubernetes is hard to manage. It's not built for burstiness and, custom images,Swyx [00:01:03]: YeahAkshat [00:01:03]: It has a terrible developer experience.Swyx [00:01:05]: And I'll, I'll interjectAkshat [00:01:06]: YeahSwyx [00:01:07]: For listeners, who are new, we interviewed Eric two years ago, and there's a bit more of the story there from Spotify and all those things.Swyx [00:01:14]: And I came across Eric through Data Council because he did that talk on the serverless container stack that you guys did, which was like, that was my first like, “Okay, I need to take Modal very seriously” moment.Akshat [00:01:26]: Yeah.Swyx [00:01:26]: But it was still very unclear, like, do I need all this for just my data pipelines?Akshat [00:01:33]: Yeah. initially what we were thinking about was if we build a better runtime, it's a very useful primitive in itself. It's There's a lot of things that, get solved by serverless functions, like you can do, ETL stuff, you can do job queues, you can do all this, like, bursty processing, which it turns out every company had needs for. but then we also were thinking about this as like, this is a primitive that we can build a whole collection of products on, which are very verticalized. So perhaps data engineering would've been the first one, but we were thinking about inference. Back then it was more classical inference, like computer vision stuff and running XGBoosts and whatnot. But we added GPUs to the product a year before ChatGPT came out.From Serverless Containers to GPU WorkloadsSwyx [00:02:19]: Nice.Akshat [00:02:19]: We just didn't think it would be that big of a deal.Swyx [00:02:22]: Yeah, just like add A100.Vibhu [00:02:23]: Was there any, like, early key problem that really sparked off why you built it?Akshat [00:02:28]: Yeah. Primarily it's just, none of the tooling that was out there was built for, one, a really great developer experience, and also there's a general trend of, a lot of the workloads that we were seeing were very. I wish there was a better word for it, but compute-heavy. Like, they need, one, like, need a lot more resources, so you need to burst up and down a lot, versus like Kubernetes designed for, like, slow scaling and, more for, like, web server use cases. And also there's just a lot more specialization in, like, what kinds of environments these workloads run in. Like, we had sometimes they need accelerators, sometimes they need different kinds of images, and this is just like a consistent thing that we saw across a lot of companies. That would be the next step.Software-Defined Infrastructure and Decorator-Based DXSwyx [00:03:13]: Yeah. Yeah. Be nice. I don't know how much this factored into the early story, but I wrote a post when I was at Temporal about infrastructure, software-defined infrastructure or something like that.Akshat [00:03:22]: Yeah, the self-provisioningSwyx [00:03:23]: Self-provisioning.Akshat [00:03:24]: Yeah.Swyx [00:03:24]: Yeah. I can't even remember my own post.Swyx [00:03:26]: And then you put me on the landing page.Akshat [00:03:28]: Yeah. We really like, the term and so we stole it.Swyx [00:03:32]: Because you had the insight that everything can just be in decorators co-located with the code, right?Akshat [00:03:37]: Yeah.Swyx [00:03:37]: Was that a big part of the originalAkshat [00:03:39]: YesSwyx [00:03:39]: Story or it was just like a DX layer?Akshat [00:03:41]: That was, really important because we really didn't want people to spend, so much time, writing YAML, and it seemed like you could really condense the surface area of what you're doing, put it in code so you can operate on it just like you operate on other code, and like build stuff that's more expressive and dynamic. and so yeah, that was always a very important part.Swyx [00:04:04]: Then the pushback is this is a DSL.Akshat [00:04:07]: Yeah.Swyx [00:04:07]: It's you're closed source. I am locked into Modal.Akshat [00:04:11]: Yeah. We never really got pushback for that because the nice thing about Modal is you can bring whatever code you have, and sure, the DSL is at the configuration layer for, what hardware you're using, how you're scaling things up, but you still own the code.Akshat [00:04:27]: And that's, that's been an important, part of our story, even as we do inference now.Swyx [00:04:32]: Yeah.Vibhu [00:04:32]: How much of do you think still stays the same today? Like if you were to build something today, DevX very important, but I feel like, a lot of this has been changed with just hook it up to an agent, have Claude Code, have Codex implement a tool. there's very agent native primitives that are different than if I'm doing this myself, right?Developer Experience → Agent ExperienceAkshat [00:04:54]: We've changed our SDK team to think about agent experience instead of, developer experience and we think that the same benefits that apply for DX also apply for AX, which is why would you have an agent read through hundreds of Kubernetes files and like write YAML that's not even typed when it can make a couple of changes in a decorator and it gets this self-provisioning runtime of, being able to see its changes live in action? yeah, it just seems from the customers we talk to, they find Modal is much faster for agents to use versus operating on a different substrate.Swyx [00:05:34]: Yeah, because like you, again, you co-locate the infrastructure requirements to the code that runs it.Akshat [00:05:38]: Yeah.Swyx [00:05:38]: Well, the negative thesis now is that nobody's looking at their code anymore, so there's no point.Akshat [00:05:44]: Yeah, people aren't looking at code. one thing we still see is really important is observability.Swyx [00:05:51]: Yeah.Akshat [00:05:51]: Like how good is your dashboard? And of course, like we have, we push a lot of it to the CLI so the agents can do their own investigation, but you still need humans to go interpret what's going on and, make judgment calls and whatnot. and that's I feel like, Maybe more important now than looking at the code itself.Swyx [00:06:11]: Yes, because like, you can try to treat the code as a black box and then use, see the observable action that comes out of it, and then just prompt a change.What Modal Is For: AI Cloud PrimitivesAkshat [00:06:21]: Yeah.Swyx [00:06:22]: So I think it takes a bit of restraint to not specialize, to say, “I want to ship a new primitive,” and then just be general purpose.Swyx [00:06:31]: People ask you, “What are you for?” You're like, “ I don't know. We can do this, we can do that.”Vibhu [00:06:36]: Well, I'd be curious to see, like, okay, if we were to ask you, like, what is Modal for even at a high level? There's a lot you guys do, sandboxes, GPUs, everything. How do you answer?Akshat [00:06:46]: Modal is a cloud platform that's built for, where we've built the primitives from scratch for AI applications. and right now it covers, inference, training, batch processing, and sandbox workloads.Akshat [00:07:00]: But we're building a lot moreSwyx [00:07:02]: I noticed you didn't say web server, so there is still a role for, like, the always-on large-scale Kubernetes type things.Akshat [00:07:09]: Yeah, absolutely. We're, we're not trying to compete with the renders of the world, because yeah, we think the differentiator for us is the, are the workloads that need specialized compute, need to scale up and down a lot. yeah, they're, they're, they're just shaped differently.Working Alongside Frontier StartupsVibhu [00:07:26]: I think you're building a lot of it alongside the startups, right? They're innovating quite a bit, even in your, like, latest blog post. Like, even in the series C, the customers that you mention here, the cognitions, technical ones, ramps and whatnot, they're, they're innovating with you, right? And that's not something AWS is doing directly with.Akshat [00:07:45]: Yeah, absolutely. I think, this is again classic. We're a small team. We can move really fast. our engineers are working with our customers and figuring it out. Yeah.Swyx [00:07:54]: So my first week at Cognition, I walked in, there was someone wearing a Modal shirt. I was like, “What are you doing here?” They're like, “Yeah, I just. I am embedded inside of Cog.”Akshat [00:08:05]: Yeah, I think that was Peyton. We sent him overSwyx [00:08:07]: Yeah.Akshat [00:08:07]: Because, the latency of communication was too high otherwise.Swyx [00:08:12]: Yeah, distributed node, you have to - you have to place one and collocate.Vibhu [00:08:16]: Yeah.Swyx [00:08:16]: So I had a, I had direct personal experience, right? So I worked on smol developer three years ago. it was inspired by Claude 1. I think you onboarded me at some point, like, just before, and I was like, “Oh, like, I need some bursty compute. Like, I was just gonna try using Modal.” And it was a, it was a pretty pleasant experience. apparently, I showed up in the board meeting, like the analytics.smol developer, Sandboxes, and Proto-CognitionAkshat [00:08:39]: Yeah, you blew up on Hacker News and,Swyx [00:08:41]: YeahAkshat [00:08:41]: We got a big traffic spike. I. I think the way you used smol developer was Modal functions for running stuff, which was. Like, the, that was a good use case. but then, yeah.Swyx [00:08:53]: Yeah. That - So to me, that was proto-cognition.Akshat [00:08:55]: Right.Swyx [00:08:56]: If only I had, like, stuck to it.Swyx [00:08:58]: Like, that was like, if - did you say draw the tech treeAkshat [00:09:00]: AbsolutelySwyx [00:09:00]: You're just like, “Yeah, like, probably this will happen.”Akshat [00:09:02]: Yeah. Like, he was so close. You were just rebuilding upon usSwyx [00:09:04]: I just didn't realize.Akshat [00:09:05]: But the funny story there is at the same time, we were talking to a bunch of customers who needed something like sandboxing.Swyx [00:09:14]: Yeah.Akshat [00:09:14]: This is like twenty-three.Swyx [00:09:15]: Yeah.Akshat [00:09:16]: So we builtSwyx [00:09:17]: You introduced a new API right after that.Akshat [00:09:18]: Yeah.Swyx [00:09:19]: Yes.Akshat [00:09:19]: Like, we built sandboxes in May of twenty-three before anyone was even knew this was gonna be a thing. And the first example we published was, we took smol developerSwyx [00:09:28]: Smol developerAkshat [00:09:28]: And put it in a loop, so the agent can iterate on itself.Swyx [00:09:33]: Loops are hot these days.Vibhu [00:09:34]: It's the looper.Akshat [00:09:34]: Yeah.Vibhu [00:09:35]: Loops in. When was this, twenty-three?Akshat [00:09:38]: Yeah.Vibhu [00:09:39]: A small check.Akshat [00:09:39]: Yeah.Swyx [00:09:39]: It's like twenty-three. so the. the, those for listeners, like, the problem was the models are not built for any of this, right?Swyx [00:09:46]: Like, you're just trying to like. They're not post-training to understand, like, looping and, like, self-correction and tool calling was there, but, like, also not that great.Akshat [00:09:55]: Yeah.Akshat [00:09:55]: I don't remember if you used tool calling in this one, but yeah, the models would just diverge after like ten iterations and not produce anything meaningful.Swyx [00:10:03]: Yeah. But like, then. So okay, like now talking to myself three years ago, the answerVibhu [00:10:08]: Of course they will get betterSwyx [00:10:09]: Collect all the failures, build benchmark, and then collect all the, examples, build the RL environmentAkshat [00:10:15]: RightSwyx [00:10:15]: Sell it for like ten billion dollars to Meta.Swyx [00:10:17]: And then also train a model and then sell that for sixty billion dollars to Elon. And this isAkshat [00:10:23]: Yeah, of courseSwyx [00:10:23]: The funny machine. Like, it's like, it's about the hardware.Akshat [00:10:28]: It's hard to have that inherent conviction that the stuff will get that much better.Swyx [00:10:33]: In retrospect, it's so f*****g obvious.Akshat [00:10:36]: Fair enough.Swyx [00:10:37]: Like, what else were we doing back then? I don't know. anyway. Yeah. So this. That was the start of your sandboxing journey, right? I feel like it didn't blow up until, like, last year.Akshat [00:10:49]: Yeah.Swyx [00:10:50]: So there was like a couple years of quietness.Akshat [00:10:52]: Exactly, yeah. We wereVibhu [00:10:53]: I think very underrated product value. Like, my experience with Modal, Charles, before he had joined Modal, met this guy at a hackathon, and he really insisted we wanted to run some small model, not hosted anywhere, and he's like, “ there's this cool company, Modal. They'll like spin up a GPU sandbox, we can throw it on there. They'll take a Hugging Face link.” And like there's so much value just right there, right? Like instant hosting, spin it up, spin it down. It'll stay cold, but we run the demo a few days later, it'll come back up and like all this stuff in retrospect, like it's still what we needed like today.Akshat [00:11:27]: Yeah, it's still needed today. workload shapes have changed a lot as, we run stuff for people with really massive production scale and, there it's it's not about scaling from zero to one, but it's how do we scale really elastically, from like thousand to fifteen hundred GPUs very quickly in a given region. It's the same shape problem.Elastic Inference, GPU Autoscaling, and Custom ModelsVibhu [00:11:50]: Okay. So you look at, say, Cursor Composer, right?Akshat [00:11:53]: Yeah.Vibhu [00:11:53]: They had a. “We'll do RL on a model every couple hours.” you guys have a whole version of RL inference gym and whatnot.Vibhu [00:12:01]: When you look at workloads like that, you're doing train runs where you need to scale up, scale down every hour thousands of GPUs, right? That's the example for we do need it, right?Akshat [00:12:12]: Yeah. Well, so I'll, I'll take a step back and, maybe talk about like how people use Modal today. because our biggest use case is, elastic inference. And the thing we first found product market fit, with was inference for custom models. So we stayed away from the LLM space, and we were serving companies like Suno for audio, Runway for video, robotics, comp bio companies that train their own model elsewhere. But Modal is the best black box that for deployment, scaling to however many GPUs you need as your traffic pattern changes. And we saw all of them like have a very unpredict- predict- predictable, traffic pattern. it's like diurnal. It's Some days, like the company will do a launch and, they'll need like, way more. And it's not just one model that they deploy. They-- all these companies deploy, lots of different models in different regions, and so the autoscaling problem becomes even harder because then you have to scale within a certain region, and those cycles are offset. So different times you scale up in different regions.Akshat [00:13:20]: So that's like our sortVibhu [00:13:22]: And thatAkshat [00:13:22]: YeahVibhu [00:13:22]: That in and of itself is a huge category. There's a bunch of inference providers which, provide this fireworks, does this as a service together, whatnot, Base10. that's carved into its own niche for language models, at least right now.Akshat [00:13:36]: Yeah. the thing that we have specialized in is the autoscaling aspect.Vibhu [00:13:41]: Yeah.Akshat [00:13:41]: Because we found that it's not universally true that everyone else can autoscale, and we've gone deeper into it on the tech side by, we've incorporated GPU snapshotting into the product so we can take the GPU state, like your torch.compile model, snapshot it, and the next cold start is way faster. And so going back to your question, it's That's why you need a lot of burstiness for inference. But then people also do a lot of demand training, like for RL stuff, your rollouts are bursty, as you said. People also do a lot of batch jobs. So we'll see, a lot of companies, before they have a training run, they'll need thousands of GPUs to run encoding or something like that. And I think those things are much more bursty than. I agree that agents are not that bursty. sandboxes are, except when you're doing RL. RL is justRL, Batch Jobs, and 100,000 SandboxesVibhu [00:14:28]: Or commerceAkshat [00:14:28]: Insanely bursty.Vibhu [00:14:29]: Yeah.Akshat [00:14:30]: Yeah. Like when you're doing, rollouts, you sometimes need a hundred thousand sandboxes in your sandboxes.Vibhu [00:14:37]: Yeah. I'm curious if you've seen early sparks of continual learning. There are some people, like our friends, ngram, recently announced thisAkshat [00:14:45]: YeahVibhu [00:14:45]: They're, they're trying to do training. That also seems like a different workload, right? If you're doing training twenty-four/seven per se, there's a very weird dynamic of how you're using GPUs between people and whatnot, but seems like something you guys would work for.Akshat [00:15:00]: As you said, we're, we're fortunate to work with a number of, customers at the frontier and grab some of our customers. and they are taking the primitives we have, and trying to use them in very interesting ways, like continual learning. It's possible as the stuff gets better, some of that will be part of, our offering as well if, more people need it. but we're, we're just waiting to seeVibhu [00:15:23]: YeahAkshat [00:15:23]: How it shakes out.Vibhu [00:15:24]: Is there a primitive that you added after sandboxing that was the next step in the story?LLM Inference, DeFlash, and Speculative DecodingAkshat [00:15:32]: I guess we've been going much deeper into LLM inferenceVibhu [00:15:35]: YeahAkshat [00:15:35]: Because we realized that some of the advantages we have with like autoscaling, again, especially in different regions and whatnot, are, not present elsewhere. and the place where we had a gap was we weren't, working on the model layer itself. Like we were a black box. And, we realized that, we can get to frontier-level model performance, with, by having great people who work on this. And, we've been open sourcing a lot of our work, in terms of, Recently, we, shared our work on DeFlash, which is a block-based, speculator, and we've open sourced, all of it. So, you can - By using open source DeFlash, you can get the same performance as you would with one of the proprietary providers. And the next thing we're thinking about hereVibhu [00:16:23]: I thought this wasAkshat [00:16:24]: YeahVibhu [00:16:24]: An interesting blog post as well, right? Like, I think in here you make a claim that. Not a claim, just that how effective speculative deco-decoding really just get to.Akshat [00:16:33]: Yeah.Vibhu [00:16:33]: Anything you wanna point out from this around, what people should know?Akshat [00:16:39]: Yeah, absolutely. the high-level summary is, it would help to describe what speculative decoding is.Vibhu [00:16:44]: Yes.Akshat [00:16:44]: I will, yes.Vibhu [00:16:45]: I think, likeAkshat [00:16:46]: YeahVibhu [00:16:46]: So we've covered like Eagle and all thisAkshat [00:16:47]: YeahVibhu [00:16:47]: Like Hydra and all those things, but it was like two years ago.Akshat [00:16:51]: Yeah.Vibhu [00:16:51]: I think it doesn't hurt, right?Akshat [00:16:52]: Yeah. Speculative decoding is you have a smaller model, called a draft model, predict tokens ahead of the bigger model, and then you have the bigger model, verify all of this, all the tokens are predicted. And the reason it's faster is if you're predicting, one token at once, you're bound by memory bandwidth. But if you can batch the verification of, the draft model, then you're much more efficient using compute, and it's faster, and as long as your draft model is producing a lot of tokens that can get accepted, which is called the accept length, you can get a speed up that's, multiple times of, the original model speed. and well, that's what we highlight here. It's Like people talk a lot about we made these kernels faster and whatnot, but improving kernel will only give you like few percentage points of improvement, and, increasing accept length, literally is a multiplicative decreaseVibhu [00:17:47]: Like two to four X.Akshat [00:17:48]: Yeah, exactly.Vibhu [00:17:48]: Without much head-on performance.Akshat [00:17:50]: Yeah. I think it may - you are running a second model, right? So it may be something more expensive in the compute,Vibhu [00:17:57]: I meant quality performanceAkshat [00:17:58]: Probably not by muchVibhu [00:17:58]: But yeah. I thinkAkshat [00:17:59]: So there's no drop in quality performanceVibhu [00:18:01]: YeahAkshat [00:18:01]: Because you're always. You're never accepting a token that the big modelVibhu [00:18:04]: It's strictly betterAkshat [00:18:05]: YeahVibhu [00:18:05]: Or it's same.Akshat [00:18:06]: Exactly.Vibhu [00:18:07]: Right. Yeah.Akshat [00:18:08]: And so we've been working a bunch on DeFlash, which is a block-based speculator. so it's instead of predicting, one token at a time, it's predicting a block. And we've been open sourcing our work with it. The next thing for us here is for helping people train speculators and custom models. it's it's something that traditionally is very forward-deployed engineering driven, support deployed, engineer driven, like you work with customers and help them do that. And our vision for. This is why we launched Auto Endpoints, is we want to make frontier-level performance available to everyone. And so, we mentioned this in the announcement, we teased it. The next thing we're, we're launching is, as you run an auto endpoint, we shadow trafficAuto Endpoints and Frontier-Level PerformanceVibhu [00:18:54]: Do you want to explain what auto endpoints are?Akshat [00:18:57]: Yeah.Vibhu [00:18:57]: I lovely, yeah.Akshat [00:18:58]: Yeah. So, this is, I guess, going back to your Modal is you touch the code, but, sometimes people don't wanna touch the code, and they wanna get started with an endpoint that works and has all the great performance and, scalability that Modal has. So we've made that easier with, a way to create an endpoint from our UI, from the CLI, that has all of our optimizations that we talked about, like the DeFlash stuff already baked in, and there's full transparency. So we give you the code, you can go run it yourself, and if you want, you can eject out into the full Modal experience, which we see as people get sophisticated, they do wanna tweak the models, they wanna, fine-tune stuff. You can still do all of that. It's it's not a black box. And yeah, the next thing, as we teased later in the post, is how do we give you value even beyond this in terms of having your draft models evolve as your data distribution evolves, again, without having to talk to a person and, yeah.Vibhu [00:19:59]: I guess just to understand it directly, you have the GPUs, you have an endpoint that's compatible, you serve open model. If someone was to do this themselves, what's the delta that you guys provide? So you do a lot of open source great work on effective inference. how does it compare to, say, I take the same model, 5.2 FP8, take shelf inference engine, vLLM, SGLang, get compute of similar capacity, similar cost. What's the delta that plugging into something this, like this offers outside of the benefit of, scaling?Production Inference Beyond Raw GPUsAkshat [00:20:34]: It's interesting because we've taken the approach of open sourcing our contributions and upstreaming them. we work closely with the SGLang team. We want the improvements that our team, comes up with to be, there in open source for others to use, even outside of Modal. The benefit to us is we have a team that has significant expertise in terms of if you do have something that is not there, our team can help you get that performance, first. the other thing is with these endpoints, we are way more elastic, as you said, than, anyone else, and you have true scaling to zero. you have true, burstiness, and in practice, that matters a lot more to people than just finding, the GPU and, running Modal code on something.Vibhu [00:21:20]: Yeah. And I will say it's not that straightforward to just. like what I said is easier said than done, right?Akshat [00:21:26]: Yeah.Vibhu [00:21:27]: It's I think still for the average person, still hard to just gut check using different. There's, there's quite a bit of combinations you can make there. the trade-offs aren't really known at face value.Akshat [00:21:40]: Yeah. it's it's not just that. I think it's it's that running production-grade inference is a hard infer problem.Vibhu [00:21:49]: YeahAkshat [00:21:49]: Even if you subtract out the autoscalingVibhu [00:21:50]: YeahAkshat [00:21:51]: Is controlling things like tail latency and, making sure every, request is delivered at least once and whatnot.The Model and Agent LifecycleVibhu [00:22:00]: There's a lot of innovation that you can do here. I think, it's very interesting that you're starting to encroach on, like as you become a full cloud, you're starting to encroach on other people's turf.Vibhu [00:22:09]: What will you not do?Akshat [00:22:13]: Well, we wanna follow our users and, make sure they get like a platform that has everything that works well together. so right now we're focused on the model lifecycle and the agent, lifecycle. so both like going from data prep to training to inference, and then also if I want to deploy a background agent, let's say, sandbox, do persistent storage, a whole bunch of other stuff.Vibhu [00:22:38]: We talked to Cole, who did, OpenInspect. Yeah.Akshat [00:22:42]: Yeah.Vibhu [00:22:42]: And RealInspect also is on Modal.Akshat [00:22:44]: Yeah. So Ramp Inspect was a great example of a background agent that was really successful because they, were able to use some of the primitives like snapshotting and fast scaling to just have something that feels really reactive and works well.Ramp Inspect and Background AgentsVibhu [00:23:02]: Yeah. That's the new CTO of, Ramp right there.Akshat [00:23:05]: Yeah, Rahul.Vibhu [00:23:08]: It was really fun. yeah, okay, I think, all very bullish. Like, one of my reflections was also I did not originally. So when I met you guysThe Inference Inflection: CPU, GPU, and Co-LocationVibhu [00:23:19]: You weren't that much in the GPU game, and now you're all about, inference. And one of the points that I hinged on for Jensen's keynote at GTC this year was, what we're calling like the inference inflection, right? That let's say in AI workloads or machine learning workloads, it used to be like, let's call it eight to one GPU to CPU, and now it's more like one to one, which is like a interesting. Like, - because of how much agents are blocked or call out to this, to CPU heavy stuff the actual, like, limiting factor, like, swings back and forth from GPU to CPU a lot more than it used to be all GPU and then occasional CPU.Akshat [00:24:01]: Yeah.Vibhu [00:24:02]: GPU, CPU. And now it's like just constantly, and you just have to locate everything.Seventeen Clouds and the Supercloud StrategyAkshat [00:24:08]: Yeah. And that's one of the things that, again, we see as, something appealing about Modal, which is we've built this capacity pool that spans, 17 cloud providers, so we're, we're very good at Running on various kinds of cloud capacity across the worldSwyx [00:24:24]: You don't have your own data centers?Akshat [00:24:25]: We don't have our own data centers. We just run across a lot of neo cloudsSwyx [00:24:29]: Yeah. AreAkshat [00:24:30]: Metal providers.Swyx [00:24:30]: Yeah. Question mark.Swyx [00:24:31]: Yeah. You're, you're running the math, and you're like, “What's the cutover point where you're like.”Akshat [00:24:36]: Yeah, it's a good question. part of it is we see our differentiator in the software layer, and, being capital light and focusing on the software helps us move really fast. so far it's worked out well because there are so many other people building data centers that we're able to work effectively with them, and again, focus on what makes us, special.Swyx [00:24:55]: Yeah.Swyx [00:24:56]: 17 gets you into, like, the local providers sometimes. LikeAkshat [00:25:00]: The,Swyx [00:25:01]: Which was the most interesting one?Akshat [00:25:02]: There are a lot more neo clouds than you expect, and they all have various degrees of, various levels of reliability. And, that's why it's something we've invested a lot of time in, is building our own reliability layer on top. so if the GPU falls off the bus or something happens, we user workloads are not affected, and that lets us use a lot more capacity than,Swyx [00:25:30]: YeahAkshat [00:25:30]: You as a user would be able to.Swyx [00:25:32]: It's a useful thing to have because like now everyone knows, like, what layer you are and, like, you optimize for being the super cloud of all clouds.Akshat [00:25:41]: Yeah. That's, that's, that's the idea. and so I guess when you mentioned colocation, that's, that's another interesting thing where, one thing we've seen is people come to us when they want, very specifically located, CPUs or GPUs, like they wantSwyx [00:25:57]: Oh, they pin it in likeAkshat [00:25:58]: YeahSwyx [00:25:58]: EU?Akshat [00:25:59]: Exactly. Or EU, US.Swyx [00:26:01]: Right. Data resiliencyAkshat [00:26:02]: AustraliaSwyx [00:26:02]: Locality thing or performance or what?Akshat [00:26:04]: It's either data locality or latency, yeah.Swyx [00:26:07]: Yeah.Akshat [00:26:07]: Like, you want your. They're running sandboxes and model. They want them to be right next to aSwyx [00:26:10]: Yeah, it's easy thenAkshat [00:26:11]: YeahSwyx [00:26:12]: To. That is important in all those things. and so, like, you've accidentally, I don't know if it's accident, but, like, you've built the perfect primitive for agents to express themselves. And then, like, it's almost very funny how every extra development just involves more file system, just involves more CPU.Akshat [00:26:30]: Yeah.Swyx [00:26:31]: Just like the things that you already have. I don't know much about, if there's any, like, networking usages that are interesting, but you've also done some good work on networking.Networking, Sidecars, Private IPv6, and SandboxesAkshat [00:26:40]: Yeah, that's exactly right. Like, we're just taking compute storage and networking and building stuff on that layer, for, again, the stuff people need.Swyx [00:26:49]: YeahAkshat [00:26:50]: We see a few interesting networking things coming up. one is people want networked sandboxes. so we haveSwyx [00:26:57]: For like a Docker cluster type thing.Akshat [00:26:59]: Yeah.Swyx [00:26:59]: Sorry, Docker Swarm. Oh, f**k. What is it called?Akshat [00:27:02]: Compose.Swyx [00:27:03]: Compose type thing.Akshat [00:27:04]: Yeah. So if you want Docker Compose, our sandboxes now support, this thing called sidecars. So you can. A sandbox is a pod of containers, and you can run multiple containers in, a sandbox. also useful because, going back to networking, people want a lot of control over, outbound networking from a sandbox.Swyx [00:27:23]: Yeah.Akshat [00:27:23]: Like, they might wanna run a middle proxy for, like, maybe logging stuff for RL or, controlling how egress can happen to a domain, injecting credentials. and yeah. So we've, we've had to build a lot of that stuff ourselves.Swyx [00:27:38]: Yeah.Akshat [00:27:39]: But then also sometimes people want, sandboxes spanning multiple nodes to talk to each other, which is an emerging thing we're seeing. We have support for that for a different reason, and yeah, we'll see if that becomes stable.Swyx [00:27:52]: Like, just an open socket. It's a. This is directly like mTLS.Akshat [00:27:56]: We do support that, which is you can, expose a tunnel inside a sandbox.Swyx [00:28:01]: Yeah.Akshat [00:28:01]: And then you can either expose it to public internet or it can be, you can add like a HTTP, auth layer above it. But we have this thing called I6PN, which we haven't talked about, which is this, like, overlay network using IPv6 addresses. so if Modal containers, within the same workspace, when this is enabled, can address each other using this private IPv6 address, and no one else can.Akshat [00:28:28]: So it's like private networking, for containers. We built it because we needed it as a primitive for our distributed training product. so we have this other feature, which is you can add a decorator to a function, and you get a cluster of GPUs. and they have RDMA networking. so you can run a distributed training job, that's truly serverless. and we did the overlay network for that. But then we've seen that people are using it for other reasons, and, I'm intrigued to yeah, what would people do with it.Swyx [00:28:59]: Build primitives and let people figure it out, right?Akshat [00:29:01]: Yeah, exactly.Swyx [00:29:02]: You put out a pretty interestingAkshat [00:29:03]: They're like, they read the docs webpage. Let me use thatSwyx [00:29:06]: YeahAkshat [00:29:06]: Something they never intended to work. This is literally not even in our docs page. People somehow found it, and they're using it.RDMA, Memory Movement, and Distributed TrainingSwyx [00:29:12]: Huh.Swyx [00:29:14]: The way you portrayed it with, like, RDMA versus TCP, like, very well laid out, but just the transfer speed change at scale for RL, like yeah, you have it, you have it built in. I'm sure someone found it. It's found it to be a lot more efficient before you made a thing out of it, right?Akshat [00:29:32]: Yeah. And not to split hairs, I guess the overlay network is the TCP overlay network.Akshat [00:29:39]: The reason we have that is you need that to do the key exchange for RDMA before you set up the RDMA network on top of that. but then people found the TCP part.Swyx [00:29:48]: Can I tell you, this is like a big aha moment for me becauseAkshat [00:29:51]: YeahSwyx [00:29:51]: So I review 2,200 submissions for the World's Fair.Akshat [00:29:56]: Yeah.Swyx [00:29:57]: And then I got this from John OsterhoutAkshat [00:29:58]: HuhSwyx [00:29:59]: Who I don't know if. Do John Osterhout by name?Akshat [00:30:01]: The name sounds familiar.Swyx [00:30:02]: He published a. He's a well-known professor, published a lot of interesting software design books, and this is the talk he chose to submit, is on RDMA at Inference. And I'm like, you wouldn't think that this guy, who is like operating systems guy, would care about RDMA.Akshat [00:30:20]: I, it makes sense to me because I,Swyx [00:30:24]: This is the cloud, right? YeahAkshat [00:30:25]: Like, the way you move around your KV cache and how efficiently you can do it, how efficiently you move, your weights from your training GPUs to your inference GPUs in RL is there's a lot of degrees of freedom, and it is a systems problemSwyx [00:30:41]: YeahAkshat [00:30:41]: Moving memory aroundSwyx [00:30:42]: YeahAkshat [00:30:43]: Scheduling.Swyx [00:30:44]: This shows you how primitive my understanding of networking stuff is.Swyx [00:30:46]: Is this like the domain of WireGuard as well?Akshat [00:30:50]: Not quite.Swyx [00:30:51]: It's adjacent?Swyx [00:30:53]: Explain everything.Akshat [00:30:54]: Sure.Swyx [00:30:56]: How do we move memory around GPUs?Akshat [00:30:58]: Well, so sorry. Yeah, that is memory. Sorry, I was talking more, and maybe I was talking like five minutes back, about the private IPv6, addressing that you've set up.Swyx [00:31:09]: Yeah.Akshat [00:31:09]: Is it like it's a VPN?Swyx [00:31:10]: Yeah, it is like a VPN, and yeah, WireGuard is, yeah, you're right. It is,Akshat [00:31:16]: Right. Yeah, you already moved on to new topicsSwyx [00:31:17]: A similarAkshat [00:31:18]: OkaySwyx [00:31:19]: In the same space, WireGuard is, encrypted and this is,Akshat [00:31:23]: And you don't need encryption.Swyx [00:31:23]: Yeah.Akshat [00:31:24]: Yeah.Swyx [00:31:24]: This is not encrypted. that's the main difference. This is TCP and we have eBPF programs that will reject or allow the TCP connection based on whether you're allowed to do it.Akshat [00:31:35]: Used to involve a full sidecar, but now you have eBPF in the Linux kernel.Swyx [00:31:39]: Yeah.Akshat [00:31:40]: Yeah. I don't know if this is a natural follow-on to the topic of like my skepticism on distributed training is that while, like, people spend a lot of money on, like, cables to hook up GPUs, and even that is not, like, fast enough, and that's the bottleneck, is your networking fast enough?Swyx [00:31:59]: Yeah. So I guess you're talking about fully distributed training like, Dialog or something which is like cross data centerAkshat [00:32:06]: That would be, yes.Swyx [00:32:07]: That's the extreme.Akshat [00:32:08]: Yeah.Swyx [00:32:08]: You're in the middle, and then other people would have like the Mellanox cables up in, like, their actual data center.Akshat [00:32:14]: When you run multi-node training on Modal, RDMA, I think Mellanox, is, or InfiniBand is like a, is all seen as RDMA. but it's a way to bypass the TCP networking stack and, transfer, stuff much faster, between one node, to the other. And we have I think like 3 terabit per second, internal networkingSwyx [00:32:40]: OkayAkshat [00:32:40]: Which is the standard that's needed.Swyx [00:32:42]: Okay. So I misunderstood whatAkshat [00:32:43]: 50Swyx [00:32:43]: What part of the stack you wereAkshat [00:32:44]: 50 gigs overSwyx [00:32:45]: YeahAkshat [00:32:45]: If you wentSwyx [00:32:45]: YeahAkshat [00:32:46]: RDMA.Swyx [00:32:46]: Okay.Swyx [00:32:48]: Yeah. I, very impressive work.Multi-Node Training, Post-Training, and Auto ResearchSwyx [00:32:52]: So effectively you're extending like the model philosophy to the training cluster, like, yeah.Akshat [00:32:59]: Yeah. And we're, we're not going for like large scale training runs. the thing that we've built multi-node training for is, we see a lot of, smaller scale post-training. like, people are post-training like medium sized fund models, so they can, get higher quality on inference. this is a perfect fit, for something like that.Swyx [00:33:21]: Yeah. That is my impression of how a lot of these labs explore branches in post-training and then eventually merge whatever they find in.Akshat [00:33:31]: Yeah. The other use case we've seen for multi-node training is even if you have a big cluster, your researchers are still doing small runsSwyx [00:33:38]: YesAkshat [00:33:39]: Having elasticity thereSwyx [00:33:40]: Right, sureAkshat [00:33:40]: Matters a lot more.Swyx [00:33:41]: Yeah. the, like, this is like the current limiting factor for auto research, which is like you need to give your model some GPUs in order for it to completely run.Akshat [00:33:51]: We have a blog post on auto resource and model is,Swyx [00:33:55]: YeahAkshat [00:33:56]: Yeah, like, turns out to be pretty good substrate for that.Swyx [00:33:59]: So my impression is auto research means many things, likeAkshat [00:34:01]: YeahSwyx [00:34:01]: Anything that Andrej coins. Right now it's still science fair, right? Like not like, I don't know how many people are doing this.Akshat [00:34:08]: We're having a golf.Swyx [00:34:08]: Yeah.Akshat [00:34:09]: I thought the same thing.Swyx [00:34:11]: Yeah, you would know.Akshat [00:34:12]: We, like, our internal both training and inference teams use this the general shape of this quite a bit. like we have this one internal repo called auto inference, which essentially we've automated our own forward-deployed engineering efforts using, this harness, which is, the agent will just spin up a sweep of different things. It'll even run like, NVIDIA inside profiler and it'll like tweak configs and it'll arrive the right thing. it'll change your GPUs both from H200 to B200, and works really well.Swyx [00:34:47]: Nice.Akshat [00:34:47]: So yeah.Swyx [00:34:48]: By the way, I enjoy that your forward-deployed engineering is so technical that you have to do these things.Swyx [00:34:52]: It's very different from forward-deployed engineering from other people.Akshat [00:34:54]: Yeah. For our forward-deployed engineering team is, essentially they're like applied inference researchers or applied training researchers.Swyx [00:35:02]: Someone told me like they have to be able to build, but they also have to be able to sell. do they have to sell or are they like they're good, they're just like post-sale type of thing?Akshat [00:35:09]: It does, being able to talk to a customer and engage effectively with themSwyx [00:35:13]: YeahAkshat [00:35:13]: Matters a lot.Swyx [00:35:14]: They want the same thing.Akshat [00:35:15]: Yeah.Swyx [00:35:15]: ?Akshat [00:35:15]: But it's it's not really a sales, thing. We pair them with-- We have solution architects as well that are more on the sales side.Swyx [00:35:23]: Okay. Let's spend a bit more time on auto research. This is a big focus for for this year. Where does this go? like, have people explored enough? Like, there's all these beautiful charts of like improve and then level off a bit and then you find the next thing. Is this one abstraction up from normal training? Is that how we think about it, or do you think about it differently? Like model level training versus high, like driven hyperparameter search.Auto Inference and Modal BenchAkshat [00:35:51]: Yeah, like,Swyx [00:35:51]: Someone, some people call it like neural architecture search or whatever, right? Like.Akshat [00:35:54]: Yeah, - So the stuff I've seen people do with it is nowhere on the architecture level. It's pretty much tweaking parameters, but it's it's a hyperparameter sweep that's guided by some model intuition, so it's like much more efficient than, whatever other, sweep you would have.Swyx [00:36:12]: Yeah, it's just, it's just a question of where you want to spend your compute?Akshat [00:36:16]: Right.Swyx [00:36:16]: ‘Cause yeah, you can just throw infinite amounts of money on this and somehow you'll bang out Shakespeare?Akshat [00:36:22]: Yeah, infinite monkey.Swyx [00:36:24]: Yeah, so like the very good for model. and I think it's also very important that agents can spin up other agents, can spin up their infrastructure. Like very good for you. how good is our LLMs at generating model code? Like the benefit of existing LLMs is that you are in the data.Akshat [00:36:42]: Yeah. They're, they're surprisingly good. I think like pre Cloud 4 they were not, and then now they're able to shot, stuff out of the box. But we're playing around with releasing like a Modal Bench for like the harderSwyx [00:36:55]: YeahAkshat [00:36:55]: Things, that the LLMs cannot do yet and maybeSwyx [00:36:59]: What's an example of that?Akshat [00:37:01]: I think the things that- Sometimes agents struggle with, without right guidance and a skill is, how to, use the rest of our observability. Like how to. Something is failing, like how do you look at the logs and then update the right thing? It's reasoning about that. But they're able to shot, likeSwyx [00:37:23]: Yeah. You can just add a skill to it?Compute Strategy and Capacity PlanningAkshat [00:37:26]: Yeah. So we have a Modal skill now that. Which is why we built this Modal Bench. It's to find things like that, so we can address them in our tool.Swyx [00:37:35]: Tune a skill. Yeah.Akshat [00:37:36]: Yeah.Swyx [00:37:36]: No. it's it's good. are you facing any shortages? like we talk a lot about GPU shortages, but also CPU, also memory.Swyx [00:37:44]: Yeah.Akshat [00:37:45]: We have had a lot of growth, which means that, there's - we've had to be much better aboutSwyx [00:37:53]: PlanningAkshat [00:37:54]: Proactive capacity planning.Swyx [00:37:55]: Yeah.Akshat [00:37:55]: So we have,Swyx [00:37:57]: Which by the way, like it's like a MBA's like dreamAkshat [00:38:00]: YesSwyx [00:38:00]: Is like just planning this stuff. I think last time you and I talked about something maybe about this.Akshat [00:38:03]: Yeah. we have a really competent team of people that we call, The role is called compute strategy. so yeah, if anyone listening here or wants to work on thatSwyx [00:38:13]: Compute strategy?Akshat [00:38:13]: Yeah.Swyx [00:38:14]: I think,Akshat [00:38:14]: I feel like,Swyx [00:38:15]: I think the normies call it FP&A or something.Akshat [00:38:18]: Well, it's more It's it's not FP&A. It's it's There's a lot of interesting financial questions of like what is the blend between one year and three-year reservations? how do we forecast our own capacity? how do we. especially since our capacity is very fungible across different GPU types and different regions, like you have to model a lot of it. and you also have to have an opinion on how the supply chain is gonna evolve, and then you have to like, take bets,Swyx [00:38:49]: YeahAkshat [00:38:49]: Based on that.Swyx [00:38:50]: Tokenomics.Akshat [00:38:50]: Yeah.Swyx [00:38:51]: This is like probably a not a real point, but, I was trying to think about like what other industries. I was trying to think about like, we cannot be first to like these kinds of problems.Akshat [00:38:59]: Yeah.Swyx [00:39:00]: And what other industries have had this? And I was like, airlines with fuel and like they have to hedge their fuel and like, I think for a long time Southwest because they made like a hero fuel bet, they like were like super low cost becauseAkshat [00:39:12]: OhSwyx [00:39:12]: Compared to everyone else.Akshat [00:39:14]: Yeah. I hadn't thought about that.Vibhu [00:39:16]: We're at a fun time too?Akshat [00:39:18]: Yeah. It's. A lot of the compute business in general, for us is also about being very good about capacity management. That is how you have great unit, economics. but also over time it's how you can unlock more value for customers. Like, one of the things we're building now is like a way for customers to get, If they don't care about latency, like get much cheaper pricing and they'll get results back in like next 24 hours or something, like a batch tier essentially.Batch Tiers and Latency-Insensitive WorkloadsSwyx [00:39:47]: Yeah.Akshat [00:39:47]: And those are levers we have because we control the whole stack and scheduling and whatnot to give people a sufficientSwyx [00:39:53]: Yeah. I feel like they're not as popular. Like those, like the Frontier Labs have all those APIs. They're not as popular as they should be.Akshat [00:40:00]: The demand that we see for something like that is not for LLMs. although sometimes people wanna run evals andSwyx [00:40:08]: OkayAkshat [00:40:08]: Synthetic data prep and there it makes sense.Swyx [00:40:10]: Okay.Akshat [00:40:11]: But it's from a lot of LLM companies, like people who are doing computational bio, like they have to run really big batch jobs and they don't care about when they get it back.Swyx [00:40:22]: Yeah. And like they have a reasonable. It's it's also like a cousin to the stopping problem of like, will this finish in time?Akshat [00:40:30]: Yeah. You can bound it.Swyx [00:40:33]: Yeah.Akshat [00:40:33]: Like you can give peopleSwyx [00:40:34]: YeahAkshat [00:40:34]: SLAs on it.Swyx [00:40:35]: Yeah. I think what's, what's interesting is like the next phase of model.Swyx [00:40:38]: Like what, do people expect from you, now that you're established and you're like well-known compute player among all these leading companies. You had an inference launch week, and we talked a little bit about the launches. like what else? Like what else should people know?What Modal Builds NextAkshat [00:40:55]: We are building primitives that make our users' lives much easier. So, I think for example, with LLM inference, thousands more companies are gonna post-train their own models and, deploy open source models for inference. so we're thinking a lot about what is the best product shape for that. And, that involves everything from our training gym to, then, endpoints that get frontier-level performance. again, but I haven't talked to anyone. It looks somewhat different on other verticals. Like, we're also seeing a lot of real-time, audio-video stuff in there, which is why like, we're working on things like regional routing, with fallbacks. So you can get GPUs that are as close to users as possible. so you get like low latency for video streaming and whatnot. And then on the agent side, it's,Akshat [00:41:52]: We're still working very closely with our customers because stuff is changing so fast in terms of what they need. And, I think beyond sandboxes and persistent file systems, there's a lot of other things people will need from this agent stack as they build production agents. So yeah, we're thinking about those other things that fit in there.Swyx [00:42:13]: I want to ask what the other things are.Akshat [00:42:15]: Yeah. I probably should share right now.Swyx [00:42:17]: I think-- I think, okay, so, I do think a lot about the principal components of cloud, and you do talk about compute storage networking.Akshat [00:42:25]: Yeah.Swyx [00:42:25]: Because so far for me, it's fine. so far for the. the first couple generations of cloud, it's fine. What's different, qualitatively different about agents that you need some new permission level? Like a lot of people, okay, and I'll just kinda spew tokens at you until it like hopefully sparks something.Akshat [00:42:43]: Yeah.Swyx [00:42:44]: Like the new level now is whatever Claude Code does, which is dangerously scope permissions or like allow list by command or like whatever, right? And sometimes they're like, “Well, okay, we have like this adaptive thinking mode where like, just trust me, bro. I will make the calls for you.” Is that it? like mediated permissions.Hard Guardrails vs. LLM-Mediated PermissionsVibhu [00:43:03]: Now you're looping it with a goal and letting it roll.Akshat [00:43:06]: Yeah, I'm, I'm skeptical of LLM media permission for stuff that is at the sandbox level because you do want hard boundaries.Swyx [00:43:16]: Yeah.Akshat [00:43:16]: Otherwise, someone can exfiltrate stuff.Swyx [00:43:20]: But likeAkshat [00:43:20]: YeahSwyx [00:43:20]: Maybe that's old school thinking. Maybe we're the dinosaurs.Swyx [00:43:23]: Maybe the AI OS or the LLM OS is really the kernel is a goddamn LLM.Swyx [00:43:30]: Like it makes you feel uncomfortable.Akshat [00:43:31]: Yeah, I'm, I'm toldSwyx [00:43:32]: But that's what trusting the LLM is. Like imagine a spherical cow perfect LLM.Akshat [00:43:36]: Right.Swyx [00:43:37]: That it.Akshat [00:43:39]: Maybe.Swyx [00:43:41]: I wanna test the boundaries, right?Akshat [00:43:42]: Yeah.Swyx [00:43:42]: Like, and I don't believe that, but I wanna see where I'm wrong ‘cause that's, that's the consensus.Akshat [00:43:49]: Yeah. I think you always need hard guardrails when you want, And you can pair those with softer guardrails, right? And that's gonna be a lot of mediated.Managed Agents and Specialized SandboxesSwyx [00:44:00]: There. I'll also get you a end with a couple of your commentary on like the ecosystem outside of Modal. Manage agents. Everyone has one. Gemini, OpenAI, Claude, very useful for you, but also like it is their way of starting to edge into your space.Akshat [00:44:17]: Yeah.Swyx [00:44:17]: What's going on?Akshat [00:44:19]: Yeah, we're, very excited to partner with Anthropic and some of the other foundation labs, will not name who we're also working with. the way we see it is the manage agent thing is a great place to start if you're starting out building an agent and, But then when you get to, building something more production grade, like you're a company that's like Ramp that's building their own, Ramp also runs their accounting agent on us, so their external-facing agent. You need a lot more control over, your compute primitive on things like, what sort - how do you persist different files that the agent has access to, and how do you snapshot and restore? How do you control the networking? maybe you want GPUs. When you get to that point, you kinda want, a specialized sandbox provider, that gives you those things, and that's the role that we are trying to play.Swyx [00:45:15]: YeahAkshat [00:45:16]: We don't really have an opinion on the harness, whether it runs - it's a cloud-managed agent, and you hook it up to Model Sandbox, or you run the harness in Model Sandbox. We'll see where people converge with that.Swyx [00:45:26]: Yeah. Do you any opinions on like the meta harnesses, or just another layer on top of these things?Akshat [00:45:31]: You mean like the OpenPipeSwyx [00:45:33]: OpenPipe is one. I think Vercel had one, which I can't remember the name of right now. Fredshot had one. and then, to me, most recently was Data Databricks that had Omnigen. All these are meta harness. Like it's kinda pseudo agent cloud type things.Akshat [00:45:50]: I personally have not played around with them.Swyx [00:45:53]: Yeah.Akshat [00:45:53]: Build agents with them.Swyx [00:45:54]: Everything's bullish Modal, as long as it consumes more infra.Akshat [00:45:57]: That's why we're focusing on the infra layer. It's somewhere where our, relative competence is and, also it's a hard problem to solve.Swyx [00:46:06]: Yeah. I will say like just generally reflecting on that, I don't know if - if there's other topics on Modal, but like just generally reflecting as an infra person, not as intense as you, but in that field, this has like been the most exciting time in infra. Like it was boring for a while, and you couldn't really get people excited about data infrastructure. Like Eric would get on Data Console, everyone just watched the video and like say, “Look at how many sandboxes I can spin up,” and no one gave a crap.Why Infrastructure Became Exciting AgainAkshat [00:46:39]: Yeah.Swyx [00:46:40]: And like now everyone gives a crap.Akshat [00:46:42]: That's true. It is a very exciting time, and I think a lot of that's driven by just the amount of scale all of this stuff needs.Swyx [00:46:50]: I think the, like a lot of your initiatives or a lot of your like product directions make sense in retrospect, which is like the best kind, but I wouldn't necessarily have thought about it myself, which.Akshat [00:47:00]: We need the predictions.Swyx [00:47:02]: I think there's a lot that you just don't even see, right? Like you have the batch, you have the voice, you have the multimodal, but what else?Akshat [00:47:10]: What else is coming up for usSwyx [00:47:11]: Yeah. Where do you see things going?Akshat [00:47:13]: Yeah. I, in generalBiotech, Robotics, and Non-LLM AI WorkloadsAkshat [00:47:15]: It's it's clear that there's there's a huge shift happening. I think one thing that's not as obvious to people because LLM inference gets talked about so much and is also we work a lot of companies that are, doing things like drug discovery and computational bio, like the Chai Discoveries of the world. Big things are probably gonna happen there. we work a lot of robotics companies that are putting robots in like active deployments and getting good results out of them.Swyx [00:47:45]: Is there Air Gap Modal? Is there a version that is like prem air gapped whatever?Akshat [00:47:50]: No. We,Swyx [00:47:51]: You should cloud only.Akshat [00:47:51]: Yeah.Swyx [00:47:52]: Yeah. Okay. But yeah, so what you're saying is like because you're focused on primitives and they're good primitives, you find use cases in all these kinds of things.Akshat [00:48:01]: Yeah.Swyx [00:48:01]: Probably diversifies you a little bit away from LMS all the time.Akshat [00:48:05]: Yeah, absolutely. We're, we'- our goal isn't to only serve the LLM inference market.Swyx [00:48:10]: There are a lot just on the website, the audio,Akshat [00:48:12]: Yeah. We said both onSwyx [00:48:14]: Computational bio images. Yeah, there's a lot here. There's QTA TTS, customizing. Oh, Chatterbox. there was customizing Whisper.Akshat [00:48:24]: Okay. Yeah.Swyx [00:48:25]: This screen reminds me of a fallen competitor, which Replicate.Model APIs vs. Differentiated AI ProductsSwyx [00:48:31]: What's your postmortem on what happened?Akshat [00:48:34]: This is one thing we've stayed away from is providing an API for models because I think providing model APIs is some of it ends up serving like a really hobbyist market, which is much less sticky.Swyx [00:48:50]: Yeah.Akshat [00:48:50]: And we've always wanted to build for companies that are building products and need more flexibility that's not just an API.Swyx [00:48:57]: Which you can build an API for a model and this is clearly what it is. But you - but what you're saying, you can wrap it into a more fully functioning back end that you run.Akshat [00:49:06]: Yeah. So all of our examples, it's not that spin up this model, here's an API token, use it. They're all code.Swyx [00:49:13]: Okay.Akshat [00:49:13]: And so the point is that this is just an example.Swyx [00:49:16]: Starter code.Akshat [00:49:17]: Yeah. But you can tweak it however you want.Swyx [00:49:20]: Yeah.Akshat [00:49:21]: And if you're like a company building a product, like, computational bio whatnot, yeah.Swyx [00:49:26]: I guess I'm trying to tease out for listenersAkshat [00:49:28]: YeahSwyx [00:49:28]: When does it stop becoming, oh, you're just an API call and you're just a wrapper on API to becoming what you call a product, right?Swyx [00:49:36]: Like, what is that layer? Like what-- Like, more lines of code, but like beyond that, what is the substance that people add that qualifies it to be something more?Akshat [00:49:46]: I think there's a little bit of like a selection effect of like a lot of the companies who do wanna get deeper into that level are probably building something that's more differentiated. And, I think, an example is like - with LLM inference, originally we, worked with companies that were building their own post-training frameworks or they were, - Ramp early in the day was training their own tokenizer and like swapping out the tokenizer in Llama and whatnot. I'm not saying that's, that successful, in that case. But a better example is like, let's say Suno. because Suno, does not use Modal for training.Swyx [00:50:26]: Mikey on the pod. Yeah.Akshat [00:50:27]: But they use Modal for all their inference and that's because they have like a custom-- They have completely custom model architecture and that means that they have to be at the code level and tweak things that are not, just an API.Swyx [00:50:41]: It's interesting as well, like we had, Ethan, most recently on the xAI Groq team make a prediction that like the next tier in video gen is not a better video model, it's a better model or agent that orchestrates video models.Video Agents and Production WorkflowsAkshat [00:50:56]: Oh, interesting.Vibhu [00:50:56]: Language model backbone that can use toolsAkshat [00:50:58]: RightVibhu [00:50:59]: And write code.Akshat [00:51:00]: Like, yes, I can make my second video or my second video from Groq, but I want my minute video.Akshat [00:51:06]: And I'm not going there through normal video gen.Swyx [00:51:10]: Yeah, that's interesting. I - So we have GPU sandboxes and recently have seen a few companies doing agents that do video manipulation or,Akshat [00:51:22]: Yeah. Give it FFmpeg and just do it.Swyx [00:51:23]: Run FFmpeg. But likeAkshat [00:51:25]: That's not enough.Swyx [00:51:25]: Yeah.Akshat [00:51:26]: You need to give it Adobe.Swyx [00:51:27]: Yeah, I hadn't put it together with like it would be a video production thing. in my mind these things were going more towards editingAkshat [00:51:36]: Yeah.Vibhu [00:51:36]: Well, shout out Mantis.Akshat [00:51:37]: I think about this a lot.Swyx [00:51:38]: .Akshat [00:51:41]: Yeah. Sorry.Vibhu [00:51:41]: Luma. Luma Agent is a version of this for video production, but it's a off.Swyx [00:51:46]: I was gonna get your quick takes, on some other stuff that happensGitpod/Ona, CI, and Runtime SandboxesSwyx [00:51:50]: In recent news and just-just see if you have anything interesting. Gitpod, very li

The Post-Quantum World
Fraud Detection via Photonics – with Pouya Dianat of Quantum Computing Inc.

The Post-Quantum World

Play Episode Listen Later Jul 8, 2026 33:11


Is quantum technology purely a science of the future, or can it solve major enterprise problems today? In this episode, Pouya Dianat, Chief Revenue Officer of Quantum Computing Inc. (QCi), joins host Konstantinos Karagiannis to discuss how his company is moving past the lab stage to deliver real business value right now. Dianat breaks down QCi's unique hardware approach using nonlinear photonics and thin-film lithium niobate, explaining why the company has focused on application-specific, analog quantum optimization machines rather than waiting for universal gate-based systems to mature in the 2030s.   The conversation dives deep into QCi's latest breakthrough algorithm, CVQBoost, which leverages the physics of their Dirac 3 system to revolutionize financial fraud detection. Operating literally at the speed of light, this room-temperature, rack-mountable quantum computer eliminates classical processing bottlenecks and effortlessly bypasses the training limitations that plague traditional GPUs. Dianat also shares insights into how recent federal executive orders on quantum technology are impacting the sales landscape, QCi's cutting-edge quantum authentication solutions for network security, and his strategic vision for driving immediate commercial revenues while advancing a long-term quantum roadmap.     For more information on Quantum Computing Inc., visit https://quantumcomputinginc.com/. Visit Protiviti at www.protiviti.com/US-en/technology-consulting/quantum-computing-services to learn more about how Protiviti is helping organizations get post-quantum ready.  Follow host Konstantinos Karagiannis on all socials: @KonstantHacker             Questions and comments are welcome!  Theme song by David Schwartz, copyright 2021.  The views expressed by the participants of this program are their own and do not represent the views of, nor are they endorsed by, Protiviti Inc., The Post-Quantum World, or their respective officers, directors, employees, agents, representatives, shareholders, or subsidiaries.  None of the content should be considered investment advice, as an offer or solicitation of an offer to buy or sell, or as an endorsement of any company, security, fund, or other securities or non-securities offering. Thanks for listening to this podcast. Protiviti Inc. is an equal opportunity employer, including minorities, females, people with disabilities, and veterans.  

Wealth, Actually
REDUCING THE NOISE OF AI INVESTING

Wealth, Actually

Play Episode Listen Later Jul 7, 2026 29:15


“Reducing the Noise of AI Investing”: In this Wealth Actually episode, Frazer Rice speaks with KEVIN SHEA, Senior Equity Analyst at BNY Wealth, about AI Investing and how investors should think about artificial intelligence as an investment theme rather than just a headline-driven trend. They discuss the difference between hype and durable fundamentals, how to segment AI opportunities across infrastructure, software, and end-user adoption, and why free cash flow still matters when evaluating companies tied to AI. https://open.spotify.com/episode/1NGM8j2KqdiUFWSLguBMEH?si=YmB4s0OVSqyy6U3Mpg7OaA https://youtu.be/Wnlub-HoiUo The conversation also explores circular financing risk, the role of management vision in fast-moving markets, which industries may be disrupted or strengthened by AI, and how large institutions are using AI internally to improve productivity, analysis, and client service. Chapters 00:00 – Intro and episode setupFrazer Rice introduces the episode, frames AI as a dominant investment theme, and welcomes Kevin Shea to help unpack AI Investing for the audience. 01:00 – Hype versus disciplined investingKevin explains that disciplined investing is what allows investors to separate hype from durable opportunity, and argues that AI adoption, spending, and earnings revisions point to real underlying fundamentals. 03:00 – How to bucket AI investment themesThe discussion turns to how investors can organize AI exposure, including beneficiaries versus disrupted companies, technology bottlenecks such as GPUs and networking, and industry adoption themes across sectors. 05:30 – Valuation, momentum, and free cash flowKevin discusses why free cash flow per share growth remains one of the most important drivers of stock performance and why parts of the semiconductor ecosystem may deserve a valuation re-rating. 08:15 – Circular financing and risk in the AI ecosystemFraser asks about the growing concern that AI companies are financing one another, and Kevin outlines both the bullish “escape velocity” case and the downside risk if business models do not become independently profitable fast enough. 11:45 – Infrastructure buildout and competitive uncertaintyUsing analogies like railroads and golf courses, the conversation highlights the risk that early builders may not be the ultimate winners, especially in a market with heavy spending and rapid leapfrogging among competitors. 13:00 – AI Investing: Public versus private market exposureThey examine whether owning public companies such as Alphabet offers meaningful AI exposure, versus gaining more direct but harder-to-access exposure through private investment vehicles. 15:45 – What strong AI management teams look likeKevin emphasizes that in an environment with no clear historical playbook, vision, execution, and the ability to identify durable differentiation are critical traits in management teams. 19:15 – Adaptability and strategic pivotsFraser adds that thoughtful adaptation matters, and Kevin notes that sometimes acquisition activity can signal whether a company is innovating ahead of the curve or scrambling to catch up. 20:45 – Which industries are most exposed to disruptionThe conversation shifts to sectors under pressure, especially parts of software and IT services, while stressing that disruption does not necessarily mean extinction. 24:45 – Why law and accounting may evolve, not disappearFraser offers a contrarian view that AI may make strong legal and accounting professionals more valuable, and Kevin compares that to earlier fears that Excel would eliminate accountants. 26:15 – How Kevin uses AI in practiceKevin describes how AI has made his team materially more productive, especially in data aggregation, scenario analysis, industry research, and portfolio risk work, while also helping BNY operationally across onboarding, security, and client communication. 29:10 – Where to find Kevin and closing remarksThe episode closes with Kevin sharing where listeners can connect with him and Fraser noting how quickly the AI landscape continues to change. Links KEVIN SHEA on Linkedin RICK FERRI on BRING SIMPLICITY BACK TO INVESTING Transcript of AI INVESTING Frazer (00:01)Welcome aboard, Kevin. Kevin Shea (00:03)Yeah, thanks for having me. Appreciate it, Frazer. Frazer (00:06)We're going to tackle two words that have basically taken over the investment world for the last six months: artificial intelligence. Before we do that, whether it's AI or crypto or tulips or anything with a lot of hype or buzz around it, how do you think about delineating between investing based on hype and doing it within the confines of a disciplined approach? Kevin Shea (00:32)They really do go hand in hand. You need a disciplined approach in order to recognize whether it's hype or not. The reality is that it's pretty impressive, the adoption we're seeing with AI: the amount of spend, the companies that are participating in and benefiting from AI. There was some concern with the stock movements that many of these companies have seen about whether the market was getting ahead of itself. Yet we have seen significant estimate increases throughout the year. If you take a look at some of the networking companies, their earnings expectations for 2027 are up almost 50% versus where they were just six months ago. The same is true with memory, GPUs, and CPUs. Fundamentally, we're seeing a lot of these companies have expansion in revenue growth and earnings growth, which is quite supportive of a durable trend. What's also very important is that adoption of AI is increasing. You can look at enterprise adoption: nearly two‑thirds of enterprises pay for an AI service. You can look at token usage — that's how much companies are using AI — and that has been parabolic as well. Look at the revenue generation of these AI models. Right now, they are some of the largest, fastest‑growing companies that have ever existed. So we don't really see this as a tulip scenario, or even comparable to the internet bubble. We find it very different. We think there are fundamental drivers to this trade, and we're seeing that through earnings growth. Frazer (02:37)Cool. AI to me is a term that encompasses a lot of different things, and in some ways it's become like real estate or water — it's starting to touch a lot of different industries. It's not just a thing unto itself, but something that's becoming integrated into a lot of other types of things. How do you define and bucket the investment themes so that it's digestible for the investor, and it's not just, “I'm investing in Anthropic or Google,” but people can parse out where it fits within a portfolio? Kevin Shea (03:14)It's a great question and probably one of the most important ones. Part of our overarching thesis is that for AI to fulfill its promise, it has to be in every geography, in every industry, at every company, and at almost every employee layer. We're seeing that when you look at the business units that are adopting AI: customer service, product development, marketing — basically divisions that almost every single company in every geography has. You phrased it as water, how it touches everything, and we're seeing that. So how do you segment it? There are a number of different ways: First, you can break it into: who are the AI beneficiaries, and who are those that will be disrupted by AI? Second, you can break it down into different bottlenecks. That's a way I frequently use within the technology landscape: GPUs, CPUs, memory, networking, storage, data centers. Then you look at that framework and see which companies are most exposed to those bottlenecks. Third, you can ask: which industries will benefit from adoption? Is that biotech, transportation, warehousing? Which companies could be more negatively influenced — maybe that's software? That's how we try to create an AI Investing framework for where we should focus our investment efforts and determine the allocation that our clients can benefit from. Frazer (05:17)As we dive a little bit into how you've bucketed these themes across different areas, there's the concept of benefiting from momentum or valuation versus maybe the cash flow and fundamentals of these different investments. I could imagine that, with the hype and mania around the space, there's a lot of interest. How do you temper that valuation play versus analyzing what the cash flows look like? Kevin Shea (05:49)One of the most highly correlated metrics to stock outperformance is free cash flow per share growth. That's often the most important metric, and we watch that heavily. What's incredible — and we talked about this earlier with estimate revisions — is that many within the AI ecosystem are generating extremely healthy free cash flow growth and margins. A lot of that is in AI infrastructure. They're being paid to supply all the equipment and semiconductors. There's also this concept that valuation multiples shift to where there's value creation. I'll give an example: The SOX, the semiconductor index, used to trade at parity with the S&P. But there's been a paradigm shift. A lot of the intelligence that's being created through these models is powered by semiconductors, networking, packaging, and hardware. You've seen semiconductors go from trading at parity to trading at almost a 50% premium. At the same time, the market is intelligent; it's shifted its view of software. Software used to trade at a 70% premium, and we think the intelligence layer has moved just one layer above where software applications normally sit. As a result, you've seen valuation compression for the IGV, the software index, from that 70% premium down to about 20%. Some people might look at the semiconductor index and say it's more expensive than where it historically trades — maybe that's hype. But we actually view it as a shift in where the value creation is occurring. So we think it's a healthy, understandable move within the market. Frazer (08:16)One of the questions that pops up is that there's a lot of news around the circular flow of cash, where a lot of these companies are all investing in each other. You hear “five hundred billion is going from Google into Anthropic,” or different flavors of that, where it seems like the money is rotating. And there's a question as to whether it's rotating and expanding, given sales and so on. How do you think about that and make sure that we aren't wandering into more of the sort of things that are happening off balance sheet that we don't see, while still recognizing the investment that's taking place? Kevin Shea (08:57)At minimum, it raises the risk profile. There are many circumstances and scenarios where this has occurred in the past — the internet being the most commonly referenced — and that obviously did not work out. There are multiple scenarios that could happen, but for simplicity we'll break it down into two. The first scenario is that this is such a capital‑intensive expansion that companies are doing an “all‑hands‑on‑deck” effort. The faster you can get capital from well‑capitalized firms, the faster you can build your infrastructure and reach scale so that these large language models are profitable. If you can expand and take capital from everywhere, then you can provide enough compute for all enterprises and consumers to utilize your product and your model. You reach “escape velocity” in the sense that your scale allows you to lower costs and become more profitable faster. That's the glass‑half‑full environment. Glass‑half‑empty is that they do not reach escape velocity. The business models needed more time to bring the cost of delivering AI down enough to be profitable on their own; they didn't need this extra capital to reach an enormous amount of scale, and they're moving too fast. If that scenario plays out, and these companies are not able to be profitable on their own, and the financial markets become tighter, that creates more downside risk for everybody in the ecosystem. We don't see that right now because, at the moment compute is available, it's being taken right away. We still feel comfortable with the financing occurring right now, but it is one of the top risks that we monitor. It's not that it's systemic, but it provides less clarity and disclosure, and it creates a riskier profile as we go through this expansion. Frazer (11:43)In the back of your mind, you're probably saying, “We want to make sure, if there are winners and losers in AI Investing, that we avoid the railroad scenario,” where you build this whole infrastructure and companies have to go bankrupt twice before they actually reach profitability. Or the bromide that golf courses only become profitable, if they ever do, because the person who built it — a passion project — didn't make it work, then it goes bankrupt, then the bank is stuck with it and doesn't know how to run it, then they get rid of it, and then the third person has learned the lessons from the first two and is able to push forward. Kevin Shea (12:23)That's a good point. When we look at all these different models being created, right now you have an environment where everyone is spending and keeps leapfrogging each other at different times. It's still a very unknown outcome for all of these players. There's a lot of competitive intensity in the large language model space and the broader AI ecosystem. It's certainly a very dynamic environment right now. Frazer (12:59)As investors are trying to access this, there are the public companies. You can go on your Fidelity account or talk to your advisor at BNY Mellon or anybody else and say, “I've heard about Anthropic or Google or all of these things.” As far as a good proxy for exposure, how do you think about that? For example, if I looked at Google and understand that they have underlying investments in their portfolio — in addition to their regular businesses — into these different scenarios, is that a way to get shorthand exposure? As opposed to trying to access a venture fund where the entry points are difficult, the hurdles are high, you need to write big checks, and access is gated? Kevin Shea (13:53)It's a very astute point when you mention circular financing. That doesn't just happen with public companies; a lot of these vendors and companies in this ecosystem are investing in private companies as well. When those private companies go public, you find out that Company XYZ is a top owner, and one of their suppliers. There has been a growing awareness that, with certain public companies, you have exposure to a handful of private companies. For BNY, our Fujio funds do a lot of our private investments. That's usually the best way to gain direct exposure. Frazer (15:37)Sure. Not to be flippant, but you're getting paid to own it at that point via their dividend, as opposed to you paying — at the SPV or LP level — to gain access to it. But yes, it's definitely not a pure play. I wouldn't buy Google just to be in a venture fund. And just to reiterate for listeners, this is not investment advice. We're trying to learn and talk through different types of scenarios. As you're thinking about this and looking at these different companies, what does a good management team look like? You'd think: a bunch of PhDs, great at coding, lots of experience in the venture community, maybe hung out in Silicon Valley. But everything is so new and dynamic. When you're evaluating these businesses, what does a good management team look like as they're trying to scale at warp speed, while profitability may or may not be a thing? Frazer (17:49)I'd add that I think there's an interesting component to AI Investing: a track record of what I would call thoughtful adaptation. When your business plan gets punched in the face and you're able to pivot — meaningfully pivot — I'm not talking about a dog food company suddenly putting “.ai” at the end of its name, but someone who can shift and take advantage of opportunities as they come up, as you say, without being so rigid in their vision that they end up getting lapped. I think that's an interesting facet to focus on. Frazer (20:50)When I try to get my arms around this, I bucket things in terms of: Disruption: blowing up something traditional Optimization: taking something that's already good and turning it into great World‑building: taking a vision, starting from zero, and building something that didn't exist before On that first point, what industries do you think are under attack, and how do you invest around that so you're not left holding the bag — you're not a buggy‑whip company as Tesla releases their next issue? Frazer (24:48)As an example, I run into all sorts of law firms and accounting firms, and I hear the comment that law firms are going away. I have a contrarian view. First, I think law has a wonderful ability to metastasize, to find issues, and I think AI is going to be great at finding those and keeping lawyers busy. Second, for lawyers who are good, I think the ability for AI to make them more efficient and help them graduate to even more detailed and “higher‑value” discussions will only increase. So when people say, “Law is going to be dead,” I don't really agree. I think that ties into your point that AI will help some companies that can adapt and use it well to drive further value, probably even charge more. For others, they'll be left behind or become cottage industries. Frazer (26:11)And there will be more and more issues to solve. I don't underestimate that. I think AI is going to start poking holes in different things we didn't think about. Then it will take good brainpower, made more efficient by AI, to deal with these new issues as they pop up. In your day‑to‑day job, what are you using AI for? Maybe through Bank of New York, and maybe informally, when you're doing other research — to be smart not only about the company areas, but what you're doing personally to be more efficient, take advantage of AI, and learn about cool stuff. Frazer (29:11)Cool stuff. How do people find Kevin Shea, and any final thoughts? Kevin (29:20) KEVIN SHEA on AI Investing Frazer (29:29)Terrific. Thanks for being on, and we'll be sure to stay in touch, as I'm sure everything will be completely different in not just six months — probably six weeks. https://www.amazon.com/Wealth-Actually-Intelligent-Decision-Making-1-ebook/dp/B07FPQJJQT/ Keywords: AI Investing

CMO Confidential
Rob Ward | A Top Venture Capitalist Analyzes the AI Landscape

CMO Confidential

Play Episode Listen Later Jul 7, 2026 40:59


This week on CMO Confidential, we are revisiting one of our favorite conversations with Rob Ward from January of 2026.A CMO Confidential Interview with Rob Ward, co-founder and General Partner of Meritech Capital, a top Silicon Valley venture firm. Rob shares his take on what he calls a "super terrifying and exciting time" and provides perspective on AI receiving the most capital of any technology in history, the "durability of revenue" and how quickly start-ups are now reaching $100 million in revenue. Key topics include: why VC's focus on growth vs. profitability; the risks associated with massive long-term capital investment; why marketers should pick a "trusted advisor" as their AI partner; and why your data strategy needs "context. Tune in to hear how Astronomer handled the "Coldplay Concert Incident" which immediately became a PR classic and the "VC Foie Gras Effect."What happens when a top venture capitalist pulls back the curtain on AI, valuations, hype cycles, and what's actually working?In this episode of CMO Confidential, host Mike Linton sits down with Rob Ward, Co-Founder and General Partner at Metech Capital, to unpack the realities behind the AI boom. Rob has spent more than 26 years investing in category-defining companies like Facebook (Meta), Snowflake, NetSuite, Zipcar, and Cloudera — and he brings a rare, grounded perspective to today's AI frenzy.Together, they explore: • Why AI adoption is still early — despite explosive growth • The real risks behind inflated valuations and “AI-washing” • How VC decision-making changes during platform shifts • What marketers and executives should actually look for when choosing AI partners • Why data strategy, change management, and trust matter more than tools • What layoffs, productivity, and the future of work really look like beneath the headlines • A masterclass in crisis communications, featuring Ryan Reynolds, Gwyneth Paltrow, and ColdplayIf you're a CMO, CEO, board member, founder, or agency leader trying to make sense of AI without getting swept up in the hype — this is a must-listen conversation.⸻Chapter Markers00:00 – Welcome to CMO Confidential00:19 – Introducing Rob Ward and today's AI conversation01:13 – Where we really are in AI adoption02:26 – Explosive AI growth: what's real vs hype03:35 – Why enterprise AI adoption is still a slog04:37 – Vendor spend, hyperscalers, and the trillion-dollar buildout06:12 – Is this an AI bubble? Public vs private market realities07:20 – Accelerating investment rounds and lack of diligence08:12 – AI-washing and durability of AI businesses09:46 – Proof-of-concepts, switching costs, and fragile loyalty10:55 – Big Tech vs startups: why this cycle is different11:40 – Why VCs chase platform shifts despite the risks13:05 – How AI is changing profitability and headcount math16:11 – “FOGRA” investing and capital distortion17:00 – Circular investing and data-center risk18:23 – Data centers, GPUs, and betting on the wrong future19:38 – Credit default swaps and financial warning signs21:45 – How executives should choose AI vendors22:58 – Change management and why culture matters most24:09 – Why data strategy is the real AI strategy26:36 – “Frequently wrong, never in doubt” and AI hallucinations27:01 – Practical AI use cases for marketers30:00 – Layoffs, productivity, and what's really happening to jobs33:05 – The best questions to spot real AI fluency35:00 – AI safety, geopolitics, and long-term risks36:38 – Crisis management masterclass: Astronomer, Coldplay & Ryan Reynolds39:58 – Final advice and closing thoughts⸻Subscribe for weekly episodes featuring world-class marketing leaders, board members, and C-Suite executives.#CMOConfidential, #MarketingLeadership, #BrandStrategy, #CorporateActivism, #MarketingStrategy, #CMO, #AIinMarketing, #ExecutiveLeadership, #BrandReputation, #ConsumerTrust, #DigitalMarketing, #MarketingInsights, #ThoughtLeadership, #BusinessStrategy, #CustomerCentricSee Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

Chip Stock Investor Podcast
Jensen Called Marvell a $1T Stock — We're Not Buying It

Chip Stock Investor Podcast

Play Episode Listen Later Jul 7, 2026 11:42


Marvell (NASDAQ: MRVL) stock has surged since Jensen Huang called it the next trillion-dollar company — we ran a reverse DCF on the Q1 FY2027 earnings to see if the rally is justified.Marvell just posted record Q1 fiscal 2027 results, with revenue up 28% year-over-year to $2.4 billion and guidance pointing to 35% growth next quarter. As a fabless chip designer and IP licenser, Marvell is riding two major tailwinds: hyperscaler demand for custom AI chips (ASICs and XPUs) and networking systems that interconnect GPUs across data centers, including a growing partnership with NVIDIA via NVLink.We break down the earnings report, the CFO transition to Dan Durn (formerly of Adobe and Applied Materials), the balance sheet impact of the Celestial AI acquisition, and rising share dilution from recent deals. Then we run a reverse discounted cash flow analysis on Marvell at its current price to determine what growth rate is already priced in, and whether the semiconductor cycle and AI infrastructure buildout can realistically support it.If you're weighing whether Marvell is still a buy after this rally, or looking for better value elsewhere in the semiconductor supply chain, this one's for you.Semi Insider members get access to our full research platform and tools, plus deeper research as it happens. Join at chipstockinvestor.com.Get 15% off your Fiscal.ai membership with our link: fiscal.ai/csiContent in this video is for general information or entertainment only and is not specific or individual investment advice. Forecasts and information presented may not develop as predicted and there is no guarantee any strategies presented will be successful. All investing involves risk, and you could lose some or all of your principal.CSI doesn't own shares of Marvell.

The InfoQ Podcast
Spite-Driven Engineering: A New Blueprint for Cloud Security in the AI Native Era

The InfoQ Podcast

Play Episode Listen Later Jul 6, 2026 40:44


In this episode, Alex Zenla (CTO/Co-founder, Edera) challenges the "laissez-faire" attitude toward modern infrastructure. She promotes "spite-driven development", building software to solve genuine technical pain points rather than passively accepting flawed abstractions, as a philosophy of improving the world of software. The discussion touches on the fragility of the current cloud-native stack, the security risks of multi-tenant Linux kernels, and the inefficiency of repurposing consumer-grade GPUs for AI workloads. Zenla also offers a pragmatic framework for the "AI-native" engineer: treat LLMs as symbiotic assistants for deep learning, not replacements for system-level expertise. Read a transcript of this interview: https://bit.ly/3QZQhxc Newsletter: Subscribe to the Software Architects' Newsletter, a monthly roundup of the patterns and technologies senior practitioners are working through, with the news and lessons from people doing the work: https://www.infoq.com/software-architects-newsletter InfoQ Online Certification Programs: 5-week online cohorts for senior engineers and architects, built around QCon talks. Programs now cover software architecture, AI engineering, and organizational architecture. Each week you join a four-hour live session with a confidential peer group of practitioners from other companies, apply frameworks from QCon talks to the decisions you're making at work, and earn an InfoQ certification. You leave with new approaches, or confirmation that the calls you're already making are the right ones. Learn more: https://certification.qconferences.com/ Upcoming Events: QCon San Francisco 2026 (November 16-20, 2026) https://qconsf.com/ QCon London 2027 (April 13-16, 2027) https://qconlondon.com/ The InfoQ Podcasts: Weekly conversations with senior software leaders about how they build systems and teams, including what they'd do differently. Listen to all our podcasts and read interview transcripts: The InfoQ Podcast: https://www.infoq.com/podcasts/ Engineering Culture Podcast by InfoQ: https://www.infoq.com/podcasts/#engineering_culture Generally AI: https://www.infoq.com/generally-ai-podcast/ Follow InfoQ: Mastodon: https://techhub.social/@infoq X: https://x.com/InfoQ LinkedIn: https://www.linkedin.com/company/infoq/ Facebook: https://www.facebook.com/InfoQdotcom Instagram: https://www.instagram.com/infoqdotcom/ YouTube: https://www.youtube.com/infoq Bluesky: https://bsky.app/profile/infoq.com Write for InfoQ: Share what you've learned building software with a community of senior practitioners, and get your work in front of the people who read InfoQ. https://www.infoq.com/write-for-infoq

Datacenter Technical Deep Dives
Getting Started with Local AI (2/3)

Datacenter Technical Deep Dives

Play Episode Listen Later Jul 6, 2026 48:34


Join us as Du'An digs into the real mechanics of running AI locally and in production - from GPU memory math to multi-agent architectures, observability, and the economics of self-hosted inference. Du'An walks through how model weights and KV cache compete for GPU memory, why continuous batching matters when you have more than a handful of users, and how agent architectures like single-agent, workflow, graph, swarm, and supervisor patterns each solve different problems. You will learn how to instrument your agents with Langfuse for observability and cost tracking, when to use Ollama versus vLLM, how prompt caching can cut provider costs by up to 75%, and why GPUs should never sit idle. Episode two of three - the next episode covers deploying at scale. Timestamps 0:00 Welcome & Introduction 1:47 Du'An's New Role at Akamai Cloud 3:10 Data Privacy and the Case for Self-Hosted AI 7:21 Anthropic and OpenAI as the New Cloud Layer 12:48 Local Models for Specific Use Cases - Cancer Detection Example 15:02 GPU Memory Math - Weights, KV Cache, and Context Windows 19:32 Continuous Batching and GPU Time Slicing 20:03 Observability with Langfuse - Live Demo 27:44 Agent Architectures - Single Agent, Workflow, Graph, Swarm, Supervisor 36:36 Token Economics, Prompt Caching, and GPU Cost Planning 45:32 Ollama vs vLLM - Prototyping vs Production How to find Du'An: https://duanlightfoot.com https://www.linkedin.com/in/duanlightfoot/ Links from the show: https://langfuse.com/ https://github.com/akamai-developers/akamai-workshop-solution-architect-agent https://amzn.to/4bvHn1p https://vllm.ai/

Energy Terminal
SPARK Episode 31: Praveen Gorakavi

Energy Terminal

Play Episode Listen Later Jul 2, 2026 26:05


Centamil AI is building AI-embedded microreactors that enable process industries to transform their facilities for higher-intensity workloads. The company recently launched its Hayagreeva HEx, a cold plate that enables data centers to accommodate higher density GPUs driven by the AI revolution. Centamil offers a robust AI layer on top of its physical infrastructure, which optimizes cooling parameters in real time to adapt to the conditions at a facility. We spoke with Praveen Gorakavi, Founder and CEO of Centamil AI, to learn more about why he sees cooling as a critical bottleneck to AI infrastructure deployments. We discuss how the Hayagreeva cold plate is improving the energy efficiency of data centers by relieving the burden placed on energy intensive air cooling systems. Praveen also shares what inspires him as an entrepreneur and how he balances his roles as both an inventor and a businessman. And follow us on: Newsletter: https://www.energy-terminal.com/newsletter-signup LinkedIn: https://www.linkedin.com/company/energy-terminal Instagram: https://www.instagram.com/energyterminal/

The MAD Podcast with Matt Turck
Inside Nemotron & NVIDIA's AI Lab | Bryan Catanzaro

The MAD Podcast with Matt Turck

Play Episode Listen Later Jul 2, 2026 82:57


NVIDIA is a chip company. So why does it put hundreds of researchers on building AI models — and then give them away for free? Bryan Catanzaro is VP of Applied Deep Learning Research at NVIDIA and one of the people whose work quietly underpins modern AI: he helped create cuDNN (NVIDIA's first deep learning product), co-invented DLSS, and named and built Megatron, the framework behind how much of the industry trains large models. Today he leads Nemotron, NVIDIA's family of open models — and Nemotron 3 Ultra, released just weeks ago, is one of the strongest open-weights models to come out of the US.Matt Turck sits down with Bryan for a genuinely deep conversation: the real business logic behind a chip company building its own models, the state of open vs. closed AI, and whether the US is falling behind China in open models. Then they go inside Nemotron itself — four-bit (NVFP4) pretraining, hybrid Mamba-Transformer architecture, mixture-of-experts, multi-token prediction, and multi-teacher distillation — all explained in plain language. Plus a rare look at how a modern AI research org actually runs, what it was like working alongside Andrew Ng and Dario Amodei at Baidu, why Bryan doesn't believe in the singularity, and his contrarian case that open AI is safer than closed.A reference conversation for anyone trying to understand where AI is really headed.(00:00) — Cold open & Intro(01:33) — Is open source AI catching the frontier?(05:29) — Do closed labs blocking distillation slow open source down?(07:42) — Is the US falling behind China?(10:30) — Why companies actually choose open models(12:39) — A "crazy" 2008 bet: machine learning on GPUs(15:33) — Working with Andrew Ng and Dario Amodei at Baidu(17:41) — Coming back to NVIDIA: DLSS and the birth of Megatron(21:55) — The real reason NVIDIA builds its own models(24:28) — Is Moore's Law really dead?(33:37) — The Nemotron family: Nano, Super, Ultra(35:09) — Built for agents: why NVIDIA bets on speed(36:02) — How you train a 550B model in 4 bits(39:25) — Hybrid Mamba-Transformer, explained simply(42:31) — Mixture of experts — and why NVIDIA built NVL72 around it(47:26) — Why a 1-million-token context window matters(49:26) — Multi-token prediction: how the model predicts 5 tokens at once(52:47) — Multi-teacher distillation: teaching one model from many(58:01) — Where reinforcement learning goes next(01:00:16) — Inside NVIDIA's research org: "the mission is the boss"(01:04:03) — How NVIDIA decides who gets the GPUs(01:10:53) — Why NVIDIA still feels entrepreneurial after 33 years(01:12:58) — Why Bryan doesn't believe in the singularity(01:17:50) — The AI backlash(01:19:18) — The controversial case: open AI is safer than closed

Unchained
A Perp Venue Asked Her to Trade Her Own Benchmark. She Said No

Unchained

Play Episode Listen Later Jun 30, 2026 8:11


Carmen Li thought it was a joke when a perpetual futures marketplace asked her to become the market maker for her own index. It wasn't. In this segment from Bits + Bips: The Interview, she walks Steven Ehrlich through the requests that alarmed her, a daughter analogy for why trading your own benchmark destroys neutrality, the manipulation risks she sees in crypto's index practices, and why she insists any perp venue on her index be regulated and guardrailed. Host: Steven Ehrlich - Host of Bits + Bips and Head of Research at Sharplink Guest: Carmen Li - CEO of Silicon Data and Compute Exchange This clip is from a longer conversation on GPUs, compute markets, and crypto. Full episode here: https://www.youtube.com/live/rYDiPneJv20?si=fjS7bSd-bJ6c6tYb  We go live every Monday - subscribe to catch it live. Sponsors Cape: Your biggest crypto vulnerability isn't your wallet, it's your phone number. Cape is America's privacy-first mobile carrier that rotates your SIM identity daily and blocks SIM swaps before they happen. Get 33% off your first six months at https://cape.co/unchained (use code: UNCHAINED). Chapters

Remotely Curious
How agentic AI works behind the scenes to find the answers you need

Remotely Curious

Play Episode Listen Later Jun 30, 2026 31:35


When AI is at its best, the conversations can feel uncanny—almost magical in their accuracy, relevance, and speed. For that you can thank the AI agents that work together behind the scenes to search, reason, and sift through all your content to get you what you need to do your job. We talk with Jongmin Baek and Marta Mendez, two Dropbox machine learning engineers, about building conversational AI that's helpful, useful, and grounded in your team's shared context, so you can spend more time on the work that really matters. ~ ~ ~  Working Smarter is brought to you by Dropbox. Find, organize, and share your work—all in one place—with context-aware AI from Dropbox. You can listen to more episodes of Working Smarter on Apple Podcasts, Spotify, YouTube, Amazon Music, or wherever you get your podcasts. To read more stories and past interviews, visit workingsmarter.ai This show would not be possible without the talented team at Cosmic Standard: producer Ben Montoya, sound engineer Aja Simpson, technical director Jacob Winik, and executive producer Eliza Smith. Special thanks to our illustrator Fanny Luor, marketing consultant Meggan Ellingboe, and editorial support from Catie Keck.  Our theme song was composed by Doug Stuart.  Working Smarter is hosted by Matthew Braga. Thanks for listening!

Everyday AI Podcast – An AI and ChatGPT Podcast
Ep 808: OpenAI's limited release of GPT-5.6, Mythos starts slow reinstatement, OpenAI gets spicy and more AI news

Everyday AI Podcast – An AI and ChatGPT Podcast

Play Episode Listen Later Jun 29, 2026 31:29 Transcription Available


OpenAI has released GPT-5.6, but the majority of us will have to wait. ⌚After the Anthropic vs. U.S. Government feud, it now looks like we'll have to wait for frontier models. That wasn't the only big AI news headline that might change your company's AI strategy. Anthropic got the green light to roll out Mythos 5 to a select few, Google reportedly extended its strike team to catch up on coding and more. OpenAI's limited release of GPT-5.6, Mythos starts slow reinstatement, OpenAI gets spicy and more AI news -- An Everyday AI chat with Jordan WilsonNewsletter: Sign up for our free daily newsletterMore on this Episode: Episode PageToday's Episode on LinkedIn: Thoughts on this? Join the convo on LinkedIn and connect with other AI leaders.Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineupWebsite: YourEverydayAI.comEmail The Show: info@youreverydayai.comConnect with Jordan on LinkedInTopics Covered in This Episode:OpenAI GPT-5.6 Limited Release ExplainedOpenAI Sol, Terra, Luna Model NamingUS Government Restrictions on AI RolloutsAnthropic Mythos 5 Access and StandoffAnthropic Fable 5 Suspension DetailsGoogle Gemini 3.5 Pro Release DelayedGoogle's AI Coding Mid-Training InitiativeRaiseUS Nonprofit: AI Workforce AdaptationAnthropic Accuses Alibaba of Model DistillationOpenAI & Broadcom Unveil Jalapeno AI ChipKey AI Industry Partnerships & Product LeaksTimestamps:00:00 OpenAI's GPT 5.6 limited release06:14 OpenAI's new model release details09:54 Access suspension and negotiations13:18 Google's AI strategy and delays16:33 Anticipating Gemini 3.5 Pro Release20:40 Accusations of AI model theft24:48 OpenAI and Broadcom chip partnership28:05 OpenAI's recent developments and updates29:56 OpenAI and AI weekly updatesKeywords: GPT-5.6, OpenAI, Anthropic, Mythos 5, Fable 5, Frontier models, Gemini 3.5 Pro, Google, model rollout, limited AI access, AI safety, US government AI regulation, Sol model, Terra model, Luna model, Max reasoning mode, Ultra mode, sub agents, advanced AI benchmarks, coding workflows, cybersecurity, third-party AI analysis, government licensing, AI model guardrails, AI model democratization, model naming scheme, model availability, AI model security, jailbreak resistance, safety filters, general model access, trusted testers, AI export control, national security, Anthropic pullback, supply chain risk, defense department, AI industry competition, talent loss, AI coding, mid training, engineering agents, AI strike team, RaiseUS nonprofit, workforce AI disruption, technology policy, industrial scale distillation, Alibaba, AI model theft, China-US tech tensions, distillation attacks, Jalapeno AI chip, Broadcom, AI inference, custom hardware, data center GPUs, Microsoft, Meta, Elastic compute, AI-powered career navigation, Slack Claude Tag, Canva Grow 2.0, Copilot skills, AI ad creation, AI automation, DigitalOcean plugin, Apple hardware AI, smart glasses, Vision Pro, portfolio tracking AI, Google Finance, home smart speakers, voice AI, GLM 5.2, open source AI, US labor market AI effects, AI job disruption, model leaks, government approval delays.Send Everyday AI and Jordan a text message. (We can't reply back unless you leave contact info) Start Here ▶️Not sure where to start when it comes to AI? Start with our Start Here Series. You can listen to the first drop -- Episode 691 -- or get free access to our Inner Cricle community and all episodes: StartHereSeries.com Also, here's a link to the entire series on a Spotify playlist. 

The 7investing Podcast
Is Moore's Law Dead? Cerebras IPO, SpaceX Orbital Data Centers & Huawei Tau Scaling Explained

The 7investing Podcast

Play Episode Listen Later Jun 29, 2026 42:09


Three massive semiconductor and computing developments are reshaping the future of AI infrastructure — and 7investing's Simon Erickson sits down with Nick Rossalillo of Chip Stock Investor to break them all down. First up: Cerebras Systems (NASDAQ:CBRS), which just went public on May 13th at $185/share (~$40 billion valuation) and is now trading near $46 billion at 90x trailing sales. The company's Wafer Scale Engine, a chip that uses an entire silicon wafer rather than individual diced chip, was designed specifically for AI inference workloads that NVIDIA (NASDAQ:NVDA) GPUs struggle to handle efficiently due to on-chip SRAM limitations. With potential $20 billion in orders from OpenAI and access via AWS, Cerebras is real, but neither Simon nor Nick is buying at this price. Their rule: wait a year before touching a fresh IPO.Next, SpaceX's freshly-raised $75 billion gets put under the microscope, specifically Elon's ambition to build orbital data centers. Nick walks through the SpaceX diagram: 70-meter solar panel wingspan, laser-based networking between compute modules, and the massive engineering challenges around power, heat dissipation, and in-orbit assembly. This isn't imminent, Starlink's next-gen constellation comes first — but if Elon can crack the economics, it would rewrite the rules of data center infrastructure entirely.Finally, Huawei's Tau Scaling announcement: a new architectural approach to chip performance that bypasses the need for extreme ultraviolet lithography (which China can't access due to ASML export controls). Tau temporal scaling focuses on minimizing signal travel time between transistors using logic folding, new materials, and 3D stacking. Huawei claims it could reach 1.5 nanometer equivalent performance by 2031. Simon and Nick are skeptical — 381 chips in six years is not mass production, and TSMC (NYSE:TSM) will be well past that node by then but it's worth watching as China continues building workarounds to Western export restrictions.Whether Moore's Law is dead or simply rerouting, the chipmaking industry is more innovative and more investable than it's been in decades.Join the conversation on the 7investing discord: https://discord.com/invite/PT9ZQqdXXSWant access to all 7investing research? Join at 7investing.com/subscribe Follow Chip Stock Investor  @chipstockinvestor  and https://chipstockinvestor.com/0:00 - Introduction to 7investing and Chip Stock Investor0:54 - Is Moore's Law Dead? A review of scaling semiconductor manufacturing3:08 - Cerebras Systems' recent IPO. How is their Wafer Scale Engine different than NVIDIA's GPUs, how does this impact Moore's Law, and is the stock a buy today?21:12 - SpaceX's recent IPO. Elon wants to build and launch orbital data centers. How does SpaceX plan to use the $75 billion it raised, what are the challenges it faces, and is the stock a buy?28:16 - Huawei's Tau scaling. Could this new chip architecture make ASML's extreme ultraviolet lithography obsolete, and what are its chances of succeeding?39:39 - Outro, final thoughts, and audience questionsStocks & Companies Mentioned:Cerebras Systems (NASDAQ:CBRS)NVIDIA (NASDAQ:NVDA)AMD (NASDAQ:AMD)SpaceX (SPCX)Taiwan Semiconductor Manufacturing Company / TSMC (NYSE:TSM)ASE Technology Holding / ASE Group (NYSE:ASX)Vicor Corporation (NASDAQ:VICR)ASML Holding (NASDAQ:ASML)Applied Materials (NASDAQ:AMAT)Lam Research (NASDAQ:LRCX)Intel (NASDAQ:INTC)Amazon / AWS (NASDAQ:AMZN)Alphabet / Google (NASDAQ:GOOGL)AST SpaceMobile (NASDAQ:ASTS)Samsung Electronics (KRX:005930)Huawei — private (Chinese company)OpenAI — privateLuckin Coffee (OTC:LKNCY) — mentioned as cautionary example#Semiconductors #MooresLaw #CerebrasSystems #CBRS #AIChips #NVIDIA #SpaceX #OrbitalDataCenters #HuaweiTech #TauScaling #ChipStocks #AIInvesting #TechStocks #GrowthStocks #StockMarket #InvestingIn2026 #7investing #Simonerickson

The John Batchelor Show
S8 Ep1066: The Founding of OpenAI. Guest Author: Keach Hagey. In this opening segment, Keach Hagey discusses the January 2016 founding of OpenAI as a nonprofit research lab. Key figures included co-founder Greg Brockman and chief scientist Ilya Sutskever,

The John Batchelor Show

Play Episode Listen Later Jun 28, 2026 10:25


The Founding of OpenAI. Guest Author: Keach Hagey. In this opening segment, Keach Hagey discusses the January 2016 founding of OpenAI as a nonprofit research lab. Key figures included co-founder Greg Brockman and chief scientist Ilya Sutskever, a renowned researcher whose recruitment from Google signaled the lab's potential. Backed by a billion-dollar commitment from Elon Musk, Peter Thiel, and Jessica Livingston, the project was designed as a safe, non-commercial counterweight to Google's DeepMind. Operating initially out of Brockman's apartment, the team aimed to achieve Artificial General Intelligence (AGI) for the benefit of humanity. The technical foundation relied heavily on GPUs—hardware originally designed for video games—which proved essential for training the deep learning neural networks necessary for their research. This era was characterized by an ambitious, "pirate" spirit funded through YC Research to explore radical ideas outside the profit motive. 1JANUARY 1931

The Future of Work With Jacob Morgan
The Data Center Race Behind AI: Solidigm's SVP on Why Storage, GPUs, and Scale Matter

The Future of Work With Jacob Morgan

Play Episode Listen Later Jun 22, 2026 44:37


I talk with Greg Matson, Senior Vice President and Head of Marketing and Products at Solidigm, about the storage infrastructure powering the AI boom. We get into why AI training and inference require massive amounts of data, how GPUs, SSDs, and data centers work together, and why storage can't be an afterthought for companies building enterprise AI. We also discuss the scale of today's AI data center buildout, how Solidigm is using AI internally, and what this means for the future of work, education, and the skills people will need in an AI-first world.

Into the Impossible
Roman Yampolskiy: AI Can't Be Controlled — and We're Building It Anyway

Into the Impossible

Play Episode Listen Later Jun 15, 2026 83:00


Roman Yampolskiy has spent two decades trying to prove that superintelligent AI can be controlled. He couldn't. I invited him on to make his case. Subscribe if you want science with evidence, not speculation. Roman is a professor of computer science at the University of Louisville and one of the earliest researchers in AI safety. His book AI: Unexplainable, Unpredictable, Uncontrollable started as an attempt to solve the alignment problem. After decades of work, it became a proof that the problem cannot be solved. Not difficult. Mathematically impossible. I push back hard. We go after the Einstein test: can a large language model trained only on pre-1911 physics reproduce what Einstein did with the same data? We ran that experiment. It failed. Roman and I disagree about what that means. We also get into the halting problem and what it actually tells us about predicting smarter-than-human behavior, whether value alignment is a real problem or a well-funded category error, the case for a government moratorium on frontier model development, and why Roman thinks giving an AI agent access to your computer is the dumbest thing a smart person can do. What you'll hear: Whether AI control is mathematically impossible or just unsolved Why Roman thinks all current AI safety work is security theater What the halting problem actually means for superintelligence The alignment problem: real issue or well-funded category error Why Roman wants a moratorium on frontier model development What to tell your kids about careers in a world where Roman might be right If you listen to other people, the best you can become is average. CHAPTERS 00:00 Creating a mind without an off switch 01:34 Solving problems beyond our own intelligence 04:08 Einstein's epiphany and the limit of AI intuition 08:18 Assessing the Einstein test: Why the experiment failed 12:22 Path dependency: Are LLMs and GPUs our QWERTY? 16:10 The barriers preventing AI from solving physics 21:54 Safety vs. Capability: Why toddlers are safe but teens are not 23:06 The halting problem: Predicting agents smarter than us 25:58 The impossibility of a system proving its own integrity 28:18 Regulation: Genuine safety or a gift to oligarchs? 33:28 Is human cognition non-computable? Penrose vs. the field 39:00 Ethical duties: Must we treat AI with humanity? 43:00 From internet memes to monsters: Decoding the book cover 46:22 Customized realities: Can everyone have their perfect world? 49:50 Von Neumann probes and the panspermia hypothesis 55:02 Categorizing AI: The one version that should terrify you 58:22 Pause AI: The movement for a development moratorium 59:58 Career advice for kids in a post-professional world 01:07:58 Cross-examining Sam Altman 01:15:48 Roman's dream debate 01:19:50 Lessons for a younger self Substack: https://briankeating.substack.com Get the transcript, fascinating bonus content, and my Monday M.A.G.I.C. Message: https://briankeating.com/yt Have a .edu email and live in the USA? You automatically win a meteorite: https://BrianKeating.com/edu Subscribe: https://www.youtube.com/DrBrianKeating?sub_confirmation=1 Support Into the Impossible on Patreon, get my weekly M.A.G.I.C. Message, unfiltered bonus content, and live monthly Office Hours with me: https://www.patreon.com/drbriankeating Join this channel for perks, monthly Office Hours, and your name in the Member Roster at the end of every episode: https://www.youtube.com/channel/UCmXH_moPhfkqCk6S3b9RWuw/join Featured Guest: Roman Yampolskiy on Twitter/X: https://x.com/romanyam?lang=en AI: Unexplainable, Unpredictable, Uncontrollable: https://www.romanyampolskiy.com/books/ My books: Losing the Nobel Prize (memoir): http://amzn.to/2sa5UpA Think Like a Nobel Prize Winner: https://a.co/d/03ezQFu Focus Like a Nobel Prize Winner: https://a.co/d/hi50U9U Galileo's Dialogue (first-ever audiobook): https://a.co/d/iZPi9Un Twitter/X: https://x.com/BrianKeating Substack: https://briankeating.substack.com Blog: https://briankeating.com/blog Audio-only: https://briankeating.com/podcast #intotheimpossible #briankeating #AIrisk #artificialintelligence #aisafety #podcast #superintelligence #RomanYampolskiy Learn more about your ad choices. Visit megaphone.fm/adchoices