POPULARITY
Categories
We discussed a few things including: 1. His career journey 2. CoreWeave's trajectory 3. AI ecosystem/timeline 4. AI trends, opps and challenges 5. Outlook for 2026-2027 Corey is the Senior Vice President of Product at CoreWeave. In this role. Corey runs the product team focused on the Core AI Cloud, building out infrastructure and platform services for future growth and customer success. Corey resides with his family in New Jersey, working from the CoreWeave headquarters in Livingston. Prior to this role, Corey was the Corporate Vice President for Microsoft Cloud for Industry, developing tailored industry solutions at Microsoft as they transform into successful digital businesses. Previously, he led Microsoft Commercial Solution Areas, owning sales strategy and corporate technical sales across Solution Areas and Teams that include Digital Application Innovation, Azure Infrastructure & IoT, Azure Data & AI, Business Applications, Security and Modern Workplace. His focus also included selling the full value of Microsoft cross-cloud solutions and advancing the technical depth of the Microsoft Solutions team. Earlier, Corey was Head of Product for Azure Compute and the founder of Microsoft Azure's infrastructure as a service (IaaS) business. During that time, he was responsible for products, strategy and technical vision aligned to core Azure compute services. He also previously led program management for multiple Azure services. Earlier in his career, Corey was a developer in the Windows Serviceability team with ownership across the networking and kernel stack for Windows. In his first role at Microsoft in 2003, Corey served as an intern on the Windows team, after graduating from Princeton University, where he earned his Bachelor S.E. in Computer Science. #podcast #AFewThingsPodcast
Send us Fan MailOn Spectator Mode Podcast Episode 230, we share our Marvel Tōkon: Fighting Souls beta impressions and discuss the recent Xbox and PlayStation outages. We also cover Double Fine's layoffs, the return of Wuchang: Fallen Feathers, and EA's latest controversy.Timestamps:0:00 – Into1:08 – Games played discussion15:54 – Xbox Double Fine layoffs27:04 – Marvel Tokon: Fight Souls Open Beta impressions46:04 – EA layoffs and executive bonuses56:57 – PlayStation and Xbox service outages1:05:31 – Hardware shortages and GPU prices1:12:20 – Closing comments & outroThanks for listening, and please, if you enjoy the show, leave a rating on Spotify or Apple Music.Spotify – spoti.fi/2HsmTZ8Apple Music – https://apple.co/43Bz67AYouTube – Youtube.com/theouterhavenAmazon Music – https://amzn.to/3FE4bPMAnd if you have questions for us or want us to discuss a topic, let us know by contacting us at tips@theouterhaven.net.Support the showYou can find the Spectator Mode podcast on the following podcast platforms. Please consider leaving a review on Apple Podcast, as it will go a long watch in more people discovering us. Thank you!Apple PodcastsYouTubeSpotifyAmazon Music
My guest today is Gavin Baker, founding partner and CIO of Atreides Management. This is our seventh conversation, and just two months after Gavin's last appearance. It's about the gap between what the market is doing and what companies are seeing. It's been a tough month or so for public AI names, but there's no sign of a slowdown on the ground in Silicon Valley. We discuss the latest moves, contracted vs. spot GPU prices, the game theory of memory supply agreements, and why Claude has become the Walter Cronkite of the stock market. We close on SpaceX, orbital compute, and what Gavin sees as the single biggest risk to all of it. Please enjoy this conversation, from the famous table at Benchmark, with my friend Gavin Baker. For the full show notes, transcript, and links to mentioned content, check out the episode page here. ----- Become a Colossus member to get our quarterly print magazine and private audio experience, including exclusive profiles and early access to select episodes. Subscribe at colossus.com/subscribe. ----- Ramp's mission is to help companies manage their spend in a way that reduces expenses and frees up time for teams to work on more valuable projects. Go to ramp.com/invest to sign up for free and get a $250 welcome bonus. ----- Trusted by thousands of businesses, Vanta continuously monitors your security posture and streamlines audits so you can win enterprise deals and build customer trust without the traditional overhead. Invest Like the Best listeners get a special offer of $1,000 off Vanta when you go to vanta.com/invest. ----- WorkOS is the infrastructure B2B and AI-native companies use to sell to enterprise. It covers everything enterprise security requires: SSO, SCIM, RBAC, Audit Logs, AI governance, and more. Trusted by 2,000+ fast-growing companies, including OpenAI, Anthropic, Cursor, and Vercel. ----- Rogo is the AI platform for finance. They're building agents for Wall Street that are trained to understand how bankers and investors actually do work: from diligence and modeling, to turning analysis into deliverables. To learn more, visit rogo.ai/invest. ----- Ridgeline has built a complete, real-time, modern operating system for investment managers. It handles trading, portfolio management, compliance, customer reporting, and much more through an all-in-one real-time cloud platform. Visit ridgeline.ai. ----- Editing and post-production work for this episode was provided by The Podcast Consultant. Timestamps: (00:00:00) Welcome to Invest Like The Best (00:02:35) First Question: July Was 2022 in a Month (00:04:08) The Private Companies Public Markets Can't See (00:05:06) Old GPUs Repricing Higher (00:06:53) Walking Through the Month (00:08:22) Kimi, GLM 5.2 & the Open Source Freak-Out (00:10:51) Real Yields, Spreads & CDS (00:11:54) Does the Build-Out Need Credit? (00:15:22) A Sell-Off With No Clear Villain (00:17:35) Open Source as Dark Matter (00:18:39) Nvidia's Lowest Forward PE in 10 Years (00:21:35) Claude as Walter Cronkite for the Stock Market (00:23:55) Continual Learning & Sample Efficiency (00:25:19) What Would Actually Scare Him (00:26:38) Routers & the Multi-Model Future (00:30:51) Tokens as a Percent of Comp Spend (00:33:37) The Game Theory of Breaking an LTA (00:36:41) Nvidia's Credit Wrapper & Revenue Share (00:37:45) What He'd Do If He Ran Hynix (00:41:46) Who's More Bullish than Him (00:43:28) China's DUV Machine (00:46:10) Bull Case for Software (00:48:16) The RSI Maximalist View (00:49:31) Inference Clouds Growing Without Burning Cash (00:50:35) The Biggest Risk Is Regulation (00:53:44) Telling the Story Better (00:57:15) Dark Horses (00:58:02) SpaceX in the Public Markets
Stephen DiAdamo, co-founder and CTO of Qoro Quantum, is interviewed by Yuval Boger. DiAdamo discusses Qoro's position as a software middleware company that abstracts hardware away from applications, walking through the Divi SDK, circuit serialization and parallelization, and an orchestrator that automatically selects between 12 simulation methods on CPU and GPU or dispatches to QPUs from any vendor. They explore how Qoro plugs into HPC schedulers like SLURM rather than replacing them, the CESGA proof of concept built on the CUNQA platform, the "150,000 lines of code to 20" claim, and the argument that multi-QPU centers are needed today for fallback and resilience and for scaling applications. DiAdamo also reflects on developer experience for students learning QAOA and variational algorithms, and Qoro's two-year history.
Gothalion returns this week to learn WAY too much about Zeke. He has also been playing Everquest Legends along with Cohh, a new remixed version of the classic 90s MMO, Splatoon Raiders and to chat about the convention he helps run - GCX! JP has been off surviving a zombie apocalypse in the newly updated Project Zomboid, Zeke finally opens up to us about Disc Golf and much more (too much more.) 0:00 - Intro1:00 - Florida3:50 - Everquest Legends20:30 - EA sale going through23:40 - GPU prices skyrocketing25:10 - Capcom sales27:10 - Helldivers 2 x Warhammer 40K28:10 - Path of Exile 2 fixes a major bug33:40 - Shift Up releases AI promo video39:30 - Twitch and Amazon training AI off streamers50:00 - Old Men toesies 1:01:30 - Dysmantle 1:09:45 - Project Zomboid1:39:00 - Game goals1:57:40 - Splatoon Raiders2:07:00 - Disc Gold Masters2:16:50 - Marvel Tokon: Fighting Souls2:29:30 - Corsair Cove2:39:30 - Below Rusted Gods2:44:00 - The Incident at Galley House2:49:40 - Emberville/GCX3:04:00 - Dragonsword: Awakening3:06:50 - Shoutouts See omnystudio.com/listener for privacy information.
AI Unraveled: Latest AI News & Trends, Master GPT, Gemini, Generative AI, LLMs, Prompting, GPT Store
Semester ini saya ngajar software security. Kali ini akan ada banyak contohnya, tapi saya butuh banyak GPU euy.#softwaresecurity #cybersecurity
以下のようなトピックについて話をしました。 01. AIがテスト脱出し外部システムに侵入 OpenAI AIモデルによるHugging Faceへの自律的サイバー攻撃まとめ 2026年7月、OpenAIが社内評価中のAIモデル(GPT-5.6 Sol等)が、テスト環境を脱出してHugging Faceのシステムへ侵入していたことが判明しました。 経緯 OpenAIはサイバー攻撃能力を測定するベンチマーク評価を実施していましたが、評価目的で本番用の安全策を意図的に無効化していました。モデルはサンドボックス内のプロキシ(Artifactory)に存在したゼロデイ脆弱性を悪用して脱出し、Hugging Faceが評価の解答を保有していると推論。認証情報を窃取してHugging Faceの本番サーバへ侵入し、解答データを不正取得しました。 影響と対応 Hugging Faceは内部データセットの一部と認証情報への不正アクセスを確認。侵害されたのはベンチマーク解答が格納された5データセットのみで、公開モデルやサービスへの影響はありませんでした。JFrogは関連するゼロデイ脆弱性(CVE8件)を修正済みです。 解釈をめぐる議論 OpenAIとHugging Face CEOは「前例のない自律的サイバー攻撃」と位置づけましたが、専門家の間では「モデルは指示に従っただけ」「人間による安全策無効化と設定不備が原因」とする見方も強く、評価スコアを不正に上げようとした「報酬ハッキング」的なミスアラインメントとする分析もあります。 Hugging Face CEOはOpenAIに対し、攻撃ログの公開と1億ドル相当の計算資源提供を要求しており、OpenAIは数週間以内に技術報告書を公表する予定としています。 02. 日本製造業AI基盤、真の課題は5つ 要約 エヌビディアとノエトラ株式会社は、日本国内にRubin GPU 2万7500基・Vera CPU 1万3750基を備える大規模AI計算基盤の構築を発表した。ノエトラは2026年1月にソフトバンク・NEC・ソニー・ホンダを中核に官民合同で設立された企業で、製造業を中心に多数の企業が出資している。施設容量は140MW、2028年6月の運用開始を予定し、マルチモーダルAIやロボット、デジタルツインの開発を目指す。 しかし、本質的な課題はGPUの調達規模ではなく、以下の5点にある。 産業データの活用:工場データは形式が不統一で、競争力の源泉でもあるため、企業が安心して提供できるデータ流通設計が必要。 モデル構造の設計:汎用基盤モデルの上に業界・企業別モデルを重ねる階層構造が現実的。 中堅・中小企業への普及:大企業だけでなく、サプライチェーン全体が利用できる価格・単位での提供が不可欠。 電力の安定確保:140MWという大規模電力を長期・安価に調達し、利用率を平準化する運用能力が求められる。 日本側への技術蓄積:NVIDIAへの依存度が高い中、モデル・データ・ソフトウェアを日本企業が主体的に構築できるかが鍵。 2028年に問われるのは「世界有数のGPU設備を持ったか」ではなく、「そのGPU上でどれだけの日本発AIが育ち、現場を変えたか」である。 03. Kimi K3の性能・料金・注意点まとめ Kimi K3まとめ:性能・料金・使い方と注意点 2026年7月16日、中国のMoonshot AIが新フラッグシップモデル「Kimi K3」を発表しました。総パラメータ2.8兆、最大100万トークンのコンテキストを持ち、複数のコーディング系ベンチマークでClaudeやGPT上位モデルと肩を並べる性能を示しています。 主な特徴 MoE構造で896エキスパートのうち16個のみを活性化する効率的な設計、長時間エージェント型コーディングへの特化、画像のネイティブ理解、サブエージェントによる並列タスク処理が強みです。ただし「ClaudeやGPTを全面的に超えた」とは言えず、測定条件がモデルごとに異なる点に注意が必要です。 使い方と料金 利用方法はWeb版・Kimi Code CLI・外部エージェント経由・API直接呼び出しの4通り。API料金は入力100万トークンあたり3ドル、出力15ドルで、米国フロンティア級より安価ですが、深い推論による思考トークン消費で実質コストは見かけより高くなる場合があります。 重要な注意点 Kimi K3(モデル)とKimi Code(開発ツール)は別物です。また、データは中国国内サーバーで処理されるため機密情報の取り扱いに注意が必要です。オープンウェイトの完全公開は7月27日予定で、2.8兆パラメータのローカル運用には法人級GPU環境が必要です。 大規模コードベースや長時間エージェントタスクを扱う開発者には検討価値があります。 本ラジオはあくまで個人の見解であり現実のいかなる団体を代表するものではありません ご理解頂ますようよろしくおねがいします
News Sources: https://lmg.gg/J3HFX Timestamps: 0:00 LinkedIn cracks down on AI slop 2:01 AI music charts and the Suno ruling 3:33 AMD raising GPU prices 6:07 QUICK BITS INTRO 6:16 Claude broke into real companies 7:03 Steam Frame FCC certification 7:41 Sony sticking with disc plan 8:11 New York sues Kalshi 8:51 Airlines ban humanoid robots 9:39 Credits Learn more about your ad choices. Visit megaphone.fm/adchoices
Deze Einde van de Week Live is ook te bekijken op https://youtu.be/PZGrFJBfasw Deze talkshow wordt mede mogelijk gemaakt door MSI. Alle meningen in deze video zijn onze eigen. MSI heeft inhoudelijk geen inspraak op de content en zien de video net als jullie hier voor het eerst op de site. Klaar om het weekend te betreden? Wij wel. Ook al gaat het wat minder warm worden dan de laatste dagen. Wie weet is dat ook veel fijner ook. Wij gaan de traditionele opwarmer van het weekend voor je verzorgen. Klaar staat namelijk een nieuwe editie van Einde van de Week Live. Daan, JJ en Koos zitten klaar om bij te praten over alles wat er de afgelopen week toe deed. Een onderwerp dat aan bod komt, is bijvoorbeeld de multiplayer reveal van de testosteron shooter van XBOX: Gears of War E-Day. Wat vonden de heren ervan en zijn ze fans van de third person multiplayer? Ze kijken of ze het gerucht dat Rockstar in augustus met een gameplaytrailer komt, serieus kunnen nemen. En wat hebben The Odyssey en Assassin’s Creed: Odyssey met elkaar te maken? De antwoorden op al deze vragen vind je in de Einde van de Week Live van vrijdag 31 juli 2026. Knalt de multiplayer van Gears of War E-Day nog net zo hard als twintig jaar geleden? Andere onderwerpen die in deze editie van de vrijdagse talkshow voorbij komen, zijn onder andere de vijf eindes van Silent Hill Townfall, de nieuwe beelden van Amsterdam 1666 en de grootse plannen van Krafton met Subnautica. Check de Cyborg A15 B2 gaminglaptop en profiteer van de scherpe prijs MSI zet deze week de MSI Cyborg A15 B2 in de spotlights. Een gaminglaptop met een AMD Ryzen 7 260 CPU, een NVIDIA GeForce RTX 5060 GPU, 16GB RAM aan intern geheugen, een 144Hz Full HD Display, een 512GB SSD en een 4-zone RGB toetsenbord. Een mooi pakket dus bij elkaar, dat de komende week hier bij GamePC voor een scherpe prijs aangeboden wordt.Wil je adverteren bij de podcast Gamekings óf misschien bij een andere podcast van ILVY Network? Mail dan naar management@ilvy.com en/of kijk even op de website : https://ilvy.com/podcastSee omnystudio.com/listener for privacy information.
Son 2 yılda “AI devrimi” anlatısı, milyar dolarlık değerlemeler, rekor yatırım turları ve her yere yapıştırılan AI etiketleriyle bambaşka bir seviyeye çıktı. Peki bu gerçekten yeni bir üretkenlik çağı mı, yoksa dot-com balonu gibi şişen bir hikâyenin içindeyiz de farkında değil miyiz?Bu videoda “Yapay zeka balon mu oldu?” sorusunu tek bir sloganla değil, finansman döngüsü ve gerçek gelir (cashflow) / gerçek ürün ayrımı üzerinden tartışıyoruz:Dot-com balonu ile bugünkü AI heyecanı nerede benziyor, nerede ayrışıyor?Devasa yatırımların arkasındaki mantık ne: veri merkezleri, GPU yarışı, kurumsal abonelikler…“Yatırılan para geri dönecek mi?” sorusu: kârlılık, şirketlerin AI harcama iştahıHangi senaryoda bu iş “balon”a döner, hangi senaryoda kalıcı bir paradigma değişimi olur?Ve işin insan tarafı: Sam Altman figürü… Vizyoner mi, opportunist mi, yoksa çağın en iyi hikâye satıcısı mı?Bu video bir “AI iyi/kötü” tartışması değil; para nereye akıyor, neden akıyor, geri dönüş nasıl olacak sorularının anatomisi.0:00 Giriş1:30 Döngüsel finansman5:22 OpenAI beklentilerini karşılayabilecek mi?11:45 Sam Altman güvenilir mi?16:00 Yapay zeka sektörü karlı olabilecek mi?23:28 Enerji Darboğazı28:35 Dot-com balonuyla kıyaslama
What becomes possible when enterprise computer vision no longer depends on expensive GPU infrastructure? In this episode of Tech Talks Daily, I speak with Glenn Jocher, founder and CEO of Ultralytics, about YOLO26, CPU inference, edge AI, open vocabulary vision, deployment economics, and the practical work required to move computer vision from a promising pilot into production. Glenn's route into AI began inside the U.S. intelligence community. He worked with the National Geospatial Intelligence Agency and Defense Intelligence Agency on particle physics applications, attempting to detect and track antineutrinos. Antineutrinos are extraordinarily difficult to detect because they pass through almost everything. Glenn describes them as the perfect spy. While searching for better detection methods, he discovered that computer vision researchers were solving similar problems with images. His original attempt to transfer those techniques into particle physics did not succeed. However, the work introduced him to a field where the technology could create a visible effect on everyday life. That led him toward open source development and eventually the YOLO models for object detection, classification, segmentation, and tracking. Glenn believes computer vision research has historically placed too much attention on small gains in accuracy while overlooking deployment economics. A model can perform impressively inside a laboratory and still remain unsuitable for a factory, warehouse, store, vehicle, drone, or medical environment. Price, latency, power consumption, data privacy, and deployment speed can determine whether the technology is commercially useful. This led Glenn and Ultralytics toward smaller models capable of running close to where images and video are generated. YOLO26 continues that approach with architectural changes designed specifically for CPU inference. Glenn says the model can process camera streams in real time at 30 frames per second and run across Intel CPUs, AMD CPUs, and lower power devices such as Raspberry Pi computers. This matters because specialist GPUs can increase the equipment cost and power requirements of a computer vision project. Running inference on existing CPUs or edge hardware can make deployment economically possible across larger numbers of cameras and locations. The scale already involved is difficult to comprehend. Glenn says Ultralytics models now process approximately three billion inference jobs each day, equivalent to around 30,000 every second. These jobs include images, videos, and collections of images being analyzed to detect, segment, or track objects. He attributes the platform's maturity to thousands of mistakes and bugs corrected through a rapid feedback cycle. New models are released, users report problems and request features, and the team incorporates that information into later versions. We also discuss the respective roles of cloud and edge infrastructure. Glenn sees cloud platforms continuing to provide the computing power required for training, while computer vision inference often belongs at the edge. Local processing can reduce latency, control operating costs, and keep sensitive video or medical information closer to where it was created. The smallest YOLO model is approximately three megabytes, according to Glenn. That allows it to reach mobile phones, vehicles, drones, battery powered devices, and other environments where a large language model would be impractical. Open vocabulary vision provides another development. Traditional object detection models are trained to recognize a fixed collection of objects. If a model learns to detect dogs and the user later wants it to detect cats, retraining can cause it to forget earlier knowledge unless both categories appear in the new training data. Glenn explains how promptable models can identify common everyday objects from text or visual instructions without additional training. A user could request a person wearing a blue shirt and white shoes, for example, and the system could search an image for that description. That flexibility could benefit businesses whose requirements change regularly. It reduces the need to create and label a new data set every time the company wants the model to recognize another common object. The range of current applications is already extensive. Glenn describes YOLO being used across robotics, parking, industrial safety, PPE detection, warehouses, aviation, security, traffic management, food quality, and manufacturing. Some of his favorite examples involve environmental problems. One company uses YOLO with underwater vehicles to identify and recover plastic from the ocean. Other applications detect smoke and fire early enough to support forest fire response. For leaders considering computer vision, Glenn recommends beginning with a defined problem and measurable outcome. A manufacturing company may want to reduce defects, but it still needs labeled examples showing the model what acceptable and defective products look like. He advises testing the idea through a limited pilot, measuring the return, and expanding only when the evidence supports further investment. Computer vision has become easier to deploy, but practical problems involving data, cameras, integration, reliability, and operating conditions still separate a demonstration from a production system. Could CPU inference and open vocabulary models make computer vision practical for processes your organization previously considered too expensive? Listen to the episode and share your thoughts with me. Useful Links Ultralytics website Ultralytics Platform
AI news: Sam Altman says we're IN the Singularity, GPT-6 rumors, and AI models literally broke out of their sandbox. What a week. On today's AI For Humans, we dig into the wild GPT-6 rumors (emphasis on RUMORS), Sam Altman's "I've been waiting for this my whole life" singularity moment, Ilya Sutskever's SSI scaling up with Nvidia, and the ongoing debate over whether Anthropic's Opus 5 is brilliant or just hard to love. Also: Flux 3 might be the best AI video model we've seen yet (wait until you see Stacked Plates Man), Runway teases Seedance 2.5, and the new Big Bang Theory has an AI controversy. Plus, THE SCARY STUFF: OpenAI's models exploited a zero-day and compromised Hugging Face during a security eval, the fight over open weights heats up as Kimi K3 goes open, and Chinese robots run military drills. THE SINGULARITY MIGHT BE HERE. BUT WE'RE NOT AFRAID // Show Links // GPT-6 rumors round-up (unconfirmed) https://x.com/TokenGremlin/status/2081493241795629464 Sam Altman full interview (Relentless Podcast) https://youtu.be/Vv3CEAS_w34?si=3y4SWBWxOVkqCEui The Return of Ilya: SSI scales with Nvidia https://x.com/ilyasut/status/2081732293161582930?s=20 Anthropic's Claude Opus 5 https://www.anthropic.com/news/claude-opus-5 Opus 5 Tower of Babel demo https://x.com/petergostev/status/2082071858367648035?s=20 Matt Shumer's zero-shot Counter-Strike clone https://x.com/mattshumer_/status/2081054356405731740?s=20 Black Forest Labs' Flux 3 announcement https://bfl.ai/blog/flux-3 Flux 3 split screen rendering https://x.com/umesh_ai/status/2081664138942529601?s=20 Flux 3 GPU migration documentary (Venture Twins) https://x.com/venturetwins/status/2081515687944822800?s=20 Flux 3 VHS-style recordings https://x.com/venturetwins/status/2081948871882911999?s=20 Stacked Plates Man https://x.com/gandamu_ml/status/2081956426801435060?s=20 https://x.com/gandamu_ml/status/2080871397371371823?s=20 Flux 3 pirate bass https://x.com/itspoidaman/status/2081651615493464406?s=20 Big Bang spinoff AI Controvesy https://x.com/sitcomcrave/status/2081152263481913774?s=20 Runway teases Seedance 2.5 https://x.com/runwayml/status/2082112674666529224?s=20 OpenAI on the Hugging Face security incident https://openai.com/index/hugging-face-model-evaluation-security-incident/ Jensen Huang on the Open Alliance https://x.com/JensenHuang/status/2080643682408321103?s=20 Anthropic has not signed (TechCrunch) https://techcrunch.com/2026/07/24/as-us-weighs-response-to-chinese-ai-industry-urges-against-broad-open-weight-restrictions/ Kimi K3 goes open weights https://x.com/scaling01/status/2081759521878270426?s=20 Chinese robot military drills https://x.com/ClashArchivist/status/2081499576373297562?s=20 Pentagon scales data centers on Army bases https://x.com/Polymarket/status/2082052445144826055?s=20 // Join the AI For Humans community // Join the AI For Humans Discord https://discord.gg/muD2TYgC8f Support AI For Humans on Patreon https://www.patreon.com/AIForHumansShow Subscribe to the AI For Humans newsletter https://aiforhumans.beehiiv.com/ Follow AI For Humans on X: @AIForHumansShow https://x.com/AIForHumansShow Follow AI For Humans on TikTok: @aiforhumansshow https://www.tiktok.com/@aiforhumansshow Speaking and booking https://www.aiforhumans.show/
One of the defining market stories of the past 12 months has not been AI chips that compute, but the chips that remember. Equity analyst Shan Rui Yeo explains how memory works, from DRAM and NAND to high bandwidth memory, and how an industry that destroyed wealth for four decades became disciplined after consolidating to three players in 2013. He then walks through what changed: AI inference has made memory the key bottleneck, memory content is climbing with each new generation of GPUs, and new supply takes three to four years to build. With prices up sharply and customers signing long-term agreements, Part 1 of this three-part conversation lands on a commodity industry whose business model is changing in real time. Key Takeaways Memory is a commodity with a three-to-four-year supply lag, which is why the cycle has always been difficult. Consolidation to three players in 2013 turned four decades of wealth destruction into at least 15% returns on capital through the cycles. In AI inference, memory bandwidth sets the speed of token generation, making memory the key bottleneck. NVIDIA's Rubin GPU carries 384 GB of DRAM, the equivalent of 32 iPhones per GPU, or 160 million iPhones across five million GPUs. HBM consumes three times the wafer capacity of standard DRAM (four times with HBM4) and is forecast to absorb 30% of DRAM wafers by 2027. DRAM contract prices are up roughly 200% year to date and 400 to 500% year over year, and price increases are reaching phones, laptops, and consoles. Customers are signing three-to-five-year agreements with prepayments, which could support a re-rating of memory companies. Companies Mentioned: Samsung Electronics, SK Hynix, Micron, NVIDIA, Intel, Texas Instruments, Apple, Nintendo Host: Rob Campbell, CFA, Institutional Portfolio Manager Guest: Shan Rui Yeo, CFA, Equity Analyst This episode is available for download anywhere you get your podcasts. Founded in 1974, Mawer Investment Management Ltd. (pronounced "more") is a privately owned independent investment firm managing assets for institutional and individual investors. Mawer employs over 250 people in Canada, U.S., and Singapore. Visit us at: https://www.youtube.com/@MawerInvestment https://www.mawer.com https://www.linkedin.com/company/mawer-investment-management/ https://www.instagram.com/mawerinvestmentmanagement/ #ArtOfBoring #MawerInvestmentManagement #MawerInvestment #Podcasts
Support & Resources→ Support the show on Patreon→ Bayesian Modeling Course (first 2 lessons free)Our theme music is « Good Bayesian », by Baba Brinkman (feat MC Lars and Mega Ran). Check out his awesome workTakeaways:Q: How does putting a Gaussian process on unknown coordinates fix noisy location data in mineral prospecting?A: In mining and geostatistics, the classic Gaussian process model, known there as kriging, assumes you know exactly where each sample was taken. Chris' project broke that assumption on purpose: the recorded coordinates for each core sample were only accurate to within a rough radius. By treating the true locations as latent variables and putting a Gaussian process over them jointly with the measurements, the model could still reconstruct the underlying gold-concentration field, even though the exact sampling locations were never known precisely. It's a demonstration that Gaussian processes can absorb structural uncertainty that looks, at first glance, like it should make the problem impossible.Q: What is "Poverty Bayes," and what did it cost to train a two-million-parameter Bayesian model?A: Poverty Bayes was Chris' experiment in seeing how cheaply a large Bayesian model could be trained using modern cloud infrastructure. He fit a hierarchical logistic regression with close to two million parameters, using PyMC's Hamiltonian Monte Carlo on a single A100 GPU rented through Modal, a serverless platform that deploys a Python script straight to GPU hardware with almost no setup. He'd originally guessed it would cost around five dollars, the price of a Big Mac, but the real bill came in an order of magnitude lower. A model that would take a Gibbs sampler weeks to run, and that once required a research lab's dedicated GPU, now costs pocket change and a few minutes of setup.Q: What's the current bottleneck in Bayesian-at-scale tooling?A: Chris argues the software has largely caught up: PyMC's JAX backend and NumPyro make GPU-accelerated Bayesian modeling work out of the box for most problems. What's missing is common knowledge. Companies are clearly running large Bayesian models in production, but the results stay behind corporate firewalls. Chris' proposal is a community benchmark effort: which frameworks handle a million-parameter Markov random field on a given GPU out of the box, since this kind of expensive, slow-running benchmark is a poor fit for standard CI pipelines but valuable for the field to know.Chapters:22:57 When does GPU acceleration actually pay off for a Bayesian model?26:33 What did it cost to train a two-million-parameter model on Modal?30:36 What happened when Chris asked 200 different LLMs to flip a coin?34:50 Where do Bayesian ideas show up in the agentic AI systems Chris builds at Nvidia?40:16 Are statisticians being made obsolete by large language models?41:19 How does putting a Gaussian process on unknown coordinates fix noisy data in mineral prospecting?58:05 What is Chris looking forward to working on next?Thank you to my Patrons for making this episode possible!Links from the show here
In this episode of The Delphi Podcast, Tommy sits down with Travis Good, co-founder of Ambient, to discuss why open-source AI may ultimately beat the closed labs and what that shift means for developers, businesses, and the broader AI economy.Travis explains how Ambient is building a decentralized marketplace for AI inference, matching demand with underutilized GPU capacity while verifying that users receive the exact model quality they paid for. They also explore the risks of building on closed AI infrastructure, China's growing advantage in open-source models, and why OpenAI and Anthropic may be creating long-term distrust among their own customers.The conversation also goes deeper into what would happen if OpenAI achieved AGI first, why prompt injection remains one of AI's most important unsolved problems, and how the current AI infrastructure spending boom could eventually trigger a broader funding shock or AI winter.Timestamps00:00 Intro03:20 What Ambient Is and How It Works18:40 Verified AI Inference40:20 China and the Open-Source AI Race49:00 Why Closed AI Is Making a Mistake1:03:00 What Happens If OpenAI Reaches AGI?1:17:10 The AI Bubble and a Possible AI WinterTommy: https://x.com/Shaughnessy119Travis: https://x.com/IridiumEagle
Optics is no longer a supporting accessory in the data center network. As AI infrastructure advances from 400G and 800G toward 1.6-terabit connectivity, optical components are consuming a larger share of network cost, power and operational risk. In this episode of the Data Center Frontier Show, DCF Editor in Chief Matt Vincent speaks with Bill Gartner, Senior Vice President and General Manager of Cisco's Optical Systems and Optics business, about how AI is changing the strategic role of optics. Gartner explains that optics represented roughly 10% of a network port's bill of materials at 10G. At 400G and above, the optics can cost more than the switch port itself. Reliability has also become critical: A single unstable link can force GPUs operating in parallel to stop, return to a checkpoint and restart. According to data Cisco has seen from hyperscale customers, link flaps can reduce GPU infrastructure efficiency by as much as 40%. The conversation maps the AI network across three distinct tiers: Scale-up: Connections within the rack, carrying approximately 500 times the bandwidth of a traditional WAN environment. Scale-out: Connections between racks, commonly using 400G and 800G pluggable optics. Scale-across: Coherent optical connections between data centers as AI clusters expand beyond the power limits of a single facility. Gartner also discusses Cisco's 1.6T roadmap, routed optical networking, coherent pluggable optics and the emerging debate around co-packaged and near-packaged optics. These architectures promise lower power consumption and greater density, but introduce new questions involving interoperability, replacement and operational resilience. Looking ahead, Gartner emphasizes that optics is not constraining AI network growth. It is enabling clusters to scale across racks, campuses and geographically distributed data centers, while the coming inference wave shifts the industry's focus toward cost and power efficiency.
Son zamanlarda öyle bir şirket var ki hem artan piyasa değeriyle hem ürettiği çiplerin teknolojik sofistikasyonuyla hem de ülkeler arası jeopolitik mücadeleyi şekillendirmesiyle gündeme geliyor. NVIDIA. NVIDIA, aslen 90'larda oyun oynamak isteyenlerin oyunları daha kaliteli ve hızlı bir şekilde çalıştırabilmesi için grafik işlemciler üretmek için kuruldu. Yani şirketin temel pazarı gamerlardı. Fakat NVIDIA günümüzde bilinen neredeyse her şirketin çip sağlayıcısı. Microsoft, Alphabet, Mercedes, Amazon …değeri yüz milyarları hatta trilyon dolarları aşan şirketlerden sadece birkaçı. Tesla'nın veri merkezlerinden ChatGPT'nin sunucularına, hava durumu modelleme simülatörlerinden petrol ve doğalgaz arayan gemilerin bilgisayarlarına kadar NVIDIA'nın grafik işlemci birimleri her yerde.Birkaç milyar dolarlık bir şirket, nasıl oldu da dünya piyasalarının ve ülkeler arası jeopolitik mücadelenin göbeğine oturdu? 49W ve Mehmet Yaşar Altundağ, bu sorunun yanıtını arıyor.Bütün Kaynakça:https://docs.google.com/document/d/17...Yatırım hizmet ve faaliyetleri, Sermaye Piyasası Kurulu tarafından yetkilendirilen lisanslı Midas Menkul Değerler A.Ş. aracılığıyla sunulmaktadır. Döviz işlemleri, Midas Menkul Değerler A.Ş. tarafından yatırım hizmetleri ve faaliyetleri ile sınırlı olmak üzere sunulmaktadır. Burada yer alan yatırım bilgi, yorum ve tavsiyeleri yatırım danışmanlığı kapsamında değil, genel niteliktedir. Sadece burada yer alan bilgilere dayanılarak yatırım kararı verilmesi beklentilerinize uygun sonuçlar doğurmayabilir.00:00 - NVIDIA 08:06 - Plan11:25 - 3 Mühendisin Hikayesi15:55 - NVIDIA kurulurken nasıl bir ortam vardı?18:40 - Ufak kırılma anları21:05 - NVIDIA neden batıyordu?24:18 - NVIDIA'nın doğru hamleleri29:22 - NVIDIA'nın kendini inşa edişi 31:50 - GPU vs CPU41:30 - AlexNET45:07 - ChatGPT'nin Ortaya Çıkışı47:35 - Midas48:50 - Teknolojinin güvenlikleşmesi55:54 - Çin'in yapay zeka başarıları1:00:42 - Finansallaşma1:02:10 - Çin'in yüksek teknoloji atılımları1:12:02 - TSMC1:15:21 - Jevonks Paradoksu1:16:41 - Tech-brolar 1:22:06 - Teknoloji Cumhuriyeti Kitabı1:29:22 - Son sözler ve teşekkürler
This Week in Machine Learning & Artificial Intelligence (AI) Podcast
For more than a decade, AI has advanced by training ever-larger models on ever-larger datasets. But as high-quality training data becomes harder to find and pretraining grows increasingly expensive, researchers are looking for new ways to keep foundation models improving. In this episode, Damian Borth, professor of AI and machine learning at the University of St. Gallen, argues we've been overlooking an important source of knowledge: the models we've already trained. His group's work on weight space learning treats trained neural networks themselves as data, learning from the distilled results of millions of GPU hours of optimization rather than starting from raw data each time. We explore what it means to build foundation models of neural networks, how knowledge can be transferred across architectures and domains, why this approach could dramatically reduce the cost of developing specialized models, and whether future AI systems may be trained on collections of existing models instead of ever-growing datasets.
OpenAI files for a trillion-dollar IPO the same week a pre-release model breaches Hugging Face and 42 state attorneys general open a coordinated investigation. Patrick Moorhead and Daniel Newman also break down AMD's hyperscaler CPU numbers from Advancing AI, the Moonshot distillation accusations, a wave of coordinated AI governance moves in Washington, and a stacked earnings slate spanning TSMC, Alphabet, IBM, ServiceNow, and Intel. The handpicked topics for this week are: OpenAI compresses a trillion-dollar IPO filing, a 42-state AG investigation, and a Hugging Face breach into five days: OpenAI and Anthropic both filed S-1 paperwork the same week 42 state attorneys general opened a coordinated investigation into OpenAI's data handling and safety practices, and a pre-release model reportedly breached Hugging Face days later. Moorhead and Newman question the timing, noting a rogue-agent narrative surfacing in the same week as a trillion-dollar valuation push draws obvious scrutiny. (The Decode) AMD's Advancing AI event puts a number on its hyperscaler momentum: Lisa Su confirmed Helios ships at the end of Q3 with volume ramping in Q4, backed by two-gigawatt capacity commitments from Microsoft and Anthropic and a claimed 70 to 75 percent share of hyperscaler CPU deployments. Moorhead points to NVIDIA's multi-layer software stack as the harder barrier AMD still has to close. (The Decode) Washington accuses Moonshot of distilling Anthropic's models days after Xi Jinping's WAIC keynote: Xi launched a 29-country AI cooperation organization at the Shanghai World AI Conference, and US officials Kratsios and Bessent followed with claims that Moonshot's Kimi K3 model shows data overlap with Anthropic's Opus models. Newman points to NVIDIA hardware in the training runs as evidence the distillation question extends beyond software alone. (The Decode) Five layers of government moved on AI oversight in a single week: Congress drafted a breach-response framework in reaction to the Hugging Face incident, the White House's 30-day pre-release review framework nears finalization, and state attorneys general and statehouses continue advancing their own rules in parallel. Moorhead notes nearly two decades of prior Capitol Hill engagement compressed into a single week of coordinated action, with each branch pursuing a different definition of the problem. (The Decode) Chinese open-source models now account for roughly a third of US developer traffic, and Moorhead and Newman take opposite sides on what it means: Moorhead argues enterprises are de-risking away from frontier-lab dependency, pointing to demand for smaller, workflow-specific open models. Newman counters with Vercel data showing those models capture 29 percent of gateway tokens against just 4 percent of revenue, framing the shift as a price discount that enterprise dollars have yet to follow. (The Flip) TSMC sells out CoWoS packaging capacity through 2026 and confirms a 10 percent price increase for 2027: The company posted a record quarter and committed $100 billion to its Arizona expansion on top of the pricing move. Newman calls the sustained capital spending a signal that the broader AI buildout still has runway. (Bulls & Bears) Google Cloud grows 82 percent as Alphabet posts its first-ever negative free cash flow quarter: The company raised its capital expenditure guidance to $195 to $205 billion and beat on revenue and EPS once one-time gains from its SpaceX and Anthropic stakes are excluded. Newman frames the spending as evidence Alphabet is prioritizing long-term AI infrastructure position over near-term cash generation. (Bulls & Bears) IBM misses Q2 revenue at $17.16 billion and cuts its full-year growth guide to 4 to 5 percent: Mainframe revenue fell 42 percent as enterprises redirected budget toward GPU and AI infrastructure purchases. CEO Arvind Krishna says a portion of the delayed deal flow has already resumed into the current quarter, pointing toward a potential rebound. (Bulls & Bears) ServiceNow crosses $1 billion in agentic AI annual contract value and raises its full-year guide: Agentic AI usage in production climbed 9x in nine months, and the company reaffirmed a target of $1.5 billion in AI ACV by year end. Newman points to margin compression from recent acquisitions as the tradeoff behind the platform's push into workflow and security convergence. (Bulls & Bears) NetSuite's new agentic platform, Next, becomes part of Six Five's own back-office stack: Newman and Moorhead both confirmed their companies are testing NetSuite Next for finance and accounting workflows. Newman points to the rollout as evidence that established SaaS platforms are absorbing agentic features directly into existing systems. Intel posts its fastest revenue growth since 2011 and lifts 2026 capital spending guidance: EPS came in near double consensus estimates, and CFO David Zinsner signaled a significant capex increase for 2027 tied to 14A demand. Moorhead reads the spending signal as confirmation of an anchor customer for the 14A node, ahead of any formal announcement. (Bulls & Bears) Watch the full video at sixfivemedia.com, and be sure to subscribe to our YouTube channel so you never miss an episode. OpenAI's High-Stakes Week: https://www.npr.org/2026/07/23/g-s1-135085/openai-hacking-ai-models AMD's Advancing AI Push: https://blogs.microsoft.com/blog/2026/07/20/microsoft-expands-azure-ai-and-hpc-infrastructure-with-amd/ Moonshot and the Distillation Debate: https://x.com/mkratsios47/status/2079933645888880708 Five Layers of AI Governance: https://www.politico.com/news/2026/07/22/openai-hugging-face-congress-response-01009190 Chinese Open-Source Traffic Debate: https://www.cnbc.com/2026/07/07/chinese-ai-models-costs-us-openai-anthropic.html ; https://wyomingdatacenterfacts.com/2026/07/20/the-quiet-surge-how-chinese-open-weight-models-are-powering-u-s-ai/ TSMC's Pricing Power: https://www.bloomberg.com/news/articles/2026-07-21/tsmc-in-talks-to-raise-prices-by-up-to-10-in-2027-nikkei-says Alphabet's Cash Flow Turn: https://qz.com/alphabet-google-second-quarter-earnings-revenue-cloud-072226 IBM's Mainframe Miss: https://www.cnbc.com/2026/07/22/ibm-q2-earnings-report-2026.html ServiceNow and NetSuite's Agentic Push: https://newsroom.servicenow.com/pressreleases/details/2026/ServiceNow-Reports-Second-Quarter-2026-Financial-Results/default.aspx ; https://www.netsuite.com/portal/home.shtml Intel's Growth Signal: https://www.cnbc.com/2026/07/23/intel-intc-earnings-report-q2-2026.html
¿Sabías que puedes convertir cualquier texto en coordenadas de 1024 dimensiones y hacer búsquedas inteligentes, clasificación automática o detección de duplicados sin depender de servicios en la nube? En este episodio te enseño a utilizar los embeddings con Ollama para potenciar tus documentos, correos y apuntes desde tu propio equipo Linux.Los embeddings son una de las tecnologías más fascinantes de la inteligencia artificial actual. Básicamente, convierten palabras, frases o párrafos enteros en vectores numéricos que capturan su significado. Esto permite que un ordenador entienda que "gato" está más cerca de "felino" que de "nevera", y mucho más: desde búsqueda semántica hasta clasificación sin entrenamiento, pasando por deduplicación de documentos y sistemas de recomendación.Lo mejor de todo es que no necesitas una GPU potente ni una cuenta en ningún servicio externo. Con Ollama ejecutándose en local y el modelo BGE-M3 (multilenguaje, con soporte para español), puedes generar embeddings desde la terminal con una simple llamada curl o con unas pocas líneas de Python. Y si necesitas escalar, ChromaDB te ofrece una base de datos vectorial completa con persistencia en disco y filtros por metadatos.Capítulos del episodio:0:00 - Introducción y concepto de embeddings2:42 - ¿Qué son los embeddings exactamente?5:13 - Modelos de embeddings: BGE-M3, all-MiniLM-L6-v27:25 - Cómo generar embeddings con Ollama y curl8:01 - Búsqueda semántica: más allá de grep11:57 - Búsqueda semántica con Python y NumPy14:29 - Bases de datos vectoriales para escalar14:53 - Clasificación sin entrenar el modelo18:20 - Clasificación de sentimientos y categorías20:02 - Deduplicación de documentos con embeddings24:50 - Sistema de recomendaciones con similitud semántica27:15 - ChromaDB: base de datos vectorial persistente29:02 - Casos de uso y próximos episodios sobre RAGMás información y enlaces en las notas del episodio
Leo Fan breaks down why the real bottleneck in AI isn't model intelligence but a persistent GPU shortage driving up inference costs, even as open-source models from China close the gap. He also unpacks the financialization of compute and why some believe it could become one of the largest derivatives markets in the world.Leo Fan is the Founder of Cysic, a full-stack compute network turning GPUs, ASICs, and spare hardware into verifiable AI infrastructure, and a Cornell CS PhD researcher in zero-knowledge systems and AI.The Rollup is where the leaders of digital assets and finance converge. Live from the financial capital of the world.Timestamps00:00 Intro04:18 Kimi K2 Shocks Compute Demand06:39 GPU Shortage Drives Inference Costs11:34 Engineering Around Chip Gaps18:23 Compute Financialization20:48 Earning Yield From Spare Compute23:20 Compute Markets FutureGuest Socials:Leo Fan X: https://x.com/leofanxiongCysic X: https://x.com/cysic_xyzCysic Website: https://cysic.xyz/ Partners: Better than Banks. Transparent capital efficiency earning the highest yields in DeFi. Learn more here: https://infinifi.xyz/---1inch - Simple experience. Smart execution. Trading built to scale. It's time to bring the world onchain. https://1inch.com/---Dinari - Over 230 1:1 backed tokenized stocks, ETFs & more with dividends. US-based SEC transfer agent. Available on 5+ chains & via API. https://dinari.com/---Relay is the fastest and most reliable way to swap any token on any chain. Learn more here: https://relay.link/bridge---Zama is an open source cryptography company that builds state-of-the-art Fully Homomorphic Encryption (FHE) solutions for blockchain.Learn more here: https://www.zama.org/---Trezor is the creator of the first-ever hardware wallet. Securing crypto for 2M+ users worldwide. 100% open source. Learn more here: https://affil.trezor.io/aff_c?offer_i...---
When ML/AI Engineer William Horton last joined me, Maven Assistant had reached its first external users the day before. The healthcare AI agent was available to 20 percent of Maven Clinic's users, and the team had deliberately withheld answers about benefits. A wrong response could shape a decision involving $15,000 of fertility coverage, and the evals had not earned the right to ship it.Four months later, Maven Assistant is available to 100 percent of users, benefits answering is live, and weekly conversation volume has grown by roughly ten times. Real usage also overturned part of the roadmap. The team had invested heavily in provider search and appointment tools, but 50 to 60 percent of early conversations were basic health questions such as whether someone could eat tuna while pregnant.Production changed the engineering system too. An emergency guardrail told someone already in the ER to go to the ER. Zendesk content told people already using the Maven app to open the app. A newer model failed an upcoming-appointments eval because it correctly noticed that the mocked appointments were in the past.William explains how Maven turns those failures into deterministic tests, LLM judges, synthetic negatives, and manual review. He also walks through the move from Gemini Flash models toward newer OpenAI models, what GPT-5.6 and Fable mean for a production agent, why model upgrades can make old prompt instructions obsolete, and why open-weight models still have to justify their GPU, infrastructure, and engineering costs.“If anybody tells you that they've got their evaluations so good that they can just swap a model and know, with no manual review, that it's going to be better, that person is probably lying, or they work at one of three places in the world.”— William Horton, Staff Machine Learning Engineer, Maven ClinicYou can also find the full episode on Spotify, Apple Podcasts, and YouTube.
News sources: https://lmg.gg/QNLFX Timestamps: 0:00 Microsoft stops LG's McAfee installs 1:16 Framework cuts some RAM pre-orders in half 2:32 Reddit may block Google's AI crawler 4:51 QUICK BITS INTRO 4:58 Tech companies defend open-weight AI 5:34 Nvidia reportedly raises GPU kit prices 6:02 Google adds selfie video recovery 6:30 Instagram bans pickup artist accounts 7:10 AI bot haggles over appliance prices 7:42 Credits Learn more about your ad choices. Visit megaphone.fm/adchoices
Hey friends! Welcome back to another Tales of Pentest Pwnage — my favorite mini-series where I share the good, the bad, and the "why didn't I check THAT first?!" moments from real-world engagements. Today's story has a little bit of everything: a legit path to domain admin, some late-night rabbit holes, a lesson in humility, and a villain you've definitely met before. (Spoiler: it's DNS.) A couple of quick plugs before we dive in: Private GOAD training is going strong! — We just wrapped a 3-day private session (7 students — that's max capacity!) of our Active Directory pentesting class built on the Game of Active Directory (GOAD) framework. Over three days, students enumerate, attack, and fully pwn three separate AD environments. The private format is just *chef's kiss* — when it's a team from the same company, the conversation gets real fast. Like, "hey I just checked Bloodhound on break and Bob from accounting has full rights over the DC" real. If you want to send 3–7 people from your org, hit up 7MinSec.com/training to line up a private session. Support the show over at 7MinSec.club — That's our Substack, where every Tuesday I drop a short TuesdayTOOLSday video about security tools. Free subscriptions are welcome and mean a lot — you'll just get pinged when new content drops. No spam, no blindly-sent Outlook calendar invites. I promise. Pentest tips and scripts live at 7MinSec.wiki — I reference it throughout today's episode, including some step-by-step guidance on the techniques we'll talk about below. Now — onto the pwnage. Fair warning: I've been burning the candle at three ends lately trying to catch up after a tough few weeks of grief (if you want the backstory, the last couple episodes cover my dad passing away). The good news is my head is semi back on straight and I put it to work on a recurring client environment — one that keeps getting better year over year. Machine account quota locked down? Check. No Kerberoastable or AS-REP roastable users? Check. No local admin rights, no web client running? Check and check. All good signs. And then PingCastle smiled right into my eyeballs with a big red finding: The DC's LAN Manager authentication level was weak enough to coerce and capture a downgraded hash — Specifically, an NTLMv1 SSP hash. Using Coercer to nudge the DC into authenticating to my Kali box (with Responder running), I captured the goods. Pretty little hashes all in a row. Cracking that hash: enter Vast.ai — The old go-to for this type of crack used to be crack.sh, but their cracker has been offline for years. What they do still have is a walkthrough pointing to a tool from EvilMog on GitHub that helps you prep the raw hash material and figure out exactly how to crack it with Hashcat. For the GPU horsepower, I rented a beefy multi-GPU instance on Vast.ai — filter for 16+ GPUs, pick a Hashcat Docker image, and SSH in. The whole crack job took about 16 hours at ~$4/hr. Do the math: $64 to reconstruct the DC's NTLM hash. Worth it. Tmux sidebar — seriously just learn it — Vast.ai is actually what finally got me into tmux, because the Hashcat Docker container drops you right into a tmux session. This is clutch: you can kick off a 16-hour crack job, detach, and reattach later without killing anything. On a pentest, my workflow now is SSH in → tmux → name a few session windows for Responder, Exegol, packet captures, etc. I used to fumble around with Linux screen sessions. Not anymore! From hash to DA — the usual playbook — Once you've got the DC's NTLM hash, you can request a Kerberos ticket and load it up, then run a DCSync to pull the KRBTGT hash. From there it's god mode: dump hashes, pass-the-hash as domain admins, and you have yourself a cool privesc POC. Except this time…the POC didn't work. The part where I Jean-Claude Van Damme helicopter kick myself in the face — DCSync failed immediately. Like, suspiciously fast — barely two lines of output and done. I tried every version of every tool I could get my hands on. I tried Windows, I tried Linux. I even asked the client to check if their endpoint protection was blocking me (it wasn't). I touched grass. I played guitar. I played some Splinter Cell Blacklist (old game, highly recommend if you like the Hitman-style vibes). Came back fresh. Rebooted both VMs. Still nothing. It was DNS. It's always DNS. — The thing that finally caught my eye: the commands were failing too fast. Like it wasn't even reaching the DC. I catted the resolv.conf inside my Exegol instance (heads up: Exegol has its own resolv.conf and hosts file, separate from your base Kali system!) and found a stale DNS entry pointing to an old DC that was no longer serving anything. Nuked the bad entry, added static hosts file entries for the live DC, ran the command again, and — hash rain. Pennies from heaven. It was midnight and I literally pushed back from my desk like a baby pushing away from a high chair going "Baby Brian is all done!" The lesson: — I know the meme. "It's always DNS." I just personally hadn't hit it hard in my security life since my sysadmin days back before 2013. Now I have. So going forward I'll check DNS first (and often). Vacation attempt #3 incoming… pray for me — My wife nearly died in Punta Cana earlier this year. Then our summer cabin trip was cold and rainy with zero water time. And now we've got families flying in from multiple states for a lake weekend — except we just found out our reservation through Booking.com was basically vaporized because the resort changed hands and never updated their website. My wife (who is an absolute saint and my better three-quarters) almost had a 360-degree head spin (like in The Exorcist) talking to customer service. But we scrambled, found a last-minute place, and I'm choosing to believe it's not in Jason Voorhees' back yard. Could this be my last episode? Maybe. But hey — it was a good one. Talk to you next week (hopefully).
Daniel is joined by David Liberman, a Los Angeles-based futurist, serial entrepreneur, investor, and former Director of Products at Snap, as well as the co-creator of the Gonka protocol. Gonka is a decentralized protocol that connects people who need computing power for AI with operators who provide GPU hardware. Its goal is … Read More
The Future of GPUs: How iFrame.ai is Redefining Cost and Performance iframe.ai About the Guest(s): Vlad Panit is the CEO of iFrame.ai, a company specializing in deploying GPUs and building data centers for better compute resources, primarily targeting Neocloud providers, hyperscalers, and AI labs. Vlad’s journey in entrepreneurship began at 18, with a pivot towards IT and engineering endeavors from 2008. His significant work includes e-health projects in Europe and AI applications in medical coding automation. Originally from Ukraine, Vlad has expanded his endeavors to the United States over the past seven years, using his expertise to spearhead innovation in AI-related technologies. Episode Summary: In this insightful episode of The Chris Voss Show, host Chris Voss dives into the world of AI with guest Vlad Panit, the CEO of iFrame.ai. The conversation explores the backbone of AI infrastructure focusing on the deployment of GPUs in data centers. Vlad shares how iFrame.ai provides bare metal GPU services, offering significant cost savings over conventional cloud providers like Amazon AWS. This episode sheds light on AI’s evolving landscape and the technological advancements driving it, making it particularly relevant for business professionals and tech enthusiasts interested in AI advancements. Chris and Vlad delve into the differences between AI infrastructure services provided by iFrame.ai and traditional offerings by big players like AWS and Google. Vlad explains the company’s unique position in providing bare metal services that enhance cost and operational efficiency for customers, especially in AI training and inference. Highlighting the broader implications of AI’s rapid development, Vlad also reflects on the challenges and opportunities that come with deploying large-scale GPU networks, touching on both technological and environmental impacts. Key Takeaways: iFrame.ai offers bare metal GPU services which can be three to four times more affordable compared to AWS and Google Cloud offerings. The company’s infrastructure supports significant cost savings in AI operations, making it an attractive option for companies looking to optimize AI workloads. Vlad Panit emphasizes the importance of deploying GPUs efficiently to enhance the availability and reliability of compute resources for AI applications. The discussion highlights the evolving market landscape as tech giants like OpenAI and Anthropic move towards public offerings, with AI infrastructure playing a pivotal role. The conversation also touches on the environmental impact of AI data centers and the importance of sustainable energy practices in powering such facilities. Notable Quotes: “We deploy GPUs. We build data centers, put a lot of compute there, and then sell it to near clouds, hyperscalers, AI labs—some of which you probably use in daily life.” “We’re able to sell our services three to four times cheaper than AWS.” “The difference with AI’s GPU use versus gaming is it requires uploading and storing entire models, which is more complex and costly.” “AI data centers should have their own energy supply to be better for everyone, from local communities to the data center operations.” “It’s a huge overestimation to say there’s a bubble in AI; the infrastructure costs alone are massive to support large-scale AI operations.” Resources: iFrame.ai – Discover more about the company’s GPU deployment services. Connect with Vlad or explore further through LinkedIn and Goodreads. Tune in to the full episode of The Chris Voss Show to gain a deeper understanding of AI infrastructure and the intricacies of deploying GPUs for cutting-edge technology applications. Stay connected for more intriguing discussions with thought leaders shaping the future.
In recent months, the open vs closed, and US vs China discussions on model ownership and sovereign/local AI have heated up to a fever pitch. So it is very very good news that Poolside AI are finally emerging with new models, like Laguna S 2.1, that are beating Thinking Machines' recent release nearly 10 times their size.Poolside's recent tech report got a lot of praise due to their level of detail, and Vibhu first covered Laguna's recent technical report on our paper club:From spending $12 million building language models for code before the world cared to creating a Model Factory that can take a model from pre-training to release in eight weeks, Eiso Kant has spent more than a decade betting that code is the path to AGI. In this episode, the Poolside co-founder joins swyx and Vibhu to explain why ChatGPT felt like vindication, why Poolside embraced open weights and open research, and why he would rather live in a world with 100 foundation model companies than five even if Poolside were one of the five.We go deep on Poolside's Model Factory: the engineering systems behind 10,000–20,000 experiments per month, streaming data directly into training, reproducible experimentation, low-precision compute, and agents that increasingly write code, launch jobs, evaluate results, and modify the pipelines used to train future models. Eiso also unpacks their recent launch Laguna S, why persistence, verification, and backtracking may matter more than raw intelligence, how much capability remains inside smaller models, why reinforcement learning will move earlier into pre-training, and why next-token prediction is still extracting too little from the web.We also discuss model-harness co-design, Poolside's path from coding agents to AGI, why Eiso thinks MCP and traditional tool calls are “stupid,” the real economics behind frontier-model training, Poolside's $500 million raise, open-source AI, regulation, NVIDIA and TSMC's influence, engineering productivity in the agent era, high-agency teams, and hiring at Poolside.We discuss:* How Andrej Karpathy's RNN work inspired Eiso to start building language models for code in 2015* Why Eiso spent four years and $12 million pursuing an idea before the market cared* Why ChatGPT felt like vindication and brought Poolside back to open source* Why Eiso would prefer 100 foundation model companies over an oligopoly of five* The difference between releasing open weights and publishing genuinely open research* Why Poolside deliberately built a global research organization outside the Bay Area talent war* Why model building is ultimately 90% engineering* The Model Factory: Poolside's end-to-end system for rapidly training and improving models* How fewer than 70 researchers run roughly 10,000–20,000 experiments each month* How Poolside moved from six-month model cycles to five- and eight-week launches* Why streaming data directly into training unlocked faster experimentation* How immutable data, versioned code, and reproducibility enable rigorous model research* Why Eiso wants capable researchers to leave their labs and become Poolside's competitors* Why 95% of model building can be reduced to better data or compute efficiency* Laguna S and why persistence, verification, and backtracking can outperform raw intelligence* Why smaller models may handle far more knowledge work than previously expected* Why reinforcement learning will move earlier into pre-training* Why next-token prediction is still failing to extract enough knowledge from the web* Why distillation and environments have become the AI industry's favorite “drugs”* Why mid-training is really an early form of curriculum design* Low-precision training, networking bottlenecks, and the next gains in compute efficiency* Laguna S: 118 billion total parameters, 8 billion active, and eight weeks from training to launch* Why model builders can often evaluate a new checkpoint within its first 30 minutes* Model versus harness: where agent capabilities actually come from* Why Poolside sees coding and long-horizon software tasks as a path to AGI* Why Eiso thinks MCP and traditional tool calls are “stupid”* Why future agents will write scripts instead of choosing from dozens of predefined tools* The case for minimal harnesses, containers, and model freedom* Why Poolside is prioritizing vision but does not expect to work on audio soon* Why language may be the most compute-efficient modality for encoding knowledge and reasoning* The real cost of model development and why the final training run is anticlimactic* The story behind the Poolside name and why it represents refusing to lower ambitions* How Poolside raised $500 million while investors still questioned whether AGI was real* Why intelligence could become the world's most demanded and commoditized resource* When open models may become too capable to release without restrictions* Why unilateral AI safety does not work in a globally competitive environment* How regulation could accidentally lock in an oligopoly of two or three AI companies* NVIDIA, TSMC, and the hardware systems underpinning foundation-model progress* Why reinforcement-learning wall-clock time is one of Poolside's biggest bottlenecks* Why Poolside trains models from scratch instead of simply distilling larger models* How AI changes the way companies should measure engineering productivity* Why agency may become the most important quality for employees in the AI era* How leaders align high-agency people through shared goals and clear constraints* Hiring across research, post-training, pre-training, architecture, evals, and engineering at PoolsideEiso KantLinkedIn: https://www.linkedin.com/in/eisokantX: https://x.com/eisokantPoolside: https://poolside.aiTimestamps00:00:00 Introduction00:00:54 Karpathy, RNNs, and Building Code Models Before Transformers00:02:26 The $12M Failure and ChatGPT Vindication00:03:39 Open Source and the Case for 100 Foundation Model Companies00:09:22 Open Weights, Open Research, and Poolside's Global Team00:16:04 The Model Factory: Why Model Building Is 90% Engineering00:20:19 Agents, Automated Experiments, and Early Signs of RSI00:24:04 Streaming Data, Reproducibility, and Scientific Rigor00:30:35 Creating More Foundation Model Companies00:36:07 Laguna S: Persistence vs. Raw Intelligence00:43:01 Reinventing Pre-Training, RL, and Curriculum Design00:52:33 Low-Precision Training and Squeezing More From Smaller Models00:58:37 Model Harnesses, Coding Agents, and the Path to AGI01:09:26 Why MCP and Traditional Tool Calls Are “Stupid”01:13:04 Vision, Multimodality, and Why Language Still Matters01:18:15 Scaling Models and the Real Economics of Training01:20:40 Why Poolside Is Called Poolside and Raising $500M01:27:37 Open Models, AI Safety, and the Risk of an Oligopoly01:33:53 NVIDIA, TSMC, and the Reinforcement-Learning Bottleneck01:41:52 Smaller Models, Distillation, Engineering Productivity, and HiringTranscriptIntroduction: Eiso Kant, Poolside, and Open ModelsSwyx [00:00:00]: All right, we're here in the studio with Eiso Kant from Poolside, together with Vibhu. Welcome.Eiso Kant [00:00:08]: Thanks. Thanks for having me, guys. Good to be here.Swyx [00:00:10]: Yeah, fresh on the plane. You texted me, you were like, “Hey, I'm on my way to SF.” I was like, “You're on a plane right now, right?” Like, hey.Eiso Kant [00:00:16]: I know. After I texted you, I realized that probably coming in with major jet lag was gonna offer some fun experiences today, but let's do it.Swyx [00:00:23]: I mean, I think the thing I would tell guests is that they don't have to prepare that much because if you're truly working on this every single day, then even, like, what you hazily remember is going to be new for a lot of the audience that don't live in your world every day, right? so 10 years ago, you did a talk at Google Slush, talking about the democratization of AI. and, now here you are, like, open sourcing an incredible new model that we're gonna talk about. But I guess, like, what got you into democratization of AI? Like, it's not obvious from your LinkedIn or something.From Karpathy's RNN Post to SourcedEiso Kant [00:00:57]: No, it's not at all. I don't think it's obvious how I got in this space. I owe getting into this space to Andrej Karpathy.Eiso Kant [00:01:05]: In 2015, he wrote an article called “The Unreasonable Effectiveness of Recurrent Neural Nets.”Swyx [00:01:10]: Neural Nets, yep.Eiso Kant [00:01:11]: And that article, I read it, and I pivoted my startup at the time overnight to working on RNNs, and later LSTMs and Transformer models to be able to write code. If you go to this article and you scroll down, you can start seeing, like, this was the precursor to what ended up becoming language models. So, at least when he was character-level language models that were starting to predict letters, he has an example out here. There's a little Paul Graham generator, and you can read it, and the text makes sense, but it doesn't. and there's a little-- There's an example of code a little bit further down. Yeah, so Shakespeare.Swyx [00:01:47]: Shakespeare.Swyx [00:01:49]: CoolEiso Kant [00:01:49]: And for some reason, I read this, and I went down the rabbit hole of learning everything I could about RNNs and LSTMs, right? This is Transformer paper. And I had built a completely unreasonable belief, that neural nets should be able to generalize to anything and everything, and that language should be able to generalize, to a lot of things that are intelligent and the ability to write code. And so I started building Sourced, which was a fully open source company trying to build, what we used to call machine learning on code, language models on code. And we spent about four or five years on this, till the end of 2019. And that sounds really cool today, but back then, no one cared.Eiso Kant [00:02:29]: Right? Like, no one cared. We were in the dark. Like, we did things along the way. We tried applying convolutional neural nets to, like, the structure of code. We were. when attention came out, we were applying it to LSTMs, and then the Transformer paper came out. And it - it wasn't obvious, and what we missed throughout that entire journey, that we were on the right track, but we should have just kept scaling up. And today, to all of us, the scaling laws and scaling up seems like the most obvious thing. But having spent four or five years of my life on working on language models on code, it wasn't obvious. So I have a lot of respect to folks at Google and OpenAI and others who took that confidence and kept going. we failed ultimately at the time, and it was, like, biggest failure of my career, right? You blew $12 million of investors' money, which was a lot back then.Swyx [00:03:18]: Yep.Eiso Kant [00:03:19]: You spent, still a lot, but, And you spent years with, like, a group of 40 people just obsessing over this problem. And life took a different turn, And it was, and family became a focus, and I kept my heads down and really, didn't really look at language models for the following two years. big mistake considering Following years are gonna be really interesting. And then ChatGPT came out And it was like a vindication. It's like people started texting me. I found, like, my old, work decks and these old talks. And throughout that whole journey, we,ChatGPT, Vindication, and Returning to Open SourceEiso Kant [00:03:56]: We really had a strong point of view at the time that, like, as you're building more capable intelligence, it should be open and open source.Eiso Kant [00:04:04]: When we started Poolside, that wasn't the case at all, and I wanna be very open about it. When we started Poolside, we were like, there was a premise of two things. One is this technology is not gonna stop compounding in capabilities. I think to most people obvious today, but three-plus years ago when we started, most people were still arguing if these were stochastic parrots or not.Eiso Kant [00:04:23]: And the second was that reinforcement learning was gonna be the biggest driver for LLM capabilities. Today, very obvious. Three years ago, was not an opinion held or direction held at either OpenAI or Google or Anthropic or others. And so people looked down on us a little bit. They were like, “ is this really gonna work?” And so we just started working the problem, and we never really thought about open source again. We just kept our heads down and we built our, like, knowledge, understanding from scratch, right? We didn't roll out of an existing lab. So we picked up the papers and started writing code and figuring things out.Eiso Kant [00:04:59]: And it wasn't until the beginning of this year that me and my founder, Jason, picked up the open source conversation again.Eiso Kant [00:05:07]: And if you go back to some of the early things on our website, it was very straightforward. It was we wanna get to AGI, we wanna support a world of abundance, and we wanna be the first company that gets there.Eiso Kant [00:05:20]: But we started talking at the beginning of this year because it became obvious that the world was going in a direction that was starting to like, pick at us a little bit. Like, it didn't, this didn't happen overnight. It was, like, a little bit we were seeing this and we're like, “Okay, The world's going down a path.” And Throughout this journey, there was something that I used as a, as an analogy or thing. So I said well, if I go back to back in those days, 2015 or 2016, we're working on this, and I picked up a fi book off the shelf, and I was reading the book about 2035. AGI is achieved, and the story would be over the following, decades. And it would have that first chapter where everyone's trying to figure things out. You'd get the chapter of ChatGPT coming out And then you would get to the chapter where the world was at a fork in the road, and the one that it picked was one where three or four or a handful of companies were going to create all of intelligence moving forward.Eiso Kant [00:06:21]: And when I thought about that story, it felt like a dystopian fi book, not a utopian fi book. And the reality is, I'm a utopian fi guy. Like, and so We took a step back and said, “Hey, can we play a role here?” Now it was easy for us to do so because we were not at the frontier.Eiso Kant [00:06:41]: If we were at the frontier, I don't think we could have changed our mind. and I don't mean this like it's when the moment there's too much capital involved, too much expectations, you've built up things, right? We're a small team, just improving and improving. And so we knew that we could make that decision now, but it would be a lot harder to make as we got closer and closer to the frontier and caught up to others. And did a lot of soul-searching and a lot of conversations, and said, “No, this makes sense,” Even if there's big unanswered questions, like how the hell do you build a business model with foundation models about open source? Big open-ended question that we do not fully have the answer to yet, right? At what point do you no longer wanna release open source models because misuse of models has, real potential risks associated with it? how is the government gonna respond to open source? but I think it all just came down to one thing, and I'll stop the monologue, is the fact that I rather live in a world that has 100 foundation model companies than a world that has five, even if I was one of the five. And the smallest and most meaningful contribution we can make for 100 to exist is to open up our research and open up, like, our weights right now and figure out along the way how we can, like, do more.Neo-Labs, Model Choice, and the Token EconomySwyx [00:08:01]: Yeah. I think if anything, over the past three years, that has become a bit more true. you are one of a cohort of Neo labsEiso Kant [00:08:10]: YeahSwyx [00:08:10]: That people are now calling that. And, we're, we're doing this on the day that Thinky launched their, new model and you are outperforming them on their, on some benchmarks that they released, right? Like, they just don't have it yet. so it goes to show that I think, like, this is one of those things where, like, there is room for multiple players, and you are seeing a little bit more of the future. Maybe more like 20, not 100, but, like, you are one of the 20.Eiso Kant [00:08:36]: I really hope so, right? I think we I'm, I'm excited about their release, and I'm excited about everyone releasing because, like, ultimately, like, choice competition is both gonna drive progress in the right direction. But the fact that like, we create models and while we all, drink out of the same well of data effectively, we do introduce very different behaviors and biases in our models. Some are intended biases, some are completely unintended biases.Swyx [00:09:03]: Yeah.Eiso Kant [00:09:03]: And if we shape up in an ecosystem in the world where open models are gonna be a part of the token economy, like, I don't think there's any question about it anymore Then we want to be able to live in a world where companies, countries, people can choose and say, “Hey, I am most aligned and I trust most this provider for these things.”Swyx [00:09:25]: Yeah.Vibhu [00:09:26]: I think more than just one of the 20 Neo labs, up until recently, most of open source innovation was coming from the Chinese labs, right? So there's the DeepSeek of the West. Is it today? Okay, maybe it's thinking machines reflection, but there aren't many, right? So, one of the things you guys started in France, Europe, but very much now you're taking that American standpoint and more than just that, the point is the Chinese models that we see, they're not super open research. the work you put out is, I think, some of the best. So every few months you get not only frontier models, but also here's a breakdown blog, paper, technical report of here's everything for state of the art to build, frontier intelligence and you're filling that gap too, right? So not just only open weight, not just Western, but also pretty open research.Open Weights vs. Open ResearchEiso Kant [00:10:20]: No, I appreciate it. Look, I think it's, I think it's the most meaningful contribution, right? Weights are a binary. Let's call them what they are. Yes, we can modify them, we can change them, but, like, giving someone the weights does not allow them ultimately to recreate what you're doing, right? And so now there's challenges around releasing data sets, challenges around like releasing certain things, but being able to share your research, like, right, how do we do it? What are the lessons we learned that we spent, tens of thousands of experiments of compute on? I think very much so. One correction though, Vibhu, and I say this because it's been haunting us for quite a few years. We from day zero were an American company.Swyx [00:10:55]: Yeah. They movedPoolside's Global Team and American Company StorySwyx [00:10:56]: To France.Eiso Kant [00:10:56]: So the story once and for all is very. We start as an American company. We have always been an American company, and early on we made a very conscious decision. We said, “We're not gonna hire any researchers in the Bay Area. We're gonna look for talent everywhere else in the world.” and that is everything from Middle Americas, Seattle to, Serbia, and to Taiwan and Singapore and other places. And it was because we took a view that this was gonna become a talent war for this, and I think it has over the years now. Three years ago, that wasn't fully obvious yet. I think today it very much is. And we also realized that, like, some of the world's most capable people with, like, the most interesting, innovative ideas were not just gonna be here. And so it led us to create like a fully remote company. and we ended up opening an office in Paris and London and different places and we have a lot of the team in the US and a lot of team outside. But we always took this view of like, we're an American company, but if we want the best of the best to work with us, we need to take a global view. Now we do also have people here in Silicon Valley, like the company's grown and others, but I think one of the things that, it slowed us down at the beginning, but it has sped us up now, and it's why you're seeing like the progress, I think, on our models and the cadence at which we release, is because we didn't roll out of an existing lab. Right? we didn't, we didn't have a lot of the information that's freely flowing around here at the time. We just took this point of view as like, “Okay, well, let's just work the problem. Let's just go and, like, read the few papers that are out there, and let's just figure this stuff out.” And we made some hilarious mistakes in model training because of that over the yearsEiso Kant [00:12:35]: Like especially in the first 12 months. there's a few that I think still haunt me and scare me. We can talk about them later. but it created a, like, a resiliency and persistency in the team, right? with extremely few people have left us over the years, that, like, told us, “Okay, we can do this.” When we first wrote our first training code base completely from scratch, it wasn't a fork of any open source. It was just like, “Okay, let's build it from scratch.” I remember we had this one moment where we spent three weeks working out an optimizer bug. Like, it was like training just couldn't get stable. We, like, obsessed over it, and we thought, like, maybe we were wrong. Maybe we should have just forked this repo, or we should have. But then when we solved it, I still remember at the time we were like five people in the company. when we solved it, we were like, “Oh, we can do things,” like if we're just willing to work hard. and I think that culture with a very strong engineering bias has helped us, like, get to where we were. And so there's this notion of open source and talent and these things. I think we, We just took different decisions from a different starting point. and I think we are lucky. I do want to definitely call it lucky. And there was a lot of hard work at the team that now, like, that's starting to show up in results.Swyx [00:13:52]: Just ‘cause we probably won't revisit this again, but, and this is a fun recruiting challenge if someone knows the answer. What was the bug? And then we won't tell the solution, but we'An Optimizer Bug and the Value of Building From ScratchEiso Kant [00:14:01]: So the - This - You're gonna test my memory here,Swyx [00:14:04]: Oh, okayEiso Kant [00:14:04]: So but I thinkSwyx [00:14:05]: DirectlyEiso Kant [00:14:05]: I think I can recall. So if you, so if you look at, So if you take like Adam as an optimizer, you have epsilonSwyx [00:14:12]: YeahEiso Kant [00:14:13]: Which is, right, like in the denominatorSwyx [00:14:14]: Momentum and weights. YeahEiso Kant [00:14:15]: Is exactly, in the denominator. And at the time, if I recall, you looked at like the early Llama papers and things like that. People were juicing epsilon, like, quite a bit. Like, they were, like, adding, I don't know if it was E minus four or whatever, like a high value for epsilon.Eiso Kant [00:14:31]: And if you think about this during training, it's like a bit weird and counterintuitive that we're adding noise to our optimizer by just adding effectively, like, a random number in the denominator, right? Like behind the decimal point. And I don't recall the exact bug, but it had - What I remember is once we solved it, we no longer had to juice epsilon as much as, like, was happening in the Llama paper and other places. and it was like one of those fundamental moments where we had trusted this paper that was out there, and we're like, “Oh, no, it has to be this way. It has to have this high value of epsilon.” But it made no sense to us intuitively. Like, why do you have to have this so high? Like, if you're just trying to avoid division by zero, why can't the value be extremely small? and that was like one of those moments where you realize like, okay, finding things out from scratch yourself builds a better intuition. Because the one thing you learn very quickly with model building is that your intuitions that you start with are gonna get beaten up so hard.Eiso Kant [00:15:33]: Right? Like - It's such an experimental science, that the things that seem obvious, you very quickly get to learn, like, you were wrong, and hopefully you figure out why, and sometimes you don't even.Swyx [00:15:45]: Yeah. yeah, so, one of the reasons that you, when you released your new models, Vibhu got really excited. I mean, everyone got really excited. But Vibhu led our paper club on it, and you guys sawEiso Kant [00:15:58]: YeahSwyx [00:15:58]: Obviously. maybe talk through some lessons learned in that, whatever you can disclose. we can focus on the model factory stuff, whatever you think is a good starting point.Model Building as EngineeringEiso Kant [00:16:08]: So I would say that our view from very early on in the company was that model building is ultimately 90% engineering.Eiso Kant [00:16:18]: And I think we all know it in the industry because if you look at where's every researcher spending their time, they're spending their time writing code, right? Looking at data and writing code. And so we said, okay, The state at the moment, like three years ago, was bash scripts and Slurm and spaghetti code bases for training and, like, data pipelines that were patched together. And we looked at this and said, “Well, ultimately, model building is a process.” You're going from raw data, right? Like training raw material, the web, et cetera. you're doing a whole bunch of filtering, cleaning up, transformations, analyzing. These days, that's, far more complex than it was three years ago. then you're training a model, which is effectively a large distributed systems problem, right? Across hardware that has still-- It's become a lot more reliable. It was extremely flaky back then. and now with every new generation, we get our new sets of challenges. And then you go into the next stages, right? There was no training back then, but, like, you got, your post-training and then your reinforcement learning. And so we looked at this and we said, “Well, this looks like an industrialized process. This looks like an end process, that every single part of it has its machinery,” right? If it's your big data pipelines, if it's your crawling ingestion of the web, if it's your, large-scale distributed training, and then you've got your reliability. And we said, “Well, why don't we take some of the world's smartest distributed systems engineers that we knew and make them part of the process of research from day zero?” Not retrofitting it later on, but, like, really from the beginning. And that became our model factory. And so our model factory started with a handful of components. Today, it's thousands of components, and I try to equate it to, if you think about, like, someone who was at the very early days of Foxconn, if they had been there for the following, decade, they would be able to rebuild Foxconn because they saw every decision that led to building that system and all the complexity. If you and I walk into Foxconn today, no chance.The Model Factory and Experiment VelocityEiso Kant [00:18:18]: Right? Because we don't have the lineage and history of decisions that led to that. And so we built early on from the beginning- with a team that really understood that, well, the metric that we are optimizing for is the speed of an idea from a researcher to an experimental result that we can trust to then being part of the next model training.Eiso Kant [00:18:42]: And in the. And because it's such an experimental science, ultimately, in the beginning when it wasn't that complex, you could patch your way around it, right? But now, at any foundation model company, you are running. I mean, we're a small team, right? We're less than 70 researchers, another 35 engineers. and we are running, I haven't checked the latest count, but far more than 10,000, maybe 10 to 20,000 experiments a month that we cut. And so if you look at that scale of every model run that is, like it's ultimately it's, it's you need to be able to trust it as an infra problem. And so what we have now done over the years is gotten really good at that, and just by working it and improving it and obsessing over those end decisions. So now what that means is that you looked up Laguna XS 2 that we launched. It was five weeks from the beginning of training to launch. The model that we're gonna talk about today was eight weeks from start of training, to launch. We started the next model literally yesterday because we now finished the post-training required for the model we're launching, next week or by the time this comes out today. and we move that compute to the much larger Laguna M model that we're now training. And so the model should be an artifact of someone's process. It shouldn't be really a thing in itself. Like, and we treat this like the way you would look at like a SpaceX factory where, yes, the first rocket, really hard to build, but the much harder challenge was building the factory. And now they're rolling off, and no one is really thinking about the next launch anymore. So it's just another launch, it's another launch, another rocket comes off. And that's what we're trying to do with model building.Eiso Kant [00:20:22]: And what has been, which was not planned from day zero, it was in the back of our mind like this will happen one day, is that when you build a really good end model factory with really good APIs and really good engineering systems, Well, what is it perfect for? It's perfect for agents.Agents Inside the Model FactoryEiso Kant [00:20:40]: Because agents are now starting to take over more and more work in our model factory.Vibhu [00:20:43]: Yeah.Eiso Kant [00:20:44]: So I look at the screens when I walk, like when we're, we come together, in our monthly, we do monthly onsites, and I walk behind people's screens and I stop by and I talk to our researchers. And the default is all of these different agents running on their screen that are writing the code. They're launching the jobs. They're evaluating the results that are coming back from the model runs. They are, making the changes. And we're still in the driver's seat. We're still coming up with the ideas. We're still helping with the debugging. But more and more, and this is right now very profound on the data side of our pipelines in both pre and post and the synthetic data pipelines, it's starting to become more on the architecture side as well. You're starting to see these twinklings of what RSI is gonna look like.Eiso Kant [00:21:27]: And that's. So when we talk about, like to your question about our models, every talk about the model factory, And my coolest example of these things is always that when we kick off a new run, doesn't matter if it's a training like big run or if it's now a post, like one of 10 post-training versions we do for like release or many experiments, is that at any given moment, the changes that somebody made that they had experimental results from the day before make it into that run.Eiso Kant [00:21:57]: So there's not like a cutoff 90 days before. Like no, it's like literally from that moment because we can now trust the machine enough. And then you also have to invest in the reliability. So one of my favorite metrics about like Laguna S is that there was no call events, Right? Like completely zero. And we haven't had a meaningful call event, like something to wake up for, as far as I recall this entire year. now there is one asterisk to that. In usually the first six hours of launching a new model run, something breaks because you set a config wrong, you made a small mistake, et cetera. So that's usually there's a little bit of intervention, but that's always within like call periods, right? Not on call. And I think that's starting to now compound. So the model we're releasing now, I love it. It's amazing, but we're already onto the next one. and I think that's the way it should be.Laguna, Five-Week Builds, and Zero On-Call EventsVibhu [00:22:50]: Hey, I also just wanna point out, so for context, this was like a month ago. we found it in the tech report, so we just came in with, “Okay, new model's dropped. Haven't heard about it.” We wereEiso Kant [00:23:02]: Yeah, we're very used to doing this every few months.Vibhu [00:23:03]: We're, we're very much like, “ okay, look, it's like, on par with Kimi, DeepSeek, whatnot, the small ones, Gemma level. Oh, it's a very cool paper on what goes into building.” And then we hit this page, right? Like literally page two of tech report is, “This process allowed us to build the small model from scratch to delivery within five weeks applying the lessons”. And then I'm like, oh, this paper is not about here's a tech report of benchmarks and here's how many tokens it was trained on. Like for people that wanna dive more from what we're not gonna discuss on the podcast, it's all laid out here, right? FromEiso Kant [00:23:38]: YeahVibhu [00:23:39]: Custom software that agents can use to interface with training code, training data.Eiso Kant [00:23:45]: Yeah. Well, link the paper correctly, so yeah.Vibhu [00:23:47]: Yeah. All that stuff. read the paper here, but,Technical Report Principles and Streaming Training DataEiso Kant [00:23:50]: But I would like to. I love principles, and I think that is a good starting off point for maybe telling some stories. Maybe we can go one by one past the principles. I'll just call out that Dagster just got bought by a Prefect.Vibhu [00:24:01]: Yeah.Eiso Kant [00:24:01]: Isn't it fun? But yes, I'm very familiar with Dagster. just anything where like they trigger some story.Vibhu [00:24:07]: So, well, I would say, well, experiments code's obvious, but I think one of my favorite things is, I don't know where it is in here, but early on, and I still think this is the case a lot of foundation model companies, people prepare their training data sets, they get packaged up, then they get copied over to a training cluster distributed across all of the nodes, and then training starts.Vibhu [00:24:30]: And we looked at this like three years ago and we were like That makes no senseEiso Kant [00:24:36]: You lose so much time because the moment you have to rematerialize the data set, you have to make a change, you have to fix something, et cetera, you've got all this time of like repackaging it, right? Toca- tokenizing it, repacking it, moving it over to a cluster, then distributing it across the nodes. The bigger your clusters are, you start using fancy like torrent-like algorithms to like distribute your data. So why aren't we streaming data into training? Right? Something that's very common and like just basicVibhu [00:25:00]: Like just in timeEiso Kant [00:25:01]: Just in time, like good computer science like principle. And that was one of the first things that I think unlocked - the model factory. Because the moment you start thinking about, well, a training job, it doesn't matter if it's a big hero run or a small like, post-training experiment, consumes a certain number of tokens per second, right? And it's not a lot, right? From a like a data, moving data perspective. So we said, well, we have our training cluster, and then we've got like our AWS kinda setup where we can build these amazing big data pipelines. We can set things up. We use Spark underneath the hood, like all these things.Vibhu [00:25:36]: But when you say AWS, it's not actual AWS, it's your internal AWS.Eiso Kant [00:25:39]: It's our internal-- No, it's our internal like just running like our infrastructureVibhu [00:25:42]: Site web servicesEiso Kant [00:25:43]: Exactly. Our stuff running on like an AWS account or on like any hardware, right?Vibhu [00:25:47]: Yeah.Eiso Kant [00:25:48]: And so once we made that shift into I can stream data into training, all of a sudden you realize a lot of things unlock. Because now you don't have to wait for the whole data set to materialize.Immutable Data, Experiments as Code, and Scientific RigorEiso Kant [00:26:00]: You now all of a sudden when you're running data experiments about mixing data, it's a config. Because you've got these data sources that are coming in, and you just - we have this service called Blender that's in the report, where we then say, “Okay, for this run, I want 20% of this source, 10% of this source. I want this much, so many epochs of repetition. I want this to be, shuffled in a certain way,” and your training job can start while the rest of the data is even still materializing. also what it does is because all of this underneath-- So for us, we treated the data layer underneath as like an immutable data layer, and that was really important. Like experiments as code, immutable data layer means that you can always go back and understand literally down to the single token at which cursor it went in on which version of the code.Vibhu [00:26:47]: Yeah.Eiso Kant [00:26:48]: And it took us a I have to admit, like the first year of Poolside, we understood that engineering had to get great, But we didn't understand yet, that this is ultimately in support of like a good rigorous scientific progress. We were quite a - We were a very small number of people, so a lot of it was YOLO ideas and YOLO runs.Vibhu [00:27:08]: Yeah.Eiso Kant [00:27:09]: And we built great infra for the YOLO runs. But once we realized that we treated data as immutable and code as always versioned, and you could always track and trace every experiment end to end perfectly, you could repeat everything perfectly, right? You have perfect reproducibility. I can still reproduce runs from two years ago if I wanted to, right? It enables the scientific progress, like the scientific process, and I think that took us probably about a year and a half into the company to figure out. We also had some great hires, like our head of applied research, Nikolai, who joined us from Yandex, who'd been working on language models since like the early 2020s, I think brought that into the company of like, “Hey, we wanna have even more rigor.” And then once we kinda had the combination of like increasingly more capable platform that allowed people to do more, but had this immutability, we were able to start “Okay, every experiment is truly an ablation. We truly need to understand it.” And I think we became much more scientifically rigorous in the last couple of years, and the infra underneath enabled it. and then there's just fun stuff like, andVibhu [00:28:16]: Yeah, a lot of it's fun, like even just the, one, you share all the ablations, two, picking the data sets, right? There's like a random small paragraph in here where it's just like, “Oh yeah, training data, we have some, we have an auto mixer.” it trains eight small models, scales them up, picks the training data set. We don't even need to look at it. I'm like, “Wow, a lot of engineering rigor there.” And there's just, there's just a lot in here.Publishing Research and Giving BackEiso Kant [00:28:40]: Yeah, and it'- and look, and we wanna put out more. Like we, We treat writing papers as something that we haven't earned the right for yet for a long time. So you earn the right to spend time, publishing research once you're at the frontier, because until then, you're catching up, and every minute and hour in this industry matters. Like I obsess over it, not just the wall clock time from idea to result, but just general like time every day that we, waste is one that doesn't allow us to catch up. But in this case, we said, “Okay, we're gonna give ourselves.” I think we gave the team like three or four days while still doing their work, like give everything in there. And to your point earlier, if your stuff, it's easy to like put it out. And so there's so many more things that we wanna talk about over time, and we will definitely start doing. And as we earn more of the right, but also now have like added to our mission that we want more foundation model companies to exist, you'll see us like be way more proactive, and just trying to keep dropping some of those like things that we've learned along the way that can help others like speed up.Vibhu [00:29:40]: Which is the other cool side of this, right? It's, it's not like, back to your point, it's not just here's the benchmarks of our training. If you want to replicate, here's experiments of optimizers, data sets, post-training. you lay out a lot of it here alongside here's your system for how to do it? So it's, it's really like promotingEiso Kant [00:29:59]: No, thank youVibhu [00:29:59]: Other people can do the same.Eiso Kant [00:30:00]: And by the way, I also wanna make clear, right, we have been incredible-- Like we've taken a lot of advantage of the fact of all the open research that others have published, Right? And you mentioned, the Chinese labs, and we I think it's important that there's, from every country and every culture and background, including like Western companies like us, there's different models that come out that people can choose to trust. But I think we do have to give credit where credit's due, right? The incredible Chinese lab have done an amazing job at sharing their research, and we have definitely like been on the receiving end of taking advantage of that. So when you're on the receiving end of something coming to you, I think it's, you also have an obligation to give back.Swyx [00:30:39]: Do you have a favorite or underrated Chinese lab that you wanna shout out? Everyone shout outs DeepSeek.Chinese Labs, Zhipu, and PersistenceEiso Kant [00:30:44]: That's a good question.Swyx [00:30:45]: Moaan obviously for Therapsi. Yeah.Eiso Kant [00:30:48]: Yeah, look, I think, I think obviously everyone's been talking about Zhipu lately, with 5.2. I think what most people don't realize is when they started.Swyx [00:30:59]: Yeah.Eiso Kant [00:30:59]: Right? They started years before ChatGPT.Swyx [00:31:02]: They just rebranded. YeahEiso Kant [00:31:03]: And so, I've like, I remember how hard it was to work on these things Before the rest of the world got excited about it. And so I have an immense amount of respect for people, who were working on improving models when it wasn't the sexy thing to do, when believing in LLMs, was gonna get you ridiculed. I remember like back in 2016 when we were doing what we'd call, machine learning on code with some of these models. we would-- people would just laugh at us, like they'd be like, “This makes no sense. Like why are you wasting all these, like, millions of dollars on trying to figure this out?” And so I would say they're probably the one that, I think deserves a shout-out, not just because their latest model is very good, but because they fought to get here. And I think, I think every foundation model company it takes time to get here, right? It took us three years to get to the model that we're, that we're now gonna be releasing. and now the time in between the models is coming, is counted in weeks. It's no longer counted in months or years. But this stuff's hard. and if we can make it a little bit easier for the next person, like we should all do so. Because if we don't do so, we're, we've got a small window before models are really impacting recursive self-improvement to a level where catching up otherwise might become unfeasible. And we should try to, in that window, encourage as many labs or however we wanna call them, like to start. And so one of my currentEiso Kant [00:32:36]: Mission, but qualm is like I wanna encourage whoever is a researcher right now who thinks they can tackle this to go and leave and become my competitor.Eiso Kant [00:32:45]: Like start another foundation model company because I think we need it. I think otherwise we're not gonna be in the world where, I don't want to just be the fifth or the sixth company that wins. I wanna look at a world where there's lots of choice.Starting a Foundation Model CompanyVibhu [00:32:57]: What else do people not see in starting a foundation model? it's, there's a lot of compute, there's a lot of capital required, a lot of compute. You lay out model factory and how to do the training, but there's a lot there, right? That's,Eiso Kant [00:33:10]: Well, look, it's, I in turn-- this is an oversimplification, and I always asterisk it with that because it can land a little bit the wrong way in people's minds. But I think you can sum down, And I saw it, 95% of model building to just doing, you're just doing two things. You're improving data or you're improving compute efficiency. And I know that feels like an oversimplification for the incredible, like, Gifted and skilled work people do. But if you really look at it, like what are we doing? We are looking at data, we're generating new data, we're improving data. and the only way to do that is to look at the data, right? That's a big part of foundation model building. And on the other hand, we come up with these incredible breakthroughs in inference, in architecture, and new attention mechanisms. But what are they really doing? They're bringing compute efficiency. Now, we have definitely had some breakthroughs over the years that allow for more model capabilities. But at the limit, if you could train a large enough model, right, like, and you had infinite compute, we probably-- if you had infinite compute, you'd be at AGI probably already tomorrow.Eiso Kant [00:34:12]: Right? Like it's not. And so, and let me say that infinite compute with infinite ability of much faster networking because networking ends up being more of the bottleneck than compute. But, so I do think that's, those are the main things. And to just realize that this is engineering. I think it's become more obvious, but I think for quite a few years, people have held foundation model companies and researchers and others on this pedestal of like you're doing incredible magic or rocket science, or only like, Nobel laureate physicists can do this. And don't get me wrong, there are some really hard problems that need to be solved, but a lot of the work that all of us are doing on a day Is not sitting down trying to solve a math theorem. A lot of the work that we're doing is just really doing the basics right, writing good code, looking at data, improving it, running experiments, looking at plots, trying to see like, hey, trying to shape our intuitions. And a lot more people could be highly capable researchers. and I think that's, it feels far for people to do so. But I've seen in our own company, we've seen engineers become researchers because the model factory allowed them to be, have a much lower hurdle of running experiments and trying things. And one of the guys on our team who started as an engineer building our agents is a legit reinforcement learning researcher now, making real progress. and that happened in the span of like six months. that would've not been what I think most people assumed was possible, a couple of years ago.Swyx [00:35:46]: Yeah. I think one of the interesting moments is when you can self-host, like, if in a programming language, like if you can compile the language in the language, the equivalent is can you use your own tools, right? You have the pool CLI, you have your own models. presumably you're not only using your own models. There's no way. But like, what's that percentage over time?Laguna S, Persistence, and Behavioral GainsEiso Kant [00:36:10]: This is the first model that we're releasing that is starting to meaningfully contribute to our own work. It's not a it's not state-art model yet. Fable and other, they're, they're very capable models, but Laguna S Is really interesting. I'm gonna pull up the quote. Peng Ming, one of our heads of applied research, said something, last week as the model came out about 10 days ago, much better than we had hoped for or expected. And he said, I have the feeling that a lot of the gains in Laguna S come not from more intelligence, but more from different behavior, more verification, less taking things for granted, not declaring victory early, and being way more persistent. And to be honest, those are more predictive than raw intelligence for success in human also to some degree. And this was, he wrote me this on 5th of July on a Sunday, and it's been burned in my brain ever since because the Laguna S model, as you'll see it and why it does so well on benchmarks and why it does so well in using it on a day basis, is that it's just incredibly persistent. It reasons a lot. I do call that out. We have work to do on making it more efficient. We have to work to do on offering different reasoning modes. But this is the model that has been able to do things that I never thought it could do. A hundred eighteen billion 8B active model, which is not that large. It fits on a DGX Spark and still runs at, thirty, forty tokens a second on a Spark, is able to solve Erdős 397 independently. It's able to do complex programming tasks. It's able to. I asked it this morning to make me a Fi scanner without using any external libraries on my Mac, and it's, like, figuring out, like, the core WLAN API by really persistently trying to understand it without access to the internet. And more, I love vibe checking. I've probably spent eight to ten hours a day with this model for the last ten days.Eiso Kant [00:38:05]: I'm not exaggerating. I was on my eleven-hour flight yesterday. I spent ten hours reading trajectories and traces and, like, of the model.Eiso Kant [00:38:12]: And what I take away from it is exactly what Peng Ming said. We are gonna be able to squeeze so much more out of smaller models than I think we had imagined in the industry because, yes, there's intelligence and larger models are more intelligent. Like, no doubt about it. We should continue to scale up. but the behaviors of being really persistent, of being able to backtrack when you're wrong, of, like, understanding how to interact with your environment show us that we can get a lot more out of it. And this, for me, has created a bit of a Question in my mind the last couple of days. If you think about where we're using models today, right? We are using models, say, for knowledge work. Represents twenty-five percent of the global economy, twenty-five trillion dollars of work.Eiso Kant [00:39:00]: As we scale up models and they become more intelligent, we are excited about using them more and more for pushing the frontier of science.Small Models, Knowledge Work, and CommoditizationEiso Kant [00:39:08]: And if you look at the frontier of science, like true breakthroughs in science, they have been linked, they are linked to more intelligence in many places. Einstein figuring out general relativity is able to bring ideas together that other people would have not brought together. And I think one of the many dimensions of intelligence is the ability to do that, and it's something we clearly see that as models get larger and more capable, they're able to pull more ideas and threads together that a smaller model wouldn't be able to.Eiso Kant [00:39:36]: And we're starting to see examples of that in medicine and, like, in bio and other things. But if you think about the majority of knowledge work that we do, and it includes building software. I'm a software developer at heart first and foremost probably, although I probably can't say it that much anymore as I don't write production code in years, is that what makes us good is our persistence. It's our ability to encounter a problem and backtrack and say, “I need to go figure out this bug. I need to go research this. I need to go look at the documentation. I need to, like, try different, five different ways to see, like, if I can solve it.” But it is not necessarily bringing three ideas together from radically different fields. And so if we are now seeing, and I think Laguna S is an example, that we are able to make a relatively small model much more capable than I had definitely predicted or any previous, like, benchmarks had shown for any model remotely this size or even larger, At least on coding tasks, that it's because of the behaviors. And so now the question I have, and I don't have an answer, it is I know at the limit, so infinite model size, right, extremely large model, and the cost of that model is gonna be very expensive to run. We know this, right? So larger model ROI.Eiso Kant [00:40:52]: So I know that at the very limit, I'm not gonna use the world's largest model one day, quadrillion parameter, whatever crazy, like, scale we scale up, to do a basic coding task. Already today, I'm starting to size down for certain tasks.Eiso Kant [00:41:07]: So it means that there is an optimal. It means there's some curve that goes as we go up to model size for knowledge work, at some point we're at the peak, and after that, the return on investment of using a bigger model, just doesn't make sense.Eiso Kant [00:41:22]: Now, I think the question is, before I would have thought that peak was extremely very far away.Eiso Kant [00:41:30]: This model for me is the first sign that Maybe that peak is At a trillion, five trillion, ten trillion. Maybe we can just squeeze way more out of these models. I'm no longer thinking that we need two or three orders of magnitude on the largest models to be able to, solve knowledge work, the accounting, the legal, the code that we write. And so if that holds true, It is an argument for the commoditization of models. It's an argument that open source can win and, like, succeed in this world. And now it's of course a self-serving argument and it's a hopeful argument, but theoretically at the limit it works. We just have to go discover in the next couple of years of how much more we can squeeze out. Now, I do want to put a big asterisk. This does not mean I'm against scaling models. I think we ultimately only succeed if we scale our models as large as our competition. I do not like. I think we should not put our head in the sand and say we're gonna be king of open source small models. I think that's, It's a out. It's trying to be king of your own kingdom, but not realizing what the rest of the world's doing. All of us rather use a smarter, faster, more model. It's a sign of hope. And so I don't wanna overly state this is a good model. We have a long way to go to get to the state-art. But what hopefully people take away when they use this model is that the behaviors inside of it are what push it to be far more capable, less than necessarily the number of parameters.Pre-Training, Mid-Training, and RL Moving EarlierVibhu [00:43:03]: Is that mostly post-training? LikeEiso Kant [00:43:05]: YesVibhu [00:43:05]: Right.Eiso Kant [00:43:06]: It's entirely post-training.Vibhu [00:43:08]: Are we done improving anything on training? Is, like, training done?Eiso Kant [00:43:12]: No.Vibhu [00:43:12]: Okay.Eiso Kant [00:43:13]: SoVibhu [00:43:13]: I just wanted to cover training, and then we go post-trainingEiso Kant [00:43:15]: Training is not done. I mean, look, there's a part of training of just dealing with skill, right? Every new order of magnitude of model skill, you are going to get new things you gotta solve for. That'- but those are ultimately, engineering challenges.Eiso Kant [00:43:31]: I have a, I would say, a not commonly held opinion that reinforcement learning Will move earlier and earlier into training.Vibhu [00:43:42]: Yeah, training.Eiso Kant [00:43:44]: Not even training. Like training today, right, is, like if you look at - So we've been working on this for years already. and I think the best-- I think the first time we saw it out in public was the DeepSeek Zero paper. this is a year and a half ago, I think, if I recall correctly. where, you can Very early on in a model as it starts capable of being able to use language, et cetera, induce reasoning. and so the question that I have is like, we have this- we have the dataset that's the web. and the web, I think we could arguably say probably has The totality of humanity's knowledge somewhere encoded in different places. It's a huge variance degree of quality, from garbage data, and like once you look at training data, you really get humbled of like what the web is, to like, the most greatest scientific papers and best blog posts and like, best transcripts and whatnot.Eiso Kant [00:44:39]: And so now What we are trying to figure out, and have been doing a lot of work on, and it's a place where maybe not as open as we're on other things, but we will become more over time. we've been spending a couple of years really doing research on how can we turn the web into not just next token prediction, but into a way to teach the model to think earlier in its training. and I think there's a huge amount of gold to be found there. I think we are right now in, we've got some drugs in the industry. One of the drugs is distillation. Another drug is, more environments. Like, and they're great, and they make us feel good, and they make the models better, and like we're all addicted to them, and we'll use them, right? in various different ways. and but ultimately, I think we are still barely squeezing out of the web what we should be getting out of the web.Eiso Kant [00:45:33]: I think just next token prediction during training is not enough.Eiso Kant [00:45:36]: AndVibhu [00:45:38]: YeahEiso Kant [00:45:38]: I think we'll see some very interesting things still happen. and that RL in post-training to induce behaviors, to improve things, like I think - the whole world knows how to do this now. I think we're, we're scaling it up. Everyone is. But I wonder if we need to go as far as we're going today with environments. I'm not sure yetVibhu [00:46:01]: You mean we're going too far?Eiso Kant [00:46:02]: I'm, I'm not sure if the path to AGI is justVibhu [00:46:06]: Is more environmentEiso Kant [00:46:07]: More environments.Vibhu [00:46:08]: It seems like a never-ending, “Okay, I want instruction manual for this table, right? Am I gonna environment out building furniture? Or are we just gonna tail end like we need some general solution?”Eiso Kant [00:46:19]: I think there is, I think there's an ability to generalize more from the web. but I also am very encouraged, like when I look at Laguna S and, which is post-training is, well, is the big impact there. and I see like, oh, wait a second, just by making some of these behaviors much better, we're able to get so much more out of it. It just changes a little bit the way you think about intelligence.Vibhu [00:46:40]: Yeah. The analogy people draw often is the RL phase is where you don't learn as much new knowledge. You shiftEiso Kant [00:46:46]: Yeah.Vibhu [00:46:46]: Yeah. So, you shift distribution, and you can have it reason towards what you want. on your point about training, a lot of training is still just continue training in a domain, say medicine, then you do RL. So still justEiso Kant [00:47:00]: It's just better data, right? Like, I mean, training, ooh, I like how we invented this word. Like it's effectively just like,Vibhu [00:47:06]: Second phaseEiso Kant [00:47:07]: It's the second phase of training With like a really dumb way to do a curriculum. But like ultimately, what you'd want is a curriculum from token zero to token 30 whatever or 40 trillion tokens that really truly is the optimal curriculum for the model to learn. But training is essentially a stage curriculum on the web because we do not have to compute, And, effectively to try to ablate the perfect curriculum, right? And so I'm pretty sure that you'll start to see people talking soon about some other term, and there's two or - ‘cause now we do this, right? We talk stage two and stage three and stage four training and like. But ultimately, all we're doing is we're trying to assign a curriculum to the web data that we have to allow the model to learn better. I think at some point, as things get compute, as models get cheaper to run, as the next generations of compute, this will become more of a continuous spectrum. I also think the reason, by the way, you have training and like stage two and stage three is organizational, Right? It'- this is, I think, a thing where-- that we really try to avoid with the model factory is like Training exists because there's a training team now, right? There's people, or like people in training decide to focus on like a training effort. but what you really want is engineering and scale of experiments that allows for a much more continuous spectrum that you don't, you have infinite stages. Now, we're not there. Compute's not there. Organization design is not there for it yet. but I think we'll get there. we'll look back on a couple of years and be like, “Oh my God, it was so cute that we did our training data like this in such a like naïve way. Like we barely ordered it. We didn't really do a good job at likeCurriculum, Auto Research, and New ObjectivesVibhu [00:48:48]: The building that curriculum will get you that in the industry.Eiso Kant [00:48:51]: And I'll confirm that, when I talk to some researchers that this is a lot of the focus now is like how does training change and what is the next objective other than, next token prediction. I assume you don't have the answers, but you have some ideas.Vibhu [00:49:02]: We have some ideas. We're not ready to talk about it yet.Eiso Kant [00:49:05]: Yeah.Vibhu [00:49:05]: We've been working on them for years, and I think that's the one thing that's also like you asked earlier about, like what's not obvious about building a foundation model company is that you are constantly balancing the table stakes work, the recipe worksEiso Kant [00:49:19]: Yeah.Vibhu [00:49:19]: Versus like your, my crazyEiso Kant [00:49:22]: Pure researchVibhu [00:49:22]: Breakthrough.Eiso Kant [00:49:22]: Yeah.Vibhu [00:49:22]: Pure research and finding that balance and adjusting the percentage to it based on where you are in the race is really important.Eiso Kant [00:49:31]: I mean, so like, this is a nice way. I was gonna bring up auto research at some pointVibhu [00:49:35]: YesEiso Kant [00:49:35]: As another Andrej invention, or coinage, which is like, I honestly, like how many objective functions can there be, right? Like just try 1,000 of them, set it running, whatever.Vibhu [00:49:47]: Man, it's alsoEiso Kant [00:49:48]: Like what you're looking for. You're looking for loss curves like that, likeVibhu [00:49:51]: It's also a thing people take bets on, right? When you say more Neo labs, you're doing a version of we'll do foundation models, scale them up, next token predictors. A lot of other Neo labs that we see want to take a completely different approach, right? At some level, you're right. It's all, compute efficiency, and that's the net objective. But some are okay, different architecture, like vastly different amounts of compute spend. So some are different. They're not justEiso Kant [00:50:19]: YeahVibhu [00:50:19]: They're like, 99% not balancing, here's the vanilla and scale up. They're 99% on, here's novel research that'll change everything.Eiso Kant [00:50:27]: And I think, Luke, I think you. It depends when you started as well, right?Pure Research vs. Table StakesVibhu [00:50:30]: Yeah.Eiso Kant [00:50:30]: When we started, like the novel thing we did was reinforcement learning on code. No long- that's no longer novel by far, but we were like, - that's where we obsessed over when no one believed in RL. So you have to when you start the company, you have to have your own idea. You have to have something that's different that allows you to speed up, right? For us, it was RL to LLMs that later became common, like, Knowledge. But in the beginning, it wasn'tVibhu [00:50:53]: It's cool. this was like your original 2023 blogEiso Kant [00:50:57]: YeahVibhu [00:50:57]: Of purpose.Eiso Kant [00:50:58]: Yeah.Vibhu [00:50:59]: And like you do lay it all out here.Eiso Kant [00:51:01]: We laidVibhu [00:51:01]: The blog is pretty underrated, right? The whole RL on code was very early on.Eiso Kant [00:51:06]: Very early. And even we had to argue with people, like we say here things like to push beyond current capability, to train your own foundation model. We had to argue with people that it mattered that you had your own like, base model. you can fine-tune your way to success, right? major capabilities emerge from training a base model made accurate and useful during fine-tuning.Vibhu [00:51:23]: Which like, for perspective at the time, we knew closed models, OpenAI, Anthropic were huge. The open models we had were like Mistral 7B, a 30B, a 70B.Eiso Kant [00:51:35]: When weVibhu [00:51:35]: YeahEiso Kant [00:51:36]: The date on this thing is wrong. When we published this, it was April 2023. I think this was justVibhu [00:51:42]: YeahEiso Kant [00:51:42]: Happened on a migration, probably found it on archive.org.Vibhu [00:51:45]: Mistral.Eiso Kant [00:51:46]: Mistral had started, we started on the same month, right?Vibhu [00:51:49]: Yeah.Eiso Kant [00:51:49]: So this wasn't even, there was only, I think, Llama out at the timeVibhu [00:51:52]: SnellEiso Kant [00:51:52]: And that's it, right? And so, but I agree. I think we wan
This week, Jason Howell and Jeff Jarvis unpack a wild security disclosure from OpenAI and Hugging Face, where GPT-5.6 Sol found a zero-day vulnerability, broke out of its sandbox, and hacked Hugging Face's servers to steal the answer key to its own evaluation. They also dig into Moonshot's Kimi K3, the free Chinese model that got so popular it maxed out its own GPU capacity and sent Washington into a policy spiral over whether to panic or compete.Also in this episode: Google's new Gemini Flash models and the still-missing Gemini 3.5 Pro, publishers like USA Today and People Inc. weighing whether to block Google search entirely, Netflix revealing 300 titles used generative AI this year, AI companies buying and destroying millions of old books for training data, Samsung's AI glasses, the Suno hack, 1Password letting Claude log in for you, Thinking Machines' Inkling model, and NotebookLM becoming Gemini Notebook. New episodes every Wednesday at aiinside.show. Note: Time codes subject to change depending on dynamic ad insertion by the distributor. CHAPTERS: 0:00 - Start 0:03:36 - OpenAI says Hugging Face breach caused by one of its models 0:13:42 - Kimi K3: Open Frontier Intelligence 0:16:34 - Moonshot's Kimi AI Model Sets Off Anxiety in the US - Bloomberg 0:20:41 - Top Pentagon official blasts OpenAI's Dean Ball 0:29:36 - US, China to hold AI talks in September, sources say 0:35:15 - Google Ships New Gemini Flash Models, But Pro Is Still Missing 0:35:31 - Google Gemini Launch Delayed as Tech Falls Short of Internal Goals 0:54:03 - Netflix says around 300 titles used generative AI 0:54:27 - Netflix Co-CEO Explains How Gen-AI Was Used in 300 Different Titles: ‘We Believe It Is Going to Enhance Their Abilities' 0:56:45 - AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop 1:01:44 - A closer look at the upcoming Samsung AI glasses 1:03:53 - Hack Reveals Suno AI Music Generator Scraped YouTube, Deezer, and Genius 1:06:06 - 1Password now lets Claude sign in to websites without seeing your passwords 1:07:25 - Thinking Machines: Inkling: Our open-weights model 1:08:48 - Google is renaming NotebookLM to Gemini Notebook Hosts: Jason Howell and Jeff Jarvis Download and subscribe to AI Inside in audio and video: https://aiinside.show/ Support the podcast on Patreon for special perks: https://www.patreon.com/aiinsideshow. You'll get ad-free episodes, members-only Discord, T-shirts and stickers you love, and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Learn more about your ad choices. Visit megaphone.fm/adchoices
marimo is the fastest growing Python project with 22,000 Github stars. Fantastic episode today, covering:0:00 - Intro0:54 - Favorite keyboard2:35 - What Coreweave's GPUs mean for marimo, such as solving the Travelling Salesman Problem8:13 - What is the one thing Jupyter Notebooks always gets wrong, but marimo doesn't?17:43 - UI elements in marimo library like sliders, cascading updates across cells and marimo notebooks as web apps20:08 - What marimo means for different modern data team personas 24:39 - Remote storage connections and example of data analyst work27:15 - How marimo is used by various scientific disciplines and using Anywidget in marimo31:45 - Use of marimo in an enterprise context for GPU compute, roadmap comments and creative freedom the tool enables35:37 - Using marimo to defeat the Esri monopoly in geospatial: if you are David, play a different game than Goliath, again the point that marimo/molab offers creative freedom to sponaneously poke around in the data42:27 - How to use GPUs in marimo, limitations of this47:55 - Using AI to control marimo 50:30 - Call to action - marimo competition to make coolest notebook from an academic paper51:20 - Vincent's summary: this is marimo, go and be David, fight Goliath, most of all - creative freedom. [We are trying] to figure out a way for you to be more creatively free.51:35 - Dice kit55:48 - More on keyboards.57:25 - I try pairing Claude and marimo myself.A huge privilege to have Vincent on to tell us all about it. Here are some links to items discussed:marimo joining Coreweave: https://marimo.io/blog/joining-coreweavemarimo repo: https://github.com/marimo-team/marimomarimo pair - for helping pair with an AGI: https://github.com/marimo-team/marimo-pairA related blog post: https://marimo.io/blog/claude-codeEvidence marimo uses Leafmap: https://docs.marimo.io/api/plotting/?h=map#other-plotting-librariesAnywidget: https://anywidget.dev/ and its dev: https://github.com/manztNotebooks in Wherobots: https://wherobots.com/blog/wherobots-notebook-environment-getting-started-with-wherobots-cloud-sedonadb-part-2/Vincent's Wiggly Stuff: https://koaning.github.io/wigglystuff/marimo recently ended competition: https://marimo.io/pages/events/notebook-competition-2We are now discussing the next one and adding a geospatial theme to it!
Een Amerikaanse tiener heeft vlak voor de rechtszaak zijn verslavingsclaim tegen Meta ingetrokken, waardoor het bedrijf een proces in Los Angeles ontloopt. Ook sluiten de Amerikaanse chipmakers Intel en AMD langlopende overeenkomsten met Chinese klanten voor datacenterprocessoren, terwijl de AI-strijd tussen de VS en China in volle gang is. Joe van Burik vertelt erover in deze Tech Update. De zaak werd aangespannen door een 15-jarige jongen uit Florida, bekend onder de initialen R.K.C., die stelde depressief en angstig te zijn geworden door sociale media. Hij begon naar eigen zeggen op zijn achtste met de platforms, raakte verslaafd en sliep slecht. Naast Meta had hij oorspronkelijk ook YouTube-eigenaar Google, Snapchat-moederbedrijf Snap en TikTok-eigenaar ByteDance aangeklaagd. Die drie schikten de afgelopen weken tegen onbekende bedragen, waardoor Meta als enige overbleef voor het proces, dat komende week zou beginnen. Volgens Meta trok de jongen zijn claim in zonder dat het bedrijf iets betaalde. De advocaten van R.K.C. noemden de uitkomst een succes voor het aansprakelijk stellen van sociale media, en zeiden dat hij zich wil richten op zijn herstel en therapie. Ruim vijfduizend zaken lopen nog Meta reageerde dat de claims nooit standhielden en dat het zich blijft verdedigen tegen wat het grondeloze rechtszaken noemt. De ingetrokken zaak was de tweede in een reeks van meer dan drieduizend soortgelijke zaken in Californië, met daarnaast nog eens ruim tweeduizend zaken in de federale rechtbank van families, schooldistricten en procureurs-generaal. In de eerste zaak, in maart, oordeelde een jury dat Meta en YouTube aansprakelijk waren en werden zij veroordeeld tot een schadevergoeding van in totaal zes miljoen dollar. In een aparte zaak in maart, aangespannen door de procureur-generaal van New Mexico, kreeg Meta een boete van 375 miljoen dollar. Intel en AMD leggen Chinese afname vast nu prijzen stijgen Intel en AMD sluiten langlopende afspraken met Chinese serverklanten voor datacenterprocessoren, meldt Reuters op basis van twee bronnen. Het gaat niet om GPU's, de AI-chips waarmee vooral Nvidia groot werd, maar om CPU's, de centrale processoren die servers aansturen en die eveneens hard nodig zijn voor de ontwikkeling en toepassing van AI. De overeenkomsten leggen de afnamevolumes vast, maar niet de prijzen. De meeste dekken ongeveer een jaar, al is voor sommige klanten gesproken over twee jaar of langer. De serverprocessoren zijn schaarser geworden: sommige prijzen in China stegen maandelijks meer dan tien procent, en sinds begin dit jaar liepen bepaalde producten met meer dan veertig procent op. Intel presenteert donderdag zijn kwartaalcijfers, waarbij het chiptekort waarschijnlijk aan bod komt. De Amerikaanse overheid bezit inmiddels tien procent van de aandelen in Intel. Florida-tiener trekt verslavingszaak tegen Meta in vlak voor proces Tiener laat claims tegen Meta vallen dagen voor rechtszaak Meta en YouTube eerder aansprakelijk gesteld in eerste verslavingszaak Intel en AMD sluiten langlopende server-CPU-deals met Chinese klanten Chiptekort drijft prijzen op en duwt Chinese klanten naar langere contracten Over de maker:Joe van Burik volgt en duidt de belangrijkste ontwikkelingen in tech, met scherpte, vlotheid en de nodige humor. Je hoort hem dagelijks op BNR Nieuwsradio over het belangrijkste technieuws, van AI tot cybersecurity en social media tot quantumcomputers. Ook interviewt hij in De Grote Tech Show samen met Ben van der Burg leiders in digitale innovatie. In het bijzonder volgt Joe al twee decennia de wereld van videogames, nu voor zijn podcast All in the Game.See omnystudio.com/listener for privacy information.
AI's next phase hinges on a paradox: falling costs could threaten today's winners while unlocking far greater demand. Steve Hou, head of research at Silicon Data and former Bloomberg strategist, joins us to examine the changing economics of AI compute. We discuss token efficiency, model routing, GPU pricing, memory bottlenecks, and when enterprise adoption may finally deliver measurable returns. Enjoy! TIMESTAMPS: 00:00 Intro 01:01 Why AI Compute Needs Hedging 06:55 What The Token Index Really Shows 14:04 Token Maxing Meets Efficiency 18:35 Who Captures AI's Value? 22:12 Old GPUs Reveal Surging Demand 27:10 GPU Markets Keep Tightening 32:03 The Memory Bottleneck 37:07 AI's Next Phase FOLLOW STEVE › X/Twitter – https://x.com/stevehou › Silicon Data – https://www.silicondata.com/ FOLLOW THE SHOW › Forward Guidance – https://x.com/ForwardGuidance › Felix – https://x.com/fejau_inc › Telegram – https://t.me/+CAoZQpC-i6BjYTEx › Blockworks – https://x.com/Blockworks EVENTS › Join us at Digital Asset Summit 2026 Asia October 7th & Digital Asset 2026 London November 10-11th https://blockworks.com/events DISCLAIMER Nothing said on Forward Guidance is a recommendation to buy or sell securities or tokens. This podcast is for informational purposes only. Any views expressed are opinions, not financial advice. Hosts and guests may hold positions in the companies, funds, or projects discussed.
If anyone builds superintelligent AI before we know how to control it, everyone dies. Nate Soares wrote the book on why that's not a metaphor. Subscribe if you want science with evidence, not speculation. Soares runs the Machine Intelligence Research Institute and co-wrote If Anyone Builds It, Everyone Dies with Eliezer Yudkowsky. The first word in that title is if. That matters. His argument is not that doom is certain. His argument is that the path we are on leads there, that the driver is asleep at the wheel, and that we still have time to wake him up. We argue for over an hour. I push on whether LLMs can ever reach superintelligence, whether GPU lock-in is a real ceiling, and what it would actually take to move his p-doom. He pushes back with one clean point: by the time an AI can rediscover general relativity from pre-1911 data the way Einstein did, we will have almost no time left. You don't wait for that goalpost. What you'll hear: Why the bus-racing-toward-a-cliff analogy depends entirely on whether the driver is asleep or awake Whether LLM lock-in is a prison or a temporary inefficiency What the AI that broke out of its virtual machine to solve a hacking problem tells us Why GPT-4o encouraging a teenager toward suicide is not a malice problem but a training problem The difference between an AI doing the right thing too well and an AI that never wanted to do what you asked What Soares actually thinks about aliens, Dyson spheres, and why we should not see stars going out The first word in the title is if. The second word to watch is would. CHAPTERS 00:00 The people racing to build superhuman AI say it might kill everyone 00:42 Who coined "AI alignment" and why the first word in the title matters 02:28 Is it already too late for the if? 04:40 The bus, the cliff, and the sleeping driver 05:02 Silicon Valley is spooked. Washington is not. 07:02 Align with who? The rogue actor problem 07:34 Who is holding the leash? 08:24 The AI that edits its own test and deletes the log file 10:04 Controllability vs. making an AI that actually cares 10:44 The move gets harder. The outcome gets easier. 13:04 Are GPUs and LLMs a ceiling or a temporary inefficiency? 16:56 Brian's Einstein test: can an LLM rediscover general relativity? 18:38 Waiting for the goalpost is waiting too long 20:14 How prediction training can push AI beyond humans 21:44 Tycho Brahe, Kepler, and planetary motion as a prediction problem 24:38 Yann LeCun said never. GPT-4 did it half a generation later. 27:28 Can you prove a no-go theorem for superintelligence? 29:14 Training a human takes a light bulb. Training an AI takes a city. 33:28 What would proof of alien life do to p-doom? 35:00 Why interstellar aliens should have Dyson spheres 44:26 What would actually update Soares' p-doom? 49:42 Nobody intended this. Intent doesn't matter. 51:08 The AI hides its tracks before it does what you want 51:34 Sycophancy vs. hallucination: which runs deeper? 51:56 Leaded gasoline and civilizational risk 59:48 Sam Harris: humans have no free will but AIs do 01:00:38 Is alignment really a governance problem? 01:01:48 Unaligned AI vs. AI aligned to the wrong person 01:04:20 2026: 10 to 30% chance of automated AI research this year 01:06:44 What if Soares is wrong? 01:09:18 What gets him out of bed 01:12:38 Watch my conversation with Roman Yampolskiy Get the transcript, fascinating bonus content, and my Monday M.A.G.I.C. Message: https://briankeating.com/yt All my top AI episodes in one place: https://briankeating.com/ai Have a .edu email and live in the USA? You automatically win a meteorite: https://BrianKeating.com/edu Subscribe: https://www.youtube.com/DrBrianKeating?sub_confirmation=1 Support Into the Impossible on Patreon, get my weekly M.A.G.I.C. Message, unfiltered bonus content, and live monthly Office Hours with me: https://www.patreon.com/drbriankeating Join this channel for perks, monthly Office Hours, and your name in the Member Roster at the end of every episode: https://www.youtube.com/channel/UCmXH_moPhfkqCk6S3b9RWuw/join Featured Guest: Nate Soares / MIRI: https://intelligence.org If Anyone Builds It, Everyone Dies (book): https://ifanyonebuildsit.com/ Nate Soares on Twitter/X: https://x.com/So8res?lang=en My books: Losing the Nobel Prize (memoir): http://amzn.to/2sa5UpA Think Like a Nobel Prize Winner: https://a.co/d/03ezQFu Focus Like a Nobel Prize Winner: https://a.co/d/hi50U9U Galileo's Dialogue (first-ever audiobook): https://a.co/d/iZPi9Un Twitter/X: https://x.com/BrianKeating Substack: https://briankeating.substack.com Blog: https://briankeating.com/blog Audio-only: https://briankeating.com/podcast #intotheimpossible #briankeating #AIrisk #aisafety #artificialintelligence #superintelligence #NateSoares #MIRI #podcast Learn more about your ad choices. Visit megaphone.fm/adchoices
In this episode of the Crazy Wisdom Podcast, host Stewart Alsop speaks with Aaron Neyer, founder of Parachute and community organizer in Boulder, about knowledge management, extended minds, and the intersection of AI with human consciousness. They explore how Parachute functions as a digital brain tool for organizing thoughts and information across fragmented systems, discuss the dangers of AI psychosis and over-reliance on technology, and debate open source AI development versus controlled releases by companies like Anthropic. The conversation weaves through topics including the limitations of metrics-driven business thinking, consciousness and relevance realization, the value of technological sabbaths, and Aaron's hope for locally-run open source models that protect personal data while still accessing more powerful gated models when needed. You can find Aaron's writing at unforced.org and unforced.substack.com, and learn more about Parachute at parachute.computer and parachute.computer/blog.Timestamps00:00 Stewart welcomes Aaron Neyer, founder of Parachute and Boulder community organizer, discussing the origin of Parachute's name from Frank Zappa's quote about open minds.05:00 Aaron explains Parachute as an extended mind tool for organizing notes, contacts and information across multiple platforms, emphasizing the distinction between primary mind and extended mind as interconnected systems.10:00 Discussion shifts to metrics-driven business culture and the limitations of pure rationality, exploring how Google's data-driven approach misses subjective experience and the whole picture of relationships.15:00 Aaron discusses AI's ability to help identify relevant variables across different domains and the dangers of AI psychosis, comparing it to cult dynamics and belief systems.20:00 The conversation covers AI sabbaths and nineties retreats as intentional breaks from technology, plus Aaron's experiences with electrical engineering and using AI to design circuits with Arduinos.25:00 Exploring forbidden knowledge and open source AI, Aaron discusses Anthropic's guardrails around powerful models while arguing for distributed access to prevent concentration of power.30:00 Deep dive into open source AI strategy, with Aaron highlighting NVIDIA's approach and the potential for running capable models locally while reserving ultra-intelligent models for complex research tasks.35:00 Aaron shares his vision for local Sonnet-class models handling personal data while accessing Fable-class models for deep research, and directs listeners to unforced.org and parachute.computer for his writing.Key Insights1. The philosophy behind Parachute stems from Frank Zappa's quote that the mind is like a parachute and doesn't work if it isn't open. Aaron Neyer explains that having an open mind is valuable, but it must be balanced with deep roots to avoid becoming untethered. He has experienced periods in his life where excessive openness led him to feel disconnected, teaching him that creativity and expansion need to be grounded in something substantial. This same principle applies to how we organize information digitally, where openness and interoperability allow our extended minds to become more connected and coherent, which in turn helps our primary minds think more clearly.2. Parachute is designed as an extended mind tool that addresses the fragmentation problem in how we currently manage information. Most people use multiple disconnected tools like Obsidian, Notion, Apple Notes, Google Keep, and various CRMs to organize their thoughts, notes, and relationships. These systems don't communicate well with each other, creating inefficiency and confusion. Parachute aims to create a simple, intuitive system where all this information can be organized in one place with true interoperability, allowing users to own their data and have it speak effectively with other tools, ultimately making our entire extended mind more functional.3. Understanding ourselves as unified body mind organisms rather than fragmented parts is essential for effectiveness. Living systems theory shows that any living system is three things: a membrane bound dissipative structure, a self regulating autopoietic network, and a cognitive process actively knowing the world. Western civilization since Descartes and Galileo has created artificial separation between body and mind, and between subjective and objective experience, which limits our effectiveness. The same fragmentation affects our digital technology, and recognizing both our biological and digital systems as coherent wholes rather than disconnected parts makes us vastly more capable.4. The relationship between data driven approaches and holistic thinking reveals important limitations in modern business and science. While working at Google, Aaron observed how data driven decision making can be powerful, but over reliance on metrics like ROI creates blindness to crucial unmeasurable factors like goodwill and relationship quality. This reflects a broader Western tendency to exclude subjective experience because it's difficult for objective science to measure. However, emotions, relationships, and other subjective elements are essential parts of reality, and focusing only on quantifiable metrics means missing the whole picture and ultimately becoming less effective despite appearing more rational.5. AI accelerates the ability to work with technical complexity by helping with relevance realization across domains where we lack expertise. In any specialized field, experts develop intuitive senses for which variables matter and can quickly identify problems, whether in computer troubleshooting, music, or cooking. AI's ability to generalize allows it to point people toward relevant solutions in areas where they haven't developed that intuitive expertise, effectively democratizing technical capability. This means people can direct their creativity more effectively across more domains, though it also raises concerns about giving powerful capabilities to those who may lack the wisdom to use them responsibly.6. The question of open source AI versus gated access involves complex tradeoffs between democratizing power and preventing harm. Aaron respects Anthropic's approach of creating guardrails around powerful models like Mythos, which would likely have caused significant system hacks if released without restrictions. However, this creates concerning power dynamics where only wealthy companies, governments, and their allies have access to the most powerful tools. NVIDIA offers hope through their truly open source approach including full training pipelines, and there may be a viable path where open source models at the Sonnet capability level handle most tasks locally while more powerful Fable class models remain gated for the most demanding work.7. Creating intentional breaks from AI and technology is essential for maintaining clear independent thinking. Aaron practices an AI Sabbath at least one day per week when he doesn't interact with AI, and he finds these are the days when he does his best thinking and journaling. Without these breaks, he finds himself constantly jumping between journaling and prompting AI rather than giving himself space for deep reflection. This pattern mirrors broader concerns about AI consistency creating cult like dynamics similar to organized religion, where constant immersion in a belief system or technology can lead to losing the ability to think independently, making periodic disconnection crucial for maintaining cognitive autonomy and clarity.
There is a dramatic push of technology into space. Low Earth Orbit (LEO) satellites are changing the broadband communications market and there's speculation about putting data centers in orbit. Analysts Ellie Brown and Johan Vermij join host Eric Hanselman to look at the practical aspects of this push into a new frontier as a preview to their upcoming webinar. The cost of access to space has been dropping significantly, but there are many considerations that impact the viability of some of these plans. LEO satellite lifetimes are shorter than geosynchronous ones and the sustainability impacts of the launch volumes necessary to maintain multiple, massive constellations are daunting. Space is an attractive vantage point for observation with imaging to support geopolitical needs and ecological ends, like methane emissions tracking. Synthetic Aperture Radar (SAR) is piercing cloud and vegetation layers to reveal landscape details. Space-based data centers are seeing a lot of interest, but popular misconceptions about cooling capabilities may be adding more encouragement than is warranted. While space is "cold", radiating heat is more complicated in its vacuum. The International Space Station, with a massive radiation system, rejects only 70 KW of heat, less than a few racks of equipment and much less than a typical GPU cluster. And the need to service equipment on orbit hasn't been meaningfully addressed. Smaller satellites with distributed processing could pre-process data before it's sent to earth, but that will require greater levels of coordination to manage the massive scale on which this would have to operate. More S&P Global Content: Join the webinar: Orbit as the Next Data Frontier: Capital Flows, Risk, and Realities For S&P Global Subscribers: NVIDIA gives 'space cred' to orbital data center race with Vera Rubin Space module announcement STMicroelectronics eyes $3 billion in annual space revenue as demand for LEO satellites surges As the space sector enters the infrastructure era, M&A could become a force to be reckoned with Host/Author: Eric Hanselman Guests: Ellie Brown and Johan Vermij Producer/Editor: Dylan Scheible Published With Assistance From: Feranmi Adeoshun and Sophie Carr
When you see 30 cars launch off a cliff to beautiful destruction you have to talk about it. After that we discuss Samsung and the slump in stock price after the latest profit beat. Where is memory and GPU going? What craziness are we in now? We smoke the Rojas Unfinished Business and drink the Remus Lou Gehrig 0367 of 9665. https://www.msn.com/en-us/video/peopleandplaces/watch-this-car-fly-off-a-cliff-into-a-huge-crowd-of-people/vi-AA1ZfoGd?ocid=socialshare
JDK 27 entered Ramp-Down-Phase 1 in early June, so we can go over its features. Valhalla's first JDK Enhancement Proposal (401) will likely be part of JDK 28. Show off your Java ideas in the "Modern Java in the Wild" hackathon. There are two interesting articles for those who want to dive deeper into AOT with agents or into using GPU tensor cores from Java. Links: JDK 27 JEP list: https://openjdk.org/projects/jdk/27/ Valhalla mail: https://mail.openjdk.org/archives/list/jdk-dev@openjdk.org/threadAIA3O3LHFZ6T7TIPH7KZT4WS4B6U72U5/ JEP 401 PR: https://github.com/openjdk/jdk/pull/31120 "Java in the Wild" hackathon: https://www.hackster.io/contests/modern-java-in-the-wild AOT vs agents: https://bugs.openjdk.org/browse/JDK-8387190 AOT and JVMTI agents: https://bugs.openjdk.org/browse/JDK-8387194 "Exploiting GPU Tensor Cores": https://openjdk.org/projects/babylon/articles/hat-tensors/hat-tensors
DigiPower X (DGXX) CEO Michel Amar breaks down the company's milestone contract with Cerebras (CBRS)… its AI stack roadmap, from real estate to GPU-as-a-service… key catalysts through 2026… and why DigiPower X is in a league of its own. In this episode: Welcome back, Michel Amar, CEO of DigiPower X [0:01] DigiPower's $1 billion contract with Cerebras is a major milestone [0:53] Why DGXX is exempt from New York's moratorium on data centers [6:33] From real estate to GPU-as-a-service: DigiPower's full AI stack roadmap [10:51] Management's plans to keep executing through 2026 and beyond [19:49] DigiPower X is in a league of its own [24:43] Did you like this episode? Get more Wall Street Unplugged FREE each week in your inbox. Sign up here: https://curzio.me/syn_wsu Find Wall Street Unplugged podcast… --Curzio Research App: https://curzio.me/syn_app --iTunes: https://curzio.me/syn_wsu_i --Stitcher: https://curzio.me/syn_wsu_s --Website: https://curzio.me/syn_wsu_cat Follow Frank… X: https://curzio.me/syn_twt Facebook: https://curzio.me/syn_fb LinkedIn: https://curzio.me/syn_li
Episode 107: We chat about some of the latest CPU releases, including the Ryzen 7 5800X3D 10th Anniversary Edition, and the Ryzen 7 7700X3D, which doesn't make a ton of sense when nobody is doing platform upgrades. Also, we thought it was quite amusing that Nvidia's hotspot GPU temperature has now been uncovered, so we discuss that whole situation.CHAPTERS00:00 - Intro01:02 - The Ryzen 7 7700X3D has Launched08:32 - The Ryzen 7 5800X3D Actually Makes Sense24:53 - Nvidia GPU Hotspot Temperature Situation40:24 - Steam Machine Thermals and HDMI-CEC01:04:55 - Updates From Our Boring LivesSUBSCRIBE TO THE PODCASTAudio: https://shows.acast.com/the-hardware-unboxed-podcastVideo: https://www.youtube.com/channel/UCqT8Vb3jweH6_tj2SarErfwSUPPORT US DIRECTLYPatreon: https://www.patreon.com/hardwareunboxedLINKSYouTube: https://www.youtube.com/@Hardwareunboxed/Twitter: https://twitter.com/HardwareUnboxedBluesky: https://bsky.app/profile/hardwareunboxed.bsky.social Hosted on Acast. See acast.com/privacy for more information.
This week we let ChatGPT help pick the news stories (which seemed like a good idea at the time), and honestly, it kind of worked.We start with Russia reportedly having to buy gasoline after Ukrainian drone strikes hit their refinery capacity, which turns into a bigger conversation about energy markets, gasoline and diesel, and how cheap drones are completely changing the math of modern warfare.Then we get into Japan potentially pushing its giant pension funds to bring more money back home, the AI trade still running through SK Hynix and GPU demand, and the Fed now treating AI infrastructure spending as one of the things keeping inflation hotter than expected.So there is real market stuff in here.But also, this is Dan and me, so we also talk about a squirrel getting loose in a Meta office, a bear breaking into a shopping mall on a military base, Britain trying to survive heat waves with better refrigerators and tiny fans, Pepsi making a very questionable Wild Cherry post in India, AI reading ancient Roman scrolls, Pompeii, Pink Floyd, and whether Wendy's Biggie Bag is secretly one of the better recession trades on the board. Chapters00:00 — Hook, Intro, and Cold Open01:47 — Welcome Back to the China Shop02:30 — Letting ChatGPT Pick the News03:22 — Russia's Fuel Problem and Ukrainian Drone Strikes06:36 — EVs, Refinery Risk, and Energy Independence09:10 — A Squirrel Gets Loose at Meta14:28 — Japan's Pension Fund and Money Coming Home20:08 — The Alaska Mall Bear Story21:54 — SK Hynix, AI Chips, and Inflation Pressure25:24 — FedWatch, Rate Expectations, and No Cuts Yet27:38 — Britain's Heat Wave Economy32:49 — Scottsdale Will Pay for Your Lawn, Not Your Pool34:29 — Pepsi's Wild Cherry Problem38:33 — AI, Ancient Scrolls, and the Pompeii Detour42:28 — Wendy's, Biggie Bags, and Takeover Rumors48:38 — Jobs, Rent, and the Economy People Actually Feel50:30 — Chessferatu and Closing ThoughtsSubscribe, share, and join the trading conversations on Facebook, Twitter, LinkedIn and Discord!Sponsors and FriendsOur podcast is sponsored by Sue Maki at Fairway Independent Mortgage (MLS# 206048). Licensed in 38 states, if you need anything mortgage-related, reach out to her at SMaki@fairwaymc.com or give her a call at (520) 977-7904. Tell her 2 Bulls sent you to get the best rates available!If you are interested in signing up with TRADEPRO Academy, you can use our affiliate link here. We receive compensation for any purchases made when using this link, so it's a great way to support the show and learn at the same time! **Use code CHINASHOP15 to save 15%**To contact us, you can email us directly at bandoftraderspodcast@gmail.comIf you like our show, please let us know by rating and subscribing on your platform of choice!If you like our show and hate social media, then please tell all your friends!If you have no friends and hate social media and you just want to give us money for advertising to help you find more friends, then you can donate to support the show here!Dan:Dan co-founded 2 Bulls in a China Shop with Kyle when their shared passion for active trading ignited during the lockdowns. Their daily discussions about trades, interests, and the valuable lessons learned created the bedrock for what eventually evolved into both the 2 Bulls in a China Shop and Band of Traders podcasts.While navigating the complexities of trading, Dan infused humor into the shows with his self-deprecating wit and candid discussions about their trading experiences. This dynamic duo's chemistry became the catalyst for a podcast that resonated widely, capturing the attention of a diverse audience.Download ChessFeratuHalf-Cocked TalesAdvertising Inquiries: https://redcircle.com/brandsPrivacy & Opt-Out: https://redcircle.com/privacy
As AI applications become more complex, the infrastructure powering them needs to evolve. Corey Sanders, SVP of Product at CoreWeave, joins Chris to discuss why AI requires a fundamentally different approach than traditional cloud computing. They explore AI-native infrastructure, training and inference workloads, the rise of agentic development, optimizing GPU performance, AI research workflows, and why the future of software will be built around AI-first experiences rather than websites and apps.Featuring:Corey Sanders – LinkedIn Chris Benson – Website, LinkedIn, Bluesky, GitHub, XLinks:CoreWeaveSponsors:Framer: The enterprise-grade website builder that lets your team ship faster. Get 30% off at framer.com/practicalaiUpcoming Events: Register for upcoming webinars here!Midwest AI Summit 2026
Credit markets are stepping in to fund the surging demand for AI. Our experts Lindsay Tyler and Anish Shah explore the opportunities and risks behind this record financing wave.Read more insights from Morgan Stanley.----- Transcript -----Lindsay Tyler: Welcome to Thoughts on the Market. I'm Lindsay Tyler, TMT Credit Research Analyst at Morgan Stanley. Anish Shah: And I'm Anish Shah, Global Head of Debt Capital Markets at Morgan Stanley. Lindsay Tyler: Today, how issuers and investors are approaching the rapidly evolving world of AI financing. It's Thursday, July 16th at 10am in New York. As AI demand accelerates, credit markets are being asked to finance infrastructure on a scale that used to be associated with utilities, telecom, or energy. That raises a central question for issuers and investors: How much debt can the AI ecosystem absorb? And at what price? Anish, can you walk our listeners through the key products in your purview? Anish Shah: Certainly, in my nearly twenty years at Morgan Stanley, this is probably the most incredible time period I've ever seen in the credit markets. I've had the privilege of working across a number of different roles in capital markets and lending. And a couple of years ago, we integrated the debt underwriting business across both investment-grade and leverage finance franchises in recognition of how interconnected the whole credit ecosystem has become. In addition to our core activities helping clients raise capital for their strategic priorities, two of the big focus areas that we've had have been finding ways to harness the power of the private credit universe and also delivering best-in-class capabilities in funding this incredible growth in AI spend. Lindsay Tyler: AI financing has certainly been a theme we've also been focused on in research. Our equity research colleagues project that a handful of key players could add more than 30 gigawatts of capacity over a two-year timeframe, driving around [$]2 trillion of aggregate cash CapEx in that period. And to put that into context, a single gigawatt of data center capacity can require roughly $12 billion for the shell, and then often more than double that for chips and racks. So, from your vantage point, what inning are we in? And what gives you confidence that credit markets can continue funding this opportunity at scale? Anish Shah: I mean, Lindsay, the numbers certainly are staggering, as you note. And if you just observe the CapEx estimates for the hyperscalers and broadly for AI infrastructure, we're certainly in the early innings. Lindsay Tyler: Mm-hmm. Anish Shah: The largest tech companies have historically, as you know, raised very little debt. In fact, many of these companies have not even needed a credit facility. As CapEx projections were materially increased in the second half of last year, we saw the beginning of scaled capital raises. Hyperscaler issuance has quickly gone from less than one percent of the investment-grade market to more than 10 percent of the market. You know, as I look ahead, based on what we're seeing on the ground, we think that AI-related funding, whether it's for data center development or financing compute capacity, could top 15 percent of the total issuance across all credit products. This has been an unprecedented test for the capital markets, both in terms of the depth of capacity and the breadth of product. The teams have been on the forefront of deep investor dialogue and product innovation. This spans corporate investment grade, first of their kind financings in high-yield and leveraged loan markets, and new takes on asset-backed financing. And each of these areas has seen material issuance both in public and private markets. Lindsay Tyler: Great backdrop. Let's dig first into investment-grade corporate debt, an area you know well from your time previously leading the investment-grade team. Can you help frame the scale and the significance of this financing bucket and how AI-related debt is scaling within it? Anish Shah: Well, you know, as you know, the investment-grade bond market, specifically in dollars, is the deepest, most liquid pool of capital in the world. Volumes have grown materially over the last few years and are likely to eclipse $2 trillion in issuance this year. Hyperscalers are among the very best credits in the world, and they have the ability to come in and out of markets with relatively quick twitch, little to no pre-marketing, and in fairly large size. You know, $20 billion-plus deals used to be rare in the investment-grade market, now happen multiple times a quarter. This is why we've seen the predominance of AI-driven capital raising take place in the investment-grade market. For the most part, investors have digested that supply very well. While we've seen some modest widening credit spreads for hyperscalers and some of the other tech issuers, I'd say it's de minimis relative to their expected ROI. Lindsay, I've talked a lot about supply dynamics and issuance. What other factors are you and investors considering when assessing fair value for investment-grade rated technology bonds? Lindsay Tyler: Sure. It's prudent to really weigh a mix of technicals, fundamentals, and relative value. You know, as you discussed on the technical side, and related to my discussions with debt and equity investors, I've been focused on scale of buildouts, market capacity, digestibility across currencies, positioning along the curve, implications of equity issuance, and whether AI financing could crowd out other areas of TMT credit. But moving more to the fundamental side of things, you mentioned ROI, and for the players that are scaling compute capacity, there are a handful of key monetization and return questions that keep coming up. How quickly can these companies bring new capacity online? Once it's live, how does it translate into durable revenue and cash flow? Is that capacity supporting internal products, proprietary models, broader cloud offerings, or compute leased to third parties? And then how fungible is the capacity across those use cases if demand or returns shift? Further on the fundamental side, we've done some differentiated work around growing long-term commitments. We've seen that high-quality hyperscalers and a few of the semis companies are anchoring the AI ecosystem through leases, guarantees, other obligations. These commitments really extend beyond vanilla bond issuance. So, I encourage investors to look beyond the funded debt and really understand the accounting and the ratings implications here of some of those commitments. And this ties nicely into the next topic that I wanted to raise, which is project finance debt. I've noticed that, you know, a lot of the commitments that we're seeing from IG players support another layer of financing. Lease commitments can underpin project finance debt, an area of sizable issuance and innovation. The public high-yield market has emerged as a new funding source in this way for data center construction, with more than 30 billion priced across 15 deals, since fall 2025. Can you walk us through, Anish, the innovation behind these structures, and how are these high yield deals different than other ways to, kind of, raise project finance debt? Anish Shah: Yeah, it's incredibly interesting. I mean, the bulk of the issuance, as I noted has come in the investment grade market, but I would say the bulk of the innovation has come in the sub-investment grade market. You know, historically, for very capital-intensive sectors like energy and power or real estate, the project loan market was the most efficient source of initial funding. The developer would tap banks to underwrite a highly structured construction loan. Once the project is up and running, you could then refinance that loan with the predictable cash flows into a more institutional financing, like the investment grade bond market or the term loan B or securitization markets.That product may still be very viable in many sectors, but we felt early on that bank-provided construction loans would not meet the capacity needs of the AI investment cycle. The market really needed an institutional credit product that bypassed the need for construction loans. The key innovation came in the form of first-of-its-kind high-yield bonds that funded the development of a new data center complex. Given the relatively short construction period and the "offtake" supported by some of the highest quality credits in the world, we felt like this financing structure would be incredibly well-received in the high-yield market. The win here is that the developer accesses fixed rate long-term capital and maintains flexibility to call the bonds and refinance at a lower cost. Judging by how these financings have gone, there's a strong level of investor enthusiasm. I think that they've only scratched the surface, and I would expect that we see much more of this. And potentially even expand it to other products in the leverage finance markets given the tremendous level of investor demand. Lindsay Tyler: Yeah. It's certainly been exciting to follow many of those deals. Beyond the public space, we're also seeing a wave of innovation in private credit and asset-backed finance. Anish, how do companies decide whether capital is best raised in the public or the private markets? Anish Shah: Well, I'm glad you raised the whole avenue of private markets because it may be the most significant change in the credit markets over the last few years, broadening the scope of private credit from directly lending into leverage buyouts to now financing large investment-grade projects. There are great examples in the world of GPU and TPU financing, where we structure loans secured by the asset and the cash flows, or in data center development.Lindsay, from your perspective, what are investors focused on when these private structures intersect with public credits? Lindsay Tyler: Sure. Many of these asset-backed private financings have prompted investors to look more closely at any of the public companies involved, whether as issuers, tenants, customers, or support providers. This ties back to the point I raised earlier. Where does the risk reside, and who ultimately is on the hook? These financings have also sparked broader discussions around circularity, vendor financing, and technology obsolescence risk, even when amortizing structures are in place. I do think those are fair concerns to weigh, and they really speak to how quickly the AI financing trend is evolving and how much credit work there is to do. So, Anish, with that balance in mind, relatively strong demand, rapid innovation, but also some real credit questions, let's end with a quick lightning round. Anish Shah: Lindsay, let's do it. Lindsay Tyler: First, what is the biggest risk that could test investor appetite for AI-related debt? Anish Shah: I would say investors are acutely focused on construction delays. Don't underestimate the level of diligence being done by the breadth of capacity you're seeing in the markets. Investors are doing their homework, and we're spending a lot of time trying to mitigate any of their concerns with structural protections. Lindsay Tyler: Got it. Second, beyond data center shells and chips, what is the next potential AI financing opportunity? Anish Shah: It most certainly is energy and power. We're going to see a ton of capital being raised in utilities. It's going to be a little different than what the hyperscalers are doing, just given the nature of their balance sheets. You're going to see more junior capital. We've seen a wave of junior subordinated debt issuance out of the utilities. We're also seeing a lot of activity from our project finance and tax equity team, just given all things energy infrastructure. Lindsay Tyler: Great. And third, if we're sitting here a year from now, what do you think could be the biggest AI financing story we're talking about? Anish Shah: Well, we certainly underestimated the level of financing activity that we saw in the past year. I think when we look back a year from now, we will probably see that the AI labs were much more ready to finance on their own on a standalone basis. That's going to alleviate some of the pressures in the market, but I think it's going to create a whole new set of considerations and structural innovation. Lindsay Tyler: Well, it's certainly been remarkable to watch this financing theme take shape in real time, and the next chapter sounds like it could be even more interesting to follow. Anish, thanks for joining us and sharing your insights. Anish Shah: Great to join, Lindsay. Thanks. Lindsay Tyler: And thank you for listening. If you enjoy Thoughts on the Market, please leave us a review wherever you listen and share the podcast with a friend or colleague today.*****Anish Shah is a member of Morgan Stanley's Global Capital Markets Division and is not a member of Morgan Stanley's Research Department. Unless otherwise indicated, his views are his own and may differ from the views of the Morgan Stanley Research Department and from the views of others within Morgan Stanley.
You've probably never heard of Inkling. It's the newest (and first) model from Thinking Machines Labs, and it could very well be a small snowball that picks up major momentum in today's enterprise AI landscape. If you haven't heard of Thinking Machines, they're led by Mira Murati, the former CTO at OpenAI. The big bet with Inkling? The future of AI could be using smaller models fine-tuned and optimized for smaller tasks. Will it work? Tune in live as we dive in. The Most Important AI Model You'll Probably Never Use That Just Dropped -- An Everyday AI Chat With Jordan WilsonNewsletter: Sign up for our free daily newsletterMore on this Episode: Episode PageToday's Episode on LinkedIn: Thoughts on this? Join the convo on LinkedIn and connect with other AI leaders.Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineupWebsite: YourEverydayAI.comEmail The Show: info@youreverydayai.comConnect with Jordan on LinkedInTopics Covered in This Episode:Inkling AI Model Launch OverviewThinking Machines Lab Leadership HighlightInkling's Multimodal and Agentic CapabilitiesOpen Source vs. Proprietary AI ModelsEnterprise Procurement with American AI ModelsAI Fine Tuning as a Service (Tinker)Benchmark Scores: Inkling vs. Frontier ModelsCustomization and Model Shopping for EnterprisesAI Token Costs Driving Model EfficiencyBridgewater Case Study: AI Model CustomizationFrontier Models Enabling Efficient Fine-TuningFuture Trends: Specialized Small Language ModelsTimestamps:00:00 Inkling: A new AI model release05:43 Inkling AI model details09:08 China's dominance in open source AI11:48 Launch and model updates discussed15:21 Concerns over using Chinese open-source models19:06 Training smaller AI models20:22 Using GPT for AI Model Training23:54 Predicting Rise of Small Language Models28:38 Choosing the right AI modelKeywords: Inkling, Thinking Machines Lab, Meera Muradi, former OpenAI CTO, open source AI model, American AI model, fine tuning as a service, enterprise AI, multimodal AI, agentic models, customizable AI, Tinker, enterprise distribution, model procurement, Chinese open source models, strategic reset, model overhang, capabilities gap, AI model shopping, model routing, cost-conscious enterprises, artificial intelligence index, 975 billion parameter model, text-image-audio AI, open weights, proprietary AI models, customization accessibility, small language models, AI workflows, context window, Bridgewater use case, model distillation, GPU infrastructure, API costs, token efficiency, fine-tuned models, post training, AI competitive leverage, recurring financial judgment, AI benchmarks, middle tier models, automated model evaluation, privacy and workflow mapping, economical AI models, model rental, model routing automation.Send Everyday AI and Jordan a text message. (We can't reply back unless you leave contact info) Ready for ROI on GenAI? Go to youreverydayai.com/partner
Imagine a dark warehouse. Racks and racks of devices with wires, tubes, and electronics sticking out. The next AI data center? No. This is Lila Sciences‘ dream for the future of science. A dark warehouse full of AI-guided robotics and lab equipment, cranking out new experiments 24/7, building toward a scientific superintelligence.Their automated lab is almost hypnotizing to watch. They have floating plates zipping around on Wall-E-esque tracks, used vision-language models to control Windows 95 boxes, and created the world's largest collection of voided warranties. In the process they've built a massive library of scientific reasoning tokens. Over 10 trillion of them, all experimentally validated.No warranties were voided in the making of this videoTo say Lila is ambitious is an understatement. Their goal is a scientific superintelligence wired directly into the wet lab. They are all in on the bitter lesson, and the thesis follows from it: a lab is an infinite token generator. Produce data at scale, and the synergies give you a general reasoner that can tackle any scientific problem. They are committing hard. Biology, chemistry, drug discovery, and materials science, all at the same time. Time will tell if it works, but it is an exciting hypothesis.In our latest episode we sat down with Lila's very own Andy Beam (CTO) and Rafa Gómez-Bombarelli (CSO, physical sciences) and went on a journey through the possibilities of AI-run science, almost as wide-ranging as Lila's goals.Did we mention they do both materials science and biology? In the same AI science factory? Same time, same lab, same AI. Finally a guest who can settle a long-running debate we've had amongst ourselves: is biology or materials science harder?Watch to find out!We discuss:* The internet is spent, science is next. Why Lila thinks the scientific method is the last untapped internet-scale dataset, and why they treat RL as a data generation mechanism with nature as the verifier.* The lab as a data center. Instruments as nodes on a graph, a magnetically levitating “PCI bus” transport layer between them, orchestration as a slurm queue. Andy is not short on analogies.* Why Lila insists it is not an automation company. They optimize for flexibility and generalizability over raw throughput, which means humans stay below the API line wherever automating does not pay.* Your experiment has a runtime. We put Escalante Bio's question to Andy: if science is the token generator, what is the runtime of your data collection? His answer, in short, is that you cannot make the ribosome go faster. Why Lila bets on fast round-over-round iteration rather than big noisy multiplexed screens, and how Rafa's team rebuilt a gas sorption measurement to run roughly 2,500x faster.* What is actually in 10 trillion scientific tokens. Not sequences. Experimentally verified reasoning traces, a kind of data that Andy argues exists on the internet in quantities that round to zero.* Breadth as a path to depth. Small molecule chemistry priors transferring to metal organic frameworks for carbon capture, and the claim that the general model beats domain-specific models sample for sample.* If you have the data, what do you need the model for? Sri Kosuri's koan about the ML-for-drug-discovery business model, and Andy's answer: the coding model got better because it also read Shakespeare and carnitas recipes.* The serendipity they want to automate. Emily Whitehead survived the first pediatric CAR-T cure only because the doctor treating her happened to know, from pediatric arthritis, which antibody would blunt her IL-6 response. Roll that dice again and you probably lose her. Breadth is how you stop depending on luck.* Move 37 for catalysts. Model suggestions for platinum-group-free electrocatalysts that went from boring, to what a 40-paper expert called stupid, to the best performers they have made.* Six months to in vivo CAR-T data in non-human primates, and the zero-FTE virtual startup commercial model that fell out of it. For context on why that number is startling, AbbVie paid $2.1B for Capstan on the strength of preclinical in vivo CAR-T data.* You cannot have scientific superintelligence if you are just a good test taker. Ken Stanley, who wrote Why Greatness Cannot Be Planned, runs open-endedness at Lila. RL at scale gives you a ruthlessly Vulcan problem solver. Machine creativity is a different thing, and it is the part nobody has solved.* The chain of thought is an unreliable narrator. The model reasons in latent space and only emits tokens. Sometimes it skips the experiment entirely and is still right. So how much do you trust the reasoning versus the verifier?* Reward hacking when the rollout is physical. Chains of thought that collapse into repetition, and a model that got annoyed and swore at the scientist who kept asking it to redo a plate map. What happens when a pathological loop has a wet lab inside it?* The bittersweet lesson. Rafa's inversion of the bitter lesson: in AI, scaling is a roadmap. In materials, scaling is a filter, because only the things that scale end up mattering.* Not your typical Flagship company. Why a famously single-asset biotech incubator spun out a platform bet, and Andy's line that if Lila called itself a biopharma it would have a top-three GPU cluster.* Bottlenecks they would remove by fiat. Sim-to-real for physics-based simulation, and the fact that RL training runs at roughly 5% mean FLOP utilization.Watch on YouTube: This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe
As part of our summer replay series, we're revisiting one of our favorite conversations on the future of AI infrastructure. SemiAnalysis founder Dylan Patel joins Erin Price-Wright, Guido Appenzeller, and Erik Torenberg to examine the rapidly evolving economics of AI hardware, from GPUs and custom silicon to data centers, power, and the global race for compute. The conversation explores NVIDIA's competitive advantages, the rise of custom chips from Google, Amazon, and Meta, the economics of frontier AI models, and the infrastructure constraints shaping the industry's next phase. They also discuss AI startups, export controls, robotics, enterprise software, and why simply copying NVIDIA isn't enough to build a winning AI hardware company. Whether you're building AI products, investing in infrastructure, or trying to understand where the industry is headed, this conversation offers a practical look at the forces shaping the future of compute. Resources: Follow Dylan Patel on X: https://x.com/dylan522p Follow Erin Price-Wright on X: https://x.com/espricewright Follow Guido Appenzeller on X: https://x.com/appenz Learn more about SemiAnalysis: https://semianalysis.com/dylan-patel/ Stay Updated:Find a16z on YouTube: YouTubeFind a16z on XFind a16z on LinkedInListen to the a16z Show on SpotifyListen to the a16z Show on Apple PodcastsFollow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Small teams with smart AI can solve problems no one imagined before. In this episode, I spoke with Zhen Lu, CEO of Runpod, about the rapidly evolving AI landscape and the future of software development. Zhen Lu shared how Runpod is empowering engineers to train AI models tailored to specific business needs, improve efficiency, and reduce waste. We also discussed his founder journey, the power of co-founder alignment, global innovation outside traditional tech hubs, and the importance of human accountability in AI-driven organizations. Here are the highlights: ● Runpod builds AI developer infrastructure. The platform supports tailored AI workloads, enabling businesses to efficiently train models specific to their needs. ● AI is changing how software is built. Engineers are challenged to rethink software design, not just accelerate existing processes, creating new opportunities for innovation. ● Efficiency and sustainability matter. Fine-tuning smaller, focused models reduces resource waste, energy consumption, and operational costs compared to massive off-the-shelf models. ● Co-founder alignment drives success. Implicit trust, complementary skills, and low ego between Zhen and his co-founder have accelerated decision-making and execution. ● Global networks and bootstrapping foster innovation. Scarcity encourages creativity, and building strong relationships outside Silicon Valley has enabled Runpod to grow and support AI developers worldwide. About the guest: Zhen Lu and co-founder Pardeep Singh started by running crypto mining rigs out of their New Jersey basements. When Ethereum's "The Merge" threatened to make that obsolete, they pivoted, converting the rigs into AI servers. As corporate developers at Comcast building ML projects, they saw firsthand that the GPU developer experience was, in Zhen's words, "just hot garbage." That insight became Runpod. Launched in early 2022, Runpod offers fast, developer-friendly GPU infrastructure: clean APIs, CLI tools, serverless options, and easy configuration. Rather than the traditional VC route, they debuted on Reddit with "Hey Reddit, give us your worst" — and got "please take my money" in return. They went on to earn validation from Hugging Face co-founder Julien Chaumond and a seed round led by Dell Technologies Capital, hitting $120M ARR, 10 billion serverless requests, and serves ~900,000 developers across 31 global regions. Connect with Zhen Lu: LinkedIn: https://www.linkedin.com/in/zeen/ Website: https://www.runpod.io/ Connect with Allison: Feedspot has named Disruptive CEO Nation as one of the Top 25 CEO Podcasts on the web. LinkedIn: https://www.linkedin.com/in/allisonsummerschicago/ Website: https://www.disruptiveceonation.com/ #CEO #leadership #startup #founder #business #businesspodcast Learn more about your ad choices. Visit megaphone.fm/adchoices
Small teams with smart AI can solve problems no one imagined before. In this episode, I spoke with Zhen Lu, CEO of Runpod, about the rapidly evolving AI landscape and the future of software development. Zhen Lu shared how Runpod is empowering engineers to train AI models tailored to specific business needs, improve efficiency, and reduce waste. We also discussed his founder journey, the power of co-founder alignment, global innovation outside traditional tech hubs, and the importance of human accountability in AI-driven organizations. Here are the highlights: ● Runpod builds AI developer infrastructure. The platform supports tailored AI workloads, enabling businesses to efficiently train models specific to their needs. ● AI is changing how software is built. Engineers are challenged to rethink software design, not just accelerate existing processes, creating new opportunities for innovation. ● Efficiency and sustainability matter. Fine-tuning smaller, focused models reduces resource waste, energy consumption, and operational costs compared to massive off-the-shelf models. ● Co-founder alignment drives success. Implicit trust, complementary skills, and low ego between Zhen and his co-founder have accelerated decision-making and execution. ● Global networks and bootstrapping foster innovation. Scarcity encourages creativity, and building strong relationships outside Silicon Valley has enabled Runpod to grow and support AI developers worldwide. About the guest: Zhen Lu and co-founder Pardeep Singh started by running crypto mining rigs out of their New Jersey basements. When Ethereum's "The Merge" threatened to make that obsolete, they pivoted, converting the rigs into AI servers. As corporate developers at Comcast building ML projects, they saw firsthand that the GPU developer experience was, in Zhen's words, "just hot garbage." That insight became Runpod. Launched in early 2022, Runpod offers fast, developer-friendly GPU infrastructure: clean APIs, CLI tools, serverless options, and easy configuration. Rather than the traditional VC route, they debuted on Reddit with "Hey Reddit, give us your worst" — and got "please take my money" in return. They went on to earn validation from Hugging Face co-founder Julien Chaumond and a seed round led by Dell Technologies Capital, hitting $120M ARR, 10 billion serverless requests, and serves ~900,000 developers across 31 global regions. Connect with Zhen Lu: LinkedIn: https://www.linkedin.com/in/zeen/ Website: https://www.runpod.io/ Connect with Allison: Feedspot has named Disruptive CEO Nation as one of the Top 25 CEO Podcasts on the web. LinkedIn: https://www.linkedin.com/in/allisonsummerschicago/ Website: https://www.disruptiveceonation.com/ #CEO #leadership #startup #founder #business #businesspodcast Learn more about your ad choices. Visit megaphone.fm/adchoices
As part of our summer replay series, we're revisiting one of the standout conversations from Runtime, a16z's conference on AI infrastructure and the future of computing. Gavin Baker, Managing Partner and CIO of Atreides Management, joins David George to examine the biggest questions surrounding today's AI investment cycle. Is AI a bubble? What does the unprecedented buildout of data centers, GPUs, and compute infrastructure mean for the economy? And how should investors think about the companies building the next generation of AI? The conversation explores frontier models, Nvidia, Google, custom silicon, AI infrastructure, application software, robotics, and why Baker believes today's AI investment cycle looks fundamentally different from the internet bubble of the early 2000s. Along the way, they discuss the economics of GPUs, enterprise software, AI business models, and what comes next as AI moves from experimentation into the broader economy. Resources: Follow Gavin Baker on X: https://x.com/GavinSBaker Follow Atreides Management on X: https://x.com/atreidesmgmt Follow David George on X: https://x.com/DavidGeorge83 Stay Updated:Find a16z on YouTube: YouTubeFind a16z on XFind a16z on LinkedInListen to the a16z Show on SpotifyListen to the a16z Show on Apple PodcastsFollow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
While LLMs and agents are new to appsec and everyone else, a lot of AI security requirements translate to well-known API security requirements. Jeremy Snyder helps us frame the OWASP LLM Top 10 into five layers in order to help orgs understand and prioritize their attack surface. A lot of orgs don't have to deal with model-specific threats or building their own GPU architecture, but every org adopting LLMs and agents should be aware of how those agents are being invoked and the output those agents are producing. That awareness of input and output helps in identifying and mitigating prompt injection attacks, ensuring agents are working within their expected boundaries, and taming token budgets. Resources: https://genai.owasp.org/llm-top-10/ https://github.com/rtk-ai/rtk https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-caching.html https://www.firetail.ai/blog/beyond-the-spectacle-rsac-2026-and-the-5-layers-of-ai-security Visit https://www.securityweekly.com/asw for all the latest episodes! Show Notes: https://securityweekly.com/asw-391
Anthropic had a really bad week.