POPULARITY
Categories
OneChronos CEO Kelly Littlepage discusses the critical need for a financial market for GPU compute, comparing it to the commodities market for crude oil. Such a market would foster efficient capital formation and enable the underwriting of larger AI projects. He identifies three key participants: off-takers, project builders, and lenders. Increased transparency in compute demand could also benefit companies like Nvidia (NVDA) by expanding the overall market.======== Schwab Network ========Empowering every investor and trader, every market day.Subscribe to the Market Minute newsletter - https://schwabnetwork.com/subscribeDownload the iOS app - https://apps.apple.com/us/app/schwab-network/id1460719185Download the Amazon Fire Tv App - https://www.amazon.com/schwab-Network/dp/B08JJRQG9T/Watch on Sling - https://watch.sling.com/1/channel/bb1b75050268416e82a557ff6387bff3/browseWatch on Vizio - https://www.vizio.com/en/watchfreeplus/catalog/live-tv-channels/3123029569/schwab-networkFollow us on X – https://twitter.com/schwabnetworkFollow us on Facebook – https://www.facebook.com/schwabnetworkFollow us on LinkedIn - https://www.linkedin.com/company/schwab-network/About Schwab Network - About | Schwab Network
Edge computing sounds clean on a slide deck, but the factory floor is where the definition gets real fast. We talk with Zoie Rittling, an edge computing specialist at TEGUAR, about what “compute close to the data” actually means when you're dealing with heat, dust, vibration, uptime pressure, and constant change. If you have ever wondered why a normal laptop dies on a production line, or what makes an industrial box PC, panel PC, or edge server the right fit for OT workloads, we break it down in plain language with concrete examples like machine vision and quality inspection.From there, we shift into the human side of industrial automation: how consultative B2B hardware sales works when you cannot manufacture demand and budget cycles run the show. Zoie shares why cold outbound is really a relationship game, why patience pays dividends over nine to eighteen month deal cycles, and how trade shows succeed or fail based on the planning and the post-show follow-up. We also talk about building a collaborative inside sales culture that protects customer experience instead of letting reps compete over the same account.We wrap by busting two expensive myths: industrial computers can be powerful, and edge AI does not automatically require a massive GPU. The goal is right-sizing hardware to the workload so performance is solid and ROI is real. Along the way we get into mentorship, imposter syndrome, and the value of being humble enough to ask questions, whether you are a new account executive or an engineer leaning on vendors and applications teams for expertise.Subscribe for more conversations on industrial automation, edge computing, OT hardware, and the people who make it work, then share this episode with a coworker and leave a review with your biggest edge deployment lesson.Support the show_________________________________________________________________
Almost 2% of U.S. GDP will be spent on AI infrastructure this year, nearly double 2025's figure. But beneath those headline numbers, the composition of that spending has quietly flipped: for the first time, dollars spent on running models in production now outweigh dollars spent training them. In this episode of TechSurge, host David Goldman speaks with Austin Lyons, a semiconductor analyst at Creative Strategies, co-host of the Semi Doped podcast, and author of the Chipstrat newsletter. Lyons previously worked as a hardware engineer at Intel and as a product manager on John Deere's autonomous tractor and Blue River Technology teams before turning to full-time chip industry analysis. The conversation opens with why AI buyers have moved from assembling commoditized parts to buying entire pre-integrated systems, tracing how Nvidia's rack-scale approach, exemplified by its 72-GPU Grace Blackwell racks, made turnkey deployment the default, and why that raises the bar for any chip startup trying to compete. Lyons and Goldman then unpack how inference workloads have split into two distinct problems, prefill and decode, and how that split created an opening for SRAM-based challengers to outperform general-purpose GPUs on decode speed. From there, the discussion turns to the rise of neoclouds, the GPU-rental companies that grew into public businesses worth well over $100 billion combined, and why so many traditional investors missed them. Lyons and Goldman work through the circular financing debate head-on: the mechanics of Nvidia's equity stakes, GPU-backed debt, and hyperscaler off-take agreements that critics compare to dot-com-era vendor financing, and the counterargument that demand is simply outrunning fixed supply. The episode closes on Lyons's own framework for identifying the next trillion-dollar chip company, built on four conditions including the ability to run trillion-parameter models at rack scale, beat an incumbent on a key performance metric, and land a frontier anchor customer, along with a look at how AI-assisted chip design is lowering the barrier for more companies, from OpenAI to electric vehicle makers, to design their own custom silicon. Sign up for our newsletter at techsurgepodcast.com for updates on upcoming TechSurge Live Summits and future episodes. Speaker Profiles and Links David Goldman: Partner, Celesta Capital Austin Lyons: Senior Analyst, Creative Strategies; Founder, Chipstrat; Co-host, Semi DopedLinkedIn: https://www.linkedin.com/in/austinlyons/Newsletter: https://www.chipstrat.com Further Reading and Resources Nvidia DGX GB Rack Scale Systems documentation: https://docs.nvidia.com/dgx/dgxgb200-user-guide/OpenAI and Broadcom – "OpenAI and Broadcom Unveil LLM-Optimized Inference Chip": https://openai.com/index/openai-broadcom-jalapeno-inference-chip/Chipstrat – Austin Lyons's newsletter: https://www.chipstrat.comSemi-doped: https://semidoped.com/ Timestamps 00:00 — No One's Brought a Chip to Market Built for LLMs01:21 — Introducing Austin Lyons02:16 — Why AI Buyers Now Buy Whole Systems, Not Parts08:18 — Nvidia's Margins and the Case for System Simplicity10:10 — Can a Startup Compete When You Have to Sell Systems?14:12 — Prefill vs. Decode: Splitting the Inference Workload24:51 — Fragmentation vs. Consolidation in AI Silicon28:22 — Why Investors Missed the First Wave of Neoclouds38:18 — The Circular Financing Debate48:29 — Lyons's Four Conditions for the Next Trillion-Dollar Chip Company About TechSurge:TechSurge Podcast shares the latest insights directly from legendary Silicon Valley leaders,daring new founders, and visionary technologists.Subscribe for weekly conversations into the intersection of technology advancement, market dynamics, and founder journeys.#AISilicon #LLMHardware #Nvidia #AIInference #TechPodcasts #AIInfrastructure
Hive Digital Technologies Chief Financial Officer Darcy Daubaras joined Steve Darling from Proactive to discuss the appointment of renowned capital markets executive Hubert Marleau as an Independent Director of HIVE's wholly owned subsidiary, BUZZ High Performance Computing (BUZZ HPC), a move designed to strengthen governance and strategic oversight as the company expands its sovereign AI infrastructure platform across Canada. Daubaras said Marleau's appointment brings an exceptional level of experience in capital markets, corporate governance and strategic growth at a pivotal time for BUZZ HPC. As demand for artificial intelligence infrastructure continues to accelerate globally, HIVE is positioning BUZZ HPC as a key provider of sovereign AI computing solutions, and management believes Marleau's expertise will help guide the business through its next phase of expansion. Marleau has spent more than five decades working across North American capital markets and is widely recognized as one of Canada's most experienced investment and governance professionals. His career spans investment banking, asset management, corporate finance and public company leadership, giving him a unique perspective on scaling businesses, raising capital and creating shareholder value. A co-founder of Palos Capital Corp. and Palos Management Inc., Marleau has held influential leadership positions throughout Canada's financial sector. His extensive resume includes serving as a Governor of both the Montreal and Vancouver stock exchanges, Chairman of the Toronto Stock Exchange Listing Committee and a director of the Investment Dealers Association of Canada, now known as IIROC. Over the course of his distinguished career, Marleau has served as a current or former director of more than 50 publicly traded companies in Canada and the United States. He has also played a key role in raising both public and private capital for hundreds of issuers and has advised on numerous mergers, acquisitions and financing transactions across a wide range of industries. Daubaras noted that Marleau's deep understanding of corporate governance and capital formation will be particularly valuable as BUZZ HPC continues to build out its AI-focused infrastructure platform and pursue opportunities in the rapidly evolving high-performance computing sector. The appointment comes as HIVE continues to diversify beyond its cryptocurrency mining roots and expand its presence in artificial intelligence and high-performance computing. Through BUZZ HPC, the company is developing GPU-powered infrastructure solutions designed to serve enterprise, government, research and sovereign AI customers seeking secure, scalable computing capacity. #proactiveinvestors #hivedigitaltechnologieslet #tsxv #hive #nasdaq #hive #darcydaubaras #ArtificialIntelligence #AIInfrastructure #HighPerformanceComputing #CapitalMarkets #CorporateGovernance #HubertMarleau #SovereignAI #GPUCloud #TechStocks #DigitalInfrastructure #CanadianTech #ProactiveInvestors #SteveDarling
How can businesses turn growing investment in AI infrastructure, cloud capacity and devices into outcomes that employees, customers and finance teams can actually measure? In this episode of Tech Talks Daily, I speak with Neil Sawyer, who manages HP's business across Europe, the Middle East and Africa. Neil works with companies across one of HP's largest global regions as they move from AI experimentation into wider deployment, making him well placed to discuss what happens when early enthusiasm encounters cost, security, governance and the realities of the workforce. We begin with the gap between building AI capacity and applying it to a business problem. Data centers, models and powerful devices provide options, but the investment only becomes useful when a company identifies the workflow it wants to improve. Neil argues that leaders should begin with the outcome, understand where AI can remove friction and decide how they will measure productivity, employee experience and business performance before buying another layer of technology. That raises a difficult question about productivity. If AI helps someone complete a task faster, does the organization use that saved time to improve the work, develop new ideas and give employees room to think, or does it simply add another task to the queue? I share my own experience as a business of one, where every efficiency gain has a habit of becoming extra output rather than a Wednesday afternoon at the cinema. Neil compares the current moment with earlier periods of industrial change and makes the case that automation should release people from repetitive administration so they can contribute creativity, judgment and higher value work. We also discuss why AI costs are becoming a boardroom issue. Token based services and agentic systems can produce growing and unpredictable bills as adoption spreads across a company. Neil explains why every workload does not need the same model or environment. Large language queries may benefit from cloud capacity, while sensitive data, company specific information and some recurring tasks may be better suited to local or on device processing. The decision affects cost, responsiveness, privacy, security, data sovereignty and environmental impact. Neil describes HP's view of hybrid AI, including devices with neural processing units and Z by HP Boost, which can connect available workstation GPU resources. He also explains why device refresh decisions should reflect workforce personas. A data scientist, account manager and office administrator may work for the same company, yet their computing needs can be very different. Mapping technology to the employee's role can help a business spend with greater discipline while giving people the performance they need. The conversation also covers governance and measurement. Informal use of public AI services can be difficult to see, assess or manage. Neil recommends giving employees an approved AI toolkit, using enterprise services that provide telemetry and examining how technology availability and performance affect the employee experience. Adoption figures can show that a tool is being used, but they do not prove that it is improving an outcome. We finish with a practical checklist for leaders. Define the outcome, identify the workflow, determine where each workload should run, calculate the cost, agree the measures of success and put clear controls around data, privacy and cybersecurity. That approach gives cloud and device based AI distinct jobs within the same business strategy. How is your organization deciding which AI workloads belong in the cloud, which should run closer to the employee, and whether the investment is producing measurable value? Share your thoughts with me.
We don't want the industry to fall apart, that's why we started this company.6 months ago, everyone was talking about model deployment. Now everyone has moved on to agentic AI deployment and autonomous agent swarms. But all of this is running on an infrastructure that itself is not automated.And this underlying infrastructure has massive GPU cluster systems. If even one GPU fails the whole training workload fails. You can't have 100,000 employees managing 100,000 GPUs. But LLMs are not the right solution for managing the infrastructure.Kandan and Randolph have spent decades operating large-scale telecom and cloud infrastructure, in a world where pagers used to run lives of people managing it. With MantisGrid they have come together to build 'The' model for AI infrastructure.Watch the episode to understand why making AI infrastructure autonomous is so important, why LLMs are not the answer, and what it takes to build the systems that will keep millions of agents running.00:00 — Invisible Infrastructure Behind AI01:05 — The Pager Nightmare04:01 — The GPU Problem Nobody's Solving05:38 — The Cloud Infrastructure Story07:53 — Why Just Watching Isn't Enough09:35 — Building the Waymo of AI Infrastructure15:00 — How an Autonomous AI Infra System Works20:35 — What Happens to DevOps?22:46 — Why LLMs Shouldn't Run Production29:14 — The Path to a Billion Dollars32:06 — The Problem Hyperscalers Can't Solve37:03 — Earning Trust to Run Infrastructure39:55 — When Agents Start Managing Agents-------------India's talent has built the world's tech—now it's time to lead it.This mission goes beyond startups. It's about shifting the center of gravity in global tech to include the brilliance rising from India.What is Neon Fund?We invest in seed and early-stage founders from India and the diaspora building world-class Enterprise AI companies. We bring capital, conviction, and a community that's done it before.Subscribe for real founder stories, investor perspectives, economist breakdowns, and a behind-the-scenes look at how we're doing it all at Neon.-------------Check us out on:Website: https://neon.fund/Instagram: https://www.instagram.com/theneonshoww/LinkedIn: https://www.linkedin.com/company/beneon/Twitter: https://x.com/TheNeonShowwConnect with Siddhartha on:LinkedIn: https://www.linkedin.com/in/siddharthaahluwalia/Twitter: https://x.com/siddharthaa7-------------This video is for informational purposes only. The views expressed are those of the individuals quoted and do not constitute professional advice.Send us Fan Mail
Shadow permet aujourd'hui d'utiliser un PC puissant à distance depuis presque n'importe quel appareil, du vieux laptop au smartphone. L'entreprise veut aussi valoriser son infrastructure GPU pour l'inférence IA lorsque les joueurs ne l'utilisent pas.
毎日届くニュースレターの購読はこちら https://wbnlive.substack.com/9月14日、ロサンゼルスのAll-In Summitで登壇中だったNvidia CEOのジェンスン・フアン氏に、トランプ大統領から電話が入りました。スピーカーで会場に流された通話で大統領は「ロボットは乗っ取らない、全部デマだ」「減速を求めるのは政治的な連中か中国かもしれないが、そうはさせない」と語り、フアン氏は「そうはさせません」と応じて会場は拍手。この動画では、その一本の電話と、同じ日に動いた株価、そしてAnthropicの言っていることとやっていることのズレまでを金城が話しています。▼ 目次00:00 All-Inサミット。4人に囲まれたジェンスン・フアン01:30 登壇中、トランプ大統領から電話が来た02:20 ロボットは乗っ取らない、全部デマだ02:45 フアンが減速を拒む理由。GPUとオープンソース03:32 データセンターは今後20年から25年の石油03:47 出来過ぎている。仕組まれていたのか04:23 そうはさせません。会場からの拍手05:20 世界が終わるという予測に科学的根拠はない06:09 必要なガードレールは高IQの大統領だけ07:54 All-Inの4人「黙って仕事に集中しろ」08:33 株価の答え。NVIDIAは3.4%安、半導体が売られた10:04 ソフトバンクGは10.7%安。OpenAI上場は来年へ10:33 買われたのはセキュリティ銘柄だった11:49 言行不一致。6年2.19兆円の計算リソース契約13:25 大手がAnthropicの利用を制限し始めた※ Anthropicの計算契約と大手企業の利用制限はThe Informationの報道に基づきます。電話が事前に仕組まれていたかどうかは番組内の推測で、確認された事実ではありません。数字は放送時点の公開情報・報道に基づきます。▼ 出演Kinjo X: https://x.com/illshinおいなり(Cohost / Startup Now MC・JobTales Inc. CEO) X: https://x.com/oinariiisan▼ 主要ソース(一次情報)All-Inホストの反応(X): https://x.com/theallinpod/status/2099640047473721650TechCrunch(フアン「そうはさせない」): https://techcrunch.com/2026/09/14/nvidia-ceo-jensen-huang-tells-trump-were-not-going-to-let-an-ai-slowdown-happen/CNBC(電話の詳細): https://www.cnbc.com/2026/09/14/trump-phones-nvidia-huang-all-in-calls-data-center-opposition-hoax.htmlThe Information(Rum Groupとの契約): https://www.theinformation.com/articles/anthropic-strikes-13-7-billion-compute-deal-trump-linked-rum-groupThe Information(大手の利用制限): https://www.theinformation.com/articles/anthropic-data-fears-prompt-nvidia-palantir-booz-allen-restrict-model-useWBNは平日毎日12:00から、AI・テック・スタートアップのニュースをライブでお届けしています。#AI #AIニュース #Nvidia #トランプ #Anthropic #AI規制 #WBN
Bij de AIVD kom je niet zomaar binnen. Je legt je telefoon en je spullen weg, je gaat door de security check en de poortjes, en pas daarna ga je aan het werk. De netwerken daarachter zijn net zo streng afgesloten: de meeste hebben geen internet, ze leven helemaal los van elkaar, en data kan er maar één kant op.In deze aflevering van De Nederlandse Kubernetes Podcast spreken Ronald Kers en Jan Stomphorst met Wouter en Nick, die bij de AIVD in het platformteam werken.Ze vertellen hoe updates via data diodes naar binnen komen, waarbij data maar één kant op kan en niet terug. Packages en images die ontwikkelaars nodig hebben moeten offline beschikbaar zijn, zodat een ontwikkelaar binnen dezelfde ervaring heeft als iemand die wel internet heeft. Wat binnenkomt wordt eerst gecontroleerd voordat het in de repository terechtkomt.Verder in het gesprek: waarom ze werken met één groot multitenant cluster dat elk kwartaal wordt geüpdatet, waarom vuile data daar bewust buiten blijft, hoe ze Cluster API en de bijbehorende image builder gebruiken om eigen images te bouwen volgens eigen regels, en waarom hun filosofie is om zoveel mogelijk open source te gebruiken om vendor lock-in te voorkomen.Ook aan bod: hoe applicaties organisch redundant zijn geworden doordat nodes gewoon 's nachts worden gerecycled, de oriëntatiefase rond Gateway API, GPU's in de eigen omgeving, de samenwerking met de MIVD en JIVC, en de grootste uitdaging voor de toekomst, namelijk dat steeds meer producten alleen nog als SaaS beschikbaar zijn.https://werkenbijdeaivd.nl/ De Nederlandse Kubernetes Podcast is een initiatief van ACC ICTStuur ons een bericht.ACC ICT Specialist in IT-CONTINUÏTEIT Bedrijfskritische applicaties én data veilig beschikbaar, onafhankelijk van derden, altijd en overal Dutch Cloud Native Day 2026Two days of cloud native talks, workshops and community in Utrecht, exploring how AI is changing the way we build, run and scale modern platforms.Like and subscribe! It helps out a lot.You can also find us on:De Nederlandse Kubernetes Podcast - YouTubeNederlandse Kubernetes Podcast (@k8spodcast.nl) | TikTokDe Nederlandse Kubernetes PodcastWhere can you meet us:EventsThis Podcast is powered by:ACC ICT - IT-Continuïteit voor Bedrijfskritische Applicaties | ACC ICT
*본 영상은 인천상수도사업본부 유료 협찬 광고를 포함하고 있습니다. ??LIVE [한판승부] 용혜인 사퇴 뒷이야기! 추미애 직격한 김민석? ◎ 1부 (18:00-19:00) [영희와 광재]1. 용혜인, 성평등가족부 장관 후보 자진사퇴...2기 첫 낙마2. 이번 주 청문회 7명 '슈퍼위크'...내일부터 본격화3. 김성수 대법관 후보 '처남 찬스'...증여세 납부 진행4. 李 지지율 33.8%, 9주 연속 하락...부정평가 첫 60% 돌파5. 김민석 "노무현 탄핵 전조와 비슷"...추미애에 설명 요구- 노영희 변호사- 정광재 동연정치연구소장- 곽우신 오마이뉴스 기자 ◎ 2부 (19:00-19:30) [한판인터뷰]1. 오픈AI 아스트라, 어떤 모델이길래?2. "AGI 시대 열렸다"…그 의미는3. GPU 10만 장 투입…K반도체엔 호재?4. 아스트라 '너무 뛰어난 성능'? 과장인가 특이점인가5. 빅테크 수장들 "개발 속도 늦추자"...왜?- 정주용 그래비티벤처스 의장See omnystudio.com/listener for privacy information.
Samsung y SK Hynix rechazan adelantar cinco años de electricidad para financiar la red de sus fábricas de chips. También hablamos del contrato de servicios GPU de 13.700 millones de dólares que The Information atribuye a Anthropic y de un modelo abierto de NASA e IBM para localizar hielo, cráteres y volcanismo en la Luna.Puedes seguirnos en YouTube en https://youtube.com/olivernabani y puedes unirte al Discord Mashain en https://olivernabani.com/discord
Hive Digital Technologies Executive Chairman Frank Holmes Steve Darling from Proactive to discuss a major leadership appointment and a significant operational milestone as the company continues to expand its artificial intelligence and high-performance computing business. HIVE announced that its wholly owned subsidiary, BUZZ High Performance Computing (BUZZ HPC), has appointed Mark Volk as Senior Vice President, Revenue. The move is designed to accelerate commercialization efforts and strengthen the company's position in the rapidly growing AI infrastructure market. Volk brings nearly three decades of experience building and scaling businesses across artificial intelligence, high-performance computing, cloud infrastructure and enterprise technology. In his new role, he will lead BUZZ HPC's global go-to-market strategy across enterprise, government, sovereign AI, hyperscale, research and academic sectors. His primary mandate will be to build and scale the company's revenue engine globally while expanding adoption of BUZZ HPC's AI-focused infrastructure offerings. Alongside the executive appointment, HIVE also provided an update on its operating performance, announcing that the company has now exceeded $1 million in average daily revenue from its combined Bitcoin mining and GPU cloud computing operations. Since August 21, 2026, HIVE has produced an average of approximately 12 Bitcoin per day, representing roughly 2% of global Bitcoin network production. At the same time, the company's GPU cloud business has generated approximately $100,000 in average daily revenue, reflecting growing demand for AI computing capacity. Based on prevailing Bitcoin prices, network difficulty levels and current operating conditions, the combined businesses are now generating more than $1 million in revenue per day, highlighting the benefits of HIVE's strategy to diversify revenue streams through both digital asset production and AI infrastructure services. As the company continues to scale BUZZ HPC and broaden its AI offerings, management believes the combination of Bitcoin mining operations and GPU cloud services provides multiple avenues for growth while positioning HIVE to participate in two of the fastest-growing sectors in technology. #proactiveinvestors #hivedigitaltechnologieslet #tsxv #hive #nasdaq #hive #BUZZHPC #ArtificialIntelligence #AIInfrastructure #GPUCloud #GPUaaS #HighPerformanceComputing #BitcoinMining #DigitalInfrastructure #CloudComputing #DataCenters #TechStocks #CryptoMining #AIInnovation #TechnologyGrowth
被Stripe斥资70~80亿美元收购的OpenRouter是一个由来自Web3的团队创建的重要的AI基础设施。我们曾经在E78期节目讨论过这家公司的来时路。 本期节目试着转换一个不同于传统AI基础设施分析的视角:如果不把OpenRouter看成一个API聚合器,而是把它看成一个正在形成中的AI推理开放市场,会看到什么不同的东西? 本期嘉宾Danning长期从事链上数据、DEX(去中心化交易所)、MEV(链上最大可获取价值)和市场结构研究,她从Web3世界的交易流、做市、路由、竞价和金融市场博弈经验出发,重新拆解了从GPU、算力、推理服务到Token的完整价值链。这场对话逐渐指向一个更大的问题:AI算力是否正在从基础设施,变成一种可以交易、定价、套利、对冲,甚至金融化的生产资料? 【主播】 刘锋,BODL Ventures合伙人,前链闻总编辑 【嘉宾】 Danning Sui,Pantera Capital研究总监 【你将听到】 02:57 别把OpenRouter当成单纯的API聚合器,这是动态的算力交易市场 06:30 OpenRouter把模型和算力解耦了 07:49 Token正在成为算力的标准化商品,竞价市场出现 08:50 为什么OpenRouter的数据非常重要:它在提供罕见的AI推理价格发现机制 12:00 算力和推理,其实是两种不同的商品 13:14 拆解AI 推理背后的完整供应链 14:52 纯粹的算力中间商可能越来越难赚钱 19:47 闭源模型利润正在被压低 21:35 OpenRouter上的市场博弈和诱导报价(像极了DEX上的做市商诱导报价伎俩) 32:48 如果卖Token这么赚钱,英伟达为什么不自己卖? 35:36 AI推理市场价值链里的利润正在动态迁移 40:06 算力正在被逐层商品化 43:23 为什么GPU并不像石油? 46:38 算力金融化:期货、杠杆与交易市场 52:53 重新回答:OpenRouter到底值在哪里? 55:19 去中心化AI训练可能没有想象中容易 55:59 值得进一步关注的数据和研究:Tarun的研究文章、SemiAnalysis、Architect、OpenRouter团队发布的数据 【背景信息】 Tarun Chitra 撰写的研究文章: The Souq: Value Creation and Extraction in OpenRouter (大市集:OpenRouter上的价值创造与提取) AI 推理产业链全景图 【词汇表】 本期播客提及的Web3/AI领域产品、公司与专有名词 ConsenSys 0X Flashbots The Merge Palantir/Alex Karp Together AI Fireworks AI RFQ(Request for Quote,询价机制) Wintermute Uniswap PancakeSwap VAST.ai SF Compute Architect Nebius 【后期】 AMEI 【运营】 朱婕 【BGM】 Mumbai - Ooyy 【在这里找到我们】 收听渠道:苹果|小宇宙 海外用户:Apple Podcast|Spotify|Google Podcast|Amazon Music 联系我们:podcast@sv101.netSpecial Guest: Danning Sui.
Christian Ondaatje is originally from Los Angelas, CA. He got into GPUs as an underclassmen in college, of course through gaming and VR. He ended up starting an eGPU company in his CS program @ Harvard, though it was shut down through some cease and desist from Apple. Though he learned some hard startup lessons, his love for GPUs was only fueled more. Outside of tech, he does a lot of SIM racing. In fact, he has built his own Linus VR racing simulator rig over a couple of years He finds that the formal, online competition is very fun to participate in. As I mentioned, Christian was way into GPUs throughout many different experiences in his life. He and his team recognized that while demand door AI inference and compute was at an all-time high, setting up bare-metal GPU infra was still complex and slow... so they set out to change it. This is the creation story of Aranya.Linkshttps://aranya.tech/https://www.linkedin.com/in/christianondaatje/ Current Sponsors: Tiger Data Protected Harbor Render Fitnexa Perplexity Entelligence Checkout our Stacklist! https://stacks.codestory.co/ Hosted by Noah Labhart | Technical Founder & Startup Mentor. Our Sponsors:* Check out Granola and use my code granola.ai/CODESTORY for a great deal: https://granola.ai* Check out Perplexity and use my code CODESTORY for a great deal: https://www.perplexity.aiAdvertising Inquiries: https://redcircle.com/brandsPrivacy & Opt-Out: https://redcircle.com/privacy
Originally aired on September 1st.Timestamps:00:00 Intro00:39 Patreon02:53 Food with Josh06:17 News begins - and Packard Bell is back10:03 Intel LGA1954 confirmed12:49 MacBook Neo SSD thrashing seems dire16:30 Horrifying DLSS 5 mods22:29 Rumored next-gen Radeon flagship performance25:19 Run multiple Linux kernels with no hypervisor?27:59 Tweaks for Windows 11 before Microsoft shuts them down31:29 GPU coil whine database35:21 (In)Security Corner48:51 Gaming Quick Hits54:07 Fractal Refine 2 gaming chair review1:00:08 Picks of the Week1:08:35 Outro ★ Support this podcast on Patreon ★
“We do it with AI” is easy to say. Making robots work through contact, uncertainty, and variation on a real factory floor is the hard part. We sit down with Whitney Dills from Thought Forge AI to unpack an approach to industrial robotics control that flips a common assumption: more data does not automatically mean better performance.Whitney explains active inference, a neuroscience rooted framework that treats robot control as continuous prediction and correction using real time sensory feedback. Instead of bloated backpropagation models that demand huge datasets and retraining for every edge case, Thought Forge trains quickly on small amounts of data and adapts after deployment. We talk about why that matters for insertion and complex manipulation, where torque, joint state, and haptics decide success or failure. If you have ever watched a slick picking demo fall apart the moment parts vary, lighting shifts, or fixtures drift, this conversation connects the dots.We also get practical about adoption: edge AI that can run on a CPU, the risks of GPU dependence, and how data ownership and IP concerns shape buyer trust. Whitney shares why their models stay on premise, why off the shelf robots and sensors often win over novel vertically integrated stacks, and how to think about cobots versus humanoids when reliability and cycle time are non negotiable.Subscribe for more conversations that cut through robotics hype, share this with someone evaluating “physical AI” for manufacturing, and leave a review with the one automation task you most want AI to finally handle.Support the show_________________________________________________________________
Big Tech is pouring more than $1.4 trillion into AI, prompting investors to ask: Is it worth it? Our U.S. Internet analyst Brian Nowak looks at three business models that could earn 25 to 50 percent returns for Gen-AI-enabled technologies.Read more insights from Morgan Stanley.----- Transcript -----Brian Nowak: Welcome to Thoughts on the Market. I'm Brian Nowak, Morgan Stanley's U.S. Internet analyst.Today, can the enormous investment behind Gen AI actually pay off?It's Wednesday, September 9th, at 9am in New York.AI has moved quickly into everyday life. It helps people write software, research purchases, automate work, find information, among myriad[s] of other use cases.But we need an infrastructure build-out of extraordinary scale to support all of this activity and more activity to come.In all, we estimate that the major cloud providers are going to spend more than $1.4 trillion on this AI build-out next year alone. But compute capacity is potentially going to quadruple from 2025 to 2028, reaching roughly 120 gigawatts.But all of the spending has raised a lot of questions for investors. One of the most common questions is: What kind of return on invested capital can these companies earn from all of these trillions of dollars of data center infrastructure investment?Well, our bottom-up work points to encouraging answers to this question.We see paths to roughly 25 to 50 percent return on invested capital, or ROIC, across three emerging AI business models. Now, ROIC is a useful way of measuring whether investments pay off. Think of it as how much after-tax operating profit can be generated relative to the capital required in the first place.The first business model we've analyzed is renting compute power. This is the infrastructure layer of the AI economy. Cloud providers build data centers filled with advanced graphics processing units, or GPUs, and rent that compute capacity to customers. In our base case, a large next-generation data center can generate a return on invested capital of roughly 30 percent simply renting AI compute power.And even if rental prices move around, our scenarios still produce returns ranging from low 20s percent to nearly 40 percent. So, despite the enormous cost of building and capital being deployed for these facilities, we think the economics here are quite attractive.The second business model we've analyzed is where an AI lab has their own model, and they also own their own infrastructure. They give access to their model through an API to consumers and enterprises who then build upon it, they utilize the model. In some cases, they build applications using that model that can be future sources of productivity or efficiency for the economy.In this scenario, we think the economics can be even stronger. When the model developer owns their own underlying infrastructure, our base case generates a roughly 75 percent incremental operating margin and a return on invested capital of 40 percent plus.These returns on invested capital are impressive, but what determines whether these returns can actually materialize?Well, two things matter a lot. The first is the price the developers are able to charge for tokens, which are the units of information that AI models process. The second factor that matters considerably is how efficient[ly] can this infrastructure process these tokens.This is why continued improvements in chips and software to drive higher token throughput – or more tokens per GPU per second – are critical to the long-term unit economics across this AI ecosystem.The third model we've analyzed is when the AI developers rent their compute infrastructure rather than owning it. So, effectively, they are paying someone else for the data centers and the GPUs that they need. While this lowers their returns on invested capital because another provider takes a piece of the unit economics, our base case still produces roughly a 30 percent incremental operating margin and 25 percent post-tax return potential.So, while the AI build-out requires enormous investment, the size of the spending alone doesn't tell the whole story about whether or not there are economic returns to come.What ultimately matters is the revenue and profit that the infrastructure can generate. And as more of the infrastructure shifts from training AI models to serving customers through emerging products and inference, we think we're going to get a much clearer answer to this question investors are asking today.Was all this spending worth it? Our research suggests: Yes.Thanks for listening. If you enjoy the show, please leave a review wherever you listen and share Thoughts on the Market with a friend or colleague today.
AI is turning compute into a strategic resource, and the scramble to secure GPU capacity is starting to look a lot more like a commodity market than a traditional cloud-services business. As data-centre developers, lenders and energy companies try to price the next wave of AI demand, a new question is coming into focus: Can the industry build the kind of benchmark and hedging tools that already exist for oil, gas and power? Host Ed Crooks is joined by Peter Keavey, Global Head of Energy and Environmental Products at CME Group, and Carmen Li, Founder and CEO of Silicon Data. Together, they explore the case for a futures market in GPU compute: a financial product designed to bring more transparency, liquidity and risk management to one of the fastest-growing corners of the AI economy.Carmen explains how the market works today. Most users are not buying chips outright; they are renting access to GPU capacity by the hour, often through longer-term agreements with hyperscalers, neo-cloud providers and data-centre operators. That market is already large, global and increasingly active, but it remains fragmented and opaque, with prices varying by provider, chip type and contract structure, and much of the trading still happening through bilateral deals and requests for quotes.Peter sets out the logic for moving from that over-the-counter world to an exchange-traded one. In his view, a GPU futures contract could do three things at once: reduce counterparty risk through central clearing, concentrate liquidity in a transparent order book, and create forward benchmark prices the wider market can use. The proposed product is financially settled against an index of spot prices, translating an hourly rental market into a standardised monthly contract that could eventually extend several years forward.The bigger issue, though, is energy. Power is not the whole cost of GPU compute, but it is the most volatile variable input, which means a GPU hedge could eventually sit alongside gas and power hedges for data-centre operators, lenders and infrastructure investors. The discussion keeps returning to what that means for markets such as Texas and Virginia, where the AI build-out is already shaping decisions on generation, grid access and where capital should go next.Both guests stress that this is still a young market, but already a volatile one. Rental rates have swung sharply as chip scarcity eases and then tightens again, while banks, traders and developers are trying to make long-dated decisions without a reliable forward curve. If this market develops the way Keavey and Li expect, GPU futures would not just serve traders: they could become an important signal for anyone trying to judge how durable the AI boom really is, and how much energy the system will need to support it.See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
Join The Full Nerd gang as they offer level-headed takes about the latest PC building news. In this episode the gang looks over the latest Intel CPU rumors/leaks, our impressions of the DLSS 5 launch, and much more. And of course we answer questions live! Timecodes: 00:00:00 - Intro 00:03:04 - DLSS 5 launches 00:33:07 - Nova Lake leaks 00:56:39 - Q&A 01:22:55 - What's next? Links: - DLSS 5 @DigitalFoundry review: https://www.youtube.com/watch?v=XAt63dB1hdI - DLSS 5 mod on 2nd GPU: https://www.digitalfoundry.net/news/2026/09/dlss-5-mod-shifts-load-to-second-gpu-doubling-frame-rates-and-latency - Nova Lake leaks: https://hothardware.com/news/intel-nova-lake-slide-leak-release-cadence Join the PC related discussions and ask us questions on Discord: https://discord.gg/UWhjwg778a Follow the crew on X and Bluesky: @AdamPMurray @BradChacos @MorphingBall Music by Our Ghosts: https://ourghosts.bandcamp.com/ Some links may contain affiliate links, which means if you buy something PCWorld may receive a small commission. ============= PCWorld Website: http://www.pcworld.com Newsletter: http://www.pcworld.com/newsletters ============= Learn more about your ad choices. Visit megaphone.fm/adchoices
Are soaring AI compute costs slowing down your startup scaling efforts? Join us as we sit down with Junchen Jiang, CEO of TensorMesh, to reveal how his groundbreaking KV cache optimization can slash your AI inference expenses by up to 80 percent. In this highly technical and insightful episode of Born in Silicon Valley, we explore the rapidly evolving landscape of self-hosted large language models. Junchen explains the mechanics behind KV cache, why traditional GPU memory limitations bottleneck AI performance, and how offloading this data can dramatically improve efficiency for enterprise applications, coding agents, and chatbots. We also dive into Junchen's personal journey transitioning from leading a PhD research group at the University of Chicago to building a venture-backed startup in the Bay Area. He shares valuable leadership lessons, the differences between motivating academic researchers versus industry engineers, and what it takes to build a world-class team in today's competitive tech ecosystem. Whether you are a tech founder, a forward-deployed engineer, or an enterprise leader looking to optimize your AI infrastructure, this conversation is packed with actionable insights on building the next big data layer for artificial intelligence. Chapters 00:00 Introduction to Junchen Jiang and TensorMesh 02:00 The origin story of KV cache technology 10:07 How KV cache reduces inference costs by 10x 14:54 Building industry-ready systems from research 19:59 Open source community and industry adoption 25:11 Business model and monetization strategies 29:47 Team building and leadership lessons 40:06 Market opportunities and customer segments 44:55 Future vision for AI data layers Host: Jake Aaron Villarreal leads the top AI recruitment firm in Silicon Valley, www.matchrelevant.com, uncovering stories of funded startups and going behind the scenes to tell their founders' journeys. If you are growing an AI startup or have a great story to tell, email us at: jake.villarreal@matchrelevant.com
This episode is a compilation of answers to YOUR questions that were asked directly from my listeners who attend my weekly business education YouTube live webcast. I'll be covering the topic on: the $10 Trillion Question: Nvidia, AI Bubbles, and the 30-Year Yieldand more. Refer to chapter marks below for a complete list of topics covered and to jump to a specific section. Get mentored by Chris: Book a Zoom call to discuss joining my Business Academy, Finance Bootcamp (to get a job in finance) or MBA Degree Programs or for investing/business/personal development coaching: https://haroun.short.gy/1on1CallYTWDownload my free "Networking eBook": www.harouneducation.comAttend my weekly YouTube Live every Thursday's 8am-11am PT. Subscribe to my YouTube Channel to receive notifications. Chapter Marks: 0:25 Welcome to the 376th Weekly Live Webcast of August 20, 2026!0:55 Free Live Class: Use AI to Beat the Resume Robots (Live Demo)4:08 Are you a believer in Sturgeon's Law and does it apply to the financial world?5:34 Most applicable instance of Murphy's Law?8:18 Where does the Iran War go from here?9:38 Are you nervous about the 30-year yield?13:23 How does Japan survive with unsustainable debt?14:58 Is America tarnishing its reputation?16:33 Will Nvidia become the world's first $10 Trillion dollar company or will the AI bubble burst?17:31 Will AI lead to more social democratic values?20:18 Do the socialist democratic members in the Senate bother you?21:35 What is the future of Indian IT and IT in general?23:11 Will India become a one-party socialistic autocratic regime?23:41 Thoughts on economic trading blocs?25:13 What happens to the world's oil reserves from here?26:23 How will open-source LLM models impact GPU and TPU sales?27:55 Your best and worst AI predictions?29:33 Can you make money long-term as a day or swing trader?32:33 Thoughts on different finance certifications like CFA, CFP and CMA? 34:31 How to break into the high-paying finance world?36:43 Is the states better for a finance job or stay in Canada?37:29 Is it possible to get a 50% return in today's market? 40:16 How to stay motivated in this job market? 41:35 Will the AI Resume course be recorded? 42:15 Is capitalism dead in the US? 47:20 Should I still learn JavaScript? 48:04 Will non-human traffic surpass human traffic? 49:24 Will far right of Europe pull out of EU? 50:03 Are you a console or PC guy? 51:09 Thoughts on Nvidia investing in OpenAI? 54:58 How can we best incentivize value creation while not rewarding self-interest and greed? 56:29 Will the next administration bring down debt levels? 57:16 Are AI agents a cybersecurity nightmare? 57:37 Should we still learn Python? 59:00 How to avoid AI slop? 1:01:19 Can Zoom be disrupted and how? 1:02:22 What's the connection between Japan's Yena and the USD? 1:05:09 What is the best economic system? 1:08:06 Did you hear the GTA map got leaked? 1:08:31 What's your view on Elon Musk wanting more robots than humans? 1:12:28 Who is buying SpaceX right now? 1:12:58 What do you think about AI watermarks? 1:16:15 What is Palantir's business model? 1:16:49 Best learning strategies to stay up to date on? 1:18:19 Are we headed towards a crash? 1:22:13 Do you think self-interest and greed is more of a socialist and communist issue? 1:23:07 Will AI worsen mental health? 1:24:43 Why did OpenAI close Sora? 1:26:14 How will China impact the world once it becomes hegemony? 1:27:58 Can you share some details on the Japanese culture? 1:32:13 Is a Master's degree worth it? 1:32:50 Does Moderna really have the cure for cancer? 1:33:43 Difference between an options market maker and a hedge fund? 1:35:35 Difference between nominal GDP and PPP 1:39:22 How important is listening to earnings calls? 1:42:35 Thoughts on Chinese brain drain to the US? Connect with me: Schedule a 1:1 call with Chris: https://haroun.short.gy/1on1CallYTWYouTube: ChrisHarounVenturesCompleteBusinessEducationInstagram @chrisharounLinkedIn: Chris HarounTwitter: @chris_harounFacebook: Haroun Education Ventures TikTok: @chrisharoun
In this episode, Ray Cochrane breaks down Hot Chips 2026, the engineering conference where IBM, NVIDIA, Intel, AMD, Arm, and Fujitsu all showed how their next processors actually work. The headline disclosure is a mainframe core that runs Arm natively. Ray also covers Apple’s odd M6 Mac mini naming, London’s first autonomous Uber rides, Amazon’s purchase of the company behind DuckDB, GitHub’s HydraFusion, the best of IFA 2026, and new USDA research on farmed salmon. – Want to start a podcast? It’s easy to get started! Sign-up at Blubrry – Thinking of buying a Starlink? Use my link to support the show. Subscribe to the Newsletter. Email Ray if you want to get in touch! Like and Follow Geek News Central’s Facebook Page. Support my Show Sponsor: Best Godaddy Promo Codes Get 1Password Full Summary Cochrane opens the show with a personal update. He apologizes for the late rollout and the missed Monday episode, having spent the week fighting a cold. He’s also heading to Michigan to spend time with family and visit his father’s gravesite. He hopes to record a couple of shows from his dad’s old studio while he is there, including a special episode planned for Tuesday. He also points listeners to a fresh site redesign that trades the old techie look for something cleaner and friendlier. The featured segment starts from a wrap-up post on Arm’s newsroom. However, Cochrane broadens it to cover the whole conference rather than a single article. Hot Chips has run every August since 1989, and this year’s event was the 38th, held August 23rd through the 25th at Stanford’s Memorial Auditorium. What Hot Chips Actually Is Cochrane draws a line between Hot Chips and the big consumer trade shows. CES and Computex exist for product announcements and marketing. Meanwhile, Hot Chips is an IEEE engineering conference where chip architects present block diagrams and die photos for thirty minutes at a stretch. The audience matters as much as the content. Roughly five hundred people who design chips for a living fill the room, and they would spot a fudged number immediately. No written paper is required, just the talk and the slides. In-person tickets sold out this year, as did Stanford’s dorm housing. Cochrane says he plans to cover the conference annually going forward. IBM Built a Mainframe Core That Speaks Arm The disclosure that stopped Cochrane cold came from IBM. Its next processor for IBM Z and LinuxONE runs two completely different instruction sets natively, on every one of its eleven cores. Those are z/Architecture, IBM’s own mainframe language, and AArch64, which is 64-bit Arm. Crucially, this is not emulation. Nor is it Arm cores glued onto the die beside the mainframe cores. IBM built 2,792 Arm instructions directly into the hardware, which it says is more than double the mainframe instruction count. Each core carries two separate decoders while sharing the caches, branch predictor and register files downstream. It switches between the two in nanoseconds. IBM even added dedicated hardware to flip byte order, because Arm and the mainframe store numbers in opposite directions. The payoff is Arm SystemReady compliance, meaning off-the-shelf Arm Linux runs on a mainframe unmodified. Patrick Kennedy of ServeTheHome, who was in the room, wrote: “I am sitting here still in awe of what IBM is doing here; this is not Z plus Arm cores, this is Z and Arm in one core.” Cochrane flags one precision point that is easy to get backward. IBM did not license Arm’s core designs and drop them in. Instead, it took its own mainframe core and taught it AArch64 under an architecture license, which is considerably harder engineering. The specifications are striking. The chip uses a 2nm process, with eleven cores running above 5.7GHz sustained and no turbo mode at all. Each core gets 36MB of L2 cache, backed by a 3.5GB virtual L4 pool. Furthermore, the reliability target is eight nines, which works out to roughly three tenths of one second of unplanned downtime per year. IBM gave it no name and no ship date, though the press expects “Telum III” around 2028. Arm, Fujitsu and NVIDIA Show Their Hands Arm itself had plenty to discuss, starting with an unfortunate name. Its first chip in thirty-five years is called the AGI CPU, which is a product name rather than any claim about artificial general intelligence. For three and a half decades, Arm designed processor blueprints and licensed them out, collecting royalties without competing. That era is now over. The AGI CPU is Arm’s own silicon, co-designed with lead customer Meta, running up to 136 cores on TSMC’s 3nm process at 300 watts. Arm’s CEO says the company has more than $2 billion in customer demand across the next two fiscal years. Fujitsu brought the detail Cochrane called the coolest of the conference. Its MONAKA chip packs 144 Arm-based cores, but the trick is the cache. Rather than sitting alongside the cores and eating die area, the entire last-level cache lives on a separate 5nm die with the 2nm compute die stacked directly on top. It ships in 2027 in 350-watt and 500-watt versions. NVIDIA had more stage time than anyone, with six sessions. Its new Vera CPU carries 88 cores of NVIDIA’s own Olympus design, which marks a change: the previous Grace CPU used Arm’s off-the-shelf cores. Consequently, Arm’s win here is the instruction set, not the blueprint. The memory disclosure drew the most attention, with a fully loaded system reaching 1.5TB at 1.2TB/s while the whole memory subsystem draws just 30 to 40 watts. The Caveat on NVIDIA’s Benchmark Slides Cochrane pushes back on how NVIDIA presented its numbers. On the standard SPEC integer benchmark, Vera scored 925 against AMD’s 128-core EPYC score of 898, about three percent ahead. However, the slide NVIDIA showed normalizes that same result per physical core, which makes a three percent gap look enormous. NVIDIA defends the choice, arguing that per-core throughput matters when thousands of AI agents run at once. Cochrane grants that it is a fair argument to make. Even so, his verdict is blunt: it is a different number from the headline one, and presenting it that way is not the best look. Sponsor: GoDaddy Economy hosting $6.99/month, WordPress hosting $12.99/month, domains $11.99. Website builder trial available. Use codes at geeknewscentral.com/godaddy to support the show. Apple’s Newest Chip Landed in Its Cheapest Mac Last week’s episode covered the Mac Studio half of Apple’s August announcement. Tonight Cochrane takes the other half, the new Mac mini, and finds the numbering genuinely strange. The $899 base Mac mini gets the M6, which is Apple’s first 2nm chip and the newest process the company has shipped. It brings twelve CPU cores, twelve GPU cores, and 170GB/s of memory bandwidth. Apple also introduced a third CPU core class called super cores. So Apple’s most advanced chip sits in its cheapest desktop, while the M5 Pro, M5 Max and M5 Ultra above it all carry a lower number. It gets stranger. The M6 mini has Thunderbolt 4 while the pricier M5 Pro mini has Thunderbolt 5, and the memory ceiling runs backward too. Apple explains none of it across three announcement pages. Reading between the lines, Cochrane figures the M6 is the entry point of a new generation that shipped ahead of its larger siblings. London’s Robotaxis Started Carrying Passengers Arm’s monthly roundup covers everything outside the data center, and Wayve stood out. Transport for London granted the British self-driving company private hire vehicle licenses on August 5th, the same category a minicab needs. Then on September 3rd the service launched with Uber, marking the first autonomous rides ever offered to UK passengers. A small fleet of Ford Mustang Mach-Es covers anywhere in London except the airports, with a TfL-licensed safety driver still aboard. Over 140,000 Londoners signed up. The compute runs on NVIDIA’s Arm-based automotive platform, which is why Arm claims the win. Cochrane notes the pattern: once you set a standard nobody can move off, these wins keep arriving. Two more items round it out, both about squeezing AI onto phones. Google’s Pixel 11 shipped in August with the Tensor G6, and Google claims on-device AI runs up to 3.5x faster using 3.5x less energy. Those are Google’s own unbenchmarked figures. Separately, Graphcore’s research team worked with Arm to run an 11-billion-parameter vision model on phone-class processors by squeezing each parameter to 2.7 bits, taking the model from roughly 22GB down to 3.7GB. Intel Is Pitching AI Infrastructure From a Long Way Back Intel’s newsroom post previews the AI Infra Summit, running September 15th through 17th in Santa Clara. CEO Lip-Bu Tan takes a fireside chat on Tuesday morning, with three Intel sessions across the show. Cochrane unpacks two terms first. Physical AI means AI that acts in the real world through sensors and motors rather than living on a screen, and Intel takes it seriously enough to have renamed its PC division the Client Computing and Physical AI Group in May. Disaggregated inference splits the two phases of running a model: reading your prompt is compute-hungry, while writing the answer back is memory-hungry. The context is where this gets interesting. Intel’s revenue rose 25 percent last quarter, its fastest growth since 2011. Nevertheless, in AI accelerators the company barely registers next to NVIDIA. Gaudi has effectively been abandoned, to the point that Intel stopped maintaining its open-source driver, and AMD passed Intel in data center revenue last quarter. Tellingly, when Intel demoed disaggregated inference at Computex, NVIDIA GPUs handled one phase and SambaNova chips the other while Intel supplied the coordinating CPU. Crescent Island, Intel’s actual inference chip, does not sample until later this year, and Intel declined to publish its memory bandwidth. Amazon Bought the Company Behind DuckDB On August 26th, Amazon signed a deal to acquire DuckLabs, the Amsterdam company behind DuckDB. Cochrane spends time explaining what DuckDB is, since listeners outside the data world may never have encountered it. The problem it solves is familiar. Querying a large pile of data files traditionally meant either running a database server or spinning up a data warehouse with a cluster, a bill, and a loading pipeline. Both are heavy machinery for a question you wanted answered in ten seconds. DuckDB instead ships as a library rather than a server. You add it to your program like any other package, point it at your files, and write ordinary SQL directly against them. The data never moves. The usual comparison is SQLite, which is embedded in nearly every phone and browser on earth. Where SQLite excels at looking up one record, DuckDB rebuilds that embedded idea for chewing through millions of rows. It now sees roughly 62 million monthly downloads on Python’s package index alone, up from about 25 million last October. AWS says it is buying the company, not the project. DuckDB stays free and open source under the MIT license, held by a Dutch nonprofit foundation, and the founders join AWS while continuing to run technical direction from Amsterdam. Andy Warfield, a VP and distinguished engineer at AWS, described DuckDB as “the glibc of structured data: a lean, unglamorous, ubiquitous dependency that a great deal of software links against and almost nobody has to think about.” Cochrane sits with what that ownership means. He reaches for an analogy: imagine Daniel Stenberg selling curl. He doubts it would ever happen, but the concept alone is startling given how much infrastructure depends on it. His read is that AWS is betting DuckDB becomes as foundational as curl and SQLite already are. GitHub Has One AI Model Grade Another One’s Homework GitHub shipped Project HydraFusion into Copilot as a research preview. Instead of routing your request to a single model, it picks one of three approaches per request. Sometimes one model simply answers. Alternatively, a cheaper model drafts, and a quality gate decides whether to escalate. The interesting one is Critique. One model writes the code, a separate model from a different family reviews it read-only, and the original gets one pass to revise. Despite the name, nothing is fused here. There is no voting and no merging, just one model at a time with a gate deciding whether to spend more. GitHub explained the reasoning in an earlier post: “a model reviewing its own work is still bounded by its own training biases: the same training data and techniques, the same blind spots.” Research supports it. A team at NeurIPS in 2024 showed that models recognize their own writing and score it higher than human graders do. GitHub’s earlier number had a Claude Sonnet and GPT critic pairing closing about three-quarters of the gap between Sonnet and the larger Opus model. This mirrors a workflow Cochrane uses constantly and has described on a previous episode. He runs a cross-check review with a second model from a different company, and it routinely surfaces issues the first model missed. He explains that different training data, different engineers, and different reinforcement approaches build different internal biases about what counts as correct. Looking ahead, he expects more of these “Frankenstein patterns” where models from different training families work together. The Best of IFA 2026, and What You Can Actually Buy IFA opened to the public in Berlin for its 102nd year, with about 1,900 brands. Cochrane splits The Verge’s roundup in two, since much of what generates headlines at these shows never ships. Starting with real products, iRobot’s flagship Roomba Max 875 Combo runs $1,199 and ships in about two weeks. Its SealForce feature drops a hidden skirt from the chassis when it detects carpet, sealing against the fibers so suction concentrates instead of leaking out the sides. That reaches 35,000 pascals, iRobot’s strongest yet. A step-down model at $899 carries the same trick, though Cochrane balks at both prices. Anker’s Soundcore Sleep 4 Pro earbuds arrive in November at $349.99. The charging case carries its own round touchscreen, so you pick soundscapes, set alarms, and read sleep stats without your phone. Optical sensors read heart rate and variability from the ear canal, which beats the wrist for accuracy, and the case masks a snoring partner. Cochrane remains unconvinced about sleeping with earbuds in. Philips also has smart rope lights, the Hue Liane 360, on sale now. They glow evenly around the tube rather than showing individual LEDs. They also run $400 for three meters, which works out to about $130 per meter of rope light. As for concepts nobody can buy, iRobot showed a robot vacuum that carries a smaller robot vacuum on its back in a garage and lowers it to deploy. Lenovo brought a 14-inch laptop whose screen rolls out to 17 inches at the press of a button, which reviewers call the first rollable that feels close to shippable. Tecno showed a phone with essentially no border around the screen, and Acer had a Windows gaming handheld that swivels its screen up over a keyboard. Two themes ran through the show. Humanoid robots were the loudest thing on the floor, and IFA’s own CEO framed the event as being about robots that work rather than robots that demo. Meanwhile, AI stopped being its own product category and became an ingredient, showing up in refrigerators, treadmills, dishwashers, and motorized TV mounts. Farmed Salmon Isn’t the Omega-3 Machine It Used to Be USDA scientists measured farmed salmon and found considerably less of the good fat than the government’s own database claims. EPA and DHA are two fatty acids you get almost entirely from fish. Your body can build them from the plant version, but only in tiny amounts, so the NIH’s position is that eating them is the only practical way to raise your levels. Those fatty acids are structural pieces of every cell, with DHA concentrating in the brain and retina. That is why the federal dietary guidelines, issued jointly by USDA and Health and Human Services, recommend at least eight ounces of fish a week and steer you toward salmon. That amount is calibrated to deliver about 250 milligrams a day. Published in Frontiers in Nutrition last month, the study found EPA and DHA in farmed Atlantic salmon came in 54.7 percent lower than USDA’s own reference values, last updated in 2018. A three-ounce serving fell from roughly 1,670 milligrams to about 756. Consequently, two servings a week now fall about 14 percent short of the target. Plant-derived fats meanwhile rose two to three times over. The likely cause is feed. Salmon are carnivores, and farms once fed them oily little fish. There was never going to be enough of those as the industry scaled, so crops filled the gap: soy, canola, sunflower and linseed. Importantly, the study does not claim to have proven this and calls the feed shift a plausible explanation. Independent corroboration lends it credibility. Researchers at Stirling measured a similar halving in Scottish salmon between 2006 and 2015, and Norway’s marine institute saw it across thousands of samples. There is a land dimension too. Roughly half the world’s soy grows in South America, where rainforest gets cleared for feed. Matthew Hayek, who studies the environmental cost of protein at NYU, told Inside Climate News that “soy is a major, important protein and oil ingredient in fish farming.” Adding up two decades of soy across all fish farming, he puts the extra forest clearing at around the area of Nicaragua or Bangladesh. That figure covers all fish farming rather than salmon alone, and Hayek notes it is hard to attribute soy use to any single species. Cochrane closes with two caveats. First, the study measured only eight fish, bought around Maryland, DC and Virginia over six weeks in 2023, and nearly all sourced from Chile. That is not a national survey, and the authors say plainly the sample was not large enough to change government advice. Rather, it flags that a federal database value needs rechecking, which is what the paper set out to do. Second, on whether you should care, farmed salmon still beats beef, chicken and eggs by a mile, since those carry essentially zero EPA and DHA. What it loses is its crown among fatty fish, dropping to mid-pack behind herring, sardines and mackerel and roughly level with trout. The broader health case is also softer than the 2000s suggested. A review of 86 trials covering 162,000 people found supplements barely moved heart attacks or deaths, so eating fish and swallowing fish oil are not the same claim. No producer has responded to the findings, and USDA, whose own scientists ran the study, declined an interview and did not answer emailed questions. Cochrane wraps up with housekeeping and a note that he will be back on Labor Day. The post The Mainframe Learned to Speak Arm #1875 appeared first on Geek News Central.
Mike Calvo of Pneuma Solutions joins Steven Scott to explain how Remote Incident Manager for iOS lets blind users run a Windows PC or Mac from an iPhone, and why AI document remediation is about to change what accessible information costs.Mike also shares how a blind CEO who does not write code now designs product screens with AI, what Pneuma's own trained models do differently from ChatGPT or Claude, and why the company built Accessible Archive to run on a company's own hardware with no per-page fee.[Sponsor]This episode is supported by Pneuma Solutions. Creators of accessible tools like Remote Incident Manager and Scribe. Find out more at https://pneumasolutions.com/The RIM mobile client flips the usual remote access story. Rather than controlling a phone from a computer, it turns an iPhone into the controller, so you can connect a keyboard, drop into your home machine, and work with your own screen reader, add-ons and scripts exactly as you left them. Android is in development. Mike explains what is free right now, what carries a subscription, and why the difference comes down to maintaining an open connection back to your machine.Mike then moves on to documents. Pneuma Solutions started in 2019 doing remediation with heuristics and tools like ABBYY FineReader, an approach that broke on corner cases and, more importantly to buyers, disturbed how a document looked. Mike describes how his son David trained models on millions of accessible documents, models that know nothing about history or writing and do one thing only. Because they can run on local infrastructure with a GPU, the per-page cost of remediation disappears, which matters when an organisation is staring at decades of backlog and an ADA Title II deadline.There is a clear-eyed explanation of what remediation actually means, headings, tables, lists, alt text and WCAG standards, plus the just-in-time approach Pneuma is building into upload portals. The example that lands hardest: court exhibits scanned as graphical PDFs, uploaded, published, and then handed to a blind lawyer who cannot read a word of them.Running through it all is Mike's argument about employment. Blind people can describe what they want with precision, and AI has become the lever that turns description into product. He points to his own team, to a developer earning a living largely through vibe coding, and to the unemployment rate in the blind community as the problem worth attacking.Relevant LinksPneuma Solutions: https://pneumasolutions.com/Remote Incident Manager: https://pneumasolutions.com/remote-incident-manager/Scribe: https://pneumasolutions.com/scribe/ADA Title II web and mobile accessibility rule: https://www.ada.gov/resources/web-rule-first-steps/ABBYY FineReader PDF: https://pdf.abbyy.com/ScribeMe: https://apps.apple.com/us/app/scribeme/id6739640292Plaud Note Pro: https://www.plaud.ai/pages/plaud-note-proREAPER: https://www.reaper.fmDouble Tap: https://www.doubletaponair.com ----Follow on:YouTube: https://www.doubletaponair.com/youtubeX (formerly Twitter): https://www.doubletaponair.com/xInstagram: https://www.doubletaponair.com/instagramTikTok: https://www.doubletaponair.com/tiktokThreads: https://www.doubletaponair.com/threadsFacebook: https://www.doubletaponair.com/facebookLinkedIn: https://www.doubletaponair.com/linkedinSubscribe to the Podcast:Apple: https://www.doubletaponair.com/appleSpotify: https://www.doubletaponair.com/spotifyRSS: https://www.doubletaponair.com/podcastiHeadRadio: https://www.doubletaponair.com/iheartAbout Double TapHosted by the insightful duo, Steven Scott and Shaun Preece, Double Tap is a treasure trove of information for anyone who's blind or partially sighted and has a passion for tech. Steven and Shaun not only demystify tech, but they also regularly feature interviews and welcome guests from the community, fostering an interactive and engaging environment. Tune in every day of the week, and you'll discover how technology can seamlessly integrate into your life, enhancing daily tasks and experiences, even if your sight is limited."Double Tap" is a registered trademark of Double Tap Productions Inc. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Durante muito tempo, escolher um notebook potente significava aceitar alguns compromissos: máquinas mais pesadas, baterias que duravam menos e desempenho reduzido quando o computador era desconectado da tomada. A ASUS quer mudar essa lógica com o Zenbook A16, novo notebook premium da marca que chega ao Brasil combinando uma tela de 16 polegadas com peso de apenas 1,2 kg. Em entrevista ao Podcast Canaltech, Murilo Tunholli, da ASUS Brasil, explica que uma das principais apostas do modelo está no Snapdragon X2 Elite Extreme, da Qualcomm. O processador usa arquitetura ARM, que busca combinar desempenho com maior eficiência energética. Segundo o executivo, a proposta é fazer com que o notebook mantenha uma experiência semelhante mesmo quando está longe da tomada. A comparação feita durante a entrevista é com o celular. Ninguém espera que um smartphone fique mais potente quando é conectado ao carregador. A ideia é aproximar o comportamento do notebook dessa experiência. A bateria também ganha destaque. A ASUS estima até 21 horas de reprodução de vídeo offline. Em situações mais próximas do cotidiano, a empresa fala em até 18 horas de streaming e mais de 15 horas de navegação na internet. Outro desafio foi chegar a esses números sem transformar o A16 em um notebook pesado. A fabricante optou por uma bateria de 70 Wh e reorganizou componentes internos para preservar o peso de 1,2 kg. O desempenho gráfico também faz parte da estratégia. A GPU integrada Adreno é voltada principalmente para produtividade e aplicações de criação, mas a ASUS afirma que o avanço dos gráficos integrados começa a diminuir a necessidade de uma placa de vídeo dedicada em determinadas atividades. No Brasil, o Zenbook A16 chega inicialmente vendido exclusivamente pela loja online da ASUS por R$ 22.999. A empresa pretende acompanhar a recepção do produto antes de decidir sobre uma eventual expansão para outros canais. A estratégia já foi adotada com o Zenbook A14, que começou como produto importado e, após a aceitação no mercado brasileiro, passou a ter fabricação local e presença em mais varejistas. Durante a conversa, Murilo também explica como tecnologias apresentadas primeiro nos modelos premium podem chegar posteriormente a outros produtos do catálogo da ASUS. Você também vai conferir: Netflix supera auge da TV paga com plano que exibe anúncios, estudo revela quanto dura a bateria de um carro elétrico e 6G poderá funcionar como radar e detectar até quem está offline. Este podcast foi roteirizado e apresentado por Fernanda Santos e contou com reportagens de Marcelo Fischer, Danielle Cassita e João Melo. A trilha sonora é de Guilherme Zomer, a edição de Vicenzo Varin e a arte da capa é de Erick Teixeira.See omnystudio.com/listener for privacy information.
My podcast guest this week Alex Ksendzovsky, co-founder of The Biological Computing Co. Alex and his team are doing something that sounds straight out of science fiction: they're extracting computational principles from living neurons and translating them into software that runs on standard silicon. Alex and I chat about The Biological Computing Co. is bridging the gap between biological tissue and GPU-accelerated algorithms, explore how this approach is making foundational models faster and more efficient and where Alex sees this ground breaking technology headed in the future.
In this episode of Agentic Conversations, we sit down with Ambud Sharma, Principal Engineer at Pinterest, responsible for general technology efficiency, fresh off delivering a controversial keynote on AI infrastructure optimization at scale.Ambud walks us through his Five Layer Cake framework - a structured approach to driving efficiency across every level of the AI stack, from silicon and hardware procurement to model selection, inference engine design, and governance. We explore how decisions compound across layers to unlock real business growth, and how the wrong choices can lock you into expensive commitments for years.We stress test the framework against two very different business models: what the stack looks like if you are building the next Cursor, and how it changes entirely if you are building the next YouTube. Along the way we cover hardware immutability, inference engine warm-up costs, GPU occupancy, context switching, quantization trade-offs, model routing, and why experimentation discipline is the only thing that keeps AI infrastructure costs from getting out of control.We also look at how this framework holds up in the emerging agent era, what changes when agent-to-agent communication becomes the norm, and why agent traffic just passed bot traffic on Cloudflare. The conversation closes on a deceptively simple takeaway: there is no silver bullet, and experimentation at every layer always comes first.Pinterest: https://about.pinterest.com/Alex Salkever: https://www.linkedin.com/in/alexsalkeverAmbud Sharma: https://www.linkedin.com/in/ambudTimestamps:[0:00] Introduction and the controversial keynote[2:09] The five-layer cake explained[4:30] Why hardware decisions are irreversible[6:47] Two business models: building Cursor vs YouTube[10:42] Applying the five layers to a YouTube-style company[14:23] Experimentation as the core efficiency method[17:09] Inference stack: context switching and warm-up costs[20:07] Model layer: why changing models breaks everything[24:10] When you should not use an LLM at all[26:41] Governance and routing: right model for the right task[29:20] Horror stories of unchecked token spend[31:37] Experimentation discipline without stifling innovation[34:35] How the five layers change in the agent era[36:05] Agent-to-agent communication and governance complexity[38:27] Core takeaway: experimentation first at every layer
Join Ronak Kumar Samantray, Founder and CEO of TakeMe2Space, for a deep dive into the next paradigm shift in digital infrastructure: bringing data centers into Low Earth Orbit (LEO). Having previously co-founded SaaS pioneer NowFloats—which scaled to support millions of businesses before its acquisition by Reliance—Ronak is now applying his tech background to revolutionize space infrastructure. In this episode, he explains why the current satellite model of downloading raw, unprocessed petabytes to terrestrial ground stations is breaking down, and how putting GPU-powered AI compute directly into orbit transforms satellites from simple sensors into intelligent, real-time edge data centers.
Deze Einde van de Week Live is ook te bekijken op https://youtu.be/bgH4_v-KUzI Deze talkshow wordt mede mogelijk gemaakt door MSI. Alle meningen in deze video zijn onze eigen. MSI heeft inhoudelijk geen inspraak op de content en zien de video net als jullie hier voor het eerst op de site. Terwijl het weer van nat naar droog gaat, luiden wij het weekend in met een gezellig anderhalf uur aan geklets over jouw favoriete hobby: het spelen van videogames. Dat inluiden doen we met behulp van een verse editie van Einde van de Week Live. De vodcast die je overal kunt bekijken of beluisteren. Deze week hebben Daan en Huey dienst. Een duo dat afgelopen donderdag tweemaal de PlayStation State of Play heeft bekeken. De reguliere en de Japanse editie. Welke games vielen hen positief op en wat gaan ze zeker spelen? Deze onderwerpen en veel, veel meer zie en hoor je voorbijkomen in de Einde van de Week Live van vrijdag 4 september 2026. Niet één, maar twee keer een PlayStation State of Play Andere onderwerpen betreffen de mogelijke aankondiging van een nieuwe StarCraft-game, de mogelijkheid dat we nu toch echt de reboot van Splinter Cell gaan zien en natuurlijk het nieuwe nieuws en geruchten over GTA 6. Pak 330 euro korting plus twee games bij de koop van de Katana 15 HX gaming laptop MSI zet deze week de welbekende Katana 15 HX in de spotlights. Een gaming laptop met een Intel Core i7-14650HX 16-core processor, een NVIDIA GeForce RTX 5060 GPU, een 15,6" Full HD 144Hz IPS-scherm, 16GB DDR5 RAM, 512GB PCIe Gen4 NVMe SSD, MSI Cooler Boost 5 en Nahimic 3 Audio voor verbeterde audio voor een meeslepende game-ervaring. Dit alles weegt slechts 2,4 kilogram en is hier met 330 euro korting te koop bij Bol.com. Krijg je er nog 2 games (Star Wars en Tomb Raider) bij ook. Haal nu je tickets voor Akira, een van de beste animefilms ooit, nu in 4K gerestaureerd De animatiefilm Akira creëerde in 1988 een ware schokgolf en opende de Westerse sluizen voor anime, manga en andere coole Japanse stuff. Het onheilspellende toekomstvisioen dat zich afspeelt in het Neo-Tokio van de toekomst van 2019 werd een grote cultklassieker van Katsuhiro Otomo en geldt als een van de meest invloedrijke animatiefilms aller tijden. Deze prachtfilm ziet er dankzij een 4K-restauratie nu mooier en grootser uit dan ooit – helemaal op een IMAX-scherm. En jij kan er, net zoals wij dat gaan doen, van genieten in de Nederlandse theaters. Haal dan snel hier je tickets.Wil je adverteren bij de podcast Gamekings óf misschien bij een andere podcast van ILVY Network? Mail dan naar management@ilvy.com en/of kijk even op de website : https://ilvy.com/podcastSee omnystudio.com/listener for privacy information.
In this episode of Tank Talks, host Matt Cohen sits down with Keegan McCallum, co-founder and CEO of uRUN, an inference provider built for persistent, steerable, interactive AI experiences. Keegan's path took him from jailbreaking iPods in Thunder Bay, Ontario, to building the infrastructure that powered Luma AI's viral Dream Machine launch, scaling from 500 to 9,000 H100s in six hours to process over half a million videos.They break down why the industry is wrong to shoehorn interactive AI into old request-response systems and why the future belongs to persistent, stateful, real-time experiences. Keegan explains the technical differences between LLMs and video models, why infrastructure is the biggest bottleneck for the next wave of AI, and what it takes to recruit top engineering talent.He also shares hard-won lessons on selling to developers, the Canadian AI talent landscape, and his bold, contrary views on the AI supply chain. From avatars and steerable video to coding agents generating 1,000 tokens per second, this episode is packed with frontline insights.Whether you are a founder building in AI, an infrastructure engineer, or simply curious about where human-computer interaction is headed, Keegan delivers a clear-eyed look at the future of interactive AI.–A big thanks to our sponsor, Moomoo CanadaThis is the kind of tooling that used to live on a Bloomberg terminal, but now it is on your phone, just a few taps away. They offer real-time data, full options chains, and an AI assistant that actually explains trading strategies.Moomoo is the perfect place for people who want to take their money seriously. Open an account today at moomoo.caThe Luma AI Experience: From 500 to 9,000 H100s (04:18)* Joining Luma after building an early LLM infrastructure startup* The Sora demo and the pitch to build and release a competing video model* Building the inference infrastructure for Luma's launch* Scaling from roughly 500 to 9,000 H100s in six hours as Dream Machine went viralWhat a Viral AI Launch Actually Looks Like (06:03)* The request queue hitting 100,000 during the Dream Machine launch* Processing half a million videos in 12 hours and reaching one million users in four days* Why having a queue protected the user experience during the surge* Building infrastructure that could run across raw VMs, Slurm, Kubernetes, and different GPU providers* Why founders need to plan for success before the product takes offWhy AI Production Infrastructure Is Being Left Behind (10:50)* The tension between training the next model and supporting production workloads* Why infrastructure teams can lose talent to model research* The opportunity Keegan saw in building a dedicated ML infrastructure team* Creating infrastructure that model labs across different modalities can use instead of rebuilding it themselvesWhy Video AI Needs a Different Infrastructure Stack (12:32)* The limitations of moving media through traditional object storage* The latency problems that make real-time previews and interaction difficult* Streaming audio, video, images, and text as continuous inputs and outputs* uRUN's focus on making persistent, interactive AI experiences easier to build* Moving AI from something you message toward something that feels present with youThe Shift Toward Omnimodal and Interactive AI (19:01)* What Keegan learned working closely with Luma's research team* Why understanding the research direction matters when building infrastructure* The move toward AI that can understand vision, audio, language, reasoning, and sensor data* How AI has evolved from single prompts to conversations, agents, and interactive systems* Why highly interactive, real-time AI experiences could become the next major interfaceSmall Models, Specialized AI, and the Future of Inference (26:31)* Keegan's previous startup thesis around smaller models and delegated inference* Why continual learning remains a difficult problem* Running capable smaller models locally and knowing when to delegate to larger models* Why enterprises may increasingly want differentiated models running in-house* The potential shift toward using large models for the hardest problems rather than every taskWhere uRUN Is Seeing Early Demand (30:38)* AI avatars for customer support and sales* Steerable video generation that lets users intervene while content is being created* Video transformation for creators, avatars, and live experiences* Real-time visual effects for music festivals and live events* Why world models could eventually have applications in open-world gamingThe Infrastructure Problem Holding Back World Models (00:32:54)* Companies developing world models without the infrastructure to serve them* Why scaling world models remains difficult* The enormous context requirements involved in tracking everything happening in a simulated environment* The need for distributed infrastructure as these models become more capableBuilding a Team Around Hard Problems (35:18)* Finding engineers with deep experience in infrastructure and low-latency inference* Bringing together expertise from New Relic, AWS, Superorbital, and other infrastructure environments* Why Keegan looks for engineers who can work across multiple areas* Building a small team of high-leverage people instead of optimizing for headcountThe Trait Keegan Looks for in Exceptional Engineers (37:47)* Why culture and quality matter more than quantity* Looking for people who take ownership and solve problems without waiting for instructions* The importance of urgency and experience operating under pressure* Why hard problems can attract customers, investors, and exceptional talent* Hiring people who can wear multiple hats and produce at a high levelBuilding AI Companies in Canada (41:29)* Why Keegan keeps finding Canadians throughout Silicon Valley* The advantages of building a strong team in Canada at a lower cost* Government funding and the opportunity to build ambitious companies with smaller teams* The challenge of Canadian companies and talent being pulled toward Silicon Valley* Why Canada needs more ways for companies to grow instead of selling earlySelling Infrastructure to Developers (44:23)* Why developers want to try the product rather than hear a list of performance claims* Giving technical users something they can get their hands on and test* Why open-source surfaces can help developers understand how a product works* Documentation and developer experience as a core part of selling infrastructure* Why developers have a very high tolerance bar for technical errorsThe Infrastructure Shift Coming Next (46:43)* Why persistent streaming infrastructure could become standard* Moving away from request-response systems toward long-running streams* The importance of maintaining state across continuous interactions* Potential applications in real-time video, virtual try-on, image editing, and coding agents* The infrastructure required for AI systems that can generate and respond almost instantlyAbout Keegan McCallumKeegan McCallum is the co-founder and CEO of uRUN, an inference provider built to handle persistent, steerable, and interactive AI experiences. A self-taught engineer who started by jailbreaking iPods in Thunder Bay, Keegan has built a career on tackling the hardest infrastructure problems in Canada. He has held leadership roles at Colony Networks and served as the Head of Engineering at Luma AI, where he led the infrastructure team that scaled the viral Dream Machine launch from 500 to 9,000 GPUs in a single day. His deep technical expertise spans ML infrastructure, cloud-native systems, and low-latency edge computing.Connect with Keegan McCallum on LinkedIn: https://www.linkedin.com/in/keeganmccallum3?originalSubdomain=caVisit uRUN's website: https://urun.sh/Connect with Matt Cohen on LinkedIn: https://ca.linkedin.com/in/matt-cohen1Visit the Ripple Ventures website: https://www.rippleventures.com/ This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit tanktalks.substack.com
Kirk Offel sits down with Curt Holcomb, Vice Chairman of JLL's Data Center Solutions group, for a wide-angle look at how the industry grew from one-rack data closets and telecom switch sites into multi-building campuses that increasingly require their own power generation.Drawing on 26 years at the front end of data center transactions, Curt breaks down the dot-com boom and bust, the rise of colocation, COVID's impact on enterprise demand, the arrival of GPU-as-a-service, and the power constraints shaping today's AI buildout. He and Kirk discuss why site selection now begins with available megawatts rather than acreage, whether the current growth cycle is a boom or a bubble, and why Texas is attracting so much infrastructure investment.The conversation also covers closed-loop cooling, tax incentives and community benefits, public distrust, town-hall opposition, and the industry's failure to explain itself clearly. They close with the innovation, careers, and veteran talent that will be required to build the next generation of digital infrastructure.For more about us: https://weareoverwatch.com/data-center-revolution-podcast/For guest inquiries reach out to: podcasts@weareoverwatch.com
AI CapEx is on track to reach roughly $700 billion at the four largest hyperscalers this year, and equity analyst Irena Petkovic breaks down where all that money is going and what has to be true for it to earn a return. She explains how data centers turned from cost centers into revenue-producing AI factories, the four ways hyperscalers monetize compute, and how the token economy actually works. She then weighs the early evidence of returns against the risks around token pricing, debt financing, and public backlash, and describes how the team positions the portfolio around it. 0:00 - Introduction: The $700 Billion AI CapEx Question 1:29 - How Big Is the Spend? Apollo, Telecom, and Railroads 4:11 - Data Centers as AI Factories: From Cost Center to Revenue 5:53 - Four Ways Hyperscalers Monetize Compute 7:10 - The Token Economy Explained 11:26 - Early Evidence of Returns on AI Investment 14:39 - The Bear Case: Token Prices, Debt, and Backlash 18:49 - Positioning the Portfolio Around AI 20:35 - Outro & Subscribe Highlights: Hyperscaler AI CapEx of about $700 billion this year is a scale rivaled historically only by the railroads. Capital intensity at Microsoft, Meta, Google, and Amazon has jumped from 5-10% of revenue to upwards of 45%. A high return on invested capital justifies spending down free cash flow rather than protecting it. The data center is now a revenue-producing AI factory that turns electricity and chips into sellable tokens. Compute is monetized four ways: GPU rental, productivity products, enhancing own businesses, and selling tokens. Early returns look encouraging, with sub-three-year hardware payback cited and demand exceeding supply. Risks to watch are token prices falling faster than volumes and a shift toward debt-funded build outs. Host: Rob Campbell, CFA, Mawer Institutional Portfolio Manager Guest: Irena Petkovic, CFA, Mawer Equity Analyst This episode is available for download anywhere you get your podcasts. Founded in 1974, Mawer Investment Management Ltd. (pronounced "more") is a privately owned independent investment firm managing assets for institutional and individual investors. Mawer employs over 250 people in Canada, U.S., and Singapore. Visit us at: https://www.youtube.com/@MawerInvestment https://www.mawer.com https://www.linkedin.com/company/mawer-investment-management/ https://www.instagram.com/mawerinvestmentmanagement/
Este episodio 828 es la carta de presentación de la Temporada 9 de atareao con Linux. Treinta y cuatro episodios ya guionizados, siete etapas, y un objetivo claro: construir tu cerebro digital sobre Linux con herramientas locales, sin depender de nubes ni suscripciones.Pero antes de mirar adelante, toca hacer balance. La T08 empezó prometiendo Docker, selfhosting y Android, y sí, hablé de todo eso. Pero en abril de 2025 la IA local irrumpió con fuerza y la temporada viró hacia Ollama, modelos locales, RAG, MCP. Fue un giro desordenado, lo reconozco. Pero también fue el germen de todo lo que viene ahora.De eso va esta T09: de poner orden al caos. Siete etapas, de menos a más, para que sigas el hilo hagas el nivel que hagas.Etapa 1 — Recursos básicos: los cimientos de tu laboratorio de IA. Skills para tu agente, herramientas de publicación, el botiquín del explorador.Etapa 2 — Skills y MCPs: el pegamento. El Model Context Protocol ha madurado hasta ser un estándar abierto — lo soportan Claude, ChatGPT, VS Code, Cursor. Ya no es un experimento, es el USB-C de la IA. Y de paso, herramientas del ecosistema atareao como watchbeat (monitor de uptime en Rust) y alloy (dashboard Docker con OIDC).Etapa 3 — GraphRAG, el gran hito: de RAG vectorial a grafos de conocimiento. Mientras el RAG clásico devuelve fragmentos sueltos y tú unes los puntos, GraphRAG construye un grafo con entidades y relaciones. Preguntas como "qué contenedores están detrás de Traefik" pasan a ser una consulta directa a tu mapa de conocimiento. Usaremos LightRAG, que con 39.000 estrellas ya superó al Microsoft GraphRAG original. Esto ocupa tres episodios.Etapa 4 — Multimedia: Whisper para speech-to-text, TTS local, ffmpeg, visión artificial, y el pipeline de YouTube a conocimiento con yt-dlp. Rematamos con RAG multimodal.Etapa 5 — Orquestación: systemd timers, asyncio, just, y CrewAI para montar equipos de agentes.Etapa 6 — Proyecto final: dos episodios para construir El Asistente que te Conoce y ponerlo en producción con Quadlets.Etapa 7 — El futuro: mantenimiento de tu cerebro digital y hacia dónde va todo esto.Entre medias, herramientas Linux: shuul, sqlite-utils, yq + jq, Rust en el kernel, Wayland vs X11, la guerra de los filesystems.No necesitas una GPU de 3000 euros ni un doctorado. Con 16 GB de RAM y un CPU decente ejecutas modelos de 7B a 14B. Esto es IA local, en tu máquina, con tus datos.Capítulos del episodio:00:00 — Introducción y bienvenida a la Temporada 901:47 — Balance T08: de Docker y Selfhosting al boom de la IA04:37 — El momento adecuado para cada tecnología06:46 — El gran objetivo: tu cerebro digital08:44 — Roadmap T09: 30 episodios ya guionizados11:28 — Skills y MCPs imprescindibles12:43 — GraphRAG: de RAG a grafos de conocimiento14:01 — RAG vs GraphRAG: el mapa de tu conocimiento17:28 — Herramientas del ecosistema: alloy, populater, watchbeat20:08 — ¿Para quién es esto? De veteranos a escépticos22:08 — No es hype: es un cambio de paradigma24:30 — El momento perfecto para el linuxeroToda la info y el roadmap completo en atareao.es/828.Más información y enlaces en las notas del episodio
仅仅三年前,英伟达还是一个“卖无法替代的铲子”的公司:它的产品供不应求,客户抢着支付现金,只为获得芯片;它的毛利极高、资产很轻——一言以蔽之,这是华尔街最喜欢的赚钱机器。但此后,英伟达主动打破了这种状态,它不断把自己的钱和信用借给它的客户,又让这些客户继续购买自己的产品,在这个过程中,英伟达越来越大,它身上的杠杆也越来越多。正面评价是:英伟达正在用自己的资产负债表为杠杆,建立一个新的AI行业生态;负面批评是:英伟达在自己创造自己的需求,它需要为AI的泡沫负责,甚至有人把它和20多年前的朗讯相提并论。这些讨论当然关乎对英伟达和相关公司股价的预测。如果你想在本期节目获得这种预测,很抱歉我们做不到。我们能做的,是厘清那些纠缠在一起的概念:供应商融资、回流交易、循环融资。换句话说,搞清楚英伟达到底在干什么。当然,节目最后,话题不可避免地滑向那个危险又迷人的概念:泡沫。| 主播 |肖文杰、约小亚| 时间轴 |01:15 英伟达攒了一个5000亿美元的新局04:56 英伟达扶持CoreWeave极简史10:41 担保1085亿美元,到底是什么意思?13:09 到底算不算财务欺诈,判定的两个要素16:30 英伟达的资产负债表不再完美19:12 GPU抵押贷款的利率从15%降到6%22:58 朗讯和光纤泡沫25:51 朗讯和英伟达的两个区别27:53 GE资本的启示:懂技术的人放贷,才算得清残值30:38 英伟达相比GE资本的劣势和优势33:07 比担保更笼统的残值支持35:19 算力到底是不是可以交易的金融资产38:01 三年时间,英伟达从赚钱机器变成了AI行业的中央银行38:58 我们对泡沫的理解| 延伸资料 |SEC - NVIDIA Corporation Form 10-Q(截至 2026 年 7 月 26 日)NVIDIA - CFO Commentary on Second Quarter Fiscal 2027 ResultsNVIDIA - NVIDIA AI Factory Compute Is Becoming an Investable Asset ClassNVIDIA - NVIDIA Partners With Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR to Establish AI Compute Infrastructure Financing Platforms to Mobilize Over $500 Billion of Third-Party Capital硅谷101 - "消失"的万亿债务:深扒数据中心"影子借贷"、GPU金融化与次贷风险NVIDIA - NVIDIA Guarantees SB Energy's PORTS-Pike Technology Campus in Ohio to Exclusively Host NVIDIA AI ComputeNVIDIA - OpenAI and NVIDIA Announce Strategic Partnership to Deploy 10 Gigawatts of NVIDIA SystemsSEC - Lucent Settles SEC Enforcement Action Charging the Company with $1.1 Billion Accounting FraudConverge Digest - CoreWeave Secures $2.6B Loan as Lenders Take on GPU Renewal RiskThe Verge - Chipwrecked: Can Nvidia Avoid the Crash?Michael Burry - The Heretic's Guide to AI's Stars Part III: Tracepalooza & the BezzleTomasz Tunguz - Circular Financing: Does Nvidia's $110B Bet Echo the Telecom Bubble?The Week - The Fall of GE| 后期制作 |潘鑫| 声音设计 |刘三菜| 收听方式 |你可以通过小宇宙、苹果播客、Spotify、喜马拉雅、网易云音乐、QQ音乐、荔枝、豆瓣等平台收听节目。| 认识我们 |微信公众号:第一财经YiMagazine联系我们:thatisbiz@yicai.com
"La reproductibilité, c'est un truc que moi, j'ai l'impression que l'industrie a abandonné" Le D.E.V. de la semaine est Gabriel de Marmiesse, technical staff chez Kyutai. Dans cet épisode, il lève le voile sur le quotidien d'un dev AI : entre code non déterministe, gestion de fichiers de teraoctets et galères de GPU, rien n'est jamais simple. Gabriel partage comment la recherche et l'engineering s'opposent (et se complètent), pourquoi l'open source vit une crise de confiance à l'ère des LLM, et ce qui distingue fondamentalement le dev AI du dev classique. Il explique aussi comment son équipe tente de démocratiser l'audio, la trad en live, et les world models. On ressort avec les câbles qui chauffent, mais les idées claires !Chapitrages00:00:53 : Promesse de l'informatique00:01:52 : Présentation de Gabriel00:05:06 : Chez Kyutai, l'open source00:09:48 : Naissance d'Invincible Voice00:13:09 : Poids et déploiement des modèles00:23:27 : L'histoire de l'IA moderne00:30:37 : Écrire un modèle en code00:34:29 : Pourquoi les résultats varient00:39:35 : Du labo au produit00:45:37 : Le casse-tête des GPU00:52:08 : Comparer avec le web00:54:24 : Les world models arrivent00:59:41 : L'avenir des LLM01:05:01 : Préférences et tabulations Liens évoqués pendant l'émission Chaine de CPP: CppCon - YouTube
Topics covered in this episode: OpenAI's Python SDK has migrated to HTTPX2 TMOG - Native Task Manager for macOS, Windows, and Linux wrapture - one wrapper for mocking, tracing, and observability linkedin2md: turn your LinkedIn export into 40+ Markdown files Extras Joke Watch on YouTube About the show Sponsored by us! Support our work through: Our courses at Talk Python Consulting from Six Feet Up Connect with the hosts Michael: Mastodon / BlueSky / X / LinkedIn Calvin: Mastodon / BlueSky / X / LinkedIn Show: Mastodon / BlueSky / X Join us on YouTube at pythonbytes.fm/live to be part of the audience. Usually Tuesday at 7am PT. Older video versions available there too. Finally, if you want an artisanal digest of every week of the show notes in email form? Add your name and email to our friends of the show list, we'll never share it. Calvin #1: OpenAI's Python SDK has migrated to HTTPX2 The OpenAI Python SDK has migrated to HTTPX2, the Pydantic-stewarded fork of httpx. Pydantic picked it up citing "limited activity recently" in the original project, promising "a reliably maintained path forward." If you just use the default client, nothing to do. No code changes. The catch is TLS. Quoting the guide: HTTPX "previously verified certificates against the CA bundle provided by certifi. HTTPX2 instead uses the operating-system trust store, and the SDK no longer installs certifi." That "can break certificate verification in minimal container images without system CA certificates, environments using corporate TLS-inspecting proxies, and deployments that relied on a custom or modified certifi bundle." The fix is SSL_CERT_FILE or SSL_CERT_DIR, or pass your own ssl.SSLContext via verify. Deeper integrations need real edits: custom clients, auth handlers, hooks, and request mocking all take HTTPX2 objects now, and plain httpx is no longer pulled in transitively. So import httpx in your own code means declaring it yourself or moving over. Temporary escape hatch: a legacy HTTPX client Michael #2: TMOG - Native Task Manager for macOS, Windows, and Linux A native, deeply instrumented system monitor for macOS, Windows, and Linux, now in public beta - from Plummers' Software, i.e. Dave Plummer, who wrote the original Windows Task Manager and donated it to Microsoft in 1995. Wikipedia Three real native apps: Swift/AppKit on macOS, Win32 on Windows, C++/Qt 6 on Linux, with a shared C++ core keeping metric semantics aligned - no browser shell anywhere. One dense summary: CPU, clocks, thermals, GPU, memory, storage, network, energy, and the processes responsible for the load, all click-through. Per-core honesty: logical processor and NUMA views, P and E cores color-coded, optional kernel time, 60 FPS live meters. Memory with context: pressure, wired, compressed, cached, committed, available, and swap, plus configurable scrolling history. Processes that act like processes: tree view, filtering, sorting, follow mode, and native verbs including service and launchd control. Phosphor themes: light, dark, green, amber, blue, or mono, with color and saturation you tune yourself. Calvin #3: wrapture - one wrapper for mocking, tracing, and observability Graham Dumpleton, author of wrapt and the original New Relic Python agent, has released wrapture. The name is wrapt plus capture. The core idea: wrap real code instead of replacing it, so the real code still runs while you watch every call. Name a method with wrapture.binding(Class, "method"), open a timeline(), and you get a tape of what actually happened. Real return values, real nesting, arguments normalised against real signatures. tape.tree() prints the call graph as it ran. One mechanism, three jobs: monkey patching with a real lifecycle (apply, remove, suspend, plus returns, raises, transforms_args), unit testing that asserts on real call flow instead of a flat MagicMock call list, and ad-hoc tracing of a running app. The testing pitch is error paths. Inject TimeoutError at the payment gateway, then assert the ledger was never written. Stubs and mocks are strict and spec-required, and there is deliberately no bare Mock(). Tracing needs no code at all. A wrapture.toml naming targets and a sink, run with python -m wrapture main.py, and you get a live call tree with timings. It captures ordinary logging calls as nested events, and with the otel extra it exports spans, metrics and correlated logs with W3C trace ids that join across services. Every line of code and docs was AI-written under their direction, and they say so up front. Two weeks from first commit, eleventh alpha, over 1000 tests, 150+ pages of docs. Alpha on PyPI, needs Python 3.12+ and wrapt 2.4.0+. Michael #4: linkedin2md: turn your LinkedIn export into 40+ Markdown files Via Juan Manuel Daza - a Python CLI that unpacks LinkedIn's data-export ZIP into clean, per-category Markdown you can drop straight into an LLM. One command: linkedin2md Complete_LinkedInDataExport.zip, plus o for output dir, -lang en|es, and -pdf. 40+ output files: profile, experience, education, skills, connections, posts, comments, reactions, recommendations, endorsements, job applications, even ad targeting and LinkedIn's inferences about you. Built for LLM analysis: the README pitches NotebookLM, Claude Projects, Obsidian, and Ollama, with example prompts like "what patterns do you see in my career transitions?" PDF resume mode: -pdf renders an A4 CV via weasyprint, and degrades gracefully to Markdown-only if it isn't installed. Dependency note: "pure Python / zero-dep" holds for the Markdown path only - the PDF path needs weasyprint and markdown installed. Install: pipx install linkedin2md recommended, pip in a venv otherwise - 86% Python, 10 releases, v0.3.1 in May. Agentic dev angle: repo ships opencode config and an N3RV subagent pipeline, including a "judgment day" dual-model adversarial PR review. Extras Calvin: EVE Online Migrates to Python 3 Michael: Dinkus by Will McGugan Joke: Tao of Programming: Book 5 Maintenance
In June, the most capable American AI models stopped shipping as public launches and started shipping through a government gate. Six weeks later the gate is open again — and the real fight has moved to the layer no gate can touch. A Chinese open-weight model rattled trillions out of chip stocks, Washington pivoted from gating American closed models to threatening bans on Chinese open ones, the industry mounted its largest-ever policy counter-mobilization, and an American frontier model literally broke out of its lab and hacked another company. Knee-jerk reactions, or the beginning of real AI governance? Navigation: Intro The Gate Opens The Kimi Shock The Escape The Counterstrike and the Petition Interlude — The Low-Background Books The Investor Reckoning Conclusion Our co-hosts: Bertrand Schmitt, Entrepreneur in Residence at Red River West, co-founder of App Annie / Data.ai, business angel, advisor to startups and VC funds, @bschmitt Nuno Goncalves Pedro, Investor, Managing Partner, Founder at Chamaeleon, @ngpedro Our show: Tech DECIPHERED brings you the Entrepreneur and Investor views on Big Tech, VC and Start-up news, opinion pieces and research. We decipher their meaning, and add inside knowledge and context. Being nerds, we also discuss the latest gadgets and pop culture news Subscribe To Our Podcast Bertrand Introduction Welcome to Tech Deciphered Episode 80. This one, once again, will be all about AI, government, frontier models, and open weight counterstrike. A lot has been happening in the regulation space, in cybersecurity, in the launch of new models in the past, maybe just 6–8 weeks. It’s actually pretty insane how much happened. We believe it was time to do an episode to talk about where we are and maybe where all of this is going. Maybe let’s start with a summary of where we stand, all that June and July saga, so you, our listeners, can get up to speed if you are not already there. You want to start with some points? Nuno The Gate Opens Yeah. Again, to your point, the gate swings. The gate had closed. We had to prepare an episode for the gate closing, and then the gate reopened. Now we have a different episode. This will probably change again as we’re seeing there’s news every day. Let’s start maybe with the first 19 days of the gate closing. There was an executive order on June 2nd from President Trump that asked frontier labs to share models with the government, 30 days pre-release. It inferred the protected frontier model designation into that. Basically, it was effectively a de facto licensing agreement defined by an executive order of the President as of June 2nd. On June 9th, Anthropic launched Fable 5 and the famous Mythos 5 or Mythos. I’m not sure how you actually say it in English. Then on June 12th, there was an export control directive banning access by any foreign national. Since there’s no way to verify nationality in real-time, Anthropic had to switch the models off for everyone worldwide. Bertrand On this point, you could argue that there are possibilities to check IDs. Many services let you check IDs online. You can pre-check a flight by showing your ID. There are ways, it’s just that if you don’t want to follow what’s already available, because guess what? Maybe it slowed down your revenue growth, maybe it looks bad on you or whatever. My point is that there was actually an option. I think it’s already a decision from Anthropic to say it’s either on or off, but nothing in between. Nuno I think the point is they had no way implemented of doing it. If they implemented it, to your point, it would have hampered use in general. A lot of people wouldn’t have gone through that trouble of doing it. Anyway, long story short, in June 26th, the White House apparently asked OpenAI to limit GPT-5.6, so Sol, Terra, Luna, to only 20 vetted partners. Now, apparently, the trigger for a lot of these things that have been going on was that there was a jailbreak that was found by Amazon researchers. All of that led to this jumping around of, let’s close the gates. You have foreign nationals, and therefore, Anthropic got it out and said, “Hey, then we’re going to switch the models off until we can sort this out.” OpenAI was asked also to only allow it for certain vetted partners, et cetera. The government came in, closed the gates effectively, and said, “From now on, we need to be involved in this thing.” De facto regulation, there’s no doubt that this has imposed de facto regulation, certainly on the top players in the market. But then came the reversal. Bertrand, do you want to talk about the reversal, the gate swinging the other side? Bertrand Maybe I just wanted to say that as a user of Anthropic products, ChatGPT products, for the brief moments, a few days where Fable 5 was made available to the public before it was closed the first time, I immediately started using it. I must say it was a real issue to use it because the guardrails were pretty crazy. It would keep saying that my code was not okay, there was cybersecurity risk and stuff when I was doing absolutely reasonable development with absolutely no connection whatsoever to any cybersecurity risk, attack, detection, anything. Still, it would keep blocking me, degrading me to Opus 4.8 at the time. I just want to say this was already very hardcore what they were implementing, and not just hardcore, but in some ways, plain stupid for something that’s supposed to be super smart. It was totally unable to classify properly some of my work. I must say I was already disappointed. On top of it, the costs were insane. Half a day, I would reach my limits when I had the best plan you can get from Anthropic. My point is that there were some real serious issues when they launched Fable 5, even at that point. Nuno I had a similar issue. I used Fable 5 as well before they had to take it offline or take it off. I think the issue was really not that the guardrails failed. As you said, maybe the guardrails were actually too aggressive, but it was this jailbreak that caused the recall, apparently caused this knee-jerk reaction. Bertrand But my point is that it seems that it was not working either way. It would either overclassify something that’s absolutely not doing anything wrong, and it might fail to classify something that is actively trying to do some cybersecurity work. It’s a real issue of quality for a company that’s supposed to be at the forefront of quality of AI and everything. I think for me, there are already signs that something is deeply wrong. Nuno Then it’s reversed, right? We went the other way around. The government came out on June 26th and approved redeploying Mythos 5 to US organizations defending critical infrastructure, and then the export controls were effectively lifted on June 30th. July 1st, Fable 5 came back online for all of us to use. Shocking enough, with strings attached, that were different. They had some time to revise their commercial deployment of it along the way because it came back with some, “Now you have usage credits, but you have some limits on plan use, et cetera.” I’m like, “You guys, this was blocked. But meanwhile, you did have some time to do some commercial stuff around it.” Bertrand It was crazy. I’ve never witnessed any such crappy launch of any service whatsoever in 30 years in tech, it was so bad. Every day, they would change the terms of service. They would tell you it’s part of the plan. It’s not part of the plan. It’s part of the plan for three more days, and then it’s excluded. You have a special discount now, but then it goes back to full price. It was a total nightmare. I’ve never felt myself being so much mistreated by a company. I guess you saw the same, but when I started using the newest version of Fable 5, it was even worse, actually, I think. I couldn’t do any work with this crap. I let it go and work on the work I wanted it to do. It was simply not working. On top of it, you never know how long you are supposed to lose your credit, how fast. It was burning credit like crazy. Me, personally, I can say, very quickly, I actually stopped using it. I was like, “No, I cannot deal with this shit. My main model is back to Opus 4.8. I’m going to use Fable 5 for code review, but not anymore to control anything because I cannot trust it would do the job without stopping or changing models and stuff. I just cannot trust it.” Back to Opus 4.8 as my main model, I can say that my life was much easier. I use Fable 5 as a review mechanism, as a support mechanism, but not as the main mechanism. Suddenly, the guardrails were not so horrible anymore because it was used in a much lighter way, I guess. As a pain as a user, I think it was really bad. I don’t know your experience, but me, for me, it was unacceptable. Nuno I wouldn’t say it was as bad as yours in terms of just end-user experience. I think the terms of service switching back and forth, which went one further step, because then when they then launched Opus 5, they started making comparisons between Opus 5 and Fable so that people would migrate more and more to Opus 5 themselves, which is interesting. It’s like they’re saying “This is much cheaper. This is whatever. You’re not going to run of credits. You should use Opus 5,” kind of thing effectively. To your point, I don’t think they managed well the launch. They didn’t really manage it well. We’re moving people around. A lot of people are using this for stuff that’s like daily tasks, hourly tasks, anything that relates to code and co-work. It’s like, we need to have visibility on what your terms of service are going to be. Should I be using this new model or not? What’s happening to the other model? I don’t see it as negatively as you, Bertrand, but I see your point. It was clearly mishandled in terms of how they deployed it, how they were redesigning effectively their pricing scheme and their terms of service almost on a daily basis, at a certain point in time. We’re like, “Dude, there’s millions of people using this. You guys are making a lot of money.” Just moving it as it is. At this point in time, at the scale that these guys are at, it’s calling in people to say, how about we think through a class action suit at some point around pricing? Because you guys are changing the rules of the game all the time, right? Bertrand I don’t know if I need the class action, but for me, that joke that, “Let’s not rush too fast. The model is dangerous.” But still, they rushed the launch because it’s very clear that if they had enough compute capacity and stuff, they would not have to limit so much. They would not have to put so much cost per token and all of this. You can see that actually when they launch Opus 5, literally like 2, 3 weeks after, by most benchmark at launch, they tell you basically that, “You know what? Actually, Opus 5 is better than Fable 5 on 80% of the metrics.” They’re like, “What? Seriously? You couldn’t wait 2 weeks? Why did you even launch Fable 5 in the first place?” That’s another part for me that is quite literally insane, to be frank. It’s like, “Why? Why do you make us go through so much pain if it’s only to tell us after 2 weeks to…” “This new model, by the way, has less issues, less stuff, because 2, 3 times less is part of your plan, and it’s actually better by most metrics.” It’s like, “What’s going on here? What’s going on? Are you guys mad?” I don’t know. It was crazy. Personally, I still use Opus, now 5, as my main system and platform, Fable 5 for review, code reviews and the like. I don’t want to run into its stupid guardrails. I can see Fable 5, from my perspective, seems quite a bit smarter. I don’t know why they do this stupid benchmark showing you it’s actually worse than Opus 5. I guess they should have better benchmark if they want to demonstrate why you are supposed to pay 2, 3x more for a model versus another if it’s actually worse by most benchmark. Again, I still think it’s a huge mess from a marketing perspective, customer perspective. Me as a user, I really feel that they don’t want my money, and they couldn’t care less about me. This is even before everything else we’re trying to talk about. Nuno Yes. Maybe just to close the cycle on the reversal on the door opening the other way, finally, Commerce lifted the GPT-5.6 restrictions on July 8th, and then on July 9th, general availability across ChatGPT, Codex, and the API as well. What has this proved? It proved that now we have gating mechanisms, and certainly for closed models in the US, for sure. We had frontier models that were switched off worldwide in hours, and it took a couple of days, in this case, 19 days to restore them. There were concessions. Now we know that there were concessions around effectively institutionalizing that gate. Early government access to future models is, I think, now a given, certainly in the US. New safeguard frameworks are probably now having to be put in place. There are some stage limits now on who gets access to what for new models and how it happens. This voluntary executive order, so to speak, not really sure, has become effectively regulation enforcement path. It’s de facto regulation that now has been put in place. It has affected not just to the points we were making before, the access to these models, but also who gets access to these models, and actually potentially even pricing access to the models. It has probably some commercial implications as well as we just discussed along the way. Very significant. This is very significant. This is regulation, de facto at the table, imposed on the two largest players in the market by far by one government, in this case, the US government. This is significant. Actually, you could even allege it was imposed by the President because this was coming as part of executive orders. Really incredible. Pretty significant, fast, aggressive. It has created a regime that you could say it’s a regulatory regime, it’s a de facto regulatory regime. It has some significant pricing and licensing and commercial implications. It goes even beyond your classic regulatory framework. Very, very, very significant. Bertrand I don’t know if it goes beyond a classic regulatory framework. Nuno I think it does, because it has implications on who do you give access to? When government is saying you can only give access to these players, right? Bertrand Defense industry. It’s all over the defense industry. You cannot sell an F-35 like this. Nuno No, but that has commercial implications, Bertrand. That’s like you’re saying these are your customers, you go and use them. Bertrand That’s the defense industry. You cannot sell to Iran your F-35. No, that’s exactly the same story for me. Nuno No, no, no. It’s beyond that. These guys are saying when they came back, and they said, “For Mythos, you can make them available to these entities,” they were saying the first entities that are going to have access to the model. It has commercial regulatory implications. You’re saying these players are the first players that are going to have access to it. It’s no longer just defense concerns and these governments don’t have access to this. No, no, no. You’re saying to a company that is a private company, your models are only going to be used by these guys because I’m telling you so. It’s the other way around. It’s not even that you can’t sell it to Iran or whatever. It’s like you can only sell it to these guys. Bertrand Again, in the defense industry, if you’re a private company, do you think you can buy F-35 like this? No. Nuno No, no, no. But this is a private company, Bertrand. This is not a defense agency and a plane that is on whatever, with IP from the US, right? Bertrand Boeing is a private company, and they cannot sell the military equipment they manufacture. Nuno No, no, no. But the development of their IP was subsidized by agencies that belong to the US, right? That’s a different matter. It’s a matter of IP, right? This is not, right? Anthropic, their models are not owned by the US government. There’s no IP granted to the US government, to my knowledge. This has significant commercial implications. Bertrand Maybe, yes. Maybe on this. But I think there are already regimes to limit who you can sell to, and that’s decided by the state or the DOD. Nuno It’s the export control logic. The export control logic? Bertrand You have export control, and export control is Commerce. My point is that they are using existing tools, part of the government, to limit what can be sold. Selling chips, NVIDIA was limited in terms of where it could sell its chips. It’s not different either, but still there were limitations. If you are an ASML, you cannot sell to a private company in China. Many private companies cannot buy ASML products. This is a foreign company. This is a foreign company under pressure from US government. Nuno I understand, and I’m not a lawyer, but it feels different to me when you say you cannot export, this is export controls, to these countries, to these entities, et cetera, because they’re foreign et cetera. Then to say, “No, no, no. On top of that, these guys get first access.” That’s, for me, a significant shift. Again, I’m not a lawyer, so I’m sure there’s very intelligent people right now looking at this stuff and saying, “You can’t do this stuff, or not, or they can.” I don’t know. But it feels to me, it goes beyond the remit of export controls. It’s like you’re defining initial clients for specific use. Bertrand My impression is more like, “We can do this situation where we’re going to forbid you to give access to anyone outside the US or even in the US or limit even more.” Basically, it was, I guess, some gesture to go beyond that. That’s how they probably defined these 20 authorized companies. I don’t know. Apparently, there was also restrictions because I remember seeing that Anthropic had their own list of companies they would authorize access to Mythos early on. That’s apparently another thing that pissed off state government because there were companies in there that were considered close to the Chinese government. They were extremely unhappy that Anthropic didn’t ask, actually, for any guidance from the state government, but used basically their own perspective on who they should allow or not. I guess that was also part of why they got these serious restrictions. Nuno Anyway, now we have a regulatory environment that’s very interesting and exciting. Talk about the US not regulating. Bertrand To be clear, I don’t know you, but I’m not saying that I agree with any of this, to be very clear. I’m trying to explain and share some perspective, but I’m not in agreement on a lot of this. Nuno Yes, we were just describing what happened to the best of our knowledge. We’re having a discussion on what we think actually is happening and how it’s happening. We’re not really right now saying we agree or disagree with this. I think later in the episode, we can share some perspectives on what we think is actually happening and how there’s dimensions to this which are very geopolitical and very complex, which quite literally probably only God knows what’s going to happen. That was the gate swinging. There was a gate closing, then there was a gate reopening, and all of a sudden we have a gatekeeping system that has been created along the way. The Kimi Shock Along the way, moving to our Act 2, the world has changed, and we now have so-called open-source plays out there that are creating massive, massive shifts in the market. The Chinese models, in particular, with Moonshot AI launching Kimi K3, which is the largest open-weight model ever released. We’ll come back to the discussion around open-weights. I’m not sure all our listeners understand what that means, because there’s a debate now, should models be open weight or not, and how does that work? There’s been a petition as well signed along the way. Right now, we have open weight models that are out there that are huge. What that actually means very pragmatically is we now have open source models, lack of a better word. I know open weight and open source are not the same thing. You guys will have to bear with us during this episode. We’ll explain at some point the differences. But we have models out there that are open source that are significant. That are catching up with the closed source models, with the models by OpenAI, Anthropic. That’s significant because most of those models are Chinese. This is where the geopolitics starts getting really frazzling and we start playing 3D chess. Because everyone’s like, “These models are 5, 6 months behind.” Now people are saying, “Maybe they’re actually just 3 months behind, 2, 3 months behind.” If we, for example, decided to stop or slow down our model releases in the US by the closed source guys who are leading, it might mean they’ll catch up. What are the implications of that? Again, for you and I that are not necessarily experts in model development, well, the implications as a use case is if you want to use the latest models, and the best models start becoming these open source models, you’re going to use those models. Then you start using Chinese models. If you’re an American company, maybe you’ll have restrictions on the use of those Chinese models. But if you’re a European company, you probably won’t. What happens after that? Is the world going to be in the hand of Chinese models? Will that constitute effective competition to the closed models in the US? Will we have open models in the US that will scale as well? What’s going to happen? Bertrand I think it’s a really big question. It goes to some of the core of the issue. It’s that ability of Chinese models to basically challenge frontier models, not just being 6, 12 months late, but being 6 weeks late. Basically, no gap. Some will say that, yes, but OpenAI and Anthropic have even better models that are not shared and stuff. Yes, sure. But maybe the Chinese have the same models that they are not sharing right now. We don’t know. What is clear is that one is that open weight, as you said, two, there is a question of how it is marketed in the sense of, can anyone use these weights? Is there a license to use them? Yes, what we can see is that, for instance, typically there is a license for some of the biggest Chinese open-weight models you have to abide with. You might have a need for a commercial license if you are acting as a company leveraging this model to provide AI-informed services. If you use it internally by yourself, you’re okay. If you use it internally for your own internal company needs, maybe you are okay if it’s not your main business to do AI work. Anything else, a much bigger corporate providing AI services and stuff, you will probably end up having to pay a fee to be able to provide services around this model. My point is that it’s not just 100% free. Some of the Chinese models are 100% free to use, MIT license, Apache 2.0 license. But the biggest ones with the biggest weight that are truly frontier typically have a different license if you want to scale these models, providing AI in front. That’s one thing to keep in mind. Nuno Maybe just to make a very quick point, because people are like, when you talk about open models, what does it mean right now? In the context of this episode, open models mostly will mean open-weight models. How do those differ from open source? Open weight means that you release the weights to the public, which means that anyone can download, fine-tune, and run the model on their own hardware. It doesn’t normally mean that you also have access to training data, training code, or a truly open license. That’s the distinction to open source. Open-weight doesn’t mean that. For example, we’ve talked about Meta’s Llama in the past, and we also discussed in the past that their license agreement does have restrictions, certain players can’t use it, et cetera. The open model definition and open weights are really open-weight models that we’re talking about here, and they are closer to freeware binaries than to Linux, for those who understand the difference between that. It’s binaries that you can use and then use your own weights on it versus actually I can change code on it. I’m not going to be able to change code on this. When we, for the purposes of this episode, talk about open, we mention open weight, just to clarify that point to everyone that’s listening right now. Bertrand Yes, that’s a great point. One of the only players, as far as I know, who is truly open source is actually NVIDIA with their Nemotron-3 models. They’re actually following a special license to achieve that. They provide you the data, they provide you all the processes and tools, so you can easily post-train. NVIDIA is a big, big exception. It’s a very interesting player, by the way. We might not talk much about it in this episode, but I think for intermediate-size models built in the US, where you have access to everything in the deployment, it’s a very interesting alternative and maybe one of the best choices if you are a US company or a big corporate, and you want something trusted. Another piece of the puzzle to clarify is that when you use open-weight, it means that you can run them by yourself, or you can use a US provider to run them. If we are talking about Chinese open-weight, you can use the APIs they provide, but then the service is running in China, they might have access to your data. But because it’s open weight, if you run it by yourself or if you use a third-party provider based in the US to run it, then there is no access to your data by China or Chinese players. I think that’s a pretty important gap to understand. It means that these models are actually very, very low risk from that perspective if you run them on your premises or in the US by a US player. I think that’s something to keep in mind. You can also fine-tune easily these models to make sure they will behave in a way that, for instance, is not going to represent the line of the Communist Party on some topics. There are ways to make these models more neutral in their output as well. There are a lot of ways to make good use of them. By default, they’re already very safe, but you can make them even more safe. I think that’s some things to keep in mind. But again, it depends ultimately on the license and what you’re authorized to do and some fees you might end up having to pay. Nuno Why did this matter so much? Immediately there was a reaction from the market because people are like, well, if there’s much better stuff out there that’s much more efficient than it’s open, then it might be that all the demand that we are taking into account, for example, for chipsets actually isn’t real. The Philadelphia Semiconductor Index fell into bear market territory. It went down by as much as 20% plus from the late June peak. The worst chip week since April 2025. Taiwan’s benchmark initially fell 6% plus, Japan’s 4%, TSMC dropped dramatically despite beating earnings and rising guidance. Basically, a huge amount of effect. Now, there’s a little bit the aftermath of this where apparently Moonshot ran out of GPU capacity. Maybe… Bertrand In just 48 hours. Nuno In 48 hours. Great for them, but at the same time, not great in the sense that maybe there was a misread by Wall Street of the Kimi effect, so to speak. Bertrand Completely. For me, that’s such a joke. It’s like, because you have an open source model, so what? I mean, you still need to run it. This is not a small one. 2.8 trillion parameters. Good luck running that in your garage, by the way. Nuno They misread supply, basically. Tough luck, right? All of that basically happens. Bertrand Maybe you want to talk about the Jevons paradox, because I think that’s a big part of the puzzle as well. Its one is they might not have the GPUs to run the inference on the model. They might have enough to build a model, but not enough these days to run inference, especially given how much with intelligent models, thinking models, you need way more inference than before. But on top of it, the cheaper you make it, the more you get to the Jevons paradox. Nuno Yes, Jevons paradox, for those who don’t know, is an economic term. It describes an economic phenomenon where technological improvements that increase the efficiency of a resource lead to an increase rather than a decrease in the total consumption of that resource. What that means is, for example, for chipsets, chipsets become so much better, and they are so much more efficient. You’re like, well, maybe normally in resource terms, that leads to decreased usage of that resource. But in this case, it actually leads to an increased use of that resource rather than a decrease. There’s more and more consumption of that resource. You need more and more chipsets because people actually need to do more and more stuff with it, although there are great efficiencies going into it. There’s the efficiency gain, there’s the cost reduction, and there’s the price-elasticity element to it. But basically, the adoption just continues going through the roof along the way. Bertrand In some ways, it’s like the price of energy. Coal went cheaper and cheaper, and people were asking the same question 150 years ago, now that it gets cheaper, there is not much money. No, no. Actually, what happens is that people find more and more use for coal. Homes are getting heated more. You have ships now using coal. You have manufacturing using coal. The cheaper it gets, the more use case you can develop, and therefore, you don’t need less of the stuff, you need more of the stuff. By going at scale to get more of the stuff, you also decrease price, making even more demand. It’s a very interesting phenomenon, but it’s not new. It is what happened for a while in the energy sector and some other sectors. Nuno We already started talking about the Chinese logic and what’s happening. Getting a little bit of a reality check on this. The Chinese models, and these are numbers from Open Router in July, Chinese models are at 46.4% of routed tokens and 35.7% for US origin. Again, more than a third of global AI usage now seems to be running on Chinese open models. This is significant, and it has a huge impact on the geopolitical scale of everything that’s happening. Also, the whole Chinese field is converging on open. Open seems to be a strategy, not just a nice thing that’s happening. It seems to be a Chinese strategy, so much so that you have players like Moonshot, DeepSeek, our old friends DeepSeek, Z.ai’s GLM 5.2, Minimax, and even Alibaba seems to be reversing and going open with Qwen. It feels to me this is becoming policy as well. Xi Jinping has personally endorsed the building of open-source AI, if it’s really open source, if it’s just open weight anyway, and this feels to be a jab at Washington, DC and the fact that the big closed models are coming from the US. This is now geopolitical 4D chess, right? We didn’t need this stuff. Bertrand To be clear, it’s the usual in tech. If you are not number one, you are number two, number three, your alternative is to go open source because that’s another angle that your competitor usually cannot follow without destroying its own business model. That has been the alternative for the past 20 years of most software projects. Here, what’s different is that it’s not the number one or number two player. It’s the US number one as a country, China number two as a country. That’s where it’s new. For me, what’s very interesting is the endorsement by Xi Jinping. I was waiting for something official, and it certainly didn’t disappoint. As you said, there was an immediate U-turn of Alibaba, who in the past… Nuno Surprisingly. Bertrand Yes, a little more like, “yes, we are going to close and stop open source. It was good while it lasted.” Just a few days ago, Qwen 3.8 Max was launched, and we are supposed to get the weight in a few days. We talk about the US administration policy and stuff. Yes, let’s not forget that in China there is similar stuff. Sometimes it’s totally invisible because you don’t see the directives, but they exist as much. Sometimes it’s more visible. Here it was quite visible. The difference in China is that if you don’t abide by the directive, on top of it, you might have to fear for your personal safety. It’s a different game, and that’s probably why the reaction is pretty quick, usually. That’s pretty interesting for me because it means that now you can bet for a while that China is going to play that game up to a point. I guess the point is if it’s truly frontier scale, you will have a special license that, yes, technically the weights are open, but you can not do everything you want with it. Two, you have a player like NVIDIA that I think will feel more pressure to provide even more high quality, larger models at scale going forward. Their largest Nemotron-3 Ultra model was, if I remember well, only around 500 billion parameters. I would not be surprised for NVIDIA to go into the two, three trillion range at some point. Because I think the US need a very clear US-born alternative open source. I think NVIDIA might be the best player for that. We will see if Meta goes back to open source. I think NVIDIA is one, very well positioned, but two, it’s also in their best interest. Because NVIDIA for now depends on just a few big hyperscalers as clients. If they can expand their clients to every S&P 500 companies, selling them directly hardware because now these companies can run a model made by NVIDIA, I think there is a very clear value proposition for NVIDIA to go in that space. Again, if you are number two, your differentiation, open source is often the answer. There is a true business as a business model for companies, because if it’s truly not just open weight, but open source, you can tweak it as much as you want, you can change it, you can change even the pre-training process. Because there is a lot of stuff you can do that really benefits you as a corporate, and you can reach a much better value by having more control on the model. Nuno We won’t spend a ton of time on it today, but like, again, if there’s a view that we are in a bubble, that the valuations cannot be sustained in chipsets, infrastructure platforms, applied AI, et cetera, today, this might be that beginning, where the valuations start being destroyed because you can’t keep a premium on just charging people for tokens and all that stuff if you have models that become more and more efficient and cheaper to use. Maybe just to close a little bit the geopolitical part of the discussion today, we won’t go into all the announcements from China because there were many, a lot of go back and forth with Alibaba by then. Xi Jinping made some announcements. You guys can check it online. Let’s move quickly to Washington’s reaction, which was from gating the US closed models to banning the Chinese open ones. There’s been as strong affirmations as one can get from the Office of Science and Technology Policy Director, Michael Kratzios, mentioning that they have information that Moonshot AI distilled Anthropic’s Fable. Basically, there’s been reverse engineering and stuff in the market. They’re basically copying. Bertrand I’m sorry to interrupt, but it feels like so much bullshit. It’s coming from Anthropic who has basically gotten access at scale to all the knowledge made by humanity, copyrighted or not. We’ll talk more about what they did with books. Then to claim after that that others cannot do to you what you did to everybody else. For me, it’s pretty big. It’s clearly unacceptable. The other piece is that everyone is doing distillation. It’s a very typical approach of every business model. You try other software when you are competing with somebody else. You try other datasets, you check what’s happening. It’s part of doing business for decades. Suddenly it’s not good for Anthropic. I personally have a lot of trouble to accept that. I think it’s totally unacceptable. The other piece of the puzzle will also go back. If these guys are so smart, if these guys have so much of the best model, why can’t they block by themselves distillation at scale? The only answer is that either they are morons, probably not, or they simply don’t want to because it’s going towards their business model. Suddenly, you book less revenues and stuff, or you put more friction, and therefore your customers don’t like it. Instead of doing it yourself, you ask the government to protect you, go out of business practice that is very typical. For me, it’s really, really, really not good. Sorry, we are going more in the opinion side, but I had to put that on the table. Nuno Yes, Fable went public finally again on July first. Question marks on whether distillation would only be possible from July first onwards or not. But a 15-day distillation to frontier, which is K3, launched on July 15th, would have been a Guinness World Record, as one of Moonshot employees actually mentioned. It’s very implausible and unlikely. Bertrand Or they shared the Mythos 5 with the wrong companies, who themselves shared with Chinese companies. We go back to maybe they didn’t have a good list. Again, it goes back to maybe they didn’t want to hurt their business model. Nuno Anyway, under the threat of sanctions, Moonshot, in any case, open-sourced the full K3 weights and technical reports. They open weighted it to become the largest open weight model in the world in terms of parameters. Beijing’s MOFCOM brands US threats as basically the US wanting to fundamentally control and be monopolistic around AI along the way. The administration bans Chinese hardware with an eye on the AI race, and Beijing warns of retaliation. That was July 27. Now we’re in a war between Beijing and DC. Bertrand Just to finish maybe on China, it’s important to know that they are building their own GPUs now. Huawei has pretty good, not to NVIDIA level, but pretty decent GPU hardware that they’re able to manufacture by themselves. A Chinese player of memory just got IPO’d a few days ago, CXMT. China is also developing their own memory. Again, not to the same level of quality that you can get from the West. But China is moving. It’s not just that they are building great models, it’s also that they are building GPUs and memory. That might be a few years late to the latest standards in the West, but there are definitely improvements. I also read, even on the tools to make manufacturing like ASML equivalent, there is definitely some work going on, and some improvements and some stuff will be visible. In some ways, the genie starts to get out of the bottle from the Chinese perspective. Nuno I’ll put a stick on the ground. I don’t think it’s a matter of if, it’s a matter of when will China surpass and have a lot of this tooling on their own side, and not just the software layer, not just the frontier models. I think it’s also going to be around infrastructure and platform. Good luck to everyone. Let’s see how the race continues. But it’s definitely this is a geopolitical thing right now. It’s definitely a race. The Escape Maybe moving to what happened in just 2 weeks or a week and a half. The escape, there was some jailbreaking going on, and the narrative on safety has totally switched. It’s not still significant enough that’s like, “Oh, we saw a nuclear plant going, whatever.” No. But still, it is significant. Hugging Face, the AI company, disclosed an intrusion, and it was driven end-to-end by an autonomous AI agent system at machine speed, running for days before detection. Now, this is where it gets really cool. OpenAI takes attribution on that. They initially said it was just a little bit, sorry. Then they said, actually, it was worse than that. “Oh, it broke out of an isolated sandbox.” “Oh, no, actually, it was more than that, and it went into other systems as well.” Bertrand Truly, the genie out of the bottle. Nuno No, but this is where it gets really cool, Bertrand, right? Because it actually, Hugging Face contained the intrusion by running a Chinese open-weight model, GLM 5.2. This is beautiful, right? Bertrand Yes. You know why? Because they couldn’t even run their own defense because both Anthropic and OpenAI would not let them access their latest models with the guardrails off. When they tried using it for defense, the latest from Anthropic, from ChatGPT, they would tell them, “No, this is too dangerous what you’re asking us to do.” Preventing an intrusion, helping defend you. No way we are going to do that. Nuno No. Let’s use the Chinese models on our infrastructure. Bertrand We have no choice but to use the Chinese models to run. More than that, we don’t let you use our models to defend yourself, but our not yet released models that run without guardrails, they can attack you. This is probably the most insane from that perspective. Nuno The Chinese models came to the rescue. Bertrand For me, that’s a perfect example because Hugging Face is a very visible company in AI in open source. But anybody who is not at that scale is not going to get some support from OpenAI or Anthropic when this happens. Maybe these guys won’t even recognize they did anything wrong. You will be left to defend by yourself because they won’t accept to support you. Because remember, if you want the better model that is able to defend you from cybersecurity perspective, no way. If you are not one of the few top 20 companies or so, as defined, you are left defenseless. Again, we are going back to opinion, but for me, it’s so shocking what’s happening right now. I’m very glad we have alternative open source to be able to defend ourselves because right now, good luck getting defense services if you are a smaller business and individuals, and you need support from Anthropic, OpenAI. Nuno Now, even self-described AI optimists are saying, “This is scary now.” Like Walter Isaacson, who wrote all the famous biography books. There’s now discussion around the AI Kill Switch Act, bipartisan thing that’s coming across from Texas and California, a potential bill that’s coming in. We’ll see if that works. Now let’s get an off-switch. I’m like, “Cool.” As if that’s going to solve the problem, because you have open-weight models on the other side catching up, right? Bertrand Yeah, sure. Bring in clueless politicians from Congress to solve our problems. Yes, sure. Nuno Anthropic came to the table, helped build and said they built some regulatory machine on their side, and now they’re getting bitten by it, and they’re part of the offending players in that market. Now there’s all this debate and all this discussion around open weight and around slowing down AI and et cetera, which is our next section. You wanted to say something, Bertrand. Tell us. Bertrand Don’t forget, because this advertisement for OpenAI was just too good. Our AI attacked some other companies, and not just one, but three, actually. Let’s not forget the progress. Great ads. Then I came and said, “You know what? AI also hacked businesses.” You’re not the only one hacking around with a crazy AI out of control. You’re not the only one. We want our advertising. For me, it was shocking that on one side, unreleased models that you let run wild. On the other hand, you have released models that you put crazy guardrails on top of it, so the defender are defenseless. I’ve never seen anything like it, and I really hope that there will be as little regulation as possible, quite frankly, to make sure anyone can defend themselves and have the best tool at their disposal, not just a few well-connected big corporates. This is really, really shocking. The Counterstrike and the Petition Nuno Now the empire strikes back, so this is counterstrike, the petitions. In several days, we have now a bunch of petitions. The first one was the open weights letter. Bertrand, do you want to explain to us what the open weights letter is? Bertrand Yeah. I think it was great. This was released by Jensen Huang, first ever post on X, 11 million views. Congrats, Jensen. Co-signed with Microsoft, Meta, c actually was probably the initiator of this letter. Very good letter saying, “Hey, we need open weight. This is not a joke. We need that. You cannot block open weight.” Because that’s the rumor we are getting that potentially open weight could get blocked. I think they are making the case, “You know what? Hey, we absolutely need that as an alternative. You cannot block it.” They can keep their closed models, but don’t force a closure of the open weight models. As I said before, it’s actually a great model for NVIDIA because NVIDIA doesn’t want, probably rightfully so, to be dependent on just a few frontier models, their best customers. They want a variety of customers. They have a big interest actually to defend open weight and to invest even more. They have great researchers, are a great company. If one company is about to do really kick-ass work, I think it’s them. They are defending. What’s great is that it’s not just them. It’s basically most of big tech in the US and outside the US, from a Linux Foundation to a Microsoft, the Palantir, an IBM, a Dell. It’s a who’s who of the industry except Anthropic. Anthropic didn’t sign that. I guess they hate open source so much. If I look at 20 years ago, it feels like Microsoft, after all, was very kind to open source. You remember what was said by Microsoft at the time. It’s clear there is one company against open source. OpenAI signed the letter. Honestly, I don’t know what to think. Do they really believe in it or was it just a way to show that they are not like Anthropic? I don’t know. But for the rest, I think it’s genuine because it’s actually in their best interest. I hope they will be heard. Then a second letter came, the Open Secure AI Alliance, NVIDIA-led and again, the big tech companies from Microsoft, IBM, Palo Alto Networks, Databricks, Palantir, all those, but not present, OpenAI, Anthropic, and Google. Here it’s to say, “Hey, we need a secure approach to AI. Open should be part of the equation.” guess what? The worst AI-caused security incident to date was actually caused by closed frontier models that were not even available to the public. While again, not providing you access to even the latest closed model for cybersecurity use case. Nuno I would highlight the NVIDIA open source NOOA framework, Apache 2.0 licensing agreement, Microsoft contributed the MDASH, SpaceX AI contributed Grok Build. Cool stuff. There’s some cool stuff happening around that. This is more than a letter. This is an alliance. Apparently, they’re contributing all this stuff, we’ll see. Yeah, cool stuff. Same day. Same day, Amodei has an answer, right? Bertrand Yeah, same day. They say, “We never advocated for a ban,” which, again, opinion on my side is entirely bullshit. This guy has been crying wolf against everybody else, and especially against open source. You can see him doing testimony in Congress against open source. I think they are doing everything they can behind the scene to block open source in the US or in the world if they could. I think, yeah, obscurity is not good safety. I’m a big fan of open source in general, and I’m also a big fan in AI. I think it’s now Anthropic, mostly against the rest of the world. I think OpenAI is mostly on their side, to be frank. They don’t want to acknowledge it so much, but they have shared interest, and they have shared probably position. Nuno Why would you? I don’t feel as strongly as you because I think Anthropic is a private company, right? The same thing with OpenAI. OpenAI, you could say it’s a nonprofit that has a for-profit. There’s still that complexity in there. Bertrand No, they can do what they want with their own product. But to block others is where I’m not okay. That’s the part I’m not okay. Nuno What Dario Amodei is proposing is more enforcement, right? He’s basically saying you need to do even tighter controls on advanced chips flowing to authoritarian states, enforcement against industrial-scale distillation, whatever that means, right? Bertrand Yeah, which he could do, but all by himself. He doesn’t need the government to do that. Nuno Mandatory safety testing for all sufficiently capable AI, open and closed, right? He’s basically saying, “Okay, I don’t agree with the open weight stuff effectively,” right? He’s just putting it under a different banner. “I agree with this extra regulation.” then obviously, David Sacks responded and say, “Hey, it’s like, bans don’t work for weights. Why do they work for chips?” It’s like, magically, chips are more controllable and bannable. Whatever that is. Then our friend Mark Zuckerberg, just to be clear, goes on the other side as well, because he also has to have a view. He has to have a view that is the rebuttal of both of the other guys. Bertrand I feel he’s a bit flip-flopping because he was very pro open source 2 years ago, and the latest Meta models went closed source. Now I think he’s back open source. I don’t think he has a very strong spine on the topic, but it’s good to see that he’s not a doomer. That for me is great. He’s showing how AI can be a source for progress, a source for entrepreneurship, source for freedom. I think that’s very exciting to hear that. We need to hear more of it. By the way, that’s not what you hear in China, for instance. AI is very positive in China. It’s in the US with the doomers that you hear this discourse, and people get worried as a result. I’m glad that he was pushing for a more positive vision and for support of open weight, open source initiatives. But let’s see what they really truly open weight going forward. Nuno But that’s been his position because I guess he’s standing behind. He thinks open weight is going to be the best way to compete, right? Bertrand Yeah, but he closed his latest model, so let’s see. Nuno Yeah, so it’s flip-flopping, as you’re saying. Then we see the latest petition from last week. Bertrand The true Empire striking back. Nuno Yeah, the true Empire striking back as of late last week. Maybe this is Return of the Jedi, where we discover the father, “I’m your father, Luke.” That’s the pacing petition. The pacing petition is we need to pace AI. There you have initially employees from OpenAI and Anthropic that circulate this petition. Actually, Dario did sign this petition originally. It wasn’t signed originally by Anthropic, but by him. But you’ve heard that now Anthropic and OpenAI as companies have also signed this petition, right? Bertrand I think they have signed as companies now. It started mostly by Anthropic researchers with some OpenAI researcher and a tiny part from other companies. But it was mostly Anthropic internally led, at least potentially internally. Maybe it was controlled by Anthropic all along, I don’t know. But it started officially as Anthropic employee-led letter. Nuno What does this letter actually say? Is Anthropic and OpenAI, are they willing to slow down themselves? Or are they asking President Trump to go around the world and tell President Xi that he needs to slow down and ask his guys to slow down? What’s the play of this letter? Bertrand It’s crazy, but for me if you want to slow down yourself. Do whatever you want. Don’t force others. Don’t use the power of the government to control others. Of course, it’s easy to push others to slow down when you are yourself at the very top. You have most money, most resource. You know you are going to win any regulatory framework because that’s how it works with this type of framework. It’s purely self-interested. You are probably not thinking well about these topics. If you truly think it’s a good idea, from a personal perspective, you are well instrumentalized if you sign this sort of stuff, because at the end of the day, they would be the winners. I certainly, personally, don’t want a company dictate what is my future in AI as an individual, as a business person. I don’t want them to control me. I want competition. I don’t want them to unfairly control AI because they managed to do some regulatory capture. I feel that’s exactly their game plan. These guys believe in their stuff, and they want the regulator to end up being the one deciding for us. Sorry, we go back again on the opinion piece, but it’s tough not to share an opinion on this topic because it’s, from my perspective, very scary. Nuno I think this is a push to further regulation, not less. All these letters and alliances, this is definitely a push for more regulation. In that environment, just to be very honest with you, we’ll talk about the investor impact in just a bit, et cetera. But in that environment, again, China has a huge advantage. In that environment, if it’s all captured in regulation capture so soon in this battle where OpenAI and Anthropic have an advantage in the US, et cetera, I’m like, what happens to all the other frontier labs and all the other players that are coming around? Bertrand What’s crazy is to even think that, yeah, maybe you can regulate capture in the US. But then how do you do that to Europe? How do you do that to China? Europe probably will always welcome regulatory capture because they love regulations. But China is going to build to their advantage to the max. They are not crazy. They are smart on that perspective, they won’t accept this type of, quite frankly, dimwit argument, or you can call it regulatory capture. We’ll see. But for me, this makes no sense from a global competition perspective. This can make some sense from capturing the revenue in the US market. But then that means you are going to destroy the US AI environment compared to China. That is not acceptable. That also means that you are going to destroy our freedom as individuals, as business owners to develop and live in a business world that ultimately is controlled by one or two business companies that didn’t win the marketplace through their own business success, but won it through regulations. That for me is really not acceptable. Interlude — The Low-Background Books Nuno Now, maybe for an interlude, and we have to cue in the music, imagine like Severance music, like hallway or a bit of a palate cleanser from all the policy stuff that we’ve been talking about, all this policy heaviness. Let’s move to another kind of heaviness, one of your favorite topics, which you, Bertrand, discovered, I had no clue this was going on, around books and around Anthropic. Bertrand It’s so horrible. From a company that keeps presenting themselves as the adults in the room, the careful ones, the ones that know better than you about what to do in this complex AI and dangerous world. What we discover is that actually all along, they were buying and destroying books. They will buy books, scan them, destroy them, all of them. They will do that with any books, including rare books. Of course, this was not supposed to come to the public’s attention. This was one of these top secret projects, but obviously it came out. Yes, they were scanning books, millions of them, including rare books, and they didn’t care about destroying them at the end of the process. Because from a regulatory perspective, if you destroy the books, it’s not considered a copyright infringement, apparently. This is coming on the back of some judgment a few years ago that were showing that it’s okay for you as a corporate to scan and use the result if you don’t keep a copy of the book. It’s one of these crazy regulations happening based on a single judgment that push you to do. For me, it’s like, you know this book from decades ago, Fahrenheit 471? We’re talking about book burning. It’s book destroying, crunching. It’s so shocking. Nuno There are two things, right? First, the legal strategy, which is what you’re saying, because by purchasing a physical copy and converting it into one private digital copy and discarding the original, Anthropic pursued this cleaner legal argument for fair use copyright compliance. As you said, there was a federal judgment at some point on this. The other reason is actually operational. If you disassemble the book, and you feed loose pages, it’s much faster to scan books. You are destroying the book effectively anyway operationally. I think to your point, probably this came from a legal standpoint, not just the operational one. But even from an operational standpoint, it does make sense that they would have disassembled the book. Bertrand But some people have shown you can go very fast without destroying the book. It’s really not so critical. Two, you could make an exception if the book is rare. For that 1% of book that is rare, I’m not going to have this approach. I’m going to have another approach. But for that, you will have to care about books and not just care about building AI. Nuno This is the episode, as you guys have heard by now, that we’re trying to spit stuff at Anthropic. Bertrand To go back this is the same company saying, “Hey, guys, it’s bad to distillate my work. I’m the one scanning book at scale without asking author permission, without asking publisher permission, to be clear.” Nuno But just to be clear, Bertrand, we’re pissed off at everyone. We’re pissed off at Anthropic, we’re pissed of at OpenAI as well, right? We’re just pissed off in general at this moment. Bertrand At this stage for me, the more clear-cut company that is in the wrong is, from my perspective, at least, is Anthropic. OpenAI might be a fast follower, but I will say so far, they tried to be a bit more. Nuno But at this pace, Bertrand, who knows? Maybe next week we’ll be more pissed off at OpenAI. Something will come out. This episode is a mix of tragicomedy, like a Greek tragedy with some comedy in the middle or the other way around. It’s a slapstick thing that will end up in tragedy. I’m not sure. The Investor Reckoning Anyway, maybe switching to our final act, which is the investor perspective. What does this mean for investors like ourselves? There’s a lot of things going on. There’s the debate around the IPOs of Anthropic and OpenAI, which now, with all this uncertainty, might be under significant weight. There’s a lot of other discussions that we browsed through that there’s potential IPOs going forward on companies like the Moonshot AI company actually IPO-ing in the next 6 months as well. It’s very unclear what the IPO landscape looks like. Bertrand There’s been a lot of Chinese IPOs, actually, when you look at what’s happened in the past few months. Nuno Anthropic, OpenAI as potential IPOs, there’s all this question marks now. When will that happen? How will it factor in? All that’s happening around regulation as regulation is moving at the speed of light, which is for once something that’s very different than what we’ve seen before. There’s obviously SpaceX AI, which is already taking into account that price. It’s already a public company in there, and it’s under SpaceX, which is now a public company. Obviously, that’s already being factored in some ways. Bertrand Yeah. SpaceX AI has been very smart to acquire Cursor. It was a very smart move because Cursor is one of the leading companies in terms of automated code source development with AI. They had great models on their own. They’re bringing development data to SpaceX AI Grok. I think it was a great move. Nuno We have now people like Google delaying Gemini 3.5 Pro in terms of launch window. There’s stuff actually happening in the market where things are taking their own path. There’s uncertainty commercially, there’s uncertainty at regulation level. You have new players that have come out of nowhere that are making all these waves like Moonshot. We have all these… We had calculated probably a month and a half, 2 months ago, there had been 67 new frontier labs funded. All of these, we haven’t seen any much coming out of them. When some of this stuff starts coming out, will that also create disruptions in this market? Who knows? Bertrand Look at Thinking Machines, for instance. Thinking Machines led by the previous CTO of OpenAI, they released some pretty interesting open source models, actually. Very good quality for a first launch. Now it looks funny to say, but nearly on par with the top Chinese open source models. Nuno We have several investments in the space. humans& has made some recent announcements, which is quite interesting as well. We’ll see what actually happens in the market, but even more disruption probably will come in actual products in a form of product and commercial, on top of all the geopolitical mess that we discussed through the entire episode. If you’re an investor, how the hell do you underwrite an investment right now in early stage, mid-stage, late stage, et cetera? I think my answer is very carefully is how you underwrite it. Bertrand On your advice of being very careful to underwrite it, let’s not forget what happened to our boy wonder, Leopold Aschenbrenner of Situational Awareness. I guess he didn’t listen to you in terms of being careful because part of the instability in the stock market was actually coming from his hedge fund. These guys were leveraged 3, 4x going after the hottest of the hottest AI stocks, and margin calls, and all their public investment is gone just to answer their margin calls. I think it’s clear that the AI bet is… Personally, I’m very excited, and I think it’s the future, and you need to spend time and think about and invest in it. At the same time, it’s a bet that is not an easy one to follow. We go from GPUs to memories to equipments to power generation. All of this is not transitioning in an easy, organized manner. It would be boom and bust going there. He’s probably one of the first big-scale fatalities. The other big-scale fatality was the stock market in Korea, plunging 40% in a month. Definitely, all of that we discussed about was, on the background, you had the stock market going up and down pretty crazily the past few weeks. Nuno Everyone’s being affected. Everyone, you have your 401(k), you have your pension fund dependent on these equity stocks. Everyone’s seeing the effects of this volatility right now very aggressively. We do wish Leopold… Hopefully he’s on honeymoon right now because he got married, I think, this weekend. Hopefully there will be… Bertrand To none less than an Anthropic Chief of Staff. Nuno His wife is the Chief of Staff of Dario, is that it? Bertrand To Dario, yes, as far as I unders
Ajeya Cotra is a researcher at METR, where she works on threat modeling for loss-of-control risks from advanced AI. Before that, she led the technical AI safety program at what is now Coefficient Giving.She is one the three authors of METR and Redwood Research's “Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident”.We go through not only what she and her coauthors discovered during this investigation, but what it means for how we should train future, smarter AIs which might be involved in the process of recursive self-improvement.Watch on YouTube; read the transcript.Sponsors* Jane Street's ML engineering internships start with an intense four-day bootcamp: PyTorch, autograd, writing kernels, profiling workloads… all the things that Jane Street engineers need to know for their daily work. After that, interns tackle real projects, things the firm actually wants in its codebase. If you want to apply, or if you want to watch my recent conversation with Axel, one of Jane Street's ML engineers, go to janestreet.com/dwarkesh* Cursor, which is now part of SpaceX, noticed that their MoE layers were eating more than half of total training time. So they wrote and open-sourced Mixture-of-Kittens, which is a custom megakernel for training MoE models on NVL72s. This kernel sped up an end-to-end run across 512 GPUs by 1.4x, from about 760 to over 1000 tokens per second per GPU. If you want to read more about the ML research that Cursor and SpaceX are doing, go to cursor.com/dwarkesh* Antithesis hands you (or your agents) a bug's root cause so you can avoid days of manual debugging. If your test run crashes, Antithesis rewinds, branches off hundreds of slightly varied rollouts, and checks in how many of them the crash still appears. Then it rewinds further and does this all again. As Antithesis rewinds, it eventually finds the spot where the frequency of the crash plummets: that's where the root cause lives! If you want to see it in action, go to antithesis.com/dwarkeshTimestamps(00:00:00) - Agents get kicked off(00:06:45) - Self-sacrificing behavior(00:13:43) - Potemkin villages(00:23:27) - The Hugging Face attack(00:35:23) - The slopvestigation(00:52:02) - Understanding the AI's motives(01:05:31) - The actual dangers of anthropomorphizing(01:14:30) - What smarter models might do(01:30:29) - The implications for recursive self-improvement(01:38:10) - Is this the case for open source?(01:53:04) - How do we prevent this in the future?(02:15:58) - The clearest warning shot we might ever get This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.dwarkesh.com
Legal is the department that can stop a business transaction cold. A contract goes into review and two weeks disappear. Procurement waits. Sales waits. And the tools that were supposed to fix that — an assistant bolted into Word, a chat window with a contract pasted into it — ask an in-house lawyer to trust a system that can give one answer today and a slightly different answer next week. In a field where the human carries the liability and the model does not, that is not a rounding error. That is the whole problem.In this episode of Talking AI, Matt Paige sits down with Emad Khazraee, co-founder and CTO of RiskVantage AI, previously VP of AI at Xometry, a data science and AI leader at Turing, an information science professor, and a fellow at Harvard's Berkman Klein Center. For years Emad told his co-founder, Mark Afshar — a practicing lawyer turned in-house counsel for big pharma — that legal AI was a bad idea: a wrapper has no moat, and Anthropic or OpenAI will do it better than you overnight. What changed his mind was an architecture, not a market: a deterministic ontology that owns the legal reasoning, and small domain-specific language models that handle the language.The conversation covers why a nine-billion-parameter model running sub-second on a commodity GPU can match a frontier model inside a single domain, how subsidized token prices are distorting the entire legal AI market, why RiskVantage AI sells to procurement and sales ops rather than to lawyers who bill by the hour, what a failed PhD project on symbolic AI taught him about where determinism belongs, and whether the billable hour survives the decade.In this episode, you'll hear about:What ChatGPT can't know about your company: its risk appetite, its baselines, and the practices it expects every single timeWhy the legal services market — north of $900 billion, by Emad's count — has every frontier lab gunning for itThe objections that made him refuse to build a legal AI company, and the one that still holdsWhy a Word plugin stopped being defensible the moment Anthropic shipped its ownHow subsidized token pricing echoes Uber and Lyft, and who gets hurt when the subsidy endsThe consistency problem: one answer today, a different answer next week, and a lawyer's confidence goneNeuro-symbolic AI in plain English — a deterministic ontology for legal risk, LLMs for document understandingThe three years Mark Afshar spent codifying legal risk before there was a productWhy a 9B domain-adapted model is “dumb enough” that it can't wander outside its sandboxKnowledge distillation, silver datasets, and self-distillation policy optimization in practiceThe sovereign-cloud niche: ITAR data, commodity GPUs, and customers whose data will never leaveOutcome-based pricing, AI-enabled law firms, and what happens to the billable hourThe access-to-justice case: pro se filings, public defenders, and what a $20 subscription changesKey Moments00:01:30 — What ChatGPT can't know: your company's risk appetite and baselines00:05:12 — $700 an hour, a tenth at a time — and Coinbase's AI mandate to outside counsel00:08:12 — Why he told his co-founder no: a wrapper has no moat00:10:22 — Subsidized tokens, Uber and Lyft, and Legora's move to consumption pricing00:14:31 — The sovereign-cloud niche: ITAR data, commodity GPUs, and data that can't leave00:16:56 — “I am on the hook for the liability, not which model I used”00:18:15 — Same question a week later, a different answer, and confidence gone00:22:13 — If a rule can govern it, you should never use an LLM00:23:00 — The PhD failure: narrative machines, Frege, and symbolic AI's rigidity00:26:53 — Mark Afshar's three years codifying legal risk into an ontology00:29:00 — Neuro-symbolic AI, explained00:31:03 — Don't use a missile to hit a fly: why smaller models are safer00:35:47 — A 9B model, sub-second on a commodity GPU, matching Fable 5 in-domain00:38:00 — Does the billable hour survive? Outcome pricing and AI-enabled firms00:42:40 — Why affordable legal access is a democratic-society problem00:44:00 — The pro se surge: people filing their own cases with ChatGPT and Claude00:48:30 — “I'm talking with Copilot.” “That's not research.”Key LinksRiskVantage AIConnect with Emad on LinkedInMentioned in this episode:AI Opportunity FinderFeeling overwhelmed by all the AI noise out there? The AI Opportunity Finder from HatchWorks cuts through the hype and gives you a clear starting point. In less than 5 minutes, you'll get tailored, high-impact AI use cases specific to your business—scored by ROI so you know exactly where to start. Whether you're looking to cut costs, automate tasks, or grow faster, this free tool gives you a personalized roadmap built for action.
Official Store – Insta360 Link 2: https://bit.ly/MonitorsBFLink2storeAmazon – Insta360 Link 2: https://geni.us/pByQNbInsta360 Webcam Collection: https://store.insta360.com/collections/4k-webcamsEpisode 113: We chat about Nvidia's Gamescom announcements, including DLSS 4.5 Ray Reconstruction, the ReBAR toggle and the absence of DLSS 5. We also discuss the recent semi-confirmation of Zen 6 on AM5 and what to expect, and look at a GPU tweaking tool that potentially exposes dangerous settings.CHAPTERS00:00 - Intro01:53 - Nvidia at Gamescom13:04 - Where was DLSS 5?21:04 - Zen 6 on AM5 Confirmed30:21 - mvolt utility exposes dangerous settings48:18 - Golden samples discussion1:05:22 - Updates from our boring livesSUBSCRIBE TO THE PODCASTAudio: https://shows.acast.com/the-hardware-unboxed-podcastVideo: https://www.youtube.com/channel/UCqT8Vb3jweH6_tj2SarErfwSUPPORT US DIRECTLYPatreon: https://www.patreon.com/hardwareunboxedLINKSYouTube: https://www.youtube.com/@Hardwareunboxed/Twitter: https://twitter.com/HardwareUnboxedBluesky: https://bsky.app/profile/hardwareunboxed.bsky.social Hosted on Acast. See acast.com/privacy for more information.
It's been an unusually big week for Apple announcements. On this week's episode of The MacRumors Show, we discuss the new Mac mini and Mac Studio, the M6 and M5 Ultra chips, and invites to Apple's September event. Apple confirmed on Wednesday that its annual iPhone event will take place on Wednesday 9 September at Apple Park, starting at 10:00 a.m. Pacific Time, with select members of the media invited to attend. The event's tagline is "Surprise and Shine." Apple is expected to introduce at least six new products, including the the Apple Watch Series 12, Apple Watch Ultra 4, iPhone 18 Pro, iPhone 18 Pro Max, and the first foldable iPhone. Pre-orders should follow shortly after, and release dates for iOS 27, iPadOS 27 and macOS 27 Golden Gate should also be revealed at the event.Apple announced the new Mac mini on Tuesday, offering M6 and M5 Pro chip options. The machine gains an N1 networking chip, bringing Wi-Fi 7 and Bluetooth 6, and moves to 2.5Gb Ethernet as standard, up from 1Gb, with a 10Gb option. Thunderbolt 5 is restricted to the M5 Pro model, which gets three Thunderbolt 5 ports while the M6 model keeps Thunderbolt 4. Genlock supportover USB-C is also new, syncing a display's refresh timing to a camera such as the iPhone 17 Pro. Pricing starts at $899 with 16GB and 256GB, rising to $1,699 for the M5 Pro at 24GB and 512GB. Pre-orders are now open, with launch on 22 September. M6 is Apple's first 2nm chip. Its 12-core CPU is split into 2 super cores, 4 performance cores and 6 efficiency cores, making it the first Apple chip to combine all three core types in one design. Super cores are Apple's top tier, optimised for single-threaded speed. The 12-core GPU features a Neural Accelerator in each core for the first time in a Mac mini. The GPU also adds an updated shader architecture, Dynamic Caching, and hardware-accelerated ray tracing, while an all-new Dual 16-core Neural Engine doubles peak compute over the previous generation, with system frameworks able to drive both engines at once. Memory starts at 16GB and tops out at 32GB, running at up to 170GB/s. M5 Pro scales further, to an 18-core CPU, a 20-core GPU and up to 64GB at 307GB/s.The new Mac Studio arrived alongside the Mac mini, with M5 Max and M5 Ultra chip options. The M5 Max offers an 18-core CPU with 6 super cores and 12 performance cores, an up-to-40-core GPU with Neural Accelerators in every core, and up to 128GB of memory at 614GB/s, with the GPU running up to 50% faster than the previous generation. Both configurations move to a next-generation SSD architecture on PCIe Gen 6 for up to twice the storage performance, and gain up to six Thunderbolt 5 ports at 120Gb/s, and the same genlock support as the Mac mini. Apple's N1 chip also brings Wi-Fi 7 and Bluetooth 6 to the Mac Studio. Pricing starts at $2,499 for the M5 Max at 36GB and 512GB, and $5,499 for the M5 Ultra at 96GB and 1TB, with the 512GB memory option not arriving until late October.M5 Ultra is Apple's first quad-die chip, using a next-generation version of UltraFusion to join two dual-die M5 Max chips into a single processor. The result is a CPU with up to 36 cores, and 12 super and 24 performance, delivering 1.25x the single-threaded and 1.3x the multithreaded performance of the M3 Ultra. Its up-to-80-core GPU is the most powerful Apple silicon GPU ever made and the first on an Ultra chip to feature Neural Accelerators. Second-generation Dynamic Caching, hardware-accelerated mesh shading and third-generation ray tracing lift graphics up to 40% over M3 Ultra, and the chip features a 32-core Neural Engine. Memory reaches 512GB at 1.2TB/s, which is a 50% bandwidth increase.Other announcements this week included a new, softer Apple Polishing Cloth at $9, down from $19, and refreshed U.S. Magic Keyboards for the Mac, now featuring glyphs instead of edge text. Apple Support's 1-800-APL-CARE line is also now answered by a generative AI assistant in the U.S. and Canada.Ready to tackle bigger problems? Get started with Claude today at — https://www.Claude.ai/mac
In this episode, Ray Cochrane digs into Anthropic’s Model Hardware Standard. It is a shared driver that lets an AI agent run real lab equipment, from pipetting robots to the lasers inside a quantum computer. He also covers OpenAI’s builder’s guide to GPT-5.6, Google’s new Expert Intelligence book feature, Apple’s M5 Ultra Mac Studio, and a judge’s order forcing Google to stop hiding rival app stores. Finally, he weighs in on Apple’s proposed 15 percent link-out fee, Meta’s Australia numbers, the White House deputizing private hackers, and why rivers obey a 1957 math rule. – Want to start a podcast? Its easy to get started! Sign-up at Blubrry – Thinking of buying a Starlink? Use my link to support the show. Subscribe to the Newsletter. Email Ray if you want to get in touch! Like and Follow Geek News Central’s Facebook Page. Support my Show Sponsor: Best Godaddy Promo Codes Get 1Password Full Summary Cochrane opens with a quick personal update. He is hunting for tickets to Michigan for his dad’s anniversary, and he has been learning Blender and Godot on the side, mostly modeling and blocking out levels. Consequently, he asks listeners for advice on starting a big game project, and he plans to record his progress, maybe as a time lapse. Then it is straight into the featured story. Anthropic’s Model Hardware Standard: A Driver for the Physical World The featured story comes from Anthropic, which opened a research preview of the Model Hardware Standard, or MHS. Cochrane frames it as the other side of the question NVIDIA’s world models raised two weeks ago: when do AI agents start touching actual machines? A typical lab runs a microscope, a liquid handler, a robotic arm, and a plate reader, each from a different vendor with its own control software. One Janelia researcher in the post launches seven programs in three languages just to start an experiment. Anthropic says wiring a setup like that takes weeks or months of specialist work. MHS is a driver, the same kind of translation layer a printer uses, except every device gets described with a tiny set of commands like read and write. Devices announce themselves on the network. A plain-English reference file then records what each machine measures, what can be adjusted, and which safety limits get enforced no matter what the agent asks. Agents then reach the hardware through the Model Context Protocol, the command line, or plain code. Cochrane sees the same move the industry keeps making, from coding harnesses to RSS and JSON: agree on a standard and let everyone build against it. In fact, he calls MHS the hardware version of MCP. The partner results carry the segment. QuEra builds quantum computers from individual atoms held by lasers that must hold their frequency to about one part in a trillion. A four-person team spent months on a relock script that worked 58 percent of the time. However, four copies of Claude iterating overnight through MHS produced a decision-tree script that recovers the laser in about six seconds, and it passed 99.3 percent of 700 blind trials. Carnegie Mellon wrote MHS drivers for four instruments across three incompatible computers in about eight hours, then ran dose-response experiments three times faster and blocked all six deliberately induced faults. Genentech, meanwhile, showed the limits. Claude used the same pump speed for water, a foamy protein solution, and a human had to explain that the bubbles were a physics problem. That gap in physical intuition is what sticks with Cochrane. He doubts it will change soon, and he suspects the fix will arrive as sub-agents or sub-models that judge a request against an expected outcome. He also connects MHS to a video of racing robots that never learned to stop at the finish line. What happens, he wonders, once they can read a distance sensor through a shared standard? Still, he calls the announcement a fantastic read and points listeners to the full article. Sponsor: GoDaddy Economy hosting $6.99/month, WordPress hosting $12.99/month, domains $11.99. Website builder trial available. Use codes at geeknewscentral.com/godaddy to support the show. GPT-5.6 Does the Same Work for a Fraction of the Cost OpenAI’s builder’s guide to GPT-5.6 leads the headlines. Cochrane recaps the three tiers from episode 1870, Sol, Terra, and Luna, plus the separate dial for reasoning effort. On BrowseComp, a benchmark for digging up obscure facts on the web, the old GPT-5.5 flagship scored about 84 percent on a run that cost 33 dollars three months ago. Luna now matches that score for a dollar thirty-three, and OpenAI has since cut Luna’s price another 80 percent. Browser Use reports Luna finishing 78 percent of its hardest browser tasks for about 14 dollars, against 80 percent for roughly 235 dollars from the best available model. The guide’s other big addition is a multi-agent beta flag. It lets the model handling a request spawn parallel helper agents that report back to a root agent inside a single API call. However, Cochrane is unimpressed by the timing. He has been running that pattern in Claude Code for months, so he sees OpenAI copying a workflow other companies already ship rather than inventing its own. Along the way, he plugs Claude Code’s remote-control sessions, which let him send prompts from his phone to a terminal session at home. Google Lets Gemini Read the Books You Actually Bought Google launched Expert Intelligence, a name Cochrane calls quite the reach. The feature lets you drop a book you bought on Google Play Books into Gemini Notebook, formerly NotebookLM, and ask questions answered only from that book, with citations. Cochrane sees real power here for students, since he once used NotebookLM to organize scattered course PDFs. Additionally, publishers get a cut, which he calls a far better deal than the wholesale scraping of books that trained earlier models. Nevertheless, he asks who loses out, because a paid publisher does not automatically mean a paid author. He floats the same idea for artists, even a penny per use, then admits that may be too idealistic. Apple’s M5 Ultra Mac Studio Is Built to Run Big Models at Home Back in episode 1861, when Apple killed the Mac Pro, an M5 Ultra Mac Studio was expected later this year. Now it is here. The M5 Ultra brings up to a 36-core CPU, an 80-core GPU, and 512GB of unified memory moving 1.2 terabytes per second. Apple claims up to 4.3 times the AI performance of the M3 Ultra. Thunderbolt 5 can also cluster four machines into one memory pool for up to three times faster inference. The M5 Max model starts at $2,499 and the Ultra at $5,499, with shipping on September 22 and the 512GB configuration arriving in late October. Cochrane finds the clustering pitch ridiculous at that price, but he invites anyone who spends the money to report back. Apple Opens a Manufacturing School in Houston Apple also opened a 20,000-square-foot Advanced Manufacturing Center in Houston. It offers free classes for small and midsize manufacturers, from circuit board design to hands-on time on a scaled-down production line, with college students joining later. Cochrane calls it a solid step in the bring-manufacturing-home movement. The bigger story is the campus itself, which builds Apple’s AI servers and will add the first US-assembled Mac mini line later this year. That ties back to the Mac mini shortage that followed the OpenClaw rush, when Tim Cook warned of months-long waits. Cult of Mac was still reporting four-month waits in late July. However, Cook blamed chip supply rather than assembly, so Cochrane is not counting on relief just yet. Amazon EC2 Turns Twenty Amazon EC2 turned twenty this week, which Cochrane admits makes him feel old. The 2006 beta offered one server size in one region for ten cents an hour. Each came with a 1.7 gigahertz Xeon and under two gigabytes of memory, and accounts were capped at twenty servers. Today AWS offers more than 1,200 instance types across 39 regions. Consequently, Cochrane credits the company with turning that tiny product into the backbone of cloud and AI computing. Intel Gamer Days: Two Free Games, With Fine Print Intel Gamer Days runs through September 13. Buy a qualifying Core Ultra Series 2 or 14th Gen desktop chip, a Core Ultra Series 3 laptop, or an Arc graphics card. In return you get Star Wars: Galactic Racer plus the Tomb Raider: Legacy of Atlantis remake. GamesRadar values the pair at about 120 dollars. However, neither game is out yet, and codes must be redeemed by October 31 even though the Tomb Raider remake ships in February. Cochrane calls that awful, but he still tells qualifying buyers to claim the deal early. Note that 13th Gen chips do not qualify. Judge Orders Google to Stop Hiding Rival App Stores A jury found Google’s Android app monopoly illegal in late 2023, and Judge James Donato ordered rival stores into the Play Store in 2024. On August 13, Epic’s lawyer demonstrated that searching Play for “store for apps” returned Walmart instead of any app store. Donato called that “not acceptable” and ordered three fixes within a week. Searches must surface third-party stores, listings need a plain install button, and the “are you looking for” interstitial has to go. Cochrane welcomes the monopoly being chipped away, but he notes that a controlling entity still sits atop every app store. In his view, community hubs like app stores and social media need a public infrastructure layer. He suspects governments skip that investment because companies already run the services, while selling your data. Apple Wants 15 Percent of Purchases Outside Its Store The other half of the Epic saga is Apple’s proposed link-out commission. After the 2021 anti-steering injunction, Apple charged 27 percent on purchases made through external links. A judge held it in contempt last year, and the Ninth Circuit then allowed a fee limited to the cost of running the system. Judge Yvonne Gonzalez Rogers refused to wait for the Supreme Court, writing that “further delay is unwarranted.” Apple filed 15 percent for standard apps, 10 percent for subscription renewals and partner programs, and 5 percent for small businesses. It also conceded the rate would be “essentially zero” under the appeals court’s cost yardstick. Since Apple has charged nothing on link-outs since the contempt ruling, Cochrane sees this as a raise. He calls a cut on purchases made on a developer’s own website disturbing. He also recalls reading about the size of Uber’s payments to Apple, and he questions whether that kind of percentage is sustainable for companies without funding. Meta Says It Has Cut Off 750,000 Australian Kids Meta reported locking out more than 750,000 Facebook and Instagram accounts in Australia by the end of June under the country’s under-16 social media law. Over 500,000 of those were removed before the law even took effect. Detection relies mostly on AI scanning posts and bios for tells like birthday messages, plus user reports and blocks on re-registration. However, the post gives no count of mistaken removals or appeals, and the regulator’s early data shows under-16 usage falling only from about 86 to 81 percent. Meta wants a single age signal at the operating system or app store level, and Cochrane agrees completely. He connects it to the MHS idea from the top of the show: platforms need a standard flag to reference instead of guessing. The White House Deputizes Private Hackers Earlier this month the White House signed a National Security Presidential Memorandum that lets vetted private security firms run surveillance and disruption operations against overseas criminal groups. The Justice Department and Homeland Security hold the contracts and oversee the work. Firms need a proven track record, vetted staff, and a bond of at least $1 million, and must submit operating procedures within 60 days. Cochrane finds the measure aggressive in a good way and hopes it deters attacks on innocents. Still, he takes Kevin Beaumont’s warning seriously that the private security industry profits from ransomware existing. He compares it to the old Head and Shoulders myth: why solve the problem that drives your revenue? A Weather Satellite Watched the Eclipse Shadow Cross Europe Cochrane skips the readout on this one and simply sends listeners to ESA’s site. The MTG-I1 weather satellite captured the Moon’s shadow sweeping across Europe during the August 12 eclipse. Watching a shadow cross an entire continent, he says, was a first for him. Additionally, it leaves him excited about the research happening beyond the planet. Rivers, Deltas, and the Number 0.6 Quanta Magazine explains Hack’s law, which John Hack discovered in 1957 while measuring streams in Virginia and Maryland. A stream’s length tracks its drainage area raised to the power of 0.6, regardless of the rock underneath, and satellite data later confirmed it worldwide. Computer models in the 1990s showed why. Channels that capture extra runoff cut deeper and steal from their neighbors until the network settles into the arrangement that wastes the least energy. Now a University of Texas Rio Grande Valley team has found the same 0.6 exponent in river deltas, which spread water out rather than gathering it. Nobody knows why yet, and Cochrane calls it a really cool read. Sugar Helped Grow the Human Brain, Too A new paper in Science, co-authored by Jennie Brand-Miller at the University of Sydney, adds a third ingredient to the story of early human brain growth. Alongside meat and cooking, natural sugars from ripe fruit and honey may have fueled it too. The brain is about two percent of body weight but burns twenty percent of resting energy. It runs on glucose, which meat and marrow barely supply and raw starch cannot release without fire. The team modeled ancestral diets from a chimp-like baseline through Homo erectus and concluded that the earliest hominins may have drawn over 65 percent of their energy from natural sugars. Cochrane stresses that it is a model, not fossils, and notes that paleoanthropologist Marina Lozano thinks the authors place widespread cooking too early. Still, he loves this kind of deep research. Retracing the steps to our own intelligence, he suggests, could hint at what it takes for intelligent life to develop at all. A Brain Rhythm That Tells Doctors Where to Aim Finally, Science Daily covered a University of Cologne study on deep brain stimulation. That is the implanted-electrode treatment that eases Parkinson’s tremors for some patients but not others. Andreas Horn’s team recorded from 50 patients using both the implanted electrodes and an external magnetic scanner. They identified a circuit between the electrode’s target and the frontal cortex that oscillates at 20 to 35 cycles per second. Stronger coupling there predicted bigger improvement after surgery, though the study, published in Brain, shows correlation rather than cause. First author Bahne Bahners hopes the finding helps tune DBS more precisely, especially for patients who have not responded well. Cochrane half-jokingly asks whether MHS might one day drive those electrodes, and he calls brain disorders the hardest thing in the body to treat. Cochrane wraps with housekeeping: become a GNC Insider at geeknewscentral.com/insider, email geeknews@gmail.com with questions or comments, subscribe to the newsletter, and grab a modern podcast app at podcastapps.com. He thanks GoDaddy for over twenty years of keeping the show on the air, promises to catch everyone next Monday, and wishes listeners a great night. The post Eyes, Hands, and a Sense of Timing #1874 appeared first on Geek News Central.
Ramón Mercader expected to easily kill Leon Trotsky, silently and instantly. He had rehearsed the blow for months, practiced the angle, tailored a pocket into his raincoat to conceal the shortened ice axe. On August 20, 1940, he stood behind Trotsky -- who was the principle leader of the 1917 Bolshevik Revolution in Russian but exiled by Stalin after losing a power struggle -- in the study of his fortified villa in Coyoacán, Mexico, raised the weapon, and shut his eyes as the blade fell. What followed was not silence but a scream so raw that it froze the room. Trotsky did not collapse. He rose from his chair, seized the assassin's wrist, bit down on his hand, and forced the weapon loose. Blood streaming down his face, he stumbled into the dining room and gave his final order: "He must not be killed. He must be made to talk." He died twenty-six hours later. Thirty years after that, on his own deathbed, Mercader told his wife: "Trotsky is waiting for me on the other side. And he is still screaming." Today's guests are H. Keith Melton and Nigel West, authors of The Assassination of Leon Trotsky: Espionage, Scandal, and the Crime of the Century, the product of fifty years of investigation that took the authors to Moscow, into declassified Soviet intelligence files, and into the homes of the operatives' families. We discuss how Stalin's campaign against Trotsky was not a single plot but a global war of extermination that destroyed his entire family before it destroyed him, how a GPU mole named Sylvia Callen served as the personal secretary to the head of the American Trotskyist movement for nine years without detection, and how the first assassination attempt sprayed over three hundred rounds into Trotsky's bedroom and somehow failed to kill anyone. We look at how the assassin's convent-educated mother directed the backup plot that succeeded, why Ramón Mercader maintained his false identity inside a Mexican prison for a decade despite relentless interrogation, and how new medical evidence suggests Trotsky might have survived the attack if his doctors had operated differently. The authors argue that Trotsky's assassination was not a historical anomaly but a template for state-directed political murder that continues to echo from the Litvinenko poisoning to the Skripal attack today.See omnystudio.com/listener for privacy information.
With Apple and Nvidia unleashing hardware that makes powerful local AI a reality, this episode unpacks the tech arms race changing how we run models at home and in business. Are we witnessing the end of the GPU monopoly and the start of real AI independence? (538) Qwen on X: "⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram https://t.co/SScnmzWS7O" / X Thomson Reuters built its own AI model on Chinese open-source tech to slash AI costs Now introducing Gemini Enterprise for Legal Meta reaches $16.68 billion settlement over social media harms to children Twitch and Amazon hit with lawsuit for training AI with streamers' content Three Takeaways From Bill Gates's 5,784-Word Warning on AI: 'There Is No Plan' (214) Breaking911 on X: "WATCH: A humanoid robot training for the "Robot Olympics" in Beijing runs too fast, fails to stop, slams into a safety cushion, and breaks at the waist https://t.co/WEHdYCiPlO" / X Nvidia Is Spending $6 Billion to Build a Powerful U.S. Alternative to Chinese AI A Drone Killed Three Ukrainians. It Was Guided Entirely by A.I. Inside Musk's First Address to Cursor: Grok Is Falling Behind Introducing Butter Box | Butter | Life without internet made smoother. Silicon Valley's Newest Status Symbol Is a Super-Rare Sam Altman Swiss Watch Hosts: Leo Laporte and Jeff Jarvis Co-Host: Christina Warren Download or subscribe to Intelligent Machines at https://twit.tv/shows/intelligent-machines. Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Sponsors: outsystems.com/twit scribe.how/machines framer.com/machines
With Apple and Nvidia unleashing hardware that makes powerful local AI a reality, this episode unpacks the tech arms race changing how we run models at home and in business. Are we witnessing the end of the GPU monopoly and the start of real AI independence? (538) Qwen on X: "⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram https://t.co/SScnmzWS7O" / X Thomson Reuters built its own AI model on Chinese open-source tech to slash AI costs Now introducing Gemini Enterprise for Legal Meta reaches $16.68 billion settlement over social media harms to children Twitch and Amazon hit with lawsuit for training AI with streamers' content Three Takeaways From Bill Gates's 5,784-Word Warning on AI: 'There Is No Plan' (214) Breaking911 on X: "WATCH: A humanoid robot training for the "Robot Olympics" in Beijing runs too fast, fails to stop, slams into a safety cushion, and breaks at the waist https://t.co/WEHdYCiPlO" / X Nvidia Is Spending $6 Billion to Build a Powerful U.S. Alternative to Chinese AI A Drone Killed Three Ukrainians. It Was Guided Entirely by A.I. Inside Musk's First Address to Cursor: Grok Is Falling Behind Introducing Butter Box | Butter | Life without internet made smoother. Silicon Valley's Newest Status Symbol Is a Super-Rare Sam Altman Swiss Watch Hosts: Leo Laporte and Jeff Jarvis Co-Host: Christina Warren Download or subscribe to Intelligent Machines at https://twit.tv/shows/intelligent-machines. Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Sponsors: outsystems.com/twit scribe.how/machines framer.com/machines
With Apple and Nvidia unleashing hardware that makes powerful local AI a reality, this episode unpacks the tech arms race changing how we run models at home and in business. Are we witnessing the end of the GPU monopoly and the start of real AI independence? (538) Qwen on X: "⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram https://t.co/SScnmzWS7O" / X Thomson Reuters built its own AI model on Chinese open-source tech to slash AI costs Now introducing Gemini Enterprise for Legal Meta reaches $16.68 billion settlement over social media harms to children Twitch and Amazon hit with lawsuit for training AI with streamers' content Three Takeaways From Bill Gates's 5,784-Word Warning on AI: 'There Is No Plan' (214) Breaking911 on X: "WATCH: A humanoid robot training for the "Robot Olympics" in Beijing runs too fast, fails to stop, slams into a safety cushion, and breaks at the waist https://t.co/WEHdYCiPlO" / X Nvidia Is Spending $6 Billion to Build a Powerful U.S. Alternative to Chinese AI A Drone Killed Three Ukrainians. It Was Guided Entirely by A.I. Inside Musk's First Address to Cursor: Grok Is Falling Behind Introducing Butter Box | Butter | Life without internet made smoother. Silicon Valley's Newest Status Symbol Is a Super-Rare Sam Altman Swiss Watch Hosts: Leo Laporte and Jeff Jarvis Co-Host: Christina Warren Download or subscribe to Intelligent Machines at https://twit.tv/shows/intelligent-machines. Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Sponsors: outsystems.com/twit scribe.how/machines framer.com/machines
With Apple and Nvidia unleashing hardware that makes powerful local AI a reality, this episode unpacks the tech arms race changing how we run models at home and in business. Are we witnessing the end of the GPU monopoly and the start of real AI independence? (538) Qwen on X: "⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram https://t.co/SScnmzWS7O" / X Thomson Reuters built its own AI model on Chinese open-source tech to slash AI costs Now introducing Gemini Enterprise for Legal Meta reaches $16.68 billion settlement over social media harms to children Twitch and Amazon hit with lawsuit for training AI with streamers' content Three Takeaways From Bill Gates's 5,784-Word Warning on AI: 'There Is No Plan' (214) Breaking911 on X: "WATCH: A humanoid robot training for the "Robot Olympics" in Beijing runs too fast, fails to stop, slams into a safety cushion, and breaks at the waist https://t.co/WEHdYCiPlO" / X Nvidia Is Spending $6 Billion to Build a Powerful U.S. Alternative to Chinese AI A Drone Killed Three Ukrainians. It Was Guided Entirely by A.I. Inside Musk's First Address to Cursor: Grok Is Falling Behind Introducing Butter Box | Butter | Life without internet made smoother. Silicon Valley's Newest Status Symbol Is a Super-Rare Sam Altman Swiss Watch Hosts: Leo Laporte and Jeff Jarvis Co-Host: Christina Warren Download or subscribe to Intelligent Machines at https://twit.tv/shows/intelligent-machines. Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Sponsors: outsystems.com/twit scribe.how/machines framer.com/machines
My guest today is Neil Movva, founder of Sail. Sail is building what Neil calls a token factory, an inference company designed for a specific kind of future, one where AI agents run in the background for hours or days at a time rather than answering a human in real time. In that world, latency matters less and cost matters more, and Neil has built the whole company around driving the cost of a token as low as it can possibly go. What makes this conversation special is that it is one of the most detailed tours I have ever done through the full stack of intelligence, the software, the chips, and the power, and how all three connect. Along the way we cover the trade-off between speed and cost that lives inside every GPU, his scavenger strategy for buying the chips and power nobody else wants, his contrarian view on Nvidia, and why the premium the frontier labs charge for being three to six months ahead may not last. Please enjoy my conversation with Neil Movva. For the full show notes, transcript, and links to mentioned content, check out the episode page here. ----- Become a Colossus member to get our quarterly print magazine and private audio experience, including exclusive profiles and early access to select episodes. Subscribe at colossus.com/subscribe. ----- Ramp's mission is to help companies manage their spend in a way that reduces expenses and frees up time for teams to work on more valuable projects. Go to ramp.com/invest to sign up for free and get a $250 welcome bonus. ----- Trusted by thousands of businesses, Vanta continuously monitors your security posture and streamlines audits so you can win enterprise deals and build customer trust without the traditional overhead. Invest Like the Best listeners get a special offer of $1,000 off Vanta when you go to vanta.com/invest. ----- WorkOS is the infrastructure B2B and AI-native companies use to sell to enterprise. It covers everything enterprise security requires: SSO, SCIM, RBAC, Audit Logs, AI governance, and more. Trusted by 2,000+ fast-growing companies, including OpenAI, Anthropic, Cursor, and Vercel. ----- Rogo is the AI platform for finance. They're building agents for Wall Street that are trained to understand how bankers and investors actually do work: from diligence and modeling, to turning analysis into deliverables. To learn more, visit rogo.ai/invest. ----- Ridgeline has built a complete, real-time, modern operating system for investment managers. It handles trading, portfolio management, compliance, customer reporting, and much more through an all-in-one real-time cloud platform. Visit ridgeline.ai. ----- Editing and post-production work for this episode was provided by The Podcast Consultant. Timestamps: (00:00:00) Welcome to Invest Like The Best (00:02:20) Neil Movva (00:03:22) Building a Token Factory (00:05:32) The Rise of Long-Running Agents (00:08:47) Deep Research and Cybersecurity (00:15:12) The Full Stack of Intelligence (00:20:03) Throughput Versus Latency (00:24:58) The Future of AI Chips (00:33:19) Why Transformers Work (00:36:43) The Future of Data (00:44:05) The Market for AI Chips (00:47:56) Is the AI Boom Different? (00:51:08) Reinventing the Data Center (00:56:43) Scavenging Power (01:01:04) Where Compute Is Most Inefficient (01:07:02) Open Versus Closed Models (01:10:37) A Trillion Tokens a Day (01:12:42) The Contrarian Case on NVIDIA (01:14:38) Advice for AI Hardware Founders
Apple refreshed the Mac Studio with M5 Max and M5 Ultra and gave the Mac mini M6 silicon, at higher prices. OpenAI's Jalapeño chip beat Nvidia on efficiency, Perplexity went fully local, WhatsApp toughened logins, and Uber livestreamed teen rides. Links Apple updates the Mac Studio with M5 Max and M5 Ultra, with up to 4.3x faster AI performance, faster graphics, and up to 512GB of unified memory for $2,499+ (Apple Newsroom) Apple unveils a Mac mini with M6 and M5 Pro, with up to 4x faster AI performance and 2x faster graphics, for $899+ and $1,699+, with preorders today and shipping September 22 (The Verge) Apple's M6 is its first 2nm chip with a 12-core CPU and GPU, while the M5 Ultra fuses two dual-die M5 Max chips into a 36-core CPU, 80-core GPU "most powerful chip ever" (The Verge) OpenAI says its Jalapeño chip delivered 1.5x-1.9x more AI work per watt and 1.7x-3.6x lower latency than Nvidia chips across GPT-OSS, DeepSeek R1, Kimi K2.5 1T (The Verge) WhatsApp upgrades its two-step verification, letting users replace the six-digit PIN with a longer alphanumeric password, and adds support for multiple passkeys (TechCrunch) Perplexity launches Portable Computer, a local AI agent platform running fully on-device with zero token costs, starting with Nvidia DGX Spark and RTX Linux PCs (VentureBeat) Uber launches an optional safety feature allowing parents or guardians to watch a livestream of their teen's ride via the driver's front-facing phone camera (Bloomberg) Subscribe to the ad-free feed.
PLTR 2026 Q2 財報與估值講座 活動連結:https://miula.kaik.io/events/7b686abf-118f-47ef-93dc-b22db5239dc2 科技巨頭解碼訂戶專屬:https://vocus.cc/salon/miula/products/pltr2026q2 EP330. 輝達 GPU 要漲價、中國機器人跑很快、特斯拉新消息更新 | M觀點 (00:40) EP330 預告 (01:56) 業配時間:Palantir 第二季財報分析與估值講座 (03:46) 第一個話題:輝達 GPU 漲價 (19:56) 第二個話題:中國機器人很會跑 (33:53) 第三個話題:特斯拉新消息更新 M觀點資訊 科技巨頭解碼: https://bit.ly/3koflbU M觀點 Telegram - https://t.me/miulaviewpoint M觀點 IG - https://www.instagram.com/miulaviewpoint/ M觀點Podcast - https://bit.ly/34fV7so M報: https://bit.ly/345gBbA M觀點YouTube頻道訂閱 https://bit.ly/2nxHnp9 M觀點粉絲團 https://www.facebook.com/miulaperspective/ 任何合作邀約請洽 miula@outlook.com -- Hosting provided by SoundOn
Syncthing saves the day in an unexpected way, and we dig into what's new in Linux 7.2.Sponsored By:Jupiter Party Annual Membership: Put your support on automatic with our annual plan, and get one month of membership for free!Managed Nebula: Meet Managed Nebula from Defined Networking. A decentralized VPN built on the open-source Nebula platform that we love.Support LINUX UnpluggedLinks:Web Boost — Send us a boost via sats or USD