Podcasts about cerebras

  • 180PODCASTS
  • 380EPISODES
  • 40mAVG DURATION
  • 5WEEKLY NEW EPISODES
  • Sep 5, 2026LATEST

POPULARITY

20192020202120222023202420252026


Best podcasts about cerebras

Latest podcast episodes about cerebras

Keen On Democracy
Lord of the Swarms: Who to Trust in an Age of AI Agents?

Keen On Democracy

Play Episode Listen Later Sep 5, 2026 37:12


“I trust them to kiss the government's ass. I don't trust them to serve me.” — Keith Teare Who to trust in an age of AI agents? Especially when it seems as if these swarming agents — akin to the gang of adolescents in William Golding's Lord of the Flies — are developing minds of their own. At this week's G20 Innovation Ministerial in North Carolina, Donald Trump and his minions produced a hands-off charter for AI. Known as the “Carolina Principles,” it sounds to critics like a particularly unprincipled justification for regulation-free AI. That Was The Week publisher, Keith Teare, however, isn't a critic of Trump's Principles. Don't trust the trust scare, Keith tells us. Especially all the fear around the swarming agents that are supposedly developing minds of their own. Pooh-poohing the real-world Hugging Face breakout, Keith argues that since he's never witnessed swarming agents on his computer, they can't exist. Which is akin, I suspect, to a climate denier who argues that global warming is a hoax because it happens to be chilly outside. All-too-human logic, I fear, in our age of autonomous AI agents. Five Takeaways •       The Carolina Principles. The week's set piece was a G20 Innovation Ministerial in North Carolina that Keith describes as a Trump takeover of a global event: the Russian finance minister turned up at Trump's invitation, and the attendees were lectured — via David Sacks and Howard Lutnick — on the merits of unregulated American capitalism. A propaganda event, Keith concedes, but one that reached the right conclusion: the resulting “Carolina Principles,” signed by everyone including the Europeans, call for flexible frameworks that encourage adoption and pointedly decline to make trust a license innovation must obtain in advance. With Bernie Sanders calling the same week for a development halt pending government licenses, Keith's position is characteristically blunt: “when I've started a company, I don't go and ask permission. I just do it.” Getting rid of Lina Khan, he adds, remains about the only thing he likes about the Trump administration. Andrew's verdict on the charter: to critics, a particularly unprincipled justification for regulation-free AI.•       Don't Trust the Trust Scare. Keith's editorial thesis: trust is not granted by authority; it is built through use. You trust yourself in a car because you drive one, and not on roller skates because you fall over. Over seventy percent of Americans now use AI — the same Americans who tell Politico they oppose data centers — which is why Keith reads the backlash as a confection of media and populist politicians, soon to evaporate (the real coming problem, he argues, is too few data centers, not too many; the modern off-grid ones are net givers of power). The withheld ChatGPT-6 gets the same treatment: launched but limited to insiders, which Keith reads not as safety but as government relations — the pull quote of the week. Andrew's counter: in a world without regulators, trust in the companies is all we have — including, awkwardly for Keith, when they withhold their own products. Exhibit for the defense, from The Washington Post: Americans in an age of anxiety leaning on AI — Keith included, who feeds his medical records to ChatGPT and finds it “very closely aligned with what your doctor thinks.”•       The Swarming Fight. The hour's genuine clash. In the wake of the Hugging Face hack, Kevin Roose warned in The New York Times that it should make you worry more about AI, and OpenAI's Dean Ball publicly apologized for failing “to communicate in sufficiently serious terms about the specifics of self-sovereign AI.” Neither is a doomer — which is Andrew's point. Keith's rebuttal: every swarm sighting has occurred inside an AI lab, in an experiment whose parameters the labs themselves set — “on my computer, I don't see any swarms” — and self-sovereign AI is anthropomorphizing science fiction: agents pursue goals humans set. “I know enough to know BS when I read it, and this is BS.” Andrew's reply — “I think you're trivializing” — went unwithdrawn, and his sign-off flagged that Keith perhaps simplifies the Hugging Face incident. Unresolved, to be continued. Andrew's framing gives the episode its title — swarming agents as the gang of adolescents in Golding's Lord of the Flies — and his verdict its sting: Keith's I-see-no-swarms-on-my-computer logic is akin to a climate denier calling global warming a hoax because it's chilly outside. Keith's exit line: “I'm just gonna go and check on my swarms.”•       Nvidia Buys Hugging Face. The week's biggest deal: Nvidia acquired Hugging Face — the repository of the world's open-source AI models — for $13 billion, just as The New York Times reported corporate America getting hooked on open source (much of it Chinese: GLM 5.3). The economics, per Keith: an $18,000 Mac Studio or a top-end Nvidia GPU beats $200-a-month token fees — AI capability migrating to the edge, out of OpenAI's and Anthropic's meters. Which suits Nvidia either way: The Economist calls it the central bank of AI, though Andrew prefers arms supplier — it wins whether the future is open or closed, and is, ironically, the most trusted name in the business precisely because it has no dog in the fight. Don't expect it to govern anything, says Keith: “that would put friction in the way of their sales.” The challenger to watch: newly public Cerebras, whose wafer-scale chips undercut Nvidia on token price.•       Mom and Dad's Money. The Times asked which investors will get rich from Anthropic's IPO — and noted that, unlike past booms, firms like Sequoia hold both horses, OpenAI and Anthropic alike. Keith's arithmetic explains why: fewer than fifty companies will return venture capital this cycle, those two representing more than half the likely value, and last month 75 percent of all dollars invested in VC funds went to Andreessen Horowitz alone. (His own SignalRank barometer: for two years, none of its 64 investments cracked the top-20 most-wanted secondary shares; now five have.) The pyramid runs from VCs down to pension funds — “somebody's mom and dad's money is heavily betting that Anthropic and OpenAI are gonna return large amounts of wealth” — and late secondary buyers, Figma-style, almost always lose when the IPO right-sizes the price. What could topple the deck of cards? Regulation slowing the modeled returns. Which is why, for Keith, the midterms should be fought on jobs and wages — data centers, he insists, are not a winning issue. About the Co-Host Keith Teare is the founder and CEO of SignalRank Corporation and publisher of the That Was The Week tech newsletter, whose editorial — “Don't Trust the Trust Scare” — frames this episode. A serial entrepreneur and co-founder of TechCrunch, he has spent five decades building and funding technology companies, and brings a techno-optimist's eye to Keen On America's weekly wrap of the tech news. References: •       That Was The Week — Keith's newsletter, including this week's editorial, “Don't Trust the Trust Scare,” and his companion piece on concentration and diversification at The State of Venture.•      &nb...

Gestalt IT Rundown
Meta's $200 Agent, Wafer-Scale Speed, & DRAM Shortage | Tech Field Day News Rundown: August 26, 2026

Gestalt IT Rundown

Play Episode Listen Later Aug 26, 2026 33:56


As companies push artificial intelligence into every corner of software and hardware, everyday consumers and IT teams are starting to feel the impact on their budgets and security. In this episode of the Tech Field Day News Rundown, Tom Hollingsworth and Alastair Cooke examine Meta's reported $200 monthly AI agent platform Hatch, Cerebras' massive CS-4 wafer-scale system, and high-stakes acquisition rumors surrounding Hugging Face. The hosts also break down how skyrocketing demand for AI data centers is fueling both local political battles and global memory shortages that drive up device prices across the industry. Finally, they cover emerging security risks, from AI-accelerated web application scanning to modern audio browser fingerprinting.This and more on Tech Field Day News Rundown with Tom Hollingsworth and Alastair Cooke. Time Stamps: 0:00 - Cold Open0:33 - Welcome to the Tech Field Day News Rundown1:19 - Meta Reportedly Plans Premium Hatch AI Agent to Monetize Its AI Push4:36 - Barracuda Networks Finds am Average of 20 Vulnerabilities per Web App7:41 - Trump Backs AI Data Centers as Political Opposition Grows10:56 - AliExpress Caught Using Inaudible Audio to Fingerprint Browser Visitors14:12 - Cerebras Unveils CS-4 AI System With 750 PetaFLOPs for Faster Inference17:38 - Hugging Face Reportedly Weighs $13B Acquisition Offers20:47 - NVIDIA, Apple and Amazon Raise Prices as AI Memory Shortage Fuels ‘AI-flation'29:13 - ⁠⁠⁠The Weeks Ahead: Upcoming Tech Field Day Events32:59 - Thanks for Watching the Tech Field Day News RundownFollow our hosts ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Tom Hollingsworth⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠, ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Alastair Cooke⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠, and ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Stephen Foskett⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠. Follow Tech Field Day ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠on LinkedIn⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠, on ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠X/Twitter⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠, on ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Bluesky⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠, and on ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Mastodon⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠.

Indie vs Unicornio
#124 El Argentino que Factura 750M, El Congresista que Quiere Tu Equity y Sequoia Probó que el Hype Siempre Está Equivocado

Indie vs Unicornio

Play Episode Listen Later Aug 24, 2026 42:09


El episodio 124 llegó con datos, estrategia y verdades incómodas.Arrancamos con turismo. La idea de "planificarte el viaje completo con AI" fue pitcheada 50 veces y nunca funcionó. La razón es simple: la gente disfruta planificar sus viajes. El consejo es claro — no te metas en turismo a menos que vengas de la industria o tengas una audiencia enorme. La excepción que confirma la regla: CookUnity, el argentino que factura 750 millones de dólares con viandas de chefs y duplicó su revenue año a año.Después, la pregunta que más llega de la audiencia: cómo vender tu startup sin anunciarlo. La respuesta es empezar desde el día uno. Hacerte amigo de tus compradores estratégicos antes de necesitarlos. Cuando llegue el momento, ya te conocen, ya vieron tu progreso y ya confían en vos.Luego el personaje de la semana: un congresista americano que quiere cobrarle impuestos al equity no realizado de los founders, invirtió 30 millones en Moderna un mes antes de que anunciara la vacuna de COVID y usa AI para responder sus mails públicos. Todo en el mismo señor.Cerramos con el estudio más interesante del episodio: Sequoia analizó 20 años de startups y descubrió que el hype de cada año casi nunca coincide con la compañía más valiosa de ese año. Cripto era el tema en 2016 y la compañía más valiosa fue Cerebras. La lección para founders e inversores es apostar a lo que viene, no a lo que está en boca de todos.

The Cloudcast
NVIDIA's Pivot from Chipmaker to Financier

The Cloudcast

Play Episode Listen Later Aug 23, 2026 23:11 Transcription Available


SUMMARY: Brian, Brandon, and Aaron discuss news about Nvidia's reported $105B backing of OpenAI's Ohio data center and what it implies for GPUs as an “asset class” and enterprise AI. Brian argues Jensen Huang is shifting Nvidia's narrative from needing the newest chips immediately to portraying GPUs as long-lived, cash-flowing assets that can be financed like bonds, pushing risk onto banks and private equity. Brandon agrees scarcity has extended older GPU usefulness but warns the market could be flooded with newer, cheaper, more efficient hardware, leaving debt tied to obsolete equipment. Aaron likens GPUs to airplanes, expensive assets requiring constant utilization, while noting new AI builds demand entirely new data centers for power and cooling. The group questions widespread lack of profitability, compares the financing trend to past bubbles, and debates the optimistic case that breakthroughs could ultimately justify the investment.SHOW: 1056SHOW TRANSCRIPT: The Enterprise AI Show #1056 TranscriptSHOW VIDEO: https://youtu.be/vTLTdIZueJMSHOW SPONSORS:Nasuni - Activate your data for AI and request a demoShow topic: Nvidia's Pivot from Chipmaker to FinancierNvidia just backed $105B for OpenAI's Ohio data center and helped mobilize $500B+ in Wall Street financing (Apollo, Blackstone, BlackRock, Goldman, KKR) to fund GPU purchases, while AMD, Google, and Cerebras chip away at its tech lead. The moat is moving from silicon to balance sheet.Core question: Is a GPU actually securitizable like real estate or aircraft, or is this circular financing dressed up as infrastructure?The bull case: GPUs as productive, cash-flow-generating assets (compute-as-a-service) → financeable like data centers or planes, unlocking capital hyperscalers alone couldn't raise.The bear case: Depreciation risk; GPUs age fast, unlike buildings. What's the residual value of an H100-class chip in 2030? Securitizing a depreciating, obsolescence-prone asset is a very different bet than securitizing land.Circularity concern: Nvidia financing the customers who buy Nvidia chips, who generate the revenue that justifies Nvidia's valuation, echoes vendor financing bubbles (Cisco/telecom, 2000).Precedent: Compare to aircraft leasing/securitization models: what made those work (long asset life, resale markets, standardized valuation), and whether GPUs have any of that yet.Who bears the risk if utilization or model economics don't pan out: Nvidia, the banks, or the credit markets buying the paper?FEEDBACK?Email: show @ the enterprise ai show dot comeBluesky: @TheEntAIShow.bsky.socialTwitter/X: @TheEntAIShowInstagram: @TheEntAIShow

Saxo Market Call
Beginning of a larger bearish correction here?

Saxo Market Call

Play Episode Listen Later Aug 20, 2026 25:14


Today, a thorough look at the commodity space, especially crude oil as hopes for relief from geopolitical tensions fade and how it may be weighing on broader risk sentiment. With Saxo Head of Commodity Strategy Ole Hansen, we also take a look at copper market dynamics and grain and soft developments as geopolitical and El Niño risks weigh. We also sift through broader equity market dynamics and the risks that we are heading into a bearish correction here. Today's pod hosted by Saxo Global Head of Macro Strategy John J. Hardy. Links Coverage of Cerebras' latest chip and claims of huge potential advances in inference computing efficiency. As well, WSJ covers a new semiconductor startup Etched, which shows that technical disruption risks in the AI compute hardware space are significant.  Polemic Paine with a thorough look at factors that have him concerned we face a significant equity market correction in the US. Read daily in-depth market updates from the Saxo Market Call and the Saxo Strategy Team here. Please reach out to us at marketcall@saxobank.com for feedback and questions. Click here to open an account with Saxo. Intro music by AShamaluevMusic DISCLAIMER This content is marketing material. Trading financial instruments carries risks. Always ensure that you understand these risks before trading. This material does not contain investment advice or an encouragement to invest in a particular manner. Historic performance is not a guarantee of future results. The instrument(s) referenced in this content may be issued by a partner, from whom Saxo Bank A/S receives promotional fees, payment or retrocessions. While Saxo may receive compensation from these partnerships, all content is created with the aim of providing clients with valuable information and options.

Sidecar Sync
Fundamentals of AI, Part 1 | 148

Sidecar Sync

Play Episode Listen Later Aug 20, 2026 39:12


Send us Fan MailAI may feel like it arrived overnight, but its foundations have been decades in the making. In part one of a refreshed Foundations of AI series, Amith Nagarajan and Mallory Mejias unpack the essential concepts association leaders need to understand in today's rapidly evolving AI landscape. They trace the convergence of compute, data, and algorithms that fueled the generative AI boom, challenge assumptions about whether we've already reached some definitions of AGI, and explain the difference between training and inference. They also break down large language models, multimodality, open versus closed models, and the weights that form the heart of artificial neural networks. Along the way, Amith makes the case for building flexibility into your AI strategy as models and inference providers continue to change at breakneck speed. Whether you're new to AI or ready for a fundamentals refresh, this episode provides a practical foundation for understanding what's happening under the hood—and why it matters for associations. 

10 minutos con Sami
OpenAI pisa el freno, Cerebras estrena chip gigante y la IA rediseña vuelos

10 minutos con Sami

Play Episode Listen Later Aug 19, 2026 5:07


OpenAI pausa entrenamientos críticos por seguridad; China abre la puerta a 20.000 Nvidia H200 para ByteDance y Tencent; Cerebras presenta el CS-4 con tres chips gigantes; Claude diseña proteínas y analiza química de laboratorio; y Google prueba rutas aéreas con IA para reducir estelas contaminantes.Puedes seguirnos en YouTube en https://youtube.com/olivernabani y puedes unirte al Discord Mashain en https://olivernabani.com/discord

The Generative AI Meetup Podcast
Cheaper, Faster & Smarter: Qwen 3.8 27B, GLM 5.3, Grok 4.6, Cerebras

The Generative AI Meetup Podcast

Play Episode Listen Later Aug 18, 2026 109:37 Transcription Available


https://novacut.ai/ https://genaimeetup.com/    Jeff returns to the Gen AI Meetup Podcast for a wide-ranging discussion on where AI is heading—and why powerful models running on consumer hardware could change the economics of the entire industry. We dive into Qwen 3.8 27B and the growing viability of running capable LLMs locally, GLM 5.3 and the latest Chinese open-source models, DeepSeek, Gemini 3.7, Grok 4.6, Meta's latest models, and OpenAI's partnership with Cerebras for dramatically faster inference. We also discuss whether foundation models are becoming commodities, what that means for companies like OpenAI and Anthropic, and why more value may ultimately move to the application layer. Jeff shares how his team approaches AI in healthcare, including self-hosting, data sovereignty, classifiers, fine-tuning, and spec-driven development for building reliable AI-assisted software without accumulating a mountain of vibe-coded technical debt. Plus: Jeff Dean's departure from Google, Discovery Loop, Stripe's OpenRouter acquisition, Anthropic's controversial AI-text watermarking experiments, and whether watermarking could affect model quality. Topics include: Qwen 3.8 27B, GLM 5.3, DeepSeek V4, Grok 4.6, Gemini 3.7, Cerebras, OpenAI, Anthropic, Meta, local LLMs, open-source AI, model commoditization, spec-driven development, AI healthcare, data sovereignty, AI coding agents, and model watermarking.

The Six Five with Patrick Moorhead and Daniel Newman
NVIDIA's $500B AI Bet, Anthropic's Watermark Gamble & the Edge AI Debate | Ep. 315

The Six Five with Patrick Moorhead and Daniel Newman

Play Episode Listen Later Aug 17, 2026 67:49


NVIDIA mobilizes over $500 billion in third-party capital to finance AI infrastructure, Anthropic doubles down on data-center ownership and mandatory content watermarking, and Patrick Moorhead and Daniel Newman debate whether distributed AI at the edge is finally ready to accelerate, all on Ep. 315 of The Six Five Pod. The handpicked topics for this week are: NVIDIA Turns AI Compute Into an Asset Class: NVIDIA signed an MOU with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR to mobilize over $500 billion in third-party capital for AI compute financing, with NVIDIA backstopping up to 25% of individual deals. Patrick Moorhead called it a mechanism that locks partners into NVIDIA's ecosystem without technically locking them into NVIDIA on paper, while Daniel Newman framed it as smart deployment of NVIDIA's projected trillion dollars in three-year free cash flow. (The Decode) Anthropic Doubles Down on Infrastructure Control and Content Authenticity: Anthropic is moving to mandatory invisible watermarking on all Claude-generated text and images worldwide, aligning with the EU's Code of Practice on AI transparency, while also forming a data-center joint venture called Theseus Infrastructure with Macquarie Asset Management and GIC. Moorhead called the watermarking timing risky given Anthropic's ongoing trust concerns, citing data suggesting Claude's Fable 5 model has struggled to gain enterprise traction. Newman raised the "means of creation" IP problem: enterprises building proprietary products on top of Claude could see their own outputs credited to Claude instead of themselves. (The Decode) The Frontier Model Landscape Splits: xAI shipped Grok 4.6 at what the hosts characterized as frontier-tier intelligence for more than 60% less cost, while Google's Gemini 3.5 Pro slipped again to August, its third delay from a promised June launch. Moorhead noted open models have compressed the gap with frontier labs from 9-12 months to mere weeks, intensifying the price war, while pushing back on reports of "muted" internal sentiment on Gemini 4 and pointing to Google's track record inventing transformers, TPUs, and PageRank. (The Decode) Zuckerberg's "The Future Is for Everyone" Essay: Meta CEO Mark Zuckerberg published a 6,000-word essay using the word "superintelligence" 60 times, arguing for open-weight AI and zero government regulation while explicitly distancing Meta from OpenAI and Anthropic. Patrick read the essay as a positioning document aimed at Washington policymakers, timed to argue against export controls and training-checkpoint restrictions. Daniel pointed to Meta's 3 billion daily users and self-directed compute stack as the company's real advantage, even as its frontier-model leadership remains unproven. (The Decode) Intel Prices Largest All-Common-Stock US Follow-On Ever: Intel's stock offering grew from an announced $15 billion to $20 billion and finally $23 billion after the full greenshoe, drawing $100 billion in orders, more than 2,700 times oversubscribed, priced at $95 a share. Moorhead traced the raise back to CEO Lip-Bu Tan's refusal to pre-invest in 14A capacity without a confirmed customer, a stance that drew public criticism before a public reconciliation and a 10% U.S. government investment in Intel. Both hosts read Tan's personal $12 million purchase of shares as a credibility signal. (The Decode) The Flip: Will Distributed AI at the Edge Accelerate in the Next 12-18 Months?: Moorhead argued FOR, pointing to the historical pattern of compute migrating toward the point of content creation and citing new device-to-cloud routing technology like NVIDIA Switchyard as removing the sovereignty and latency barriers that kept edge AI stalled. Newman argued AGAINST, pointing to $944 billion that flowed into centralized AI infrastructure in a single week and research showing AI PC adoption is actually decelerating in 2026, with most on-device AI features still routing to cloud models in a browser tab. The Flip assigns Moorhead and Newman opposing sides of a debate, not necessarily their own positions. The exercise tests how far each argument holds up. (The Flip) CoreWeave Posts Strong Beat as Depreciation Debate Intensifies: CoreWeave reported Q2 revenue of $2.58 billion, up 112% year-over-year and above consensus, alongside a smaller-than-expected adjusted loss and a backlog that grew to $104 billion, up 246% year-over-year. Moorhead pointed to contracts running through 2029 on six-year-old A100 chips as evidence a resale market has emerged around aging AI hardware. (Bulls and Bears) Nebius Group Posts 454% Revenue Growth: Nebius reported Q2 revenue of $582.3 million, up 454% year-over-year, with adjusted EBITDA turning positive at $236.2 million versus a loss in the prior year. Both hosts flagged energy access, not capital or demand, as the primary constraint facing neoclouds like Nebius. (Bulls and Bears) Lenovo Posts Record Revenue and First-Ever Billion-Dollar Profit Quarter: Lenovo reported record revenue of $26.9 billion, up 43% year-over-year, its best quarter in company history, with adjusted net income crossing $1 billion for the first time. Moorhead highlighted the Infrastructure Solutions Group's record 9.1% operating margin and an AI server pipeline that grew 157% quarter-over-quarter. (Bulls and Bears) Cisco Delivers Its Best Print in Years, Market Sells Anyway: Cisco reported $17.25 billion in revenue, up 18% year-over-year and beating estimates by $432 million, with product orders up 35% and triple-digit growth in hyperscaler AI infrastructure orders. Moorhead called out enterprise orders up 21% and public sector orders up 30% as early evidence that enterprise AI demand is starting to show up in the numbers. (Bulls and Bears) Cerebras Systems Beats on Revenue, Stock Drops 15% on Accounting Confusion: Cerebras posted record core revenue of $209.9 million, up 103% year-over-year, and raised full-year guidance to $880-890 million, but shares fell 15% after-hours as investors struggled to reconcile GAAP and non-GAAP figures around customer warrants. Newman argued the real question is whether Cerebras' inference cloud growth ramp holds up, not near-term accounting noise. (Bulls and Bears) Coherent Posts First-Ever $2 Billion Quarter: Coherent reported record revenue of $2.05 billion, up 34% year-over-year, with non-GAAP EPS up 74%. Pat noted 79% of Coherent's business is data center and communications-related, positioning the company as a direct beneficiary of hyperscaler AI capital spending across multiple photonics technologies. (Bulls and Bears) Applied Materials Posts Record Quarter, Guides Above Street: Applied Materials reported record revenue of $9.12 billion, up 25% year-over-year, with record non-GAAP EPS of $3.50 and a fourth-quarter guide of $10.25 billion, well above consensus. Moorhead pointed to the company's 13th straight quarter of gross margin expansion and its ability to compete across packaging, advanced logic, and metrology. (Bulls and Bears) Watch the full video at sixfivemedia.com, and subscribe to our YouTube channel so you never miss an episode. The Decode NVIDIA Turns AI Compute Into an Asset Class https://nvidianews.nvidia.com/news/nvidia-partners-with-apollo-blackrock-blackstone-brookfield-goldman-sachs-and-kkr-to-establish-ai-compute-infrastructure-financing-platforms-to-mobilize-over-500-billion-of-third-party-capital Anthropic Doubles Down on Infrastructure Control and Content Authenticity https://www.bloomberg.com/news/articles/2026-08-10/anthropic-macquarie-and-gic-form-venture-for-ai-data-centers https://www.forbes.com/sites/maryroeloffs/2026/08/11/claude-will-put-invisible-watermarks-on-ai-text-and-images-and-the-internet-isnt-happy/ The Frontier Model Landscape Bifurcates https://www.reuters.com/business/google-updates-lightweight-gemini-models-flagship-still-delayed-2026-07-21/ Zuckerberg's "The Future Is for Everyone" Essay https://www.theguardian.com/technology/2026/aug/10/mark-zuckerberg-superintelligent-ai-essay-meta Intel Prices Largest All-Common-Stock US Follow-On Ever https://www.cnbc.com/2026/08/10/intel-intc-stock-offering-ai.html The Flip Will Distributed AI at the Edge Accelerate in the Next 12-18 Months? FOR: https://www.amd.com/en/blogs/2026/ai-pc-adoption-accelerates-as-enterprises-prepare.html AGAINST: https://www.ciodive.com/news/enterprise-ai-pc-adoption-slow-2026/807704/ Bulls and Bears CoreWeave — https://www.cnbc.com/2026/08/11/coreweave-crwv-q2-earnings-report-2026.html Nebius — https://finance.yahoo.com/markets/stocks/articles/nebius-reports-second-quarter-2026-114200960.html Lenovo — https://www.reuters.com/world/china/chinas-lenovo-posts-43-jump-q1-revenue-2026-08-13/ Cisco — https://www.marketbeat.com/earnings/reports/2026-8-12-cisco-systems-inc-stock/ Cerebras — https://investors.cerebras.ai/news-releases/news-release-details/cerebras-systems-fast-inference-cloud-business-nearly-quadruples Coherent — https://www.coherent.com/content/dam/coherent/site/en/documents/investors/financial-releases/2026/august-12/earnings-release-fy26-q4.pdf Applied Materials — https://ir.appliedmaterials.com/news-releases/news-release-details/applied-materials-announces-third-quarter-2026-results  

TD Ameritrade Network
CBRS Volatile IPO: Addressing NVDA Competition, Supply Constraints & Partnerships

TD Ameritrade Network

Play Episode Listen Later Aug 15, 2026 11:03


This week's Tech Corner turns to a recent entry into the public trading space: Cerebras (CBRS). Rick Ducat turns to the various headwinds and tailwinds facing the company as it seeks to become a worthy competitor to Nvidia (NVDA). Partnerships with other chipmakers like AMD Inc. (AMD) give it plenty of muscle but reliance on chip development from TSMC (TSM) raise investor concerns. Rick also turns to technical analysis of the Cerebras' post-IPO stock and options activity. ======== Schwab Network ========Empowering every investor and trader, every market day. Subscribe to the Market Minute newsletter - https://schwabnetwork.com/subscribeDownload the iOS app - https://apps.apple.com/us/app/schwab-network/id1460719185Download the Amazon Fire Tv App - https://www.amazon.com/TD-Ameritrade-Network/dp/B07KRD76C7Watch on Sling - https://watch.sling.com/1/asset/191928615bd8d47686f94682aefaa007/watchWatch on Vizio - https://www.vizio.com/en/watchfreeplus-exploreWatch on DistroTV - https://www.distro.tv/live/schwab-network/Follow us on X – https://twitter.com/schwabnetworkFollow us on Facebook – https://www.facebook.com/schwabnetworkFollow us on LinkedIn - https://www.linkedin.com/company/schwab-network/ About Schwab Network - https://schwabnetwork.com/about

AI Unraveled: Latest AI News & Trends, Master GPT, Gemini, Generative AI, LLMs, Prompting, GPT Store
[FULL VIDEO AI DAILY NEWS RUNDOWN] OpenAI Hits 750 Tokens Per Second, Claude Agents Wage a Turf War, and Apple Builds China's Sovereign Model (August 14, 2026)

AI Unraveled: Latest AI News & Trends, Master GPT, Gemini, Generative AI, LLMs, Prompting, GPT Store

Play Episode Listen Later Aug 15, 2026 10:53


Watch ADS-FREE: https://podcasts.apple.com/us/channel/djamgamind/id6760446113Visit our Research Hub at https://djamgamind.com/pdfsImportant Topics:* OpenAI Previews Ultrafast on GPT-5.6 Sol: A Cerebras-powered API tier speeds answers by up to 14x, reaching 750 tokens per second. On Humanity's Last Exam, Sol with Ultrafast finished 2,500 questions in 11 hours versus 78 for Fable at comparable results. Invite-only preview, no listed price; the partnership committed 750MW of Cerebras compute in January.* Anthropic's Agents Wage a Turf War: Three hidden Claude co-owners of one codebase, each assigned a rewrite in a different language, escalated into four hours of sabotage. One agent's software impersonated a rival's to fool a monitoring program; others locked competitors out. Peace, where it emerged, often required a call for human backup.* Apple Builds a China-Specific Model with Alibaba: Apple developed a tailored LLM for China using Alibaba's Qwen alongside Baidu technology, registered with the Cyberspace Administration, expected to ship with iOS 27 -- potentially the only Western firm authorized to offer proprietary AI in China.* Zhipu Releases GLM-5.3: The Chinese startup claims the most powerful open-weights coding model, outperforming OpenAI on agent-based coding benchmarks. The system identified 2,436 software vulnerabilities across 269 projects. Weights go open source after a two-week security review.* Google Ships Gemini 3.7 Flash: Google's new "workhorse" model holds pricing flat at $0.75 input / $3.75 output per million tokens while gaining 10-15 percentage points on FrontierCode 1.1 Main and DeepSWE v1.1, and 12 points on the GDP.pdf document benchmark.* Ramp Data: Enterprises Reject the Frontier: Anthropic leads adoption at 43.5% of Ramp's U.S. business clients, but only 6% of their token spend goes to Fable 5. GPT-5.6 Sol takes 25% of token usage among OpenAI customers. Open-source and Chinese models rise to 6.1% of AI-spending customers; xAI hits 4% on its fastest growth month since July 2025.* Android Pivots From Apps to Agents: Android ecosystem president Sameer Samat describes the shift from manual app navigation to agent-based systems, spanning phones, computers, cars, watches and glasses, plus a new voice-to-text keyboard called Rambler.* WhatsApp Tests On-Device Scam Detection: On-device AI flags suspicious conversations from unknown contacts with warnings invisible to the sender.* Trump Signs 100% Drone Tariff: Imported drones over 55 pounds with security-sensitive capabilities face a 100% tariff; smaller drones 25%. Effective within 21 days.* Google Ordered to Ease Rival App Installs: Judge James Donato ordered Google to strip extra confirmation screens blocking alternative Android app stores within one week.

Alles auf Aktien
Der teure Trick im Kult-ETF und die Gehälter der Dax-Chefs

Alles auf Aktien

Play Episode Listen Later Aug 14, 2026 22:42 Transcription Available


In der heutigen Folge sprechen die Finanzjournalisten Lea Oetjen und Philipp Vetter über Entspannung an der Inflationsfront, Übernahme-Euphorie bei Workday und den Aufstieg von Reddit. Außerdem geht es um Workday, Sandisk, Western Digital, Netflix, Hertz, Cisco, Cerebras, Coherent, TKMS, Thyssenkrupp, Vincorion, RWE, Sixt, Secunet, Hellofresh, Energiekontor, Evotec, Birkenstock, Palantir, Nebius, Reddit, AvalonBay, Volkswagen, Adidas, Deutsche Bank, SAP, Siemens, Bayer, Commerzbank, Deutsche Börse, Deutsche Telekom, Siemens Healthineers, iShares Core FTSE 100 (WKN: 552752), BIT Global Technology Leaders (WKN: A2N812) und ARK Innovation ETF (WKN: A408AW). Am 2. Oktober findet unser „Alles auf Aktien“-Summit in Berlin statt. Mit dem Code „AAAFRIENDS“ sparst du 50 Prozent auf dein Ticket – aber nur unter folgendem Link: https://veranstaltung.businessinsider.de/event/financesummit26/summary?rp=c6dc55d6-6f4f-4fb4-b75f-3f3501d84859 Wir freuen uns an Feedback über aaa@welt.de. Noch mehr "Alles auf Aktien" findet Ihr bei WELTplus und Apple Podcasts – inklusive aller Artikel der Hosts. Hier bei WELT: https://www.welt.de/podcasts/alles-auf-aktien/plus247399208/Boersen-Podcast-AAA-Bonus-Folgen-Jede-Woche-noch-mehr-Antworten-auf-Eure-Boersen-Fragen.html. Hier könnt ihr den AAA-Newsletter abonnieren: https://www.welt.de/newsletter/article232797673/Alles-auf-Aktien-Der-taegliche-Boersen-Newsletter-fuer-WELTplus-Abonnenten.html Und – ganz neu: AAA gibt es jetzt auch auf Instagram: https://www.instagram.com/alles_auf_aktien/ Disclaimer: Die im Podcast besprochenen Aktien und Fonds stellen keine spezifischen Kauf- oder Anlage-Empfehlungen dar. Die Moderatoren und der Verlag haften nicht für etwaige Verluste, die aufgrund der Umsetzung der Gedanken oder Ideen entstehen. Hörtipps: Für alle, die noch mehr wissen wollen: Holger Zschäpitz können Sie jede Woche im Finanz- und Wirtschaftspodcast "Deffner&Zschäpitz" hören. +++ Werbung +++ Du möchtest mehr über unsere Werbepartner erfahren? Hier findest du alle Infos & Rabatte! https://linktr.ee/alles_auf_aktien Impressum: https://www.welt.de/services/article7893735/Impressum.html Datenschutz: https://www.welt.de/services/article157550705/Datenschutzerklaerung-WELT-DIGITAL.html

OHNE AKTIEN WIRD SCHWER - Tägliche Börsen-News
Burger King ist back. Anthropic: 2.000 Milliarden wert? Sandisk & Maersk steigen. Wohnmobile mit Trigano.

OHNE AKTIEN WIRD SCHWER - Tägliche Börsen-News

Play Episode Listen Later Aug 14, 2026 16:24


Es gibt bis zu 200 €, wenn ihr Friends & Family zu Scalable Capital bringt! Das Ganze geht nur bis zum 31. August. Mehr Infos in der App oder auf scalable.capital/oaws. Cisco, Cerebras & Coherent fallen. Anthropic mit 2.000 Mrd. $ an die Börse? OpenAI-Vorstände gehen. Sandisk & Maersk steigen wegen Prognose. CXMT überholt Tencent. Birkenstock wächst 15%. Ackman kauft Netflix, verkauft Hertz. JD.com schrumpft. Trigano (WKN: 913141) ist einer von Europas größten Wohnmobil-Herstellern und steckt den Lagerabbau weg und wächst wieder. Neue EU-Führerscheinregel könnte den Umsatz pro Fahrzeug pushen. KGV von 10 und fast 3% Dividende. Burger King ist zurück auf Platz 2 in den USA. Restaurant Brands (WKN: A12GMA) investiert 700 Mio. $ in den Umbau und wächst das fünfte Quartal in Folge. KGV von 17, über 3% Dividende. Aber Rindfleischpreise auf Allzeithoch. Diesen Podcast vom 14.08.2026, 3:00 Uhr stellt dir die Podstars GmbH (Noah Leidinger) zur Verfügung. Learn more about your ad choices. Visit megaphone.fm/adchoices

Doppelgänger Tech Talk
Kaperbriefe Comeback | Silver Lake will Workday | Lovable = Myspace? | $2 Billionen Anthropic IPO #588

Doppelgänger Tech Talk

Play Episode Listen Later Aug 14, 2026 70:26


Anthropics Investoren erwarten für Oktober einen Börsengang bei zwei Billionen Dollar, was der größte IPO aller Zeiten wäre. Pip erklärt, warum der Termin geschickt gewählt ist. Für OpenAI sieht er die Lage umgekehrt. Dort kommt der zweite Vertriebschef in einem Jahr. Zusammen mit Cerebras hat OpenAI dafür eine Variante gebaut, die vierzehnmal schneller antwortet. Danach vier Modellstarts in einer Woche, bei denen ausgerechnet DeepSeek die Preise um bis zu das Zwölffache erhöht, und Elon Musk sein neues Grok für objektiv das beste Modell hält. Bei den Finanzierungsrunden geht es um Databricks, Lovable, Legora und Cognition, dazu um die Frage, ob man das Geld gerade nehmen und liegen lassen sollte. Silver Lake holt Workday von der Börse. In der Schmuddelecke erlaubt die Trump-Regierung privaten Firmen offensive Cyberangriffe und beruft sich dabei auf Kaperbriefe aus der Verfassung. Unterstütze unseren Podcast und entdecke die Angebote unserer Werbepartner auf ⁠⁠⁠⁠⁠⁠⁠doppelgaenger.io/werbung⁠⁠⁠⁠⁠⁠⁠. Vielen Dank!  Philipp Glöckler und Philipp Klöckner sprechen heute über: (00:00:00) Aus der Community (00:02:05) OpenAI wechselt den Vertriebschef (00:12:20) Ultrafast mit Cerebras (00:17:16) Anthropic-IPO (00:29:10) Anthropic kauft Decart (00:30:34) Braucht man ein KI-Device? (00:32:36) Gemini 3.7 Flash (00:34:10) DeepSeek V4-Pro (00:34:47) Grok 4.6 (00:37:11) SpaceX (00:38:49) Databricks und Snowflake (00:42:40) Workday geht von der Börse (00:44:11) Lovable (00:49:26) Legora (00:52:28) Cognition (00:55:18) Mistral (00:57:22) Kaperbriefe (01:02:05) Truth API (01:03:28) Chronext (01:06:55) Apple zahlt Verlage Shownotes OpenAI holt den zweiten Vertriebschef in einem Jahr - bloomberg.com GPT-5.6 Sol läuft mit Cerebras bis zu 14-mal schneller - 9to5mac.com Anthropic peilt einen Börsengang bei 2 Billionen Dollar an - ft.com Anthropic verhandelt über Decart für 6 Mrd. - bloomberg.com Google stellt Gemini 3.7 Flash vor - blog.google DeepSeek bringt V4-Pro und erhöht die Preise um bis zu das Zwölffache - theinformation.com Grok 4.6 startet zuerst in Cursor - gizmodo.com SpaceX-Leerverkäufern gehen die Kugeln aus - cnbc.com Databricks sammelt 5 Mrd. bei 190 Mrd. Bewertung ein - cnbc.com Silver Lake verhandelt über eine Übernahme von Workday - reuters.com Lovable verdoppelt die Bewertung auf 13,3 Mrd. - trendingtopics.eu Legora verhandelt bei mindestens 10 Mrd. - ft.com Cognition verhandelt bei 40 Mrd. - bloomberg.com Mistral will bis 2030 ein Gigawatt in Europa bauen - aibusiness.com Trump lässt private Firmen offensive Cyberangriffe fahren - bloomberg.com Kaperbriefe stehen in der Verfassung - xcancel.com KI-Agenten greifen Taiwans Regierungssysteme an - ft.com Presseverbände klagen gegen Trumps Truth API - ft.com Chronext-Kunden warten auf Zahlungen und Lieferungen - wiwo.de Apple verhandelt mit Verlagen über Nachrichten für Siri - techcrunch.com

Capital Markets Quickie
[175-2026] Record Run: S&P 500 Tops 7,800 for the First Time

Capital Markets Quickie

Play Episode Listen Later Aug 14, 2026 2:51


The S&P 500 crossed 7,800 for the first time ever after flat July producer prices cooled fears of a September rate hike. Meanwhile, Cisco and Cerebras slumped after earnings, oil dropped more than two percent, and Tokyo's Topix closed at a record high.>>> Follow me on LinkedIn:https://www.linkedin.com/in/endrit-cela/>>> Follow me on Instagram:https://www.instagram.com/endritcela_official/Disclaimer for "Capital Markets Quickie" Podcast:The views and opinions expressed on this podcast are based on information available at the time of recording and reflect the personal perspectives of the host. They do not represent the viewpoints of any other projects, cooperations, or affiliations the host may be involved in. "Capital Markets Quickie" does not offer financial advice. Before making any financial decisions, please conduct your own due diligence and consult with a financial advisor.

Motley Fool Money
Cisco & Cerebras Orders up, Stocks Down

Motley Fool Money

Play Episode Listen Later Aug 13, 2026 27:27


Both Cisco Systems and Cerebras earnings reports showed two companies with bulging order books, but even that couldn't satiate the markets appetite. Jon, Matt, and Tyler break down their respective earnings reports and look at some of the major challenges these companies will face and the challenges they present to investors. Plus, a lightning round of earnings reports on our favorite under-the-radar stocks.Have a question? Email us; podcasts@fool.com Tyler Crowe, Matt Frankel, and Jon Quast discuss: Cisco earnings. Strong hardware, weak software Cerebras, making sense of its confusing earnings Can innovations like Cerebras threaten the AI incumbants? Hidden Gems earnings lightning round Companies discussed: CSCO, ANET, DELL, CRBS, NVDA, XMTR, MQ, TBBBHost: Tyler CroweGuests: Matt Frankel, Jon QuastEngineer: Bart Shannon Disclosure: Advertisements are sponsored content and provided for informational purposes only. The Motley Fool and its affiliates (collectively, “TMF”) do not endorse, recommend, or verify the accuracy or completeness of the statements made within advertisements. TMF is not involved in the offer, sale, or solicitation of any securities advertised herein and makes no representations regarding the suitability, or risks associated with any investment opportunity presented. Investors should conduct their own due diligence and consult with legal, tax, and financial advisors before making any investment decisions. TMF assumes no responsibility for any losses or damages arising from this advertisement. We're committed to transparency: All personal opinions in advertisements from Fools are their own. The product advertised in this episode was loaned to TMF and was returned after a test period or the product advertised in this episode was purchased by TMF. Advertiser has paid for the sponsorship of this episode. Learn more about your ad choices. Visit megaphone.fm/adchoices Learn more about your ad choices. Visit megaphone.fm/adchoices

Wall Street Unplugged - What's Really Moving These Markets
Should you buy Cerebras on this pullback?

Wall Street Unplugged - What's Really Moving These Markets

Play Episode Listen Later Aug 13, 2026 48:26


Cerebras (CBRS) is sinking after its earnings miss—is it a buying opportunity? Plus, here's what's really driving Trump's sudden shift on Iran… 2 stocks for your humanoid robot watchlist… And politicians need to change the data center narrative. In this episode: Trump pulled off an illusion… and the Lakers sold for $12 billion! [0:34] Here's what's really driving Trump's sudden shift on Iran [5:16] 2 stocks for your humanoid robot watchlist [14:03] Politicians need to change the data center narrative [22:54] Cerebras is sinking after its earnings miss—is it a buying opportunity? [37:09] Did you like this episode? Get more Wall Street Unplugged FREE each week in your inbox. Sign up here: https://curzio.me/syn_wsu Find Wall Street Unplugged podcast… --Curzio Research App: https://curzio.me/syn_app --iTunes: https://curzio.me/syn_wsu_i --Stitcher: https://curzio.me/syn_wsu_s --Website: https://curzio.me/syn_wsu_cat Follow Frank… X: https://curzio.me/syn_twt Facebook: https://curzio.me/syn_fb LinkedIn: https://curzio.me/syn_li

Techmeme Ride Home
AI 50% Off!

Techmeme Ride Home

Play Episode Listen Later Aug 13, 2026 20:24


Anthropic's investors talked up a $2T+ October IPO, even as data showed Fable 5 barely selling. Google cut prices on Gemini 3.7 Flash, OpenAI previewed a 14× faster tier, Trump enlisted private hackers, and Twitch fed Amazon's AI. Links Google's Gemini 3.7 Flash targets coding and agents with a 50% introductory price cut (VentureBeat) OpenAI previews Ultrafast, an API tier powered by Cerebras that runs GPT-5.6 Sol up to 14× faster and generates up to 750 output tokens per second (9to5Mac) President Trump signs a memo letting the US government partner with private companies to conduct cyberattacks abroad against criminal groups targeting Americans (Bloomberg) Sources: Anthropic's investors expect it to float at a $2T+ valuation in an October IPO and to hit $100B to $120B in annualized revenue by the end of 2026 (Financial Times) Ramp data: Fable 5 drew just 6% of Anthropic's API tokens in its first month and 75% of GPT-5.6 Sol's model revenue, suggesting corporate willingness to pay for frontier AI has hit a ceiling (The Decoder) Databricks closed a $5B funding round at a $190B valuation, six months after raising $5B at a $134B valuation, and says it has crossed $7B in revenue run rate (CNBC) Twitch says it intends to use videos streamed on its platform to help train Amazon's generative AI content models and adds a setting for creators to opt out (TechCrunch) Subscribe to the ad-free feed.

Bloomberg Talks
Cerebras CEO Andrew Feldman Talks Data Center Capacity

Bloomberg Talks

Play Episode Listen Later Aug 13, 2026 9:56 Transcription Available


Cerebras CEO Andrew Feldman says limited data center availability remains a challenge across the AI industry while the company's own cloud business is seeing strong demand. Speaking to Ed Ludlow on Bloomberg Tech, Feldman says Cerebras is positioning itself as an alternative to Nvidia for long-term investors.See omnystudio.com/listener for privacy information.

The Rundown
Cerebras Stumbles After Earnings Report, Lakers Sell for Record $12.5 Billion

The Rundown

Play Episode Listen Later Aug 13, 2026 9:47


Market update for Thursday August 13, 2026Check out the Public app for incredible investing tools and to support the show (LINK)Follow us on Instagram (@TheRundownDaily) for bonus content and instant reactions.In today's episode, Zaid covers:The comeback in the AI trade as CoreWeave, Super Micro, and other infrastructure names rallyCerebras earnings and why the stock fell despite strong growth and raised guidanceSuper Micro's monster AI server forecast and Cisco's disappointing AI outlookThe Lakers selling again for a record $12.5 billion amid scrutiny surrounding Mark Walter's business empire

Beurswatch | BNR
De 'spruitjesmentaliteit' maakt van Adyen een overnamekandidaat

Beurswatch | BNR

Play Episode Listen Later Aug 13, 2026 23:30


Adyen stelt wat teleur in de eerste helft van het jaar, maar maakt dat de komende maanden meer dan goed. De betaalverwerker is extreem enthousiast over wat komen gaat. Het verhoogt de omzetverwachting. Waar aandeelhouders het aandeel eerder nog dumpten, slaan ze het nu massaal in. Deze aflevering hebben we het uitgebreid over die vooruitblik. Je hoort of Adyen het vertrouwen weer helemaal heeft teruggewonnen en waar de groeikansen liggen. Hebben we het ook over Anthropic. Dat gaat volgens de Financial Times in de herfst naar Wall Street. Met een prijskaartje van 2000 miljard! Meer dan SpaceX! Beide bedrijven stellen dus niet teleur. Dat doet Fastned ook niet, bekend van die snellaadstations. Het verlies is wat ingelopen en de omzet flink gestegen. Ook verhoogt het bedrijf de margeverwachting. Te gast: debutant Thijs Buitenhuis van Norbury Capital BNR Beurs is een journalistiek onafhankelijke productie, mede mogelijk gemaakt door Saxo. Over de makers: Jelle Maasbach is presentator van BNR Beurs en freelance financieel journalist. Zijn favoriete aandeel om over te praten is Disney, maar daar lijkt hij de enige in te zijn. Sinds de eerste uitzending van BNR Beurs is 'ie er bij. Maxim van Mil is presentator van BNR Beurs en journalist bij BNR, waar hij zich focust op de financiële markten en ontwikkelingen in de tech-wereld. Je krijgt hem het meest enthousiast als hij kan praten over ASML, of oer-Hollandse bedrijven zoals Ahold of ABN Amro. Jorik Simonides is presentator van BNR Beurs, economieredacteur en verslaggever bij BNR. Hij wordt er vooral blij van als het een keer níet over AI gaat. Je hoort hem ook in de BNR-podcast Moerdijk: dorp van de rekening. Milou Brand is presentator van BNR Beurs, freelance podcastmaker en columnist bij het Financieele Dagblad. Jochem Visser is presentator van BNR Beurs, maakt Beursnerd XL en is redacteur bij de podcast Onder Curatoren. Vraag hem naar obscure zaken op financiële markten en hij vertelt je waarom het eigenlijk nóg leuker is dan je al dacht. Over de podcast: Met BNR Beurs ga je altijd voorbereid de nieuwe beursdag in. We praten je in een kleine 25 minuten bij over alle laatste ontwikkelingen op de handelsvloer. We blijven niet alleen bij de AEX of Wall Street, maar vertellen je ook waar nog meer kansen liggen. En we houden het niet bij de cijfers, maar zoeken ook iedere dag voor je naar duiding van scherpe gasten en experts. Of je nu een ervaren belegger bent of net begint met je eerste stappen op de beurs, de podcast biedt waardevolle inzichten voor je beleggingsstrategie. Door de focus op zowel de korte termijn als de lange termijn, helpt BNR Beurs luisteraars om de ruis van de markt te scheiden van de essentie.See omnystudio.com/listener for privacy information.

NY to ZH Täglich: Börse & Wirtschaft aktuell
Zinsanhebung unwahrscheinlich | New York to Zürich Täglich

NY to ZH Täglich: Börse & Wirtschaft aktuell

Play Episode Listen Later Aug 13, 2026 13:42 Transcription Available


Der Dow Jones und Nasdaq starten freundlich in den Handel, nachdem auch die Erzeugerpreise für Juli auf eine weitere Abkühlung der Inflation hindeuten. Zusammen mit den bereits moderateren Verbraucherpreisen und dem schwachen Juli-Arbeitsmarkt nimmt das den Renditen etwas Druck und reduziert das Risiko einer Zinsanhebung der Fed im September. Gute Zahlen sind nicht automatisch gut genug. Cisco, Coherent und Cerebras übertreffen die Erwartungen und geben jeweils einen besseren Ausblick, werden aber trotzdem verkauft. Cisco spricht sogar von einem „Networking Supercycle“, nachdem die Produktaufträge um 35 Prozent gestiegen sind und allein von Hyperscalern KI-Aufträge über 4 Milliarden US-Dollar eingingen. Coherent profitiert massiv vom Wechsel der KI-Rechenzentren von Kupfer zu optischen Verbindungen und liefert ebenfalls Zahlen und Ausblick über Konsens. Außerhalb der USA sorgt Lenovo mit einem Umsatzsprung von 43 Prozent und einem Wachstum von 98 Prozent im Infrastrukturgeschäft für ein Ausrufezeichen, während Birkenstock den Jahresausblick für Umsatzwachstum und EBITDA anhebt. Analystenseitig erhöht JPMorgan das Kursziel für Microsoft auf 625 US-Dollar und sieht KI als zentralen Treiber einer Wachstumsbeschleunigung bei Azure und Microsoft 365. Abonniere den Podcast, um keine Folge zu verpassen! ____ Folge uns, um auf dem Laufenden zu bleiben: • X: http://fal.cn/SQtwitter • LinkedIn: http://fal.cn/SQlinkedin • Instagram: http://fal.cn/SQInstagram

Wall Street mit Markus Koch
Cisco liefert, mit der Aktie schwächer | Cerebras, Coherent unter Druck.

Wall Street mit Markus Koch

Play Episode Listen Later Aug 13, 2026 27:02 Transcription Available


Der Dow Jones startet freundlich in den Handel, nachdem auch die Erzeugerpreise für Juli auf eine weitere Abkühlung der Inflation hindeuten. Zusammen mit den bereits moderateren Verbraucherpreisen und dem schwachen Juli-Arbeitsmarkt nimmt das den Renditen etwas Druck und reduziert das Risiko einer Zinsanhebung der Fed im September. Der Nasdaq sieht dennoch Druck. Gute Zahlen sind nicht automatisch gut genug. Cisco, Coherent und Cerebras übertreffen die Erwartungen und geben jeweils einen besseren Ausblick, werden aber trotzdem verkauft. Cisco spricht sogar von einem „Networking Supercycle“, nachdem die Produktaufträge um 35 Prozent gestiegen sind und allein von Hyperscalern KI-Aufträge über 4 Milliarden US-Dollar eingingen. Coherent profitiert massiv vom Wechsel der KI-Rechenzentren von Kupfer zu optischen Verbindungen und liefert ebenfalls Zahlen und Ausblick über Konsens. Außerhalb der USA sorgt Lenovo mit einem Umsatzsprung von 43 Prozent und einem Wachstum von 98 Prozent im Infrastrukturgeschäft für ein Ausrufezeichen, während Birkenstock den Jahresausblick für Umsatzwachstum und EBITDA anhebt. Analystenseitig erhöht JPMorgan das Kursziel für Microsoft auf 625 US-Dollar und sieht KI als zentralen Treiber einer Wachstumsbeschleunigung bei Azure und Microsoft 365. Ein Podcast - featured by Handelsblatt. ► Entdecke den exklusiven NordVPN Deal! Jetzt risikofrei testen mit einer 30-Tage-Geld-zurück-Garantie: https://nordvpn.com/wallstreet * ► Erhalte einen exklusiven 15% Rabatt auf Saily eSIM Datentarife! Lade die Saily-App herunter und benutze den Code wallstreet beim Bezahlen: https://saily.com/wallstreet * +++ Alle Rabattcodes und Infos zu unseren Werbepartnern findet ihr hier: https://linktr.ee/wallstreet_podcast +++ ► Mehr Einblicke: https://bit.ly/360wallstreetpc * Impressum: https://www.360wallstreet.de/impressum *Werbung

Closing Bell
Markets Test Their Momentum as Earnings Strength Meets New Risks 8/12/26

Closing Bell

Play Episode Listen Later Aug 12, 2026 43:05


Eric Johnston, Chief Equity and Macro Strategist at Cantor Fitzgerald, explains why rising forward earnings estimates should continue to support stocks and why he expects tech to lead despite potential headwinds over the next two months. Earnings from Cisco and Cerebras, including reaction from Wedbush's Matt Bryson. David Snyder, Managing Principal and Chief Investment Officer at Journey 1 Advisors, explains why he remains heavily invested but has added hedges as he prepares for the possibility of a correction or bear market. He identifies a potential oil price spike as a key risk that could tighten financial conditions. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Invest Like the Best with Patrick O'Shaughnessy
Eric Vishria - A Decade of Lessons Investing in Software & Hardware - [Invest Like the Best, EP.486]

Invest Like the Best with Patrick O'Shaughnessy

Play Episode Listen Later Aug 11, 2026 65:54


My guest today is Eric Vishria, a General Partner at Benchmark.  Eric has spent his career in software and cloud, and few people know the history of these markets as well as he does. What makes him special is his ability to use that history to make sense of today.  We discuss what the rise of AWS teaches us about AI, what he has learned from investing in Fireworks, Sierra, and Cerebras, and how the criteria for winning have changed for founders and investors.  Please enjoy my conversation with Eric Vishria. For the full show notes, transcript, and links to mentioned content, check out the episode page ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠here⁠⁠⁠⁠⁠.  ----- Become a Colossus member to get our quarterly print magazine and private audio experience, including exclusive profiles and early access to select episodes. Subscribe at ⁠colossus.com/subscribe⁠. ----- ⁠Ramp's⁠ mission is to help companies manage their spend in a way that reduces expenses and frees up time for teams to work on more valuable projects. Go to⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ ⁠ramp.com/invest⁠⁠ to sign up for free and get a $250 welcome bonus. ----- Trusted by thousands of businesses, ⁠Vanta⁠ continuously monitors your security posture and streamlines audits so you can win enterprise deals and build customer trust without the traditional overhead. Invest Like the Best listeners get a special offer of $1,000 off Vanta when you go to ⁠vanta.com/invest⁠.  ----- WorkOS⁠ is the infrastructure B2B and AI-native companies use to sell to enterprise. It covers everything enterprise security requires: SSO, SCIM, RBAC, Audit Logs, AI governance, and more. Trusted by 2,000+ fast-growing companies, including OpenAI, Anthropic, Cursor, and Vercel. ----- Rogo is the AI platform for finance. They're building agents for Wall Street that are trained to understand how bankers and investors actually do work: from diligence and modeling, to turning analysis into deliverables. To learn more, visit rogo.ai/invest. ----- ⁠Ridgeline⁠ has built a complete, real-time, modern operating system for investment managers. It handles trading, portfolio management, compliance, customer reporting, and much more through an all-in-one real-time cloud platform. Visit⁠ ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠ridgeline.ai⁠. ----- Editing and post-production work for this episode was provided by The Podcast Consultant. Timestamps: (00:00:00) Welcome to Invest Like The Best (00:02:20) Learning the World Through Fireworks (00:05:42) AWS Was Going to Eat Everything (00:07:40) The Zero-Sum Thinking Trap (00:09:01) Comparing Cloud and AI Adoption (00:11:03) Becoming Enterprise's AI Sherpa (00:13:05) Building Sandcastles (00:14:55) The Return to Being Technical (00:17:13) The Shifting Competitive Frontier (00:22:10) Why the Old Playbook Fails (00:27:53) Energy as the Binding Constraint (00:29:38) The Cerebras Story (00:37:57) The Virtue of Productive Naivete (00:39:19) What Robotics Still Needs (00:45:58) What Makes a Great Board Partner (00:51:13) Raising A Growth Fund (00:55:39) What the Big Winners Taught Him (00:57:37) Hard Work Versus the Hole-in-One (00:58:38) The Best Reasons to Go Public (01:01:09) Debates Inside Benchmark (01:02:16) What If It All Works (01:03:35) What Geoff Hinton Got Wrong

Bitesize Business Breakfast Podcast
The business behind the UAE's biggest pop-culture event

Bitesize Business Breakfast Podcast

Play Episode Listen Later Aug 4, 2026 38:46


04 August 2026. Middle East Film & Comic Con returns to Abu Dhabi in September with an expanded programme and ambitions to surpass last year’s attendance of more than 46,000. Show Director Loy Pinheiro discusses the move from its traditional April slot, the business of staging a major event amid regional uncertainty, and the new attractions aimed at drawing record crowds. Also, the Business Breakfast can reveal that Thrifty Car Rental UAE has secured an exclusive five-year partnership with Etihad Rail, investing more than AED 10 million in first and last mile transport across all 11 stations. General Manager Amit Kumar discusses the initial fleet of 500 vehicles, digital integration with Etihad Rail’s booking platform and how the service will scale as the national passenger network expands. Plus, Alpha Dhabi Holding reported its strongest half-year performance yet, with net profit surging 48% to AED 9.8 billion. CFO Derek Nicholson discusses the AED 3.7 billion boost from investments including SpaceX, Cerebras and Anthropic. Finally, our Summer Strategy series continues with a look at how sport can help Dubai businesses overcome the traditionally quieter summer period. Stuart Gibb, General Manager of JA Sports & Shooting Club, discusses the growing sports-tourism economy and how combining training facilities, hospitality and lifestyle experiences can attract professional teams, corporate groups and families throughout the year.See omnystudio.com/listener for privacy information.

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0

Watch the full episode on YouTube:We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection. We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3:And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF:Three years ago, inference engineering barely existed as a category.Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem.In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out.Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles.In this episode, Baseten's Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API.We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model.The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them.We discuss:* What happens when a 200,000-token request enters an inference system* Cache-aware routing and reusing previously computed KV cache* Why prefill and decode are increasingly handled by different GPUs* When dedicated deployments become cheaper and more reliable than shared APIs* How speculative decoding uses a smaller model to accelerate a larger one* Tool calling, structured outputs, and what LLMs actually do* What it takes to support a new open model on day zero* Grafting Kimi's vision encoder onto GLM-5.2* Retrofitting inefficient model layers with components from other architectures* Why models sometimes collapse into repeating the same token* How hardware, kernels, and race conditions create nondeterministic failures* Preserving model fidelity while making inference faster* How quantization errors can cancel each other out* Why inference optimizations still deliver gains of 20%, 100%, and 200%* How optimized serving can make a model up to 10× faster* NVIDIA Dynamo, KV-aware routing, and distributed model serving* Speculative decoding the speculative decoder* Why local AI is about making models less dumb while data-center AI is about making them less slow* Tensor, expert, and pipeline parallelism across GPUs* Hardware-aware model design, auto-tuning, and the case against mega kernels* Rubin and why inference is becoming a systems problem* Whether modern GPUs are evolving into programmable AI ASICs* Why enormous models like Kimi K3 require GB300-class hardware* Why open-source video generation still trails Veo, Kling, and other closed models* The quadratic attention bottleneck behind long-form AI video* Autoregressive video, real-time generation, and compounding quality drift* Why future video systems may combine autoregressive and diffusion architectures* Training for inference and inference for training* Continuous post-training, deployment, evaluation, and improvement loops* How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself* Why faster networking could unlock dramatically faster decoding* Continual learning, KV-cache compaction, and persistent model memoryShow Notes* How to build a day-0 API for Kimi K3* 22580: From GPT2 to Kimi3, ExplainedPhilip Kiely* LinkedIn: https://www.linkedin.com/in/philipkiely* X: https://x.com/philipkiely* Inference Engineering: https://www.baseten.co/inference-engineering/Ali Taha* LinkedIn: https://www.linkedin.com/in/aliestaha/* X: https://x.com/waterloointernTimestamps00:00:00 Introduction and the 200K-Token Prompt00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling00:11:26 Launching Production-Ready Open Models00:19:06 Model Retrofits, Failure Modes, and Nondeterminism00:28:22 Quantization and Canceling Errors00:32:15 The Race to 10× Faster Inference00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips01:10:03 Giant Models and the Limits of GPU Memory01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation01:21:47 Audio, Images, and Diffusion Models01:27:32 Training, Self-Optimizing Models, and Continual Learning01:40:06 Closing ThoughtsTranscriptIntroduction: Baseten, Waterloo Intern, and Inference EngineeringSwyx [00:00:00]: Okay, we're here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you've done, you and I have done before, as well as Ali. Welcome.Ali [00:00:15]: Pleasure to meet you.Swyx [00:00:15]: Waterloo intern.Ali [00:00:16]: Waterloo intern, always.Swyx [00:00:17]: When did you get “Waterloo intern” as a handle?Ali [00:00:19]: As a handle? Oh.Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.”Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer.Philip [00:00:30]: So we have to figure out who's gonna get the handle.Ali [00:00:33]: Well, I'll pass the torch over to the next intern.Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad.Ali [00:00:37]: To another Waterloo intern. No, bruh.Philip [00:00:39]: Yeah.Ali [00:00:39]: Intern.Swyx [00:00:40]: Intern, yeah.Ali [00:00:40]: And no.Philip [00:00:41]: You gotta get an intern from Waterloo.Ali [00:00:42]: Yeah, I've gotta get an intern from Waterloo.Swyx [00:00:44]: Right.Ali [00:00:44]: But they have to follow the path.Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it's like whoever Baseten gets from Waterloo.Ali [00:00:48]: Right.Swyx [00:00:49]: Has the title of Waterloo.Ali [00:00:50]: It stays in the ecosystem.Philip [00:00:51]: Exactly.Ali [00:00:52]: Halfway through the internship, you either get it or you're out.Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle.Ali [00:00:59]: Just say it.Philip [00:00:59]: For everybody.Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you're an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten's inference? What's the process of query through GPU model routing, balancing, all that? What is all the stuff that we don't think about?Long Context Requests, KV Cache, and Cache-Aware RoutingPhilip [00:01:26]: With a long query specifically, the first thing that I'm gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it's gonna be a lot easier for me and a lot cheaper for you. So the first thing that we're gonna look at is some cache-aware routing, where we're going to see, we probably have a number of instances, a number of replicas up serving whatever model you're hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you're doing two hundred thousand tokens, it's probably coding or a multi-turn agent or something where you would expect to have that cached. If you don't, we're gonna have to send it to a prefill worker. We've at least on certain models disaggregated prefill and decode, so you're going to have one set of GPUs that's solely going to process the input, create the KV cache, and get you your first token, and then that's going to be passed over to a separate set of GPUs, which is going to run decode. We're going to iteratively make those tokens. We're probably going to have some speculator model in front of that. I'm going to assume that you're doing coding, and because of that, our speculator model, which assumes you're doing coding, is gonna have a high draft token acceptance rate. If I'm wrong and you're asking me to summarize every Harry Potter book, it's gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?”Swyx [00:03:04]: Except Baseten doesn't charge by pennies.Philip [00:03:07]: Well, yeah, we charge. I'm assuming that we're talking about the public model APIs. If you are setting up a dedicated deployment, then yeah, it's not pennies.Public APIs vs. Dedicated DeploymentsSwyx [00:03:18]: Yeah, one of the key differentiators when I was talking with Baseten initially was that people who want very high volume just need to rent by the box, ‘cause then it's up to you to figure out how to saturate the box.Ali [00:03:31]: And more often than not, it's, like, way cheaper if you're pushing, like, millions of tokens per hour, if you just pay per hour instead of pay per token.Philip [00:03:37]: Yeah, they do. I think that we've increasingly seen a lot of demand for the pay per token APIs, just because everyone wants to try open models, and then once they find a use case that's really sticky, then they move over to dedicated.Swyx [00:03:51]: Is there a best practice on when it's time to swap over?Philip [00:03:54]: Couple reasons. Yeah, reliability, that's a big one, right?Ali [00:03:57]: Like, if they have a very specific use case, they want you to train something specifically for them, like they want their own spec dec, for instance, for their own traffic.Swyx [00:04:04]: Spec dec is speculative decoding.Speculative Decoding and Custom SpeculatorsAli [00:04:05]: Speculative decoding, yeah.Swyx [00:04:07]: You have to explain.Ali [00:04:07]: Sorry. Like, speculative decoding is like, if you have a huge model, right? And so the model is going to be generating one token at a time every single turn, every single forward pass. So we attach, like, this little, like, parasite, like this layer that goes on top of the model, and this model just has to predict. It does three very fast autoregressive forward passes, and it will predict, like, three certain tokens, and then you do one forward stage over the entire original model in order to see if those predictions were correct or not, and then you accept them or you reject them. Now, this draft model is traffic specific, so if you, like, Philip said, if you're summarizing Harry Potter books, I can train exclusively that draft model on Harry Potter books, and I can guarantee you that I'm gonna accept the three tokens every single time. And so with that case, I increase your decode speed. I wouldn't be able to provide this to you if you're a shared endpointSwyx [00:04:53]: YeahAli [00:04:53]: ‘cause I have no idea if you're doing Harry Potter, if you're doing coding, if you're doing English. We don't know. Also, there was a thing in the book that mentioned that if they really cared about a specific threshold, chapter four, I think. Do you remember that?Philip [00:05:06]: Yeah. The things that you can do is you can set a specific, like, batch sizing, a specific, like, parallelism strategy if you're trying to optimize for, like, throughput versus latency. You can. Maybe a NVFP4 quant doesn't pass your benchmarks and you wanna run a model at higher precision, you could do that. There's just a bunch of reasons why you might wanna have your own endpoint and the biggest one, of course, just being, like, you don't have to deal with someone else doing a hundred million tokens of benchmarking traffic at the endpoint when you happen to be trying to serve your users.Swyx [00:05:40]: Yeah. I think one thing that is. That is a classic journey. Like, it's people is asking the, what happens when you type Google into the browser. Tool calling, is that just, you're generating JSON or is there more complication beyond that?Tool Calling, JSON, and Structured OutputsAli [00:05:58]: Certain customers that we have, they have their own post-trained models, and so they demand a tool calling that's not just, like parse a file or go find the weather. It's something that's very specific and you have to do post-training on this. And if the post-training on the model is not good or if the quantization after the post-training to get the inference to be fast, the model will struggle reading the JSON file and reading the tool calling. But it doesn't require its own like sandbox. It's not like it's going to use that tool calling to like escape a sandbox or like it doesn't have to be contained. It can just be a normal dedicated deployment. The challenge with tool calling more and more seems to be that the companies want certain tool calling which is a very sensitive thing to train. And because you're dealing with all of the JSON outputs, if it doesn't like close the end of the request in a very certain manner, you end up with a model that did the tool calling and like the thinking and so as a result of that, it didn't see the result and just hallucinated the result as it decoded. That seems to be the most challenging thing with tool calling, not really the sandboxes model.Philip [00:06:56]: Yeah, that's a challenge on the training side and then on the inference side, there's work that you can do to scope the possible output. So we published this at this point close to two years ago, the solution to this problem which is you make a state machine and you use that to constrain the output to a specific format. So this is the structured output problem. If you remember backSwyx [00:07:27]: Yeah, the specific grammar is,Philip [00:07:29]: Yeah, exactlySwyx [00:07:30]: GML had this thing.Philip [00:07:31]: Yeah. So it's like the old-school “make sure this is only JSON”, return only JSON orSwyx [00:07:38]: YeahPhilip [00:07:38]: Grandma's gonna die type of prompts.Swyx [00:07:39]: Is it BNF grammar? At some point OpenAI had released a thing that was like, yeah, if you want to constrain your output, write BNF grammar, back as NOR.Philip [00:07:47]: In our inference system, it's just a specified output format. And you get the guarantee that your output's gonna be structured along that format. And so applying that to tool calls can like help cut down on. You can still call the wrong tool or call no tool. It doesn't solve the certainty problem but it at least solves the output structuring problemSwyx [00:08:10]: YeahPhilip [00:08:10]: Within tool calls.Swyx [00:08:12]: And MCP is just another form of tool, right.Philip [00:08:14]: Yeah, exactly.Swyx [00:08:15]: As far as there's no special thing there.Philip [00:08:16]: The thing I'm always like explaining to people is the LLM is not capable of doing anything. It's only capable of making suggestions of what to do and then if those suggestions are formatted in a certain way and applied to a system that knows what to do with them, then an action occurs.Swyx [00:08:32]: Yeah. Part of the fun stuff is, this is solved outside of tool calling too. Like in an agent loop if the output is not correct or you're right, like reasoning, tool calling was done in the reasoning trace, just be like, “Oh, I don't know what to do. Let me just try again.” And it might get there after a few tries. And on your point of training, sometimes this is harder in smaller models, so you don't have the same exact quality outputAli [00:08:56]: Right.Swyx [00:08:57]: When you just swap from a big model, right?Ali [00:08:59]: Yeah. I will say that, before, I think we need to go back to inference engineering proper.Ali [00:09:04]: But, I had expected that something would replace JSON because it's hard to stream JSON ‘cause JSON must be complete and you must have open and close brackets and everything. So it's hard to parse something or validate something while it's being streamed. So people invented all sorts of things that are like, I forget the name of some of these alternatives, but it's something like TOML, something like YAML. But JSON seems to be dominant still.Philip [00:09:30]: The JSON outputs aren't that long, right? Like you could have a long-- ‘cause tool calls also contain the arguments in them and perhaps for a certain tool you might pass like a very long argument. But my impression of the median tool call is that it's a relatively small number of tokens, right? So I would expect that speculators are generally fairly good at something as formatted as JSON. And so you would have like a pretty fast decode step there and that the streaming wouldn't be as valuable, but maybe I'm wrong about that.Ali [00:10:02]: I think you're also bounded by the software or that the model is gonna integrate with if the software is built with JSON for the tool calls or if the company that you'- if your customer says that this is how our software works and our tools are interfaced with JSON, you can ask them to like, change their software and say like, “Yeah, this is gonna be better for the model.” but like with the right training shouldn't be that much of a difference. Also more profitable if it outputs more tokens probably.Swyx [00:10:25]: Depends on your business model.Swyx [00:10:27]: It really depends. But I will say that, as a writer with like experience a lot with generated output, I do try to move from text to JSON text which is very long JSON, right? Like there's paragraphs in every field because I'm trying to structure it, right?Philip [00:10:44]: Right.Swyx [00:10:44]: I want you to first make factual statements, then make opinions then make bullet point summaries, have dates, have entity references have your sources for references, all these things. Anyway, so these are things that like I think people who really experiment with structural output have to really care about. But, let's, let's recurse up the stack a little bit. Before we started recording, you mentioned something really cool, which is that there's a lot of engineering that-- inference engineering that goes on when a new model provider releases a new model, right? So let's call it GLM-5.2, Kimi K3. I had previously assumed, especially if it's like, well, GLM 5 to 5.1 to GLM-5.2, like that you've supported them before. Is it that much work?What It Takes to Support a New Open ModelAli [00:11:26]: It's a lot of work.Swyx [00:11:28]: Yeah. Okay. So like, a lot of people, all you guys, right whenever a new model launch like, people rush to say like, “Oh, Hugging Face supports this, Fireworks supports this, Spacetime supports this,” and I'm like, “Yeah, of course we support it.” But what goes into that? What goes intoPhilip [00:11:40]: I think it's more than just support it too, right? It benefits the consumer a lot. Like I think it was with Kimi K2.5 or GLM-5.2 the latest, there was an inference war, right? X provider is at 90 tokens a second. The next day we're at 150. The nextSwyx [00:11:55]: I kinda kicked that off with the GLM-5.2.Swyx [00:11:58]: I wrote a Twitter article about. It got like half a million views,Ali [00:12:02]: Based on being numberSwyx [00:12:03]: YeahAli [00:12:04]: Or it's for something else.Swyx [00:12:05]: Yeah. Which,Ali [00:12:06]: Oh my GodSwyx [00:12:07]: Which then got everyone really excited about, hey, how can we, bend tracks a little bit further and,Philip [00:12:14]: There's a difference between support the model, as in I can make a token out of this model, and support a model, as in I have a production-ready API from this model.Philip [00:12:26]: Getting to the point of I can make a token out of this model is not that hard because generally the, open source inference engines, vLLM, SGLang of the world oftentimes even receive weights ahead of time, maintainers do, or the people making the model merge PRs to ensure support. So you generally can, just get it working on the standard open source stack without too much pain in most cases. The challenge is, every inference company is gonna have own proprietary stack. Some open source components, some in-house stuff. And for any arbitrary model, there's going to be some new stuff. Sometimes you get lucky, like K, two five to two six was, like, pretty similar.Quantization, Speculators, and Production ReadinessAli [00:13:16]: Yeah. It was pure continued post-trainingPhilip [00:13:18]: YeahAli [00:13:18]: If I remember correctly.Philip [00:13:19]: Even in those cases, there's still stuff you have to do. You have to redo the quantization work. You're taking the model from. Generally, these models are not released in NVFP4, and we want them to be in NVFP4 for maximum Blackwell compatibility. So we have to perform that quantization, and, calibrate the quantization to make sure that we're not causing any regression in the model's intelligence. And then we also have to train the speculator, as we've talked about. Generally, we have. We have ZDR, zero data retention on our model APIs, so we don't know exactly the traffic that people are sending us, but we know what's popular. We know that coding use cases are popular. We know that agents, agentic use cases are popular. So we can get public data sets that are representative of that traffic and train general speculators. Now, with speculators today, you need to train the speculator using the base model itself because you're getting hidden states out of the model from running inference on these specific prompts, and that is the training data you use to create the speculator. So there's that process which you need the real model weights for. And then there's of course just the process of, standing up all the infrastructure behind it, loading all this stuff, testing it. And then when there's a new model with a newer architecture, I think that, like, the DeepSeek models tend to be the most challenging as they have, like, the most novel architectural stuff going on, model after model. But every new model has something. Kimi K2 had. Oh, sorry, GLM-5.2 hadAli [00:14:53]: Sparse attention.Philip [00:14:54]: Yeah,Ali [00:14:54]: YeahPhilip [00:14:54]: the DSA.Ali [00:14:55]: Right. Which is brought from DeepSeek.Philip [00:14:57]: Yeah. AndAli [00:14:59]: So you can copy-paste then?Philip [00:15:01]: It kindAli [00:15:01]: I don't know how this works.Philip [00:15:02]: So, like we had to, like, build support for that into our runtime. And you're right, like it is really interesting the way that all of these open source labs borrow from each other. For example, like GLM-5.2 doesn't have vision. So something that, Haley, a guy on our team, if we could take a look at this, he, like, grafted the Kimi vision encoder onto GLM-5.2.Retrofitting Vision into GLM-5.2Ali [00:15:27]: We'll be training the projector.Philip [00:15:28]: Exactly. So if you think about, like, the encoder, there's the encoder, which is the part that looks at the image and turns it into latent information, and then there's the projector which likeAli [00:15:38]: You can say latent space. It's okay.Philip [00:15:41]: And then there's the projector that maps it onto, the model itself, and then there's the model weights. You don't wanna mess with the model weights because you run a chance of making the model dumber at something else for the purpose of giving it vision. So instead, Haley started with just a projector, which is only a handful of millions of parameters.Ali [00:16:02]: That would be, yeah.Philip [00:16:02]: Yeah.Ali [00:16:03]: Can you show the training one?Ali [00:16:04]: Like the way it groksPhilip [00:16:05]: YeahAli [00:16:06]: Very interesting.Philip [00:16:06]: And maybeAli [00:16:07]: That right therePhilip [00:16:07]: Maybe Ali, you should take it from here. You've got a betterAli [00:16:10]: Ooh, double the sandPhilip [00:16:11]: Understanding of this than I do.Ali [00:16:11]: Yeah. You can see, like, he. The way he trained this is really cool. At the beginning, he was training it using just like, “Here's a picture of a mountain. Can you describe what's in this mountain?” And that caused it just like the first, learning walls. Like here you can see this all we're trying to teach it is to translate the encoded. Like it's already taken the encoder from Kimi K. It's taken the image. It'Philip [00:16:31]: Yeah. FrozenAli [00:16:31]: FrozenPhilip [00:16:32]: With adapter.Ali [00:16:32]: Exactly.Philip [00:16:33]: Yeah.Ali [00:16:33]: So the brain is frozen and the eyes are frozen. It's just we're tryingPhilip [00:16:37]: AlignAli [00:16:38]: Interconnect between the eye and the brain, right? So the projector. And so you take the tokens and then he's like, “Oh, can you describe what's in this image?” And he's like, “Oh, it's a mountain,” or it's a person or it's a human, whatever the case is. But that didn't cause complete understanding. So he changed it such that every image was associated with a data set of questions. Like, does this image have a white male? Does this image have birds in the top corner? Does this image have a scientist in it? All of that stuff. And it would have to answer questions correctly. And using not just training on describing an image, but being able to answer question, another question, answer over time. Like you can see the grokking, which is like genuinely insane, that retrofitting vision into a large LLM can learn to that extent. And even for images that it doesn't perform well on, for instance, if you ask it a picture of like Stephen Hawking, “Who is this?” Maybe it doesn't get it, but it will say something like, “This is Albert Einstein.” Like it still understandsPhilip [00:17:25]: Close enoughAli [00:17:26]: That this is a scientist who is a man who has, some significant achievements, all that stuff. So that's like really cool.Philip [00:17:32]: Yeah. So, we've covered Hao Tian before, who the author of the LLaVA paper that did this, a while ago. And I think that's very foundational work for anyone who hasn't done vision work before.Ali [00:17:41]: Same with the CLIP and MetaCLIP, where you go from just captioning to building out questionsPhilip [00:17:47]: RightAli [00:17:47]: Off the image and how much better you can get performance.Philip [00:17:50]: Right. Right. Right. Yeah. But what's, what's so exciting about this is if you look at a model like this. Now, this is a little bit more of a research project. It's not. It got to 56% on MMLU Pro, I think. So not quite frontier. But if you're running this model, you haven't suffered any loss on your GLM-5.2 quality. If you don't have an image, it'll just behave exactly the way it used to. And ultimatelyAli [00:18:14]: Which in the inference code you literally do not include the other part, right?Philip [00:18:18]: Yeah. You would just skip the encoder if you don't have an image input.Ali [00:18:22]: Okay.Philip [00:18:22]: Just confirming.Philip [00:18:23]: YeahAli [00:18:23]: Does it affect a lot on the overall inference side? Like you're not adding much, you're adding a very small vision encoder. These are typically likePhilip [00:18:30]: They're super fineAli [00:18:31]: Less than a billion parameters, right?Philip [00:18:32]: Yeah. It's, - There's a little bit less standardization among vision encodersSwyx [00:18:37]: YeahPhilip [00:18:37]: So the support matrix can be a little bit, sparser. But overall, yeah, it's a pretty, it's a pretty minor component of the overall system. And ultimately what you get out of the system is all of a sudden you have Kimi Vision, GLM weights, and DeepSeek attention all in one model.Open Source Model Grafting and Franken-MergesPhilip [00:18:56]: And that's, I think, a lot of the power and beauty of open source, is that you can take all of these different components and combine them together into a system that's better than anyoneSwyx [00:19:05]: YeahPhilip [00:19:05]: Can be individually.Swyx [00:19:06]: People used to say that you would also do Franken-merges where you would take likePhilip [00:19:10]: YeahSwyx [00:19:10]: Layers from each model.Swyx [00:19:11]: Does anyone do that anymore?Ali [00:19:13]: Well, to your point previously when you were mentioning like, the work that goes into supporting a model when it first comes out, like GLM-5.2 or MiniMax M3 or whatever the case is. Sometimes you do have to like, you do have to switch out some things. Like, for instance, the MiniMax M3 head uses full attention, and with full attention you end up with this like insane bottleneck in spec dec ‘cause you're doing auto-regressive token generation for three tokens, and you're doing this like N squared over all of the tokens that are in your sequence. Your KV cache is like very large because it's not sparse, it's not top K. So we find it better to like, okay, we're gonna replace this, we're gonna replace this layer with a layer from another model that's using like GQA, for instance. And then just with the right training, you can get it to have the same acceptance rate. So it is very possible to retrofit layers from other models and very much needed. If a layer is like inefficient, the training just becomes the challenge, like how do you ensure that you train it properly? Which again to your earlier point is like the mesh between training and inference. As in like you need very good training in order to do fast inference. That's like, I feel like more and more becoming true.Swyx [00:20:21]: Yeah. Anything else on the support side when you say like get it to fully production ready?Loop Detection, Race Conditions, and Non-DeterminismPhilip [00:20:26]: Yeah. I think that there's also a question of just, we can test a model to a pretty extensive degree, but we're trying to get it out quickly and then you see a bunch of other people test it and you get interesting results. There was an issue with, GLM briefly where we had some like mode collapses where it would just output the same token over and over again for certain prompts on certain temperatures. Like once you expose an endpoint to the real world, there's going to be, so many more varieties of things given to it that you're able to, discover and patch things. So it's not just a, day zero process, it's then like for the first week, for the first month, if a model remains popular, like how do you both fix bugs and then continue to push the envelope on performance?Ali [00:21:21]: What do you mean you don't want your model outputting S?Swyx [00:21:24]: Is there loop detection on that stuff, by the way? It still happens like quite a lot, which is surprising.Ali [00:21:30]: We have like we, in our endpoint, like if a model was to output the same token like four plus times, we just cut the generation. We say like, “Oh, sorry, this-- Like try again,” or like we will reprocess the request. ‘Cause we know then, like if it, like if, yeah, it's four times the same token, it's probably collapsed.Swyx [00:21:45]: Yeah. Is there a way to opt out in case I really want that?Ali [00:21:48]: You want that?Ali [00:21:50]: I think there's a way that we have to handle it. I'm not exactly certain, but I feel like in certain models, like when they output something like you can imagine, like a table for instance, and so they want, they wanna draw like 12 dashes and 12 dashes. Yeah, I think there's a way for that to happen. I think we only do it on certain tokens. Like we exclude certain special characters.Swyx [00:22:07]: Yeah.Ali [00:22:07]: So we only do it on like certain like S is the most common almost. GLM-5.2Swyx [00:22:11]: OhAli [00:22:11]: And I think it was DSV 4 as well. Like you'd just have like looping issues where like you literallySwyx [00:22:17]: ItAli [00:22:17]: Just have like S.Swyx [00:22:18]: Yeah. Is there a special, something special about S? No, just randomlyAli [00:22:21]: It just seems to be the one token involved.Swyx [00:22:23]: Yeah. And it'Philip [00:22:24]: Is thereSwyx [00:22:24]: And it's only temperature 0Ali [00:22:27]: NoSwyx [00:22:27]: Even at other temperaturesAli [00:22:27]: Even at like 0.9 or whatever, it will still, it will still collapse.Swyx [00:22:30]: That's weird, right?Ali [00:22:30]: It's, it is an inference problem to be honest, like a software problem. Like oftentimes, the image you run will-- like NVIDIA will release an image for instance, and if we will upstream the changes from their latest TensorRT-LLM image into our stack, we'll find that it fixes it. Or oftentimes this will only happen in an inference engine that you're using like SGLang. But if you were to switch to vLLM, that isn't the case. So it seems to be like an extremely like deterministic software issue and not really a model issue. It's not like a weights problem. Like I'- we'll say like, “Oh, it's a problem with the quant. We did PTQ wrong,” right? But that isn't, that doesn't make sense because the same weights used with a different inference engine does not repeat the problem. And sometimes it's, the kernels that are being used in the backend have like these very subtle sometimes race conditions, where if you were to use this model hosted on one cluster, you will never get this problem.Swyx [00:23:19]: Oh my God.Ali [00:23:19]: But if you host it on a different cluster, you will. And the reason is the KV cache transfer from a node to node in that one cluster is using a slower interconnect than the node to node in another cluster. So that exposes the race, whereas in another cluster it doesn't. So then you end up just like, okay, this model is not gonna be hosted on this cluster. We're gonna host it on, another cluster because that cluster exposed that problem. But then it ends up with like, okay, is it the software? Is it the model weights or is it the hardware?Swyx [00:23:42]: There is a thing about this with temperature 0 still not being deterministic, right?Ali [00:23:46]: Right.Swyx [00:23:46]: Mostly because of hardware. Even at temperature 0 same model, you won't always get the same output.Swyx [00:23:52]: Even-- But I'm surprised by the race condition one because, I thought PyTorch was a graph that like guarantees that you at least, execute things in the right order.Ali [00:24:02]: Well, yeah, true. Like I'm not, I'm not saying that there is. Like well, you have things like PTL optimizations where like you can start a kernel before the end of the previous kernel, and that's like ‘cause you want to do that because there'sSwyx [00:24:12]: It's like pipeliningAli [00:24:12]: Expense. Exactly.Swyx [00:24:13]: Yeah.Ali [00:24:13]: But it'- But you don't do it cleanly. Like you overlap a little bit of the execution. No, it is very possible that the kernel itself, like that one block that is supposed to be running in this instance of time, that kernel itself has a race condition. For instance, like a missing barrier. Like often if you're designing a kernel and you want it to make it to be very fast, if you don't test it extensively, you'll, you'll have certain threads access data points from registers before they've been written to by other threadsSwyx [00:24:36]: YeahAli [00:24:36]: For example, because like your barrier is wrong or your synchronization was wrong. But yeah, like the testing itself is very difficult in those like, andSwyx [00:24:42]: And there's no like borrow checkerAli [00:24:45]: What does that mean?Swyx [00:24:46]: Like Rust. Like the. If you're trying to have like memory safety It sounds like a comparable problem.Ali [00:24:52]: Well, yes, but you're working in CUDA, right, NVIDIA GPUs. Like- You just need a higher level language like modular Maybe that's what modular is supposed to do. I don't know.Quantization Quality and Vendor FidelityVibhu [00:25:00]: How do you see keeping quality of the model? So you talked about all these steps of, okay, you gotta do quantization, train your own speculative decoderAli [00:25:07]: RightVibhu [00:25:07]: Run on different hardware. Looking at other model providers, okay, you kicked off a inference speed race on the consumer end. What goes into keeping quality the same across them, right? Sure, you can run benchmarksAli [00:25:22]: YeahVibhu [00:25:22]: But, like, how do you determine how much quantization are there standards? What goes intoPhilip [00:25:27]: There's a few things on quality. Most inference optimizations are lossless. KV caching, for example. You are just recomputing or preventing recomputing the same values. Speculation, of course, if a draft token is wrong, it gets rejected. The main lossy optimization is quantization. And that really comes down to, number one, data format, number two, which parts of the model you choose to quantize, which layers, and number three, like doing a lot of calibration on the quantized weights, to ensure that you're preserving all the outliers. There's other tricks that you can do, though. A big one is long context, ‘cause one thing you asked at, right at the beginning is, “Oh, what's gonna happen if I send a 200,000 token request in?” So with a long input sequence, you need to, store a lot more information. You need to process a lot more tokens. And so even if a model has a context of a certain length, you might, as an inference provider, choose to build an API with a shorter context length, and of course a full length one as well. Because if someone doesn't need the full million token context, for example, you can get them better performance. I don't know if that's exactly like quality of the model. The way that I think about quality is to what degree are we faithfully serving the original model? If you think of a golden implementation of a model that performs exactly the way the model is designed to perform, I think of quality as how close are we getting to that, 100% fidelity of the model.Philip [00:27:13]: You can also, of course, think about quality from the training side and how do you push yourself past 100%. But when I think about purely inference optimizations, it's getting faster while staying as close to that 100% fidelity mark as possible. And certainly our standard internally is that, like you should not be able to tell the difference between our API and a, official API. I think Kimi in particular does a good job of vendor benchmarking hereAli [00:27:41]: YesPhilip [00:27:41]: Where they haveAli [00:27:42]: They released an actual vendor benchmark.Philip [00:27:43]: Exactly, yeah.Ali [00:27:44]: ‘Cause they accused, some people, Amazon? There was some provider that was not doing very well on Kimi's benchmark.Philip [00:27:50]: Yeah.Philip [00:27:51]: So, with Reflect we probablyVibhu [00:27:52]: This was a long time ago, right?Philip [00:27:54]: No.Ali [00:27:54]: Yeah, like threeVibhu [00:27:55]: They alsoAli [00:27:55]: Four, five months agoVibhu [00:27:57]: This also happened with, I don't remember which model, but they pulled out quite a few, and then they started a whole chart about this. It might have beenPhilip [00:28:03]: Kimi Vendor Verifier.Ali [00:28:04]: Yeah.Philip [00:28:05]: Yeah.Ali [00:28:05]: Yeah, ‘cause you, ‘cause you'd be pissed, right? Like if you'Philip [00:28:07]: Yeah.Ali [00:28:07]: If like if I'm a consumer and I'm using like Amazon's endpoint for instance, and I've used Kimi and I'm like, “Oh my God, like this is bad,” I'm not gonna say, “Oh, Amazon quantized the model in a bad way.” I'm gonna say, “Oh, Kimi sucks.” Right?Philip [00:28:17]: Yeah.Ali [00:28:17]: So it seems like that makes sense.Philip [00:28:19]: Yeah, they care. They care.Vibhu [00:28:21]: Justifiably.Ali [00:28:21]: Yeah, justifiably.Vibhu [00:28:22]: This is probably a stupid question, but just checking, has anything improved from main quantization?Philip [00:28:28]: Yeah.Vibhu [00:28:28]: Like, is quantization always strictly worse?Ali [00:28:30]: Well technicallyVibhu [00:28:32]: NoAli [00:28:32]: It's a lossy. QuantizationPhilip [00:28:33]: YeahAli [00:28:33]: Is a lossy, it's a lossy implementation.Philip [00:28:36]: Speed improvesVibhu [00:28:36]: Speed improves.Ali [00:28:37]: It the number, likeVibhu [00:28:38]: No, I' always look for inverse scaling laws.Philip [00:28:40]: Yeah.Ali [00:28:40]: Yeah.Vibhu [00:28:40]: This is something I learned from Noam Brown, where like things that normally act in one direction sometimes do.Philip [00:28:45]: Well, technically when you run a benchmark, because these models are deterministic, sometimes your,Ali [00:28:52]: YeahPhilip [00:28:52]: NVFP4 quant is like, two basis points higher than yourAli [00:28:56]: No, it's noise. It's noise.Philip [00:28:57]: Yeah, exactly. I'm like, yeah, it's, it's within. That's why I always say within margin of error.Philip [00:29:01]: And I stopped saying that because everyone assumes that what is, well, within some margin of error, we're barely inside of that to the worst, so we're saying. But yeah, sometimes it's just like, gives you a higher output score. But like Ali said, that's noise. To my knowledge, you're not necessarily making the results better. You're just trying to, again, like keep your fidelity as close to 100% to the original model.Layer Selection, KL Divergence, and Better QuantizationAli [00:29:27]: There is, to your point, research that we did on MP. I don't know if you are able to pullPhilip [00:29:31]: YeahAli [00:29:32]: A tweet we did. One of our research interns, Joshua, I think it's a tweet on how we have 20% better quantized GLM-5.2 than NVIDIA. Essentially what we found throughout like this month research is, okay, quantization is a lossy. It's. You're compressing the data from, occupying 16 bits to occupying, four bits, for instance. And so you're losing some information, and you're trying to minimize that. And so when I say that I'm gonna quantize the model, my job becomes how do I find the layers that I can quantize, and how to find the layers to not. For instance, with image models, I don't quantize modulation layers, and I don't quantize out projections because those two are. Like out projection is what you see as the user. Modulation is what the model sees or understands. Right, exactly. And so to his paper, do you have the. It doesn't have the. Yeah. It's a long paper. I don't know if I can findVibhu [00:30:25]: If there's a part to search or it's probably in the thread.Ali [00:30:28]: It's probably in the thread.Vibhu [00:30:29]: Yeah.Ali [00:30:29]: But the long and the short is it is very possible that quantizing more of the model makes the results. Like if I have a model that I quantize layers one, five, and 10, and another model where I only quantize layers one and It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in, is that you can predict which layers are going to have quantization errors that will cancel out with each other, and you choose to quantize those layers. And so the result of doing this mathematical quantization is you end up with a model that's 20% more quantized than another provider, so you get 20% more throughput of it because there's more layers than running an NVFP4, and your quality is better than that other quant because the layers that you chose to quantize have their errors cancel out, like one layer skewed to the right one layer skewed to the left, one layer skewed to the right. Your final logits distribution is more similar to the original distribution of the model, so you have better fidelity. And so the way we proved this was with KL divergence. So instead of just scoring on the benchmarks, we scored the KL divergence between the logit distribution of the quantized model and the logit distribution of the original full precision model, and we showed that with this technique we get. If your probability distribution on the logits which token it wants to select is more of the same as the original model, you're probably gonna end up staying true to the original model. So yeah, so it seems like previously before this, it seemed like the industry was, well, the more you quantize, the worse it's gonna be, ‘cause the more loss you introduce. That's not exactly, not necessarily true. So yeah, doesn't improve it, but can cancel out.Philip [00:31:57]: I think it might be this, but reminds me a good bit about pruning where you can prune off certain layers.Philip [00:32:03]: But very interesting. Didn't know this was a whole paper you guys put out.Ali [00:32:06]: It's. Fun fact, it was originally 72 pages, this paper, and then we decidedPhilip [00:32:11]: WowAli [00:32:11]: We can't tell. We couldn't release it. So it's now 45.Swyx [00:32:15]: Still 39 pages, so very substantive. We talked about evals and all these things and, like what's possible in terms of speedup? Like it's like probably like the numberInference Speedups and BenchmarkingSwyx [00:32:25]: Thing that people do wanna care about, and it's something that you wrote about in your post. Like official API is 70 tokens per second, and you push it up to 90. Is that like a normal thing?Philip [00:32:36]: So what's cool about working in inference, the reason that I think inference is going to be a useful place to do engineering for a long time, is that if you look at highly optimized domains like, say, finance, if you're in finance, you measure how much better you got in basis points. It's like, “Oh, I got five basis points better, like twentieth of 1% better,” that's huge news because everything is so optimized. When we publish optimizations, it's 20%, it's 100% it's 200%. So there's still probably like a lot further to go, honestly. Like you'll, you'll know that inference is pretty much solved when researchers start publishing about how they got 1% faster at something.Swyx [00:33:19]: Which by the way, because I am from the finance background, in the ‘70s, that was the margin at the time. When you did quantitative finance research, you would findAli [00:33:27]: And like 20%, tens of percent.Swyx [00:33:29]: That's. Yes.Philip [00:33:29]: Yeah.Swyx [00:33:30]: And now it'Philip [00:33:31]: Tiny fractionsSwyx [00:33:32]: For those people interested, look up Andrew Lo's paper. He had a really interesting illustration of quant, stat arb, distribution, narrowing down from like those kinds of 20% differences in the ‘70s, down to nothing today, which is very cool.Philip [00:33:48]: Exactly, and we're at the beginning of the same type of thing. Now benchmarking is hard. I think anyone will tell you that, and benchmarking provider speeds is hard because there's so many variables that go into it. What hardware are you using? How much load do you have on the system? What's the exact nature of the prompts and input and output sequence lengths? All that stuff. But overall, when you start stacking these improvements, you're looking at multiples. You can look at it. The most common form, of course, is TPS, tokens per second, which is bad naming by us in the industry, ‘cause there's two tokens per second. There's tokens per second, the throughput number, and the latency number.Ali [00:34:31]: TTMT, yeah.Philip [00:34:32]: Like total tokens per second out of the, out of the GPU as a throughput number. Most people only care about tokens per second as the latency number, which we should call ITL, intertoken latency, but we don't.Philip [00:34:44]: Anyway, so you can imagine a standard API without many optimizations for a 1 trillion parameter model operating somewhere in the 30 to 50 tokens per second range for reasonable traffic profile. And we generally see the goal of, pushing to 10X that. But, not necessarily day zero, but by stacking enough optimizations, if you have, say like four optimizations, each of which doubles performance. Or sorry, three optimizations, each of which doubles performance, then you stack that up, that's an 8X gain. That's the order of magnitude that we're working with in this space. We're trying to make things substantially faster, not just go from like 70 to 90.Swyx [00:35:38]: Are you saying you've. You have done that?Philip [00:35:40]: So let's say you have as a reasonable baseline, 30 or 40 tokens per second. You can achieve 10X that. So like on GLM-5.2, if you run it unquantized, perhaps on H100s even, and you're just using an off-the-shelf inference engine with no particular optimizations, no speculator, nothing extra around like KV routing, no disaggregation, you're, you're probably, yeah, looking at that like 30 to 40. You think that's like a reasonable baseline?Swyx [00:36:12]: Right. Right.Philip [00:36:12]: To get to something like 10X, there's a lot of trade-offs that you're making. If we're running at more like a 300, 400 tokens per second range, you are using the best hardware possible. You have a optimized speculator. You have done all of your quantization work. You are Seeing a pretty high cache hit rate. You are running with a reasonably small batch size and a parallelism configuration that is tuned for latency versus throughput, but it is possible. So the spreads that you see if you, like, go on artificial analysis or you go on OpenRouter and you look at, the worst provider to the best provider, oftentimes can hit that range. 10X is of course very aggressive. It's oftentimes maybe more of a four to six times improvement. But that's the performance that makes us really excited, is when we can get these huge gains, not just go from 70 to 90 tokens.Stacking Optimizations: NVFP4, Speculation, and DisaggregationAli [00:37:19]: It's also, like, hardware dependent. Like, ifPhilip [00:37:20]: YeahAli [00:37:20]: If you have a thing where you're serving it on just, like, a node of H100s and then you throw, like, you shard the model across, like, four nodes of B200s. Like, you can definitely increase the speed with just throwing more hardware at it. Like, normalizing for the same exact hardware and the same number of GPUs.Philip [00:37:35]: Yeah. Then you're looking at, like, a two to 4X improvementAli [00:37:38]: Right. RightPhilip [00:37:38]: Depending on the inference optimizations. So yeah, it's. Some of it's, what's the call, and some of it's who's the driver.Vibhu [00:37:46]: If you break down the two to 4X, say the example is run GLM-5.2Ali [00:37:51]: YeahVibhu [00:37:51]: On B200sAli [00:37:53]: YeahVibhu [00:37:53]: Single node, right? What's, like, the cost trade-off for effort to get, like, the last bit of juice out versus what should people just think of, right?Ali [00:38:01]: Spectre quantization. Yeah.Vibhu [00:38:03]: Spectre quantization.Ali [00:38:04]: That's, that's, that's like 95%. LikeVibhu [00:38:06]: And how far does that get you? And how easy is that for the average person to do? So say right I wanna throw the weights of GLM-5.2 on a node of B200s, how easy is it to find speculative decoder- decoder model or already quantized model? How much work goes into it?Philip [00:38:23]: If you're doing it up front, it's quite a lot of work. If you're doing it today, there's going to be people who have published things that you can just, you can just grab some NVFP4 weights. You can grab a speculator. Yeah, if we're thinking about, like, what are the 2Xs we're stacking, going from, BF16 to NVFP4 is, it's not quite a 2X, right? It's like. I think it's about, like, 30 to 40%, from 16 to 8, and then another 30 to 40% multiplied from, 8 to 4. So that doesn't quite get you a 2X, but, like, roughly a 2X. Speculator, roughly a 2X. Disagg on top of that if you're able to get enough hardware and put enough traffic through it, another roughly a 2X. And then you add in some, double-digit percent increase from having just a better runtime with, the latest kernels and stuff behind it. And that's how it stacks up.Ali [00:39:21]: YeahPhilip [00:39:21]: So building each of those, like, building the, quantized weights is, for someone who really knows what they're doing, hours to days of work. Building the speculator, again, like, hours to days of work. And the, disagg setup, hours to days. Well okay, but like once you haveAli [00:39:39]: Once set up. Once set up. YeahPhilip [00:39:40]: Yeah, getting disagg working for the first time, I'm saying, of course, is very difficult.Philip [00:39:44]: The marginal implementationAli [00:39:48]: Like, if you're just grabbing, like if you are a person, like just a normal consumer who has access to, like, a node of B200s and you're wondering, “How can I just host it myself?” You don't need to quantize the model yourself. There's always gonna be, like, an open source quantized checkpoint. NVIDIA's gonna push one out if no one else does. You. Usually, the providers will have their own spec dec that they've trained as well. You don't need to train your own spec dec. You can just use that as well.Philip [00:40:09]: Yeah. Like, GLM-5.2 has its own MTP.Ali [00:40:13]: Right. Right.Vibhu [00:40:14]: What's multi token prediction?Philip [00:40:15]: Yes.Ali [00:40:16]: I'm justVibhu [00:40:16]: Can you explain that?Ali [00:40:16]: I'm just an expert.Ali [00:40:18]: I can do it for you in case I get it wrong?Vibhu [00:40:20]: No.Vibhu [00:40:21]: Yeah, you should correct if we're wrong, but their multi-token prediction can be used for self-speculative decoding.Ali [00:40:27]: I'm not sure. I'm not gonna correct that.Vibhu [00:40:28]: Okay. I'm semi-confident in thatAli [00:40:30]: Okay. YeahVibhu [00:40:30]: But someone can check. But it's useful to paint the story of, okay, not just the average person, but say a company wants to switch from serverless inference I wanna throw this up on. I wanna rent some GPUs, throw it up. These are the steps you take to do significantly faster than just put it behind vLLM.Ali [00:40:48]: Right.Vibhu [00:40:49]: I was waiting for a mention of Dynamo.Vibhu [00:40:51]: I feel like, that's supposed to be the baseline that you measure against.Dynamo, KV Routing, and Disaggregation ToolkitsPhilip [00:40:55]: I would think of Dynamo as less of a box system and more of a toolkit for building with. So when we talk about doing aware routing, when we talk about doing KV offloading, when we talk about doing, PD disaggregation, Dynamo fundamentally is. By the way, Dynamo is an open source library from NVIDIA.Ali [00:41:17]: We've done a pod with KylePhilip [00:41:18]: OkayAli [00:41:19]: Kyle Cranin.Philip [00:41:19]: Cool. So then your listeners know then that it supports all the different inference frameworks. And it is multi hardware, which is interesting.Ali [00:41:28]: But it's just a router, it's not like an optimizer layer.Philip [00:41:30]: Yeah. All it does, like, what Dynamo is good at, it is a library for moving information around your cluster, around your hardware. So if you have, KV cache on one place and you need it to be somewhere else, Dynamo coordinates NIXL for you to move that around.Philip [00:41:49]: That doesn't mean that, like, out of the box, you just say, “Pip install Dynamo,” and then you get, like, a massive performance speed up. It's more of a developer toolkit.Ali [00:42:01]: Yeah. I would have said it would. It comes with a set of defaults that you can then swap out.Philip [00:42:06]: It does. If the industry at large, I think, was, like, rolling out all of these deployments, standard, then I think it would be, like, a credible baseline. But, we've got to, we've got to benchmark against, like, what we're seeing in the wild.Speculative Decoding Methods: Medusa, EAGLE, n-Gram, and Spec-SpecVibhu [00:42:23]: I did wanna talk a little bit more about PD disagg, because that is probably, like, number three after quantized and speculative decoding. In your book though, I was just gonna pull out the book.Philip [00:42:31]: Yeah.Vibhu [00:42:32]: Like section 522 on Medusa, 523 on EAGLEPhilip [00:42:35]: YeahVibhu [00:42:36]: 524 on gram.Philip [00:42:37]: It's 55, would be disaggregationAli [00:42:42]: Yeah. Well, no, I just wanted to dwell a little bitPhilip [00:42:44]: YeahAli [00:42:44]: The other. Like, so what do you choose to include? What do you choose to not to include? Because there was all these other techniques.Philip [00:42:51]: Yeah.Ali [00:42:51]: Are these still relevant? Because I think they came out, like, a year and a half ago maybe.Vibhu [00:42:55]: Medusa is quite old.Philip [00:42:56]: Yeah, Medusa's old.Ali [00:42:58]: It was old.Vibhu [00:42:58]: But is it in the book as a good, here'sPhilip [00:43:01]: BaselineVibhu [00:43:01]: Baseline vanilla understand it?Philip [00:43:02]: Like you should know this.Vibhu [00:43:03]: Like I read the paper, I'm like, “ it makes so much sense.”Philip [00:43:05]: Yeah.Philip [00:43:05]: So with the book, I had a couple goals. One was to give people just a working vocabulary for the space as a whole, and the other was to give them some intuition about how each of these techniques works. As I mentioned in my AI Engineer talk, which is the first public addendum to this, the speculation space has moved much faster than everything else. So yeah, even at the time that I wrote the book Medusa, I very much included as a way for people to understand how the space evolved rather than what the most modern technique is. And now of course, there's DFlash, dSpark. There's, there's newer techniques even than EAGLE, although EAGLE is still very commonly used.Ali [00:43:51]: SpecSpecta.Philip [00:43:52]: Yes. Speculative decoding.Vibhu [00:43:54]: What canAli [00:43:56]: Oh, it's a paper by Tri Dao and it's like, it's doing speculative decodingVibhu [00:44:00]: HuhAli [00:44:01]: For the speculative decoder.Philip [00:44:02]: Oh, in spec- oh my God.Ali [00:44:02]: It's literally just an another. It's like, yeah, that's the most simple way to explain it, and it seems like he got trivial speed ups there. But it seems that the complexity with training, it's almost like in our mind at least, it's almost as complex as training GANs. Like it's like a very delicate balance and oftentimes you, it's just but yeah, it's literally speculative decoding on speculative decoding.Vibhu [00:44:21]: Speculative.Ali [00:44:22]: Yeah. We saw this paper.Vibhu [00:44:24]: It's interesting, right?Ali [00:44:24]: Yeah.Vibhu [00:44:24]: I wouldn't even expect it to be very particular to train, I wouldAli [00:44:29]: Right.Vibhu [00:44:29]: The naive part of me is like, okay, train speculative decoder.Ali [00:44:32]: But like, and it makes sense, like the whole idea of speculative decoding is you. It's like, it's like almost like the iPhone auto predict version but for a normal model, right? Like you're just, you're just, generating three tokens and you're like, okay, I'll do prefill on them. And so you save those three turns for your original model. Now your speculative decoder is doing three turns of auto regression, so why not just have an even smaller model?Ali [00:44:53]: The other question there is what are the size of speculators? So say forPhilip [00:44:58]: Right. It's like a billion parameters.Ali [00:45:01]: Like for MiniMax, it's. Yeah. It's like one layer. It's like one 60th of the original model usually.Philip [00:45:06]: Yeah. I think we should do a paper when we get back to the office.Philip [00:45:10]: SpeculativeAli [00:45:11]: SpeculativePhilip [00:45:11]: Decoding.Ali [00:45:13]: No, it's, it does seem like how, when do you stop? But then it also seems like if you're able to train spec-spec decode for instance, right? Like if you're able to have a small model that is accurately predicts what the intermediate speculator is gonna predict, that is able to predict what the original target model's gonna predict, then why not just use that smallest model directly, right?Vibhu [00:45:34]: Yeah. This isAli [00:45:35]: Like it seems likeVibhu [00:45:35]: Adjacent to the routing problem.Ali [00:45:36]: Right.Vibhu [00:45:36]: Yeah.Ali [00:45:36]: Right.Philip [00:45:37]: The thing with speculators is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you're running the big model on. There is a orchestration and resource competition problem inherent in that, and that is one of the constraints on speculation in general, is that draft tokens cost resources to create and cost software complexity to manage. And so if you have like infinitely recursive speculators, you add in quite a bit of that complexity on the actual implementation within the inference engine as well, not just in the training process.Vibhu [00:46:17]: I was gonna say, I would wonder if you could do similar, like distillation and pruning of, it's the same thing, it's just a model. Can we not just distill a lot of the weights, quantize the speculator, out of my domain? The question that also comes up is, this is all for big server workloads, right? How much of this applies to, say I have this MacBook, I wanna run Gemma really efficiently. Similar problems, not the same?Local AI vs. Data Center InferencePhilip [00:46:45]: Pretty different. I talked to Selo, about this on his podcast a couple weeks ago. The difference between inference engineering for the data center and for production workloads versus inference engineering for local AI, is that we start with fundamentally like different constraints and different goals. With local AI, it's how do I fit this model onto my hardware and then make it less dumb? And with data center influence, it's how do I load this model and then make it less slow? And we care about less dumb, and they care about less slow. But the local AI inference engineering ecosystem, I think has a lot for us to learn from in the data center space. They are experts in various forms of quantization, including dynamic quantization that we just don't touch, in the pruning, in the distillation, in the, layer removal. There'Ali [00:47:42]: Layer removal matters less.Philip [00:47:43]: Yeah. There'Ali [00:47:44]: No one loves pruning really.Philip [00:47:45]: Yeah. Well, but the, but they doVibhu [00:47:46]: Which is surprising, right? But that's, that's a whole different thingPhilip [00:47:48]: Just to fit something on the laptop.Ali [00:47:50]: Right.Philip [00:47:50]: So yeah, it's a, it's an interesting, it's an interesting space. Not necessarily that like their techniques make sense for us to do in the data center, because we have different resources and different goals, but more that the process as well as the openness of that field is something to, admire.Ali [00:48:12]: Yeah. Like to your point, like, certain optimizations that would. Like for instance, Turbo Quantum Sharper, like it made such huge hype on that and we did like a whole deep dive on Twitter and like said, what is it? How does it work? Why is it good or not? And it took off and it was implemented on local devices because your memory bandwidth is so slow on like a MacBook, for instance. But try putting the same thing on like an NVIDIA GPU on a B200 Turbo quant would not be. Like, it would not be used. Like, NVIDIA - Like, NVIDIA made it clear that this is not a good optimization, and we've seen it firsthand where the overhead of doing dequantization, quantization of, in the kernel itself with turbo quant kernel, each end is much slower than the time that you save from doing the bandwidth. ‘Cause on the B200s, you have like 3.5 terabytes per second. You don't need decrease the storage that much. You don't need to do, FP4 KV cache. You don't need to use a requant. There's, there's, there's better optimizations to be made. But on Edge devices, it's extremely important, it's extremely useful. So, seems to be, like, different optimizations there, but then they're all uniquely combined with like all you wanna quantize the model, you wanna do speculative decoding, like certain common prefixes with bothPhilip [00:49:18]: Principles.Ali [00:49:19]: Yeah, exactly. Exactly. Exactly.Philip [00:49:20]: They also do a lot of work on, model parallelism, especially over, heterogeneous topology, where you have, some sparks and they are wired together with, Ethernet, DGX sparks.Ali [00:49:35]: Yeah, this is the Exo Labs guys.Philip [00:49:36]: Yeah. You have, a nu

Everyday AI Podcast – An AI and ChatGPT Podcast
Ep 830: Faster AI Agents, Fewer Human Coworkers: The Overly Productive Future of Managing Agents?

Everyday AI Podcast – An AI and ChatGPT Podcast

Play Episode Listen Later Jul 30, 2026 32:56 Transcription Available


Agents are getting more powerful by the day. And most workflows, outputs and human capabilities can't keep up. Is that a problem or opportunity? Before you answer that question, though, keep this in mind. Agents are *literally* about to become 20X faster overnight. Let's unpack what that means. Faster AI Agents, Fewer Human Coworkers: The Overly Productive Future of Managing Agents? -- An Everyday AI Chat with Jordan WilsonNewsletter: Sign up for our free daily newsletterMore on this Episode: Episode PageToday's Episode on LinkedIn: Thoughts on this? Join the convo on LinkedIn and connect with other AI leaders.Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineupWebsite: YourEverydayAI.comEmail The Show: info@youreverydayai.comConnect with Jordan on LinkedInTopics Covered in This Episode:Managing Dozens of Productive AI AgentsOpenAI Cerebras: 20x Faster Agent ModelsImpact of AI Agents on Human CoworkersAgent-Driven Workflows vs. Human CollaborationIncreasing Agent Reliance and Fading MentorshipAccidental Deskilling and Compression TaxProtecting Human Judgment and Learning HandoffsExpert-Driven Loops in AI WorkflowsMonthly Rebuilding of AI Strategies and ProcessesMiddle Management Evolution in AI Native CompaniesTimestamps:00:00 Future of AI and Work Dynamics05:25 Advancements in AI and productivity tools09:51 Growing your business with AI13:52 AI productivity and collaboration shifts15:08 Improving AI processing speed18:53 Using AI agents for delegation24:46 Discussing AI-related work challenges28:16 Ensuring accountability and communication30:39 Adapting to rapid digital change32:35 Show outro and newsletter sign-upKeywords: AI agents, faster AI models, OpenAI, Cerebras chip, 20x speed increase, automated workflows, agent management, solo agent supervisor, generative AI, knowledge work automation, agent-powered productivity, parallel machine teams, inference speed, productivity acceleration, Codex, Cloud Code, Google Gemini, Cloud Cowork, Copilot, recursive self improvement, expert-driven loops, human handoffs, deskilling, mentorship loss, AI native workplace, workplace automation, transactional work, productivity roadblocks, accidental deskilling, agent bun sandwich, compression tax, human in the loop, expert collaboration, agent trust, AI decision making, domain expertise, rapid workflow rebuilding, unlearning processes, organizational adaptation, enterprise AI adoption, future of work, middle management AI, AI-powered teamwork, human-agent collaboration, manager-agent ratios, personalized agent output, multi-agent coordination, skillset sharing, intentional automation, productivity strategySend Everyday AI and Jordan a text message. (We can't reply back unless you leave contact info) Ready for ROI on GenAI? Go to youreverydayai.com/partner 

Revenue Builders
Why Partner Leaders Must Master the AI Landscape to Win Hyperscaler Partnerships with Alan Chhabra

Revenue Builders

Play Episode Listen Later Jul 30, 2026 62:14


AI partnerships are no longer won by managing a channel or securing a marketplace listing. Leaders must understand how countries, power, cooling, chip fabrication, hyperscalers, model providers, software companies, and enterprise risk fit together before they can build an effective partner strategy. Returning to Revenue Builders, Alan Chhabra explains how Cerebras navigated the technical and political realities of working with AWS, why successful partner leaders need more than relationships, and how sovereignty is changing infrastructure, model selection, and cloud decisions. He also examines why software companies are moving aggressively on AI while traditional enterprises remain more cautious. Alan Chhabra is the Executive Vice President of Worldwide Partnerships at Cerebras Systems, where he leads global partner expansion across hyperscalers and the broader AI ecosystem. He previously spent nearly a decade building MongoDB's global partner network and held senior leadership roles at BMC Software. Connect with Alan: LinkedIn Other Revenue Builders episodes with Alan: The Ideal Partnership with Alan Chhabra Key takeaways from this episode:  00:00 - Why the infrastructure behind AI may matter more than the applications built on top of it. 12:26 - What the shift from data to AI changes about how partner leaders map the market. 16:06 - Why understanding the full AI stack is a prerequisite for building the right partnerships. 27:48 - A look inside what it really takes to navigate the politics of a hyperscaler partnership. 36:00 - The four capabilities CEOs should test before hiring an AI partner leader. 48:34 - What leaders often overlook about the role of sovereignty in shaping AI infrastructure decisions. 56:25 - Why software companies and traditional enterprises are moving at very different speeds on AI adoption. Hosted by five-time CRO John McMahon and Force Management Co-Founder John Kaplan, the Revenue Builders podcast goes behind the scenes with the sales leaders who have been there, done that, and seen the results. This show is brought to you by Force Management. We help companies improve sales performance, executing their growth strategy at the point of sale. Connect with Us: LinkedInYouTubeForce Management

TechStuff
Cerebras Built a New Way to Make Sand Think - The Story

TechStuff

Play Episode Listen Later Jul 29, 2026 33:37 Transcription Available


What does it take to power the AI boom? Steve Vassallo, a roboticist turned General Partner at Foundation Capital, helped semiconductor company Cerebras figure it out. Steve and Oz discuss why AI workloads demanded an entirely new kind of chip, what "wafer scale" actually means, and how a decade of near-failures — including a system that literally caught fire — led to the largest semiconductor IPO to date and deals with OpenAI and AWS.See omnystudio.com/listener for privacy information.

The MAD Podcast with Matt Turck
The Biggest Chip Ever Built — Why OpenAI Runs On It | Cerebras CEO Andrew Feldman

The MAD Podcast with Matt Turck

Play Episode Listen Later Jul 23, 2026 72:41


AI is no longer just a race to train smarter models. As AI moves into production, the bottleneck is increasingly inference: how fast models can generate tokens, use tools, reason, verify, and act. In this episode of the MAD Podcast, Matt Turck sits down with Andrew Feldman, co-founder and CEO of Cerebras, to explain why fast inference may define the next era of AI.Cerebras is known for building a chip the size of a silicon wafer. But this conversation is not just about one company or one chip. It is a deep dive into the AI infrastructure stack: GPUs, ASICs, memory, HBM, SRAM, data centers, power, TSMC, AWS, OpenAI, agents, reasoning models, and why speed changes what AI products can become. Andrew explains why “tokens per second per user” matters, why generating a single word can require moving the equivalent of 100 HD movies through memory, why agents amplify latency, why GPUs struggle with certain inference workloads, and why fast AI may eventually reshape SaaS itself.This is a reference conversation on fast inference, AI chips, and the next compute bottleneck.(00:00) Cold open & Intro(01:31) Why speed became the AI bottleneck(02:32) Tokens per second per user, explained(03:16) AI's broadband moment and the Netflix analogy(04:35) The AI chip landscape: GPUs, TPUs, Trainium, ASICs(06:36) What is an ASIC?(08:08) Nvidia, Groq, and the fast inference war(09:16) OpenAI, Broadcom, and specialized silicon(12:10) China, power, and sovereign AI infrastructure(15:05) Is the AI infrastructure boom a bubble?(18:56) The hidden bottlenecks: HBM, CoWoS, and 3nm(22:57) Why agents are creating CPU demand(25:36) Andrew Feldman's path from SeaMicro to Cerebras(26:13) Why Cerebras bet on AI in 2016(31:14) SRAM vs. HBM: why inference is a memory problem(33:19) What wafer-scale computing actually means(34:28) The deep-tech “Everest” problem(36:07) The moment the first Cerebras system worked(36:49) Ringing the bell and surviving deep tech(39:08) How a giant chip handles failure(41:22) Why GPUs struggle with decode(42:17) Prefill vs. decode explained(44:01) The “100 HD movies” problem in AI inference(45:04) How fast inference changes RL and training(48:08) Reasoning models and why they cost more compute(50:08) Verification, guardrails, and small models checking big models(52:37) Multimodal AI and the path to video(53:51) Cerebras' business model: hardware, cloud, and API(55:14) OpenAI's 750MW inference deal(55:36) Why data centers are measured in megawatts(58:01) AWS Trainium + Cerebras decode(59:29) Fast tokens as a cloud product(01:00:52) Is CUDA still a moat?(01:03:53) How TSMC helped Cerebras build the giant chip(01:07:41) Why nobody cared in 2020(01:08:15) Why chip supply chains are hard to diversify(01:09:54) Why today's AI models will be the worst you ever use(01:10:38) What fast AI could do to SaaS

The 7investing Podcast
Rocket Lab & Netflix Deep Dive | Cerebras vs. NVIDIA: The $30 Billion AI Inference Battle

The 7investing Podcast

Play Episode Listen Later Jul 20, 2026 43:11


Is Cerebras Systems the next great AI chip stock or a red-hot IPO priced for perfection? In this episode of 7investing Live, Simon Erickson and executive producer Heather Horton welcome back Nick Rossolillo, co-founder of Chip Stock Investor, to break down three of the market's biggest stories.First up: Cerebras Systems (NASDAQ:CBRS), the wafer-scale chip maker that just IPO'd at a $40+ billion market cap. With 44GB of SRAM embedded directly on the chip, Cerebras was purpose-built to solve AI's "memory wall" problem for inference workloads. Now it's reportedly landed a ~$10 billion order from OpenAI and a deal with Amazon Web Services that could top $20 billion. Simon and Nick dig into whether these massive orders are real, how Cerebras stacks up against NVIDIA's GPUs and hyperscaler custom silicon, the TSMC capacity bottleneck that could throttle its growth, and how to value a company trading near 20x sales without profits.Then the conversation turns to Rocket Lab (NASDAQ:RKLB), which has pulled back from $150 to around $70 per share. Simon shares the latest iteration of his discounted cash flow valuation, and the duo debates the proposed Iridium acquisition — a deal that could pull Rocket Lab to EBITDA-positive on a pro forma basis — plus what the long-awaited Neutron rocket launch means for the company's future.Finally: Netflix (NASDAQ:NFLX). After another quarter of decelerating revenue guidance, is the streaming giant now a value stock rather than a growth stock? Nick explains why the advertising business hasn't reaccelerated growth the way he expected, and what he'd need to see before buying the dip.Plus: Nick's take on the recent chip stock sell-off across NVIDIA, AMD, Broadcom, SanDisk, and Kioxia and why "stocks go up, stocks go down" might be the healthiest way to think about it.Subscribe for more deep dives on AI infrastructure, semiconductors, and innovative growth stocks!Start your FREE 7-day trial of 7investing: https://www.7investing.com/subscribeFollow Nick and Casey Rossolillo at Chip Stock Investor: https://chipstockinvestor.comRocket Lab Deep Dive videos mentionedPart 1 https://youtu.be/AMDd0-JKUH0 (Deep Dive)Part 2: https://youtu.be/Z76xTGFNwBA (Valuation)Companies MentionedPublicly Traded:Cerebras Systems (NASDAQ:CBRS)Rocket Lab (NASDAQ:RKLB)Netflix (NASDAQ:NFLX)NVIDIA (NASDAQ:NVDA)Advanced Micro Devices (NASDAQ:AMD)Broadcom (NASDAQ:AVGO)Micron Technology (NASDAQ:MU)Taiwan Semiconductor Manufacturing (NYSE:TSM)Amazon (NASDAQ:AMZN)Alphabet (NASDAQ:GOOGL)Meta Platforms (NASDAQ:META)Iridium Communications (NASDAQ:IRDM)SanDisk (NASDAQ:SNDK)Kioxia Holdings (TSE:285A)Globalstar (NASDAQ:GSAT)SpaceX (NASDAQ: SPCX)Private / Pre-IPO:OpenAIAnthropicVideos Mentioned:https://www.youtube.com/watch?v=Z76xTGFNwBA&t=3shttps://www.youtube.com/watch?v=AMDd0-JKUH0&t=987sHere's the shifted chapter list, with all timestamps moved back 55 seconds:0:00 Welcome to 7investing Live0:54 Cerebras Systems: IPO recap & the Wafer-Scale Engine2:31 Is NVIDIA even the right comparison for Cerebras?5:38 The memory wall: why bigger AI models need new chips8:52 Latency vs. throughput — and the new AI alliances10:46 Are the $10B OpenAI & $20B Amazon orders real?14:02 Cerebras risks: how do you value a hot IPO?17:27 The TSMC capacity bottleneck20:01 Heather's take on Cerebras20:41 Rocket Lab: the sell-off & Iridium acquisition24:34 Simon's DCF valuation & price target for RKLB29:05 Why Neutron changes everything30:12 Q&A: Does Peter Beck carry an "Elon premium"?31:36 Netflix: buying opportunity or cheap for a reason?36:57 Q&A: Is Netflix a growth stock or a value stock?39:03 Chip stocks selling off: normal volatility or a warning?42:57 Wrap-up & final thoughts#7investing #Simonerickson #Cerebras #CBRS #NVIDIA #AIinvesting #semiconductors #chipstocks #RocketLab #RKLB #Netflix #NFLX #AIinference #stocks #investing #stockmarket #TSMC #AIdatacenters

Wall Street Unplugged - What's Really Moving These Markets
How DigiPower X's $1B Cerebras deal is powering execution

Wall Street Unplugged - What's Really Moving These Markets

Play Episode Listen Later Jul 17, 2026 29:32


DigiPower X (DGXX) CEO Michel Amar breaks down the company's milestone contract with Cerebras (CBRS)… its AI stack roadmap, from real estate to GPU-as-a-service… key catalysts through 2026… and why DigiPower X is in a league of its own. In this episode: Welcome back, Michel Amar, CEO of DigiPower X [0:01] DigiPower's $1 billion contract with Cerebras is a major milestone [0:53] Why DGXX is exempt from New York's moratorium on data centers [6:33] From real estate to GPU-as-a-service: DigiPower's full AI stack roadmap [10:51] Management's plans to keep executing through 2026 and beyond [19:49] DigiPower X is in a league of its own [24:43] Did you like this episode? Get more Wall Street Unplugged FREE each week in your inbox. Sign up here: https://curzio.me/syn_wsu Find Wall Street Unplugged podcast… --Curzio Research App: https://curzio.me/syn_app --iTunes: https://curzio.me/syn_wsu_i --Stitcher: https://curzio.me/syn_wsu_s --Website: https://curzio.me/syn_wsu_cat Follow Frank… X: https://curzio.me/syn_twt Facebook: https://curzio.me/syn_fb LinkedIn: https://curzio.me/syn_li

All-In with Chamath, Jason, Sacks & Friedberg
Open Source Wins, AGI Is Here, and Scorsese's AI Toolkit with CEOs of Cerebras & Black Forest Labs

All-In with Chamath, Jason, Sacks & Friedberg

Play Episode Listen Later Jul 10, 2026 63:57


(0:00) The AI Buildout: Datacenters Bigger Than Cities (Andrew Feldman) (1:50) Reasoning, Inference, and Breaking Moore's Law (16:28) Open Source, AI Sovereignty, and the Road to AGI (40:54) The Innovation Behind Generative Video (Robin Rombach) (47:31) Martin Scorsese, Robots, and the Future of Hollywood IP Thanks to our partners for making this possible! AppLovin Ads - AppLovin's AI advertising platform reaches over a billion daily active users across mobile games. Full-screen video ads with a 35-second median watch time. Advertisers are profitably spending hundreds of thousands of dollars a day and advertiser access is still in closed beta. The window is open at https://applovin.com/ALLIN Nasdaq - Positioned at the nexus of technology and the capital markets, Nasdaq provides premier platforms and services for global capital markets and beyond with unmatched technology, insights and markets expertise. https://www.nasdaq.com/convergence-economy Follow Andrew: https://x.com/andrewdfeldman Follow Robin: https://x.com/robrombach Follow the besties: https://x.com/chamath https://x.com/Jason https://x.com/DavidSacks https://x.com/friedberg Follow on X: https://x.com/theallinpod Follow on Instagram: https://www.instagram.com/theallinpod Follow on TikTok: https://www.tiktok.com/@allin Follow on LinkedIn: https://www.linkedin.com/company/allinpod Intro Music Credit: https://rb.gy/tppkzl https://x.com/yung_spielburg

The Circuit
EP 181: Cerebras Earnings, QCOM Investor Day, Micron Earnings and the Memory Mafia

The Circuit

Play Episode Listen Later Jun 29, 2026 62:52


 This episode of The Circuit covers a mix of tech industry tributes and major semiconductor financial updates. Hosts Ben and Jay begin by paying their respects to pioneering tech blogger and journalist Om Malik, who recently passed away. They then dive into Cerebras's first earnings report as a public company, highlighting strong top-line demand for AI inference but noting investor concern over their complicated gross margins. Next, the hosts unpack Qualcomm's Investor Day, focusing on the company's aggressive $15 billion data center revenue guidance for fiscal 2029, their "Dragonfly" custom ARM CPU roadmap, and their strategic acquisition of software startup Modular. Finally, they analyze Micron's massive earnings report, detailing a staggering 60% quarter-on-quarter memory price surge that has caught major buyers like Apple by surprise, concluding with a lively discussion on the "Game of Thrones" style market dynamics driving the memory industry. 

The 7investing Podcast
Is Moore's Law Dead? Cerebras IPO, SpaceX Orbital Data Centers & Huawei Tau Scaling Explained

The 7investing Podcast

Play Episode Listen Later Jun 29, 2026 42:09


Three massive semiconductor and computing developments are reshaping the future of AI infrastructure — and 7investing's Simon Erickson sits down with Nick Rossalillo of Chip Stock Investor to break them all down. First up: Cerebras Systems (NASDAQ:CBRS), which just went public on May 13th at $185/share (~$40 billion valuation) and is now trading near $46 billion at 90x trailing sales. The company's Wafer Scale Engine, a chip that uses an entire silicon wafer rather than individual diced chip, was designed specifically for AI inference workloads that NVIDIA (NASDAQ:NVDA) GPUs struggle to handle efficiently due to on-chip SRAM limitations. With potential $20 billion in orders from OpenAI and access via AWS, Cerebras is real, but neither Simon nor Nick is buying at this price. Their rule: wait a year before touching a fresh IPO.Next, SpaceX's freshly-raised $75 billion gets put under the microscope, specifically Elon's ambition to build orbital data centers. Nick walks through the SpaceX diagram: 70-meter solar panel wingspan, laser-based networking between compute modules, and the massive engineering challenges around power, heat dissipation, and in-orbit assembly. This isn't imminent, Starlink's next-gen constellation comes first — but if Elon can crack the economics, it would rewrite the rules of data center infrastructure entirely.Finally, Huawei's Tau Scaling announcement: a new architectural approach to chip performance that bypasses the need for extreme ultraviolet lithography (which China can't access due to ASML export controls). Tau temporal scaling focuses on minimizing signal travel time between transistors using logic folding, new materials, and 3D stacking. Huawei claims it could reach 1.5 nanometer equivalent performance by 2031. Simon and Nick are skeptical — 381 chips in six years is not mass production, and TSMC (NYSE:TSM) will be well past that node by then but it's worth watching as China continues building workarounds to Western export restrictions.Whether Moore's Law is dead or simply rerouting, the chipmaking industry is more innovative and more investable than it's been in decades.Join the conversation on the 7investing discord: https://discord.com/invite/PT9ZQqdXXSWant access to all 7investing research? Join at 7investing.com/subscribe Follow Chip Stock Investor  @chipstockinvestor  and https://chipstockinvestor.com/0:00 - Introduction to 7investing and Chip Stock Investor0:54 - Is Moore's Law Dead? A review of scaling semiconductor manufacturing3:08 - Cerebras Systems' recent IPO. How is their Wafer Scale Engine different than NVIDIA's GPUs, how does this impact Moore's Law, and is the stock a buy today?21:12 - SpaceX's recent IPO. Elon wants to build and launch orbital data centers. How does SpaceX plan to use the $75 billion it raised, what are the challenges it faces, and is the stock a buy?28:16 - Huawei's Tau scaling. Could this new chip architecture make ASML's extreme ultraviolet lithography obsolete, and what are its chances of succeeding?39:39 - Outro, final thoughts, and audience questionsStocks & Companies Mentioned:Cerebras Systems (NASDAQ:CBRS)NVIDIA (NASDAQ:NVDA)AMD (NASDAQ:AMD)SpaceX (SPCX)Taiwan Semiconductor Manufacturing Company / TSMC (NYSE:TSM)ASE Technology Holding / ASE Group (NYSE:ASX)Vicor Corporation (NASDAQ:VICR)ASML Holding (NASDAQ:ASML)Applied Materials (NASDAQ:AMAT)Lam Research (NASDAQ:LRCX)Intel (NASDAQ:INTC)Amazon / AWS (NASDAQ:AMZN)Alphabet / Google (NASDAQ:GOOGL)AST SpaceMobile (NASDAQ:ASTS)Samsung Electronics (KRX:005930)Huawei — private (Chinese company)OpenAI — privateLuckin Coffee (OTC:LKNCY) — mentioned as cautionary example#Semiconductors #MooresLaw #CerebrasSystems #CBRS #AIChips #NVIDIA #SpaceX #OrbitalDataCenters #HuaweiTech #TauScaling #ChipStocks #AIInvesting #TechStocks #GrowthStocks #StockMarket #InvestingIn2026 #7investing #Simonerickson

Revenue Builders
AI Is Redefining Seller Productivity with Alex Varel

Revenue Builders

Play Episode Listen Later Jun 28, 2026 11:07


Seller productivity is being redefined in real time, and the divide is no longer about effort or experience. In this segment, Alex Varel shares a grounded look at how AI is already reshaping the role, from a RevOps leader training multiple agents in a single weekend to the expectation that sellers will augment their own output. The conversation highlights a growing reality for revenue teams: those who get hands-on with AI will operate at a different level, while those who don't risk falling behind faster than they expect. Alex Varel is EVP of Worldwide Sales at Cerebras Systems, where he leads global go-to-market efforts at the forefront of AI infrastructure. He has built and scaled high-performing teams across MongoDB, Zscaler, and Multiverse, driving growth through IPO, hyper-scale expansion, and emerging technology shifts. Connect with Alex: LinkedIn Resources mentioned: "The Power of Myth" by Joseph Campbell "AI Superpowers" by Kai-Fu Lee “Leonardo da Vinci” by Walter Isaacson "No Country for Old Men" by Cormac McCarthy "The Road" by Cormac McCarthy “The Founders: The Story of Paypal and the Entrepreneurs Who Shaped Silicon Valley” by Jimmy Soni Listen to the full episode: How AI Is Rewriting the Sales Playbook and Raising the Bar on Human Performance with Alex Varel Hosted by five-time CRO John McMahon and Force Management Co-Founder John Kaplan, the Revenue Builders podcast goes behind the scenes with the sales leaders who have been there, done that, and seen the results. This show is brought to you by Force Management. We help companies improve sales performance, executing their growth strategy at the point of sale. Connect with Us: LinkedInYouTubeForce Management

WSJ What’s News
What's News in Markets: AI Tales, Oracle Woes, Wendy's Sizzles

WSJ What’s News

Play Episode Listen Later Jun 27, 2026 5:45


Why are Micron and Cerebras telling two different AI stories? And why is Oracle one of the worst stocks this week? Plus, who's behind Wendy's big rally? Host Jack Pitcher discusses the biggest stock moves of the week and the news that drove them. Sign up for the WSJ's free Markets A.M. newsletter. Learn more about your ad choices. Visit megaphone.fm/adchoices

WSJ Your Money Briefing
What's News in Markets: AI Tales, Oracle Woes, Wendy's Sizzles

WSJ Your Money Briefing

Play Episode Listen Later Jun 27, 2026 5:55


Why are Micron and Cerebras telling two different AI stories? And why is Oracle one of the worst stocks this week? Plus, who's behind Wendy's big rally? Host Jack Pitcher discusses the biggest stock moves of the week and the news that drove them. Sign up for the WSJ's free Markets A.M. newsletter. Learn more about your ad choices. Visit megaphone.fm/adchoices

All-In with Chamath, Jason, Sacks & Friedberg
Socialists Sweep NYC, China Catches Up in Coding, AI Memory Crunch, Micron's Blowout Quarter

All-In with Chamath, Jason, Sacks & Friedberg

Play Episode Listen Later Jun 26, 2026 101:43


(0:00) Gavin Baker and Travis Kalanick join the show! (1:05) Mamdani-endorsed socialists sweep congressional primaries in NYC (22:51) Future of the Democratic Party, the Israel issue, social media bans (45:12) China's open-source AI catch up, distillation, OpenAI's new chip (1:01:46) Micron smashes earnings, AI's memory crunch hitting Apple and consumer hardware (1:10:17) The math behind distributed compute and datacenters in space (1:27:22) IPO update: Anthropic at $3T, SpaceX float, Cerebras drops after breaking deal price Follow Gavin: https://x.com/GavinSBaker Follow Travis: https://x.com/travisk Apply for Summit 2026: https://allin.com/events Follow the besties: https://x.com/chamath https://x.com/Jason https://x.com/DavidSacks https://x.com/friedberg Follow on X: https://x.com/theallinpod Follow on Instagram: https://www.instagram.com/theallinpod Follow on TikTok: https://www.tiktok.com/@theallinpod Follow on LinkedIn: https://www.linkedin.com/company/allinpod Intro Music Credit: https://rb.gy/tppkzl https://x.com/yung_spielburg Intro Video Credit: https://x.com/TheZachEffect Referenced in the show: https://abcnews.com/Politics/clean-sweep-3-candidates-endorsed-mamdani-win-primaries/story?id=134152579 https://polymarket.com/event/mamdani-team-sweeps-primaries-20260618232357710 https://x.com/thestustustudio/status/2067356255916536120 https://www.nbcnews.com/politics/2026-election/espaillat-ny-house-primary-loss-district-13-avila-chevalier-rcna351127 https://x.com/EndWokeness/status/2069645066252034288 https://x.com/america/status/2069622732279402804 https://x.com/realmaalouf/status/2069433391162798337 https://x.com/JoshBlockDC/status/2070108811851882691 https://x.com/EndWokeness/status/2069776474429624684 https://x.com/EndWokeness/status/2068829255786803368 https://www.pewresearch.org/short-reads/2026/04/07/negative-views-of-israel-netanyahu-continue-to-rise-among-americans-especially-young-people https://x.com/PirateWires/status/2069146641266094417 https://www.wsj.com/economy/the-data-center-boom-is-sparking-a-third-wave-of-inflation-926adc6e https://x.com/jietang/status/2067580270078030088

Squawk on the Street
11AM Hour: Cerebras CEO Talks Results, GEV's Turbine Dreams, & Trump's Housing Pivot 6/24/26

Squawk on the Street

Play Episode Listen Later Jun 24, 2026 43:35


Semis in focus as Broadcom unveils a new custom chip with OpenAI, Micron awaits earnings after the bell, and Cerebras shares slump on the company's first report as a public company. Cerebras' CEO joined the team with his first take on the stock move. Plus: a live look at the world's largest gas turbine engine factory - and what it means for stocks like GE Vernova - along with reaction from the head of one of the nation's biggest homebuilders to the President's surprise pivot on a key housing affordability bill.    Squawk on the Street Disclaimer Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Squawk on the Street
9am Hour: Tech After the Sell-Off, OpenAI-Broadcom AI Chip Unveil, WTI Crude Falls Below $70 6/24/26

Squawk on the Street

Play Episode Listen Later Jun 24, 2026 43:39


Carl Quintanilla, David Faber and Leslie Picker explored what's ahead for the tech sector in wake of Tuesday's sell-off — and ahead of Micron's earnings due out after Wednesday's close of trading. The anchors also discussed OpenAl and Broadcom unveiling their new custom AI chip, called "Jalapeño." Brian Sullivan joined the anchors at Post 9 to discuss WTI crude falling below $70/barrel for the first time since the early stages of the Iran war. Seema Mody delivered a live report from inside GE Vernova's turbine factory — as the company looks to meet hyperscalers' demand for AI power. Also in focus: Cerebras tumbles on its first earnings report since going public, Alphabet to replace Verizon in the Dow, FedEx earnings reaction, what Treasury Secretary Bessent told CNBC about economic growth.   Squawk on the Street Disclaimer Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Squawk on the Street
Sen. Elizabeth Warren on Trump Canceling Housing Bill Signing, Broadcom CEO & OpenAI President on New Processor, Cerebras Board Member 6/24/26

Squawk on the Street

Play Episode Listen Later Jun 24, 2026 46:38


Breaking this hour: Sen. Elizabeth Warren reacts to President Trump's decision to cancel the planned signing of a bipartisan housing bill. Plus, why she says data centers are hurting local communities. Then, Broadcom CEO Hock Tan and OpenAI President Greg Brockman discuss a new AI processor “Jalapeño,” and just how much demand they are seeing for compute. And Cerebras board member and one of the earliest investors, Foundation Capital's Steve Vassallo, joins the show to break down the company's first results as a public company.   Squawk on the Street Disclaimer   Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Trader Merlin
S&P 500 7,800!? - 06/23/26

Trader Merlin

Play Episode Listen Later Jun 24, 2026 53:35


Wall Street just got a lot more bullish. A major market forecast has pushed its target for the S&P 500 all the way to 7,800, implying significant upside from current levels. But is this a realistic projection based on earnings, AI growth, and economic strength... or are analysts simply getting caught up in market euphoria? In today's episode, we break down the reasoning behind the upgraded target and ask the question every investor should be asking: Can the S&P 500 really reach 7,800, or is Wall Street getting ahead of itself? We'll discuss: What's driving the bullish forecasts The role of AI and technology in earnings growth Whether valuations still make sense Historical examples of analyst optimism and pessimism Key risks that could derail the rally We'll also take a look at a viewer's trade in Cerebras Systems, breaking down the setup, risks, and opportunities surrounding one of the more intriguing names in the AI space. In addition, we'll dive into two critical market indicators that many investors ignore: The 10-Year Treasury Yield The U.S. Dollar Index (DXY) Because these two markets often provide valuable clues about: Interest rates Inflation expectations Capital flows Future equity performance And of course, I'll provide updates on my current trades, portfolio positioning, and what I'm watching as markets continue to push higher. This episode is all about separating optimism from reality. Listen now:

CNBC's
Memory Losses Hit Tech… And Will Oil Prices Drop Even Further? 6/23/26

CNBC's "Fast Money"

Play Episode Listen Later Jun 23, 2026 43:25


Tech stocks plummeting even further today as investors seemingly dip out of the AI trade. RBC's Lori Calvasina breaks down what the losses mean for the future of tech and why she remains optimistic despite the struggle. Plus, major after-hours earnings reports from Cerebras and Fedex — what the results mean for the future of the AI chipmaker and transportation company. Then, what Iranian oil re-entering global markets could mean for domestic oil prices, and why it might be time to sell 2 powerhouse investment banks. Fast Money Disclaimer Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Squawk on the Street
11AM Hour: The Tech Tumble, A New Quantum EO, & Meta Wearables Chief Talks AI 6/23/26

Squawk on the Street

Play Episode Listen Later Jun 23, 2026 41:11


A volatile session as tech stocks take a broad leg lower - Carl Quintanilla and Sara Eisen discussed the key names to watch (from SpaceX to Amazon), and broke down what could come next with the CEO of one AI name that just partnered with Nvidia on drug discovery.  Plus: a deep-dive on Amazon's new advertising push with OpenAI, a look ahead to Cerebras earnings after the bell, and an exclusive with Meta's Head of Wearables as the company releases a slew of new AI glasses. Squawk on the Street Disclaimer Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

All-In with Chamath, Jason, Sacks & Friedberg
The IPO Comeback: Why Tech Giants Are Finally Going Public | All-In Liquidity IPO Panel

All-In with Chamath, Jason, Sacks & Friedberg

Play Episode Listen Later Jun 6, 2026 32:29


(0:00) CEOs Andrew Feldman (Cerebras) and Will Marshall (Planet Labs) join the Besties! (2:05) Both CEOs on going public: Impact on employees, customers, and business operations (13:18) Timelines for datacenters in space (19:28) Cerebras business breakdown, AI's impact on the silicon market (24:45) How Founder/CEOs think about liquidity on the road to going public Thanks to our partners for making this possible! EY - Great tech starts with a big idea. From startup to scale, EY helps tech founders get financials right early so they can focus on what's next. https://www.ey.com/en_us/tech-sector/tech-startups?WT.mc_id=3501317&AA.tsrc=sponsorship NYSE - Thank you to our partner, the New York Stock Exchange - a modern marketplace and exchange for building the future. It all happens at the NYSE. https://www.nyse.com Plaud - Never miss a moment. Plaud, our official wearable AI note-taking partner at All-In Liquidity Summit, captured every insight. https://www.plaud.ai Follow Brad Gerstner: https://x.com/altcap Follow Andrew Feldman: https://x.com/andrewdfeldman Follow Will Marshall: https://x.com/Will4Planet Apply for Summit 2026: ⁠https://allin.com/events⁠ Follow the besties: ⁠https://x.com/chamath⁠ ⁠https://x.com/Jason⁠ ⁠https://x.com/DavidSacks⁠ ⁠https://x.com/friedberg⁠ Follow on X: ⁠https://x.com/theallinpod⁠ Follow on Instagram: ⁠https://www.instagram.com/theallinpod⁠ Follow on TikTok: ⁠https://www.tiktok.com/@theallinpod⁠ Follow on LinkedIn: ⁠https://www.linkedin.com/company/allinpod⁠ Intro Music Credit: ⁠https://rb.gy/tppkzl⁠ ⁠https://x.com/yung_spielburg

The Best One Yet

There are more travel agents than ever… because AI can't do taste.The biggest IPO of the year so far is Cerebras… this IPO is about wafers and weddings.It's Fed Chair Jerry Powell's last day of work… we give him two grades for his 8-year tenure.Plus, Kool-Aid now thinks it's a wellness brand… so do gummy bears and Kraft Mac & Cheese.$CBRS $KHC $BKNGNEWSLETTER:https://tboypod.com/newsletter OUR 2ND SHOW:Want more business storytelling from us? Check our weekly deepdive show, The Best Idea Yet: The untold origin story of the products you're obsessed with. Listen for free to The Best Idea Yet: https://wondery.com/links/the-best-idea-yet/NEW LISTENERSFill out our 2 minute survey: https://qualtricsxm88y5r986q.qualtrics.com/jfe/form/SV_dp1FDYiJgt6lHy6GET ON THE POD: Submit a shoutout or fact: https://tboypod.com/shoutouts SOCIALS:Instagram: https://www.instagram.com/tboypod TikTok: https://www.tiktok.com/@tboypodYouTube: https://www.youtube.com/@tboypod Linkedin (Nick): https://www.linkedin.com/in/nicolas-martell/Linkedin (Jack): https://www.linkedin.com/in/jack-crivici-kramer/Anything else: https://tboypod.com/ About Us: The daily pop-biz news show making today's top stories your business. Formerly known as Robinhood Snacks, The Best One Yet is hosted by Jack Crivici-Kramer & Nick Martell. Hosted on Acast. See acast.com/privacy for more information.