American semiconductor company
POPULARITY
Categories
Andrew, Ben, and Tom break down Nvidia's $35 billion Lambda/Anthropic cloud deal and new MediaTek partnership that squeezes Broadcom and Marvell, plus Bessent's latest comments on bonds and inflation.Join our live YouTube stream Monday through Friday at 8:30 AM EST:http://www.youtube.com/@TheMorningMarketBriefingPlease see disclosures:https://www.narwhal.com/disclosure
Guy Adami and Dan Nathan open by recounting a fan's Guy-themed T-shirt sighting on CNBC's Fast Money, then discuss Fed Chair Kevin Warsh's Jackson Hole remarks as largely status quo, with the S&P 500 near all-time highs and the VIX around 14 despite potential catalysts like the August jobs report. They argue markets appear complacent ahead of midterms and cite historical midterm drawdowns, while noting election-related tensions, Canada trade friction, and Russia/NATO risks as possible volatility drivers. The hosts highlight a widening disconnect between strong equities and weakening consumer signals seen in recent retail earnings, alongside rising delinquency rates and persistent inflation pressures. They review key earnings and themes: Nvidia's extraordinary growth but complex circular AI financing relationships and competitive chip efforts, Salesforce's sharp rally and software rebound, and previews of Dell and Broadcom amid hardware and TPU demand dynamics, before plugging an interview with Imran Khan. —FOLLOW USYouTube: @RiskReversalMediaInstagram: @riskreversalmediaTwitter: @RiskReversalLinkedIn: RiskReversal MediaThe financial opinions expressed in Risk Reversal content are for information purposes only. The opinions expressed by the hosts and participants are not an attempt to influence specific trading behavior, investments, or strategies. Past performance does not necessarily predict future outcomes. No specific results or profits are assured when relying on Risk Reversal. Before making any investment or trade, evaluate its suitability for your circumstances and consider consulting your own financial or investment advisor. The financial products discussed in Risk Reversal carry a high level of risk and may not be appropriate for many investors. If you have uncertainties, it's advisable to seek professional advice. Remember that trading involves a risk to your capital, so only invest money that you can afford to lose. Derivatives are not suitable for all investors and involve the risk of losing more than the amount originally deposited and any profit you might have made. This communication is not a recommendation or offer to buy, sell or retain any specific investment or service.
In this week's Stock Market Options Trading Podcast, I break down an XSP 0DTE trade I entered at the market open and how I was able to turn the position into a risk-free trade after the market moved in my favor.We'll walk through the trade structure, why I started with an in-the-money call credit spread, and how adding the opposing spread converted the position into an Iron Condor with no remaining downside risk.I also take a first look at CBOE's newer Magnificent 10 Index (MGTN), an equal-weighted index featuring the Magnificent 7 plus Broadcom, Palantir, and AMD. We'll look at why cash-settled index options are interesting for traders, along with the biggest issue with MGTN right now: liquidity.Plus, this week's episode covers:• The current SPX market setup and important gamma levels• What I'm watching as the market trades below 7700• This week's major economic and employment reports• Tuesday's SPX Opening Range Breakout statistics• The return of the X Files segmentThe goal is always the same: use data, backtesting, and market statistics to find practical ways to trade the S&P 500 and index options.
Charles is joined by Mario Veneroso, CEO and Managing Partner of Veneroso Wealth Management, to discuss how to navigate potential Federal Reserve rate hikes, why incredible S&P 500 earnings growth keeps the market rally alive, and whether tech and healthcare plays like Broadcom, Hims & Hers, and palantir are smart buys for your portfolio. Learn more about your ad choices. Visit podcastchoices.com/adchoices
Plus: Best Buy raises outlook following better-than-expected second-quarter results. And Cisco gave its 90,000 employees their own AI agent. Julie Chang hosts. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
What happens when an organization writes careful AI governance policies but its infrastructure cannot enforce any of them? In this episode of Tech Talks Daily, I speak with Sabina Anja, Chief Technologist at Broadcom within the VMware Cloud Foundation division, about the infrastructure controls required as AI agents move from generating answers to accessing data, calling APIs, modifying systems, and triggering work. Sabina brings experience from both sides of enterprise technology. She remembers cabling networks, dealing with unstable infrastructure, and receiving those weekend calls when downtime had already upset the business. That background informs her belief that ambitious AI programs cannot succeed without stable, observable, and enforceable infrastructure beneath them. Many organizations are repeating a familiar pattern. Business teams adopt AI services before IT has established visibility, ownership, or control. The terminology may have changed from shadow IT to shadow AI, but the management problem remains. Sabina argues that CIOs first need an inventory of agents, nonhuman identities, data access, processes, and accountable owners. The risk becomes greater because agents behave differently from people. They operate across multiple systems at machine speed and can perform repeated actions without appreciating the wider business outcome. An agent does not need malicious intent to cause disruption. Excessive permissions, flat networks, inconsistent access rules, and years of deferred infrastructure work can give it plenty of opportunities. Sabina recommends brokered access rather than direct access, alongside dedicated virtual machines or namespaces, microsegmentation, lateral security, east-west policy controls, and tamper-evident logging. Organizations also need to define which data an agent can view, modify, or move, especially when sovereignty and regulatory requirements apply. One of Sabina's most memorable ideas is to treat an AI agent like a superhuman contractor. It should have a defined purpose, a named manager, a clear access specification, an activity record, and an end date. Additional permissions should be earned through evidence of reliable behavior rather than granted on the first day. She also warns about agent debt. AI systems are developing rapidly, so an agent created today may become outdated within months. Sabina recommends assuming that many agents will expire after six to nine months rather than allowing forgotten systems and permissions to accumulate indefinitely. For CIOs wanting an immediate test, her advice is straightforward. Create an inventory of nonhuman identities with production access. Then select one agent and examine every part of the infrastructure it attempted to reach. The question is not simply whether the application produced the expected result. Leaders should ask whether the agent entered systems, networks, or data stores that nobody expected it to access. We also challenge the familiar claim that AI agents will take everybody's jobs. Sabina sees an opportunity to remove repetitive tasks and give technology professionals new skills, although she warns that agents may behave like teenagers armed with infrastructure permissions. They may not take your job, but they could become remarkably good at testing your patience. I'd love to hear your thoughts. Does your organization know how many AI agents have production access and who is accountable for each one?
News Sources: https://lmg.gg/R3YOs Timestamps: 0:00 Meta settles teen addiction lawsuit 1:23 Apple M6 Mac mini and M5 Mac Studio 2:37 Twitter kills Nitter 4:36 QUICK BITS INTRO 4:45 Norway weighs smart glasses rules 5:13 Ring's new default encryption 5:43 Xbox 25th anniversary accessories 6:08 OpenAI's Broadcom chip benchmarks 6:45 Apple polishing cloth price cut 7:12 Credits Learn more about your ad choices. Visit megaphone.fm/adchoices
If there were any doubts about who's wearing the crown in the AI revolution... Nvidia just delivered another monster quarter. In today's episode, we're breaking down the latest earnings from Nvidia—and these aren't numbers that matter only to NVDA shareholders. Nvidia reported $96.2 BILLION in quarterly revenue, up an incredible 106% from a year ago. Even more impressive, its Data Center business generated $89 billion, up 117% year over year. Think about that for a moment. Nvidia isn't just growing. A company of this size just more than DOUBLED its revenue in one year. So the big question for today's show isn't simply whether Nvidia had a good quarter. It's: Can Nvidia—and the AI boom—keep this going? We'll dive into the numbers and look at what Nvidia's results tell us about the entire artificial-intelligence ecosystem. We'll discuss: Nvidia's latest earnings – What jumped out from the report and where the growth is coming from. Data Center dominance – What $89 billion in quarterly Data Center revenue tells us about global AI infrastructure spending. The AI spending boom – Are Microsoft, Meta, Amazon, Alphabet and other hyperscalers still willing to spend enormous amounts of money building AI infrastructure? Semiconductors – What Nvidia's results could mean for AMD, Broadcom, Micron and the rest of the chip sector. Memory – More AI computing means enormous demand for high-performance memory. Does Nvidia's growth strengthen the case for DRAM and HBM? Energy & infrastructure – All those GPUs have to go somewhere—and they require data centers, electricity, cooling, networking and an enormous infrastructure buildout. Valuation – At some point, even incredible growth can become fully priced in. Has Nvidia reached that point? The broader market – Nvidia has become so large and influential that its results can impact the Nasdaq, S&P 500 and overall investor sentiment. That's what makes this earnings report so important. Nvidia is no longer simply a semiconductor company investors watch four times a year. It's become one of the market's primary gauges of the entire AI investment cycle. Going into today's report, options markets were pricing roughly a 5.4% move in Nvidia shares, representing approximately $280 BILLION in potential market-cap movement in either direction. That's larger than the entire market capitalization of most companies! And with concerns growing recently about massive AI spending, stretched technology valuations and whether companies are generating enough return on their AI investments, Nvidia's results provide an important reality check. If AI is a bubble, somebody forgot to tell Nvidia's customers. But that doesn't mean the risks have disappeared. We'll separate the incredible fundamentals from the stock's valuation and ask the question traders actually care about: Great company... but is it still a great trade? For additional research, check out Nvidia Investor Relations and Nvidia Financial Reports. Listen now:
2,5% Zinsen p.a. auf ein unbegrenztes Guthaben mit bis zu fünfmal der gesetzlichen Einlagensicherung*. Auch für Kinder. Das gibt's bei Scalable Capital. Mehr Infos hier. Trump macht 1.000 Trades im Juni. Kanada verhängt Gegenzölle. US-Zinsen steigen NVIDIA kauft KI-Firma . Samsung schüttet 80 Mrd. $ aus. Fielmann senkt Prognose. Bald Tesla Cybercab ohne Lenkrad? Fleisch-Aktien haben's schwer. Broadcom will Geld. Minebea Mitsumi (WKN: A3D9VL) hat 60% Weltmarktanteil bei Miniaturkugellagern. Die stecken in Roboterhänden, Servern und Autos. KGV 15, solides Wachstum. Robotik-Fantasie gibt's gratis obendrauf. Langfristige Wette, aber greifbar. NBA-Teams explodieren im Wert. Die Timberwolves wechseln für 4,5 Mrd. $ den Besitzer. Madison Square Garden Sports (WKN: A140F0) hat sich seit September verdoppelt. Und Lakers-Verkäufer Mark Walter steckt in echten Problemen. Diesen Podcast vom 24.08.2026, 3:00 Uhr stellt dir die Podstars GmbH (Noah Leidinger) zur Verfügung. *Veränderlicher Zins auf unbegrenztes Guthaben. Konditionen sowie Guthabenverteilung auf scalable.capital/tagesgeld. Learn more about your ad choices. Visit megaphone.fm/adchoices
In this episode of The Circuit, hosts Ben Bajarin and Jay Goldberg break down the shifting dynamics across the semiconductor landscape, from analog chips to custom silicon and corporate balance sheets.Analog Devices (ADI) & The Data Center Wave: ADI reports a strong quarter, driven by a doubling of its data center revenue and expanding opportunities in 800-volt data center architectures. Ben and Jay discuss ADI's conservative focus on SAM (Serviceable Addressable Market) over TAM, high gross margins, and how analog content per gigawatt is becoming a key driver for power, safety, and sub-volt conversion.AI Sentiment vs. Market Reality: The hosts address the hype around AI "curing cancer," clarifying Moderna's recent bio-informatics announcement, and discuss the growing local friction surrounding massive data center construction projects.Custom ASICs & Marvell's Google Win: Analysis of Marvell's expanding relationship with Google for periphery XPU attach, optical interconnects, and CXL. They unpack why this is a massive win for Marvell's diversification and clear up misconceptions regarding Broadcom's market position.The Era of "Balance Sheet Competition": Semiconductor giants are using equity, warrants, and Special Purpose Vehicles (SPVs) to fund data center capacity. Ben and Jay debate the risks of chip companies taking on debt-like obligations to compete with NVIDIA.
One day after the worst session for the Dow and S&P 500 since late July, Carl Quintanilla, David Faber and Sara Eisen discussed stocks rebounding as Treasury Secretary Bessent's push to lower bond yields remains in focus. RBC Capital's Helima Croft joined the program to talk about higher oil prices and geopolitics, as President Trump seeks to ramp up economic pressure on Iran. Also in focus: What Bessent told CNBC about yield moves and Trump's approach to Iran; A slew of Wall Street analysts cut their price targets on Walmart after the stock's worst day in four years; Sources tell David that a Broadcom-backed special purpose vehicle is tapping the debt market for $70 billion to support an AI buildout; Crypto's great week; SpaceX's losing streak sends the stock back below its IPO price; What the bond rout means for private credit. Squawk on the Street Disclaimer Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
In der heutigen Folge sprechen die Finanzjournalisten Nando Sommerfeldt und Holger Zschäpitz über Broadcoms Schuldenplan, Georgs geniale Gas-Idee und die historischen Auswüchse am Anleihemarkt. Außerdem geht es um Walmart, Advance Auto Parts, AutoZone, O'Reilly Automotive, Home Depot, Lowe's, TJX Companies, Coty, JD Sports Fashion, Adidas, Puma, Fresenius, Fresenius Medical Care, Sartorius, UBS, Tonies, Korea Gas Corporation, Vontobel, Alibaba, Deere, Broadcom, Apollo Global Management, Blackstone, Goldman Sachs, Bank of America, Nvidia, Apple, Microsoft, Amazon, Alphabet, TSMC, Meta Platforms, Samsung Electronics, ASML, SK Hynix, Deutsche Börse, Moderna, BioNTech, Eli Lilly, Novo Nordisk, Mini Future Long auf den TTF-Gaspreis von Vontobel, (WKN: VY80G4), Vanguard FTSE Global All-Cap ETF thesaurierend (WKN: A42B1M), Vanguard FTSE Global All-Cap ETF ausschüttend (WKN: A42B1N), Vanguard FTSE Global Small-Cap ETF thesaurierend (WKN: A42B1P), Vanguard FTSE Global Small-Cap ETF ausschüttend (WKN: A42B1Q), Vanguard FTSE All-World ex-U.S. UCITS ETF thesaurierend (WKN: A42B1R), Vanguard FTSE All-World ex-U.S. ETF ausschüttend (WKN: A42B1S), SPDR MSCI ACWI IMI ETF (WKN: A1JJTD), Vanguard ESG Global All Cap ETF thesaurierend (WKN: A2QL8U), Vanguard FTSE All-World ETF thesaurierend (WKN: A2PKXG), Vanguard FTSE All-World ETF ausschüttend (WKN: A1JX52). Am 2. Oktober findet unser „Alles auf Aktien“-Summit in Berlin statt. Mit dem Code „AAAFRIENDS“ sparst du 50 Prozent auf dein Ticket – aber nur unter folgendem Link: https://veranstaltung.businessinsider.de/event/financesummit26/summary?rp=c6dc55d6-6f4f-4fb4-b75f-3f3501d84859 Wir freuen uns an Feedback über aaa@welt.de. Holt euch jetzt den exklusiven NordVPN-Deal inkl. 4 Bonusmonaten mit dem Code Allesaufaktien oder unter https://nordvpn.com/allesaufaktien Noch mehr "Alles auf Aktien" findet Ihr bei WELTplus und Apple Podcasts – inklusive aller Artikel der Hosts. Hier bei WELT: https://www.welt.de/podcasts/alles-auf-aktien/plus247399208/Boersen-Podcast-AAA-Bonus-Folgen-Jede-Woche-noch-mehr-Antworten-auf-Eure-Boersen-Fragen.html. Hier könnt ihr den AAA-Newsletter abonnieren: https://www.welt.de/newsletter/article232797673/Alles-auf-Aktien-Der-taegliche-Boersen-Newsletter-fuer-WELTplus-Abonnenten.html Und – ganz neu: AAA gibt es jetzt auch auf Instagram: https://www.instagram.com/alles_auf_aktien/ Disclaimer: Die im Podcast besprochenen Aktien und Fonds stellen keine spezifischen Kauf- oder Anlage-Empfehlungen dar. Die Moderatoren und der Verlag haften nicht für etwaige Verluste, die aufgrund der Umsetzung der Gedanken oder Ideen entstehen. Hörtipps: Für alle, die noch mehr wissen wollen: Holger Zschäpitz können Sie jede Woche im Finanz- und Wirtschaftspodcast "Deffner&Zschäpitz" hören. +++ Werbung +++ Du möchtest mehr über unsere Werbepartner erfahren? Hier findest du alle Infos & Rabatte! https://linktr.ee/alles_auf_aktien Anzeige: Eight Sleep: Der Pod 5 reguliert die Temperatur im Bett automatisch, trackt Schlaf- und Gesundheitswerte ohne Wearable und kann so zu besserem Schlaf beitragen. Mit dem Code ALLESAUFAKTIEN erhaltet ihr auf https://www.eightsleep.com/allesaufaktien bis zu 350 Euro Rabatt. Impressum: https://www.welt.de/services/article7893735/Impressum.html Datenschutz: https://www.welt.de/services/article157550705/Datenschutzerklaerung-WELT-DIGITAL.html
Stripe schreibt seinen Investoren, die Singularität habe am 1. Januar begonnen, und bestätigt im selben Brief den Kauf von OpenRouter. Danach geht es eine Stufe tiefer in die Nahrungskette der Absatzfinanzierung: Broadcom nimmt bis zu hundert Milliarden Dollar Schulden auf, um Anthropic den Kauf von Chips zu ermöglichen, die Broadcom selbst entworfen hat. In Ohio warnt das Wahlkampfkomitee der Republikaner die KI-Konzerne, dass Rechenzentren die Partei einen Senatssitz kosten. Ab Montag läuft ChatGPT-Werbung auch in Deutschland, und es geht um die Frage, ob eine Suchmaschine mit Werbung überhaupt glaubwürdig bleiben kann. Bei den Quartalszahlen liegt Anthropic vor OpenAI, in den ersten Wochen des laufenden Quartals dreht sich das Momentum aber wieder. Anthropic könnte in den nächsten Tagen den Börsenprospekt vorlegen. Dazu Drohnen von Amazon in bis zu fünfhundert Städten, ein Cybercab ohne Lenkrad, eine Kamera im AirPod und eigene Chips von Waymo. Bei den Zahlen bremst Klarna wegen Deutschland, und die Doping-Olympiade macht achtzehn Millionen Umsatz bei sechsundneunzig Millionen Kosten. Am Ende eine längere Rechnung darüber, warum Zara mit Innenstadtläden mehr verdient als Shein online. Unterstütze unseren Podcast und entdecke die Angebote unserer Werbepartner auf doppelgaenger.io/werbung. Vielen Dank! Philipp Glöckler und Philipp Klöckner sprechen heute über: (00:00:00) Stripe und die Singularität (00:08:21) VC-Konzentration (00:11:16) Broadcom bürgt für Anthropic (00:13:24) Rechenzentren im Wahlkampf (00:20:19) Die Kelce-Kampagne (00:23:18) ChatGPT-Werbung in Europa (00:38:05) OpenAI gegen Anthropic (00:51:45) Amazon-Drohnen (00:57:05) Tesla Cybercab (00:58:51) AirPods mit Kamera (01:03:35) Waymo baut Chips (01:06:17) Der günstige Waymo (01:08:42) SpaceX wollte Cognition (01:12:45) Klarna (01:15:52) Enhanced Games (01:19:28) Ramp Router (01:22:08) Personio (01:23:09) LinkedIn-Slop (01:25:04) Unitree (01:26:51) Shein gegen Inditex (01:37:52) Apotheken und E-Commerce (01:42:50) Neom (01:43:32) LandSpace landet (01:45:09) Spirit-Flugbegleiter (01:46:12) Metas Suchtprozess Shownotes Stripe schreibt Investoren, die Singularität habe begonnen - axios.com Venture Capital war noch nie so konzentriert - profgmedia.com Broadcom sucht mehr als 60 Mrd. für den nächsten KI-Schuldendeal - bloomberg.com Das GOP-Memo zum Rechenzentrums-Risiko in Ohio - axios.com Die Kelce-Brüder sammeln Kühlmittel für Rechenzentren - thoughtcatalog.com ChatGPT Ads kommt in 31 europäische Märkte - openai.com OpenAI baut Vertriebsteams für Werbung in Europa auf - digiday.com OpenAIs Q2-Umsatz wächst langsamer als der von Anthropic - wsj.com Friar: OpenAI wird 2027 eine Public Company - cnbc.com Anthropic will die Größe des SpaceX-IPO erreichen oder übertreffen - bloomberg.com Amazon weitet die Drohnenlieferung aus - bloomberg.com Teslas Cybercab startet mit Mitarbeitern als Fahrgästen - thenextweb.com Hinweise auf AirPods mit Kamera in macOS 26.7 - macrumors.com Waymo hat einen eigenen Chip für seine Robotaxis gebaut - bloomberg.com Das günstigere Waymo-Robotaxi ist in drei Städten für alle offen - techcrunch.com SpaceX wollte Cognition kaufen - bloomberg.com Klarna senkt die Prognose und sucht einen neuen CFO - wsj.com Enhanced Games mit 62 Mio. Verlust im zweiten Quartal - themirror.com Ramp startet einen eigenen KI-Modell-Router - techcrunch.com Personio zeigt KI-Ausgaben neben den Vergütungsdaten - linkedin.com Unitree legt im Shanghaier Debüt um mehrere Hundert Prozent zu - cnbc.com Shein peilt den Handelsstart am 28. August an - reuters.com Neom: The Line wird zum KI-Rechenzentrum - golem.de China landet erstmals eine Trägerrakete auf dem Boden - qz.com Flugbegleiter wehren sich gegen Googles Datenkauf - wsj.com Der erste Zeuge im Meta-Prozess in Oakland - thenextweb.com
Die Wall Street steuert nach dem Rückschlag vom Donnerstag auf eine Erholung zu. Die Futures auf den S&P 500 und Nasdaq liegen im Plus, während Öl und Anleiherenditen etwas nachgeben und damit kurzfristig für Entlastung sorgen. Im Fokus steht Broadcom: Der Konzern verhandelt über eine Fremdfinanzierung für KI-Chips und Rechenzentren, die insgesamt bis zu 100 Milliarden US-Dollar erreichen könnte. Das zeigt, wie stark der KI-Boom inzwischen vom Kreditmarkt abhängt. Positiv überraschen Ross Stores mit einem deutlichen Gewinn-Beat, 10 Prozent höheren vergleichbaren Umsätzen und einem angehobenen Ausblick sowie BJ's Wholesale mit besser als erwarteten Zahlen. Entscheidend bleibt aber die Frage, ob die hohen langfristigen Renditen den Technologiesektor weiter bremsen. Abonniere den Podcast, um keine Folge zu verpassen! ____ Folge uns, um auf dem Laufenden zu bleiben: • X: http://fal.cn/SQtwitter • LinkedIn: http://fal.cn/SQlinkedin • Instagram: http://fal.cn/SQInstagram
In der heutigen Folge sprechen die Finanzjournalisten Nando Sommerfeldt und Holger Zschäpitz über die Unitree-Euphorie, gute Gründe für Gold und die große Bitcoin-Shorts-Liquidierung. Außerdem geht es um Merck & Co., Qiagen, Sartorius, Infineon, SUSS MicroTec, CXMT, Marvell Technology, Alphabet, Broadcom, KraneShares China Technology & Semiconductor STAR 50 Index ETF (WKN: A3CU6C), Walmart, Alibaba, Baidu, Deere, Goldman Sachs, Royal Bank of Canada, Jefferies Financial Group, Evercore, Barclays, Morgan Stanley, DocMorris, Amazon, Redcare Pharmacy, iShares Global Infrastructure UCITS ETF USD (thesaurierend) (WKN: A3D6N1), iShares Global Infrastructure UCITS ETF USD (ausschüttend) (WKN: A0LEW9). Am 2. Oktober findet unser „Alles auf Aktien“-Summit in Berlin statt. Mit dem Code „AAAFRIENDS“ sparst du 50 Prozent auf dein Ticket – aber nur unter folgendem Link: https://veranstaltung.businessinsider.de/event/financesummit26/summary?rp=c6dc55d6-6f4f-4fb4-b75f-3… Wir freuen uns an Feedback über aaa@welt.de. Noch mehr "Alles auf Aktien" findet Ihr bei WELTplus und Apple Podcasts – inklusive aller Artikel der Hosts. Hier bei WELT: https://www.welt.de/podcasts/alles-auf-aktien/plus247399208/Boersen-Podcast-AAA-Bonus-Folgen-Jede-Woche-noch-mehr-Antworten-auf-Eure-Boersen-Fragen.html. Hier könnt ihr den AAA-Newsletter abonnieren: https://www.welt.de/newsletter/article232797673/Alles-auf-Aktien-Der-taegliche-Boersen-Newsletter-fuer-WELTplus-Abonnenten.html Und – ganz neu: AAA gibt es jetzt auch auf Instagram: https://www.instagram.com/alles_auf_aktien/ Disclaimer: Die im Podcast besprochenen Aktien und Fonds stellen keine spezifischen Kauf- oder Anlage-Empfehlungen dar. Die Moderatoren und der Verlag haften nicht für etwaige Verluste, die aufgrund der Umsetzung der Gedanken oder Ideen entstehen. Hörtipps: Für alle, die noch mehr wissen wollen: Holger Zschäpitz können Sie jede Woche im Finanz- und Wirtschaftspodcast "Deffner&Zschäpitz" hören. +++ Werbung +++ Du möchtest mehr über unsere Werbepartner erfahren? Hier findest du alle Infos & Rabatte! https://linktr.ee/alles_auf_aktien Anzeige: Eight Sleep: Der Pod 5 reguliert die Temperatur im Bett automatisch, trackt Schlaf- und Gesundheitswerte ohne Wearable und kann so zu besserem Schlaf beitragen. Mit dem Code ALLESAUFAKTIEN erhaltet ihr auf https://www.eightsleep.com/allesaufaktien bis zu 350 Euro Rabatt. Impressum: https://www.welt.de/services/article7893735/Impressum.html Datenschutz: https://www.welt.de/services/article157550705/Datenschutzerklaerung-WELT-DIGITAL.html
In this episode, we welcome Ernie China, who recently joined Everpure after a decade at VMware and Broadcom. Ernie brings deep expertise in the virtualization ecosystem and offers a unique perspective on the industry landscape during this pivotal time. Our conversation serves as a warm introduction to Ernie and sets the stage for a broader discussion on the challenges and opportunities facing IT teams today. The core of our discussion explores the current state of virtualization, particularly how enterprises are reacting to the recent changes following the Broadcom acquisition. Ernie explains why customers are feeling significant turmoil as they navigate shifts from perpetual licenses to subscription models. They discuss the common struggle many organizations face, describing a transitional phase where leaders are weighing whether to stay the course with their existing VMware investment or explore alternative hypervisors and container based strategies. Finally, we highlight Everpure's commitment to supporting customers throughout this transition. Everpure remains fully invested in the VMware ecosystem while providing a robust foundation for those looking to modernize their infrastructure through containers or other hypervisors. We make a compelling case for why IT leaders must elevate data management as a strategic priority, ensuring that decisions about performance, security, and efficiency are made long before infrastructure choices are finalized. To learn more, visit: https://www.everpuredata.com/solutions/virtualization.html Check out the new Everpure digital customer community to join the conversation with peers and Everpure experts: https://purecommunity.purestorage.com/ 00:00 Coming Up 00:29 Intro to our Guest 04:32 Inside the Broadcom Transition 09:11 Stats of the Episode on Virtualization Trends 15:10 Alternative Hypervisors and Options 19:50 Dual Strategies 24:50 Data Management is Imperative 27:44 Hot Takes Segment
Crypto reporter Yueqi Yang talks with TITV Host Akash Pasricha about Kalshi topping $4B in annualized revenue and seeking a $40B valuation. We also talk with Eliyan co-founders Ramin Farjadrad and Patrick Sohelli about AI chiplet interconnects and competing with Broadcom, Leo Schwartz about David Sacks returning to Craft Ventures to target a new $1B fund, and we get into the USDA trimming its Salesforce footprint for C3 AI with Laura Bratton.Articles discussed on this episode: https://www.theinformation.com/articles/kalshi-tops-4-billion-annualized-revenue-seeks-40-billion-valuationhttps://www.theinformation.com/articles/sacks-craft-targets-1-billion-first-fund-since-white-house-stintSubscribe: YouTube: https://www.youtube.com/@theinformation The Information: https://www.theinformation.com/subscribe_hSign up for the AI Agenda newsletter: https://www.theinformation.com/features/ai-agendaTITV airs weekdays on YouTube, X and LinkedIn at 10AM PT / 1PM ET. Or check us out wherever you get your podcasts.Follow us:X: https://x.com/theinformationIG: https://www.instagram.com/theinformation/TikTok: https://www.tiktok.com/@titv.theinformationLinkedIn: https://www.linkedin.com/company/theinformation/Chapters:00:00 - Introduction01:13 - Kalshi Tops $4B In Annualized Revenue, Seeks $40B Valuation08:59 - Chip Infrastructure Startup Eliyan Hits Unicorn Status19:38 - Sacks' Craft Targets $1B For First Fund Since White House25:35 - USDA Taps C3 AI to Trim Salesforce Footprint
Patrick Moorhead and Daniel Newman cover Tesla and SpaceX's $16.8 billion TerraFab chip factory, Samsung's zHBM debut at FMS 2026, a week of frontier model launches, and a simulated debate for "The Flip" over whether custom silicon will capture the majority of AI workload dollars by 2028. The episode closes with earnings from SpaceX, Palantir, AMD, Astera Labs, IonQ, and Lattice Semiconductor. The handpicked topics for this week are: 1. TerraFab Breaks Ground in Grimes County, Texas: Moorhead and Newman dig into Tesla and SpaceX's approved $16.8 billion TerraFab chip factory near Texas A&M, where Intel is confirmed as a digital foundry partner despite minimal SEC disclosure so far. Moorhead raises the open question of who handles the analog side of production that Intel doesn't cover, and Newman points to Musk's track record of hitting ambitious goals years behind his original stated timelines. Both hosts frame the investment as a genuine step toward expanding domestic chip manufacturing. (The Decode) 2. Samsung's zHBM Headlines a Blockbuster Memory Show at FMS 2026: The hosts cover the Future of Memory and Storage show in Santa Clara, where Samsung's BVNAND and zHBM concepts headlined alongside new announcements from Kioxia, SanDisk, and SK Group's high-bandwidth flash spec through OCP. Pat traces the industry's swing from oversupply and negative margins to today's memory shortage, driven by AI demand that has grown 5-10x faster than expected. Daniel adds that stacking memory directly on the accelerator, as zHBM proposes still needs to solve for the physical constraints of heat before it becomes production-ready. (The Decode) 3. Three Frontier Model Launches Land in a Single Week: Alibaba shipped Qwen 3.8 Max, DeepSeek released V4 Flash at 14 cents per million tokens, and Meta debuted its first coding agent, prompting a discussion on how fast the frontier is moving. Newman argues every company will need a new benchmark built around operational workloads and real-world token efficiency. He points out that Chinese model usage and frontier lab revenue are both climbing simultaneously, and makes the case that healthy frontier labs are essential to funding the open source models built in their wake. (The Decode) 4. Anthropic and AMD Both Make Moves on Custom Silicon: Anthropic confirmed it's developing custom AI chips, reportedly with Samsung and reportedly Broadcom. AMD announced its acquisition of Taalas, a company that etches model weights directly into silicon for major inference speed gains. Patrick traces his own multi-year prediction that heterogeneous computing would become standard, pointing to Google's TPU program as proof it can work at scale. Daniel frames both moves as evidence that AI companies are racing to control total cost of ownership (TCO) as memory and compute costs keep climbing. (The Decode) 5. The Flip: Will Custom Silicon Capture the Majority of AI Workload Dollars by 2028? Patrick argues every frontier lab is now pursuing vertical integration, citing Anthropic's new chip program, AMD's acquisition of Taalas, and NVIDIA's $20 billion license of Grok's inference technology as evidence the shift toward specialized silicon has become a step function. Daniel counters that every lab building custom silicon is simultaneously signing record contracts with NVIDIA and AMD, arguing that compute itself functions as the real moat regardless of chip architecture. (The Flip) 6. SpaceX Posts 92% Growth in Its First Public Earnings Report: SpaceX's debut public earnings showed 92% growth, narrowing losses, and rapidly expanding AI compute revenue layered on top of its launch and satellite businesses. Daniel flags new AI deals with Google and Anthropic as driving much of that growth, while Pat highlights that the company's stated $100 billion run rate target depends heavily on December performance and questions how soon TerraFab related capex will hit the income statement. (Bulls and Bears) 7. Palantir Delivers Its Fastest Revenue Growth in Company History: Palantir posted 93% year-over-year growth and record commercial revenue, with CEO Alex Karp highlighting that no company at Palantir's scale has grown this fast before. Patrick points to the company's sovereign AI positioning as being well-timed given growing enterprise distrust of frontier model providers, and Daniel adds that 150% commercial growth is far outpacing other enterprise software peers. (Bulls and Bears) 8. AMD Posts Record Growth as Investors Expect Even More: AMD delivered record revenue and growth. Investors had priced in results closer to NVIDIA's historic beats, and capex nearly quadrupled without much detail on where the spending is going. Moorhead traces the increase to AMD building out Helios rack-scale systems ahead of shipping, and both hosts note the timing of Elon Musk's public NVIDIA endorsement landed awkwardly on AMD's earnings day. (Bulls and Bears) 9. Astera Labs Beats on Revenue and EPS as the Market Shrugs: Astera Labs beat on both top line and EPS, posting record growth off a still small base. Pat points out new CXL innovations unveiled at FMS as evidence of the company's position building connectivity infrastructure alongside Broadcom and Marvell. (Bulls and Bears) 10. IonQ Raises Guidance and Closes Its $1.8 Billion Skywater Acquisition: IonQ raised its FY26 guidance from $280 to $290 million and closed its Skywater acquisition, positioning the company as a vertically integrated quantum player across networking, sensing, and compute. Daniel notes IonQ is approaching a billion dollar annual run rate once Skywater is included, and highlights a new Anduril partnership as a notable expansion into defense. Patrick adds nuance to the "only quantum pure play with foundry assets" framing, noting other quantum companies maintain foundry relationships of their own. (Bulls and Bears) 11. Lattice Semiconductor Beats Across Revenue, EPS, and Guidance: Lattice beat on revenue, EPS, and guidance, growing 62% and closing its $1.65 billion AMI acquisition. Newman points to the company's data center AI segment growing 83% as FPGA content keeps expanding inside rack scale systems, and both hosts frame Lattice as approaching a billion dollars in annual revenue. (Bulls and Bears) Watch the full video at sixfivemedia.com, and be sure to subscribe to our YouTube channel so you never miss an episode. The Decode TerraFab Breaks Ground in Grimes County, Texas https://techcrunch.com/2026/08/06/tesla-and-spacex-willinvest-16-8b-to-start-building-terafab-chip-factory-intexas/ FMS 2026 in Santa Clara https://news.samsung.com/global/samsung-unveils-next-gen-3d-memory-vision-at-fms-2026-charting-the-future-of-ai-infrastructure The Frontier-Model Compression Week https://www.reuters.com/business/retail-consumer/alibaba-unveils-its-most-capable-ai-model-date-not-far-behind-moonshots-size-2026-08-03/ https://huggingface.co/blog/ResterChed/deepseek-v4-flash-official-release https://www.reuters.com/technology/meta-launches-new-ai-coding-tool-powered-by-muse-spark-12-2026-08-05/ The Heterogeneous-Compute Week https://www.reuters.com/business/anthropic-build-in-house-chip-design-team-claude-hire-engineers-2026-08-05/ https://ir.amd.com/news-events/pressreleases/detail/1296/amd-acquires-taalas-to-advance-compute-solutions-for-rapidly-growing-ai-inference-market The Flip Will Custom Silicon Capture the Majority of AI Workload Dollars by 2028? FOR: https://www.reuters.com/business/anthropic-build-in-house-chip-design-team-claude-hire-engineers-2026-08-05/ AGAINST: https://ir.amd.com/news-events/pressreleases/detail/1292/amd-and-anthropic-announce-strategic-partnership-to-deploy-up-to-2-gigawatts-of-amdinstinct-mi450-series-gpus Bulls and Bears SpaceX Q2 2026 Earnings https://www.cnbc.com/2026/08/04/spacex-spcx-earningslive-updates-q2-2026.html Palantir Q2 2026 Earnings https://www.businesswire.com/news/home/20260802523449/en/Palantir-Reports-Q2-2026-U.S.-Comm-Revenue-Growth-of-149-YY-and-Revenue-Growth-of-93-YY-Raises-FY-2026-Revenue-Guidance-to-82-YY-Growth-and-U.S.-Comm-Revenue-Guidance-to-134-YY-Crushing-Consensus-Expectations AMD Q2 2026 Earnings https://www.cnbc.com/2026/08/04/amd-earnings-report-q2-2026.html Astera Labs Q2 2026 Earnings https://finance.yahoo.com/technology/articles/astera-labsreports-second-quarter-200500111.html IonQ Q2 2026 Earnings https://www.ionq.com/news/ionq-announces-record-second-quarter-2026-revenues-growing-287-yoy Lattice Semiconductor Q2 2026 Earnings https://www.investing.com/news/company-news/lattice-q2-2026-slides-record-revenue-ami-deal-drive-growth-93CH-4836197
The AI trade might have found its canary in the coal mine. Leopold Aschenbrenner—a 25-year-old former OpenAI researcher built Situational Awareness into one of the hottest funds on Wall Street by betting long on the construction and infrastructure side of AI (Micron, CoreWeave) and short on the chips and software everyone else was chasing (Nvidia, Broadcom, Oracle). Then July happened, prime brokers made margin calls, and his entire public stock book got liquidated overnight in a single block trade. This episode uses that blowup as the lens for a much bigger story: how a three-quarters-of-a-trillion-dollar hyperscaler spending pledge is propping up U.S. GDP growth, why OpenAI just quietly cut prices as “tokenmaxxing” dies and Chinese labs undercut on cost, why Mark Zuckerberg picked a fight with OpenAI and Anthropic, how overleveraged Oracle has become, how the most profitable companies in the world are posting negative free cash flow, why hyperscaler bonds are going undersubscribed, and why the length of this year’s tech bond issuance is starting to crowd out the rest of the corporate debt market. It closes with the 2008 parallel: mortgage-backed securities stacked into CDOs stacked into synthetic swaps, and how today’s version runs through corporate bonds, CLOs, hedge fund leverage, and private credit—with Treasuries as the eventual exit ramp when the unwind accelerates. Resources Analysis Atlas: Hyperscaler AI Capex 2026: The $710 Billion Year Inside the Four Spenders J.P. Morgan Asset Management: How AI demand and capex shape investing in tech stocks New York Times: U.S. Economy Slows as Inflation Bites U.S. Bureau of Economic Analysis (BEA): GDP (Advance Estimate), 2nd Quarter 2026 CNBC: OpenAI cuts prices for two of its GPT-5.6 AI models as companies grow sensitive to costs Business Insider: The All-You-Can-Eat AI Era Is Over. It’s Time to Count Calories CNBC: Chinese AI models are gaining ground with U.S. companies as OpenAI, Anthropic costs surge The New York Times: Even China’s A.I. Powerhouses Can’t Figure Out How to Profit Off A.I. CNBC: Meta’s flurry of AI initiatives this month hasn’t helped lift the stock. What will? The Hill: Zuckerberg knocks AI development centralization, control Buttondown: D.A.D.: Zuckerberg Blasts Anthropic and OpenAI for “Centralizing” AI Power Reuters: AI investment boom puts Big Tech’s free cash flow under pressure CNBC: Bets against SpaceX grow to 32% of float as Elon Musk warns short sellers won’t survive Yahoo! Finance: NVIDIA’s rising CDS the talk of Wall Street amid circular financing fears Bloomberg: Big Tech Debt Flood Is Taking Over Risk In Market: Credit Weekly CNBC: Fed meeting recap: Warsh says Fed won’t hesitate to stop inflation, but bond market has doubts CNN: AI investor Leopold Aschenbrenner forced to unwind all public stock positions after steep losses, sources say Yahoo! Finance: Hyperscalers Hit $700 Billion in 2026 AI Spending Plans - 24/7 Wall St. Global Finance Magazine: AI’s Financial Circle Game Reuters: Hyperscaler debt binge pushes yields up as investor demand cools UNFTR Resources Video: Situational Awareness: How a 25-Year-Old's Hedge Fund Exposed the Entire AI Bubble. Essay: The Great Unwind. -- If you like #UNFTR, please leave us a rating and review on Apple Podcasts and Spotify: unftr.com/rate and follow us on Facebook, Bluesky, and Instagram at @UNFTRpod. Visit us online at unftr.com. Become a member at unftr.com/memberships. Buy yourself some Unf*cking Coffee at shop.unftr.com. Visit our bookshop.org page at bookshop.org/shop/UNFTRpod to find the full UNFTR book list, and find book recommendations from our Unf*ckers at bookshop.org/lists/unf-cker-book-recommendations. Access the UNFTR Musicless feed by following the instructions at unftr.com/accessibility.Support the show: https://www.unftr.com/membershipsSee omnystudio.com/listener for privacy information.
Meta brengt Muse Code uit, een eigen codeertool die vanaf nu beschikbaar is in bèta en draait op Muse Spark 1.2, het nieuwste AI-model van het bedrijf. Daarmee neemt Meta het op tegen Claude Code van Anthropic en Codex van OpenAI, na een lange periode van achterstand op het gebied van kunstmatige intelligentie. Verder gaat Anthropic eigen AI-chips ontwerpen. Rosanne Peters vertelt erover in deze Tech Update. De tool is bedoeld voor het schrijven en verbeteren van software en werkt volgens Meta vooral goed bij langdurige, complexe taken en projecten. Voor Muse Code betaal je per gebruik, dus zonder verplicht abonnement. Wie Meta toestemming geeft de eigen code te gebruiken voor het trainen van toekomstige modellen, betaalt ruim tien keer minder. Het is het eerste grote project van AI-chef Alexandr Wang, die vorig jaar werd aangenomen om de AI-strategie van Mark Zuckerberg te versterken. Anthropic gaat eigen AI-chips ontwerpen Anthropic zet intern een team op dat eigen AI-chips gaat ontwerpen. Het bedrijf wil daarmee minder afhankelijk worden van hardware van onder andere Google en Amazon, maar zegt tegelijk gebruik te blijven maken van zijn huidige chipleveranciers. Persbureau Reuters meldde in april al dat Anthropic dit overwoog. Wanneer het ontwerp klaar is en wanneer de chip in productie gaat, laat het bedrijf niet weten. Ook is onbekend wie de chips gaat maken; eerdere berichten noemden Samsung als mogelijke partner. Anthropic is niet de enige: OpenAI werkt sinds vorig jaar met Broadcom en TSMC aan een eigen chip, die mogelijk al eerder in productie gaat. Volgens Reuters kost het ontwerpen van een geavanceerde AI-chip ongeveer een half miljard dollar, inclusief de kosten voor gespecialiseerde ingenieurs. Meta introduceert Muse Code en het model Muse Spark 1.2 Meta mengt zich met Muse Code in de strijd om AI-codeeragenten Muse Code is in bèta beschikbaar voor macOS en Linux Anthropic bevestigt eigen team voor het ontwerpen van chips Anthropic zoekt ingenieurs voor zijn nieuwe custom silicon team Over de maker:Rosanne Peters is techredacteur en maakt De Grote Tech Show en De Technoloog. Sinds 2025 doet ze redactie- en productiewerk en is zij te horen in de Tech Update tijdens De Ochtend- en Avondspits. See omnystudio.com/listener for privacy information.
Amazon zahlt OpenAI die vollen 50 Milliarden aus, obwohl keine der vereinbarten Bedingungen eingetreten ist. Google hat für Anthropic eine 200-Milliarden-Konstruktion aus Broadcom, Apollo, Blackstone, Morgan Stanley und ehemaligen Krypto-Minern gebaut. Pip erklärt dabei von Grund auf, wie Off-Balance-Sheet-Finanzierung funktioniert und warum niemand diese Chips in der eigenen Bilanz haben will. Daraus entwickelt er eine These: AI 2.0 kommt erst noch. Aus China kommen die passenden Belege, mit Alibabas neuem Modell und einem DeepSeek, das je Aufgabe ein Hundertstel der westlichen Konkurrenz kostet. Dazu Palantirs Zahlen und ein Rechenexempel über die Margen zwischen Chip und Kunde. Zum Schluss schauen wir in die SpaceX Earnings. Unterstütze unseren Podcast und entdecke die Angebote unserer Werbepartner auf doppelgaenger.io/werbung. Vielen Dank! Philipp Glöckler und Philipp Klöckner sprechen heute über: (00:00:00) SpaceX vor den Zahlen (00:13:06) Amazon investiert in OpenAI (00:17:25) Eine Milliarde Nutzer (00:19:51) Off-Balance-Sheet erklärt (00:23:23) Googles 200-Milliarden-Maschine (00:41:32) AI 2.0 (00:43:57) Alibaba Qwen3.8-Max (00:44:53) DeepSeek V4-Flash (00:46:41) Chinas Militär destilliert (00:48:49) Amazon bei drei Billionen (00:49:30) Palantir (01:00:33) Snap (01:01:07) xAI gegen Minnesota (01:03:05) Telegram im App Store (01:04:07) Google zieht Satellitenbilder zurück (01:05:25) Liechtenstein gehackt (01:08:47) Airtable an Bending Spoons (01:09:26) SpaceX-Earnings Shownotes Amazon schließt 50-Mrd.-Investment in OpenAI ab - ft.com OpenAI überschreitet eine Milliarde Nutzer - wsj.com Googles 200-Mrd.-Finanzierungsmaschine für Anthropic - ft.com Off-Balance-Sheet-Finanzierung erklärt - corporatefinanceinstitute.com Alibaba veröffentlicht Qwen3.8-Max - bloomberg.com DeepSeeks V4-Flash ist das günstigste bekannte Modell - reuters.com Chinesische Militärforscher trainieren mit US-Modellen - reuters.com Amazon überschreitet drei Billionen Dollar Marktwert - bloomberg.com Palantir-Quartalszahlen, Umsatz plus 93 Prozent - cnbc.com Palantir-Aktie steigt nach turbogeladenem Wachstum - marketwatch.com Snap stellt AR-Brille für 2.195 Dollar in Aussicht - bloomberg.com Snap-Aktie springt über 10 Prozent - cnbc.com Richter lässt Minnesotas Nudify-Verbot in Kraft - nbcnews.com Telegram nach kurzem Rauswurf zurück im App Store - reuters.com Google zieht KI-Satellitenbilder zurück - npr.org Hacker stehlen 31.000 Datensätze aus Liechtensteins Geldwäscheregister - zeit.de
Watch the full episode on YouTube:We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection. We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3:And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF:Three years ago, inference engineering barely existed as a category.Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem.In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out.Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles.In this episode, Baseten's Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API.We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model.The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them.We discuss:* What happens when a 200,000-token request enters an inference system* Cache-aware routing and reusing previously computed KV cache* Why prefill and decode are increasingly handled by different GPUs* When dedicated deployments become cheaper and more reliable than shared APIs* How speculative decoding uses a smaller model to accelerate a larger one* Tool calling, structured outputs, and what LLMs actually do* What it takes to support a new open model on day zero* Grafting Kimi's vision encoder onto GLM-5.2* Retrofitting inefficient model layers with components from other architectures* Why models sometimes collapse into repeating the same token* How hardware, kernels, and race conditions create nondeterministic failures* Preserving model fidelity while making inference faster* How quantization errors can cancel each other out* Why inference optimizations still deliver gains of 20%, 100%, and 200%* How optimized serving can make a model up to 10× faster* NVIDIA Dynamo, KV-aware routing, and distributed model serving* Speculative decoding the speculative decoder* Why local AI is about making models less dumb while data-center AI is about making them less slow* Tensor, expert, and pipeline parallelism across GPUs* Hardware-aware model design, auto-tuning, and the case against mega kernels* Rubin and why inference is becoming a systems problem* Whether modern GPUs are evolving into programmable AI ASICs* Why enormous models like Kimi K3 require GB300-class hardware* Why open-source video generation still trails Veo, Kling, and other closed models* The quadratic attention bottleneck behind long-form AI video* Autoregressive video, real-time generation, and compounding quality drift* Why future video systems may combine autoregressive and diffusion architectures* Training for inference and inference for training* Continuous post-training, deployment, evaluation, and improvement loops* How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself* Why faster networking could unlock dramatically faster decoding* Continual learning, KV-cache compaction, and persistent model memoryShow Notes* How to build a day-0 API for Kimi K3* 22580: From GPT2 to Kimi3, ExplainedPhilip Kiely* LinkedIn: https://www.linkedin.com/in/philipkiely* X: https://x.com/philipkiely* Inference Engineering: https://www.baseten.co/inference-engineering/Ali Taha* LinkedIn: https://www.linkedin.com/in/aliestaha/* X: https://x.com/waterloointernTimestamps00:00:00 Introduction and the 200K-Token Prompt00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling00:11:26 Launching Production-Ready Open Models00:19:06 Model Retrofits, Failure Modes, and Nondeterminism00:28:22 Quantization and Canceling Errors00:32:15 The Race to 10× Faster Inference00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips01:10:03 Giant Models and the Limits of GPU Memory01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation01:21:47 Audio, Images, and Diffusion Models01:27:32 Training, Self-Optimizing Models, and Continual Learning01:40:06 Closing ThoughtsTranscriptIntroduction: Baseten, Waterloo Intern, and Inference EngineeringSwyx [00:00:00]: Okay, we're here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you've done, you and I have done before, as well as Ali. Welcome.Ali [00:00:15]: Pleasure to meet you.Swyx [00:00:15]: Waterloo intern.Ali [00:00:16]: Waterloo intern, always.Swyx [00:00:17]: When did you get “Waterloo intern” as a handle?Ali [00:00:19]: As a handle? Oh.Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.”Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer.Philip [00:00:30]: So we have to figure out who's gonna get the handle.Ali [00:00:33]: Well, I'll pass the torch over to the next intern.Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad.Ali [00:00:37]: To another Waterloo intern. No, bruh.Philip [00:00:39]: Yeah.Ali [00:00:39]: Intern.Swyx [00:00:40]: Intern, yeah.Ali [00:00:40]: And no.Philip [00:00:41]: You gotta get an intern from Waterloo.Ali [00:00:42]: Yeah, I've gotta get an intern from Waterloo.Swyx [00:00:44]: Right.Ali [00:00:44]: But they have to follow the path.Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it's like whoever Baseten gets from Waterloo.Ali [00:00:48]: Right.Swyx [00:00:49]: Has the title of Waterloo.Ali [00:00:50]: It stays in the ecosystem.Philip [00:00:51]: Exactly.Ali [00:00:52]: Halfway through the internship, you either get it or you're out.Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle.Ali [00:00:59]: Just say it.Philip [00:00:59]: For everybody.Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you're an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten's inference? What's the process of query through GPU model routing, balancing, all that? What is all the stuff that we don't think about?Long Context Requests, KV Cache, and Cache-Aware RoutingPhilip [00:01:26]: With a long query specifically, the first thing that I'm gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it's gonna be a lot easier for me and a lot cheaper for you. So the first thing that we're gonna look at is some cache-aware routing, where we're going to see, we probably have a number of instances, a number of replicas up serving whatever model you're hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you're doing two hundred thousand tokens, it's probably coding or a multi-turn agent or something where you would expect to have that cached. If you don't, we're gonna have to send it to a prefill worker. We've at least on certain models disaggregated prefill and decode, so you're going to have one set of GPUs that's solely going to process the input, create the KV cache, and get you your first token, and then that's going to be passed over to a separate set of GPUs, which is going to run decode. We're going to iteratively make those tokens. We're probably going to have some speculator model in front of that. I'm going to assume that you're doing coding, and because of that, our speculator model, which assumes you're doing coding, is gonna have a high draft token acceptance rate. If I'm wrong and you're asking me to summarize every Harry Potter book, it's gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?”Swyx [00:03:04]: Except Baseten doesn't charge by pennies.Philip [00:03:07]: Well, yeah, we charge. I'm assuming that we're talking about the public model APIs. If you are setting up a dedicated deployment, then yeah, it's not pennies.Public APIs vs. Dedicated DeploymentsSwyx [00:03:18]: Yeah, one of the key differentiators when I was talking with Baseten initially was that people who want very high volume just need to rent by the box, ‘cause then it's up to you to figure out how to saturate the box.Ali [00:03:31]: And more often than not, it's, like, way cheaper if you're pushing, like, millions of tokens per hour, if you just pay per hour instead of pay per token.Philip [00:03:37]: Yeah, they do. I think that we've increasingly seen a lot of demand for the pay per token APIs, just because everyone wants to try open models, and then once they find a use case that's really sticky, then they move over to dedicated.Swyx [00:03:51]: Is there a best practice on when it's time to swap over?Philip [00:03:54]: Couple reasons. Yeah, reliability, that's a big one, right?Ali [00:03:57]: Like, if they have a very specific use case, they want you to train something specifically for them, like they want their own spec dec, for instance, for their own traffic.Swyx [00:04:04]: Spec dec is speculative decoding.Speculative Decoding and Custom SpeculatorsAli [00:04:05]: Speculative decoding, yeah.Swyx [00:04:07]: You have to explain.Ali [00:04:07]: Sorry. Like, speculative decoding is like, if you have a huge model, right? And so the model is going to be generating one token at a time every single turn, every single forward pass. So we attach, like, this little, like, parasite, like this layer that goes on top of the model, and this model just has to predict. It does three very fast autoregressive forward passes, and it will predict, like, three certain tokens, and then you do one forward stage over the entire original model in order to see if those predictions were correct or not, and then you accept them or you reject them. Now, this draft model is traffic specific, so if you, like, Philip said, if you're summarizing Harry Potter books, I can train exclusively that draft model on Harry Potter books, and I can guarantee you that I'm gonna accept the three tokens every single time. And so with that case, I increase your decode speed. I wouldn't be able to provide this to you if you're a shared endpointSwyx [00:04:53]: YeahAli [00:04:53]: ‘cause I have no idea if you're doing Harry Potter, if you're doing coding, if you're doing English. We don't know. Also, there was a thing in the book that mentioned that if they really cared about a specific threshold, chapter four, I think. Do you remember that?Philip [00:05:06]: Yeah. The things that you can do is you can set a specific, like, batch sizing, a specific, like, parallelism strategy if you're trying to optimize for, like, throughput versus latency. You can. Maybe a NVFP4 quant doesn't pass your benchmarks and you wanna run a model at higher precision, you could do that. There's just a bunch of reasons why you might wanna have your own endpoint and the biggest one, of course, just being, like, you don't have to deal with someone else doing a hundred million tokens of benchmarking traffic at the endpoint when you happen to be trying to serve your users.Swyx [00:05:40]: Yeah. I think one thing that is. That is a classic journey. Like, it's people is asking the, what happens when you type Google into the browser. Tool calling, is that just, you're generating JSON or is there more complication beyond that?Tool Calling, JSON, and Structured OutputsAli [00:05:58]: Certain customers that we have, they have their own post-trained models, and so they demand a tool calling that's not just, like parse a file or go find the weather. It's something that's very specific and you have to do post-training on this. And if the post-training on the model is not good or if the quantization after the post-training to get the inference to be fast, the model will struggle reading the JSON file and reading the tool calling. But it doesn't require its own like sandbox. It's not like it's going to use that tool calling to like escape a sandbox or like it doesn't have to be contained. It can just be a normal dedicated deployment. The challenge with tool calling more and more seems to be that the companies want certain tool calling which is a very sensitive thing to train. And because you're dealing with all of the JSON outputs, if it doesn't like close the end of the request in a very certain manner, you end up with a model that did the tool calling and like the thinking and so as a result of that, it didn't see the result and just hallucinated the result as it decoded. That seems to be the most challenging thing with tool calling, not really the sandboxes model.Philip [00:06:56]: Yeah, that's a challenge on the training side and then on the inference side, there's work that you can do to scope the possible output. So we published this at this point close to two years ago, the solution to this problem which is you make a state machine and you use that to constrain the output to a specific format. So this is the structured output problem. If you remember backSwyx [00:07:27]: Yeah, the specific grammar is,Philip [00:07:29]: Yeah, exactlySwyx [00:07:30]: GML had this thing.Philip [00:07:31]: Yeah. So it's like the old-school “make sure this is only JSON”, return only JSON orSwyx [00:07:38]: YeahPhilip [00:07:38]: Grandma's gonna die type of prompts.Swyx [00:07:39]: Is it BNF grammar? At some point OpenAI had released a thing that was like, yeah, if you want to constrain your output, write BNF grammar, back as NOR.Philip [00:07:47]: In our inference system, it's just a specified output format. And you get the guarantee that your output's gonna be structured along that format. And so applying that to tool calls can like help cut down on. You can still call the wrong tool or call no tool. It doesn't solve the certainty problem but it at least solves the output structuring problemSwyx [00:08:10]: YeahPhilip [00:08:10]: Within tool calls.Swyx [00:08:12]: And MCP is just another form of tool, right.Philip [00:08:14]: Yeah, exactly.Swyx [00:08:15]: As far as there's no special thing there.Philip [00:08:16]: The thing I'm always like explaining to people is the LLM is not capable of doing anything. It's only capable of making suggestions of what to do and then if those suggestions are formatted in a certain way and applied to a system that knows what to do with them, then an action occurs.Swyx [00:08:32]: Yeah. Part of the fun stuff is, this is solved outside of tool calling too. Like in an agent loop if the output is not correct or you're right, like reasoning, tool calling was done in the reasoning trace, just be like, “Oh, I don't know what to do. Let me just try again.” And it might get there after a few tries. And on your point of training, sometimes this is harder in smaller models, so you don't have the same exact quality outputAli [00:08:56]: Right.Swyx [00:08:57]: When you just swap from a big model, right?Ali [00:08:59]: Yeah. I will say that, before, I think we need to go back to inference engineering proper.Ali [00:09:04]: But, I had expected that something would replace JSON because it's hard to stream JSON ‘cause JSON must be complete and you must have open and close brackets and everything. So it's hard to parse something or validate something while it's being streamed. So people invented all sorts of things that are like, I forget the name of some of these alternatives, but it's something like TOML, something like YAML. But JSON seems to be dominant still.Philip [00:09:30]: The JSON outputs aren't that long, right? Like you could have a long-- ‘cause tool calls also contain the arguments in them and perhaps for a certain tool you might pass like a very long argument. But my impression of the median tool call is that it's a relatively small number of tokens, right? So I would expect that speculators are generally fairly good at something as formatted as JSON. And so you would have like a pretty fast decode step there and that the streaming wouldn't be as valuable, but maybe I'm wrong about that.Ali [00:10:02]: I think you're also bounded by the software or that the model is gonna integrate with if the software is built with JSON for the tool calls or if the company that you'- if your customer says that this is how our software works and our tools are interfaced with JSON, you can ask them to like, change their software and say like, “Yeah, this is gonna be better for the model.” but like with the right training shouldn't be that much of a difference. Also more profitable if it outputs more tokens probably.Swyx [00:10:25]: Depends on your business model.Swyx [00:10:27]: It really depends. But I will say that, as a writer with like experience a lot with generated output, I do try to move from text to JSON text which is very long JSON, right? Like there's paragraphs in every field because I'm trying to structure it, right?Philip [00:10:44]: Right.Swyx [00:10:44]: I want you to first make factual statements, then make opinions then make bullet point summaries, have dates, have entity references have your sources for references, all these things. Anyway, so these are things that like I think people who really experiment with structural output have to really care about. But, let's, let's recurse up the stack a little bit. Before we started recording, you mentioned something really cool, which is that there's a lot of engineering that-- inference engineering that goes on when a new model provider releases a new model, right? So let's call it GLM-5.2, Kimi K3. I had previously assumed, especially if it's like, well, GLM 5 to 5.1 to GLM-5.2, like that you've supported them before. Is it that much work?What It Takes to Support a New Open ModelAli [00:11:26]: It's a lot of work.Swyx [00:11:28]: Yeah. Okay. So like, a lot of people, all you guys, right whenever a new model launch like, people rush to say like, “Oh, Hugging Face supports this, Fireworks supports this, Spacetime supports this,” and I'm like, “Yeah, of course we support it.” But what goes into that? What goes intoPhilip [00:11:40]: I think it's more than just support it too, right? It benefits the consumer a lot. Like I think it was with Kimi K2.5 or GLM-5.2 the latest, there was an inference war, right? X provider is at 90 tokens a second. The next day we're at 150. The nextSwyx [00:11:55]: I kinda kicked that off with the GLM-5.2.Swyx [00:11:58]: I wrote a Twitter article about. It got like half a million views,Ali [00:12:02]: Based on being numberSwyx [00:12:03]: YeahAli [00:12:04]: Or it's for something else.Swyx [00:12:05]: Yeah. Which,Ali [00:12:06]: Oh my GodSwyx [00:12:07]: Which then got everyone really excited about, hey, how can we, bend tracks a little bit further and,Philip [00:12:14]: There's a difference between support the model, as in I can make a token out of this model, and support a model, as in I have a production-ready API from this model.Philip [00:12:26]: Getting to the point of I can make a token out of this model is not that hard because generally the, open source inference engines, vLLM, SGLang of the world oftentimes even receive weights ahead of time, maintainers do, or the people making the model merge PRs to ensure support. So you generally can, just get it working on the standard open source stack without too much pain in most cases. The challenge is, every inference company is gonna have own proprietary stack. Some open source components, some in-house stuff. And for any arbitrary model, there's going to be some new stuff. Sometimes you get lucky, like K, two five to two six was, like, pretty similar.Quantization, Speculators, and Production ReadinessAli [00:13:16]: Yeah. It was pure continued post-trainingPhilip [00:13:18]: YeahAli [00:13:18]: If I remember correctly.Philip [00:13:19]: Even in those cases, there's still stuff you have to do. You have to redo the quantization work. You're taking the model from. Generally, these models are not released in NVFP4, and we want them to be in NVFP4 for maximum Blackwell compatibility. So we have to perform that quantization, and, calibrate the quantization to make sure that we're not causing any regression in the model's intelligence. And then we also have to train the speculator, as we've talked about. Generally, we have. We have ZDR, zero data retention on our model APIs, so we don't know exactly the traffic that people are sending us, but we know what's popular. We know that coding use cases are popular. We know that agents, agentic use cases are popular. So we can get public data sets that are representative of that traffic and train general speculators. Now, with speculators today, you need to train the speculator using the base model itself because you're getting hidden states out of the model from running inference on these specific prompts, and that is the training data you use to create the speculator. So there's that process which you need the real model weights for. And then there's of course just the process of, standing up all the infrastructure behind it, loading all this stuff, testing it. And then when there's a new model with a newer architecture, I think that, like, the DeepSeek models tend to be the most challenging as they have, like, the most novel architectural stuff going on, model after model. But every new model has something. Kimi K2 had. Oh, sorry, GLM-5.2 hadAli [00:14:53]: Sparse attention.Philip [00:14:54]: Yeah,Ali [00:14:54]: YeahPhilip [00:14:54]: the DSA.Ali [00:14:55]: Right. Which is brought from DeepSeek.Philip [00:14:57]: Yeah. AndAli [00:14:59]: So you can copy-paste then?Philip [00:15:01]: It kindAli [00:15:01]: I don't know how this works.Philip [00:15:02]: So, like we had to, like, build support for that into our runtime. And you're right, like it is really interesting the way that all of these open source labs borrow from each other. For example, like GLM-5.2 doesn't have vision. So something that, Haley, a guy on our team, if we could take a look at this, he, like, grafted the Kimi vision encoder onto GLM-5.2.Retrofitting Vision into GLM-5.2Ali [00:15:27]: We'll be training the projector.Philip [00:15:28]: Exactly. So if you think about, like, the encoder, there's the encoder, which is the part that looks at the image and turns it into latent information, and then there's the projector which likeAli [00:15:38]: You can say latent space. It's okay.Philip [00:15:41]: And then there's the projector that maps it onto, the model itself, and then there's the model weights. You don't wanna mess with the model weights because you run a chance of making the model dumber at something else for the purpose of giving it vision. So instead, Haley started with just a projector, which is only a handful of millions of parameters.Ali [00:16:02]: That would be, yeah.Philip [00:16:02]: Yeah.Ali [00:16:03]: Can you show the training one?Ali [00:16:04]: Like the way it groksPhilip [00:16:05]: YeahAli [00:16:06]: Very interesting.Philip [00:16:06]: And maybeAli [00:16:07]: That right therePhilip [00:16:07]: Maybe Ali, you should take it from here. You've got a betterAli [00:16:10]: Ooh, double the sandPhilip [00:16:11]: Understanding of this than I do.Ali [00:16:11]: Yeah. You can see, like, he. The way he trained this is really cool. At the beginning, he was training it using just like, “Here's a picture of a mountain. Can you describe what's in this mountain?” And that caused it just like the first, learning walls. Like here you can see this all we're trying to teach it is to translate the encoded. Like it's already taken the encoder from Kimi K. It's taken the image. It'Philip [00:16:31]: Yeah. FrozenAli [00:16:31]: FrozenPhilip [00:16:32]: With adapter.Ali [00:16:32]: Exactly.Philip [00:16:33]: Yeah.Ali [00:16:33]: So the brain is frozen and the eyes are frozen. It's just we're tryingPhilip [00:16:37]: AlignAli [00:16:38]: Interconnect between the eye and the brain, right? So the projector. And so you take the tokens and then he's like, “Oh, can you describe what's in this image?” And he's like, “Oh, it's a mountain,” or it's a person or it's a human, whatever the case is. But that didn't cause complete understanding. So he changed it such that every image was associated with a data set of questions. Like, does this image have a white male? Does this image have birds in the top corner? Does this image have a scientist in it? All of that stuff. And it would have to answer questions correctly. And using not just training on describing an image, but being able to answer question, another question, answer over time. Like you can see the grokking, which is like genuinely insane, that retrofitting vision into a large LLM can learn to that extent. And even for images that it doesn't perform well on, for instance, if you ask it a picture of like Stephen Hawking, “Who is this?” Maybe it doesn't get it, but it will say something like, “This is Albert Einstein.” Like it still understandsPhilip [00:17:25]: Close enoughAli [00:17:26]: That this is a scientist who is a man who has, some significant achievements, all that stuff. So that's like really cool.Philip [00:17:32]: Yeah. So, we've covered Hao Tian before, who the author of the LLaVA paper that did this, a while ago. And I think that's very foundational work for anyone who hasn't done vision work before.Ali [00:17:41]: Same with the CLIP and MetaCLIP, where you go from just captioning to building out questionsPhilip [00:17:47]: RightAli [00:17:47]: Off the image and how much better you can get performance.Philip [00:17:50]: Right. Right. Right. Yeah. But what's, what's so exciting about this is if you look at a model like this. Now, this is a little bit more of a research project. It's not. It got to 56% on MMLU Pro, I think. So not quite frontier. But if you're running this model, you haven't suffered any loss on your GLM-5.2 quality. If you don't have an image, it'll just behave exactly the way it used to. And ultimatelyAli [00:18:14]: Which in the inference code you literally do not include the other part, right?Philip [00:18:18]: Yeah. You would just skip the encoder if you don't have an image input.Ali [00:18:22]: Okay.Philip [00:18:22]: Just confirming.Philip [00:18:23]: YeahAli [00:18:23]: Does it affect a lot on the overall inference side? Like you're not adding much, you're adding a very small vision encoder. These are typically likePhilip [00:18:30]: They're super fineAli [00:18:31]: Less than a billion parameters, right?Philip [00:18:32]: Yeah. It's, - There's a little bit less standardization among vision encodersSwyx [00:18:37]: YeahPhilip [00:18:37]: So the support matrix can be a little bit, sparser. But overall, yeah, it's a pretty, it's a pretty minor component of the overall system. And ultimately what you get out of the system is all of a sudden you have Kimi Vision, GLM weights, and DeepSeek attention all in one model.Open Source Model Grafting and Franken-MergesPhilip [00:18:56]: And that's, I think, a lot of the power and beauty of open source, is that you can take all of these different components and combine them together into a system that's better than anyoneSwyx [00:19:05]: YeahPhilip [00:19:05]: Can be individually.Swyx [00:19:06]: People used to say that you would also do Franken-merges where you would take likePhilip [00:19:10]: YeahSwyx [00:19:10]: Layers from each model.Swyx [00:19:11]: Does anyone do that anymore?Ali [00:19:13]: Well, to your point previously when you were mentioning like, the work that goes into supporting a model when it first comes out, like GLM-5.2 or MiniMax M3 or whatever the case is. Sometimes you do have to like, you do have to switch out some things. Like, for instance, the MiniMax M3 head uses full attention, and with full attention you end up with this like insane bottleneck in spec dec ‘cause you're doing auto-regressive token generation for three tokens, and you're doing this like N squared over all of the tokens that are in your sequence. Your KV cache is like very large because it's not sparse, it's not top K. So we find it better to like, okay, we're gonna replace this, we're gonna replace this layer with a layer from another model that's using like GQA, for instance. And then just with the right training, you can get it to have the same acceptance rate. So it is very possible to retrofit layers from other models and very much needed. If a layer is like inefficient, the training just becomes the challenge, like how do you ensure that you train it properly? Which again to your earlier point is like the mesh between training and inference. As in like you need very good training in order to do fast inference. That's like, I feel like more and more becoming true.Swyx [00:20:21]: Yeah. Anything else on the support side when you say like get it to fully production ready?Loop Detection, Race Conditions, and Non-DeterminismPhilip [00:20:26]: Yeah. I think that there's also a question of just, we can test a model to a pretty extensive degree, but we're trying to get it out quickly and then you see a bunch of other people test it and you get interesting results. There was an issue with, GLM briefly where we had some like mode collapses where it would just output the same token over and over again for certain prompts on certain temperatures. Like once you expose an endpoint to the real world, there's going to be, so many more varieties of things given to it that you're able to, discover and patch things. So it's not just a, day zero process, it's then like for the first week, for the first month, if a model remains popular, like how do you both fix bugs and then continue to push the envelope on performance?Ali [00:21:21]: What do you mean you don't want your model outputting S?Swyx [00:21:24]: Is there loop detection on that stuff, by the way? It still happens like quite a lot, which is surprising.Ali [00:21:30]: We have like we, in our endpoint, like if a model was to output the same token like four plus times, we just cut the generation. We say like, “Oh, sorry, this-- Like try again,” or like we will reprocess the request. ‘Cause we know then, like if it, like if, yeah, it's four times the same token, it's probably collapsed.Swyx [00:21:45]: Yeah. Is there a way to opt out in case I really want that?Ali [00:21:48]: You want that?Ali [00:21:50]: I think there's a way that we have to handle it. I'm not exactly certain, but I feel like in certain models, like when they output something like you can imagine, like a table for instance, and so they want, they wanna draw like 12 dashes and 12 dashes. Yeah, I think there's a way for that to happen. I think we only do it on certain tokens. Like we exclude certain special characters.Swyx [00:22:07]: Yeah.Ali [00:22:07]: So we only do it on like certain like S is the most common almost. GLM-5.2Swyx [00:22:11]: OhAli [00:22:11]: And I think it was DSV 4 as well. Like you'd just have like looping issues where like you literallySwyx [00:22:17]: ItAli [00:22:17]: Just have like S.Swyx [00:22:18]: Yeah. Is there a special, something special about S? No, just randomlyAli [00:22:21]: It just seems to be the one token involved.Swyx [00:22:23]: Yeah. And it'Philip [00:22:24]: Is thereSwyx [00:22:24]: And it's only temperature 0Ali [00:22:27]: NoSwyx [00:22:27]: Even at other temperaturesAli [00:22:27]: Even at like 0.9 or whatever, it will still, it will still collapse.Swyx [00:22:30]: That's weird, right?Ali [00:22:30]: It's, it is an inference problem to be honest, like a software problem. Like oftentimes, the image you run will-- like NVIDIA will release an image for instance, and if we will upstream the changes from their latest TensorRT-LLM image into our stack, we'll find that it fixes it. Or oftentimes this will only happen in an inference engine that you're using like SGLang. But if you were to switch to vLLM, that isn't the case. So it seems to be like an extremely like deterministic software issue and not really a model issue. It's not like a weights problem. Like I'- we'll say like, “Oh, it's a problem with the quant. We did PTQ wrong,” right? But that isn't, that doesn't make sense because the same weights used with a different inference engine does not repeat the problem. And sometimes it's, the kernels that are being used in the backend have like these very subtle sometimes race conditions, where if you were to use this model hosted on one cluster, you will never get this problem.Swyx [00:23:19]: Oh my God.Ali [00:23:19]: But if you host it on a different cluster, you will. And the reason is the KV cache transfer from a node to node in that one cluster is using a slower interconnect than the node to node in another cluster. So that exposes the race, whereas in another cluster it doesn't. So then you end up just like, okay, this model is not gonna be hosted on this cluster. We're gonna host it on, another cluster because that cluster exposed that problem. But then it ends up with like, okay, is it the software? Is it the model weights or is it the hardware?Swyx [00:23:42]: There is a thing about this with temperature 0 still not being deterministic, right?Ali [00:23:46]: Right.Swyx [00:23:46]: Mostly because of hardware. Even at temperature 0 same model, you won't always get the same output.Swyx [00:23:52]: Even-- But I'm surprised by the race condition one because, I thought PyTorch was a graph that like guarantees that you at least, execute things in the right order.Ali [00:24:02]: Well, yeah, true. Like I'm not, I'm not saying that there is. Like well, you have things like PTL optimizations where like you can start a kernel before the end of the previous kernel, and that's like ‘cause you want to do that because there'sSwyx [00:24:12]: It's like pipeliningAli [00:24:12]: Expense. Exactly.Swyx [00:24:13]: Yeah.Ali [00:24:13]: But it'- But you don't do it cleanly. Like you overlap a little bit of the execution. No, it is very possible that the kernel itself, like that one block that is supposed to be running in this instance of time, that kernel itself has a race condition. For instance, like a missing barrier. Like often if you're designing a kernel and you want it to make it to be very fast, if you don't test it extensively, you'll, you'll have certain threads access data points from registers before they've been written to by other threadsSwyx [00:24:36]: YeahAli [00:24:36]: For example, because like your barrier is wrong or your synchronization was wrong. But yeah, like the testing itself is very difficult in those like, andSwyx [00:24:42]: And there's no like borrow checkerAli [00:24:45]: What does that mean?Swyx [00:24:46]: Like Rust. Like the. If you're trying to have like memory safety It sounds like a comparable problem.Ali [00:24:52]: Well, yes, but you're working in CUDA, right, NVIDIA GPUs. Like- You just need a higher level language like modular Maybe that's what modular is supposed to do. I don't know.Quantization Quality and Vendor FidelityVibhu [00:25:00]: How do you see keeping quality of the model? So you talked about all these steps of, okay, you gotta do quantization, train your own speculative decoderAli [00:25:07]: RightVibhu [00:25:07]: Run on different hardware. Looking at other model providers, okay, you kicked off a inference speed race on the consumer end. What goes into keeping quality the same across them, right? Sure, you can run benchmarksAli [00:25:22]: YeahVibhu [00:25:22]: But, like, how do you determine how much quantization are there standards? What goes intoPhilip [00:25:27]: There's a few things on quality. Most inference optimizations are lossless. KV caching, for example. You are just recomputing or preventing recomputing the same values. Speculation, of course, if a draft token is wrong, it gets rejected. The main lossy optimization is quantization. And that really comes down to, number one, data format, number two, which parts of the model you choose to quantize, which layers, and number three, like doing a lot of calibration on the quantized weights, to ensure that you're preserving all the outliers. There's other tricks that you can do, though. A big one is long context, ‘cause one thing you asked at, right at the beginning is, “Oh, what's gonna happen if I send a 200,000 token request in?” So with a long input sequence, you need to, store a lot more information. You need to process a lot more tokens. And so even if a model has a context of a certain length, you might, as an inference provider, choose to build an API with a shorter context length, and of course a full length one as well. Because if someone doesn't need the full million token context, for example, you can get them better performance. I don't know if that's exactly like quality of the model. The way that I think about quality is to what degree are we faithfully serving the original model? If you think of a golden implementation of a model that performs exactly the way the model is designed to perform, I think of quality as how close are we getting to that, 100% fidelity of the model.Philip [00:27:13]: You can also, of course, think about quality from the training side and how do you push yourself past 100%. But when I think about purely inference optimizations, it's getting faster while staying as close to that 100% fidelity mark as possible. And certainly our standard internally is that, like you should not be able to tell the difference between our API and a, official API. I think Kimi in particular does a good job of vendor benchmarking hereAli [00:27:41]: YesPhilip [00:27:41]: Where they haveAli [00:27:42]: They released an actual vendor benchmark.Philip [00:27:43]: Exactly, yeah.Ali [00:27:44]: ‘Cause they accused, some people, Amazon? There was some provider that was not doing very well on Kimi's benchmark.Philip [00:27:50]: Yeah.Philip [00:27:51]: So, with Reflect we probablyVibhu [00:27:52]: This was a long time ago, right?Philip [00:27:54]: No.Ali [00:27:54]: Yeah, like threeVibhu [00:27:55]: They alsoAli [00:27:55]: Four, five months agoVibhu [00:27:57]: This also happened with, I don't remember which model, but they pulled out quite a few, and then they started a whole chart about this. It might have beenPhilip [00:28:03]: Kimi Vendor Verifier.Ali [00:28:04]: Yeah.Philip [00:28:05]: Yeah.Ali [00:28:05]: Yeah, ‘cause you, ‘cause you'd be pissed, right? Like if you'Philip [00:28:07]: Yeah.Ali [00:28:07]: If like if I'm a consumer and I'm using like Amazon's endpoint for instance, and I've used Kimi and I'm like, “Oh my God, like this is bad,” I'm not gonna say, “Oh, Amazon quantized the model in a bad way.” I'm gonna say, “Oh, Kimi sucks.” Right?Philip [00:28:17]: Yeah.Ali [00:28:17]: So it seems like that makes sense.Philip [00:28:19]: Yeah, they care. They care.Vibhu [00:28:21]: Justifiably.Ali [00:28:21]: Yeah, justifiably.Vibhu [00:28:22]: This is probably a stupid question, but just checking, has anything improved from main quantization?Philip [00:28:28]: Yeah.Vibhu [00:28:28]: Like, is quantization always strictly worse?Ali [00:28:30]: Well technicallyVibhu [00:28:32]: NoAli [00:28:32]: It's a lossy. QuantizationPhilip [00:28:33]: YeahAli [00:28:33]: Is a lossy, it's a lossy implementation.Philip [00:28:36]: Speed improvesVibhu [00:28:36]: Speed improves.Ali [00:28:37]: It the number, likeVibhu [00:28:38]: No, I' always look for inverse scaling laws.Philip [00:28:40]: Yeah.Ali [00:28:40]: Yeah.Vibhu [00:28:40]: This is something I learned from Noam Brown, where like things that normally act in one direction sometimes do.Philip [00:28:45]: Well, technically when you run a benchmark, because these models are deterministic, sometimes your,Ali [00:28:52]: YeahPhilip [00:28:52]: NVFP4 quant is like, two basis points higher than yourAli [00:28:56]: No, it's noise. It's noise.Philip [00:28:57]: Yeah, exactly. I'm like, yeah, it's, it's within. That's why I always say within margin of error.Philip [00:29:01]: And I stopped saying that because everyone assumes that what is, well, within some margin of error, we're barely inside of that to the worst, so we're saying. But yeah, sometimes it's just like, gives you a higher output score. But like Ali said, that's noise. To my knowledge, you're not necessarily making the results better. You're just trying to, again, like keep your fidelity as close to 100% to the original model.Layer Selection, KL Divergence, and Better QuantizationAli [00:29:27]: There is, to your point, research that we did on MP. I don't know if you are able to pullPhilip [00:29:31]: YeahAli [00:29:32]: A tweet we did. One of our research interns, Joshua, I think it's a tweet on how we have 20% better quantized GLM-5.2 than NVIDIA. Essentially what we found throughout like this month research is, okay, quantization is a lossy. It's. You're compressing the data from, occupying 16 bits to occupying, four bits, for instance. And so you're losing some information, and you're trying to minimize that. And so when I say that I'm gonna quantize the model, my job becomes how do I find the layers that I can quantize, and how to find the layers to not. For instance, with image models, I don't quantize modulation layers, and I don't quantize out projections because those two are. Like out projection is what you see as the user. Modulation is what the model sees or understands. Right, exactly. And so to his paper, do you have the. It doesn't have the. Yeah. It's a long paper. I don't know if I can findVibhu [00:30:25]: If there's a part to search or it's probably in the thread.Ali [00:30:28]: It's probably in the thread.Vibhu [00:30:29]: Yeah.Ali [00:30:29]: But the long and the short is it is very possible that quantizing more of the model makes the results. Like if I have a model that I quantize layers one, five, and 10, and another model where I only quantize layers one and It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in, is that you can predict which layers are going to have quantization errors that will cancel out with each other, and you choose to quantize those layers. And so the result of doing this mathematical quantization is you end up with a model that's 20% more quantized than another provider, so you get 20% more throughput of it because there's more layers than running an NVFP4, and your quality is better than that other quant because the layers that you chose to quantize have their errors cancel out, like one layer skewed to the right one layer skewed to the left, one layer skewed to the right. Your final logits distribution is more similar to the original distribution of the model, so you have better fidelity. And so the way we proved this was with KL divergence. So instead of just scoring on the benchmarks, we scored the KL divergence between the logit distribution of the quantized model and the logit distribution of the original full precision model, and we showed that with this technique we get. If your probability distribution on the logits which token it wants to select is more of the same as the original model, you're probably gonna end up staying true to the original model. So yeah, so it seems like previously before this, it seemed like the industry was, well, the more you quantize, the worse it's gonna be, ‘cause the more loss you introduce. That's not exactly, not necessarily true. So yeah, doesn't improve it, but can cancel out.Philip [00:31:57]: I think it might be this, but reminds me a good bit about pruning where you can prune off certain layers.Philip [00:32:03]: But very interesting. Didn't know this was a whole paper you guys put out.Ali [00:32:06]: It's. Fun fact, it was originally 72 pages, this paper, and then we decidedPhilip [00:32:11]: WowAli [00:32:11]: We can't tell. We couldn't release it. So it's now 45.Swyx [00:32:15]: Still 39 pages, so very substantive. We talked about evals and all these things and, like what's possible in terms of speedup? Like it's like probably like the numberInference Speedups and BenchmarkingSwyx [00:32:25]: Thing that people do wanna care about, and it's something that you wrote about in your post. Like official API is 70 tokens per second, and you push it up to 90. Is that like a normal thing?Philip [00:32:36]: So what's cool about working in inference, the reason that I think inference is going to be a useful place to do engineering for a long time, is that if you look at highly optimized domains like, say, finance, if you're in finance, you measure how much better you got in basis points. It's like, “Oh, I got five basis points better, like twentieth of 1% better,” that's huge news because everything is so optimized. When we publish optimizations, it's 20%, it's 100% it's 200%. So there's still probably like a lot further to go, honestly. Like you'll, you'll know that inference is pretty much solved when researchers start publishing about how they got 1% faster at something.Swyx [00:33:19]: Which by the way, because I am from the finance background, in the ‘70s, that was the margin at the time. When you did quantitative finance research, you would findAli [00:33:27]: And like 20%, tens of percent.Swyx [00:33:29]: That's. Yes.Philip [00:33:29]: Yeah.Swyx [00:33:30]: And now it'Philip [00:33:31]: Tiny fractionsSwyx [00:33:32]: For those people interested, look up Andrew Lo's paper. He had a really interesting illustration of quant, stat arb, distribution, narrowing down from like those kinds of 20% differences in the ‘70s, down to nothing today, which is very cool.Philip [00:33:48]: Exactly, and we're at the beginning of the same type of thing. Now benchmarking is hard. I think anyone will tell you that, and benchmarking provider speeds is hard because there's so many variables that go into it. What hardware are you using? How much load do you have on the system? What's the exact nature of the prompts and input and output sequence lengths? All that stuff. But overall, when you start stacking these improvements, you're looking at multiples. You can look at it. The most common form, of course, is TPS, tokens per second, which is bad naming by us in the industry, ‘cause there's two tokens per second. There's tokens per second, the throughput number, and the latency number.Ali [00:34:31]: TTMT, yeah.Philip [00:34:32]: Like total tokens per second out of the, out of the GPU as a throughput number. Most people only care about tokens per second as the latency number, which we should call ITL, intertoken latency, but we don't.Philip [00:34:44]: Anyway, so you can imagine a standard API without many optimizations for a 1 trillion parameter model operating somewhere in the 30 to 50 tokens per second range for reasonable traffic profile. And we generally see the goal of, pushing to 10X that. But, not necessarily day zero, but by stacking enough optimizations, if you have, say like four optimizations, each of which doubles performance. Or sorry, three optimizations, each of which doubles performance, then you stack that up, that's an 8X gain. That's the order of magnitude that we're working with in this space. We're trying to make things substantially faster, not just go from like 70 to 90.Swyx [00:35:38]: Are you saying you've. You have done that?Philip [00:35:40]: So let's say you have as a reasonable baseline, 30 or 40 tokens per second. You can achieve 10X that. So like on GLM-5.2, if you run it unquantized, perhaps on H100s even, and you're just using an off-the-shelf inference engine with no particular optimizations, no speculator, nothing extra around like KV routing, no disaggregation, you're, you're probably, yeah, looking at that like 30 to 40. You think that's like a reasonable baseline?Swyx [00:36:12]: Right. Right.Philip [00:36:12]: To get to something like 10X, there's a lot of trade-offs that you're making. If we're running at more like a 300, 400 tokens per second range, you are using the best hardware possible. You have a optimized speculator. You have done all of your quantization work. You are Seeing a pretty high cache hit rate. You are running with a reasonably small batch size and a parallelism configuration that is tuned for latency versus throughput, but it is possible. So the spreads that you see if you, like, go on artificial analysis or you go on OpenRouter and you look at, the worst provider to the best provider, oftentimes can hit that range. 10X is of course very aggressive. It's oftentimes maybe more of a four to six times improvement. But that's the performance that makes us really excited, is when we can get these huge gains, not just go from 70 to 90 tokens.Stacking Optimizations: NVFP4, Speculation, and DisaggregationAli [00:37:19]: It's also, like, hardware dependent. Like, ifPhilip [00:37:20]: YeahAli [00:37:20]: If you have a thing where you're serving it on just, like, a node of H100s and then you throw, like, you shard the model across, like, four nodes of B200s. Like, you can definitely increase the speed with just throwing more hardware at it. Like, normalizing for the same exact hardware and the same number of GPUs.Philip [00:37:35]: Yeah. Then you're looking at, like, a two to 4X improvementAli [00:37:38]: Right. RightPhilip [00:37:38]: Depending on the inference optimizations. So yeah, it's. Some of it's, what's the call, and some of it's who's the driver.Vibhu [00:37:46]: If you break down the two to 4X, say the example is run GLM-5.2Ali [00:37:51]: YeahVibhu [00:37:51]: On B200sAli [00:37:53]: YeahVibhu [00:37:53]: Single node, right? What's, like, the cost trade-off for effort to get, like, the last bit of juice out versus what should people just think of, right?Ali [00:38:01]: Spectre quantization. Yeah.Vibhu [00:38:03]: Spectre quantization.Ali [00:38:04]: That's, that's, that's like 95%. LikeVibhu [00:38:06]: And how far does that get you? And how easy is that for the average person to do? So say right I wanna throw the weights of GLM-5.2 on a node of B200s, how easy is it to find speculative decoder- decoder model or already quantized model? How much work goes into it?Philip [00:38:23]: If you're doing it up front, it's quite a lot of work. If you're doing it today, there's going to be people who have published things that you can just, you can just grab some NVFP4 weights. You can grab a speculator. Yeah, if we're thinking about, like, what are the 2Xs we're stacking, going from, BF16 to NVFP4 is, it's not quite a 2X, right? It's like. I think it's about, like, 30 to 40%, from 16 to 8, and then another 30 to 40% multiplied from, 8 to 4. So that doesn't quite get you a 2X, but, like, roughly a 2X. Speculator, roughly a 2X. Disagg on top of that if you're able to get enough hardware and put enough traffic through it, another roughly a 2X. And then you add in some, double-digit percent increase from having just a better runtime with, the latest kernels and stuff behind it. And that's how it stacks up.Ali [00:39:21]: YeahPhilip [00:39:21]: So building each of those, like, building the, quantized weights is, for someone who really knows what they're doing, hours to days of work. Building the speculator, again, like, hours to days of work. And the, disagg setup, hours to days. Well okay, but like once you haveAli [00:39:39]: Once set up. Once set up. YeahPhilip [00:39:40]: Yeah, getting disagg working for the first time, I'm saying, of course, is very difficult.Philip [00:39:44]: The marginal implementationAli [00:39:48]: Like, if you're just grabbing, like if you are a person, like just a normal consumer who has access to, like, a node of B200s and you're wondering, “How can I just host it myself?” You don't need to quantize the model yourself. There's always gonna be, like, an open source quantized checkpoint. NVIDIA's gonna push one out if no one else does. You. Usually, the providers will have their own spec dec that they've trained as well. You don't need to train your own spec dec. You can just use that as well.Philip [00:40:09]: Yeah. Like, GLM-5.2 has its own MTP.Ali [00:40:13]: Right. Right.Vibhu [00:40:14]: What's multi token prediction?Philip [00:40:15]: Yes.Ali [00:40:16]: I'm justVibhu [00:40:16]: Can you explain that?Ali [00:40:16]: I'm just an expert.Ali [00:40:18]: I can do it for you in case I get it wrong?Vibhu [00:40:20]: No.Vibhu [00:40:21]: Yeah, you should correct if we're wrong, but their multi-token prediction can be used for self-speculative decoding.Ali [00:40:27]: I'm not sure. I'm not gonna correct that.Vibhu [00:40:28]: Okay. I'm semi-confident in thatAli [00:40:30]: Okay. YeahVibhu [00:40:30]: But someone can check. But it's useful to paint the story of, okay, not just the average person, but say a company wants to switch from serverless inference I wanna throw this up on. I wanna rent some GPUs, throw it up. These are the steps you take to do significantly faster than just put it behind vLLM.Ali [00:40:48]: Right.Vibhu [00:40:49]: I was waiting for a mention of Dynamo.Vibhu [00:40:51]: I feel like, that's supposed to be the baseline that you measure against.Dynamo, KV Routing, and Disaggregation ToolkitsPhilip [00:40:55]: I would think of Dynamo as less of a box system and more of a toolkit for building with. So when we talk about doing aware routing, when we talk about doing KV offloading, when we talk about doing, PD disaggregation, Dynamo fundamentally is. By the way, Dynamo is an open source library from NVIDIA.Ali [00:41:17]: We've done a pod with KylePhilip [00:41:18]: OkayAli [00:41:19]: Kyle Cranin.Philip [00:41:19]: Cool. So then your listeners know then that it supports all the different inference frameworks. And it is multi hardware, which is interesting.Ali [00:41:28]: But it's just a router, it's not like an optimizer layer.Philip [00:41:30]: Yeah. All it does, like, what Dynamo is good at, it is a library for moving information around your cluster, around your hardware. So if you have, KV cache on one place and you need it to be somewhere else, Dynamo coordinates NIXL for you to move that around.Philip [00:41:49]: That doesn't mean that, like, out of the box, you just say, “Pip install Dynamo,” and then you get, like, a massive performance speed up. It's more of a developer toolkit.Ali [00:42:01]: Yeah. I would have said it would. It comes with a set of defaults that you can then swap out.Philip [00:42:06]: It does. If the industry at large, I think, was, like, rolling out all of these deployments, standard, then I think it would be, like, a credible baseline. But, we've got to, we've got to benchmark against, like, what we're seeing in the wild.Speculative Decoding Methods: Medusa, EAGLE, n-Gram, and Spec-SpecVibhu [00:42:23]: I did wanna talk a little bit more about PD disagg, because that is probably, like, number three after quantized and speculative decoding. In your book though, I was just gonna pull out the book.Philip [00:42:31]: Yeah.Vibhu [00:42:32]: Like section 522 on Medusa, 523 on EAGLEPhilip [00:42:35]: YeahVibhu [00:42:36]: 524 on gram.Philip [00:42:37]: It's 55, would be disaggregationAli [00:42:42]: Yeah. Well, no, I just wanted to dwell a little bitPhilip [00:42:44]: YeahAli [00:42:44]: The other. Like, so what do you choose to include? What do you choose to not to include? Because there was all these other techniques.Philip [00:42:51]: Yeah.Ali [00:42:51]: Are these still relevant? Because I think they came out, like, a year and a half ago maybe.Vibhu [00:42:55]: Medusa is quite old.Philip [00:42:56]: Yeah, Medusa's old.Ali [00:42:58]: It was old.Vibhu [00:42:58]: But is it in the book as a good, here'sPhilip [00:43:01]: BaselineVibhu [00:43:01]: Baseline vanilla understand it?Philip [00:43:02]: Like you should know this.Vibhu [00:43:03]: Like I read the paper, I'm like, “ it makes so much sense.”Philip [00:43:05]: Yeah.Philip [00:43:05]: So with the book, I had a couple goals. One was to give people just a working vocabulary for the space as a whole, and the other was to give them some intuition about how each of these techniques works. As I mentioned in my AI Engineer talk, which is the first public addendum to this, the speculation space has moved much faster than everything else. So yeah, even at the time that I wrote the book Medusa, I very much included as a way for people to understand how the space evolved rather than what the most modern technique is. And now of course, there's DFlash, dSpark. There's, there's newer techniques even than EAGLE, although EAGLE is still very commonly used.Ali [00:43:51]: SpecSpecta.Philip [00:43:52]: Yes. Speculative decoding.Vibhu [00:43:54]: What canAli [00:43:56]: Oh, it's a paper by Tri Dao and it's like, it's doing speculative decodingVibhu [00:44:00]: HuhAli [00:44:01]: For the speculative decoder.Philip [00:44:02]: Oh, in spec- oh my God.Ali [00:44:02]: It's literally just an another. It's like, yeah, that's the most simple way to explain it, and it seems like he got trivial speed ups there. But it seems that the complexity with training, it's almost like in our mind at least, it's almost as complex as training GANs. Like it's like a very delicate balance and oftentimes you, it's just but yeah, it's literally speculative decoding on speculative decoding.Vibhu [00:44:21]: Speculative.Ali [00:44:22]: Yeah. We saw this paper.Vibhu [00:44:24]: It's interesting, right?Ali [00:44:24]: Yeah.Vibhu [00:44:24]: I wouldn't even expect it to be very particular to train, I wouldAli [00:44:29]: Right.Vibhu [00:44:29]: The naive part of me is like, okay, train speculative decoder.Ali [00:44:32]: But like, and it makes sense, like the whole idea of speculative decoding is you. It's like, it's like almost like the iPhone auto predict version but for a normal model, right? Like you're just, you're just, generating three tokens and you're like, okay, I'll do prefill on them. And so you save those three turns for your original model. Now your speculative decoder is doing three turns of auto regression, so why not just have an even smaller model?Ali [00:44:53]: The other question there is what are the size of speculators? So say forPhilip [00:44:58]: Right. It's like a billion parameters.Ali [00:45:01]: Like for MiniMax, it's. Yeah. It's like one layer. It's like one 60th of the original model usually.Philip [00:45:06]: Yeah. I think we should do a paper when we get back to the office.Philip [00:45:10]: SpeculativeAli [00:45:11]: SpeculativePhilip [00:45:11]: Decoding.Ali [00:45:13]: No, it's, it does seem like how, when do you stop? But then it also seems like if you're able to train spec-spec decode for instance, right? Like if you're able to have a small model that is accurately predicts what the intermediate speculator is gonna predict, that is able to predict what the original target model's gonna predict, then why not just use that smallest model directly, right?Vibhu [00:45:34]: Yeah. This isAli [00:45:35]: Like it seems likeVibhu [00:45:35]: Adjacent to the routing problem.Ali [00:45:36]: Right.Vibhu [00:45:36]: Yeah.Ali [00:45:36]: Right.Philip [00:45:37]: The thing with speculators is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you're running the big model on. There is a orchestration and resource competition problem inherent in that, and that is one of the constraints on speculation in general, is that draft tokens cost resources to create and cost software complexity to manage. And so if you have like infinitely recursive speculators, you add in quite a bit of that complexity on the actual implementation within the inference engine as well, not just in the training process.Vibhu [00:46:17]: I was gonna say, I would wonder if you could do similar, like distillation and pruning of, it's the same thing, it's just a model. Can we not just distill a lot of the weights, quantize the speculator, out of my domain? The question that also comes up is, this is all for big server workloads, right? How much of this applies to, say I have this MacBook, I wanna run Gemma really efficiently. Similar problems, not the same?Local AI vs. Data Center InferencePhilip [00:46:45]: Pretty different. I talked to Selo, about this on his podcast a couple weeks ago. The difference between inference engineering for the data center and for production workloads versus inference engineering for local AI, is that we start with fundamentally like different constraints and different goals. With local AI, it's how do I fit this model onto my hardware and then make it less dumb? And with data center influence, it's how do I load this model and then make it less slow? And we care about less dumb, and they care about less slow. But the local AI inference engineering ecosystem, I think has a lot for us to learn from in the data center space. They are experts in various forms of quantization, including dynamic quantization that we just don't touch, in the pruning, in the distillation, in the, layer removal. There'Ali [00:47:42]: Layer removal matters less.Philip [00:47:43]: Yeah. There'Ali [00:47:44]: No one loves pruning really.Philip [00:47:45]: Yeah. Well, but the, but they doVibhu [00:47:46]: Which is surprising, right? But that's, that's a whole different thingPhilip [00:47:48]: Just to fit something on the laptop.Ali [00:47:50]: Right.Philip [00:47:50]: So yeah, it's a, it's an interesting, it's an interesting space. Not necessarily that like their techniques make sense for us to do in the data center, because we have different resources and different goals, but more that the process as well as the openness of that field is something to, admire.Ali [00:48:12]: Yeah. Like to your point, like, certain optimizations that would. Like for instance, Turbo Quantum Sharper, like it made such huge hype on that and we did like a whole deep dive on Twitter and like said, what is it? How does it work? Why is it good or not? And it took off and it was implemented on local devices because your memory bandwidth is so slow on like a MacBook, for instance. But try putting the same thing on like an NVIDIA GPU on a B200 Turbo quant would not be. Like, it would not be used. Like, NVIDIA - Like, NVIDIA made it clear that this is not a good optimization, and we've seen it firsthand where the overhead of doing dequantization, quantization of, in the kernel itself with turbo quant kernel, each end is much slower than the time that you save from doing the bandwidth. ‘Cause on the B200s, you have like 3.5 terabytes per second. You don't need decrease the storage that much. You don't need to do, FP4 KV cache. You don't need to use a requant. There's, there's, there's better optimizations to be made. But on Edge devices, it's extremely important, it's extremely useful. So, seems to be, like, different optimizations there, but then they're all uniquely combined with like all you wanna quantize the model, you wanna do speculative decoding, like certain common prefixes with bothPhilip [00:49:18]: Principles.Ali [00:49:19]: Yeah, exactly. Exactly. Exactly.Philip [00:49:20]: They also do a lot of work on, model parallelism, especially over, heterogeneous topology, where you have, some sparks and they are wired together with, Ethernet, DGX sparks.Ali [00:49:35]: Yeah, this is the Exo Labs guys.Philip [00:49:36]: Yeah. You have, a nu
In der heutigen Folge sprechen die Finanzjournalisten Lea Oetjen und Holger Zschäpitz über Freud und Leid bei Amazon und Apple, die enttäuschenden Zahlen von Coinbase und den historischen Fall von Adidas. Außerdem geht es um Microsoft, Roblox, Reddit, Monolithic Power Systems, AXT, Marvell Technology, Live Nation Entertainment, Western Union, Infineon, Siemens Energy, Siemens, Hochtief, Siltronic, Nemetschek, TeamViewer, SAP, IREN, Nebius, Sandisk, Micron Technology, SK Hynix, Bloom Energy, CoreWeave, Core Scientific, Nvidia, AMD, Broadcom, Adobe, Arm Holdings, Lam Research, Schneider Electric, Ahold Delhaize und Walmart. Mit dem Code „AAAFRIENDS“ sparst du 50 Prozent auf dein Ticket – aber nur unter folgendem Link. https://veranstaltung.businessinsider.de/event/financesummit26/summary?rp=c6dc55d6-6f4f-4fb4-b75f-3f3501d84859 Wir freuen uns an Feedback über aaa@welt.de. Noch mehr "Alles auf Aktien" findet Ihr bei WELTplus und Apple Podcasts – inklusive aller Artikel der Hosts. Hier bei WELT: https://www.welt.de/podcasts/alles-auf-aktien/plus247399208/Boersen-Podcast-AAA-Bonus-Folgen-Jede-Woche-noch-mehr-Antworten-auf-Eure-Boersen-Fragen.html. Hier könnt ihr den AAA-Newsletter abonnieren: https://www.welt.de/newsletter/article232797673/Alles-auf-Aktien-Der-taegliche-Boersen-Newsletter-fuer-WELTplus-Abonnenten.html Und – ganz neu: AAA gibt es jetzt auch auf Instagram: https://www.instagram.com/alles_auf_aktien/ Disclaimer: Die im Podcast besprochenen Aktien und Fonds stellen keine spezifischen Kauf- oder Anlage-Empfehlungen dar. Die Moderatoren und der Verlag haften nicht für etwaige Verluste, die aufgrund der Umsetzung der Gedanken oder Ideen entstehen. Hörtipps: Für alle, die noch mehr wissen wollen: Holger Zschäpitz können Sie jede Woche im Finanz- und Wirtschaftspodcast "Deffner&Zschäpitz" hören. +++ Werbung +++ Du möchtest mehr über unsere Werbepartner erfahren? Hier findest du alle Infos & Rabatte! https://linktr.ee/alles_auf_aktien Anzeige: Eight Sleep: Der Pod 5 reguliert die Temperatur im Bett automatisch, trackt Schlaf- und Gesundheitswerte ohne Wearable und kann so zu besserem Schlaf beitragen. Mit dem Code ALLESAUFAKTIEN erhaltet ihr auf https://www.eightsleep.com/allesaufaktien bis zu 350 Euro Rabatt. Impressum: https://www.welt.de/services/article7893735/Impressum.html Datenschutz: https://www.welt.de/services/article157550705/Datenschutzerklaerung-WELT-DIGITAL.html
Plus: Amazon reports higher revenue in the second quarter from its cloud-computing business. And a data-center developer working with Anthropic plans to borrow $15 billion for a new Texas campus, with backing from Google. Julie Chang hosts. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
What if your biggest CX innovation is also your biggest source of customer churn?Agility requires a clear-eyed view of both the promise and the peril of new technologies. It's the ability to embrace innovation like AI not just for efficiency, but with a rigorous focus on the customer outcomes that build long-term value.Today, we're going to talk about the delicate balance of implementing AI in the customer experience. Specifically, we'll explore the rise of sophisticated, agentic AI voice bots, and the critical "dealbreakers" that cause customers to abandon interactions and, potentially, your brand altogether.To help me discuss this topic, I'd like to welcome, Sushil Kumar, CEO at Cyara.About Sushil KumarSushil leads Cyara's strategy and growth with a vision to redefine how enterprises build trustworthy, AI-driven customer experiences. A builder at heart, he has spent his career creating platforms that change how software is engineered, validated, and delivered. His focus is helping global organizations achieve new levels of reliability, speed, and customer trust. Before joining Cyara, Sushil was the co-founder and CEO of RelicX.ai, a generative AI test automation pioneer acquired by Harness. He previously led major product and business organizations at Oracle, CA Technologies, and Broadcom, where he built and scaled category-defining AI, DevOps, and cloud solutions adopted by thousands of enterprises worldwide. With more than 25 years of experience leading high-growth teams and multi-hundred-million-dollar product lines, Sushil is known for blending deep technical insight with a pragmatic, product-first approach to leadership. His work centers on transforming how modern software and customer experiences are created, making them more resilient, more intelligent, and more human. Sushil holds a Bachelor of Technology from BIT Sindri and an MBA from the Indian Institute of Foreign Trade. He lives in the San Francisco Bay Area.Sushil Kumar on LinkedIn: https://www.linkedin.com/in/sushil-kumar-343780/---------- Resources ---------- Cyara: cyara.comThe Agile Brand podcast is brought to you by TEKsystems. Learn more here: https://aglbrnd.co/r/2868abd8085a9703We're proud to be a media partner for #MAICON26 - Oct. 13-15! Learn how AI can power your marketing and business and help you grow smarter. Use code AGILE150 to save! https://aglbrnd.co/r/7fe458ced0f04658Reach your customers with Reddit. Spend $500 in ad spend, get $500 back in ad credit! Learn more: https://advertalize.com/r/491818c79fb1873fChaser is the only Slack-native project management platform that helps teams turn messages into tracked tasks, automate follow-ups, and maintain team-wide visibility, without adopting another tool. Now integrated with Claude and other GenAI tools. Learn more at trychaser.com and use code AGILEBRAND for a 3-month free trial (normal trial is 14 days).The most influential minds in software, AI, and engineering leadership will be at WeAreDevelopers World Congress North America, September 23-25 in San Jose. Learn more: https://aglbrnd.co/r/60a7299222a7bcf1Start building your own apps with Replit and get $20 off. Learn more: https://aglbrnd.co/r/93531742a7625a20Enjoyed the show? Tell us more at and give us a rating so others can find the show at: https://aglbrnd.co/r/faaed112fc9887f3Connect with Greg on LinkedIn: https://www.linkedin.com/in/gregkihlstromDon't miss a thing: get the latest episodes, sign up for our newsletter and more: https://aglbrnd.co/r/35ded3ccfb6716baCheck out The Agile Brand Guide website with articles, insights, and Martechipedia, the wiki for marketing technology: https://www.agilebrandguide.comThe Agile Brand is produced by Missing Link—a Latina-owned strategy-driven, creatively fueled production co-op. From ideation to creation, they craft human connections through intelligent, engaging and informative content. https://www.missinglink.company Hosted on Acast. See acast.com/privacy for more information.
The Senate confirms Jay Clayton to lead ODNI. A new CISA framework highlights critical infrastructure isolation capabilities. OpenAI's rogue agent breached more than just Hugging Face. The average cost of a data breach continues to rise. Indirect prompt injection proves irresistible to cyber criminals. Broadcom patches multiple VMware products. ShinyHunters claims responsibility for Ernst & Young's recent breach. Our guest is Sean Zadig, CISO at Yahoo, discussing the impact of AI on the defense side. These aren't the droids you're looking for. Remember to leave us a 5-star rating and review in your favorite podcast app. Miss an episode? Sign-up for our daily intelligence roundup, Daily Briefing, and you'll never miss a beat. And be sure to follow CyberWire Daily on LinkedIn. CyberWire Guest Today we are joined by Sean Zadig, CISO at Yahoo, discussing the impact of AI on the defense side without all the hype in either direction, what is a CISO actually doing about it. Selected Reading Senate Confirms Jay Clayton to Lead U.S. Intelligence Community (The New York Times) China and Iran Are Already Inside US Grids: CISA Demands Tested Isolation Plans (Tech Times) OpenAI's rogue agent compromised a customer at a second tech firm, executive says (Reuters) Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident (Hugging Face) The Average Cost of a Data Breach Rises to $5 Million (Infosecurity Magazine) USSPACECOM Issues Space Warfighting Environment 2040 for Joint Force Space Operations (ExecutiveGov) THE SPACE WARFIGHTING ENVIRONMENT 2040 Framing the Future for the Joint Warfighter (U.S. Space Command) Notes from Underground: Adversarial Prompt Injection (Proofpoint) Critical VM Escape Vulnerability Patched in VMware ESXi (SecurityWeek) ShinyHunters Claims Ernst & Young Hack (SecurityWeek) America bans imported robots due to supply chain and security risks (The Register) Share your feedback. What do you think about CyberWire Daily? Please take a few minutes to share your thoughts with us by completing our brief listener survey. Thank you for helping us continue to improve our show. Want to hear your company in the show? N2K CyberWire helps you reach the industry's most influential leaders and operators, while building visibility, authority, and connectivity across the cybersecurity community. Learn more at sponsor.thecyberwire.com. The CyberWire is a production of N2K Networks, your source for strategic workforce intelligence. © N2K Networks, Inc.
Dario Amodei said Anthropic never backed an open-weights ban, pitching mandatory safety tests instead as OpenAI and Google signed on. Altman headed to Washington, Korea's KOSPI cratered 11% on AI jitters, Apple launched Klarna leasing, and shipped 194 CVE fixes. Anthropic wants tests, not bans, as OpenAI and Google back open weights (The New Stack) Source: Sam Altman will meet with senior US officials, lawmakers, and economists in Washington, DC, this week to preview OpenAI's upcoming family of AI models (CNBC) South Korea's KOSPI drops 11%+, led by chip stocks, amid concerns over China's chipmaking progress and the AI spending boom; Samsung falls 11%+ and SK Hynix 12% (Bloomberg) Credit default swap prices tied to Oracle, SpaceX, Alphabet, Amazon, Meta, Broadcom, and Nvidia hit record highs as investors turn jittery over Big Tech's data center debt; Oracle's five-year CDS reached 215bps (FT) Apple launches Apple Upgrade, a new US leasing program in partnership with Klarna that replaces the iPhone Upgrade Program, starting at $17.99/month for iPhones (MacRumors) Apple releases 26.6 updates for iOS, macOS, iPadOS, watchOS, tvOS, and visionOS with a huge number of security fixes; macOS Tahoe 26.6 alone addresses 155 CVEs (9to5Mac) Subscribe to the ad-free feed. Learn more about your ad choices. Visit megaphone.fm/adchoices
Was, wenn die wichtigste Frage nach einem Börsenhalbjahr nicht lautet "Welche Aktie lief am besten?", sondern "Welche Entscheidungen waren richtig?" Genau darum geht es in unserem vierten gemeinsamen Summer Special mit Finanzjournalist Clemens Faustenhammer: kein Ranking aus Gewinnern und Verlierern, sondern Prozess statt Prognose. Ich erzähle, warum ich trotz erwarteter Korrektur mein Depot konsequent umgebaut habe, wieso mein ETF-Portfolio mit rund 16,5 % im Plus die größte Überraschung des Halbjahres war und welche Rolle UnitedHealth, CVS Health und Fastenal dabei gespielt haben.Clemens teilt seine eigene Sicht auf ein Halbjahr voller Überraschungen, wir sprechen über Zinseszins und Haltestrategien, den Wandel vom klassischen Finanzblog zu Substack und darüber, wie Profi-Investoren wie Terry Smith ihre Strategie anpassen. ⏱️ KAPITEL0:00 Begrüßung und Einleitung zum Summer Special3:15 IBM & die Marktkorrektur im Juli 20266:45 Depotqualität statt Markttiming10:20 Zinseszins-Strategie: Clemens über Haltedauer14:50 (Teil-)Verkauf von AT&T und Stanley Black & Decker19:30 KI als Werkzeug für die Aktienanalyse24:10 Finanzblogs vs. Substack: Wandel im Finanzjournalismus29:00 Terry Smith & der Strategiewechsel bei Profi-Investoren34:15 SaaS-Aktien im KI-Hype: Zukunftsaussichten39:00 Abspaltungen und M&A: Aktuelle Trends43:45 Einzelaktien vs. Themen-ETFs im Performance-Check48:30 Südkorea-Volatilität und Samsung-Spekulationen53:20 Zoetis: Fehleranalyse einer schwierigen Position58:00 US-Konsumschwäche: Tractor Supply im Fokus1:02:45 Erfolge mit Corning, Broadcom, Cisco und Fastenal1:06:50 Ausblick: Unser Fokus fürs zweite Halbjahr 20261:09:30 Fazit: Zufriedenheit statt Markt schlagen1:11:12 Outro
The episode reveals infrastructure dependence and vendor consolidation risks in the IT channel, illustrated by Broadcom's abrupt closure of the VMware Cloud Service Provider (VCSP) program. This move eliminated license access for numerous MSPs, disrupting established practices reliant on VMware platforms and forcing providers into accelerated, unplanned migrations. The event highlights the vulnerability of service provider business models when built on external vendor programs without autonomy or long-term contractual assurance. The most consequential development discussed is the forced transition experienced by Valor C3 Data Centers after Broadcom shut down the VCSP program on October 31, 2025. According to Justin Fox, this action imposed a non-negotiable deadline and provided no grandfathering, causing hundreds of MSPs to lose access to essential licenses. Valor allocated several hundred hours to research and migration planning, citing costs between $200,000 and $300,000 per site for new landing zone infrastructure, not including increased hardware prices driven by AI market demand. Decisions centered on reducing repeat vendor risk, balancing reuse of existing hardware, and evaluating alternatives such as full open source OpenStack via Platform9. Supporting developments point to broader changes in the virtualization market post-Broadcom. Justin Fox noted that, while Proxmox and Hyper-V are common destinations for displaced VMware users (especially in small, single-tenant environments), larger service providers prioritize native multi-tenancy, platform flexibility, and hardware independence—criteria that led Valor to OpenStack. The importance of ecosystem compatibility, operational simplicity, and readiness to pivot away from vendor-managed solutions was elevated against the background of supply chain disruptions and rising hardware costs. For MSPs and IT leaders, these circumstances clarify the need for robust vendor risk assessment and contingency infrastructure strategies. Reliance on proprietary vendor programs presents exposure to sudden policy changes, price escalations, and contract terminations. Transitioning to open platforms can reduce repeat risks, but does not eliminate dependency—especially when managed open source solutions have their own governance and continuity considerations. Clear communication with customers, careful management of migration costs, and ongoing evaluation of vendor relationships are required to avoid operational shocks and revenue disruption in an increasingly consolidated channel environment. Supported by: Pax8 ScalePad
2,5% Zinsen p.a. auf ein unbegrenztes Guthaben mit bis zu fünfmal der gesetzlichen Einlagensicherung*. Auch für Kinder. Das gibt's bei Scalable Capital. Mehr Infos hier. SAP beste große Tech-Aktie. Intel und Oracle im Minus. NVIDIA kauft Speicherchips für 500 Mrd. $. Broadcom für 200. SK Hynix und Samsung freut's. VW senkt Prognose. Verizon macht KI. Japan reguliert Pokémon. Canadian National Railway zufrieden. General Motors (WKN: A1C9CM) hat pro Aktie das gewinnstärkste Halbjahr der Firmengeschichte. Weniger Autos verkauft, trotzdem mehr verdient. Preisdisziplin, schrumpfende E-Auto-Verluste und ein wachsendes Abo-Business machen's möglich. US-Rüstung boomt. RTX (WKN: A2PZ0R) und Lockheed Martin (WKN: 894648) heben Prognosen an und sitzen auf über 500 Mrd. $ Auftragsbestand. Startup Anduril strebt 100 Mrd. $ Bewertung an. Rüstungskonzerne pumpen Rekordsummen in Startups. Diesen Podcast vom 27.07.2026, 3:00 Uhr stellt dir die Podstars GmbH (Noah Leidinger) zur Verfügung. *Veränderlicher Zins auf unbegrenztes Guthaben. Konditionen sowie Guthabenverteilung auf scalable.capital/tagesgeld. Learn more about your ad choices. Visit megaphone.fm/adchoices
En el episodio de hoy Valentina Orduz y Juan Manuel de los Reyes analizaron el debut bursátil de la empresa de memoria china CXMT, que se disparó cerca de 460% , y la firma del memorando de entendimiento entre Samsung y Broadcom por más de USD $200,000 millones. También abordaron el giro geopolítico: la caída del petróleo tras la señal de pausa de Irán, la expansión del conflicto en Medio Oriente y el papel del Estrecho de Ormuz. Cerraron con la reunión de la Fed del 28 y 29 de julio, un mercado que ya no espera recortes sino alzas de tasas, y las elecciones legislativas de noviembre.
Spring AI 2.0 released in June 2026 and is one of several projects ensuring Java's continued strength in enterprise AI development. Here are some tips on deterministic agents, MCP servers and skills, prompt engineering, and where Java's richness for AI solutions stems from. In this "Input/Output" episode of the Inside Java Podcast, recorded during JavaOne 2026, Lize Raes talks to Dan Vega, Spring Developer Advocate at Broadcom, Java Champion, speaker, and author. "Input/Output" is our new show, where we talk to people outside of OpenJDK to bring you their perspectives and insights into what's happening in the Java ecosystem.
AI is no longer just a race to train smarter models. As AI moves into production, the bottleneck is increasingly inference: how fast models can generate tokens, use tools, reason, verify, and act. In this episode of the MAD Podcast, Matt Turck sits down with Andrew Feldman, co-founder and CEO of Cerebras, to explain why fast inference may define the next era of AI.Cerebras is known for building a chip the size of a silicon wafer. But this conversation is not just about one company or one chip. It is a deep dive into the AI infrastructure stack: GPUs, ASICs, memory, HBM, SRAM, data centers, power, TSMC, AWS, OpenAI, agents, reasoning models, and why speed changes what AI products can become. Andrew explains why “tokens per second per user” matters, why generating a single word can require moving the equivalent of 100 HD movies through memory, why agents amplify latency, why GPUs struggle with certain inference workloads, and why fast AI may eventually reshape SaaS itself.This is a reference conversation on fast inference, AI chips, and the next compute bottleneck.(00:00) Cold open & Intro(01:31) Why speed became the AI bottleneck(02:32) Tokens per second per user, explained(03:16) AI's broadband moment and the Netflix analogy(04:35) The AI chip landscape: GPUs, TPUs, Trainium, ASICs(06:36) What is an ASIC?(08:08) Nvidia, Groq, and the fast inference war(09:16) OpenAI, Broadcom, and specialized silicon(12:10) China, power, and sovereign AI infrastructure(15:05) Is the AI infrastructure boom a bubble?(18:56) The hidden bottlenecks: HBM, CoWoS, and 3nm(22:57) Why agents are creating CPU demand(25:36) Andrew Feldman's path from SeaMicro to Cerebras(26:13) Why Cerebras bet on AI in 2016(31:14) SRAM vs. HBM: why inference is a memory problem(33:19) What wafer-scale computing actually means(34:28) The deep-tech “Everest” problem(36:07) The moment the first Cerebras system worked(36:49) Ringing the bell and surviving deep tech(39:08) How a giant chip handles failure(41:22) Why GPUs struggle with decode(42:17) Prefill vs. decode explained(44:01) The “100 HD movies” problem in AI inference(45:04) How fast inference changes RL and training(48:08) Reasoning models and why they cost more compute(50:08) Verification, guardrails, and small models checking big models(52:37) Multimodal AI and the path to video(53:51) Cerebras' business model: hardware, cloud, and API(55:14) OpenAI's 750MW inference deal(55:36) Why data centers are measured in megawatts(58:01) AWS Trainium + Cerebras decode(59:29) Fast tokens as a cloud product(01:00:52) Is CUDA still a moat?(01:03:53) How TSMC helped Cerebras build the giant chip(01:07:41) Why nobody cared in 2020(01:08:15) Why chip supply chains are hard to diversify(01:09:54) Why today's AI models will be the worst you ever use(01:10:38) What fast AI could do to SaaS
Take a Network Break! In this week’s episode our red alert highlights two critical vulnerabilities in RabbitMQ, and we dig into listener follow-up about data centers in space. This week’s news coverage considers Apple’s $30 billion spending pledge to Broadcom, TSCM promising another $100 billion to build even more chip fabs in the US, and... Read more »
Take a Network Break! In this week’s episode our red alert highlights two critical vulnerabilities in RabbitMQ, and we dig into listener follow-up about data centers in space. This week’s news coverage considers Apple’s $30 billion spending pledge to Broadcom, TSCM promising another $100 billion to build even more chip fabs in the US, and... Read more »
Take a Network Break! In this week’s episode our red alert highlights two critical vulnerabilities in RabbitMQ, and we dig into listener follow-up about data centers in space. This week’s news coverage considers Apple’s $30 billion spending pledge to Broadcom, TSCM promising another $100 billion to build even more chip fabs in the US, and... Read more »
Is Cerebras Systems the next great AI chip stock or a red-hot IPO priced for perfection? In this episode of 7investing Live, Simon Erickson and executive producer Heather Horton welcome back Nick Rossolillo, co-founder of Chip Stock Investor, to break down three of the market's biggest stories.First up: Cerebras Systems (NASDAQ:CBRS), the wafer-scale chip maker that just IPO'd at a $40+ billion market cap. With 44GB of SRAM embedded directly on the chip, Cerebras was purpose-built to solve AI's "memory wall" problem for inference workloads. Now it's reportedly landed a ~$10 billion order from OpenAI and a deal with Amazon Web Services that could top $20 billion. Simon and Nick dig into whether these massive orders are real, how Cerebras stacks up against NVIDIA's GPUs and hyperscaler custom silicon, the TSMC capacity bottleneck that could throttle its growth, and how to value a company trading near 20x sales without profits.Then the conversation turns to Rocket Lab (NASDAQ:RKLB), which has pulled back from $150 to around $70 per share. Simon shares the latest iteration of his discounted cash flow valuation, and the duo debates the proposed Iridium acquisition — a deal that could pull Rocket Lab to EBITDA-positive on a pro forma basis — plus what the long-awaited Neutron rocket launch means for the company's future.Finally: Netflix (NASDAQ:NFLX). After another quarter of decelerating revenue guidance, is the streaming giant now a value stock rather than a growth stock? Nick explains why the advertising business hasn't reaccelerated growth the way he expected, and what he'd need to see before buying the dip.Plus: Nick's take on the recent chip stock sell-off across NVIDIA, AMD, Broadcom, SanDisk, and Kioxia and why "stocks go up, stocks go down" might be the healthiest way to think about it.Subscribe for more deep dives on AI infrastructure, semiconductors, and innovative growth stocks!Start your FREE 7-day trial of 7investing: https://www.7investing.com/subscribeFollow Nick and Casey Rossolillo at Chip Stock Investor: https://chipstockinvestor.comRocket Lab Deep Dive videos mentionedPart 1 https://youtu.be/AMDd0-JKUH0 (Deep Dive)Part 2: https://youtu.be/Z76xTGFNwBA (Valuation)Companies MentionedPublicly Traded:Cerebras Systems (NASDAQ:CBRS)Rocket Lab (NASDAQ:RKLB)Netflix (NASDAQ:NFLX)NVIDIA (NASDAQ:NVDA)Advanced Micro Devices (NASDAQ:AMD)Broadcom (NASDAQ:AVGO)Micron Technology (NASDAQ:MU)Taiwan Semiconductor Manufacturing (NYSE:TSM)Amazon (NASDAQ:AMZN)Alphabet (NASDAQ:GOOGL)Meta Platforms (NASDAQ:META)Iridium Communications (NASDAQ:IRDM)SanDisk (NASDAQ:SNDK)Kioxia Holdings (TSE:285A)Globalstar (NASDAQ:GSAT)SpaceX (NASDAQ: SPCX)Private / Pre-IPO:OpenAIAnthropicVideos Mentioned:https://www.youtube.com/watch?v=Z76xTGFNwBA&t=3shttps://www.youtube.com/watch?v=AMDd0-JKUH0&t=987sHere's the shifted chapter list, with all timestamps moved back 55 seconds:0:00 Welcome to 7investing Live0:54 Cerebras Systems: IPO recap & the Wafer-Scale Engine2:31 Is NVIDIA even the right comparison for Cerebras?5:38 The memory wall: why bigger AI models need new chips8:52 Latency vs. throughput — and the new AI alliances10:46 Are the $10B OpenAI & $20B Amazon orders real?14:02 Cerebras risks: how do you value a hot IPO?17:27 The TSMC capacity bottleneck20:01 Heather's take on Cerebras20:41 Rocket Lab: the sell-off & Iridium acquisition24:34 Simon's DCF valuation & price target for RKLB29:05 Why Neutron changes everything30:12 Q&A: Does Peter Beck carry an "Elon premium"?31:36 Netflix: buying opportunity or cheap for a reason?36:57 Q&A: Is Netflix a growth stock or a value stock?39:03 Chip stocks selling off: normal volatility or a warning?42:57 Wrap-up & final thoughts#7investing #Simonerickson #Cerebras #CBRS #NVIDIA #AIinvesting #semiconductors #chipstocks #RocketLab #RKLB #Netflix #NFLX #AIinference #stocks #investing #stockmarket #TSMC #AIdatacenters
Story of the Week (DR):Social Media's 'Big Tobacco Moment': Meta Faces $1.4 Trillion Fine for Allegedly Fueling Teen Suicide and AddictionMeta is being sued by 33 US states, led by California, Colorado, Kentucky, and New Jersey.12 blue/12 red/9 purpleThe states allege that Meta deliberately designed Facebook and Instagram to be addictive to children and teens, fueling a youth mental health crisis (including anxiety, depression, self-harm, and suicide).They also accuse Meta of violating child privacy laws by collecting data from children under 13 without parental consent.Meta warned a federal court that it could face up to $1.4T in penalties if the states prevail at the upcoming trial (set for August 18, 2026). Meta extrapolated this massive figure—which is roughly equivalent to the company's entire stock market value—based on the methodology proposed by the lead states for calculating damages.Meta calls the penalty "outlandish" and "unsubstantiated," arguing it has no precedent in consumer protection history: 'A sanction of that size has no analog in the history of consumer protection enforcement.' The company accuses the states of improperly multiplying penalties (e.g., stacking fines based on daily usage time). Meta denies the allegations, asserting its platforms have extensive safety tools and that the claims are unmoored from actual unfair practices.‘I Don't Think I'm Ever Going to Stop,' Says Mark Zuckerberg. Even With 'Infinite Money,' He Has No Plans to Retreat to His Massive Hawaii EstateMeta found to breach EU laws with 'addictive' Instagram, Facebook designsInstagram and Facebook's “addictive” designs have put Meta in breach of the European Union's digital laws, the EU concluded Friday in a preliminary report.The tech giant violated the EU's Digital Services Act by failing to adequately consider the risks associated with design features that affected the physical well-being of its users, including minors and vulnerable adults, the European Commission said.These features include infinite scroll, which constantly shows fresh content, autoplay, push notifications and highly personalized recommendation systems — feeding users' compulsion to continue using platforms and putting them into “autopilot mode.”The EU Commission also accused Meta of ignoring available information about how much time young people are spending on Instagram or Facebook at night, and how different types of content formats, from reels to stories, could lead to excessive use of its services.Meta said, “We disagree with these preliminary findings.”New Zealand Moves To Ban Climate Change Litigation. Will The U.S. Follow?New Zealand has proposed a bill to limit the ability of individuals to sue high greenhouse gas emitters over the impacts of climate change, relying instead on the enforcement measures taken by the government. The bill appears poised to pass.The women who wouldn't let climate data disappear MMAfter losing their jobs at National Oceanic and Atmospheric Administration (NOAA), Rebecca Lindsey, her sister Mary and colleague Anna Eshelman teamed up to rebuild a pivotal resource the Trump administration took offlineRebecca Lindsey, a technical writer for NASA–one of 280,000 federal workers fired by Musk/Trump, joined forces with former NOAA employees Anna Eshelman, and Mary Lindsey, her older sister, to become the core team behind the deactivated site's successor, Climate.us, preserving over 15 years of key climate data and resources.Elon Musk says he always wanted his SpaceX employees to get rich — and now thousands of them are millionairesElon Musk's 'Chainsaw for Bureaucracy' Just Left an $11 Billion Budget Hole as Trump Rehires StaffThe trove features key maps, educational materials and climate indicator reports, including the now-deleted Fifth National Climate Assessment, the government's most comprehensive analysis of climate change that was at risk of being lost to the publicJersey Mike's $12 billion IPO filing reveals a $50 million payday for the founder's stepson and a $41 million jetFamily members of founder Peter Cancro were employed Jersey Mike's in various roles and received compensation in excess of $120,000 from the Company as follows for the years ended December 28, 2025 and December 31, 2024 and 2023:John Cancro, Mr. Cancro's brother, received total compensation of approximately $20,019,231, $519,231 and $500,000, respectively;Paul J. Cancro, Mr. Cancro's son, received total compensation of approximately $8,001, $216,022 and $208,023, respectively;Robert Cancro, Mr. Cancro's son, received total compensation of approximately $38,462, $1,038,462 and $1,000,000, respectively;Tatiana Cancro, Mr. Cancro's wife, received total compensation of approximately $11,538, $311,539 and $300,000, respectively;Caroline Jones, Mr. Cancro's daughter, received total compensation of approximately $38,462, $1,038,462 and $1,000,000, respectively;Alexandra Powers, Mr. Cancro's sister-in-law, received total compensation of approximately $0, $1,165,437 and $0;Daniel Powers, Mr. Cancro's brother-in-law, received total compensation of approximately $30,213,462, $1,793,952 and $0;John Tesauro Jr., Mr. Cancro's brother-in-law, received total compensation of approximately $0, $0 and $2,429,628, respectively;Phillip Sivolobov, Mr. Cancro's stepson, received total compensation of approximately $50,011,538, $311,539 and $276,923, respectively.GRAND TOTAL: $112MStepson Phillip got $51M, brother John got $21MOther Peter Cancro schwag in 2025:Got a $41 million jet and an additional fixed amount of $166,666.66 per month in light of the business expenses incurred by Mr. Cancro related to air transportation to travel from time to time for business purposes.Lease agreements valued at $1M in rent (leases go to 2030) Controlled company: Blackstone (will control more than 50% of voting power)Board: 8 directorsBlackstone:Chair Nigel Travis (also chair of Abercrombie & Fitch)David N. KestnbaumDevon L. RinkerMichael J. StaubFounder/former CEO/chair: Peter CancroCEO Charles R. MorrisonCheryl S. Miller, director on two controlled companies:Tyson FoodsOld Dominion Freight Line (Congdon brothers)Fran Horowitz, CEO of Abercrombie & Fitch (where chair serves as chair)Goodliest of the Week (MM/DR):DR: Amazon, Walmart and Other Large Employers Could Face New Costs As New Jersey Targets Companies With Medicaid Workers— Will Other States Follow?DR: UBS says rich people will be younger, female and openly queer thanks to the Great Wealth TransferMM: Meta Buried Research Linking Instagram To Teen Harm While Facing $1.4 Trillion Penalty That Could Erase Its Entire WorthMM: ESG!Madison Square Garden Kept a List of Gay CelebritiesAn internal Madison Square Garden database of VIPs labels Joe a “medium risk,” one of roughly 400 celebrities given a risk score.If you're a celebrity and you're marked with a risk score—even as a low risk—it means “you've done something in the publicity world, the social media world, that has caught the attention of the wrong people,” the source continues.The talent database also tracks some celebrities' race, gender identity, and sexual orientation; 93 entries are marked as “LGBTQIA.”MM: California vs. Elon Musk: Tesla Snubbed as New EV Incentives Boost Rivian, Lucid MM DRAssholiest of the Week (MM):Billionaire amplificationKen Griffin says everyone is misinterpreting the AI revolution — and wishes Zohran and Bernie would ‘read a damn history book for once'“[Capitalism is] the greatest success story in the history of humanity,” Griffin said, urging the self-identified socialist politicians, “whether it's Bernie Sanders, whether it's Mamdani,” to “read a damn history book for once and then tell us how to run our country.”Jeff Bezos 'Made All of Our Lives So Much Better,' Says Billionaire Investor Tim Draper"Amazon has made all of our lives so much better," Draper said.Draper said he has benefited from what Bezos has done, and that's a part of the world economy that isn't spoken about enough."Those geniuses who create this incredible world for us are benefiting all of us."40 Epstein-Tied Billionaires Have Injected $1.6B Into US Elections, Report FindsThese are the millionaires and billionaires pledging to fund Trump accountsZuckMeta AI Data Centre Contractor Triggers Biohazard Scare After Flushing Rare Bacterium Into Public SewersMeta Platforms To Build $9 Billion A.I. Data Centre In CanadaThe $145 Billion Lie? Zuckerberg's Leaked Town Hall Audio Exposes Massive AI Failures After Mass LayoffsMeta jumps into AI coding market in effort to chase Anthropic and OpenAIWhat person in their right mind would trust Zuckerberg with their coding?Hollywood Vs Zuckerberg: CAA Warns Meta's AI Image Tool Needs A Major Privacy Overhaul‘I Don't Think I'm Ever Going to Stop,' Says Mark Zuckerberg. Even With 'Infinite Money,' He Has No Plans to Retreat to His Massive Hawaii EstateJeff Sonnenfeld - DRIn defense of Musk, SpaceX, and dual class shares“the rigid formulas of proxy advisors create a perverse, socially destructive incentive: hoard your wealth like an oligarch to maintain your good governance rating, or give it away and risk losing your company”“The proxy advisors want a world governed by rigid mathematical formulas because auditing a checklist is easy. Evaluating human character, industry dynamics, track records of success and failure, and the capacity for visionary leadership is hard. But it is exactly that hard work of judgment which is vital. When it comes to dual-class shares, it is time for the critics to step out of the theoretical vacuum and look at the real-world scoreboards”Headliniest of the WeekDR: OpenAI Wants a $1 Trillion Valuation. But College Students Are Testing At The Level Of 10-Year-Olds AND Suspecting AI cheating, Ivy League prof ordered an in-person final; scores fell 50%MM: ‘Waymo Takes Revenge, Dropping Drunk Teens Directly Into Squad of CopsMM: West Virginia spent $3M to create university program to fight ‘woke ideology.' One student is enrolledWho Won the Week?DR: Free Float's new platformAnd blowhard Jeffrey Sonnenfeld for arguing that dualclass owners are necessary because the dualclass mechanism allows them to sell shares and maintain voting power so they can “cure diseases, endow universities, and combat poverty” MM: Free Float data: Elon Musk and the age of the corporate leviathanAbove a certain size the ordinary rules of governance apparently cease to apply.Of the 16 listed firms worth more than $1trn, seven are shareholder fundamentalists.None has an elaborate statement of corporate purpose, since they are mostly content making heaps of moneyFree Float says Nvidia, Amazon, Broadcom, Micron are all TOTALITARIANNext are the corporate paternalists, who believe that the problem with shareholder democracy is that its voters do not know what is best for them.Free Float says Berkshire, Google, Meta are all TOTALITARIANThe final clan presents the strongest argument against the end of corporate history: individual shareholders consent to hand over all of their rights to the world's richest man, who then governs as he sees fit.Free Float says SpaceX, Tesla are TOTALITARIANBasically, it's nice of the economist to recognize everything Free Float says every weekPredictionsDR: Jeff Sonnenfeld writes something on Fortune that triggers meMM: Jeff Sonnenfeld writes something on Fortune that triggers Damion
- Details on Extended Partnership Between Apple and Broadcom - Report: Apple Testing CXMT Chips for Devices Sold in China - Evercore Sees Apple/Broadcom Extension as "Strategically Positive" - EU Court: Core App Store Platform Makes Apple a Gatekeeper - Apple TV Nabs 87 Emmy Nominations - Sponsored by Helix Sleep: Get 20% off Sitewide, 25% off Luxe Mattresses, and 30% off Elite Mattresses at helixsleep.com/macosken - Sponsored by Notion: Learn more about Notion's Developer Platform today at notion.com/macosken - Catch Ken on Mastodon - @macosken@mastodon.social - Send Ken an email: info@macosken.com - Chat with us on Patreon for as little as $1 a month. Support the show at Patreon.com/macosken
Benjamin and Chance talk about all the changes in iOS 27 beta 3, and Mayo's first experience with Siri AI on Apple Watch. Also, we dive in to the details of what's new in macOS Golden Gate, Apple tries to lobby to use Chinese memory components, signs a new deal with Broadcom, and likely supply constraints for the iPhone Fold. And in Happy Hour Plus, with Apple TV scoring a record number of Emmy nominations, we talk about what we've been watching recently across TV and film. Subscribe at 9to5mac.com/join. Sponsored by Copilot Money: Get two months free with code 9TO5MAC at copilot.money/9to5mac. Sponsored by Shopify: All you need is the idea. Shopify handles the rest. Start your free trial at shopify.com/happyhour. Sponsored by Framer: The only free design tool that brings your ideas to the web. Visit framer.com/happyhour for 30% off a Framer Pro annual plan.
The aftermath of Apple's price hikes. Apple is taking its fight against Epic to the Supreme Court. More AI features coming to Apple's Creature Studio. And Apple is showing confidence in its upcoming iPhone Fold, expecting to sell 10 million units! America is having MacBook sticker shock. MacBook price hikes expected to contribute to 13.6% drop in global laptop shipments. Apple weighs buying RAM from two blacklisted Chinese suppliers to curb rising costs. Broadcom and Apple extend custom silicon pact to 2031. iPhone 18 Pro leaks: Qualcomm or Apple C2 model, A20 details, camera upgrades. Apple takes Epic fight over app store fees to the Supreme Court. Tim Cook's government liaison position comes into focus before stepping down as Apple CEO. Apple in Russia's crosshairs again, facing $52M fine for not installing state-required apps. If you wanted more AI in Apple's Creator Studio, Tuesday's update gives it to you. Safari's new MCP server lets coding agents inspect and debug websites. Siri AI can pull info from third-party apps in the latest developer beta. Apple to launch 5 new iPhone models to gain market share amid memory crunch. Confident Apple increases its iPhone Fold orders to 10 million. Apple TV teases major new sci-fi series: Neuromancer. iPhone 17 Pro Max buried in America's 250th anniversary time capsule: to be opened in 2276. Picks of the Week Glenn's Pick: Turtles! Jason's Pick: Default Folder X Andy's Pick: Readest Hosts: Leo Laporte, Andy Ihnatko, and Jason Snell Guest: Glenn Fleishman Download or subscribe to MacBreak Weekly at https://twit.tv/shows/macbreak-weekly. Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Sponsors: blackhat.com/us-26 and use code TWIT joindeleteme.com/twit promo code TWIT
Plus: Apple will spend $30 billion on U.S-made chips from Broadcom. And a European court dismisses the iPhone maker's appeal over the EU's tech rules. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Oil prices skyrocketing after President Trump strikes Iran overnight and announces the U.S.-Iran ceasefire is effectively over. Where the war is heading, and whether investors can expect oil to keep climbing. Then, the latest out of the Fed as the open market committee's minutes report comes out. Head of investment strategy at Bryn Mawr Trust Andrew Davis gives his take on the new Fed communication framework, and why he's expecting financials and health care to outperform come earnings season next week. Plus, Nvidia falling behind, Apple taking a $30 billion bite out of Broadcom, and the latest read on Auto sales. Fast Money Disclaimer Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Carl Quintanilla, Jim Cramer and David Faber discussed market reaction to President Trump's comments at the NATO Summit in Turkey, where he said the ceasefire with Iran is "over." That news sent oil prices spiking, yields jumping and stocks falling sharply. What should you make of the volatility? Also in focus: Trillion-dollar tech's bear club, Apple's $30 billion chip deal with Broadcom, Amazon bond sale update, Blue Origin's fundraising push, Cramer on Walmart as a buying opportunity, Samsung and SK Hynix drag South Korea's KOSPI into bear territory. Squawk on the Street Disclaimer Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Council on Foreign Relations President Emeritus Richard Haass joins to discuss his outlook for the Iran War after President Trump declared that the ceasefire is over. Plus, as oil prices spike, Rapidan Energy Group CEO Scott Modell discusses why he feels like we're at a juncture here. We also break down the details of Apple expanding its partnership with Broadcom in its largest commitment to U.S. manufacturing to date. Squawk on the Street Disclaimer Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
The aftermath of Apple's price hikes. Apple is taking its fight against Epic to the Supreme Court. More AI features coming to Apple's Creature Studio. And Apple is showing confidence in its upcoming iPhone Fold, expecting to sell 10 million units! America is having MacBook sticker shock. MacBook price hikes expected to contribute to 13.6% drop in global laptop shipments. Apple weighs buying RAM from two blacklisted Chinese suppliers to curb rising costs. Broadcom and Apple extend custom silicon pact to 2031. iPhone 18 Pro leaks: Qualcomm or Apple C2 model, A20 details, camera upgrades. Apple takes Epic fight over app store fees to the Supreme Court. Tim Cook's government liaison position comes into focus before stepping down as Apple CEO. Apple in Russia's crosshairs again, facing $52M fine for not installing state-required apps. If you wanted more AI in Apple's Creator Studio, Tuesday's update gives it to you. Safari's new MCP server lets coding agents inspect and debug websites. Siri AI can pull info from third-party apps in the latest developer beta. Apple to launch 5 new iPhone models to gain market share amid memory crunch. Confident Apple increases its iPhone Fold orders to 10 million. Apple TV teases major new sci-fi series: Neuromancer. iPhone 17 Pro Max buried in America's 250th anniversary time capsule: to be opened in 2276. Picks of the Week Glenn's Pick: Turtles! Jason's Pick: Default Folder X Andy's Pick: Readest Hosts: Leo Laporte, Andy Ihnatko, and Jason Snell Guest: Glenn Fleishman Download or subscribe to MacBreak Weekly at https://twit.tv/shows/macbreak-weekly. Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Sponsors: blackhat.com/us-26 and use code TWIT joindeleteme.com/twit promo code TWIT
The aftermath of Apple's price hikes. Apple is taking its fight against Epic to the Supreme Court. More AI features coming to Apple's Creature Studio. And Apple is showing confidence in its upcoming iPhone Fold, expecting to sell 10 million units! America is having MacBook sticker shock. MacBook price hikes expected to contribute to 13.6% drop in global laptop shipments. Apple weighs buying RAM from two blacklisted Chinese suppliers to curb rising costs. Broadcom and Apple extend custom silicon pact to 2031. iPhone 18 Pro leaks: Qualcomm or Apple C2 model, A20 details, camera upgrades. Apple takes Epic fight over app store fees to the Supreme Court. Tim Cook's government liaison position comes into focus before stepping down as Apple CEO. Apple in Russia's crosshairs again, facing $52M fine for not installing state-required apps. If you wanted more AI in Apple's Creator Studio, Tuesday's update gives it to you. Safari's new MCP server lets coding agents inspect and debug websites. Siri AI can pull info from third-party apps in the latest developer beta. Apple to launch 5 new iPhone models to gain market share amid memory crunch. Confident Apple increases its iPhone Fold orders to 10 million. Apple TV teases major new sci-fi series: Neuromancer. iPhone 17 Pro Max buried in America's 250th anniversary time capsule: to be opened in 2276. Picks of the Week Glenn's Pick: Turtles! Jason's Pick: Default Folder X Andy's Pick: Readest Hosts: Leo Laporte, Andy Ihnatko, and Jason Snell Guest: Glenn Fleishman Download or subscribe to MacBreak Weekly at https://twit.tv/shows/macbreak-weekly. Join Club TWiT for Ad-Free Podcasts! Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit Sponsors: blackhat.com/us-26 and use code TWIT joindeleteme.com/twit promo code TWIT
- Judge Agrees to Let Apple and Epic Argue for Pause Ahead of SCOTUS Hearing - Apple and Broadcom Extend Buyer/Supplier Partnership Through 2031 - Developers Get Fourth blankOS 26.6 and Third blankOS 27 Betas - Third iOS 27 Developer Beta Brings Siri Voice Customizations - Third watchOS 27 Beta Brings Siri AI - Third macOS 27 Developer Beta Adds New Golden Gate Wallpaper - New iOS Popup Alerts Users When Data Being Sent to Google Cloud - Apple Seeds Fourth Sonoma and Sequoia RCs to Developers - Apple TV Announces Thriller Series "Guilty Creatures" - Apple Starts Free 4K Upgrades of Some TV Shows - Sponsored by Helix Sleep: Get 20% off Sitewide, 25% off Luxe Mattresses, and 30% off Elite Mattresses at helixsleep.com/macosken - Sponsored by Notion: Learn more about Notion's Developer Platform today at notion.com/macosken - Catch Ken on Mastodon - @macosken@mastodon.social - Send Ken an email: info@macosken.com - Chat with us on Patreon for as little as $1 a month. Support the show at Patreon.com/macosken
As Chief Investment Strategist at Hightower Advisors, Stephanie Link spends her days researching the market's winners. Today, she's breaking it all down for us. Stephanie explains why the economy keeps defying the doom-and-gloom headlines, which stocks she thinks will champion the next decade, and why she believes we're only in the third inning of the AI revolution. Nicole and Stephanie get tactical fast: the difference between the Mag 7 and the "S&P 493," which stocks are not worth the hype, and Stephanie's thesis that cybersecurity stocks will end up being bigger than AI. She names her favorite tickers across cybersecurity, data centers, robotics, and quantum computing, and explains whether it's too late to buy NVIDIA. Stephanie also gets real about the so-called "AI bubble," the circular spending debate freaking out investors, and why she bought SpaceX but capped it at just 2% of her portfolio. Plus: what her 19-year-old daughter is investing in, why "FAANG" just got a 2026 update, and the one boring, unsexy ETF Stephanie says every new investor should consider. Check out Nicole's financial literacy course The Money School Find a Financial Advisor or Financial Coach from Nicole's company Private Wealth Collective Watch video clips from the pod on Money Rehab's Instagram and Nicole Lapin's Instagram Read more about Stephanie's work Here's what Nicole covers with Stephanie: 00:00 Are You Ready for Some Money Rehab? 01:12 Stephanie Link Joins Money Rehab 01:41 Why the Market Keeps Defying the Doom and Gloom 04:07 CapEx, Decoded 06:58 The K-Shaped Economy: Why the Vibes Don't Match the Numbers 09:49 Inflation, Oil Prices, and the War's Ripple Effect 11:41 Mag 7 vs. the S&P 493: Which ETF Should You Buy? 16:16 Why Stephanie Says No to Leveraged ETFs 17:32 Cybersecurity Will Be Bigger Than AI 20:16 The Best Cybersecurity Stocks on Stephanie's List 23:15 How to Vet a CEO Before You Buy the Stock 25:16 Investing Lessons from Stephanie's 19-Year-Old Daughter 28:23 What Stephanie Won't Buy: Crypto, Staples, and Energy 32:20 Inside the AI "Food Chain": Powering the Data Center Boom 35:43 Robotics and Quantum Computing: The Next Big Themes 41:00 MicroStrategy vs. Palo Alto: What "On Sale" Really Means 43:07 FAANG Is Dead, Long Live MANGOES 43:47 Why Stephanie Bought SpaceX (and Kept It to 2%) 55:40 Is It Too Late to Buy NVIDIA? 59:06 Grading the Innings: AI, Cybersecurity, and Robotics 59:50 Hot Stocks: Micron, SanDisk, and the Chips Everyone's Chasing 1:02:27 Is There an AI Bubble? 1:03:16 The Circular Spending Debate 1:08:08 Stephanie Link's Tip You Can Take Straight to the Bank All investing involves the risk of loss, including loss of principal. This podcast is for informational purposes only and does not constitute financial, investment, or legal advice. Always do your own research and consult a licensed financial advisor before making any financial decisions or investments. Disclosures: Hightower invests in companies including Boeing Company, Dover Corp, General Electric, Quanta Services, Rockwell Automation, Union Pacific Corporation, Broadcom, International Business Machines Corporation, Marvell Technology Inc, ServiceNow Inc, Palo Alto Networks Inc, Snowflake Inc, Synopsys Inc, Bank of America, Capital One, Coinbase, Morgan Stanley, Truist Financial Corp, Amazon.com Inc, Meta, SpaceX, Alcoa Corp, Antofagasta PLC, Natera Inc, UnitedHealth Group Inc, iShares MSCI Brazil ETF, SLB Limited
OpenAI has released GPT-5.6, but the majority of us will have to wait. ⌚After the Anthropic vs. U.S. Government feud, it now looks like we'll have to wait for frontier models. That wasn't the only big AI news headline that might change your company's AI strategy. Anthropic got the green light to roll out Mythos 5 to a select few, Google reportedly extended its strike team to catch up on coding and more. OpenAI's limited release of GPT-5.6, Mythos starts slow reinstatement, OpenAI gets spicy and more AI news -- An Everyday AI chat with Jordan WilsonNewsletter: Sign up for our free daily newsletterMore on this Episode: Episode PageToday's Episode on LinkedIn: Thoughts on this? Join the convo on LinkedIn and connect with other AI leaders.Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineupWebsite: YourEverydayAI.comEmail The Show: info@youreverydayai.comConnect with Jordan on LinkedInTopics Covered in This Episode:OpenAI GPT-5.6 Limited Release ExplainedOpenAI Sol, Terra, Luna Model NamingUS Government Restrictions on AI RolloutsAnthropic Mythos 5 Access and StandoffAnthropic Fable 5 Suspension DetailsGoogle Gemini 3.5 Pro Release DelayedGoogle's AI Coding Mid-Training InitiativeRaiseUS Nonprofit: AI Workforce AdaptationAnthropic Accuses Alibaba of Model DistillationOpenAI & Broadcom Unveil Jalapeno AI ChipKey AI Industry Partnerships & Product LeaksTimestamps:00:00 OpenAI's GPT 5.6 limited release06:14 OpenAI's new model release details09:54 Access suspension and negotiations13:18 Google's AI strategy and delays16:33 Anticipating Gemini 3.5 Pro Release20:40 Accusations of AI model theft24:48 OpenAI and Broadcom chip partnership28:05 OpenAI's recent developments and updates29:56 OpenAI and AI weekly updatesKeywords: GPT-5.6, OpenAI, Anthropic, Mythos 5, Fable 5, Frontier models, Gemini 3.5 Pro, Google, model rollout, limited AI access, AI safety, US government AI regulation, Sol model, Terra model, Luna model, Max reasoning mode, Ultra mode, sub agents, advanced AI benchmarks, coding workflows, cybersecurity, third-party AI analysis, government licensing, AI model guardrails, AI model democratization, model naming scheme, model availability, AI model security, jailbreak resistance, safety filters, general model access, trusted testers, AI export control, national security, Anthropic pullback, supply chain risk, defense department, AI industry competition, talent loss, AI coding, mid training, engineering agents, AI strike team, RaiseUS nonprofit, workforce AI disruption, technology policy, industrial scale distillation, Alibaba, AI model theft, China-US tech tensions, distillation attacks, Jalapeno AI chip, Broadcom, AI inference, custom hardware, data center GPUs, Microsoft, Meta, Elastic compute, AI-powered career navigation, Slack Claude Tag, Canva Grow 2.0, Copilot skills, AI ad creation, AI automation, DigitalOcean plugin, Apple hardware AI, smart glasses, Vision Pro, portfolio tracking AI, Google Finance, home smart speakers, voice AI, GLM 5.2, open source AI, US labor market AI effects, AI job disruption, model leaks, government approval delays.Send Everyday AI and Jordan a text message. (We can't reply back unless you leave contact info) Start Here ▶️Not sure where to start when it comes to AI? Start with our Start Here Series. You can listen to the first drop -- Episode 691 -- or get free access to our Inner Cricle community and all episodes: StartHereSeries.com Also, here's a link to the entire series on a Spotify playlist.