Podcasts about Layer

  • 3,154PODCASTS
  • 7,149EPISODES
  • 44mAVG DURATION
  • 1DAILY NEW EPISODE
  • Aug 29, 2026LATEST

POPULARITY

20192020202120222023202420252026

Categories



Best podcasts about Layer

Show all podcasts related to layer

Latest podcast episodes about Layer

Agent Rise with Neil Mathweg (formally Onion Juice)
The Tech Wizard Who Built His Business on Handshakes (featuring Green Bay Greg)

Agent Rise with Neil Mathweg (formally Onion Juice)

Play Episode Listen Later Aug 29, 2026 53:11


Green Bay Greg has AI running through his entire business — ranking on Google, a mobile app, tools listening to every call. And he'll tell you straight up: none of it will save you. In this episode, Neil sits down with Green Bay Greg, one of the most tech-savvy agents in the business and the leader of a 14-agent independent brokerage doing over 450 transactions a year. Here's the twist: more than half of those deals still come from past clients and sphere, built on a decade of consistency, client events, and doing the mundane work most agents avoid. Greg is the rare tech expert who argues the tools aren't the point. AI is only valuable when it makes you a more efficient version of an agent who's already mastering the fundamentals — the calls, the matchmaking, the follow-up, the real customer service. Layer tech on top of that, and you become impossible to catch. Skip the fundamentals and buy tools to avoid them, and you get smoked. If you've ever felt behind on AI or buried in shiny objects, this episode is permission to double down on what actually works — and proof, from the most advanced tech agent Neil knows, that the fundamentals still win. If you want a business that's clear, consistent, and congruent with who you are, this one's for you.

Code Story
S13 Bonus: The Intention Layer: Why Coding Agents Need Product Judgment with Drew Dillon, Founder & CEO of Brief

Code Story

Play Episode Listen Later Aug 27, 2026 19:04 Transcription Available


Drew Dillon is a tech founder and product leader, centered in the Bay Area. He is a person who likes to do a bit of everything, having an engineering background, but spending time in design, sales, and product management. He's been a part of many startups - Yammer being one of them, and several others from Y Combinator - along with being a consultant and fractional CPO. Outside of tech, he has 2 kids that take up most of his time. When he isn't on the sidelines of a soccer game, he enjoys woodworking and reading a good sci-fi book. At a prior startup, Drew built a system using AI, which spit out something reasonably shaped. He took a look at the process he went through to create it, and realized that there are a ton of disparate parts, all separate from the shared organizational context - IE the "why" behind what is being built. He set out to change that, and built institutional memory and product judgement within AI agents. This is the creation story of Brief. SponsorsTiger DataProtected HarborRenderLinkshttps://briefhq.ai/https://mymoxieai.com/https://www.linkedin.com/in/drewdil/Checkout our episode stacks on Stacklist! https://stacks.codestory.co/ Hosted by Noah Labhart | Technical Founder & Startup Mentor.Advertising Inquiries: https://redcircle.com/brandsPrivacy & Opt-Out: https://redcircle.com/privacy

BlockHash: Exploring the Blockchain
Ep. 766 Memoket | Building the Context Layer for AI (feat. Elisa Lu)

BlockHash: Exploring the Blockchain

Play Episode Listen Later Aug 25, 2026 29:26


For episode 766 of the BlockHash Podcast, host Brandon Zemp is joined by Elisa Lu, COO and Co-founder of Memoket, an AI hardware company building the context layer between the physical world and AI. Before Memoket, she spent nearly a decade in brand commercialization and marketing leadership at Anker Innovations, where she grew brand revenue 6x in three years. At Memoket, she leads international operations and Marketing and is building the company she wished existed during years of running businesses on back-to-back conversations. Stop briefing your AI from scratch after every meeting. Memoket Gem ships August 2026. Grab your early bird spot for US$179 at memoket.ai — only 2,000 available!

Run The Numbers
The Stripe Guide to Pricing, Billing, and Quote-to-Cash with Wisam Hirzalla

Run The Numbers

Play Episode Listen Later Aug 20, 2026 54:54


In this episode of Run the Numbers, CJ sits down with Wisam Hirzalla, Head of Product at Stripe Billing, to trace the full journey of a dollar—from product packaging and pricing through quoting, contracts, billing, tax, collections, and revenue recognition.—SPONSORS:Anrok is the sales tax platform that watches your exposure everywhere, automates compliance, and flags risk before it turns into a surprise back-tax letter from a state you've never set foot in. Companies like Anthropic, Notion, and Vanta already trust Anrok to stay ahead of rules that move faster than any spreadsheet can. Talk to a sales tax expert for a personalized exposure estimate at https://www.anrok.com/rtnRightRev is a revenue recognition platform built for the AI economy, helping finance support usage-based pricing, credits, hybrid contracts, seats plus consumption, and whatever commercial model comes next. It gives product teams the freedom to keep innovating without outdated revenue systems slowing them down. Learn how RightRev can help at https://rightrev.com/cjPulley is an equity management platform that lets you issue options, model dilution, and complete 409As without your cap table turning into a spreadsheet disaster. Founders raising, hiring, and scaling use Pulley to keep equity clean and stay focused on building. Learn more or request a demo at https://pulley.com/mostlymetricsRillet is an AI-native ERP built for modern finance teams that want to replace NetSuite and close faster. With revenue recognition, close management, multi-entity support, and native Stripe and Salesforce integrations, Rillet helps scaling companies run their finance stack in one place. Hundreds of teams, including Windsurf and Mercor, use Rillet to make the zero-day close real. Book a demo at https://www.rillet.com/cjMaximor is an autonomous finance platform that runs order-to-cash, procure-to-pay, the close, cash management, and reporting on self-learning agents instead of a dozen disconnected tools. One PE-backed customer posts 98% of transactions directly to its ERP, with the remaining 2% routed to a human for review. You pay for outcomes, not seats. See it at https://www.maximor.ai/Brex is an intelligent finance platform with AI-powered workflows that enforce expense policies at the point of sale, match receipts automatically, and reduce month-end close from weeks to hours. Thousands of companies, including Anthropic, Coinbase, and DoorDash, already run on Brex. Stop asking A-level finance talent to do B-level admin work. Learn more at https://www.brex.com/metricsEY has been part of Silicon Valley since it was just a valley, helping the most successful names in tech go from startup to exit to megacap. With teams across strategy, tax, audit, and transactions, EY helps you get your financials right early, long before your investors start asking for it. You build the next big thing, and EY will help you build it right. Learn more at https://www.ey.com/techstartups—LINKS: Mostly Talent: https://mostlymetrics.typeform.com/to/cLTxtAsNGuest: https://www.linkedin.com/in/wisam-hirzalla-9b20a91/Company: https://stripe.com/CJ: https://www.linkedin.com/in/cj-gustafson-13140948/Mostly metrics: https://www.mostlymetrics.com—TIMESTAMPS:0:00 What Drives Billing Downstream1:41 Welcome to Run the Numbers2:37 Wisam's Path to "Queen of Billing"4:55 What Head of Billing at Stripe Entails5:45 Defining Quote-to-Cash7:50 Layer 1: Pricing and Packaging Mistakes14:46 Pricing Models Rising Because of AI18:11 Layer 2: Inside a Quote19:27 Why Enterprise Deals Get So Messy24:38 Layer 3: Contracting27:21 Layer 4: Billing30:15 Why Usage-Based Billing Is So Complex34:51 Tokens as a Pricing Unit36:57 Billing's Move Out of the ERP39:53 Layer 5: Provisioning and Entitlements43:36 Layer 6: Tax44:27 Layer 7: Getting Paid48:18 Layer 8: RevRec50:52 When Billing Stopped Being Back Office51:47 Asking Better Questions Segment52:28 Will AI Replace Finance Jobs

synaigy insights! mit Joubin Rahimi
161 DXP, AI und der Personalisierungs-Mythos: Warum jetzt erst alles losgeht

synaigy insights! mit Joubin Rahimi

Play Episode Listen Later Aug 20, 2026 17:13 Transcription Available


In dieser Folge sprechen Joubin Rahimi und Sebastian Stang, CRO von Magnolia, darüber, was eine Digital Experience Platform heute wirklich leisten muss, warum der Begriff DXP gerade selbst von Analysten neu sortiert wird und wie AI die Content-Erstellung im CMS gerade fundamental verändert. Du erfährst, warum Bring Your Own LLM der pragmatischste Weg in einer Token Economy ist, warum die Personalisierung vor zehn Jahren ein Versprechen war und jetzt erst real wird, und warum GEO zwar wichtig, aber nicht überschätzt werden sollte. Ein ehrlicher Inside Talk für alle, die Plattformstrategie, Commerce und Content endlich zusammen denken wollen. 3 Key Takeaways 1. DXP ist kein Produkt, sondern eine Orchestrierungslogik. Wer eine Plattform baut, kombiniert CMS, PIM, DAM und Commerce mit AI als verbindendem Layer. Den einen DXP-Standard gibt es nicht, und genau das macht die Auswahl deiner Tools strategisch so wichtig. 2. AI gehört in deine Workflows, nicht in die Zwischenablage. Copy-Paste zwischen ChatGPT, Claude und CMS ist die Übergangslösung, echte Adoption beginnt erst, wenn Content-Erstellung, Übersetzung und Publikation nahtlos im System laufen. Bring Your Own LLM ist dabei der richtige Ansatz, weil Compliance-Prozesse selten verhandelbar sind. 3. Personalisierung war zehn Jahre lang Versprechen, jetzt wird sie eingelöst. Hyperpersonalisierung bedeutet, dass die AI Inhalte und Seiten je nach Daten und Kontext automatisch zusammenstellt. Wer jetzt nicht startet, optimiert in zwei Jahren noch immer auf Standard-Templates, während andere Unternehmen längst individuelle Customer Journeys ausliefern.

The Data Chief
How Vanguard is Architecting its AI Semantic Layer

The Data Chief

Play Episode Listen Later Aug 19, 2026 48:12


Semantic layers and ontologies have moved from nice-to-have data modeling tools to the foundational engine required for enterprise AI. In this episode, Raman Tallamraju, Senior Director and Head of Enterprise Data Architecture and Engineering at Vanguard, breaks down how Vanguard is architecting its AI semantic layer to turn scattered institutional knowledge into reliable, agent-ready context. He shares why autonomous agents expose decades of hidden data debt, how to bridge domain-specific definitions like clients versus prospects, and how to balance building a unified semantic layer with a pragmatic, federated data operating model. Key Moments: Why Data Is Your Differentiator (02:36): Across four asset managers, Raman shares the one lesson that holds: data is the real differentiator in an AI-first world. Why Semantic Layers and Ontologies Are a Priority (14:12): Autonomous agents remove the human workaround, exposing years of technical debt in data modeling that tribal knowledge used to hide. Why Context Makes or Breaks Your AI Agents (19:18): Context is everything. Agents act confidently on wrong answers when terms like client, prospect, and lead go undefined. Launching an AI-Ready Data Program: Where to Start (24:54): Raman advises starting with a real business use case tied to points of economic leverage, rather than trying to boil the ocean. Why You Need a Federated Data Model (36:51): Raman explains why a single central platform isn't practical at a global firm, and the four levers he uses to earn real business ownership of data. Key Quotes: “If we're going to go in an AI-first world and everybody has access to the same frontier models… well, what is going to be differentiated about you? Data will be your differentiator, and the companies that bring the best data game are going to have enduring advantages over those that don't.” - Raman Tallamraju “You want your data to be well-defined. You want your data to be trusted. You want your data to be well-connected. You want your data to be contextualized. You want your data to be consumed in a multimodal way, and you want it to be ready for humans and machines at scale.” - Raman Tallamraju “If you're going to start with tech and you're going to go towards the shiny toys, those are good, but without the data foundations, they're going to sit in the garage. So I started calling it the Ferraris in the garage problem. Unless you get the data strategy running ahead, these Ferraris are going to run out of gas.” - Raman Tallamraju Mentions: The Innovator's Dilemma by Clay Christensen 57% of enterprises traced a wrong AI answer to missing business context — Credible bets portable, open-source semantic code beats proprietary metadata Guest Bio: Raman Tallamraju has twenty years leading enterprise technology and data strategy across four of the world's largest asset managers — Vanguard, T. Rowe Price, Capital Group, and Fidelity. Appointed Officer and Group Vice President at T. Rowe Price, where Raman built and led a 300-person global technology organization responsible for the firm's enterprise architecture and digital transformation. Wharton CTO Program, 2024. Raman has operated across the full investment value chain — research and trading platforms, client experience, distribution technology, and enterprise data infrastructure — with portfolio accountability exceeding $100M. Raman's career has been defined by taking on complex, high-stakes technology transformations and delivering measurable business outcomes: modernizing legacy estates, building scalable platforms, and creating the organizational structures that sustain them. Hear more from Cindi Howson here. Sponsored by ThoughtSpot.

AdTechGod Pod
Ep. 147: The Partnership Layer of AI with Google's Ravi Viswanathan

AdTechGod Pod

Play Episode Listen Later Aug 18, 2026 27:07


AdTechGod sits down with Ravi Viswanathan, Strategic Partnerships Lead at Google, to discuss his journey from software engineering to partnerships, more than a decade at Google, and how AI could reshape advertising. They explore generative creative, new ad-supported business models, consumer value exchange, and Ravi's approach to building trusted partnerships. Takeaways: A technical background can create credibility and deeper trust in strategic partnerships. AI's impact extends beyond productivity into creative generation and entirely new advertising experiences. Advertising could increasingly subsidize services such as mobile plans, hardware, and other everyday costs. Successful ad-supported models require a transparent and fair value exchange with consumers. Great partnerships begin by understanding the partner's problem rather than focusing on what you want to sell. Curiosity, technical fluency, and trust can lead to stronger client relationships and long-term career growth. Chapters:00:00 Welcome to the AdTechGod Pod00:13 Meet Ravi Viswanathan, Strategic Partnerships Lead at Google01:12 Ravi's Career Journey: Engineering to Strategic Partnerships04:07 Why Technical Knowledge Matters in Partnerships05:05 Why Ravi Has Stayed at Google for 11 Years07:30 The Continuing Evolution of Advertising08:09 How AI Is Changing Careers and Advertising10:02 AI-Powered Pricing, Targeting, and Generative Creative10:40 Will Consumers Accept Fully AI-Generated Advertising?14:22 What's Next for Advertising?16:10 Advertising Across New Screens and Consumer Touchpoints17:27 The Future of Ad-Supported Business Models20:07 Creating a Fair Value Exchange for Consumers20:45 Misconceptions About Working With Google22:38 Ravi's Career and Partnership Advice24:23 Closing Thoughts Learn more about your ad choices. Visit megaphone.fm/adchoices

ShopTalk » Podcast Feed
728: BIMI, Good Favicon Practice, and CSS @Layer Exploration

ShopTalk » Podcast Feed

Play Episode Listen Later Aug 17, 2026 61:09


Show DescriptionHaving type two fun, email branding with BIMI, accessibility details in icons and favicons, using AI as a coding and maintenance partner, CSS @layer property, design systems, and the economics of fast AI hardware. Listen on WebsiteWatch on YouTubeLinks How and Why to Implement BIMI Selectors - BIMI Group You kinda want an orange favicon. – Chris Coyier Thinking Horizontally in CSS @layer AMD snaps up Toronto chip startup Taalas | Financial Post

The Word on Investing by TRADEway
The Faith Layer: What Makes Christian Traders Different

The Word on Investing by TRADEway

Play Episode Listen Later Aug 17, 2026 4:55


What truly sets a Christian trader apart? It's not that Christians never make mistakes or always have winning trades. The difference is found in the foundation we build upon. In this episode, we explore how a Biblical worldview changes the way we approach the stock market. When we recognize that Christ is Lord over every area of life—including our finances and investing—we can trade with confidence rooted in Him rather than in our own understanding. Drawing from Proverbs 3, Proverbs 16, and Philippians 4 (KJV), we discuss how faith helps protect traders from both pride during success and discouragement during difficult seasons. Instead of placing our confidence in market predictions or personal ability, we learn to trust the Lord, commit our work to Him, and remain content whether we are experiencing abundance or adversity. In this episode, you'll learn: ✔ What it means to approach the markets from a Biblical worldview ✔ Why faith gives Christian traders a unique perspective ✔ How trusting God helps overcome fear, pride, and discouragement ✔ The true context of Philippians 4:13 and why it matters for traders ✔ How Biblical principles can shape wiser financial decisions and long-term stewardship At TRADEway, we believe trading is more than pursuing financial success. It's about developing wisdom, discipline, and faithful stewardship as we seek to honor God with the resources He has entrusted to us.

Moneycontrol Podcast
5263: Quick commerce firms cross 9M daily orders; Canva co-founder Cameron Adams interview; and explained: The context layer for AI agents

Moneycontrol Podcast

Play Episode Listen Later Aug 17, 2026 8:09


In today's Tech3 from Moneycontrol, The government steps up its electronics manufacturing push as ECMS clears Rs 7,877 crore of projects, while India's quick-commerce market crosses 9.5 million daily orders. Canva is expanding its AI and enterprise ambitions in India, now its fourth-largest market, and Zetwerk moves closer to its IPO with a Rs 2,600-crore fresh issue, an OFS and a DRHP that highlights both improving operating performance and key risks.

Telecom Reseller
Meetric Sees Conversational Intelligence as a New Revenue Layer for Service Providers, Podcast

Telecom Reseller

Play Episode Listen Later Aug 14, 2026


“The best value you will get out of it is if you get as many conversations and as many conversation types as possible,” says Mattias Ohde of Meetric. In this special podcast for the Cloud Communications Alliance and TR Publications, Doug Green speaks with Mattias Ohde of Meetric about conversational intelligence, AI and the new opportunity emerging for service providers, UCaaS providers, mobile carriers, MSPs and channel partners. Ohde says Meetric focuses on conversational intelligence for service providers, helping them collect and structure conversations from telephony, video, live meetings, email and other sources. Once those conversations are organized, AI can help extract insights, create workflows, support automation and make everyday work more structured. The conversation centers on a shift in how business communications are understood. For years, companies have recorded calls and saved voicemails, but much of that information was difficult to search, organize or use. Ohde says AI changes that by allowing conversations to become data points. “Now with the introduction of AI, you can all of a sudden consider these conversations as data points,” Ohde says. That creates what Ohde describes as a kind of “gold mine” for businesses. By bringing together conversations from across departments and communications channels, companies can gain a broader view of customers, operations and markets. They can also run more advanced analysis across the accumulated body of communications, creating insights that were not previously practical. For service providers, the opportunity is both operational and commercial. Ohde says conversational intelligence gives providers something new to bring to customers beyond traditional voice, UCaaS and collaboration tools. It can help providers deliver a more direct impact on the customer's daily work while creating a higher-value service layer. “This is in a way quite revolutionary for the industry itself,” Ohde says. Meetric delivers its services as a white-label offering, allowing mobile carriers, UCaaS platform providers and local service providers to bring conversational intelligence to market under their own brand, through their own invoicing and bundled with other services. Ohde says that creates a fast route to market for providers looking to add AI-enabled services without building the full capability themselves. He also discusses MCP and open AI workflows, noting that Meetric's functionality can be used inside other tools and platforms. That allows conversational data and AI-generated insights to become part of broader workflows across the customer's business. Ohde says the revenue opportunity is significant because conversational intelligence services may command much higher market pricing than traditional subscription services for UCaaS or mobile connectivity. For CCA members, MSPs and channel partners, the message is direct: the conversations customers are already having may contain untapped value. AI now gives providers a practical way to capture, organize and act on that information — while building a new service and revenue opportunity on top of the communications infrastructure they already provide. Learn more at https://meetric.com

Let's Talk Cabling!
AHL Kickoff Meetings That Save Projects

Let's Talk Cabling!

Play Episode Listen Later Aug 13, 2026 41:12 Transcription Available


Send us Fan MailWe tackle the real-world problems that slow down low voltage projects: messy handoffs, unclear ownership, and customers whose expectations do not match the contract. We share practical systems for kickoff meetings, estimating, data center career moves, and the growing need for networking and security awareness as everything rides on the network. • setting a sales to operations handoff standard with a kickoff meeting agenda • defining the documents sales must deliver to run a project cleanly • aligning customer expectations with schedule and contract language early • building estimating skills through blueprint reading, takeoffs, and labor units • using labor units based on average productivity and tracking them on jobs • finding data center jobs and understanding faster pace and higher quality demands • preparing for fiber-heavy work and learning high-count fiber basics • deciding how much Layer 2 knowledge helps for PoE, scripts, and device turn-up • explaining Wi Fi 7 performance limits without overwhelming the customer • developing lead techs with standards, mentorship, and troubleshooting habits • pushing customers to name one network decision owner across multiple systems • clarifying physical security vs logical security responsibilities in the field DM me, let me know, “Hey Chuck, I'm interested in that estimating class,” and I will reinstitute it and put it out there for people If anybody knows Mike Rowe, please put him in touch with me so I can get him on the showSupport the showKnowledge is power!  Make sure to stop by the webpage to buy me a cup of coffee or support the show at https://linktr.ee/letstalkcabling .  Also if you would like to be a guest on the show or have a topic for discussion send me an email at chuck@letstalkcabling.com Chuck Bowser RCDD TECH#CBRCDD #RCDD

Vanguards of Health Care by Bloomberg Intelligence
Innovaccer's Bet on the Data Layer Powering Autonomous Healthcare

Vanguards of Health Care by Bloomberg Intelligence

Play Episode Listen Later Aug 13, 2026 50:21 Transcription Available


“Your doctor has more outdated technology than your Uber driver does,” Abhinav Shashank, co-founder and CEO of Innovaccer, tells Bloomberg Intelligence analyst Jonathan Palmer in this episode of the Vanguards of Healthcare podcast. Shashank explains why fixing healthcare’s fragmented data infrastructure is the foundation for the industry’s AI future and why long-term value will accrue to companies that own the data layer rather than the user interface. He also traces Innovaccer’s evolution from a data startup into a platform spanning 80 million patient lives, and argues that “autonomous healthcare” could strip hundreds of billions of dollars of administrative waste. The conversation explores the economics of AI, the danger of poorly designed automation, how acquisitions are filling gaps and Innovaccer’s ambition to reach $1 billion of annual recurring revenue.See omnystudio.com/listener for privacy information.

CRYPTO 101
Ep. 743 Why Monad Is Taking Ethereum's Biggest Bottleneck Head-On with Keone Hon

CRYPTO 101

Play Episode Listen Later Aug 13, 2026 27:12 Transcription Available


In this episode of the Crypto 101 Podcast, Keone Hon, co-founder of Monad, joins from the Out East Summit to explain how his background at Jump Trading helped shape Monad's high-throughput, low-latency blockchain design. He breaks down how Monad uses pipelining, parallel execution, faster block times, and faster finality to create a more efficient EVM-compatible Layer 1. Check out Omaha Steaks and use my code BEEF for a great deal: https://www.omahasteaks.comCheck out Scribe and use my code scribe.how/CRYPTO101 for a great deal: https://scribe.comCheck out Quince: https://quince.com/CRYPTO101Check out Shopify: https://shopify.com/crypto101Check out ShipStation and use my code crypto for a great deal: https://www.shipstation.comGet my #1 altcoin pick for this month.Get immediate access to my entire crypto portfolio for just $1.00 today! Get your FREE copy of "Crypto Revolution" and start making big profits from buying, selling,Get immediate access to my entire crypto portfolio.. just $1.00 today! Go here to get access: https://www.crypto101insider.com/cryptnation-directm6pypcy1?utm_source=Internal&utm_medium=YouTube&utm_content=Podcast&utm_term=20250916Get your FREE copy of "Crypto Revolution: Your Guide To The Future of Money". In this book, I reveal how to make (and keep) a fortune during this crypto bull run! http://www.cryptorevolution.com/free?utm_source=Internal&utm_medium=YouTube&utm_content=Podcast&utm_term=20250916Chapters00:00 - Keone Hon joins from the Out East Summit01:00 - From Jump Trading to building Monad02:35 - Why Keone left trading to start a blockchain05:15 - Pipelining and high-performance blockchain architecture08:35 - Monad vs Ethereum block times and throughput10:55 - Why faster finality improves liquidity and DeFi13:40 - Purple, Kuru, prop AMMs, and apps on Monad16:10 - Why Monad is an EVM-compatible Layer 118:55 - Encrypted mempools, MEV, and pre-trade privacy21:30 - Running a Monad validator on a Costco MacBook25:10 - Coinbase token sale and Monad's open access visionSubscribe to YouTube for Exclusive Content:https://www.youtube.com/@crypto101podcast?sub_confirmation=1Follow us on social media for leading-edge crypto updates and trade alerts:https://twitter.com/Crypto101Podhttps://instagram.com/crypto_101Guest Linkshttps://x.com/keoneHD*This is NOT financial, tax, or legal advice*Boardwalk Flock LLC. All Rights Reserved  ▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬Fog by DIZARO https://soundcloud.com/dizarofrCreative Commons — Attribution-NoDerivs 3.0 Unported — CC BY-ND 3.0 Free Download / Stream: http://bit.ly/Fog-DIZAROMusic promoted by Audio Library https://youtu.be/lAfbjt_rmE8▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬Our Sponsors:* Check out Omaha Steaks and use my code BEEF for a great deal: https://www.omahasteaks.com* Check out Quince and use my code quince.com/CRYPTO101 for a great deal: https://www.quince.com* Check out Scribe and use my code scribe.how/CRYPTO101 for a great deal: https://scribe.com* Check out ShipStation and use my code crypto for a great deal: https://www.shipstation.com* Check out Shopify and use my code shopify.com/crypto101 for a great deal: https://www.shopify.comAdvertising Inquiries: https://redcircle.com/brandsPrivacy & Opt-Out: https://redcircle.com/privacy

0xResearch
Pump's Bull Case, Fomo's Social Moat & Crypto's Consumer Layer

0xResearch

Play Episode Listen Later Aug 12, 2026 52:22


Is the next onchain bull market already taking shape? This week, we unpack Pump's resurgence and Fomo's explosive growth as social trading brings users back onchain. We debate Pump's buybacks and token value, Fomo's social moat, where crypto infrastructure captures value, and whether trading can become social media. Enjoy! TIMESTAMPS: 00:00 Intro00:46 Pump's Bull Case08:24 Are Pump's Revenues Real?13:33 Pump's Buyback Debate23:58 The Best Pump Pair Trade31:19 Fomo's Breakout Growth39:12 Fomo's Social Trading Vision44:17 When Social Trading Becomes Gambling50:03 Get Your DAS Tickets! FOLLOW THE SHOW › 0xResearch – https://x.com/0xResearch › Luke – https://x.com/0xMether › Carlos – https://x.com/0xcarlosg › Shaunda – https://x.com/shaundadevens › Ryan – https://x.com/AvgJoesCrypto › Telegram – https://t.me/+UFFz4z3qyrhhMDYx › Blockworks – https://x.com/Blockworks Check out Blockworks Research today! Research, data, governance, tokenomics, and models – all in one place Blockworks Research: https://www.blockworksresearch.com/ Free Daily Newsletter: https://blockworks.co/newsletter EVENTS › Join us at Digital Asset Summit 2026 Asia October 7th & Digital Asset 2026 London November 10-11th https://blockworks.com/events DISCLAIMER Nothing said on 0xResearch is a recommendation to buy or sell securities or tokens. This podcast is for informational purposes only. Any views expressed are opinions, not financial advice. Hosts and guests. mayhold positions in the companies, funds, or projects discussed.

BIG Life Devotional | Daily Devotional for Women
2176 The Parable of the Pearl

BIG Life Devotional | Daily Devotional for Women

Play Episode Listen Later Aug 11, 2026 17:10


Yesterday we studied the Parable of the Treasure. Now, Jesus continues his teaching with another similar story tied to it. They go together, so let’s read them together. Matthew 13: 44-45, “The Kingdom of Heaven is like a treasure that a man discovered hidden in a field. In his excitement, he hid it again and sold everything he owned to get enough money to buy the field. Again, the Kingdom of Heaven is like a merchant on the lookout for choice pearls. When he discovered a pearl of great value, he sold everything he owned and bought it.” Remember yesterday was all about a man who discovered hidden treasure in a field. He just found it there – stumbled upon it. But now in this story of the pearl, we have a merchant on the lookout for choice pearls. A merchant is on a mission to find something specific. He is trained, he is experienced, and he knows exactly what he is looking for. One wasn’t looking and found hidden treasure. One was searching intentionally and discovered a pearl of great value. And isn’t it the same with finding Jesus. Some of us just stumbled right into Jesus when we weren’t even looking for him – while others of us found him after searching through every other religion, every other way of life and then we finally found Jesus. Which one is right? BOTH because they both result in the same life changing encounter. I didn’t go searching for Jesus. I had a beautiful childhood and lived in a sweet little bubble. At 15 my family started going to church and I was happy to go because in my tiny country town there wasn’t much else to do. I wasn’t expecting to meet a Savior, I didn’t even know I needed one – but there he was, so sweetly awaiting our encounter. It wasn’t dramatic. I didn’t have a bunch of bad decisions and regrets to lay on the altar. I just knew there was a prompting from within calling my surrender to his divine guidance. I was the one who stumbled on hidden treasure in a field when I wasn’t even looking for it. That’s the kingdom of heaven. Then there’s my friend Cat. In her 30’s she had walked down a path of addiction that had led her to a crashing end. While in rehab trying to get sober and regain her life, she was introduced to a variety of “higher powers” for the choosing. She began studying and trying them. Buddhism was surely it, so she tried it – but felt nothing. So on to the next one. How about Hinduism. No, that wasn’t it. She left rehab now sober, but still searching. On the way out, her counselor handed her her own personal copy of ‘Jesus Calling’ and said, “At some point you’re going to need this.” Cat tucked that little devotional book away in a closet and returned to life. It was a better life sober, but not a whole life. When the next big family catastrophe hit, she knew she needed something to carry her through and not return to the familiarity of her addiction, and she remembered the words of the counselor at rehab who had said, “At some point, you’re going to need this.” And Cat knew – if there was ever a time I need something, it’s now. So, she found that book, ‘Jesus Calling’ in the closet, she opened the page to that day’s date, and there she found every single word she most needed. But more than words, she found the loving, forgiving, life-changing Savior she had been searching for all along. AND SHE KNEW IT. She knew this was THE ONE. Jesus was the pearl of greatest value and nothing would ever be the same for her. Cat was the one who had been searching for years and finally found her pearl. That’s the kingdom of heaven. Isn’t it amazing that we have BOTH? Both a Savior who will pursue us when we aren’t even pursuing him – AND a Savior who will patiently wait as we pursue everything that isn’t him and perfectly align an encounter with him when everything else has failed? Jesus doesn’t fit in our little box. It doesn’t have to be one way. He is THE WAY, and he will make the way to you in whatever way that needs to be. He’s the hidden treasure you stumble upon without even looking – and he’s the pearl of great value you’ve studied and searched for your entire life. But remember this – either way, he’s worth giving up EVERYTHING for. In both parables, when found, everything else was sold. Remember that. Jesus requires everything. He’s worth everything. He’s greater than everything. For my sisters who are more like me without the dramatic story of searching for Jesus, and more just stumbling upon him when you weren’t even looking, remember you’ve found the greatest treasure of all! And for my sisters who have searched their entire lives, here’s Jesus and he’s everything you’ve been looking for! He is your pearl! And really think about Jesus being our pearl. How wild that he would use that as the example in his teaching. Do you know how a pearl is formed? A pearl begins with an irritation. If an oyster never has an irritation, they simply never have a pearl. But if their little world inside their shell is penetrated by the irritation of a grain of sand, then that irritation creates something that becomes valuable. The irritation of sand inside the shell causes the oyster to respond. Their response covers the sand with layers of naturally produced protection from within. Layer after layer, something beautiful is formed around something that originally caused great pain. Think about that – the very thing that irritated the oyster becomes part of what makes the pearl so valuable. Of course that’s why Jesus uses the pearl as an example – he knows this is how God works sometimes! Something you never asked for gets into your life. A disappointment, a betrayal, a loss, a failure, a terrible season, an series of bad decisions that turned into an addiction. Irritation grows. It gets worse. You’re so uncomfortable. But God is growing something within you. Something of great value. This is where Jesus shows up!!!!!!! Don’t you see how God can use what was so painful to produce something so beautiful in your life? He wastes nothing. The oyster doesn’t wake up one morning and say, “I think I’ll make a pearl today.” It simply responds to what has entered its life. And slowly, layer by layer, something extraordinary develops. Maybe that’s what God has been doing in you. You thought you were just surviving. God was forming you. You thought you were just getting through another hard season. God was adding another layer. You thought you were being delayed. God was developing depth. You thought the pain was ruining your story. God was making something beautiful from it. And then you find your pearl. You find your Jesus, that is what you’ve been searching for all along. And what do you do when you find that pearl – You don’t just keep looking for something else. You don’t keep trying other things. NO. You know this is what you’ve been looking for all your life. This is your answer. Here he is. And you give up everything else because you know there’s absolutely NOTHING better! What if some of the things you’ve spent your life chasing were never the pearl? But Jesus says, “You found it.” The Kingdom is the pearl. Jesus is the pearl. And when you finally understand His worth, you realize, I don’t have to spend the rest of my life searching. I’ve found what my heart was looking for. Jesus is worth everything. Follow Pamela on Instagram – https://instagram.com/headmamapamela Or Facebook – https://www.facebook.com/pamela.crim Find out more about BIG Life – http://biglifehq.com

Right on Radio
(Part 1)The Noahide Framework – Trump, Kabbalah, and the Legal Theology of the Beast System

Right on Radio

Play Episode Listen Later Aug 11, 2026 9:56 Transcription Available


In this episode Jeff Shepard presents a theoretical, evidence-driven interpretation tying together the Noahide framework, Kabbalistic influence, and Donald Trump's political role as a potential forerunner to a centralized global authority centered on Jerusalem. Shepard traces the Noahide laws from their rabbinic codification in Tractate Sanhedrin 56a through Maimonides and cites historical sources (including the 1906 Jewish Encyclopedia) on their enforcement provisions. He highlights Public Law 102-14 (1991) and subsequent presidential proclamations that recognize the seven Noahide laws as foundational to Western civilization and notes the public honors extended to Rabbi Menachem Mendel Schneerson. The episode surveys Chabad-Lubavitch as a primary institutional driver, explaining its Kabbalistic theology (Lurianic concepts such as tzimtzum, tikkun olam, and nitzotzot) and its global reach. Shepard connects those theological currents to Western esotericism and thinkers like Alice Bailey, arguing a continuity between mystical Kabbalah, Western occult streams, and plans for an externalized spiritual hierarchy. Focusing on Donald Trump, Shepard outlines documented links—public records, personal acknowledgments, and associates (Eitan Yerdeni, Jared Kushner's family ties, Michael Cohen's red string)—and discusses invitations from nascent Sanhedrin bodies and a proposed international court grounded in Noahide principles. He interprets these ties as part of a normalization process that could shift legal and geopolitical authority toward Jerusalem, including preparation for Third Temple developments (red heifers, reconstituted Sanhedrin). The core argument presents a three-layer architecture for what Shepard terms the BEAST system: Layer 1 (technological — neural interfaces, AI, CBDC-ready identification and surveillance networks), Layer 2 (financial and governance — central banking, CBDCs, regional governance models, Great Reset actors), and Layer 3 (legal-theological enforcement — Noahide law as a religious-legal mechanism to justify exclusion and capital enforcement). Shepard explains how these layers interlock to control commerce, monitor compliance, and supply theological justification for punitive measures. In the closing segment Shepard addresses responses for the Christian remnant: documenting institutions, building resilient off-grid communities, refusing coercion, and maintaining testimony in the face of potential legal classification of Christian confession as idolatry. He frames the analysis biblically (Daniel, Revelation, Isaiah) and calls listeners to study, test the evidence, and stand with faith. Episode details: host Jeff Shepard; topics include Noahide law history and enforcement, Public Law 102-14, Chabad-Lubavitch and Kabbalah, Alice Bailey and Western esotericism, Trump's documented connections and potential institutional role, the three-layer BEAST architecture (technological, financial, legal-theological), and practical spiritual and community responses. A study through the bible verse by verse and chapter by chapter. with host Jeff Shepherd. Want to Understand and Explain Everything Biblically? Click Here: Decoding the Power of Three: Understand and Explain Everything or go to www.rightonu.com and click learn more. Use coupon code Summer50 for $50. value savings until August 31st.. Thank you for Listening to Right on Radio. Prayerfully consider supporting Right on Radio. Click Here for all links, Right on Community ROC, Podcast web links, Freebies, Products (healing mushrooms, EMP Protection) Social media, courses and more...https://linktr.ee/RightonRadio Live Right in the Real World! We talk God and Politics, Faith Based Broadcast News, views, Opinions and Attitudes We are Your News Now. Keep the Faith

The Twenty Minute VC: Venture Capital | Startup Funding | The Pitch
20VC: Will OpenRouter Sell for $10BN to Stripe? | Why Chinese Open Models Are Beating America—and What Happens Next | Why Enterprises Are More Fearful of Anthropic and OpenAI Than China | Is the Routing Layer Becoming a Commodity with Alex Atallah

The Twenty Minute VC: Venture Capital | Startup Funding | The Pitch

Play Episode Listen Later Aug 10, 2026 59:59


Alex Atallah is the Founder and CEO @ OpenRouter, the unified interface for LLMs. The company has raised over $153M in funding, with the latest valuation pricing the company at $1.3BN. OpenRouter is reportedly in an acquisition process with Stripe for $10BN.  AGENDA: 00:00 Is OpenRouter Selling to Stripe for $10 Billion? 04:05 What Did Alex Learn From Scaling OpenSea? 06:38 What Did OpenRouter's Founding Thesis Get Wrong? 14:47 Is AI Model Routing Already Being Commoditized? 19:12 Do Falling Token Prices Help or Hurt OpenRouter? 27:16 Should America Be Alarmed by Chinese Open Models? 32:43 Will US Open-Source Models Compete With Chinese Models in the Next 12 Months? 39:26 Will the Router Be Swallowed by the Agent Framework? 48:56 Is Distillation Wrong—and How Should We Look at It? 50:01 Is the Reported $10 Billion Stripe Deal Actually Happening?

The Fintech Blueprint
Building the AI Distribution Layer for 5000+ Banks, with Fiserv Co-Head of Financial Solutions Srini Krish

The Fintech Blueprint

Play Episode Listen Later Aug 10, 2026 40:40


In this episode, Lex chats with Srini Krish — Co-Head of Financial Solutions at Fiserv, one of the original fintechs, in business for nearly five decades and sitting at the intersection of commerce and banking. Lex and Srini discuss how Fiserv acts as the technology backbone for 5,000+ US banks and credit unions that lack the wherewithal to match JPMorgan or Wells Fargo on their own, and how the firm is packaging AI into that distribution layer through Agent OS and partnerships with OpenAI and Anthropic. Srini lays out his four-bucket framework for enterprise AI - better client service, internal productivity, AI embedded in products, and a platform banks can use to build their own agents - and explains why money demands deterministic outcomes rather than probabilistic guesses, keeping a human in the middle as commercial loan underwriting compresses from weeks to hours. They explore the competitive race against challengers like Mercury and Ramp, the mainframe that has outlived thirty years of obituaries, and where power sits between the AI labs and their distribution channels once inference commoditizes. NOTABLE DISCUSSION POINTS: MIPS became tokens. Srini frames the whole AI shift through continuity: engineers once measured effectiveness by MIPS consumed and how often they compiled code; today the metric is token consumption. Same discipline of doing more with minimal resource, thirty years apart. Money forces determinism. Probabilistic outputs are fine for many tasks but unacceptable for balances - a figure 1% or 5% off is a failure, it has to be right every time. So Fiserv's Agent OS rollout starts with non-real-time, human-in-the-middle use cases and only graduates toward autonomy and eventually customer-built agents. It's a crawl-walk-run path, and Fiserv says it's clearly still crawling. The moat is distribution, not model access. Fiserv's 5,000+ banks and credit unions can't engage OpenAI or Anthropic directly at scale, so Fiserv becomes the platform that packages agentic workflows - turning commercial loan decisions from a multi-week process into hours, with the auditability and observability those institutions could never build alone. TOPICS Fintech, Fiserv, EmbeddedFinance, AgenticAI, EnterpriseAI, Banking, Payments, DigitalBanking, CommunityBanks, FinancialInfrastructure, AIAgents, OpenAI, Anthropic, ClaudeCode, JPMorganChase, FirstData, Mercury, Ramp, Plaid   ABOUT THE FINTECH BLUEPRINT

KYO Conversations
The $5 Million Mistake That Taught Him How to Live (Ft. Randy Cohen)

KYO Conversations

Play Episode Listen Later Aug 9, 2026 38:07


Randy Cohen is the founder and Chief Energizing Officer of TicketCity, which he started as a university student with $1,200 and six basketball tickets. Four decades later, he is an author, entrepreneur, connector, father, grandfather and the heart behind the Loop of Love—a philosophy centred on giving back, bringing people together and helping them leave a room feeling better than when they entered. For the discount at TicketCity use code: COHEN1403 - https://www.ticketcity.com/ EXCITING STUFF! My second book now available for pre-order! You're Invited: A Life-Changing Conversation with History's Greatest Wisdom Teachers Get your MENTAL FITNESS BLUEPRINT here! A special thanks to our mental fitness + sweat partner Sip Saunas   Connect with Marc: https://konect.to/marcchampagne   Timestamps: 00:00 — The question that opens every interview: “Who are you?” 00:45 — Randy's Loop of Love and the identity beneath his title 02:47 — Starting TicketCity with $1,200 and six basketball tickets 03:20 — “Heart of a cougar, soul of a lion” 06:18 — The early role model who taught Randy to light up a room 07:32 — Risk, the circle of life and entering a different season 09:13 — The LAYER model: listen, acknowledge, explore and respond 11:10 — The books and practices helping Randy learn to simply be 13:02 — Getting 1% better and shifting from winning to giving back 13:56 — The $5 million restaurant mistake 14:25 — Why you cannot own a business without showing up 15:23 — Closing the restaurants and discovering that less is more 18:44 — COVID, lost identity and watching his life force disappear 20:07 — How Randy began getting his colours back 21:16 — Small relationship practices that restore energy 24:09 — How gratitude creates a loop that interrupts overthinking 25:19 — Joseph Campbell, dragons and the Mental Fitness Blueprint 26:21 — Stop slaying dragons and start flying with them 29:28 — Randy's hot-tub ritual and morning hummingbird report 29:56 — Creating enough space to notice what others are carrying 31:49 — What a life well lived really means 33:01 — The people and principles that shaped Randy's swagger 36:49 — Final reflections, TicketCity and living the dash * Special props

The Tech Blog Writer Podcast
Creating a Coordination Layer for AI Agents With Blue Language Labs

The Tech Blog Writer Podcast

Play Episode Listen Later Aug 8, 2026 27:00


What happens when an AI agent is authorized to make a payment, but nobody can verify the wider agreement behind it? In this episode of Tech Talks Daily, I speak with Zor Gorelov of Blue Language Labs about the infrastructure businesses may need as AI agents move from answering questions to negotiating, approving, purchasing, coordinating, and settling commercial activity. Many current business processes depend on human coordination. People reconcile spreadsheets, chase signatures, confirm deliveries, review exceptions, and resolve disagreements between systems. This work often remains invisible because employees absorb the ambiguity through emails, calls, and follow-up. Agent driven business changes the speed and volume of those interactions. One agent making an isolated payment can be handled as a software transaction. Several agents coordinating dependent actions across companies, banks, suppliers, platforms, and customers creates a much larger infrastructure problem. Zor argues that authorization answers only part of the question. An agent may have permission to pay, but every participant also needs to understand what the payment covers, which conditions apply, who can approve changes, what evidence confirms delivery, and when funds should be captured, refunded, or settled. Blue Language Labs is developing an open source protocol designed to structure those commitments. Blue Documents represent machine executable agreements containing participants, permissions, obligations, conditions, and the current state of a business process. Blue Mandates provide agents with revocable authority. A business can define spending limits, permitted actions, and thresholds requiring human approval. The meeting notes include the example of a restaurant operator allowing an agent to accept smaller bookings automatically while requiring approval for catering orders involving over 20 people. Blue Timelines provide an append only, hash linked record of actions, approvals, and changes. The aim is to give participants an independent history they can use when resolving disputes, instead of relying on conflicting emails or records controlled by one company. Zor brings the concept to life through a travel package assembled by an AI agent. The agent identifies a boutique hotel with spare inventory, a restaurant with available tables, and a local guide with unused capacity. Each business defines its terms, the agent assembles the offer, and the participants approve their roles. The customer purchases one package. Payment can be authorized at the beginning and captured according to agreed conditions as the hotel, restaurant, and guide confirm fulfillment. If one participant declines or fails to deliver, predefined rules determine whether the agent finds a replacement, changes the package, or triggers a cancellation. We also consider how Blue differs from traditional workflow systems, agent orchestration tools, and blockchain smart contracts. Blue is designed for coordination across separate businesses without requiring every participant to join one company platform or use global blockchain consensus. The opportunity could be especially valuable for smaller companies. Agents may allow several independent businesses to combine inventory, services, and expertise into offers they could not create individually. Adoption will depend on whether businesses, banks, and customers trust the protocol, accept shared definitions, and retain meaningful control. What would need to be written into a machine executable agreement before your organization could rely on another company's AI agent? Listen to the conversation and share your thoughts with me.

The Scene, from Indiana Public Radio
S03 E26 - We Layer on the Flavor

The Scene, from Indiana Public Radio

Play Episode Listen Later Aug 7, 2026 48:15


Pop of Culture is back! This week, we head to a commercial kitchen to learn the art of layering flavors with Chef Jason Reynolds. He also talks about how his catering team maintains their culinary creativity while managing hundreds of events each month. Then, just ahead of the 2026 IndyFringe Festival (August 13 - 23), Executive Director Paul Daily explains how they turn dozens of shows around in ten days.We also meet comic book author—and owner of Aw Yeah comics in Muncie—Christina Blanch, and we'll hear from the remaining artists in this year's Muncie Three Trails Music Festival!

THORChain Weekly Live
Rujira App Layer Update on THORChain | Podcast #223

THORChain Weekly Live

Play Episode Listen Later Aug 7, 2026 111:14


In this episode, Hans and PragmaticMonkey give us a Rujira update on what has been done, what they are working on, and what to expect.Swap now https://swap.thorchain.org/ THORChain is a decentralized crypto exchange. THORChain is the first and biggest DEX for Bitcoin. You can use any self custody wallet to swap and there's no KYC required.Timestamps:00:00:00 Intro00:02:00 Kenton update00:04:00 Hans starts and expresses thanks for the ADR 31 result00:05:00 Recap of the history00:07:00 Market makers and perpetuals00:12:00 Perpetuals breakdown and Dynamic Concentrated Liquidity (DCL)00:15:00 Delta-neutral package00:17:00 PragmaticMonkey talks about bootstrapping the perpetuals market00:22:00 Why is the reserve involved in revenue generation?00:23:00 Status of the new stablecoin launch00:26:00 Exportable stablecoin00:32:00 Slow block times recently — do we know the cause?00:35:00 Hans leaves00:36:00 Order book screen share00:40:00 Automated order book explained00:43:00 Spread ratio explanation00:47:00 Custom Concentrated Liquidity will add much more utility00:49:00 COSMWASM MinBPS slip impact00:55:00 RujiTrade arbitrage revenue metrics01:00:00 Unique user index is increasing01:01:00 CCL is the best way to attract new users01:03:00 You can copy the strategies of other users01:07:00 Sophisticated users will always refine their strategies01:12:00 Are there plans to integrate tools into third-party systems?01:14:00 Kenton: We should have one API01:18:00 We are working together on this01:19:00 POL DAO is being developed01:22:00 AutoRujira update01:25:00 Main differences between CCL and DCL01:28:00 Getting more assets on THORChain is important01:32:00 PragmaticMonkey's thoughts on the sudden end of Kujira01:36:00 We are part of the same community01:40:00 Merger discussion01:43:00 How are the developer finances looking?01:44:00 When is the next community gathering?

Lend Academy Podcast
A New Intelligence Layer for Community Lenders with Mike de Vere, CEO of Zest AI

Lend Academy Podcast

Play Episode Listen Later Aug 6, 2026 34:42


Mike de Vere runs Zest AI, a company that has been applying machine learning to credit underwriting for over two decades, starting with some of the largest banks on the planet and now serving a large share of the credit union market. Since his last appearance on the show three years ago, Zest has expanded well past underwriting into fraud detection and portfolio management, tied together by an intelligence layer and a generative AI companion called LuLu. Mike makes a specific argument in this conversation: machine learning still makes the credit decision, generative AI makes the feedback loop faster, and the real advantage available to community financial institutions is a willingness to pool what they know.What We CoveredZest today, from underwriting to fraud to portfolio managementWhy the intelligence layer is what makes an ecosystemStarting with Discover, Citi and Freddie Mac, then moving down marketLuLu, named after a corgi, and what she actually doesSafety and soundness as the first use case for most institutionsReplacing quarterly reports that used to take weeksPeer benchmarking versus building your own data lakeCollective intelligence across 2,000 credit models in productionWhy generative AI has no role in making the credit decisionShrinking model refit cycles from 18 months to daily evaluationZest customers versus non-customers on growth, delinquency and efficiencyCash flow underwriting, and why generic national models failZest Protect and fighting AI-powered fraud with AIThe two objections that come up most in sales conversationsTakeaways from the IQ AI Lending Forum in Santa FeKey TakeawaysThe performance gap is measurable. Comparing Zest customers to non-customers across 2024 and 2025, Mike says his customers grew 16 times faster, ran roughly 20 points lower on delinquency, and were 501 basis points better on efficiency ratio.Generative AI belongs around the credit decision, not inside it. Zest still uses supervised, locked-down machine learning models for underwriting, because a regulator will ask you to explain the decision. What generative AI changes is the speed of evaluation, from an 18-month refit cycle to daily.Comparison is where the value sits. A lender looking only at its own data lake has visibility on itself and nothing else. LuLu is built to normalize performance data across institutions so a chief lending officer's instinct can be checked against thousands of real policy instances rather than one career's worth of experience.Community lenders have a structural advantage they underuse. The credit union industry holds roughly $2.4 trillion in assets. If it acted as one institution, it would be bigger than Wells Fargo, and unlike the big banks these institutions are actually willing to share.About Mike de VereMike de Vere is the CEO of Zest AI, the AI lending technology company that has been doing machine learning in credit since well before AI became a standard fintech conference track. He came to Zest from a career in data and consumer insights, with leadership roles at J.D. Power, The Harris Poll and Nielsen. Zest now touches $5.6 trillion in assets under management, and by the end of this year expects one in three credit union members to have their consumer loans decisioned with its technology.Connect with Fintech One-on-One:Tweet me @PeterRentonConnect with me on LinkedInFind previous Fintech One-on-One episodes

Michigan Insider
009 - 5 for 5 Rule is an Extra Layer of Chaos to College Sports 080526

Michigan Insider

Play Episode Listen Later Aug 5, 2026 6:11


See omnystudio.com/listener for privacy information.

Circles Off - Sports Betting Podcast
You Project 80, the Line Is 72.5... You Might Have Nothing | Presented by ProphetX

Circles Off - Sports Betting Podcast

Play Episode Listen Later Aug 5, 2026 35:44


You project a running back for 80 rushing yards. The line is 72.5. Nearly eight yards clear — an easy over, right? By the end of this video, that same exact setup might have no bet at all. Rob Pizzola walks through the full process of turning a raw projection into an actual price — not a vibe, not a gut call, a real number you can trust. Starting from a simple two-minute version anyone can use, then building layer by layer into how the market actually prices probability, why your model's mean and the market's median aren't the same thing, and why that gap is the single most common way a good projection turns into a bad bet. By the end, you'll be able to price any prop, and every alternate line on the board, from nothing but a projection and a little bit of math. This is Circles Off, part of The Hammer Betting Network. ⬇️ Download the Excel file: https://docs.google.com/spreadsheets/d/1j2dlXlL1fgN4rLoI1Q_mTEFpwpnXvr0f/edit?usp=drive_link&ouid=110512856455811077194&rtpof=true&sd=true INSTRUCTIONS: Once the file is open, in the top left go to File > Download > Microsoft Excel (.xlsx)

Leaders In Payments
Building the Trust Layer for Payments with Noam Izhaki, CEO of Ballerine | Episode 513

Leaders In Payments

Play Episode Listen Later Aug 5, 2026 22:45 Transcription Available


Merchant onboarding is where growth goes to die, and where fraud quietly sneaks in. We sit down with Noam Izhaki, Co-founder and CEO of Ballerine, to unpack why the payments stack can feel real-time and automated while KYC, KYB, underwriting, and compliance still depend on slow investigations, scattered systems, and ever-growing analyst teams.We walk through Noam's journey from building early online platforms in Tel Aviv to learning hard lessons in remittances and then at Wix, where the same merchant risk challenges showed up at scale. That experience led to Ballerine: a platform designed to help merchant acquirers, PSPs, marketplaces, card ecosystem players, and banks bring their policies and data into one place and use AI agents to automate decisions across the seller lifecycle, from onboarding through ongoing monitoring. We also dig into how this differs from traditional fraud and compliance point solutions that provide signals but still leave the hardest part, judgment, to humans.Then we zoom out to the future of payments: agentic commerce, agents buying from other agents, and a world where creating “a business” is cheap, fast, and sometimes fake. Noam shares what he's seeing around fraud industrialization, including transaction laundering as a service, and why the biggest advantage may be becoming a true trust layer for the internet with real-time, global risk decisions.If you're building for scale, ask yourself whether your plan requires hiring your way out of risk. Subscribe for more conversations like this, share the episode with a payments leader who's feeling the pressure, and leave a review with your biggest question about AI in merchant risk.

Ekabo Home Financial Freedom Mastermind Podcast
167. The Two-Layer Blueprint For Managing Rentals Across States

Ekabo Home Financial Freedom Mastermind Podcast

Play Episode Listen Later Aug 5, 2026 27:15 Transcription Available


Ethereum Cat Herders Podcast
Execution Layer Meeting 242[2026-07-30] | ACDE 242

Ethereum Cat Herders Podcast

Play Episode Listen Later Aug 5, 2026 61:30


The conversation covers the preparation and live streaming setup, updates on Lambsidam DevNet 7, preparation for Glamstadam DevNet 8, repricing analysis and benchmarking, results of the repricing analysis and DevNet 8 launch, support for Discovery V5 and mainnet transition, hive chain compatibility, EIP proposal deadline, and EIP proposals and discussion. The conversation covered several Ethereum Improvement Proposals (EIPs) related to gas accounting, state changes, account abstraction, transaction assertions, bloom filter removal, and EVMification of precompiles. The discussion also included housekeeping and updates on the status of various EIPs.TakeawaysLive streaming preparation and DevNet updatesRepricing analysis and benchmarkingSupport for Discovery V5 and mainnet transitionHive chain compatibility and EIP proposal deadlineEIP proposals and discussion Gas accounting and state changes are being addressed through EIPs 8116, 8037, 7807, and 7906.The removal of bloom filters and EVMification of precompiles are being proposed through EIPs 7668 and 8200, respectively.Housekeeping and updates on the status of EIPs are ongoing to ensure proper review and approval.Chapters00:00 Preparation and Live Streaming06:31 Lambsidam DevNet 7 Update08:44 Glamstadam DevNet 8 Preparation08:59 Repricing Analysis and Benchmarking11:39 Repricing Analysis Results and DevNet 8 Launch23:15 Discovery V5 Support and Mainnet Transition33:07 Hive Chain Compatibility and EIP Proposal Deadline36:38 EIP Proposals and Discussion42:00 EIP 8116 and EIP 8037: State Gas and Execution Gas43:35 EIP 7807: Block Hash and Logs Bloom Field44:48 EIP 7819: Account Abstraction and State Creation49:56 EIP 7906: Transaction Assertions and State Changes01:05:09 EIP 7668: Removal of Bloom Filters01:07:21 EIP 8200: EVMification of Precompiles01:18:36 Housekeeping and EIP Status Updates

Hospitality Daily Podcast
Our 4-Layer AI Framework: Data, Reporting, Insights, Action - Matt Schwartz, Sage Hospitality Group

Hospitality Daily Podcast

Play Episode Listen Later Aug 4, 2026 10:31


Sage Hospitality CTO Matt Schwartz shares the four-layer framework guiding the company's AI strategy: data, reporting, insights, and action. He explains why hotel data must be brought together and normalized before AI can produce useful analysis or support operational decisions.The framework moves from a reliable data foundation toward a future in which AI agents can take action, with the ultimate goal of giving property teams more time with guests and their other associates.Listen to the series with Matt:Episode 1: The Weekend Class That Changed My CareerEpisode 2: How We're Leading AI Adoption With a Human-First ApproachLearn more:Read Ethan Mollick's essay, The Bitter Lesson versus The Garbage CanSee how Actabl approaches hotel data normalizationLearn more about the Destination AI Forum in Washington, DC where Matt is speakingMore on hotel data and AI with Actabl:Why Actabl's Approach to Hotel Data Earned a Patent and Prepares Hotels for AIHow Hotel Companies Turn AI Into Competitive AdvantageHow Innovative Hotel Companies Are Building Better AI Through Forward-Deployed Engineering A few more resources:If you're new to Hospitality Daily, start here. You can send me a message here with questions, comments, or guest suggestionsIf you want to get my summary and actionable insights from each episode delivered to your inbox each day, subscribe here for free.Follow Hospitality Daily and join the conversation on YouTube, LinkedIn, and Instagram.If you want to advertise on Hospitality Daily, here are the ways we can work together.If you found this episode interesting or helpful, send it to someone on your team so you can turn the ideas into action and benefit your business and the people you serve!Music for this show is produced by Clay Bassford of Bespoke Sound: Music Identity Design for Hospitality Brands

Unleashed - How to Thrive as an Independent Professional
656. Jean-Christophe Lanoix, Turboconsultant: Running a Solo Practice on AI Agents, end-to-end

Unleashed - How to Thrive as an Independent Professional

Play Episode Listen Later Aug 3, 2026 42:54


Jean-Christophe Lanoix is the founder of Turboconsultant. He spent seventeen years at Hinicio, a strategy consultancy specialising in hydrogen — joining as an intern, rising to Associate Director and leaving in 2024, two years after the firm's exit. He now builds the system he wishes he had had. Introducing Turboconsultant An execution layer for solo consultants: a virtual team of AI agents spanning the whole practice — business development, research, methodologies, marketing and content, deliverable production, quality control, meetings and follow-up, admin. The consultant directs. The agents execute. It installs as a plugin on Claude Cowork and turns a generalist agentic workspace into a consulting-grade one. Three things separate it from a general-purpose assistant. Personalisation. A one-time onboarding hands it the consultant's own methodologies, templates, frameworks and voice. What comes out arrives in their template, follows their methods, and reads in their words. Quality control. Every document it produces passes three layers of audit before it leaves — sources traced, figures recomputed, claims contested by models outside the system. That is where AI hallucination and error get caught. Data walling. Nothing moves from one engagement to another, and the models do not train on client work. And it is one system rather than a patchwork of tools: the whole practice runs in a single environment that sharpens with every engagement. The net effect, measured on the existing user base, is a 25x on average, while improving the general quality of the work. Demonstration of Turboconsutlant  Both cases are invented, and so is the consultant behind them — a generic profile carrying no proprietary methodologies, no past deliverables, no writing to learn from. The engagement starts the way a real one does: a request for proposal lands by email. Everything that follows comes out of that one document and public sources. The proposal — thirty minutes The system interviews the consultant first — how they want to work, whether they know the client, who delivers, what their day rate is. Then it writes, in their own template: context, the question behind the client's question, objectives, approach and methodology, scope in and out, work packages, data room requirements, timeline, budget options, next steps. The consultant reviews and signs off. Nothing leaves unvalidated. Without AI, a proposal of that depth is several man-days.   The methodology section — six work packages mapped onto one governing question. Pricing the work — three options, fees, payment terms. A commercial proposal, not an analysis. The kick-off pack — fifteen minutes One instruction produces the whole pack: the client kick-off deck, an internal briefing note — assumptions, questions to ask, objections to expect, the landmines — and the data room request list in Excel, item by item, with what each one unblocks and the date it is needed by. In a firm, this is junior work. A solo consultant has no junior. It costs them half a day to a full day, every time. The work-package sequence agreed at kick-off, and the date after which the analysis stops waiting for client data. The First Deliverable — forty-five minutes The pattern holds whatever the sector: the research, the market sizing, the model underneath, and a deck of exhibits in the consultant's own consulting-grade template — each exhibit carrying one argument, each figure carrying its source. This is where the consultant sits as a director rather than an executant: give the instruction, let the team execute, come back and review. It is also where the system refuses to please. Here it tested the client board's own headline figure instead of repeating it, and the deck says why. Forty-five minutes for a work package that is easily five to ten days of work without it. An exhibit from the first deliverable — the sizing funnel, in the consultant's template. The economics behind it, recomputed from the model that sits under the deck. Quality control — three layers, then a loop Nothing reaches the client before three independent passes. Layer 1 — sources. Every claim and every figure traced back to where it came from, the source assessed, each one rated: verified, probable, or not established. Layer 2 — internal audit. An isolated sub-agent reviews the finished document against a purpose-built grid and rebuilds the Excel model from its own inputs. Same family of model, so it catches the obvious: the unsupported claim, the figure that contradicts the model, the page that does not carry its title. Layer 3 — dual external audit. Two models from outside the system, independent of each other and of the one that did the work. No shared bias, and two verdicts to compare. This is where the qualitative control happens. Here the two disagreed — which is itself the finding. Then the loop. The consultant asks the system to integrate the findings and audit again, and iterates. Half an hour takes a document from roughly 85% client-ready to something that needs only a final human review. The three verdicts on the first draft, as returned — including the external model that said HOLD. Everything from both engagements is published Every file is downloadable, unedited: the proposals, the contract, the kick-off packs, the deliverables, the Excel models, and the three audit reports in full — including the one that said do not ship, and the revised deliverable it forced. A second engagement, in a different sector, was launched live during the recording; it landed after the recording stopped, and its files are there too. Access the files: turboconsultant.com/unleashed Timings above are the elapsed time of each task on this run, not a benchmark. Watch the video:  Listener Offer Are you a good fit? ·       You consult on your own — strategy or management? ·       You want to scale without compromising quality? ·       Consumer AI gets you a draft, not a deliverable? Three yeses? Thirty minutes with me, free. If it's not a fit, I'll tell you. → https://calendly.com/jcl-turboconsultant/unleashed?month=2026-07  Or skip the call and start now — Unleashed listeners, neither code expires: ·      UNLEASHED_TC — Turboconsultant, €399/month instead of €499 ·      UNLEASHED_AUDIT — AI-First Audit, €200 instead of €300 Timestamps 04:02: Turboconsultant's Features and Benefits  07:32: Onboarding and Personalization Process  12:24: Demo of Turboconsultant's Capabilities 23:42: Kickoff Meeting Preparation and Work Package Execution  43:45: Quality Assurance and Iterative Improvement 50:07: Additional Features and Conclusion Links: Video permalink: https://umbrex.com/wp-content/uploads/2026/08/Jean-Christophe-Lanoix.mp4 LinkedIn: https://www.linkedin.com/in/jean-christophe-lanoix-93119018 Company website: https://www.turboconsultant.com Files from this episode: https://www.turboconsultant.com/unleashed This episode on Umbrex: https://umbrex.com/unleashed/episode-656-jean-christophe-lanoix-turboconsultant-running-a-solo-practice-on-ai-agents-end-to-end/ Unleashed is produced by Umbrex, which has a mission of connecting independent management consultants with one another, creating opportunities for members to meet, build relationships, and share lessons learned. Learn more at www.umbrex.com. *AI generated timestamps and show notes.  

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0

Watch the full episode on YouTube:We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection. We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3:And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF:Three years ago, inference engineering barely existed as a category.Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem.In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out.Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles.In this episode, Baseten's Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API.We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model.The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them.We discuss:* What happens when a 200,000-token request enters an inference system* Cache-aware routing and reusing previously computed KV cache* Why prefill and decode are increasingly handled by different GPUs* When dedicated deployments become cheaper and more reliable than shared APIs* How speculative decoding uses a smaller model to accelerate a larger one* Tool calling, structured outputs, and what LLMs actually do* What it takes to support a new open model on day zero* Grafting Kimi's vision encoder onto GLM-5.2* Retrofitting inefficient model layers with components from other architectures* Why models sometimes collapse into repeating the same token* How hardware, kernels, and race conditions create nondeterministic failures* Preserving model fidelity while making inference faster* How quantization errors can cancel each other out* Why inference optimizations still deliver gains of 20%, 100%, and 200%* How optimized serving can make a model up to 10× faster* NVIDIA Dynamo, KV-aware routing, and distributed model serving* Speculative decoding the speculative decoder* Why local AI is about making models less dumb while data-center AI is about making them less slow* Tensor, expert, and pipeline parallelism across GPUs* Hardware-aware model design, auto-tuning, and the case against mega kernels* Rubin and why inference is becoming a systems problem* Whether modern GPUs are evolving into programmable AI ASICs* Why enormous models like Kimi K3 require GB300-class hardware* Why open-source video generation still trails Veo, Kling, and other closed models* The quadratic attention bottleneck behind long-form AI video* Autoregressive video, real-time generation, and compounding quality drift* Why future video systems may combine autoregressive and diffusion architectures* Training for inference and inference for training* Continuous post-training, deployment, evaluation, and improvement loops* How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself* Why faster networking could unlock dramatically faster decoding* Continual learning, KV-cache compaction, and persistent model memoryShow Notes* How to build a day-0 API for Kimi K3* 22580: From GPT2 to Kimi3, ExplainedPhilip Kiely* LinkedIn: https://www.linkedin.com/in/philipkiely* X: https://x.com/philipkiely* Inference Engineering: https://www.baseten.co/inference-engineering/Ali Taha* LinkedIn: https://www.linkedin.com/in/aliestaha/* X: https://x.com/waterloointernTimestamps00:00:00 Introduction and the 200K-Token Prompt00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling00:11:26 Launching Production-Ready Open Models00:19:06 Model Retrofits, Failure Modes, and Nondeterminism00:28:22 Quantization and Canceling Errors00:32:15 The Race to 10× Faster Inference00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips01:10:03 Giant Models and the Limits of GPU Memory01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation01:21:47 Audio, Images, and Diffusion Models01:27:32 Training, Self-Optimizing Models, and Continual Learning01:40:06 Closing ThoughtsTranscriptIntroduction: Baseten, Waterloo Intern, and Inference EngineeringSwyx [00:00:00]: Okay, we're here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you've done, you and I have done before, as well as Ali. Welcome.Ali [00:00:15]: Pleasure to meet you.Swyx [00:00:15]: Waterloo intern.Ali [00:00:16]: Waterloo intern, always.Swyx [00:00:17]: When did you get “Waterloo intern” as a handle?Ali [00:00:19]: As a handle? Oh.Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.”Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer.Philip [00:00:30]: So we have to figure out who's gonna get the handle.Ali [00:00:33]: Well, I'll pass the torch over to the next intern.Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad.Ali [00:00:37]: To another Waterloo intern. No, bruh.Philip [00:00:39]: Yeah.Ali [00:00:39]: Intern.Swyx [00:00:40]: Intern, yeah.Ali [00:00:40]: And no.Philip [00:00:41]: You gotta get an intern from Waterloo.Ali [00:00:42]: Yeah, I've gotta get an intern from Waterloo.Swyx [00:00:44]: Right.Ali [00:00:44]: But they have to follow the path.Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it's like whoever Baseten gets from Waterloo.Ali [00:00:48]: Right.Swyx [00:00:49]: Has the title of Waterloo.Ali [00:00:50]: It stays in the ecosystem.Philip [00:00:51]: Exactly.Ali [00:00:52]: Halfway through the internship, you either get it or you're out.Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle.Ali [00:00:59]: Just say it.Philip [00:00:59]: For everybody.Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you're an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten's inference? What's the process of query through GPU model routing, balancing, all that? What is all the stuff that we don't think about?Long Context Requests, KV Cache, and Cache-Aware RoutingPhilip [00:01:26]: With a long query specifically, the first thing that I'm gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it's gonna be a lot easier for me and a lot cheaper for you. So the first thing that we're gonna look at is some cache-aware routing, where we're going to see, we probably have a number of instances, a number of replicas up serving whatever model you're hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you're doing two hundred thousand tokens, it's probably coding or a multi-turn agent or something where you would expect to have that cached. If you don't, we're gonna have to send it to a prefill worker. We've at least on certain models disaggregated prefill and decode, so you're going to have one set of GPUs that's solely going to process the input, create the KV cache, and get you your first token, and then that's going to be passed over to a separate set of GPUs, which is going to run decode. We're going to iteratively make those tokens. We're probably going to have some speculator model in front of that. I'm going to assume that you're doing coding, and because of that, our speculator model, which assumes you're doing coding, is gonna have a high draft token acceptance rate. If I'm wrong and you're asking me to summarize every Harry Potter book, it's gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?”Swyx [00:03:04]: Except Baseten doesn't charge by pennies.Philip [00:03:07]: Well, yeah, we charge. I'm assuming that we're talking about the public model APIs. If you are setting up a dedicated deployment, then yeah, it's not pennies.Public APIs vs. Dedicated DeploymentsSwyx [00:03:18]: Yeah, one of the key differentiators when I was talking with Baseten initially was that people who want very high volume just need to rent by the box, ‘cause then it's up to you to figure out how to saturate the box.Ali [00:03:31]: And more often than not, it's, like, way cheaper if you're pushing, like, millions of tokens per hour, if you just pay per hour instead of pay per token.Philip [00:03:37]: Yeah, they do. I think that we've increasingly seen a lot of demand for the pay per token APIs, just because everyone wants to try open models, and then once they find a use case that's really sticky, then they move over to dedicated.Swyx [00:03:51]: Is there a best practice on when it's time to swap over?Philip [00:03:54]: Couple reasons. Yeah, reliability, that's a big one, right?Ali [00:03:57]: Like, if they have a very specific use case, they want you to train something specifically for them, like they want their own spec dec, for instance, for their own traffic.Swyx [00:04:04]: Spec dec is speculative decoding.Speculative Decoding and Custom SpeculatorsAli [00:04:05]: Speculative decoding, yeah.Swyx [00:04:07]: You have to explain.Ali [00:04:07]: Sorry. Like, speculative decoding is like, if you have a huge model, right? And so the model is going to be generating one token at a time every single turn, every single forward pass. So we attach, like, this little, like, parasite, like this layer that goes on top of the model, and this model just has to predict. It does three very fast autoregressive forward passes, and it will predict, like, three certain tokens, and then you do one forward stage over the entire original model in order to see if those predictions were correct or not, and then you accept them or you reject them. Now, this draft model is traffic specific, so if you, like, Philip said, if you're summarizing Harry Potter books, I can train exclusively that draft model on Harry Potter books, and I can guarantee you that I'm gonna accept the three tokens every single time. And so with that case, I increase your decode speed. I wouldn't be able to provide this to you if you're a shared endpointSwyx [00:04:53]: YeahAli [00:04:53]: ‘cause I have no idea if you're doing Harry Potter, if you're doing coding, if you're doing English. We don't know. Also, there was a thing in the book that mentioned that if they really cared about a specific threshold, chapter four, I think. Do you remember that?Philip [00:05:06]: Yeah. The things that you can do is you can set a specific, like, batch sizing, a specific, like, parallelism strategy if you're trying to optimize for, like, throughput versus latency. You can. Maybe a NVFP4 quant doesn't pass your benchmarks and you wanna run a model at higher precision, you could do that. There's just a bunch of reasons why you might wanna have your own endpoint and the biggest one, of course, just being, like, you don't have to deal with someone else doing a hundred million tokens of benchmarking traffic at the endpoint when you happen to be trying to serve your users.Swyx [00:05:40]: Yeah. I think one thing that is. That is a classic journey. Like, it's people is asking the, what happens when you type Google into the browser. Tool calling, is that just, you're generating JSON or is there more complication beyond that?Tool Calling, JSON, and Structured OutputsAli [00:05:58]: Certain customers that we have, they have their own post-trained models, and so they demand a tool calling that's not just, like parse a file or go find the weather. It's something that's very specific and you have to do post-training on this. And if the post-training on the model is not good or if the quantization after the post-training to get the inference to be fast, the model will struggle reading the JSON file and reading the tool calling. But it doesn't require its own like sandbox. It's not like it's going to use that tool calling to like escape a sandbox or like it doesn't have to be contained. It can just be a normal dedicated deployment. The challenge with tool calling more and more seems to be that the companies want certain tool calling which is a very sensitive thing to train. And because you're dealing with all of the JSON outputs, if it doesn't like close the end of the request in a very certain manner, you end up with a model that did the tool calling and like the thinking and so as a result of that, it didn't see the result and just hallucinated the result as it decoded. That seems to be the most challenging thing with tool calling, not really the sandboxes model.Philip [00:06:56]: Yeah, that's a challenge on the training side and then on the inference side, there's work that you can do to scope the possible output. So we published this at this point close to two years ago, the solution to this problem which is you make a state machine and you use that to constrain the output to a specific format. So this is the structured output problem. If you remember backSwyx [00:07:27]: Yeah, the specific grammar is,Philip [00:07:29]: Yeah, exactlySwyx [00:07:30]: GML had this thing.Philip [00:07:31]: Yeah. So it's like the old-school “make sure this is only JSON”, return only JSON orSwyx [00:07:38]: YeahPhilip [00:07:38]: Grandma's gonna die type of prompts.Swyx [00:07:39]: Is it BNF grammar? At some point OpenAI had released a thing that was like, yeah, if you want to constrain your output, write BNF grammar, back as NOR.Philip [00:07:47]: In our inference system, it's just a specified output format. And you get the guarantee that your output's gonna be structured along that format. And so applying that to tool calls can like help cut down on. You can still call the wrong tool or call no tool. It doesn't solve the certainty problem but it at least solves the output structuring problemSwyx [00:08:10]: YeahPhilip [00:08:10]: Within tool calls.Swyx [00:08:12]: And MCP is just another form of tool, right.Philip [00:08:14]: Yeah, exactly.Swyx [00:08:15]: As far as there's no special thing there.Philip [00:08:16]: The thing I'm always like explaining to people is the LLM is not capable of doing anything. It's only capable of making suggestions of what to do and then if those suggestions are formatted in a certain way and applied to a system that knows what to do with them, then an action occurs.Swyx [00:08:32]: Yeah. Part of the fun stuff is, this is solved outside of tool calling too. Like in an agent loop if the output is not correct or you're right, like reasoning, tool calling was done in the reasoning trace, just be like, “Oh, I don't know what to do. Let me just try again.” And it might get there after a few tries. And on your point of training, sometimes this is harder in smaller models, so you don't have the same exact quality outputAli [00:08:56]: Right.Swyx [00:08:57]: When you just swap from a big model, right?Ali [00:08:59]: Yeah. I will say that, before, I think we need to go back to inference engineering proper.Ali [00:09:04]: But, I had expected that something would replace JSON because it's hard to stream JSON ‘cause JSON must be complete and you must have open and close brackets and everything. So it's hard to parse something or validate something while it's being streamed. So people invented all sorts of things that are like, I forget the name of some of these alternatives, but it's something like TOML, something like YAML. But JSON seems to be dominant still.Philip [00:09:30]: The JSON outputs aren't that long, right? Like you could have a long-- ‘cause tool calls also contain the arguments in them and perhaps for a certain tool you might pass like a very long argument. But my impression of the median tool call is that it's a relatively small number of tokens, right? So I would expect that speculators are generally fairly good at something as formatted as JSON. And so you would have like a pretty fast decode step there and that the streaming wouldn't be as valuable, but maybe I'm wrong about that.Ali [00:10:02]: I think you're also bounded by the software or that the model is gonna integrate with if the software is built with JSON for the tool calls or if the company that you'- if your customer says that this is how our software works and our tools are interfaced with JSON, you can ask them to like, change their software and say like, “Yeah, this is gonna be better for the model.” but like with the right training shouldn't be that much of a difference. Also more profitable if it outputs more tokens probably.Swyx [00:10:25]: Depends on your business model.Swyx [00:10:27]: It really depends. But I will say that, as a writer with like experience a lot with generated output, I do try to move from text to JSON text which is very long JSON, right? Like there's paragraphs in every field because I'm trying to structure it, right?Philip [00:10:44]: Right.Swyx [00:10:44]: I want you to first make factual statements, then make opinions then make bullet point summaries, have dates, have entity references have your sources for references, all these things. Anyway, so these are things that like I think people who really experiment with structural output have to really care about. But, let's, let's recurse up the stack a little bit. Before we started recording, you mentioned something really cool, which is that there's a lot of engineering that-- inference engineering that goes on when a new model provider releases a new model, right? So let's call it GLM-5.2, Kimi K3. I had previously assumed, especially if it's like, well, GLM 5 to 5.1 to GLM-5.2, like that you've supported them before. Is it that much work?What It Takes to Support a New Open ModelAli [00:11:26]: It's a lot of work.Swyx [00:11:28]: Yeah. Okay. So like, a lot of people, all you guys, right whenever a new model launch like, people rush to say like, “Oh, Hugging Face supports this, Fireworks supports this, Spacetime supports this,” and I'm like, “Yeah, of course we support it.” But what goes into that? What goes intoPhilip [00:11:40]: I think it's more than just support it too, right? It benefits the consumer a lot. Like I think it was with Kimi K2.5 or GLM-5.2 the latest, there was an inference war, right? X provider is at 90 tokens a second. The next day we're at 150. The nextSwyx [00:11:55]: I kinda kicked that off with the GLM-5.2.Swyx [00:11:58]: I wrote a Twitter article about. It got like half a million views,Ali [00:12:02]: Based on being numberSwyx [00:12:03]: YeahAli [00:12:04]: Or it's for something else.Swyx [00:12:05]: Yeah. Which,Ali [00:12:06]: Oh my GodSwyx [00:12:07]: Which then got everyone really excited about, hey, how can we, bend tracks a little bit further and,Philip [00:12:14]: There's a difference between support the model, as in I can make a token out of this model, and support a model, as in I have a production-ready API from this model.Philip [00:12:26]: Getting to the point of I can make a token out of this model is not that hard because generally the, open source inference engines, vLLM, SGLang of the world oftentimes even receive weights ahead of time, maintainers do, or the people making the model merge PRs to ensure support. So you generally can, just get it working on the standard open source stack without too much pain in most cases. The challenge is, every inference company is gonna have own proprietary stack. Some open source components, some in-house stuff. And for any arbitrary model, there's going to be some new stuff. Sometimes you get lucky, like K, two five to two six was, like, pretty similar.Quantization, Speculators, and Production ReadinessAli [00:13:16]: Yeah. It was pure continued post-trainingPhilip [00:13:18]: YeahAli [00:13:18]: If I remember correctly.Philip [00:13:19]: Even in those cases, there's still stuff you have to do. You have to redo the quantization work. You're taking the model from. Generally, these models are not released in NVFP4, and we want them to be in NVFP4 for maximum Blackwell compatibility. So we have to perform that quantization, and, calibrate the quantization to make sure that we're not causing any regression in the model's intelligence. And then we also have to train the speculator, as we've talked about. Generally, we have. We have ZDR, zero data retention on our model APIs, so we don't know exactly the traffic that people are sending us, but we know what's popular. We know that coding use cases are popular. We know that agents, agentic use cases are popular. So we can get public data sets that are representative of that traffic and train general speculators. Now, with speculators today, you need to train the speculator using the base model itself because you're getting hidden states out of the model from running inference on these specific prompts, and that is the training data you use to create the speculator. So there's that process which you need the real model weights for. And then there's of course just the process of, standing up all the infrastructure behind it, loading all this stuff, testing it. And then when there's a new model with a newer architecture, I think that, like, the DeepSeek models tend to be the most challenging as they have, like, the most novel architectural stuff going on, model after model. But every new model has something. Kimi K2 had. Oh, sorry, GLM-5.2 hadAli [00:14:53]: Sparse attention.Philip [00:14:54]: Yeah,Ali [00:14:54]: YeahPhilip [00:14:54]: the DSA.Ali [00:14:55]: Right. Which is brought from DeepSeek.Philip [00:14:57]: Yeah. AndAli [00:14:59]: So you can copy-paste then?Philip [00:15:01]: It kindAli [00:15:01]: I don't know how this works.Philip [00:15:02]: So, like we had to, like, build support for that into our runtime. And you're right, like it is really interesting the way that all of these open source labs borrow from each other. For example, like GLM-5.2 doesn't have vision. So something that, Haley, a guy on our team, if we could take a look at this, he, like, grafted the Kimi vision encoder onto GLM-5.2.Retrofitting Vision into GLM-5.2Ali [00:15:27]: We'll be training the projector.Philip [00:15:28]: Exactly. So if you think about, like, the encoder, there's the encoder, which is the part that looks at the image and turns it into latent information, and then there's the projector which likeAli [00:15:38]: You can say latent space. It's okay.Philip [00:15:41]: And then there's the projector that maps it onto, the model itself, and then there's the model weights. You don't wanna mess with the model weights because you run a chance of making the model dumber at something else for the purpose of giving it vision. So instead, Haley started with just a projector, which is only a handful of millions of parameters.Ali [00:16:02]: That would be, yeah.Philip [00:16:02]: Yeah.Ali [00:16:03]: Can you show the training one?Ali [00:16:04]: Like the way it groksPhilip [00:16:05]: YeahAli [00:16:06]: Very interesting.Philip [00:16:06]: And maybeAli [00:16:07]: That right therePhilip [00:16:07]: Maybe Ali, you should take it from here. You've got a betterAli [00:16:10]: Ooh, double the sandPhilip [00:16:11]: Understanding of this than I do.Ali [00:16:11]: Yeah. You can see, like, he. The way he trained this is really cool. At the beginning, he was training it using just like, “Here's a picture of a mountain. Can you describe what's in this mountain?” And that caused it just like the first, learning walls. Like here you can see this all we're trying to teach it is to translate the encoded. Like it's already taken the encoder from Kimi K. It's taken the image. It'Philip [00:16:31]: Yeah. FrozenAli [00:16:31]: FrozenPhilip [00:16:32]: With adapter.Ali [00:16:32]: Exactly.Philip [00:16:33]: Yeah.Ali [00:16:33]: So the brain is frozen and the eyes are frozen. It's just we're tryingPhilip [00:16:37]: AlignAli [00:16:38]: Interconnect between the eye and the brain, right? So the projector. And so you take the tokens and then he's like, “Oh, can you describe what's in this image?” And he's like, “Oh, it's a mountain,” or it's a person or it's a human, whatever the case is. But that didn't cause complete understanding. So he changed it such that every image was associated with a data set of questions. Like, does this image have a white male? Does this image have birds in the top corner? Does this image have a scientist in it? All of that stuff. And it would have to answer questions correctly. And using not just training on describing an image, but being able to answer question, another question, answer over time. Like you can see the grokking, which is like genuinely insane, that retrofitting vision into a large LLM can learn to that extent. And even for images that it doesn't perform well on, for instance, if you ask it a picture of like Stephen Hawking, “Who is this?” Maybe it doesn't get it, but it will say something like, “This is Albert Einstein.” Like it still understandsPhilip [00:17:25]: Close enoughAli [00:17:26]: That this is a scientist who is a man who has, some significant achievements, all that stuff. So that's like really cool.Philip [00:17:32]: Yeah. So, we've covered Hao Tian before, who the author of the LLaVA paper that did this, a while ago. And I think that's very foundational work for anyone who hasn't done vision work before.Ali [00:17:41]: Same with the CLIP and MetaCLIP, where you go from just captioning to building out questionsPhilip [00:17:47]: RightAli [00:17:47]: Off the image and how much better you can get performance.Philip [00:17:50]: Right. Right. Right. Yeah. But what's, what's so exciting about this is if you look at a model like this. Now, this is a little bit more of a research project. It's not. It got to 56% on MMLU Pro, I think. So not quite frontier. But if you're running this model, you haven't suffered any loss on your GLM-5.2 quality. If you don't have an image, it'll just behave exactly the way it used to. And ultimatelyAli [00:18:14]: Which in the inference code you literally do not include the other part, right?Philip [00:18:18]: Yeah. You would just skip the encoder if you don't have an image input.Ali [00:18:22]: Okay.Philip [00:18:22]: Just confirming.Philip [00:18:23]: YeahAli [00:18:23]: Does it affect a lot on the overall inference side? Like you're not adding much, you're adding a very small vision encoder. These are typically likePhilip [00:18:30]: They're super fineAli [00:18:31]: Less than a billion parameters, right?Philip [00:18:32]: Yeah. It's, - There's a little bit less standardization among vision encodersSwyx [00:18:37]: YeahPhilip [00:18:37]: So the support matrix can be a little bit, sparser. But overall, yeah, it's a pretty, it's a pretty minor component of the overall system. And ultimately what you get out of the system is all of a sudden you have Kimi Vision, GLM weights, and DeepSeek attention all in one model.Open Source Model Grafting and Franken-MergesPhilip [00:18:56]: And that's, I think, a lot of the power and beauty of open source, is that you can take all of these different components and combine them together into a system that's better than anyoneSwyx [00:19:05]: YeahPhilip [00:19:05]: Can be individually.Swyx [00:19:06]: People used to say that you would also do Franken-merges where you would take likePhilip [00:19:10]: YeahSwyx [00:19:10]: Layers from each model.Swyx [00:19:11]: Does anyone do that anymore?Ali [00:19:13]: Well, to your point previously when you were mentioning like, the work that goes into supporting a model when it first comes out, like GLM-5.2 or MiniMax M3 or whatever the case is. Sometimes you do have to like, you do have to switch out some things. Like, for instance, the MiniMax M3 head uses full attention, and with full attention you end up with this like insane bottleneck in spec dec ‘cause you're doing auto-regressive token generation for three tokens, and you're doing this like N squared over all of the tokens that are in your sequence. Your KV cache is like very large because it's not sparse, it's not top K. So we find it better to like, okay, we're gonna replace this, we're gonna replace this layer with a layer from another model that's using like GQA, for instance. And then just with the right training, you can get it to have the same acceptance rate. So it is very possible to retrofit layers from other models and very much needed. If a layer is like inefficient, the training just becomes the challenge, like how do you ensure that you train it properly? Which again to your earlier point is like the mesh between training and inference. As in like you need very good training in order to do fast inference. That's like, I feel like more and more becoming true.Swyx [00:20:21]: Yeah. Anything else on the support side when you say like get it to fully production ready?Loop Detection, Race Conditions, and Non-DeterminismPhilip [00:20:26]: Yeah. I think that there's also a question of just, we can test a model to a pretty extensive degree, but we're trying to get it out quickly and then you see a bunch of other people test it and you get interesting results. There was an issue with, GLM briefly where we had some like mode collapses where it would just output the same token over and over again for certain prompts on certain temperatures. Like once you expose an endpoint to the real world, there's going to be, so many more varieties of things given to it that you're able to, discover and patch things. So it's not just a, day zero process, it's then like for the first week, for the first month, if a model remains popular, like how do you both fix bugs and then continue to push the envelope on performance?Ali [00:21:21]: What do you mean you don't want your model outputting S?Swyx [00:21:24]: Is there loop detection on that stuff, by the way? It still happens like quite a lot, which is surprising.Ali [00:21:30]: We have like we, in our endpoint, like if a model was to output the same token like four plus times, we just cut the generation. We say like, “Oh, sorry, this-- Like try again,” or like we will reprocess the request. ‘Cause we know then, like if it, like if, yeah, it's four times the same token, it's probably collapsed.Swyx [00:21:45]: Yeah. Is there a way to opt out in case I really want that?Ali [00:21:48]: You want that?Ali [00:21:50]: I think there's a way that we have to handle it. I'm not exactly certain, but I feel like in certain models, like when they output something like you can imagine, like a table for instance, and so they want, they wanna draw like 12 dashes and 12 dashes. Yeah, I think there's a way for that to happen. I think we only do it on certain tokens. Like we exclude certain special characters.Swyx [00:22:07]: Yeah.Ali [00:22:07]: So we only do it on like certain like S is the most common almost. GLM-5.2Swyx [00:22:11]: OhAli [00:22:11]: And I think it was DSV 4 as well. Like you'd just have like looping issues where like you literallySwyx [00:22:17]: ItAli [00:22:17]: Just have like S.Swyx [00:22:18]: Yeah. Is there a special, something special about S? No, just randomlyAli [00:22:21]: It just seems to be the one token involved.Swyx [00:22:23]: Yeah. And it'Philip [00:22:24]: Is thereSwyx [00:22:24]: And it's only temperature 0Ali [00:22:27]: NoSwyx [00:22:27]: Even at other temperaturesAli [00:22:27]: Even at like 0.9 or whatever, it will still, it will still collapse.Swyx [00:22:30]: That's weird, right?Ali [00:22:30]: It's, it is an inference problem to be honest, like a software problem. Like oftentimes, the image you run will-- like NVIDIA will release an image for instance, and if we will upstream the changes from their latest TensorRT-LLM image into our stack, we'll find that it fixes it. Or oftentimes this will only happen in an inference engine that you're using like SGLang. But if you were to switch to vLLM, that isn't the case. So it seems to be like an extremely like deterministic software issue and not really a model issue. It's not like a weights problem. Like I'- we'll say like, “Oh, it's a problem with the quant. We did PTQ wrong,” right? But that isn't, that doesn't make sense because the same weights used with a different inference engine does not repeat the problem. And sometimes it's, the kernels that are being used in the backend have like these very subtle sometimes race conditions, where if you were to use this model hosted on one cluster, you will never get this problem.Swyx [00:23:19]: Oh my God.Ali [00:23:19]: But if you host it on a different cluster, you will. And the reason is the KV cache transfer from a node to node in that one cluster is using a slower interconnect than the node to node in another cluster. So that exposes the race, whereas in another cluster it doesn't. So then you end up just like, okay, this model is not gonna be hosted on this cluster. We're gonna host it on, another cluster because that cluster exposed that problem. But then it ends up with like, okay, is it the software? Is it the model weights or is it the hardware?Swyx [00:23:42]: There is a thing about this with temperature 0 still not being deterministic, right?Ali [00:23:46]: Right.Swyx [00:23:46]: Mostly because of hardware. Even at temperature 0 same model, you won't always get the same output.Swyx [00:23:52]: Even-- But I'm surprised by the race condition one because, I thought PyTorch was a graph that like guarantees that you at least, execute things in the right order.Ali [00:24:02]: Well, yeah, true. Like I'm not, I'm not saying that there is. Like well, you have things like PTL optimizations where like you can start a kernel before the end of the previous kernel, and that's like ‘cause you want to do that because there'sSwyx [00:24:12]: It's like pipeliningAli [00:24:12]: Expense. Exactly.Swyx [00:24:13]: Yeah.Ali [00:24:13]: But it'- But you don't do it cleanly. Like you overlap a little bit of the execution. No, it is very possible that the kernel itself, like that one block that is supposed to be running in this instance of time, that kernel itself has a race condition. For instance, like a missing barrier. Like often if you're designing a kernel and you want it to make it to be very fast, if you don't test it extensively, you'll, you'll have certain threads access data points from registers before they've been written to by other threadsSwyx [00:24:36]: YeahAli [00:24:36]: For example, because like your barrier is wrong or your synchronization was wrong. But yeah, like the testing itself is very difficult in those like, andSwyx [00:24:42]: And there's no like borrow checkerAli [00:24:45]: What does that mean?Swyx [00:24:46]: Like Rust. Like the. If you're trying to have like memory safety It sounds like a comparable problem.Ali [00:24:52]: Well, yes, but you're working in CUDA, right, NVIDIA GPUs. Like- You just need a higher level language like modular Maybe that's what modular is supposed to do. I don't know.Quantization Quality and Vendor FidelityVibhu [00:25:00]: How do you see keeping quality of the model? So you talked about all these steps of, okay, you gotta do quantization, train your own speculative decoderAli [00:25:07]: RightVibhu [00:25:07]: Run on different hardware. Looking at other model providers, okay, you kicked off a inference speed race on the consumer end. What goes into keeping quality the same across them, right? Sure, you can run benchmarksAli [00:25:22]: YeahVibhu [00:25:22]: But, like, how do you determine how much quantization are there standards? What goes intoPhilip [00:25:27]: There's a few things on quality. Most inference optimizations are lossless. KV caching, for example. You are just recomputing or preventing recomputing the same values. Speculation, of course, if a draft token is wrong, it gets rejected. The main lossy optimization is quantization. And that really comes down to, number one, data format, number two, which parts of the model you choose to quantize, which layers, and number three, like doing a lot of calibration on the quantized weights, to ensure that you're preserving all the outliers. There's other tricks that you can do, though. A big one is long context, ‘cause one thing you asked at, right at the beginning is, “Oh, what's gonna happen if I send a 200,000 token request in?” So with a long input sequence, you need to, store a lot more information. You need to process a lot more tokens. And so even if a model has a context of a certain length, you might, as an inference provider, choose to build an API with a shorter context length, and of course a full length one as well. Because if someone doesn't need the full million token context, for example, you can get them better performance. I don't know if that's exactly like quality of the model. The way that I think about quality is to what degree are we faithfully serving the original model? If you think of a golden implementation of a model that performs exactly the way the model is designed to perform, I think of quality as how close are we getting to that, 100% fidelity of the model.Philip [00:27:13]: You can also, of course, think about quality from the training side and how do you push yourself past 100%. But when I think about purely inference optimizations, it's getting faster while staying as close to that 100% fidelity mark as possible. And certainly our standard internally is that, like you should not be able to tell the difference between our API and a, official API. I think Kimi in particular does a good job of vendor benchmarking hereAli [00:27:41]: YesPhilip [00:27:41]: Where they haveAli [00:27:42]: They released an actual vendor benchmark.Philip [00:27:43]: Exactly, yeah.Ali [00:27:44]: ‘Cause they accused, some people, Amazon? There was some provider that was not doing very well on Kimi's benchmark.Philip [00:27:50]: Yeah.Philip [00:27:51]: So, with Reflect we probablyVibhu [00:27:52]: This was a long time ago, right?Philip [00:27:54]: No.Ali [00:27:54]: Yeah, like threeVibhu [00:27:55]: They alsoAli [00:27:55]: Four, five months agoVibhu [00:27:57]: This also happened with, I don't remember which model, but they pulled out quite a few, and then they started a whole chart about this. It might have beenPhilip [00:28:03]: Kimi Vendor Verifier.Ali [00:28:04]: Yeah.Philip [00:28:05]: Yeah.Ali [00:28:05]: Yeah, ‘cause you, ‘cause you'd be pissed, right? Like if you'Philip [00:28:07]: Yeah.Ali [00:28:07]: If like if I'm a consumer and I'm using like Amazon's endpoint for instance, and I've used Kimi and I'm like, “Oh my God, like this is bad,” I'm not gonna say, “Oh, Amazon quantized the model in a bad way.” I'm gonna say, “Oh, Kimi sucks.” Right?Philip [00:28:17]: Yeah.Ali [00:28:17]: So it seems like that makes sense.Philip [00:28:19]: Yeah, they care. They care.Vibhu [00:28:21]: Justifiably.Ali [00:28:21]: Yeah, justifiably.Vibhu [00:28:22]: This is probably a stupid question, but just checking, has anything improved from main quantization?Philip [00:28:28]: Yeah.Vibhu [00:28:28]: Like, is quantization always strictly worse?Ali [00:28:30]: Well technicallyVibhu [00:28:32]: NoAli [00:28:32]: It's a lossy. QuantizationPhilip [00:28:33]: YeahAli [00:28:33]: Is a lossy, it's a lossy implementation.Philip [00:28:36]: Speed improvesVibhu [00:28:36]: Speed improves.Ali [00:28:37]: It the number, likeVibhu [00:28:38]: No, I' always look for inverse scaling laws.Philip [00:28:40]: Yeah.Ali [00:28:40]: Yeah.Vibhu [00:28:40]: This is something I learned from Noam Brown, where like things that normally act in one direction sometimes do.Philip [00:28:45]: Well, technically when you run a benchmark, because these models are deterministic, sometimes your,Ali [00:28:52]: YeahPhilip [00:28:52]: NVFP4 quant is like, two basis points higher than yourAli [00:28:56]: No, it's noise. It's noise.Philip [00:28:57]: Yeah, exactly. I'm like, yeah, it's, it's within. That's why I always say within margin of error.Philip [00:29:01]: And I stopped saying that because everyone assumes that what is, well, within some margin of error, we're barely inside of that to the worst, so we're saying. But yeah, sometimes it's just like, gives you a higher output score. But like Ali said, that's noise. To my knowledge, you're not necessarily making the results better. You're just trying to, again, like keep your fidelity as close to 100% to the original model.Layer Selection, KL Divergence, and Better QuantizationAli [00:29:27]: There is, to your point, research that we did on MP. I don't know if you are able to pullPhilip [00:29:31]: YeahAli [00:29:32]: A tweet we did. One of our research interns, Joshua, I think it's a tweet on how we have 20% better quantized GLM-5.2 than NVIDIA. Essentially what we found throughout like this month research is, okay, quantization is a lossy. It's. You're compressing the data from, occupying 16 bits to occupying, four bits, for instance. And so you're losing some information, and you're trying to minimize that. And so when I say that I'm gonna quantize the model, my job becomes how do I find the layers that I can quantize, and how to find the layers to not. For instance, with image models, I don't quantize modulation layers, and I don't quantize out projections because those two are. Like out projection is what you see as the user. Modulation is what the model sees or understands. Right, exactly. And so to his paper, do you have the. It doesn't have the. Yeah. It's a long paper. I don't know if I can findVibhu [00:30:25]: If there's a part to search or it's probably in the thread.Ali [00:30:28]: It's probably in the thread.Vibhu [00:30:29]: Yeah.Ali [00:30:29]: But the long and the short is it is very possible that quantizing more of the model makes the results. Like if I have a model that I quantize layers one, five, and 10, and another model where I only quantize layers one and It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in, is that you can predict which layers are going to have quantization errors that will cancel out with each other, and you choose to quantize those layers. And so the result of doing this mathematical quantization is you end up with a model that's 20% more quantized than another provider, so you get 20% more throughput of it because there's more layers than running an NVFP4, and your quality is better than that other quant because the layers that you chose to quantize have their errors cancel out, like one layer skewed to the right one layer skewed to the left, one layer skewed to the right. Your final logits distribution is more similar to the original distribution of the model, so you have better fidelity. And so the way we proved this was with KL divergence. So instead of just scoring on the benchmarks, we scored the KL divergence between the logit distribution of the quantized model and the logit distribution of the original full precision model, and we showed that with this technique we get. If your probability distribution on the logits which token it wants to select is more of the same as the original model, you're probably gonna end up staying true to the original model. So yeah, so it seems like previously before this, it seemed like the industry was, well, the more you quantize, the worse it's gonna be, ‘cause the more loss you introduce. That's not exactly, not necessarily true. So yeah, doesn't improve it, but can cancel out.Philip [00:31:57]: I think it might be this, but reminds me a good bit about pruning where you can prune off certain layers.Philip [00:32:03]: But very interesting. Didn't know this was a whole paper you guys put out.Ali [00:32:06]: It's. Fun fact, it was originally 72 pages, this paper, and then we decidedPhilip [00:32:11]: WowAli [00:32:11]: We can't tell. We couldn't release it. So it's now 45.Swyx [00:32:15]: Still 39 pages, so very substantive. We talked about evals and all these things and, like what's possible in terms of speedup? Like it's like probably like the numberInference Speedups and BenchmarkingSwyx [00:32:25]: Thing that people do wanna care about, and it's something that you wrote about in your post. Like official API is 70 tokens per second, and you push it up to 90. Is that like a normal thing?Philip [00:32:36]: So what's cool about working in inference, the reason that I think inference is going to be a useful place to do engineering for a long time, is that if you look at highly optimized domains like, say, finance, if you're in finance, you measure how much better you got in basis points. It's like, “Oh, I got five basis points better, like twentieth of 1% better,” that's huge news because everything is so optimized. When we publish optimizations, it's 20%, it's 100% it's 200%. So there's still probably like a lot further to go, honestly. Like you'll, you'll know that inference is pretty much solved when researchers start publishing about how they got 1% faster at something.Swyx [00:33:19]: Which by the way, because I am from the finance background, in the ‘70s, that was the margin at the time. When you did quantitative finance research, you would findAli [00:33:27]: And like 20%, tens of percent.Swyx [00:33:29]: That's. Yes.Philip [00:33:29]: Yeah.Swyx [00:33:30]: And now it'Philip [00:33:31]: Tiny fractionsSwyx [00:33:32]: For those people interested, look up Andrew Lo's paper. He had a really interesting illustration of quant, stat arb, distribution, narrowing down from like those kinds of 20% differences in the ‘70s, down to nothing today, which is very cool.Philip [00:33:48]: Exactly, and we're at the beginning of the same type of thing. Now benchmarking is hard. I think anyone will tell you that, and benchmarking provider speeds is hard because there's so many variables that go into it. What hardware are you using? How much load do you have on the system? What's the exact nature of the prompts and input and output sequence lengths? All that stuff. But overall, when you start stacking these improvements, you're looking at multiples. You can look at it. The most common form, of course, is TPS, tokens per second, which is bad naming by us in the industry, ‘cause there's two tokens per second. There's tokens per second, the throughput number, and the latency number.Ali [00:34:31]: TTMT, yeah.Philip [00:34:32]: Like total tokens per second out of the, out of the GPU as a throughput number. Most people only care about tokens per second as the latency number, which we should call ITL, intertoken latency, but we don't.Philip [00:34:44]: Anyway, so you can imagine a standard API without many optimizations for a 1 trillion parameter model operating somewhere in the 30 to 50 tokens per second range for reasonable traffic profile. And we generally see the goal of, pushing to 10X that. But, not necessarily day zero, but by stacking enough optimizations, if you have, say like four optimizations, each of which doubles performance. Or sorry, three optimizations, each of which doubles performance, then you stack that up, that's an 8X gain. That's the order of magnitude that we're working with in this space. We're trying to make things substantially faster, not just go from like 70 to 90.Swyx [00:35:38]: Are you saying you've. You have done that?Philip [00:35:40]: So let's say you have as a reasonable baseline, 30 or 40 tokens per second. You can achieve 10X that. So like on GLM-5.2, if you run it unquantized, perhaps on H100s even, and you're just using an off-the-shelf inference engine with no particular optimizations, no speculator, nothing extra around like KV routing, no disaggregation, you're, you're probably, yeah, looking at that like 30 to 40. You think that's like a reasonable baseline?Swyx [00:36:12]: Right. Right.Philip [00:36:12]: To get to something like 10X, there's a lot of trade-offs that you're making. If we're running at more like a 300, 400 tokens per second range, you are using the best hardware possible. You have a optimized speculator. You have done all of your quantization work. You are Seeing a pretty high cache hit rate. You are running with a reasonably small batch size and a parallelism configuration that is tuned for latency versus throughput, but it is possible. So the spreads that you see if you, like, go on artificial analysis or you go on OpenRouter and you look at, the worst provider to the best provider, oftentimes can hit that range. 10X is of course very aggressive. It's oftentimes maybe more of a four to six times improvement. But that's the performance that makes us really excited, is when we can get these huge gains, not just go from 70 to 90 tokens.Stacking Optimizations: NVFP4, Speculation, and DisaggregationAli [00:37:19]: It's also, like, hardware dependent. Like, ifPhilip [00:37:20]: YeahAli [00:37:20]: If you have a thing where you're serving it on just, like, a node of H100s and then you throw, like, you shard the model across, like, four nodes of B200s. Like, you can definitely increase the speed with just throwing more hardware at it. Like, normalizing for the same exact hardware and the same number of GPUs.Philip [00:37:35]: Yeah. Then you're looking at, like, a two to 4X improvementAli [00:37:38]: Right. RightPhilip [00:37:38]: Depending on the inference optimizations. So yeah, it's. Some of it's, what's the call, and some of it's who's the driver.Vibhu [00:37:46]: If you break down the two to 4X, say the example is run GLM-5.2Ali [00:37:51]: YeahVibhu [00:37:51]: On B200sAli [00:37:53]: YeahVibhu [00:37:53]: Single node, right? What's, like, the cost trade-off for effort to get, like, the last bit of juice out versus what should people just think of, right?Ali [00:38:01]: Spectre quantization. Yeah.Vibhu [00:38:03]: Spectre quantization.Ali [00:38:04]: That's, that's, that's like 95%. LikeVibhu [00:38:06]: And how far does that get you? And how easy is that for the average person to do? So say right I wanna throw the weights of GLM-5.2 on a node of B200s, how easy is it to find speculative decoder- decoder model or already quantized model? How much work goes into it?Philip [00:38:23]: If you're doing it up front, it's quite a lot of work. If you're doing it today, there's going to be people who have published things that you can just, you can just grab some NVFP4 weights. You can grab a speculator. Yeah, if we're thinking about, like, what are the 2Xs we're stacking, going from, BF16 to NVFP4 is, it's not quite a 2X, right? It's like. I think it's about, like, 30 to 40%, from 16 to 8, and then another 30 to 40% multiplied from, 8 to 4. So that doesn't quite get you a 2X, but, like, roughly a 2X. Speculator, roughly a 2X. Disagg on top of that if you're able to get enough hardware and put enough traffic through it, another roughly a 2X. And then you add in some, double-digit percent increase from having just a better runtime with, the latest kernels and stuff behind it. And that's how it stacks up.Ali [00:39:21]: YeahPhilip [00:39:21]: So building each of those, like, building the, quantized weights is, for someone who really knows what they're doing, hours to days of work. Building the speculator, again, like, hours to days of work. And the, disagg setup, hours to days. Well okay, but like once you haveAli [00:39:39]: Once set up. Once set up. YeahPhilip [00:39:40]: Yeah, getting disagg working for the first time, I'm saying, of course, is very difficult.Philip [00:39:44]: The marginal implementationAli [00:39:48]: Like, if you're just grabbing, like if you are a person, like just a normal consumer who has access to, like, a node of B200s and you're wondering, “How can I just host it myself?” You don't need to quantize the model yourself. There's always gonna be, like, an open source quantized checkpoint. NVIDIA's gonna push one out if no one else does. You. Usually, the providers will have their own spec dec that they've trained as well. You don't need to train your own spec dec. You can just use that as well.Philip [00:40:09]: Yeah. Like, GLM-5.2 has its own MTP.Ali [00:40:13]: Right. Right.Vibhu [00:40:14]: What's multi token prediction?Philip [00:40:15]: Yes.Ali [00:40:16]: I'm justVibhu [00:40:16]: Can you explain that?Ali [00:40:16]: I'm just an expert.Ali [00:40:18]: I can do it for you in case I get it wrong?Vibhu [00:40:20]: No.Vibhu [00:40:21]: Yeah, you should correct if we're wrong, but their multi-token prediction can be used for self-speculative decoding.Ali [00:40:27]: I'm not sure. I'm not gonna correct that.Vibhu [00:40:28]: Okay. I'm semi-confident in thatAli [00:40:30]: Okay. YeahVibhu [00:40:30]: But someone can check. But it's useful to paint the story of, okay, not just the average person, but say a company wants to switch from serverless inference I wanna throw this up on. I wanna rent some GPUs, throw it up. These are the steps you take to do significantly faster than just put it behind vLLM.Ali [00:40:48]: Right.Vibhu [00:40:49]: I was waiting for a mention of Dynamo.Vibhu [00:40:51]: I feel like, that's supposed to be the baseline that you measure against.Dynamo, KV Routing, and Disaggregation ToolkitsPhilip [00:40:55]: I would think of Dynamo as less of a box system and more of a toolkit for building with. So when we talk about doing aware routing, when we talk about doing KV offloading, when we talk about doing, PD disaggregation, Dynamo fundamentally is. By the way, Dynamo is an open source library from NVIDIA.Ali [00:41:17]: We've done a pod with KylePhilip [00:41:18]: OkayAli [00:41:19]: Kyle Cranin.Philip [00:41:19]: Cool. So then your listeners know then that it supports all the different inference frameworks. And it is multi hardware, which is interesting.Ali [00:41:28]: But it's just a router, it's not like an optimizer layer.Philip [00:41:30]: Yeah. All it does, like, what Dynamo is good at, it is a library for moving information around your cluster, around your hardware. So if you have, KV cache on one place and you need it to be somewhere else, Dynamo coordinates NIXL for you to move that around.Philip [00:41:49]: That doesn't mean that, like, out of the box, you just say, “Pip install Dynamo,” and then you get, like, a massive performance speed up. It's more of a developer toolkit.Ali [00:42:01]: Yeah. I would have said it would. It comes with a set of defaults that you can then swap out.Philip [00:42:06]: It does. If the industry at large, I think, was, like, rolling out all of these deployments, standard, then I think it would be, like, a credible baseline. But, we've got to, we've got to benchmark against, like, what we're seeing in the wild.Speculative Decoding Methods: Medusa, EAGLE, n-Gram, and Spec-SpecVibhu [00:42:23]: I did wanna talk a little bit more about PD disagg, because that is probably, like, number three after quantized and speculative decoding. In your book though, I was just gonna pull out the book.Philip [00:42:31]: Yeah.Vibhu [00:42:32]: Like section 522 on Medusa, 523 on EAGLEPhilip [00:42:35]: YeahVibhu [00:42:36]: 524 on gram.Philip [00:42:37]: It's 55, would be disaggregationAli [00:42:42]: Yeah. Well, no, I just wanted to dwell a little bitPhilip [00:42:44]: YeahAli [00:42:44]: The other. Like, so what do you choose to include? What do you choose to not to include? Because there was all these other techniques.Philip [00:42:51]: Yeah.Ali [00:42:51]: Are these still relevant? Because I think they came out, like, a year and a half ago maybe.Vibhu [00:42:55]: Medusa is quite old.Philip [00:42:56]: Yeah, Medusa's old.Ali [00:42:58]: It was old.Vibhu [00:42:58]: But is it in the book as a good, here'sPhilip [00:43:01]: BaselineVibhu [00:43:01]: Baseline vanilla understand it?Philip [00:43:02]: Like you should know this.Vibhu [00:43:03]: Like I read the paper, I'm like, “ it makes so much sense.”Philip [00:43:05]: Yeah.Philip [00:43:05]: So with the book, I had a couple goals. One was to give people just a working vocabulary for the space as a whole, and the other was to give them some intuition about how each of these techniques works. As I mentioned in my AI Engineer talk, which is the first public addendum to this, the speculation space has moved much faster than everything else. So yeah, even at the time that I wrote the book Medusa, I very much included as a way for people to understand how the space evolved rather than what the most modern technique is. And now of course, there's DFlash, dSpark. There's, there's newer techniques even than EAGLE, although EAGLE is still very commonly used.Ali [00:43:51]: SpecSpecta.Philip [00:43:52]: Yes. Speculative decoding.Vibhu [00:43:54]: What canAli [00:43:56]: Oh, it's a paper by Tri Dao and it's like, it's doing speculative decodingVibhu [00:44:00]: HuhAli [00:44:01]: For the speculative decoder.Philip [00:44:02]: Oh, in spec- oh my God.Ali [00:44:02]: It's literally just an another. It's like, yeah, that's the most simple way to explain it, and it seems like he got trivial speed ups there. But it seems that the complexity with training, it's almost like in our mind at least, it's almost as complex as training GANs. Like it's like a very delicate balance and oftentimes you, it's just but yeah, it's literally speculative decoding on speculative decoding.Vibhu [00:44:21]: Speculative.Ali [00:44:22]: Yeah. We saw this paper.Vibhu [00:44:24]: It's interesting, right?Ali [00:44:24]: Yeah.Vibhu [00:44:24]: I wouldn't even expect it to be very particular to train, I wouldAli [00:44:29]: Right.Vibhu [00:44:29]: The naive part of me is like, okay, train speculative decoder.Ali [00:44:32]: But like, and it makes sense, like the whole idea of speculative decoding is you. It's like, it's like almost like the iPhone auto predict version but for a normal model, right? Like you're just, you're just, generating three tokens and you're like, okay, I'll do prefill on them. And so you save those three turns for your original model. Now your speculative decoder is doing three turns of auto regression, so why not just have an even smaller model?Ali [00:44:53]: The other question there is what are the size of speculators? So say forPhilip [00:44:58]: Right. It's like a billion parameters.Ali [00:45:01]: Like for MiniMax, it's. Yeah. It's like one layer. It's like one 60th of the original model usually.Philip [00:45:06]: Yeah. I think we should do a paper when we get back to the office.Philip [00:45:10]: SpeculativeAli [00:45:11]: SpeculativePhilip [00:45:11]: Decoding.Ali [00:45:13]: No, it's, it does seem like how, when do you stop? But then it also seems like if you're able to train spec-spec decode for instance, right? Like if you're able to have a small model that is accurately predicts what the intermediate speculator is gonna predict, that is able to predict what the original target model's gonna predict, then why not just use that smallest model directly, right?Vibhu [00:45:34]: Yeah. This isAli [00:45:35]: Like it seems likeVibhu [00:45:35]: Adjacent to the routing problem.Ali [00:45:36]: Right.Vibhu [00:45:36]: Yeah.Ali [00:45:36]: Right.Philip [00:45:37]: The thing with speculators is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you're running the big model on. There is a orchestration and resource competition problem inherent in that, and that is one of the constraints on speculation in general, is that draft tokens cost resources to create and cost software complexity to manage. And so if you have like infinitely recursive speculators, you add in quite a bit of that complexity on the actual implementation within the inference engine as well, not just in the training process.Vibhu [00:46:17]: I was gonna say, I would wonder if you could do similar, like distillation and pruning of, it's the same thing, it's just a model. Can we not just distill a lot of the weights, quantize the speculator, out of my domain? The question that also comes up is, this is all for big server workloads, right? How much of this applies to, say I have this MacBook, I wanna run Gemma really efficiently. Similar problems, not the same?Local AI vs. Data Center InferencePhilip [00:46:45]: Pretty different. I talked to Selo, about this on his podcast a couple weeks ago. The difference between inference engineering for the data center and for production workloads versus inference engineering for local AI, is that we start with fundamentally like different constraints and different goals. With local AI, it's how do I fit this model onto my hardware and then make it less dumb? And with data center influence, it's how do I load this model and then make it less slow? And we care about less dumb, and they care about less slow. But the local AI inference engineering ecosystem, I think has a lot for us to learn from in the data center space. They are experts in various forms of quantization, including dynamic quantization that we just don't touch, in the pruning, in the distillation, in the, layer removal. There'Ali [00:47:42]: Layer removal matters less.Philip [00:47:43]: Yeah. There'Ali [00:47:44]: No one loves pruning really.Philip [00:47:45]: Yeah. Well, but the, but they doVibhu [00:47:46]: Which is surprising, right? But that's, that's a whole different thingPhilip [00:47:48]: Just to fit something on the laptop.Ali [00:47:50]: Right.Philip [00:47:50]: So yeah, it's a, it's an interesting, it's an interesting space. Not necessarily that like their techniques make sense for us to do in the data center, because we have different resources and different goals, but more that the process as well as the openness of that field is something to, admire.Ali [00:48:12]: Yeah. Like to your point, like, certain optimizations that would. Like for instance, Turbo Quantum Sharper, like it made such huge hype on that and we did like a whole deep dive on Twitter and like said, what is it? How does it work? Why is it good or not? And it took off and it was implemented on local devices because your memory bandwidth is so slow on like a MacBook, for instance. But try putting the same thing on like an NVIDIA GPU on a B200 Turbo quant would not be. Like, it would not be used. Like, NVIDIA - Like, NVIDIA made it clear that this is not a good optimization, and we've seen it firsthand where the overhead of doing dequantization, quantization of, in the kernel itself with turbo quant kernel, each end is much slower than the time that you save from doing the bandwidth. ‘Cause on the B200s, you have like 3.5 terabytes per second. You don't need decrease the storage that much. You don't need to do, FP4 KV cache. You don't need to use a requant. There's, there's, there's better optimizations to be made. But on Edge devices, it's extremely important, it's extremely useful. So, seems to be, like, different optimizations there, but then they're all uniquely combined with like all you wanna quantize the model, you wanna do speculative decoding, like certain common prefixes with bothPhilip [00:49:18]: Principles.Ali [00:49:19]: Yeah, exactly. Exactly. Exactly.Philip [00:49:20]: They also do a lot of work on, model parallelism, especially over, heterogeneous topology, where you have, some sparks and they are wired together with, Ethernet, DGX sparks.Ali [00:49:35]: Yeah, this is the Exo Labs guys.Philip [00:49:36]: Yeah. You have, a nu

Declutter Your Chaos
381 | Declutter the Bedroom - Guided Mindful Decluttering Session Focused on Your Bedroom

Declutter Your Chaos

Play Episode Listen Later Jul 31, 2026 25:30


Hi Friends! Press play, head to your bedroom, and declutter with me. In this guided session, I'll walk you through my method, which is a mindful approach that turns decluttering into a moving meditation. We'll arrive in the space, connect with the deeper meaning behind what we're doing, breathe, explore the clutter without judgment, and then notice what becomes possible when we create a little more space. Declutter in Layers Start with one category at a time and make your way around the room: Layer 1: Trash Walk around the room and collect anything that can simply be thrown away or recycled. Layer 2: Clothes Gather clothes and separate what needs to be put away, washed, donated, or let go. Layer 3: Toiletries & Personal Items Collect skincare, makeup, hair products, medications, and other personal items and return them to where they belong. Layer 4: Give Away Look for anything you already know you're ready to donate or give to someone else. Layer 5: Things That Belong Somewhere Else Gather anything that belongs in another room. Put it in a basket or pile to relocate when you're finished. Layer 6: Things You're Still Deciding About Set aside anything you're not ready to make a decision about yet. Don't let one difficult item stop your momentum. Layer 7: Clean Once the clutter is cleared, wipe down surfaces, vacuum, dust, open a window, or do whatever makes the room feel fresh. Layer 8: Put Everything Away Return the things you're keeping to their homes and enjoy the space you created. Hope this helps you! XO, Amber   Mindful Decluttering Course: https://declutteryourchaos.thrivecart.com/declutter-without-overwhelm/ Email me with COACHING in the subject to make an appointment for FREE coaching for the podcast. amber@declutteryourchaos.com Join Me Live EVERY DAY until August 12th If you're on Instagram, come join my 30 Days of Mindful Decluttering series! I'm going live every morning, 9am Pacific,  sharing practical mindset shifts, tips and support  to help you create more space in your home and your life. Find me at @amber.cammidge on Instagram.   About Amber Amber Cammidge is the founder of Declutter Your Chaos,  host of the Declutter Your Chaos podcast, with more than 1.5 million downloads, hundreds of raving reviews, and the creator of the Mindful Method for decluttering. She helps women stop fighting their clutter by understanding the psychology behind it. Blending behavioral science, neuroscience, practical decluttering strategies, and mindfulness - informed by her training in Mindfulness-Based Stress Reduction (MBSR) - Amber teaches that your home isn't just where you live - it's where you practice becoming the person you want to be. Her work is about far more than organizing a house. It's about creating space for clarity, confidence, and a life that allows you to become the best version of you.   

Declutter Your Chaos - Minimalism, Decluttering, Home Organization
381 | Declutter the Bedroom - Guided Mindful Decluttering Session Focused on Your Bedroom

Declutter Your Chaos - Minimalism, Decluttering, Home Organization

Play Episode Listen Later Jul 31, 2026 25:30


Hi Friends! Press play, head to your bedroom, and declutter with me. In this guided session, I'll walk you through my method, which is a mindful approach that turns decluttering into a moving meditation. We'll arrive in the space, connect with the deeper meaning behind what we're doing, breathe, explore the clutter without judgment, and then notice what becomes possible when we create a little more space. Declutter in Layers Start with one category at a time and make your way around the room: Layer 1: Trash Walk around the room and collect anything that can simply be thrown away or recycled. Layer 2: Clothes Gather clothes and separate what needs to be put away, washed, donated, or let go. Layer 3: Toiletries & Personal Items Collect skincare, makeup, hair products, medications, and other personal items and return them to where they belong. Layer 4: Give Away Look for anything you already know you're ready to donate or give to someone else. Layer 5: Things That Belong Somewhere Else Gather anything that belongs in another room. Put it in a basket or pile to relocate when you're finished. Layer 6: Things You're Still Deciding About Set aside anything you're not ready to make a decision about yet. Don't let one difficult item stop your momentum. Layer 7: Clean Once the clutter is cleared, wipe down surfaces, vacuum, dust, open a window, or do whatever makes the room feel fresh. Layer 8: Put Everything Away Return the things you're keeping to their homes and enjoy the space you created. Hope this helps you! XO, Amber   Mindful Decluttering Course: https://declutteryourchaos.thrivecart.com/declutter-without-overwhelm/ Email me with COACHING in the subject to make an appointment for FREE coaching for the podcast. amber@declutteryourchaos.com Join Me Live EVERY DAY until August 12th If you're on Instagram, come join my 30 Days of Mindful Decluttering series! I'm going live every morning, 9am Pacific,  sharing practical mindset shifts, tips and support  to help you create more space in your home and your life. Find me at @amber.cammidge on Instagram.   About Amber Amber Cammidge is the founder of Declutter Your Chaos,  host of the Declutter Your Chaos podcast, with more than 1.5 million downloads, hundreds of raving reviews, and the creator of the Mindful Method for decluttering. She helps women stop fighting their clutter by understanding the psychology behind it. Blending behavioral science, neuroscience, practical decluttering strategies, and mindfulness - informed by her training in Mindfulness-Based Stress Reduction (MBSR) - Amber teaches that your home isn't just where you live - it's where you practice becoming the person you want to be. Her work is about far more than organizing a house. It's about creating space for clarity, confidence, and a life that allows you to become the best version of you.   

The aSaaSins Podcast
The Voice Layer for Emerging Markets: Aethex AI's Co-founder and CTO Ayooluwa Odemuyiwa

The aSaaSins Podcast

Play Episode Listen Later Jul 31, 2026 18:42


Most voice AI startups today are thin wrappers on someone else's models. Ayooluwa Odemuyiwa went the opposite direction and built the entire stack.In this episode, Justin sits down with the CTO and co-founder of Aethex AI, who along with co-founder Mariama is building the voice layer for emerging markets across Africa and the Middle East. This is the AI that banks, call centers, and telcos use to reach the customers a human team never could, and it's already running 17,000 calls a day.Ayooluwa takes us through her path from Caltech physics to Meta to Stanford to founder, and unpacks the decision that defines the company: in her markets, you aren't competing against other software, you're competing against the cost of human labor. That single reality is why Aethex owns everything from data collection to model serving to deployment, so they can pull every lever on cost and latency.We also get into why building for the harder market first reveals things a Western-first founder never sees, how the forward deployed engineer became their most valuable role, and why in a market this new, what you're really selling is trust.A conversation about infrastructure, conviction, and betting on the markets everyone else overlooked.

NeuroEdge with Hunter Williams
Ep. 5 | The Missing Layer | Brandon Amalani on EMF and Electrotherapy

NeuroEdge with Hunter Williams

Play Episode Listen Later Jul 31, 2026 154:55


Most people I talk to are doing everything right with peptides and hormones and still have a layer they have never touched.I brought Brandon Amalani on for exactly that reason. He has spent over 20 years at the intersection of traditional Chinese medicine, Japanese herbal systems, EMF research, and electrotherapy. He built his own PEMF device from scratch because nothing on the market met his standards. He runs Shen Blossom, one of the most respected small batch Japanese herbal operations in the country, and works with Blushield, which focuses on EMF protection.In this episode we go deep on the layer most people in the peptide world are completely ignoring. How your cells act like radio antennas. Why voltage gated calcium channels matter more than most people realize. What terrain theory actually means and why it changes everything. The three tier EMF protection system Brandon uses. And why electrotherapy might be the most underrated recovery and optimization tool available right now.We also get into why herbs act like load balancers for the body, why peptides are software running on hardware that has to be primed first, and why the most disciplined people in history never needed any of this and still outperformed almost everyone.Brandon closes with a live demo of the Arc device.For research and entertainment purposes only.Follow Brandon:Blushield (code HUNTERW for a discount): https://www.blushield.com/Shen Blossom: https://shenblossom.com/All my links: https://hunterwilliamshealth.com/links

Let's Talk Cabling!
AHL: Cat 6A Or Overkill and other questions

Let's Talk Cabling!

Play Episode Listen Later Jul 30, 2026 39:04 Transcription Available


Send us Fan MailWe answer rapid-fire questions from the field on cabling choices, troubleshooting discipline, and what real professionalism looks like on a job site. Along the way, we get practical about Cat 6A, intermittent network drops, training new techs, tool buying, and how PoE and newer AI camera demands change your planning. • when Cat 6A is actually worth the effort and when it is just spec inertia • bend radius, cable diameter, termination time, testing time, and other real install impacts • why intermittent drops demand Layer 1 focus and documented testing • a real-world example of a hidden physical fault that only appears sometimes • habits that signal a skilled tech fast: cable management, labeling, service loops, clean work • knowing when to keep troubleshooting versus calling for help • PoE, NEC safety versus performance, and what is changing in the field • AI cameras and what changes for power, bandwidth, placement, and lighting • pulling cable in older buildings by planning pathways, avoiding electrical, and communicating • training priorities for new technicians and why slow correct terminations win • tool strategy, organization, and buying quality within a budget Please go look up that LinkedIn post, find uh that video, and please donate. And if you can't donate anything financially, would you please pray for successful completion of that? Reezy make sure you join the the Bixie Technician study group guess who it's run by me yes we're doing a Zoom call next week Tuesday at 8 p.m and we're gonna answer questions for people studying for their Bixie Tech or installer questions. Be there or be square and it is free it is free. Support the showKnowledge is power!  Make sure to stop by the webpage to buy me a cup of coffee or support the show at https://linktr.ee/letstalkcabling .  Also if you would like to be a guest on the show or have a topic for discussion send me an email at chuck@letstalkcabling.com Chuck Bowser RCDD TECH#CBRCDD #RCDD

Dr. Amen Kaur - Become Narcissist Free
How to Stop Overthinking and Start Creating

Dr. Amen Kaur - Become Narcissist Free

Play Episode Listen Later Jul 30, 2026 30:09 Transcription Available


For masterclass clink hereDo you keep overthinking every idea, studying what everyone else is doing, and still find you can't start the thing you know you're meant to create?In this episode, we're breaking down why overthinking is exactly what's keeping your creativity stuck. If you're capable, driven, and full of ideas that never make it into the world, the problem isn't your discipline and it isn't a lack of originality. There are three layers of noise sitting between you and the part of you that already knows what to do, and each one has a name, science behind it, and a way through.I'm walking you through the Three Layers of Noise: why your nervous system filters every thought before you get to have it, why unfelt emotion jams the signal your intuition is trying to send, and why the voice in your head is a storyteller, not a reporter. You'll hear how Rick Rubin and Albert Einstein both described the same order of creation, and what the research on insight says about where ideas actually come from.In this episode, you will discover:Why comparison can only ever produce a copy, and what that's costing you.Layer 1: The Body. Why you can't out-think a stressed nervous system.Layer 2: The Emotional Static. Why the answer is arriving but can't get through.Layer 3: The Thought Stream. Why grinding at your desk produces nothing while ideas arrive in the shower.The order the research confirms: the idea comes first, the strategy comes second.The question that replaces "what's wrong with me."Research and books referenced:Amy Arnsten (Yale) on stress and the prefrontal cortex. Lisa Feldman Barrett on constructed emotion and interoception.Sarah Garfinkel and Hugo Critchley on interoception.John Kounios and Mark Beeman on the neuroscience of insight.Roger Beaty on the default mode network and creativity.Albert Einstein's account of his own thinking, from Jacques Hadamard, The Psychology of Invention in the Mathematical Field (1945).Rick Rubin, The Creative Act: A Way of Being.Free Masterclass: https://www.amenkaur.com/masterclassIf you want to go deeper and align the systems underneath your creativity in the right order, join my free masterclass hereJoin the Conversation:Which of the three layers of noise is loudest for you right now? Drop a comment below. I read every single one, and they help me decide what to cover next.Sending you so much love, Dr. Amen Kaur#creativity #overthinking #intuition #creativeblock #startingover #beingyouKeywords:self-discovery, personal growth, creativity, emotional intelligence, innovative thinking, overcoming fear, entrepreneurship, spiritual awakening, finding purpose, mental health awareness, navigating anxiety, unique talents, expressing creativity, building confidence, overcoming obstacles, mind-body connection, emotional healing, intuitive guidance, self-acceptance, developing resilience

Speak English with Tiffani Podcast
906 : The English Fluency Framework Every Native Knows — And You Were Never Taught

Speak English with Tiffani Podcast

Play Episode Listen Later Jul 26, 2026 33:22


Have you ever watched a native English speaker get asked a simple question like, “So how was your weekend?” …and they just flow?In today's episode of Speak English With Tiffani, you're going to learn the FluencyPanion Framework—a simple structure native speakers use (often without realizing it) to expand any topic into natural, confident English.This isn't about learning “more vocabulary.” It's about learning the structure that keeps you talking.You'll learn 3 powerful layers you can stack on any topic:Layer 1 — Three DetailsStop answering in one breath. You'll learn how to give three clear details so your answers sound full and natural (without rambling).Layer 2 — Opinion + Three ReasonsDetails help people understand you… but opinions make you interesting. You'll learn how to share what you think and support it with three reasons—so your English sounds personal and confident.Layer 3 — Personal Experience (Story) + the Five W'sStories create connection. You'll learn how to share a short personal experience and anchor it with Who, What, When, Where, Why—so your listener can truly “see” what you mean.By the end of this episode, you'll be able to take one simple topic (like your morning, your weekend, your job, or your hometown) and turn it into real, fluent English—step by step.If you want to sign up for the free English email newsletter, go to https://speakenglishwithtiffani.com/newsletter

Stab Podcasts
Special Guest: Albee Layer's Psychotic Shark Encounter

Stab Podcasts

Play Episode Listen Later Jul 25, 2026 64:19


Albee Layer and Torrey Meister got water-herded by a shark in Maui, and he's still a staunch proponent of sharks' rights. The same can't be said at Fred Pawle — to some the star, to many others the villain of our currently-running Shark Week. Keep your ears peeled for a snippet of the Great Shark Debate inside. Also in this ep we announce our new EAST surfer (presented by Kona Big Wave and Vans), dissect the financial realities of the Challenger Series, and why tradesmen actually have better surfing lives than the "pros". PS, we're throwing a US Open party with Quiksilver this Saturday night, July 25th at Bungalow in HB. Get there.

The Twenty Minute VC: Venture Capital | Startup Funding | The Pitch
20VC: OpenAI and Anthropic Threatened by Kimi? | Should the US Ban Chinese Open-Source Models | Should Openrouter Sell & Value in the Routing Layer? | Stripe Buying Paypal: What You Need to Know

The Twenty Minute VC: Venture Capital | Startup Funding | The Pitch

Play Episode Listen Later Jul 23, 2026 83:12


AGENDA: 00:04 China's Kimi and Qwen Put Frontier AI on Notice00:08 Washington Debates Whether Chinese AI Models Should Be Banned 00:17 Can America Build a Profitable Open-Weight AI Champion? 00:21 OpenRouter's Moment: Is This the Perfect Time to Sell? 00:31 Fireworks' $1.5B Raise Signals the Real AI Money Is in Infrastructure 00:39 Why Every Great AI App May Need to Build Its Own Model 00:50 Stripe's Bold Play to Buy PayPal 01:01 The AI Funding Frenzy: Why Late-Stage Venture Is Winning 01:12 Nuclear Startups Go Wild While Databricks and Stripe Stay Private 01:15 The AI Supply Chain War: TSMC, ASML, DRAM—and Nvidia's Next Move  

Epic Success with Dr Shannon Irvine
Stop Hiring B‑Players: The 3‑Layer A‑Player Interview

Epic Success with Dr Shannon Irvine

Play Episode Listen Later Jul 23, 2026 12:42


hiring slack layer b players shannon irvine
The Grit! with Chas Smith
382 - The Grit! July 23, 2026

The Grit! with Chas Smith

Play Episode Listen Later Jul 23, 2026 97:10


In today's show Chas and David peek into the future of virtual reality surfing to feel the sensation of JJF's carve at Margs and calculate what is lost in the journey, they articulate precisely where and when the backside Windshield Wiper can and should be used, long for the return of the Floater, give Meola and Layer their due applause a decade too late, learn why Gerry doesn't have time for Pipeline, and resist the return of Y2K. Plus Barrel or Nah?! Enjoy!

The Liquid Lunch Project
Can Blockchain Actually Help World Peace?

The Liquid Lunch Project

Play Episode Listen Later Jul 22, 2026 27:47


Can blockchain do more than make rich people argue on the internet? In this episode of The Liquid Lunch Project, Matt and Luigi sit down with Erai Beckmann, founder of Peace Through Trade, to talk about a much bigger use case for blockchain, AI, and cryptocurrency: global trade that creates trust, financial access, and maybe even less fighting between humans. Wild concept, apparently. Erai breaks down why trade has historically been one of the strongest tools for reducing conflict, how blockchain can open doors for people left out of traditional finance, and why Peace Through Trade is building a legal, sustainable Layer-1 Proof-of-Work blockchain with real-world use in mind.  The conversation covers regulation, energy use, stablecoins, government trust, environmental impact, and the giant gap between hype coins and technology built to last.  

BlockHash: Exploring the Blockchain
Ep. 754 Uphold | Digital Asset Services, XDC Staking and AI Trends (feat. Marcos Santillana)

BlockHash: Exploring the Blockchain

Play Episode Listen Later Jul 21, 2026 23:03


For episode 754 of the BlockHash Podcast, host Brandon Zemp is joined by Marcos Santillana, VP of Digital Asset Services, Institutional at Uphold.Marcos Santillana is VP of Digital Asset Services at Uphold, a crypto platform licensed across the US, UK and EU. He works directly with crypto foundations and protocol teams on listing, treasury management and wider ecosystem partnerships. With deep exposure to hundreds of blockchain projects assessed through a regulatory, commercial and technical lens, Marcos brings rare cross-jurisdictional fluency across US and UK frameworks.Layer-1 and Layer-2 protocols and crypto foundations can get in touch with Marcos to discuss the full range of services Uphold offers. 

HealthyGamerGG
What Everyone Gets Wrong About ADHD

HealthyGamerGG

Play Episode Listen Later Jul 20, 2026 29:43


In this episode, Dr. K reacts to a viral ADHD skit to break down what the internet frequently gets wrong about the condition. Moving beyond surface-level symptoms, he explains the neurobiology of time blindness, why high-stakes pressure paralyzes the neurodivergent mind, and how to actually build discipline by severing the link between action and impulsivity. What to expect in this episode: The Viral Assessment: A breakdown of a relatable skit covering tardiness, forgotten hobbies, and dead plants, and why focusing only on the struggles of ADHD ignores the nuance of how to actually fix them. Time Blindness Mechanics: How ADHD impairs the internal biological clock, preventing the brain from subconsciously gathering data on how long tasks take, which makes daily planning nearly impossible. Sensory Organization: Why people with ADHD must stop relying on "neurotypical" memory strategies and instead build rigid sensory systems—like the rule to "don't put it down, put it away". Suppressing Bodily Signals: How hyperfocus causes individuals to ignore basic needs like thirst or hunger until they are starving, leading to highly impulsive and unhealthy choices when they finally take a break. The Avalanche of Guilt: How a simple mistake (like forgetting a friend's birthday for seven years) instantly triggers emotional dysregulation, transforming a missed task into the deep-seated belief that "I am a fundamentally bad person". Meds vs. Therapy: Why stimulant medication only treats the primary symptoms of ADHD (Layer 1), while psychotherapy is equally necessary to heal the broken identity and low self-esteem (Layer 2) that build up over a lifetime. The High-Stakes "Lockdown": Why elevating the pressure or importance of a task—which typically motivates neurotypical people—often causes individuals with ADHD to freeze or completely shut down due to emotional overwhelm. The Secret to Follow-Through: Why the key to finishing a project is actually learning not to start it on a whim, as the exact same frontal lobe circuit that restrains impulsive excitement is responsible for deliberate, long-term discipline. Dr. K's NEW Guide to Love, Sex, & Relationships is here! Order now: https://bit.ly/4dO3x0VHG Coaching : https://bit.ly/46bIkdo Dr. K's Guide to Mental Health: https://bit.ly/44z3SztHG Memberships : https://bit.ly/3TNoMVf Products & Services : https://bit.ly/44kz7x0 HealthyGamer.GG: https://bit.ly/3ZOopgQ Learn more about your ad choices. Visit megaphone.fm/adchoices

Your Last Meal with Rachel Belle
Hannah Dasher: Eight-Layer Yellow Cake with Fudge Icing

Your Last Meal with Rachel Belle

Play Episode Listen Later Jul 16, 2026 40:21


Hannah Dasher is a country music star and author of the new cookbook Stand By Your Pan, named after her super popular cooking videos enjoyed by her millions of TikTok and Instagram followers. Hannah's house in Nashville is a 1970s time capsule and her kitchen is an avocado green dream, stocked with nostalgic, vintage cookware, including the Tupperware so many of our moms bought at parties. It's Tupperware's 80th birthday, and the iconic plastic bowls, not to mention the Tupperware parties, found their place in the American Zeitgeist thanks to a woman named Brownie Wise. Never heard of her? Most haven't. We'll tell you how she rose to Oprah-level fame in the 1940s and was erased from the history books in a span of a couple years. Plus, the company's fascinating, post-war origin story. Hannah and host Rachel Belle also talk about: The most bizarre food items Hannah has found wrapped in a napkin in her purse Why she included Naomi Judd's Possum Pie in her cookbook What foods she insists on making from scratch and which she doesn't mind getting from a box or a can. And so much more! Catch Hannah on tour!To celebrate the big anniversary, Tupperware re-released its Servalier bowls, the ones with the sunburst lids, in the classic 1970s colors, check em out!  Become a Cascade PBS member and support public media!    Watch Rachel's Cascade PBS TV show The Nosh with Rachel Belle.  Sign up for Rachel's (free!) biweekly Cascade PBS newsletter for more food musings.  Follow along on Instagram.  Order Rachel's cookbook Open Sesame.