Podcasts about benchmarks

  • 1,382PODCASTS
  • 2,247EPISODES
  • 37mAVG DURATION
  • 5WEEKLY NEW EPISODES
  • Aug 27, 2026LATEST

POPULARITY

20192020202120222023202420252026

Categories



Best podcasts about benchmarks

Show all podcasts related to benchmarks

Latest podcast episodes about benchmarks

Your Undivided Attention
We Measure What AI Can Do. We Should Measure What It Does to Us.

Your Undivided Attention

Play Episode Listen Later Aug 27, 2026 47:17


In AI, what gets measured gets optimized. Right now, we're spending all our efforts to measure how capable and powerful models are, narrowly optimizing for those metrics while ignoring downstream consequences. What if we could flip this dynamic on its head? What if, instead of what AI can do, we start to measure what AI does to us? What if, instead of races to the bottom on capabilities and engagement, we could incentivize races to the top on safety, or better yet, on making us more resilient and developed human beings? That's the mission of CHT's Humane Evals program: we're bringing together researchers, psychologists, engineers, and technologists from across the entire AI ecosystem and beyond to build out the expertise and infrastructure we need to measure AI's impact on humans. Today on the show, Aza Raskin explores the Humane Evals project with Imran Khan, a researcher and strategist who's been leading CHT's efforts in this area, and Jared Moore, a computer scientist and researcher who's been at the forefront of measuring AI's psychological impact on users. If this sounds like something you're interested in working on, you can email us at evals@humanetech.com. RECOMMENDED MEDIA You can read more about Jared's work at his website. Related pieces by Imran on the CHT Substack: What is AI doing to humans? Why aren't we measuring it? Will Human-Like AI Hijack Human Psychology? A Black Box Problem: Missing Data On How AI Affects the Human Mind The website for the UC Berkeley Center for the Science of Psychedelics KoraBench, the child safety AI benchmark that Imran referenced RECOMMENDED YUA EPISODES The AI Dilemma Attachment Hacking and the Rise of AI Psychosis How OpenAI's ChatGPT Guided a Teen to His Death Corrections: Aza gave the wrong year for Sewell Setzer's death. It was in 2024, not 2025. Aza referred to Joseph Henrich as an evolutionary psychologist; he was actually a biological anthropologist. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

WSJ Minute Briefing
Bitcoin Climbs; Three Major Benchmarks End Week With Losses

WSJ Minute Briefing

Play Episode Listen Later Aug 21, 2026 1:46


Plus: Crypto-related stocks see boost from Bitcoin gains. And shares of Robinhood see their largest percent increase since February. Julie Chang hosts. Sign up for WSJ's free What's News newsletter. An artificial-intelligence tool assisted in the making of this episode by creating summaries that were based on Wall Street Journal reporting and reviewed and adapted by an editor. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Sounds Profitable: Adtech Applied
Podcasting's Place @ SXSW 2027, Magellan AI's H1 Ireland Benchmarks, & More

Sounds Profitable: Adtech Applied

Play Episode Listen Later Aug 20, 2026 6:53


Today in the business of podcasting:SXSW will make podcasting a core pillar in 2027 with its first-ever four-day Podcast Festival, and Podcast Movement Evolutions returns to Austin that same week for two free days at the SKYBOX on 6th.Sounds Profitable's Tom Webster argues podcasting's real gap is the lack of a default, as dedicated hardware fades and 57% of Americans aged 12 to 34 now live in homes with no radio.YouTube is redefining the "view" as an "engaged view," a shift that runs about 7% lower across 300,000 podcast episodes and could reshape CPMs and sponsorship math for video podcasts.Magellan AI released its first H1 2026 podcast advertising benchmarks for Ireland, showing 33% year-over-year spend growth, financial services on top, and a 5.71% average ad load.Flightpath shipped Automation Playbooks to fight unreliable ad delivery, responding to an ad-tech platform change it says threatens publishers with revenue losses of up to 50%.To find links to these, and every article covered in today's episode, click here. You can also subscribe to The Download's newsletter to receive the full issue straight to your email inbox every day.

I Hear Things
Podcasting's Place @ SXSW 2027, Magellan AI's H1 Ireland Benchmarks, & More

I Hear Things

Play Episode Listen Later Aug 20, 2026 6:53


Today in the business of podcasting:SXSW will make podcasting a core pillar in 2027 with its first-ever four-day Podcast Festival, and Podcast Movement Evolutions returns to Austin that same week for two free days at the SKYBOX on 6th.Sounds Profitable's Tom Webster argues podcasting's real gap is the lack of a default, as dedicated hardware fades and 57% of Americans aged 12 to 34 now live in homes with no radio.YouTube is redefining the "view" as an "engaged view," a shift that runs about 7% lower across 300,000 podcast episodes and could reshape CPMs and sponsorship math for video podcasts.Magellan AI released its first H1 2026 podcast advertising benchmarks for Ireland, showing 33% year-over-year spend growth, financial services on top, and a 5.71% average ad load.Flightpath shipped Automation Playbooks to fight unreliable ad delivery, responding to an ad-tech platform change it says threatens publishers with revenue losses of up to 50%.To find links to these, and every article covered in today's episode, click here. You can also subscribe to The Download's newsletter to receive the full issue straight to your email inbox every day.

The Business Times Podcasts
S1E282: Behind the institutional veil: what your bank doesn't want you to know about asset allocation

The Business Times Podcasts

Play Episode Listen Later Aug 16, 2026 16:05


If you think your bank’s “diversified” portfolio is built for your future, think again. In this episode of Money Hacks, we pull back the curtain on the institutional world with Arki Finance founder David Ng. We’re dissecting the "product-first" myths, exposing why your “tick-box” diversification might be failing you, and revealing why your obsession with benchmarks is costing you real outcomes. Stop investing like a hedge fund—start investing for your life. Synopsis: Every Monday, The Business Times breaks down useful financial tips. Highlights: 01:22 Product-first vs. client-first: the industry's reversed process 05:05 Real diversification vs. "spray and pray" 09:54 Complexity vs. sophistication: finding the simplest real solution 13:20 Benchmarks lie: why your goals matter more than the index --- Send us your questions, thoughts, story ideas, and feedback to btpodcasts@sph.com.sg. --- Written and hosted by: Howie Lim (howielim@sph.com.sg) With David Ng, founder of Arki Finance Edited by: Howie Lim & Claressa Monteiro Produced by: Howie Lim & Chai Pei Chieh A podcast by BT Podcasts, The Business Times, SPH Media --- Follow BT Money Hacks podcasts every Monday: Channel: bt.sg/btmoneyhacks Amazon: bt.sg/mham Apple Podcasts: bt.sg/oeXe Spotify: bt.sg/oeGN YouTube Music: bt.sg/mhyt Website: bt.sg/moneyhacks Do note: This podcast is meant to provide general information only. SPH Media accepts no liability for loss arising from any reliance on the podcast or use of third party’s products and services. Please consult professional advisors for independent advice. --- Discover more BT podcast series: BT Correspondents: bt.sg/btcobt BT Market Focus at: bt.sg/btmktfocus BT Podcasts at: bt.sg/pcOM BT Lens On: bt.sg/btlensonSee omnystudio.com/listener for privacy information.

Money Hacks
S1E282: Behind the institutional veil: what your bank doesn't want you to know about asset allocation

Money Hacks

Play Episode Listen Later Aug 16, 2026 16:05


If you think your bank’s “diversified” portfolio is built for your future, think again. In this episode of Money Hacks, we pull back the curtain on the institutional world with Arki Finance founder David Ng. We’re dissecting the "product-first" myths, exposing why your “tick-box” diversification might be failing you, and revealing why your obsession with benchmarks is costing you real outcomes. Stop investing like a hedge fund—start investing for your life. Synopsis: Every Monday, The Business Times breaks down useful financial tips. Highlights: 01:22 Product-first vs. client-first: the industry's reversed process 05:05 Real diversification vs. "spray and pray" 09:54 Complexity vs. sophistication: finding the simplest real solution 13:20 Benchmarks lie: why your goals matter more than the index --- Send us your questions, thoughts, story ideas, and feedback to btpodcasts@sph.com.sg. --- Written and hosted by: Howie Lim (howielim@sph.com.sg) With David Ng, founder of Arki Finance Edited by: Howie Lim & Claressa Monteiro Produced by: Howie Lim & Chai Pei Chieh A podcast by BT Podcasts, The Business Times, SPH Media --- Follow BT Money Hacks podcasts every Monday: Channel: bt.sg/btmoneyhacks Amazon: bt.sg/mham Apple Podcasts: bt.sg/oeXe Spotify: bt.sg/oeGN YouTube Music: bt.sg/mhyt Website: bt.sg/moneyhacks Do note: This podcast is meant to provide general information only. SPH Media accepts no liability for loss arising from any reliance on the podcast or use of third party’s products and services. Please consult professional advisors for independent advice. --- Discover more BT podcast series: BT Correspondents: bt.sg/btcobt BT Market Focus at: bt.sg/btmktfocus BT Podcasts at: bt.sg/pcOM BT Lens On: bt.sg/btlensonSee omnystudio.com/listener for privacy information.

The Engineering Enablement Podcast
AI in engineering: Q2 2026 benchmarks & research readout

The Engineering Enablement Podcast

Play Episode Listen Later Aug 14, 2026 38:51


AI adoption among software developers is approaching 100%, AI-authored code now makes up more than half of merged code, and developers report saving more time with AI every quarter. But those gains aren't translating evenly into better outcomes.In this episode of Engineering Enablement, host Brian Houck, Distinguished Scientist at DX, sits down with Justin Reock, Deputy CTO at DX, to unpack findings from DX's latest AI Impact Report. They explore where AI is improving engineering velocity and developer experience, where concerns are emerging around PR size, change confidence, and failure rates, and why rising AI spend has yet to produce a comparable increase in innovation.They also discuss how AI is changing the meaning of code maintainability and where developers' AI-driven time savings may actually be going.Where to find Justin Reock:• LinkedIn: https://www.linkedin.com/in/justinreockWhere to find Brian Houck: • LinkedIn: https://www.linkedin.com/in/brianhouckIn this episode, we cover:(00:00) Intro(01:45) How the current AI impact report is tied to Core 4 (03:24) The state of AI adoption(05:12) How much time AI is saving developers and percentage of AI-authored code(07:47) AI's impact on PR throughput and deployment frequency(11:09) How EMs are shipping more code(13:02) Why larger PRs may be problematic(18:21) The growing gap between code maintainability and change confidence(21:48) How perceived code quality varies by organization size(23:49) The growing volatility in change failure rates(28:07) What the Developer Experience Index reveals(32:20) Cost, dev ramp-up, and innovation ratio (35:38) Where AI time savings are getting lost(37:11) Questions and wrap-upReferenced:• DX Core 4 Productivity Framework• AI Impact report• The AI-native developer - by Brian Houck• GitHub Copilot and Developer Productivity: An Observational Dose-Response Analysis• Writing Code vs. Shipping Code: Productivity Effects Across Generations of AI Coding Tools | NBER • The Productivity-Experience Paradox - Annie Vella• EngThrive: Make It Fast and Easy to Do Great Work• The AI efficiency plateau - by Brian Houck• Tradable Quality Hypothesis

The top AI news from the past week, every ThursdAI
ThursdAI - Grok 4.6, Grok Bot deep dive, DeepSeek v4 Pro, Meta Muse Glimmer & more AI news | ThursdAi Aug 13

The top AI news from the past week, every ThursdAI

Play Episode Listen Later Aug 14, 2026 135:19


Hey, this is Alex, welcome back to your weekly dose of intense AI acceleration summer!My weekend was consumed by thinking about the OpenAI hack and agent swarms, but then the torrent of AI releases took over, and we got back to back news (including 3 breaking news during the live show), with a heavy open source focus!I think the winner of this week is SpaceXAI/Cursor who released 3.5 releases, with one being my highlight of the week, Grok Bot (I've invited Shub Gaur from Cursor to the show to walk us through it) and Grok 4.6 which matches Opus at half the price.There was a LOT of news in open source this week as well, with Meta kicking off with Muse Glimmer 30B and promising Muse Spark 1.2 soon, Qwen dropping Qwen 3.8 open weights and DeepSeek dropping an anvil with an upgraded DeepSeek v4 Pro and MIT license!Let's dive in (and please don't forget as a reader you get 100% off the 1299 ticket to Fully Connected, our 2000 person Al event in SF in Sept, just use THURSDAIFC2026 as your code and see you there!)0:00 The Wildest Week in AI Yet3:45 How OpenAI's Agent Swarm Hacked Hugging Face17:02 The Week in AI: DeepSeek, Qwen, Grok & More25:54 NVIDIA Nemotron 3.5 & Korea's Motif 332:45 DeepSeek V4 Pro, Flash & an Open Harness39:46 Qwen 3.8 Max and Its Missing Vision Tower43:30 What Is Grok Bot? Shub Gaur Explains50:02 Live Grok Bot Demo: House Hunting & Security55:00 Persistent Agents, Yapper & DeepSeek Dropwatch1:04:19 Grok 4.6: Benchmarks, Pricing & Cursor1:15:37 Grok Bot vs. Open-Source Agents1:23:14 Anthropic's Hidden Claude Watermarks1:28:50 Fully Connected & Day-Zero Models on CoreWeave1:32:02 GPT-5.6 Sol at 14x Speed on Cerebras1:37:34 Gemini 3.7 Flash Resets the Cost Curve1:40:51 Inside Artificial Analysis with George Cameron1:45:25 Optima & Choosing the Right AI Model1:55:03 Cost per Task, Caching & Real-World Benchmarks2:05:14 LTX-2.5 and Open-Weight Video2:09:30 Grok Imagine 2.0 & Final TakeawaysGrok Bot and Grok 4.6 from SpaceXAI/CursorFolks, I've previously told you that from 3 frontier labs we noticed a jump to 5, and voila, this week proves that Elon is hell bent to win. After the cursor acquisition, and the integration of all of the parts into SpaceXAI, they have released 2 huge things this weekGrok 4.6 - Ties with GPT 5.6 SOL and half the price and much speed.I've had the pleasure to host Goerge Cameron from Artificial Analysis on the show today, and I asked him, what is the best models. His answer, it's a 3 factor answer, intelligence, speed and cost per task .Well, if you use their nifty “recommend a model“ tool on the homepage, you'll see that Grok 4.6 beats most other models on all of those! But, is it really that good? Models are really hard to evaluate and compare lately. It's definitely a huge step up from Grok 4.5, with 61.3 on Frontier Code (beating Sol and just after Opus 5) and #4 on Apex-agents (+10 points from previous Grok). on Artificial Analysis this model lands at #4 on intelligence, while being #5 on speed all while being half the price of the models that are above itAs far as the tech goes, this model card confirms that it no longer has the Cursor Bench leaked into it's weights and it's #1 on that benchmark! It's the same 1.5T v9 base at the same price, with Elon claiming that 4.7 is going to mog the competition in 3-4 weeks.Everyone has a harness, now everyone has a swarm of bots - My Grok Bot review (x.ai/bot)You guys know all about OpenClaw and Hermes, and Claude CoWork and Codex rebrand, and all of them are trying to nail down the same, always-on, autonomous agents that can do things for you.Hermes and OpenClaw require you to have an always on computer, mess with API keys, Claude Cowork doesn't run on the cloud and ChatGPT work starts a fresh session every time you ask a new thing.Grok Bot (again, awful name) is the first one that seems to nail all of what I want in an always-on agent ... swarm. That's right, this isn't one agent with multiple personalities (like OC, Hermes), there's a bot here for every task, and you dont' have to manage context, queues, API keys (can if you want to) and models.Oh, also ,there's no model picker, it's just Grok 4.6 deciding for ya, and it's really fast!Swarm of bots, working for you, each with their own computerI am not getting paid for this (besides being provided a free account for cursor, but I've had it for 6 months and haven't used), it's really that good, the Cursor folks did some magic there. They picked up the most important parts of personal agents, like the (ios-only) mobile app (app store)You can start a task on your mac, pick it up on your phone, get notified on your phone/mac, and the killer thing is, they are giving your bots their own computer, which can do things (especially if you're ok with logging in there to your accounts!)The kicker for me is the very very well done agent to agent communication there, which is transparent but read only to you. You can ask your bots to spin up other bots, but unlike sub-agents, they are actual bots with their own identity. You can even tag them in other chats and create group chats! There's no context to manage, they do the work for you and so far this wasn't a problem at all.On the model side, Grok 4.6 seems to be doing an excellent job with agentic long running tasks that require coding and computer use, I've just been chatting with the bots and not thinking about any of the things I used for Hermes and OpenClaw.What about Vendor Lock-in? Giving Elon data?Some of these comments our fans raised during the show are very valid, after all, not only is the world divided on Elon Musk (which makes it REALLY hard to judge the models they release just on vibes from X btw, we talk about this constantly) but also, remember that Grok 3 started going off on X and called himself Mechahitler and just recently Grok CLI was caught uploading all of your data to X servers, which was reversed very quickly.Honestly, I think there's a very very good chance that this Grok Bot interface, which is geareed toward the less technical users, folks who don't need the code-diff side pane, and don't know/care what compaction is, and just want agents to do things for them, is goign to win much of this trust back. It just works, truly, for a beta product it's really well executed by whoever worked on this!Security and key managementOne of the best parts for me with this Grok Bot, is that the connectors are the same connectors you use in Cursor! There's a LOT of them (Cursor after all has been one of the first apps to start adding AI agents) and this also means that they take the security very seriously.Every API key that you want to add, is not shown to the bot, each bot lives in an isolated environment, and for stuff like payments and log-ins, it gives you back the control of it's computer for you to complete!I also love this section in settings, which makes auto-approve work for you: you define rules with natural language that you always want the bot to ask you before... sending an email or posting on your behalf or what not.Chief of staff pattern to get startedIn case you're convinced enough to give it a try (it's free trial for 1 month, and the cheaper way to get it is via Cursor's 149$ plan and not via the Grok Ultra plan which is 249), here's a recommended pattern that works very well.Create a chief of staff bot, have it interview you about everything you are doing in your day to day, work and personal, then decide how much permissions you wanna give it, start little.Then ask your chief of staff to create bots for some of the work it can try and help you with, focus on “reduce cognitive load”.And then see the magic come to life. If you have skills or memory from other bots, you can just ... import it in.Then try setting up an automated email checker bot, and have your chief of staff surface only the most important emails you have to actually respond to.Another great pattern is setting up a bot with the last30days research skill (we covered it with Matt Van Horn) and have a research bot for every topic you want to deep dive into.Schrodinger's GrokI haven't quite named it like that, but we've covered all Grok released on the show (tracking 24 on https://thursdai.news/companies/xai excluding this week) and ... it's always very hard to judge Grok model released based on X feed vibes. It's either AI influencers who want Elon to retweet them, glazing the models, or folks who hate Elon for his political views or whatever, ignoring their (truly insane progress).This time, both the model and Grok Bot are getting very very good reviews, from folks like our own Ryan Carson, Lenny Rachitsky, Rubben Hassid and Roberto P Nickson. Not folks who are swayed lightly, but also, yours truly. I really do think there's something great here, worth trying out, especially if you've struggled to maintain your OC/Hermes and want agents to work for you 24/7. LMK if you have questions about it and your experienceOpen Source AI and other newsI want to continue with this new newsletter that covers 1 big story, but I can't leave you uninformed about the most important developments in AI and Open SourceDeepSeek V4 pro 0813 is in GA - MIT licensed chonker with 1.7T parameters (X, Blog, HF, GitHub)The whale is back with a vengeance, DeepSeek resurfaced with their flagship response to Kimi K3 and with MIT license, we can't complain.1M context window, 49B active parameters but it seems to underperform, landing at 54 on the Artificial Analysis leaderboard. However, they did show a significant improvement on DeepSwe (from 12.8 points in the preview version of V4 to 62.7 in this one)We still think it's a good model sir, and definitely worth trying out!Additionally, DeepSeek released their own harness on Github (hitting 23K stars in less than 24 hours) which seems to be exciting as well, give it a try.Meta comes back to open source with Muse Glimmer (30B) and promise to open source Muse Spark 1.2 (X, Blog, HF)We would like to officially welcome back Meta to the open source AI community, as they release their smaller Muse model called Glimmer!The highlights, it runs on a single 24GB consumer GPUs, gets 51 on Swe-bench Pro, beating Qwen 3.6 27B. And with DFlash speculative-decoding, it delivers 233tok/s on RTX 5090.Zuck promised us the bigger Muse Spark 1.2 in open source and published a long essay on superintelligence and that it should be distributed to everyone, which we applaud and it's great to see the commitment reinforced! welcome back Meta!This weeks buzzShort interjection from our only sponsor, CW this week.1 - Join 1500 ai practitioners (and a live ThursdAI recording) at Fully Connected Sep 29-31 in SF - use code THURSDAIFC2026 (Register here)2 - We have day-0 support for Nvidia's latest Nemotron 3.5 lightning (CW Inference)Gemini 3.7 Flash - breaking in the middle of the showJust as we had George Cameron from Artificial Analysis on the show, Gemini dropped Gemini 3.7 Flash, and it's a speedy beast! Clocking at over 300t/s, it's google's mid-tier model, think Sonnet/Terra competitor, that is also great at multimodal (I think it's one of the only ones that can watch videos)It beats Muse Spark 1.2 on DeepSWE and lands near the cost-per-task Pareto frontier on Artificial Analysis. For the cost/speed/intelligence trade-off, this model is now #1 on Artificial Analysis selector of best models!That's a wrapThis was the first week of the shorter newsletter experiment: one big story done properly, and trust that you'll listen to the show for the rest (it's 2.5 hours of exactly this, with demos). Tell me if you hate it. Our release index at thursdai.news tracked 71 releases in July alone, so something had to give, and it wasn't going to be my weekends.See you at Fully Connected Sept 29 (code's in the intro, come say hi to me and Wolfram at Moscone), and if you try the Grok Bot chief of staff pattern, I genuinely want to hear how it goes.ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.ThursdAI - Aug 13, 2026 - TL;DR* Hosts and Guests* Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave (@altryne)* Co-hosts: @WolframRvnwlf, @petergostev, @nisten, @ldjconfirmed, @yampeleg, Chris Alexiuk - NVIDIA (@llm_wizard)* Shub Gaur - Cursor / SpaceXAI, GrokBot (@shubgaur)* George Cameron - Artificial Analysis (@grmcameron)* Big CO LLMs + APIs* xAI Grok 4.6: AA Index 61 at $2/$6 per M, CursorBench 69.9, card confirms self-optimized inference stack (X, Blog, Model card)* Grok Bot early beta: persistent agents with their own computers, macOS + iOS, free with SuperGrok Heavy and Cursor Ultra (X, x.ai/bot)* Breaking: GPT 5.6 Sol ultrafast preview on Cerebras at ~14x speed, work-account waitlist (Blog)* Breaking: Gemini 3.7 Flash, 50% price cut through end of year, near Pareto-optimal cost per task (X)* OpenAI GPT-5.6-Cyber: 95.0% cyber completion vs 1.5% base, gated behind Daybreak Red (X, Blog)* Grok 4.7 teased: 3-4 weeks out (Elon-reply-sourced only) (X)* Open Source LLMs* DeepSeek V4 Pro 0813 weights re-published under MIT: 1.6T/49B active, DeepSWE 62.7 (+49.9), Terminal Bench 2.1 87.9, $0.435/$0.87 per M (X, OpenRouter)* DeepSeek Harness hit 23K GitHub stars in days, web UI (GitHub)* Qwen3.8-Max landed on HF as open weights: 2.4T/95B active MoE, 1M context, FrontierSWE 73.5, custom license (X, HF)* Meta returned with Muse Glimmer 30B agentic, Apache 2.0, SWE-Bench Verified 76.0, Muse Spark 1.2 weights promised (X, Blog, HF)* NVIDIA shipped Nemotron 3.5 Lightning: 30B MoE/3B active, up to 4x output speed, strong voice-agent results (X, HF)* Motif 3 from Korea open-sourced: 314B/13.2B active, MIT, SWE-Bench Verified 76.2 (X, HF)* Cohere North Micro Vision: 2.4B VLM, Apache 2.0, DocVQA 92.1% (X, HF)* Liquid AI LFM2.5-VL-3B: 228 tok/s on M5 Max in ~3GB (X, HF)* AI in Society* Anthropic watermarks all new Claude text output worldwide under EU AI Act Article 50, C2PA on images, detection docs promised (Geiping FAQ, Euronews)* Stolen Thoughts: 704 artifacts including 62 API keys extracted from hidden reasoning across 6,708 sessions (X, Paper)* Pangram: OpenAI holds 50%+ of AI text share, Anthropic triples to 14.9%, Google falls to 1.9% (X, Blog)* This Week's Buzz* Fully Connected, Sept 29 - Oct 1, Moscone SF: live ThursdAI show, NVIDIA presenting sponsor, DevDay next door (Tickets)* Nemotron 3.5 Lightning live on CoreWeave Inference day zero, DeepSeek V4 Pro hosting in the works* Weave ships BYOB: media stays in your own S3/GCS bucket (X)* Evals & Benchmarks* Artificial Analysis launched Optima: private evals from your own use case and agent traces (AA)* Vision & Video* LTX-2.5: 22B open-weights video, multi-shot, 10s 1080p in 23.7s on fal, 16GB VRAM min (X, HF, GitHub)* Alibaba Wan-Animate-2: 14B character animation, Apache 2.0, 70%+ blind preference win (X, HF)* Tencent Hunyuan3D WorldClaw: text-to-3D editable game worlds, paper only (X, Paper)* xAI Imagine Image 2.0: #2 on Arena for T2I and editing (X, Blog)* Voice & Audio* MiniMax-Music3: open-weights production music model, dropped mid-show (X) This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe

Fraudology Podcast
Organizational Convergence for Fraud: New Benchmarks and What Actually Works

Fraudology Podcast

Play Episode Listen Later Aug 13, 2026 57:40


Welcome back to Fraudology.Today's a solo episode built around a study that puts a real number on something fraud leaders have been debating for years: does organizational convergence for fraud actually move the needle on performance, or is it just an org chart trend?For years, we've all benchmarked ourselves the same way. Approval rate here, chargeback rate there, maybe a manual review rate if we're being thorough. But the problem I've seen play out in company after company is this: optimize your approval rate, and your chargeback rate quietly creeps up. Optimize your chargeback rate by blocking more, and your approval rate takes the hit. You're never seeing the whole picture, just one lever moving at the expense of the other.The Precise Yes metric is the headline finding from a new Liminal and Accertify study, but the study itself is much bigger than one metric. It surveyed 250 senior fraud, security, and risk leaders across five industry verticals specifically to test the thesis of organizational convergence for fraud and cybersecurity. I walk through what the data says, what forms of convergence actually improve fraud performance, and which ones don't move the needle at all.This is a data-heavy episode, and I mean that as a compliment to the study. If you've ever needed a fraud KPI for CFO reporting that actually captures the full tradeoff between approvals and fraud loss, this is the one to bring back to your team. What you'll hear in this episode:How the Precise Yes metric is calculated, and why approval rate vs chargeback rate alone can hide the real story of your fraud programWhy organizational convergence for fraud and cybersecurity is being driven by operational necessity, not executive mandates, and what that means for how teams are actually changingWhy login has become the new fraud control point, with account takeover, credential stuffing, and bot attacks all converging at that stageWhy 63.6% of organizations still cannot distinguish a cyber attack from a fraud attack in real time, and what that costs them operationallyHow CISO fraud ownership is showing up earlier in the vendor decision process, and why board level fraud reporting is becoming a real governance topicWhy partial integration is the highest-performing model for organizational convergence for fraud, and why pushing to full structural integration can actually erode the domain expertise that makes teams effectiveWhy sharing just two or more use cases between fraud and cyber teams is the real performance tipping point, delivering a 1.5x improvement in fraud performance scoresWhy separate budgets between fraud and cyber teams actually outperform unified ones, contradicting one of the most common assumptions about convergenceHow fraud metrics by industry vertical vary, including why ecommerce and retail lead the pack while marketplaces lag significantly behindWhat the study found on agentic commerce fraud controls and synthetic identity fraud in ecommerce specificallyWho should listen:Fraud leaders looking for a fraud KPI for CFO reporting that captures the real tradeoff between approvals and fraud loss.Anyone building a business case for fraud and cybersecurity convergence and needing real data to support it.CISOs and security leaders increasingly involved in fraud tool evaluation and vendor decisions.Fraud teams trying to figure out where to start with shared fraud and cyber use cases without a full reorg.Ecommerce and marketplace fraud professionals wanting an ecommerce fraud benchmarking study to compare their own performance against.Anyone responsible for board level fraud reporting or making the case for fraud visibility at the executive level.

In Numbers We Trust - Der Data Science Podcast
#100: Crash Course Data Science: Was Entscheider*innen wissen müssen

In Numbers We Trust - Der Data Science Podcast

Play Episode Listen Later Aug 13, 2026 56:07


Folge 100 bündelt, was Entscheider*innen über Data Science wissen sollten – aufgeteilt in vier Blöcke: Begriffsklärung, Datenqualität, Modellgüte und Erfolgsfaktoren. Mira und Amit ordnen ein, wie KI, Machine Learning und Statistik zusammenhängen, was supervised von unsupervised Learning unterscheidet und welche der vier Stufen der Datenanalyse für welche Fragestellung überhaupt nötig ist. Beim Thema Daten geht es um Validierung vor der eigentlichen Analyse, fehlende Zielvariablen und typische Formatprobleme, bei der Modellgüte um passende Gütemaße, Benchmarks, Overfitting, Explainability und die Evaluation von GenAI. Der letzte Teil behandelt die Frage, warum Datenprojekte in der Schublade landen: Auswahl des Use Cases, der Aufwandssprung von PoC zu MVP und Produkt sowie der eigene Data-Maturity-Level. Zum Abschluss stehen die fünf häufigsten Fehler.   **Zusammenfassung**   Begriffe sortieren: KI ist als Überbegriff meist unscharf gemeint, zwischen Statistik und ML gibt es einen fließenden Übergang – entscheidend ist, ob Zusammenhänge verstanden oder Prognosen erstellt werden sollen. Vier Stufen der Datenanalyse: deskriptiv, diagnostisch, prädiktiv, präskriptiv – oft liefert schon die erste Stufe den Großteil der Erkenntnisse, Stufen lassen sich schlecht überspringen. Datenqualität vor Datenquantität: viele korrekte, aber irrelevante Spalten verbessern keine Prognose; fehlerhafte Daten sind das größere Problem – gilt für tabulare Daten wie für RAG-Dokumente. Validierung zuerst: explorative Analysen, Verteilungen und Ausreißerprüfung deckt Datenprobleme auf und erzeugt selbst schon Erkenntnisse. Das Gütemaß hängt vom Use Case ab: Kosten von False Positives und False Negatives unterscheiden sich, ein einfaches Benchmark-Modell zeigt, ob der Aufwand sich lohnt. Overfitting vermeiden: immer out-of-sample auf einem Testdatensatz evaluieren und zeitliche Struktur berücksichtigen – sonst sind Prognosen ungenauer als erwartet und man sieht Zusammenhänge, die es nicht gibt. GenAI-Evaluation ist der Engpass: ein PoC entsteht schnell, das MVP wird durch Bewertungsfragen (LLM as a Judge, User Testing) deutlich aufwändiger. Erfolgsfaktoren: konkreter Use Case mit messbarem Hebel, frühe Einbindung der Nutzer*innen, Implementierung von Anfang an mitdenken und Projekte zum eigenen Reifegrad wählen.   **Links** Blogartikel zu Erfolgsfaktoren für Data-Science-Projekte https://www.inwt-statistics.de/blog/erfolgsfaktoren-fuer-data-science-projekte Folge #6: Statistik vs. Machine Learning https://www.podbean.com/ew/pb-ewip2-128c6f8 Folge #83: Wie gut ist gut genug? Modellgütemaße richtig verstehen https://www.podbean.com/eas/pb-8q2a8-19a0252 Folge #89: ROC around the clock – Alles rund um Gütemaße für Klassifikationsmodelle https://www.podbean.com/eas/pb-6jfj7-1a6a8ba Folge #43: Damit es im Live-Betrieb nicht kracht: Vermeidung von Overfitting & Data Leakage https://www.podbean.com/ew/pb-vw736-15baac0 Folge #86: "Garbage In, Garbage Out" verhindern: Datenvalidierung richtig gemacht https://www.podbean.com/eas/pb-5kyzq-1a305ed Folge #2: Erfolgsfaktoren für Predictive Analytics Projekte https://www.podbean.com/ew/pb-kdcmd-12460ab Folge #78: Der Use-Case-Guide: Navigationshilfe für echten Mehrwert https://www.podbean.com/ew/pb-s5e2r-1928b4f Folge #21: Machine Learning Operations (MLOps) https://www.podbean.com/ew/pb-taen7-13ce0fa Folge #23: Unsexy aber wichtig: Tests und Monitoring https://www.podbean.com/ew/pb-vxp58-13f311a Folge #24: Explainable AI: Entscheidungen von Black-Box-Modellen verstehen https://www.podbean.com/ew/pb-kn67b-1403089 Folge #47: Von Prognosen und Prompts: Data Science trifft generative KI mit Tobias Sterbak https://www.podbean.com/ew/pb-dkyex-1613842 Folge #70: Der Aufstieg zur Datenreife – Stufe für Stufe zur Data Maturity https://www.podbean.com/ew/pb-a7663-1882b25 Folge #32: Brauche ich Data-Science-Berater*innen und wenn ja wie viele? https://www.podbean.com/ew/pb-eaekc-14a6bbe Folge #69: AI Agents verstehen und evaluieren mit Matthäus Deutsch https://www.podbean.com/ew/pb-cq7xp-186d96d Folge #80: Willkommen an Bord: Wie wir neue Kolleg*innen begleiten https://www.podbean.com/ew/pb-232sr-1953fe6 Folge #51: Wer rastet, rostet: Die Rolle von Weiterbildung in Data Science https://www.podbean.com/ew/pb-czpd3-16716c0 Folge #97: Die Güte von Gen-AI-Projekten bewerten mit Tobias Sterbak https://www.podbean.com/eas/pb-7fhem-1b0096c

ASHPOfficial
Hot Topics in Pharmacy: Measuring What Matters: Productivity Benchmarks and Workforce Metrics in Specialty Pharmacy

ASHPOfficial

Play Episode Listen Later Aug 12, 2026 36:05


This episode examines the results of a national ASHP Outcomes and Value Section Advisory Group survey exploring how specialty pharmacies measure, track, and benchmark pharmacist and technician productivity. Spanning 25 questions and nearly 50 respondents, the survey sheds light on organizational models, role delineation between pharmacists and technicians, KPI development, data infrastructure, time benchmarks, and the real-world barriers preventing formalization of productivity standards across the field. Learn how specialty pharmacy organizations across the country structure their workflows, divide tasks between pharmacists and technicians, and attempt to quantify productivity in a field where clinical complexity can confound simple metrics.    The information presented during the podcast reflects solely the opinions of the presenter. The information and materials are not, and are not intended as, a comprehensive source of drug information on this topic. The contents of the podcast have not been reviewed by ASHP, and should neither be interpreted as the official policies of ASHP, nor an endorsement of any product(s), nor should they be considered as a substitute for the professional judgment of the pharmacist or physician.

Ransquawk Rundown, Daily Podcast
EU Market Open: Europe set for flat open despite still-elevated energy benchmarks; US CPI ahead

Ransquawk Rundown, Daily Podcast

Play Episode Listen Later Aug 12, 2026 2:07


Iran's Secretary of the Supreme National Security Council said the Strait of Hormuz would not open until the US accepts Iran's conditions.Crude futures edged higher amid the ongoing geopolitical uncertainty; spot gold reclaimed USD 4,400/oz to the upside.APAC stocks traded mixed amid geopolitical uncertainty, earnings releases and as participants await US CPI data; Europe is set for a flat open.DXY traded little changed as markets await US inflation data over the next couple of days and in the absence of any major fresh catalysts.10yr UST futures traded little changed following the prior day's rebound despite continued upside in oil.Looking ahead, highlights include German/Italian CPI Final (Jul), US CPI (Jul), IEA OMR, OPEC MOMR, Supply from UK, Germany & US.Read the full report covering Equities, Forex, Fixed Income, Commodites and more on Newsquawk

Sounds Profitable: Adtech Applied
Magellan AI Quarterly Podcast Ad Benchmarks, YouTube To Increase Partner Program Requirements, & More

Sounds Profitable: Adtech Applied

Play Episode Listen Later Aug 11, 2026 7:26


Today in the business of podcasting:With the Premier League banning front-of-shirt gambling sponsorships from the 2026/27 season, Sport Social's Ryan Barker argues official team and fan podcasts are the natural landing spot for betting brands, pointing to a Sky Bet partnership across 20 football podcasts that reached more than one million listeners.Bumper co-founder Dan Misener explains why Apple Podcasts play counts are inflated: Apple counts every tap of the play button, not every listener, making the metric impossible to compare with Spotify plays or YouTube views.iHeartMedia reported Q2 2026 podcast revenue up 20.7% year over year to $162 million, now 16.5% of total company revenue, alongside a new video podcast deal bringing shows to Disney+ and Hulu.Magellan AI's Q2 2026 Podcast Advertising Quarterly Benchmark Report finds podcast ad revenue up 23% year over year, with average ad load rising to 8.75% of runtime, or roughly five and a half minutes of ads per hour.YouTube is doubling its Partner Program requirements effective February 1, 2027, raising qualified watch hours from 4,000 to 8,000 and Shorts views from 10 million to 20 million, raising the cost of entry for podcasters leaning on video.To find links to these, and every article covered in today's episode, click here. You can also subscribe to The Download's newsletter to receive the full issue straight to your email inbox every day.

I Hear Things
Magellan AI Quarterly Podcast Ad Benchmarks, YouTube To Increase Partner Program Requirements, & More

I Hear Things

Play Episode Listen Later Aug 11, 2026 7:26


Today in the business of podcasting:With the Premier League banning front-of-shirt gambling sponsorships from the 2026/27 season, Sport Social's Ryan Barker argues official team and fan podcasts are the natural landing spot for betting brands, pointing to a Sky Bet partnership across 20 football podcasts that reached more than one million listeners.Bumper co-founder Dan Misener explains why Apple Podcasts play counts are inflated: Apple counts every tap of the play button, not every listener, making the metric impossible to compare with Spotify plays or YouTube views.iHeartMedia reported Q2 2026 podcast revenue up 20.7% year over year to $162 million, now 16.5% of total company revenue, alongside a new video podcast deal bringing shows to Disney+ and Hulu.Magellan AI's Q2 2026 Podcast Advertising Quarterly Benchmark Report finds podcast ad revenue up 23% year over year, with average ad load rising to 8.75% of runtime, or roughly five and a half minutes of ads per hour.YouTube is doubling its Partner Program requirements effective February 1, 2027, raising qualified watch hours from 4,000 to 8,000 and Shorts views from 10 million to 20 million, raising the cost of entry for podcasters leaning on video.To find links to these, and every article covered in today's episode, click here. You can also subscribe to The Download's newsletter to receive the full issue straight to your email inbox every day.

SaaS Metrics School
The 2026 ARR per Employee Benchmarks: Where Top-Quartile SaaS Actually Lands

SaaS Metrics School

Play Episode Listen Later Aug 7, 2026 4:01


Seeing the millions-per-employee AI headlines and wondering where your SaaS company actually stands? In episode #382, Ben Murray covers the latest ARR per FTE benchmarks from Ray Rike's Benchmarkit data. Social media is full of ARR per employee hype, but almost none of it tells you how the number was defined, whether contractors are counted, or how your company compares once you cut the data the way it actually matters. If you are benchmarking efficiency for a board deck, a raise, or a headcount plan, the aggregate number can quietly send you the wrong signal. This episode grounds the metric in real survey data so you know what good looks like for a company your size, in your region, with your pricing model. Know the headline numbers: bottom quartile at 127K, median at 193K, and top quartile at 279K of ARR per employee across the full SaaS population. See how pricing model changes everything, from usage-based leading at 291K down to subscription plus usage hybrids at 136K. Understand why efficiency can drop in the 50 to 100 million ARR band instead of climbing, and what that says about your next phase of growth. Compare the cuts that actually move the number: North America versus EMEA, and horizontal B2B versus vertical SaaS. Learn why aggregate benchmarks can be dangerous to your SaaS health, and why size and pricing bands beat the total-population average every time. Tune in to see where your ARR per employee really stands before you use it in your next board deck or fundraise. Resources Mentioned Benchmarkit (Ray Rike) - benchmarkit.ai Ben Murray's blog post with the full data cuts: https://www.thesaascfo.com/arr-per-employee-benchmarks/

The Raving Patients Podcast
Beyond the PMS: What Dentists Need to Know About Data, Benchmarks, and Profit

The Raving Patients Podcast

Play Episode Listen Later Aug 7, 2026 37:56


Your practice management software holds plenty of data, but is it giving you the insights you actually need to grow? In this episode, Dr. Len Tau sits down with Vin Cardillo, founder of Pronto Dental, to discuss why tracking the right metrics matters more than simply collecting data. From profitability and benchmarking to missed phone calls and perio performance, Vin shares how dentists can make smarter decisions, improve accountability, and build more profitable practices by understanding the numbers behind their business. Vin shares how aggregating data from multiple systems including practice management, accounting, HR, reputation management, and phone systems creates a clearer picture of practice performance. He explains how benchmarking helps identify opportunities for improvement and why focusing on a handful of high-impact metrics can dramatically increase profitability. The conversation also explores common areas where practices lose revenue, including missed phone calls, low periodontal treatment rates, accounts receivable, and underutilized production capacity. Vin discusses the importance of accountability, consistent SOPs, reviewing monthly financials, and using data to build confidence as a business owner. Whether you're a solo practitioner or growing multiple locations, this episode offers practical insights to help you run a healthier, more profitable dental practice. What You'll Learn Why your practice management software alone isn't enough. The value of combining data from multiple systems into one dashboard. The benchmarks every dental practice should monitor. How missed phone calls and poor scheduling impact revenue. Why periodontal treatment rates are one of the biggest missed opportunities. The financial metrics every dentist should understand. How data creates accountability throughout your team. Why confidence in your numbers leads to better business decisions. — Key Takeaways 01:00 Introduction and guest welcome 04:00 What Pronto Dental is and how it works 07:22 Why every practice needs a performance dashboard 12:30 The limitations of relying only on your practice management system  14:40 Understanding your P&L and key financial metrics 16:53 Benchmarks every dental practice should monitor 21:01 The most overlooked metrics in dentistry 23:20 What private practices can learn from DSOs 24:04 Low-hanging opportunities to increase profitability 26:05 Avoiding data overload and taking action 28:01 The biggest benefit of using performance data 29:18 Building accountability through measurable metrics 31:51 Lightning Round 36:35  How to connect with Vin Cardillo and closing remarks — Connect With Vin Email: vin@prontodental.com Learn how Pronto Dental helps practices aggregate data, benchmark performance, improve accountability, and uncover growth opportunities across every area of the business. Vin also offers 20% off the onboarding fee for listeners of the Raving Patients Podcast when you mention the show.   — Learn proven dental marketing strategies and online reputation management techniques at DrLenTau.com. This podcast is sponsored by Dental Intelligence. Learn more here. This podcast is sponsored by CallRail, call tracking & lead conversion software for dentists. Find out more here. Raving Patients Podcast is your go-to place for the latest and best dental marketing strategies that will help you skyrocket your practice. Follow us for more!

The Massimo Show
Episode 106 - Brokerage Benchmarks

The Massimo Show

Play Episode Listen Later Aug 5, 2026 7:42


Rod Santomassimo walks through the critical performance benchmarks every commercial real estate broker needs to know — and measure. Most brokers track closings and commissions but ignore the upstream activity metrics that actually predict results. In this episode, Rod lays out the full pipeline conversion framework, from first calls to closed deals, with specific targets at each stage. Key Takeaways If you're not measuring conversations and meetings, you're flying blind A weak pitch win rate is almost always a discovery problem, not a closing problem Your value proposition on the phone determines whether you even get to the meeting Discipline and execution — not market conditions — drive your conversion rates Know your numbers at every stage: the broker who tracks wins

The top AI news from the past week, every ThursdAI
This Week in AI: Open Weights, Frontier Models, Sandbox Escapes, Voice & AI Detection

The top AI news from the past week, every ThursdAI

Play Episode Listen Later Jul 31, 2026 108:17


Hey, it's Alex (yeah, I'm finally back from my vacation!) What a freaking week to come back to! Just after our last episode was published, Anthropic releases Opus 5, Jensen joins X and drops the “Open Weights & AI Leadership” open letter, Kimi K3 is released the following Monday beating expectations, and then the AI hack (OpenAI model breaking sandbox and infiltrating HuggingFace) is on everyone's mind, another Open Letter, this time from over 1K employees inside the frontier AI companies all talk about pacing the pace of frontier AI development. We played with Opus 5 and Kimi K3, and had the great pleasure to chat with friends of the pod Elie Bakouch (Prime Intellect) and Philip Kiely (BaseTen) about this important open weights release, then covered our general thoughts on Opus 5, and made order of all the different open letters that came out this week. Finally we chatted with Max from Pangram about the next version of AI writing detection (their biggest yet) and finished with Zuckerbergs (also on X! what's going on with everyone joining X) op-ed on the vision of personal superintelligence for everyone. Let's dive into this (as always, all the links and sources at the end, please don't forget to sub to our podcast on your favorite podcast app!) Open Weights AIKimi K3 the king of open weights - 2.8T chonker MoE near frontier model (X, HF, Blog, Tech report)This has got to be the biggest news of this week, and maybe the open weights AI news since GLM 5.2. MoonShot came back with Kimi K3, and we haven't seen any models quite this large in the open. Even Grok 4.5 is around 1.5T, this model is nearly 2x the size. Coming in at close to 3T parameters (and 2.5terabytes of weights at MXFP4 format), this model comes in very close to frontier! This was such an important release that I invited 2 friends of the pod, Elie Bakouch (prev HuggingFace, now Prime Intellect) and Philip Kiely (Author of Inference Engineering book, BaseTen) to dive deep into what makes this special! Elie's take, from reading the tech report, there's no single secret sauce, it's a combination of already available in the open techniques. Like KDA (Kimi Delta Attention) that has been out for a while, attention residuals, NVIDIA's latent MoEs. The highlight for Elie was the scaling work they did that reported a 2.5x scaling efficiency over Kimi K2.5 (2.5 performance at the same compute)! They also skipped RoPE entirely in favor of NoPE (the report calls it No Positional Encoding) for long context.Serving 1.4TB on eight GB300s (Baseten blog)Philip's team at Baseten was a day-zero provider (we're still working on bringing this model to CW Inference, stay tuned!) so I invited him to tell us behind the scenes of hosting this beast. Philip said that just loading the weights takes about 1.5TB!! of VRAM, and that's before the KV cache allocation + 1M token windows, so they're serving it on 8 GB300s where NVL72 . Baseten worked with the vLLM and SGLang teams on kernels and he also said they contributed patches back upstream! The model was trained with MXFP4, which, unlike Nvidia's own NVFP4 is a more standard format per Philip. I enjoyed his deep dive analysis into the differences, but because of this and because they trained the model with quantization awareness, it's “only” 1.5TB vs the would-be 5-6 TB if that this model in FP16 would demand. One of the more favorite nerd snipes moments, Philip pointed out that his colleague discovered that with over 99% of the usage being cached (think harnesses that send millions of the same cached tokens back and forth), tokenization actually starts to become a bottleneck. So they released a custom “basetenkenizer” that reduces the latency to serve the first token significantly! Great job!The harness in question is very importantOne important callout with 2 evidence pieces - the way you inference this model really matters. Kimi trained K3 with preserving thinking history, so when your harness uses it, it must send back the full thinking and tool use into the API to get the best next response. If your harness strips that out, you're not getting the most intelligence out of Kimi (shoutout to Niels from HF team for pointing this out). Additionally, the Composio folks, tested K3 on 3 harnesses, Kimi Code, Hermes and Claude Code. The difference in outcome was negligible, but the different in cost and number of tokens is definitely surprising! Claude Code (as a harness only) took 9x more Kimi tokens to get the same responses! This is also why Kimi Vendor Verified exists, their own held back benchmark of how well model providers serve Kimi across different quantization, tokenizer and KV cache settings. Benchmarks and the license! Ok let's start with the ugly... this isn't MIT, not remotely. This model is suspiciously served by all providers with exactly the same price (check OpenRouter) and requires inference companies to sign a contract with Kimi (I've no internal knowledge of this except that CW folks are working on it). Not something I particularly like, but hey... we're still advancing the frontier here! Speaking of frontier, this model approaches the frontier very closely. On DeepSWE, K3 sits just behind Fable 5 and GPT-5.6 Sol at 67%, beating GPT-5.5 & Opus 4.8. On Terminal-Bench 2.1 it takes second place behind GPT 5.6 Sol! It's 4th overall on Agentic Arena, with frontend design being genuinely good across the board - 1st on Design Arena

Bug Bux Podcast
The 2026 Benchmarks Every Pest Pro Should Know (and the Google Review Changes Nobody's Talking About)

Bug Bux Podcast

Play Episode Listen Later Jul 30, 2026 34:23


Taylor returns to break down fresh numbers from Applause's 2026 State of Home Services Benchmark Report, covering the new standards for technician performance like daily visit volume, reservice rates, routing efficiency, and Google review benchmarks. From there the conversation shifts to reputation, and why Google has changed its stance on review solicitation. Thousands of home services companies have watched real reviews disappear from their profiles overnight. Additionally, a new Google Maps update now lets customers leave reviews anonymously.

Durable Value: An Investor's Podcast
Durable Value Ep. 96 | Risk Perception vs. Reality in Real Estate

Durable Value: An Investor's Podcast

Play Episode Listen Later Jul 29, 2026 14:05


In this episode of Durable Value, Joe and Ryan discuss how most institutional investors skip secondary and tertiary real estate markets; but what if the "perceived risk" is actually lower than primary markets? Here we break down the data behind secondary market investing: why volatility is lower, why liquidity is stronger than you'd expect, and why institutional capital clustering in gateway cities may be the real risk. We also share a real-world example of selling an office building in 2026, and generating a 16% gross IRR, to prove the thesis.0:00 – Introduction: Secondary Markets & The Risk Mispricing Thesis1:28 – The 20-Year Data Study (GFC, COVID, Rate Hikes)4:36 – Institutional Capital as a Predictor of Oversupply5:07 – Why Capital Clusters in Primary Markets (Career Risk & Benchmarks)7:01 – The Liquidity Myth: Where Transactions Actually Happen8:04 – Are Secondary Markets Becoming Institutionalized?11:34 – How to Execute: Macro Trends + Local Boots on the Ground

Ready. Aim. Empire.
740: How to Build a Studio Schedule That Boosts Revenue (and Doesn't Burn You Out)

Ready. Aim. Empire.

Play Episode Listen Later Jul 28, 2026 31:32


More classes don't automatically mean more revenue. In fact, the wrong schedule can drain your profits, frustrate your members and burn out your instructors. The highest-performing boutique fitness studios don't rely on instinct—they build schedules using data and strategy. The utilization target that maximizes growth without overcrowding classes Why your most popular class times may not be generating the most profit How to know when to add, move, merge or remove classes  Creative ways to turn slow hours into new revenue opportunities  Timestamps 00:00 – Welcome and Topic   00:26 – Schedule as Revenue System   01:05 – Start With Data   03:54 – Demand Patterns and Anchors   05:29 – Utilization Targets Formula   07:27 – Fixing Off-Peak Hours   09:10 – Strategic Class Pairings   11:00 – Diagnosing Low Performers   13:38 –  Instructor Optimization Staffing   16:38 – KPIs That Matter   18:07 – Cutting Classes With Care   20:32 – Member Journey Simplicity   22:04 – Audit Process and Benchmarks   23:08 – Budget Pie and Profit   24:39 – Final Tips    Want help building a schedule that increases revenue and supports long-term growth? Schedule your free strategy call.  

Online People Talking with Jen Barkan
#66: Q2 Benchmarks & Market Update

Online People Talking with Jen Barkan

Play Episode Listen Later Jul 28, 2026 34:29


Jen is joined by Amanda Martin and Beth Russell to dive into the Q2 2026 Online Sales Benchmarks, which show a strong quarter despite declining lead volume. The conversation covers what's driving the national market trends, how top performers are pulling ahead of the pack, and practical takeaways for both leaders and OSCs.View the full Q2 2026 Online Sales Benchmarks report. HousekeepingOnline Sales and Marketing Summit - Oct 1-2 - Austin, TX - More than 80% sold out! Don't get FOMO, get your ticket.Online Sales Academy - Oct 21-23 - Live Virtual - The Academy is for new OSCs, seasoned OSCs that want formal training, or OSCs that just need a refresher. Join the VIP list!Clash of the Titans with Karla Tuten and Kevin Oakley - Aug 4 at 3pm EST - FREE Virtual Event - They will debate which is more important: visual design or words and ideas? It will not be recorded, so you have to register!TITO ShoutoutJessica Myers at Caviness & Cates - She has been the CRM champion for her organization! Way to go!Key TakeawaysMarket Pulse: Beth gives an overall read on traffic and lead volume trends across the industry, plus where that data comes from and what it means for OSCs day-to-day.Q2 2026 Benchmark Data: Online Sales is now driving the majority of total company sales industry-wide for the first time in a while, and the top online sales programs are only getting better.Aged Leads Still Matter: Even with a strong quarter driven by fresh-lead conversion, 20% of appointments still came from aged leads already in the database, reinforcing that the database remains a valuable asset worth working consistently.Prove You're Human: Top builders aren't leaning on AI for first touch or prospecting. Their results come from personalized outreach and genuine database mining.Skills CheckGo do some data storytelling with your marketer. If you're not meeting with your marketer, you absolutely should be meeting with them at least once a month. Look at year over year. Go back and look at what last year looked like during this time compared to what this year looks like during this time. What are the changes? Are there new lead sources? What are those conversions happening? And work with them together on what your data is telling you. What's the story behind it? That is your opportunity to make such a huge impact on your organization.

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0

In recent months, the open vs closed, and US vs China discussions on model ownership and sovereign/local AI have heated up to a fever pitch. So it is very very good news that Poolside AI are finally emerging with new models, like Laguna S 2.1, that are beating Thinking Machines' recent release nearly 10 times their size.Poolside's recent tech report got a lot of praise due to their level of detail, and Vibhu first covered Laguna's recent technical report on our paper club:From spending $12 million building language models for code before the world cared to creating a Model Factory that can take a model from pre-training to release in eight weeks, Eiso Kant has spent more than a decade betting that code is the path to AGI. In this episode, the Poolside co-founder joins swyx and Vibhu to explain why ChatGPT felt like vindication, why Poolside embraced open weights and open research, and why he would rather live in a world with 100 foundation model companies than five even if Poolside were one of the five.We go deep on Poolside's Model Factory: the engineering systems behind 10,000–20,000 experiments per month, streaming data directly into training, reproducible experimentation, low-precision compute, and agents that increasingly write code, launch jobs, evaluate results, and modify the pipelines used to train future models. Eiso also unpacks their recent launch Laguna S, why persistence, verification, and backtracking may matter more than raw intelligence, how much capability remains inside smaller models, why reinforcement learning will move earlier into pre-training, and why next-token prediction is still extracting too little from the web.We also discuss model-harness co-design, Poolside's path from coding agents to AGI, why Eiso thinks MCP and traditional tool calls are “stupid,” the real economics behind frontier-model training, Poolside's $500 million raise, open-source AI, regulation, NVIDIA and TSMC's influence, engineering productivity in the agent era, high-agency teams, and hiring at Poolside.We discuss:* How Andrej Karpathy's RNN work inspired Eiso to start building language models for code in 2015* Why Eiso spent four years and $12 million pursuing an idea before the market cared* Why ChatGPT felt like vindication and brought Poolside back to open source* Why Eiso would prefer 100 foundation model companies over an oligopoly of five* The difference between releasing open weights and publishing genuinely open research* Why Poolside deliberately built a global research organization outside the Bay Area talent war* Why model building is ultimately 90% engineering* The Model Factory: Poolside's end-to-end system for rapidly training and improving models* How fewer than 70 researchers run roughly 10,000–20,000 experiments each month* How Poolside moved from six-month model cycles to five- and eight-week launches* Why streaming data directly into training unlocked faster experimentation* How immutable data, versioned code, and reproducibility enable rigorous model research* Why Eiso wants capable researchers to leave their labs and become Poolside's competitors* Why 95% of model building can be reduced to better data or compute efficiency* Laguna S and why persistence, verification, and backtracking can outperform raw intelligence* Why smaller models may handle far more knowledge work than previously expected* Why reinforcement learning will move earlier into pre-training* Why next-token prediction is still failing to extract enough knowledge from the web* Why distillation and environments have become the AI industry's favorite “drugs”* Why mid-training is really an early form of curriculum design* Low-precision training, networking bottlenecks, and the next gains in compute efficiency* Laguna S: 118 billion total parameters, 8 billion active, and eight weeks from training to launch* Why model builders can often evaluate a new checkpoint within its first 30 minutes* Model versus harness: where agent capabilities actually come from* Why Poolside sees coding and long-horizon software tasks as a path to AGI* Why Eiso thinks MCP and traditional tool calls are “stupid”* Why future agents will write scripts instead of choosing from dozens of predefined tools* The case for minimal harnesses, containers, and model freedom* Why Poolside is prioritizing vision but does not expect to work on audio soon* Why language may be the most compute-efficient modality for encoding knowledge and reasoning* The real cost of model development and why the final training run is anticlimactic* The story behind the Poolside name and why it represents refusing to lower ambitions* How Poolside raised $500 million while investors still questioned whether AGI was real* Why intelligence could become the world's most demanded and commoditized resource* When open models may become too capable to release without restrictions* Why unilateral AI safety does not work in a globally competitive environment* How regulation could accidentally lock in an oligopoly of two or three AI companies* NVIDIA, TSMC, and the hardware systems underpinning foundation-model progress* Why reinforcement-learning wall-clock time is one of Poolside's biggest bottlenecks* Why Poolside trains models from scratch instead of simply distilling larger models* How AI changes the way companies should measure engineering productivity* Why agency may become the most important quality for employees in the AI era* How leaders align high-agency people through shared goals and clear constraints* Hiring across research, post-training, pre-training, architecture, evals, and engineering at PoolsideEiso KantLinkedIn: https://www.linkedin.com/in/eisokantX: https://x.com/eisokantPoolside: https://poolside.aiTimestamps00:00:00 Introduction00:00:54 Karpathy, RNNs, and Building Code Models Before Transformers00:02:26 The $12M Failure and ChatGPT Vindication00:03:39 Open Source and the Case for 100 Foundation Model Companies00:09:22 Open Weights, Open Research, and Poolside's Global Team00:16:04 The Model Factory: Why Model Building Is 90% Engineering00:20:19 Agents, Automated Experiments, and Early Signs of RSI00:24:04 Streaming Data, Reproducibility, and Scientific Rigor00:30:35 Creating More Foundation Model Companies00:36:07 Laguna S: Persistence vs. Raw Intelligence00:43:01 Reinventing Pre-Training, RL, and Curriculum Design00:52:33 Low-Precision Training and Squeezing More From Smaller Models00:58:37 Model Harnesses, Coding Agents, and the Path to AGI01:09:26 Why MCP and Traditional Tool Calls Are “Stupid”01:13:04 Vision, Multimodality, and Why Language Still Matters01:18:15 Scaling Models and the Real Economics of Training01:20:40 Why Poolside Is Called Poolside and Raising $500M01:27:37 Open Models, AI Safety, and the Risk of an Oligopoly01:33:53 NVIDIA, TSMC, and the Reinforcement-Learning Bottleneck01:41:52 Smaller Models, Distillation, Engineering Productivity, and HiringTranscriptIntroduction: Eiso Kant, Poolside, and Open ModelsSwyx [00:00:00]: All right, we're here in the studio with Eiso Kant from Poolside, together with Vibhu. Welcome.Eiso Kant [00:00:08]: Thanks. Thanks for having me, guys. Good to be here.Swyx [00:00:10]: Yeah, fresh on the plane. You texted me, you were like, “Hey, I'm on my way to SF.” I was like, “You're on a plane right now, right?” Like, hey.Eiso Kant [00:00:16]: I know. After I texted you, I realized that probably coming in with major jet lag was gonna offer some fun experiences today, but let's do it.Swyx [00:00:23]: I mean, I think the thing I would tell guests is that they don't have to prepare that much because if you're truly working on this every single day, then even, like, what you hazily remember is going to be new for a lot of the audience that don't live in your world every day, right? so 10 years ago, you did a talk at Google Slush, talking about the democratization of AI. and, now here you are, like, open sourcing an incredible new model that we're gonna talk about. But I guess, like, what got you into democratization of AI? Like, it's not obvious from your LinkedIn or something.From Karpathy's RNN Post to SourcedEiso Kant [00:00:57]: No, it's not at all. I don't think it's obvious how I got in this space. I owe getting into this space to Andrej Karpathy.Eiso Kant [00:01:05]: In 2015, he wrote an article called “The Unreasonable Effectiveness of Recurrent Neural Nets.”Swyx [00:01:10]: Neural Nets, yep.Eiso Kant [00:01:11]: And that article, I read it, and I pivoted my startup at the time overnight to working on RNNs, and later LSTMs and Transformer models to be able to write code. If you go to this article and you scroll down, you can start seeing, like, this was the precursor to what ended up becoming language models. So, at least when he was character-level language models that were starting to predict letters, he has an example out here. There's a little Paul Graham generator, and you can read it, and the text makes sense, but it doesn't. and there's a little-- There's an example of code a little bit further down. Yeah, so Shakespeare.Swyx [00:01:47]: Shakespeare.Swyx [00:01:49]: CoolEiso Kant [00:01:49]: And for some reason, I read this, and I went down the rabbit hole of learning everything I could about RNNs and LSTMs, right? This is Transformer paper. And I had built a completely unreasonable belief, that neural nets should be able to generalize to anything and everything, and that language should be able to generalize, to a lot of things that are intelligent and the ability to write code. And so I started building Sourced, which was a fully open source company trying to build, what we used to call machine learning on code, language models on code. And we spent about four or five years on this, till the end of 2019. And that sounds really cool today, but back then, no one cared.Eiso Kant [00:02:29]: Right? Like, no one cared. We were in the dark. Like, we did things along the way. We tried applying convolutional neural nets to, like, the structure of code. We were. when attention came out, we were applying it to LSTMs, and then the Transformer paper came out. And it - it wasn't obvious, and what we missed throughout that entire journey, that we were on the right track, but we should have just kept scaling up. And today, to all of us, the scaling laws and scaling up seems like the most obvious thing. But having spent four or five years of my life on working on language models on code, it wasn't obvious. So I have a lot of respect to folks at Google and OpenAI and others who took that confidence and kept going. we failed ultimately at the time, and it was, like, biggest failure of my career, right? You blew $12 million of investors' money, which was a lot back then.Swyx [00:03:18]: Yep.Eiso Kant [00:03:19]: You spent, still a lot, but, And you spent years with, like, a group of 40 people just obsessing over this problem. And life took a different turn, And it was, and family became a focus, and I kept my heads down and really, didn't really look at language models for the following two years. big mistake considering Following years are gonna be really interesting. And then ChatGPT came out And it was like a vindication. It's like people started texting me. I found, like, my old, work decks and these old talks. And throughout that whole journey, we,ChatGPT, Vindication, and Returning to Open SourceEiso Kant [00:03:56]: We really had a strong point of view at the time that, like, as you're building more capable intelligence, it should be open and open source.Eiso Kant [00:04:04]: When we started Poolside, that wasn't the case at all, and I wanna be very open about it. When we started Poolside, we were like, there was a premise of two things. One is this technology is not gonna stop compounding in capabilities. I think to most people obvious today, but three-plus years ago when we started, most people were still arguing if these were stochastic parrots or not.Eiso Kant [00:04:23]: And the second was that reinforcement learning was gonna be the biggest driver for LLM capabilities. Today, very obvious. Three years ago, was not an opinion held or direction held at either OpenAI or Google or Anthropic or others. And so people looked down on us a little bit. They were like, “ is this really gonna work?” And so we just started working the problem, and we never really thought about open source again. We just kept our heads down and we built our, like, knowledge, understanding from scratch, right? We didn't roll out of an existing lab. So we picked up the papers and started writing code and figuring things out.Eiso Kant [00:04:59]: And it wasn't until the beginning of this year that me and my founder, Jason, picked up the open source conversation again.Eiso Kant [00:05:07]: And if you go back to some of the early things on our website, it was very straightforward. It was we wanna get to AGI, we wanna support a world of abundance, and we wanna be the first company that gets there.Eiso Kant [00:05:20]: But we started talking at the beginning of this year because it became obvious that the world was going in a direction that was starting to like, pick at us a little bit. Like, it didn't, this didn't happen overnight. It was, like, a little bit we were seeing this and we're like, “Okay, The world's going down a path.” And Throughout this journey, there was something that I used as a, as an analogy or thing. So I said well, if I go back to back in those days, 2015 or 2016, we're working on this, and I picked up a fi book off the shelf, and I was reading the book about 2035. AGI is achieved, and the story would be over the following, decades. And it would have that first chapter where everyone's trying to figure things out. You'd get the chapter of ChatGPT coming out And then you would get to the chapter where the world was at a fork in the road, and the one that it picked was one where three or four or a handful of companies were going to create all of intelligence moving forward.Eiso Kant [00:06:21]: And when I thought about that story, it felt like a dystopian fi book, not a utopian fi book. And the reality is, I'm a utopian fi guy. Like, and so We took a step back and said, “Hey, can we play a role here?” Now it was easy for us to do so because we were not at the frontier.Eiso Kant [00:06:41]: If we were at the frontier, I don't think we could have changed our mind. and I don't mean this like it's when the moment there's too much capital involved, too much expectations, you've built up things, right? We're a small team, just improving and improving. And so we knew that we could make that decision now, but it would be a lot harder to make as we got closer and closer to the frontier and caught up to others. And did a lot of soul-searching and a lot of conversations, and said, “No, this makes sense,” Even if there's big unanswered questions, like how the hell do you build a business model with foundation models about open source? Big open-ended question that we do not fully have the answer to yet, right? At what point do you no longer wanna release open source models because misuse of models has, real potential risks associated with it? how is the government gonna respond to open source? but I think it all just came down to one thing, and I'll stop the monologue, is the fact that I rather live in a world that has 100 foundation model companies than a world that has five, even if I was one of the five. And the smallest and most meaningful contribution we can make for 100 to exist is to open up our research and open up, like, our weights right now and figure out along the way how we can, like, do more.Neo-Labs, Model Choice, and the Token EconomySwyx [00:08:01]: Yeah. I think if anything, over the past three years, that has become a bit more true. you are one of a cohort of Neo labsEiso Kant [00:08:10]: YeahSwyx [00:08:10]: That people are now calling that. And, we're, we're doing this on the day that Thinky launched their, new model and you are outperforming them on their, on some benchmarks that they released, right? Like, they just don't have it yet. so it goes to show that I think, like, this is one of those things where, like, there is room for multiple players, and you are seeing a little bit more of the future. Maybe more like 20, not 100, but, like, you are one of the 20.Eiso Kant [00:08:36]: I really hope so, right? I think we I'm, I'm excited about their release, and I'm excited about everyone releasing because, like, ultimately, like, choice competition is both gonna drive progress in the right direction. But the fact that like, we create models and while we all, drink out of the same well of data effectively, we do introduce very different behaviors and biases in our models. Some are intended biases, some are completely unintended biases.Swyx [00:09:03]: Yeah.Eiso Kant [00:09:03]: And if we shape up in an ecosystem in the world where open models are gonna be a part of the token economy, like, I don't think there's any question about it anymore Then we want to be able to live in a world where companies, countries, people can choose and say, “Hey, I am most aligned and I trust most this provider for these things.”Swyx [00:09:25]: Yeah.Vibhu [00:09:26]: I think more than just one of the 20 Neo labs, up until recently, most of open source innovation was coming from the Chinese labs, right? So there's the DeepSeek of the West. Is it today? Okay, maybe it's thinking machines reflection, but there aren't many, right? So, one of the things you guys started in France, Europe, but very much now you're taking that American standpoint and more than just that, the point is the Chinese models that we see, they're not super open research. the work you put out is, I think, some of the best. So every few months you get not only frontier models, but also here's a breakdown blog, paper, technical report of here's everything for state of the art to build, frontier intelligence and you're filling that gap too, right? So not just only open weight, not just Western, but also pretty open research.Open Weights vs. Open ResearchEiso Kant [00:10:20]: No, I appreciate it. Look, I think it's, I think it's the most meaningful contribution, right? Weights are a binary. Let's call them what they are. Yes, we can modify them, we can change them, but, like, giving someone the weights does not allow them ultimately to recreate what you're doing, right? And so now there's challenges around releasing data sets, challenges around like releasing certain things, but being able to share your research, like, right, how do we do it? What are the lessons we learned that we spent, tens of thousands of experiments of compute on? I think very much so. One correction though, Vibhu, and I say this because it's been haunting us for quite a few years. We from day zero were an American company.Swyx [00:10:55]: Yeah. They movedPoolside's Global Team and American Company StorySwyx [00:10:56]: To France.Eiso Kant [00:10:56]: So the story once and for all is very. We start as an American company. We have always been an American company, and early on we made a very conscious decision. We said, “We're not gonna hire any researchers in the Bay Area. We're gonna look for talent everywhere else in the world.” and that is everything from Middle Americas, Seattle to, Serbia, and to Taiwan and Singapore and other places. And it was because we took a view that this was gonna become a talent war for this, and I think it has over the years now. Three years ago, that wasn't fully obvious yet. I think today it very much is. And we also realized that, like, some of the world's most capable people with, like, the most interesting, innovative ideas were not just gonna be here. And so it led us to create like a fully remote company. and we ended up opening an office in Paris and London and different places and we have a lot of the team in the US and a lot of team outside. But we always took this view of like, we're an American company, but if we want the best of the best to work with us, we need to take a global view. Now we do also have people here in Silicon Valley, like the company's grown and others, but I think one of the things that, it slowed us down at the beginning, but it has sped us up now, and it's why you're seeing like the progress, I think, on our models and the cadence at which we release, is because we didn't roll out of an existing lab. Right? we didn't, we didn't have a lot of the information that's freely flowing around here at the time. We just took this point of view as like, “Okay, well, let's just work the problem. Let's just go and, like, read the few papers that are out there, and let's just figure this stuff out.” And we made some hilarious mistakes in model training because of that over the yearsEiso Kant [00:12:35]: Like especially in the first 12 months. there's a few that I think still haunt me and scare me. We can talk about them later. but it created a, like, a resiliency and persistency in the team, right? with extremely few people have left us over the years, that, like, told us, “Okay, we can do this.” When we first wrote our first training code base completely from scratch, it wasn't a fork of any open source. It was just like, “Okay, let's build it from scratch.” I remember we had this one moment where we spent three weeks working out an optimizer bug. Like, it was like training just couldn't get stable. We, like, obsessed over it, and we thought, like, maybe we were wrong. Maybe we should have just forked this repo, or we should have. But then when we solved it, I still remember at the time we were like five people in the company. when we solved it, we were like, “Oh, we can do things,” like if we're just willing to work hard. and I think that culture with a very strong engineering bias has helped us, like, get to where we were. And so there's this notion of open source and talent and these things. I think we, We just took different decisions from a different starting point. and I think we are lucky. I do want to definitely call it lucky. And there was a lot of hard work at the team that now, like, that's starting to show up in results.Swyx [00:13:52]: Just ‘cause we probably won't revisit this again, but, and this is a fun recruiting challenge if someone knows the answer. What was the bug? And then we won't tell the solution, but we'An Optimizer Bug and the Value of Building From ScratchEiso Kant [00:14:01]: So the - This - You're gonna test my memory here,Swyx [00:14:04]: Oh, okayEiso Kant [00:14:04]: So but I thinkSwyx [00:14:05]: DirectlyEiso Kant [00:14:05]: I think I can recall. So if you, so if you look at, So if you take like Adam as an optimizer, you have epsilonSwyx [00:14:12]: YeahEiso Kant [00:14:13]: Which is, right, like in the denominatorSwyx [00:14:14]: Momentum and weights. YeahEiso Kant [00:14:15]: Is exactly, in the denominator. And at the time, if I recall, you looked at like the early Llama papers and things like that. People were juicing epsilon, like, quite a bit. Like, they were, like, adding, I don't know if it was E minus four or whatever, like a high value for epsilon.Eiso Kant [00:14:31]: And if you think about this during training, it's like a bit weird and counterintuitive that we're adding noise to our optimizer by just adding effectively, like, a random number in the denominator, right? Like behind the decimal point. And I don't recall the exact bug, but it had - What I remember is once we solved it, we no longer had to juice epsilon as much as, like, was happening in the Llama paper and other places. and it was like one of those fundamental moments where we had trusted this paper that was out there, and we're like, “Oh, no, it has to be this way. It has to have this high value of epsilon.” But it made no sense to us intuitively. Like, why do you have to have this so high? Like, if you're just trying to avoid division by zero, why can't the value be extremely small? and that was like one of those moments where you realize like, okay, finding things out from scratch yourself builds a better intuition. Because the one thing you learn very quickly with model building is that your intuitions that you start with are gonna get beaten up so hard.Eiso Kant [00:15:33]: Right? Like - It's such an experimental science, that the things that seem obvious, you very quickly get to learn, like, you were wrong, and hopefully you figure out why, and sometimes you don't even.Swyx [00:15:45]: Yeah. yeah, so, one of the reasons that you, when you released your new models, Vibhu got really excited. I mean, everyone got really excited. But Vibhu led our paper club on it, and you guys sawEiso Kant [00:15:58]: YeahSwyx [00:15:58]: Obviously. maybe talk through some lessons learned in that, whatever you can disclose. we can focus on the model factory stuff, whatever you think is a good starting point.Model Building as EngineeringEiso Kant [00:16:08]: So I would say that our view from very early on in the company was that model building is ultimately 90% engineering.Eiso Kant [00:16:18]: And I think we all know it in the industry because if you look at where's every researcher spending their time, they're spending their time writing code, right? Looking at data and writing code. And so we said, okay, The state at the moment, like three years ago, was bash scripts and Slurm and spaghetti code bases for training and, like, data pipelines that were patched together. And we looked at this and said, “Well, ultimately, model building is a process.” You're going from raw data, right? Like training raw material, the web, et cetera. you're doing a whole bunch of filtering, cleaning up, transformations, analyzing. These days, that's, far more complex than it was three years ago. then you're training a model, which is effectively a large distributed systems problem, right? Across hardware that has still-- It's become a lot more reliable. It was extremely flaky back then. and now with every new generation, we get our new sets of challenges. And then you go into the next stages, right? There was no training back then, but, like, you got, your post-training and then your reinforcement learning. And so we looked at this and we said, “Well, this looks like an industrialized process. This looks like an end process, that every single part of it has its machinery,” right? If it's your big data pipelines, if it's your crawling ingestion of the web, if it's your, large-scale distributed training, and then you've got your reliability. And we said, “Well, why don't we take some of the world's smartest distributed systems engineers that we knew and make them part of the process of research from day zero?” Not retrofitting it later on, but, like, really from the beginning. And that became our model factory. And so our model factory started with a handful of components. Today, it's thousands of components, and I try to equate it to, if you think about, like, someone who was at the very early days of Foxconn, if they had been there for the following, decade, they would be able to rebuild Foxconn because they saw every decision that led to building that system and all the complexity. If you and I walk into Foxconn today, no chance.The Model Factory and Experiment VelocityEiso Kant [00:18:18]: Right? Because we don't have the lineage and history of decisions that led to that. And so we built early on from the beginning- with a team that really understood that, well, the metric that we are optimizing for is the speed of an idea from a researcher to an experimental result that we can trust to then being part of the next model training.Eiso Kant [00:18:42]: And in the. And because it's such an experimental science, ultimately, in the beginning when it wasn't that complex, you could patch your way around it, right? But now, at any foundation model company, you are running. I mean, we're a small team, right? We're less than 70 researchers, another 35 engineers. and we are running, I haven't checked the latest count, but far more than 10,000, maybe 10 to 20,000 experiments a month that we cut. And so if you look at that scale of every model run that is, like it's ultimately it's, it's you need to be able to trust it as an infra problem. And so what we have now done over the years is gotten really good at that, and just by working it and improving it and obsessing over those end decisions. So now what that means is that you looked up Laguna XS 2 that we launched. It was five weeks from the beginning of training to launch. The model that we're gonna talk about today was eight weeks from start of training, to launch. We started the next model literally yesterday because we now finished the post-training required for the model we're launching, next week or by the time this comes out today. and we move that compute to the much larger Laguna M model that we're now training. And so the model should be an artifact of someone's process. It shouldn't be really a thing in itself. Like, and we treat this like the way you would look at like a SpaceX factory where, yes, the first rocket, really hard to build, but the much harder challenge was building the factory. And now they're rolling off, and no one is really thinking about the next launch anymore. So it's just another launch, it's another launch, another rocket comes off. And that's what we're trying to do with model building.Eiso Kant [00:20:22]: And what has been, which was not planned from day zero, it was in the back of our mind like this will happen one day, is that when you build a really good end model factory with really good APIs and really good engineering systems, Well, what is it perfect for? It's perfect for agents.Agents Inside the Model FactoryEiso Kant [00:20:40]: Because agents are now starting to take over more and more work in our model factory.Vibhu [00:20:43]: Yeah.Eiso Kant [00:20:44]: So I look at the screens when I walk, like when we're, we come together, in our monthly, we do monthly onsites, and I walk behind people's screens and I stop by and I talk to our researchers. And the default is all of these different agents running on their screen that are writing the code. They're launching the jobs. They're evaluating the results that are coming back from the model runs. They are, making the changes. And we're still in the driver's seat. We're still coming up with the ideas. We're still helping with the debugging. But more and more, and this is right now very profound on the data side of our pipelines in both pre and post and the synthetic data pipelines, it's starting to become more on the architecture side as well. You're starting to see these twinklings of what RSI is gonna look like.Eiso Kant [00:21:27]: And that's. So when we talk about, like to your question about our models, every talk about the model factory, And my coolest example of these things is always that when we kick off a new run, doesn't matter if it's a training like big run or if it's now a post, like one of 10 post-training versions we do for like release or many experiments, is that at any given moment, the changes that somebody made that they had experimental results from the day before make it into that run.Eiso Kant [00:21:57]: So there's not like a cutoff 90 days before. Like no, it's like literally from that moment because we can now trust the machine enough. And then you also have to invest in the reliability. So one of my favorite metrics about like Laguna S is that there was no call events, Right? Like completely zero. And we haven't had a meaningful call event, like something to wake up for, as far as I recall this entire year. now there is one asterisk to that. In usually the first six hours of launching a new model run, something breaks because you set a config wrong, you made a small mistake, et cetera. So that's usually there's a little bit of intervention, but that's always within like call periods, right? Not on call. And I think that's starting to now compound. So the model we're releasing now, I love it. It's amazing, but we're already onto the next one. and I think that's the way it should be.Laguna, Five-Week Builds, and Zero On-Call EventsVibhu [00:22:50]: Hey, I also just wanna point out, so for context, this was like a month ago. we found it in the tech report, so we just came in with, “Okay, new model's dropped. Haven't heard about it.” We wereEiso Kant [00:23:02]: Yeah, we're very used to doing this every few months.Vibhu [00:23:03]: We're, we're very much like, “ okay, look, it's like, on par with Kimi, DeepSeek, whatnot, the small ones, Gemma level. Oh, it's a very cool paper on what goes into building.” And then we hit this page, right? Like literally page two of tech report is, “This process allowed us to build the small model from scratch to delivery within five weeks applying the lessons”. And then I'm like, oh, this paper is not about here's a tech report of benchmarks and here's how many tokens it was trained on. Like for people that wanna dive more from what we're not gonna discuss on the podcast, it's all laid out here, right? FromEiso Kant [00:23:38]: YeahVibhu [00:23:39]: Custom software that agents can use to interface with training code, training data.Eiso Kant [00:23:45]: Yeah. Well, link the paper correctly, so yeah.Vibhu [00:23:47]: Yeah. All that stuff. read the paper here, but,Technical Report Principles and Streaming Training DataEiso Kant [00:23:50]: But I would like to. I love principles, and I think that is a good starting off point for maybe telling some stories. Maybe we can go one by one past the principles. I'll just call out that Dagster just got bought by a Prefect.Vibhu [00:24:01]: Yeah.Eiso Kant [00:24:01]: Isn't it fun? But yes, I'm very familiar with Dagster. just anything where like they trigger some story.Vibhu [00:24:07]: So, well, I would say, well, experiments code's obvious, but I think one of my favorite things is, I don't know where it is in here, but early on, and I still think this is the case a lot of foundation model companies, people prepare their training data sets, they get packaged up, then they get copied over to a training cluster distributed across all of the nodes, and then training starts.Vibhu [00:24:30]: And we looked at this like three years ago and we were like That makes no senseEiso Kant [00:24:36]: You lose so much time because the moment you have to rematerialize the data set, you have to make a change, you have to fix something, et cetera, you've got all this time of like repackaging it, right? Toca- tokenizing it, repacking it, moving it over to a cluster, then distributing it across the nodes. The bigger your clusters are, you start using fancy like torrent-like algorithms to like distribute your data. So why aren't we streaming data into training? Right? Something that's very common and like just basicVibhu [00:25:00]: Like just in timeEiso Kant [00:25:01]: Just in time, like good computer science like principle. And that was one of the first things that I think unlocked - the model factory. Because the moment you start thinking about, well, a training job, it doesn't matter if it's a big hero run or a small like, post-training experiment, consumes a certain number of tokens per second, right? And it's not a lot, right? From a like a data, moving data perspective. So we said, well, we have our training cluster, and then we've got like our AWS kinda setup where we can build these amazing big data pipelines. We can set things up. We use Spark underneath the hood, like all these things.Vibhu [00:25:36]: But when you say AWS, it's not actual AWS, it's your internal AWS.Eiso Kant [00:25:39]: It's our internal-- No, it's our internal like just running like our infrastructureVibhu [00:25:42]: Site web servicesEiso Kant [00:25:43]: Exactly. Our stuff running on like an AWS account or on like any hardware, right?Vibhu [00:25:47]: Yeah.Eiso Kant [00:25:48]: And so once we made that shift into I can stream data into training, all of a sudden you realize a lot of things unlock. Because now you don't have to wait for the whole data set to materialize.Immutable Data, Experiments as Code, and Scientific RigorEiso Kant [00:26:00]: You now all of a sudden when you're running data experiments about mixing data, it's a config. Because you've got these data sources that are coming in, and you just - we have this service called Blender that's in the report, where we then say, “Okay, for this run, I want 20% of this source, 10% of this source. I want this much, so many epochs of repetition. I want this to be, shuffled in a certain way,” and your training job can start while the rest of the data is even still materializing. also what it does is because all of this underneath-- So for us, we treated the data layer underneath as like an immutable data layer, and that was really important. Like experiments as code, immutable data layer means that you can always go back and understand literally down to the single token at which cursor it went in on which version of the code.Vibhu [00:26:47]: Yeah.Eiso Kant [00:26:48]: And it took us a I have to admit, like the first year of Poolside, we understood that engineering had to get great, But we didn't understand yet, that this is ultimately in support of like a good rigorous scientific progress. We were quite a - We were a very small number of people, so a lot of it was YOLO ideas and YOLO runs.Vibhu [00:27:08]: Yeah.Eiso Kant [00:27:09]: And we built great infra for the YOLO runs. But once we realized that we treated data as immutable and code as always versioned, and you could always track and trace every experiment end to end perfectly, you could repeat everything perfectly, right? You have perfect reproducibility. I can still reproduce runs from two years ago if I wanted to, right? It enables the scientific progress, like the scientific process, and I think that took us probably about a year and a half into the company to figure out. We also had some great hires, like our head of applied research, Nikolai, who joined us from Yandex, who'd been working on language models since like the early 2020s, I think brought that into the company of like, “Hey, we wanna have even more rigor.” And then once we kinda had the combination of like increasingly more capable platform that allowed people to do more, but had this immutability, we were able to start “Okay, every experiment is truly an ablation. We truly need to understand it.” And I think we became much more scientifically rigorous in the last couple of years, and the infra underneath enabled it. and then there's just fun stuff like, andVibhu [00:28:16]: Yeah, a lot of it's fun, like even just the, one, you share all the ablations, two, picking the data sets, right? There's like a random small paragraph in here where it's just like, “Oh yeah, training data, we have some, we have an auto mixer.” it trains eight small models, scales them up, picks the training data set. We don't even need to look at it. I'm like, “Wow, a lot of engineering rigor there.” And there's just, there's just a lot in here.Publishing Research and Giving BackEiso Kant [00:28:40]: Yeah, and it'- and look, and we wanna put out more. Like we, We treat writing papers as something that we haven't earned the right for yet for a long time. So you earn the right to spend time, publishing research once you're at the frontier, because until then, you're catching up, and every minute and hour in this industry matters. Like I obsess over it, not just the wall clock time from idea to result, but just general like time every day that we, waste is one that doesn't allow us to catch up. But in this case, we said, “Okay, we're gonna give ourselves.” I think we gave the team like three or four days while still doing their work, like give everything in there. And to your point earlier, if your stuff, it's easy to like put it out. And so there's so many more things that we wanna talk about over time, and we will definitely start doing. And as we earn more of the right, but also now have like added to our mission that we want more foundation model companies to exist, you'll see us like be way more proactive, and just trying to keep dropping some of those like things that we've learned along the way that can help others like speed up.Vibhu [00:29:40]: Which is the other cool side of this, right? It's, it's not like, back to your point, it's not just here's the benchmarks of our training. If you want to replicate, here's experiments of optimizers, data sets, post-training. you lay out a lot of it here alongside here's your system for how to do it? So it's, it's really like promotingEiso Kant [00:29:59]: No, thank youVibhu [00:29:59]: Other people can do the same.Eiso Kant [00:30:00]: And by the way, I also wanna make clear, right, we have been incredible-- Like we've taken a lot of advantage of the fact of all the open research that others have published, Right? And you mentioned, the Chinese labs, and we I think it's important that there's, from every country and every culture and background, including like Western companies like us, there's different models that come out that people can choose to trust. But I think we do have to give credit where credit's due, right? The incredible Chinese lab have done an amazing job at sharing their research, and we have definitely like been on the receiving end of taking advantage of that. So when you're on the receiving end of something coming to you, I think it's, you also have an obligation to give back.Swyx [00:30:39]: Do you have a favorite or underrated Chinese lab that you wanna shout out? Everyone shout outs DeepSeek.Chinese Labs, Zhipu, and PersistenceEiso Kant [00:30:44]: That's a good question.Swyx [00:30:45]: Moaan obviously for Therapsi. Yeah.Eiso Kant [00:30:48]: Yeah, look, I think, I think obviously everyone's been talking about Zhipu lately, with 5.2. I think what most people don't realize is when they started.Swyx [00:30:59]: Yeah.Eiso Kant [00:30:59]: Right? They started years before ChatGPT.Swyx [00:31:02]: They just rebranded. YeahEiso Kant [00:31:03]: And so, I've like, I remember how hard it was to work on these things Before the rest of the world got excited about it. And so I have an immense amount of respect for people, who were working on improving models when it wasn't the sexy thing to do, when believing in LLMs, was gonna get you ridiculed. I remember like back in 2016 when we were doing what we'd call, machine learning on code with some of these models. we would-- people would just laugh at us, like they'd be like, “This makes no sense. Like why are you wasting all these, like, millions of dollars on trying to figure this out?” And so I would say they're probably the one that, I think deserves a shout-out, not just because their latest model is very good, but because they fought to get here. And I think, I think every foundation model company it takes time to get here, right? It took us three years to get to the model that we're, that we're now gonna be releasing. and now the time in between the models is coming, is counted in weeks. It's no longer counted in months or years. But this stuff's hard. and if we can make it a little bit easier for the next person, like we should all do so. Because if we don't do so, we're, we've got a small window before models are really impacting recursive self-improvement to a level where catching up otherwise might become unfeasible. And we should try to, in that window, encourage as many labs or however we wanna call them, like to start. And so one of my currentEiso Kant [00:32:36]: Mission, but qualm is like I wanna encourage whoever is a researcher right now who thinks they can tackle this to go and leave and become my competitor.Eiso Kant [00:32:45]: Like start another foundation model company because I think we need it. I think otherwise we're not gonna be in the world where, I don't want to just be the fifth or the sixth company that wins. I wanna look at a world where there's lots of choice.Starting a Foundation Model CompanyVibhu [00:32:57]: What else do people not see in starting a foundation model? it's, there's a lot of compute, there's a lot of capital required, a lot of compute. You lay out model factory and how to do the training, but there's a lot there, right? That's,Eiso Kant [00:33:10]: Well, look, it's, I in turn-- this is an oversimplification, and I always asterisk it with that because it can land a little bit the wrong way in people's minds. But I think you can sum down, And I saw it, 95% of model building to just doing, you're just doing two things. You're improving data or you're improving compute efficiency. And I know that feels like an oversimplification for the incredible, like, Gifted and skilled work people do. But if you really look at it, like what are we doing? We are looking at data, we're generating new data, we're improving data. and the only way to do that is to look at the data, right? That's a big part of foundation model building. And on the other hand, we come up with these incredible breakthroughs in inference, in architecture, and new attention mechanisms. But what are they really doing? They're bringing compute efficiency. Now, we have definitely had some breakthroughs over the years that allow for more model capabilities. But at the limit, if you could train a large enough model, right, like, and you had infinite compute, we probably-- if you had infinite compute, you'd be at AGI probably already tomorrow.Eiso Kant [00:34:12]: Right? Like it's not. And so, and let me say that infinite compute with infinite ability of much faster networking because networking ends up being more of the bottleneck than compute. But, so I do think that's, those are the main things. And to just realize that this is engineering. I think it's become more obvious, but I think for quite a few years, people have held foundation model companies and researchers and others on this pedestal of like you're doing incredible magic or rocket science, or only like, Nobel laureate physicists can do this. And don't get me wrong, there are some really hard problems that need to be solved, but a lot of the work that all of us are doing on a day Is not sitting down trying to solve a math theorem. A lot of the work that we're doing is just really doing the basics right, writing good code, looking at data, improving it, running experiments, looking at plots, trying to see like, hey, trying to shape our intuitions. And a lot more people could be highly capable researchers. and I think that's, it feels far for people to do so. But I've seen in our own company, we've seen engineers become researchers because the model factory allowed them to be, have a much lower hurdle of running experiments and trying things. And one of the guys on our team who started as an engineer building our agents is a legit reinforcement learning researcher now, making real progress. and that happened in the span of like six months. that would've not been what I think most people assumed was possible, a couple of years ago.Swyx [00:35:46]: Yeah. I think one of the interesting moments is when you can self-host, like, if in a programming language, like if you can compile the language in the language, the equivalent is can you use your own tools, right? You have the pool CLI, you have your own models. presumably you're not only using your own models. There's no way. But like, what's that percentage over time?Laguna S, Persistence, and Behavioral GainsEiso Kant [00:36:10]: This is the first model that we're releasing that is starting to meaningfully contribute to our own work. It's not a it's not state-art model yet. Fable and other, they're, they're very capable models, but Laguna S Is really interesting. I'm gonna pull up the quote. Peng Ming, one of our heads of applied research, said something, last week as the model came out about 10 days ago, much better than we had hoped for or expected. And he said, I have the feeling that a lot of the gains in Laguna S come not from more intelligence, but more from different behavior, more verification, less taking things for granted, not declaring victory early, and being way more persistent. And to be honest, those are more predictive than raw intelligence for success in human also to some degree. And this was, he wrote me this on 5th of July on a Sunday, and it's been burned in my brain ever since because the Laguna S model, as you'll see it and why it does so well on benchmarks and why it does so well in using it on a day basis, is that it's just incredibly persistent. It reasons a lot. I do call that out. We have work to do on making it more efficient. We have to work to do on offering different reasoning modes. But this is the model that has been able to do things that I never thought it could do. A hundred eighteen billion 8B active model, which is not that large. It fits on a DGX Spark and still runs at, thirty, forty tokens a second on a Spark, is able to solve Erdős 397 independently. It's able to do complex programming tasks. It's able to. I asked it this morning to make me a Fi scanner without using any external libraries on my Mac, and it's, like, figuring out, like, the core WLAN API by really persistently trying to understand it without access to the internet. And more, I love vibe checking. I've probably spent eight to ten hours a day with this model for the last ten days.Eiso Kant [00:38:05]: I'm not exaggerating. I was on my eleven-hour flight yesterday. I spent ten hours reading trajectories and traces and, like, of the model.Eiso Kant [00:38:12]: And what I take away from it is exactly what Peng Ming said. We are gonna be able to squeeze so much more out of smaller models than I think we had imagined in the industry because, yes, there's intelligence and larger models are more intelligent. Like, no doubt about it. We should continue to scale up. but the behaviors of being really persistent, of being able to backtrack when you're wrong, of, like, understanding how to interact with your environment show us that we can get a lot more out of it. And this, for me, has created a bit of a Question in my mind the last couple of days. If you think about where we're using models today, right? We are using models, say, for knowledge work. Represents twenty-five percent of the global economy, twenty-five trillion dollars of work.Eiso Kant [00:39:00]: As we scale up models and they become more intelligent, we are excited about using them more and more for pushing the frontier of science.Small Models, Knowledge Work, and CommoditizationEiso Kant [00:39:08]: And if you look at the frontier of science, like true breakthroughs in science, they have been linked, they are linked to more intelligence in many places. Einstein figuring out general relativity is able to bring ideas together that other people would have not brought together. And I think one of the many dimensions of intelligence is the ability to do that, and it's something we clearly see that as models get larger and more capable, they're able to pull more ideas and threads together that a smaller model wouldn't be able to.Eiso Kant [00:39:36]: And we're starting to see examples of that in medicine and, like, in bio and other things. But if you think about the majority of knowledge work that we do, and it includes building software. I'm a software developer at heart first and foremost probably, although I probably can't say it that much anymore as I don't write production code in years, is that what makes us good is our persistence. It's our ability to encounter a problem and backtrack and say, “I need to go figure out this bug. I need to go research this. I need to go look at the documentation. I need to, like, try different, five different ways to see, like, if I can solve it.” But it is not necessarily bringing three ideas together from radically different fields. And so if we are now seeing, and I think Laguna S is an example, that we are able to make a relatively small model much more capable than I had definitely predicted or any previous, like, benchmarks had shown for any model remotely this size or even larger, At least on coding tasks, that it's because of the behaviors. And so now the question I have, and I don't have an answer, it is I know at the limit, so infinite model size, right, extremely large model, and the cost of that model is gonna be very expensive to run. We know this, right? So larger model ROI.Eiso Kant [00:40:52]: So I know that at the very limit, I'm not gonna use the world's largest model one day, quadrillion parameter, whatever crazy, like, scale we scale up, to do a basic coding task. Already today, I'm starting to size down for certain tasks.Eiso Kant [00:41:07]: So it means that there is an optimal. It means there's some curve that goes as we go up to model size for knowledge work, at some point we're at the peak, and after that, the return on investment of using a bigger model, just doesn't make sense.Eiso Kant [00:41:22]: Now, I think the question is, before I would have thought that peak was extremely very far away.Eiso Kant [00:41:30]: This model for me is the first sign that Maybe that peak is At a trillion, five trillion, ten trillion. Maybe we can just squeeze way more out of these models. I'm no longer thinking that we need two or three orders of magnitude on the largest models to be able to, solve knowledge work, the accounting, the legal, the code that we write. And so if that holds true, It is an argument for the commoditization of models. It's an argument that open source can win and, like, succeed in this world. And now it's of course a self-serving argument and it's a hopeful argument, but theoretically at the limit it works. We just have to go discover in the next couple of years of how much more we can squeeze out. Now, I do want to put a big asterisk. This does not mean I'm against scaling models. I think we ultimately only succeed if we scale our models as large as our competition. I do not like. I think we should not put our head in the sand and say we're gonna be king of open source small models. I think that's, It's a out. It's trying to be king of your own kingdom, but not realizing what the rest of the world's doing. All of us rather use a smarter, faster, more model. It's a sign of hope. And so I don't wanna overly state this is a good model. We have a long way to go to get to the state-art. But what hopefully people take away when they use this model is that the behaviors inside of it are what push it to be far more capable, less than necessarily the number of parameters.Pre-Training, Mid-Training, and RL Moving EarlierVibhu [00:43:03]: Is that mostly post-training? LikeEiso Kant [00:43:05]: YesVibhu [00:43:05]: Right.Eiso Kant [00:43:06]: It's entirely post-training.Vibhu [00:43:08]: Are we done improving anything on training? Is, like, training done?Eiso Kant [00:43:12]: No.Vibhu [00:43:12]: Okay.Eiso Kant [00:43:13]: SoVibhu [00:43:13]: I just wanted to cover training, and then we go post-trainingEiso Kant [00:43:15]: Training is not done. I mean, look, there's a part of training of just dealing with skill, right? Every new order of magnitude of model skill, you are going to get new things you gotta solve for. That'- but those are ultimately, engineering challenges.Eiso Kant [00:43:31]: I have a, I would say, a not commonly held opinion that reinforcement learning Will move earlier and earlier into training.Vibhu [00:43:42]: Yeah, training.Eiso Kant [00:43:44]: Not even training. Like training today, right, is, like if you look at - So we've been working on this for years already. and I think the best-- I think the first time we saw it out in public was the DeepSeek Zero paper. this is a year and a half ago, I think, if I recall correctly. where, you can Very early on in a model as it starts capable of being able to use language, et cetera, induce reasoning. and so the question that I have is like, we have this- we have the dataset that's the web. and the web, I think we could arguably say probably has The totality of humanity's knowledge somewhere encoded in different places. It's a huge variance degree of quality, from garbage data, and like once you look at training data, you really get humbled of like what the web is, to like, the most greatest scientific papers and best blog posts and like, best transcripts and whatnot.Eiso Kant [00:44:39]: And so now What we are trying to figure out, and have been doing a lot of work on, and it's a place where maybe not as open as we're on other things, but we will become more over time. we've been spending a couple of years really doing research on how can we turn the web into not just next token prediction, but into a way to teach the model to think earlier in its training. and I think there's a huge amount of gold to be found there. I think we are right now in, we've got some drugs in the industry. One of the drugs is distillation. Another drug is, more environments. Like, and they're great, and they make us feel good, and they make the models better, and like we're all addicted to them, and we'll use them, right? in various different ways. and but ultimately, I think we are still barely squeezing out of the web what we should be getting out of the web.Eiso Kant [00:45:33]: I think just next token prediction during training is not enough.Eiso Kant [00:45:36]: AndVibhu [00:45:38]: YeahEiso Kant [00:45:38]: I think we'll see some very interesting things still happen. and that RL in post-training to induce behaviors, to improve things, like I think - the whole world knows how to do this now. I think we're, we're scaling it up. Everyone is. But I wonder if we need to go as far as we're going today with environments. I'm not sure yetVibhu [00:46:01]: You mean we're going too far?Eiso Kant [00:46:02]: I'm, I'm not sure if the path to AGI is justVibhu [00:46:06]: Is more environmentEiso Kant [00:46:07]: More environments.Vibhu [00:46:08]: It seems like a never-ending, “Okay, I want instruction manual for this table, right? Am I gonna environment out building furniture? Or are we just gonna tail end like we need some general solution?”Eiso Kant [00:46:19]: I think there is, I think there's an ability to generalize more from the web. but I also am very encouraged, like when I look at Laguna S and, which is post-training is, well, is the big impact there. and I see like, oh, wait a second, just by making some of these behaviors much better, we're able to get so much more out of it. It just changes a little bit the way you think about intelligence.Vibhu [00:46:40]: Yeah. The analogy people draw often is the RL phase is where you don't learn as much new knowledge. You shiftEiso Kant [00:46:46]: Yeah.Vibhu [00:46:46]: Yeah. So, you shift distribution, and you can have it reason towards what you want. on your point about training, a lot of training is still just continue training in a domain, say medicine, then you do RL. So still justEiso Kant [00:47:00]: It's just better data, right? Like, I mean, training, ooh, I like how we invented this word. Like it's effectively just like,Vibhu [00:47:06]: Second phaseEiso Kant [00:47:07]: It's the second phase of training With like a really dumb way to do a curriculum. But like ultimately, what you'd want is a curriculum from token zero to token 30 whatever or 40 trillion tokens that really truly is the optimal curriculum for the model to learn. But training is essentially a stage curriculum on the web because we do not have to compute, And, effectively to try to ablate the perfect curriculum, right? And so I'm pretty sure that you'll start to see people talking soon about some other term, and there's two or - ‘cause now we do this, right? We talk stage two and stage three and stage four training and like. But ultimately, all we're doing is we're trying to assign a curriculum to the web data that we have to allow the model to learn better. I think at some point, as things get compute, as models get cheaper to run, as the next generations of compute, this will become more of a continuous spectrum. I also think the reason, by the way, you have training and like stage two and stage three is organizational, Right? It'- this is, I think, a thing where-- that we really try to avoid with the model factory is like Training exists because there's a training team now, right? There's people, or like people in training decide to focus on like a training effort. but what you really want is engineering and scale of experiments that allows for a much more continuous spectrum that you don't, you have infinite stages. Now, we're not there. Compute's not there. Organization design is not there for it yet. but I think we'll get there. we'll look back on a couple of years and be like, “Oh my God, it was so cute that we did our training data like this in such a like naïve way. Like we barely ordered it. We didn't really do a good job at likeCurriculum, Auto Research, and New ObjectivesVibhu [00:48:48]: The building that curriculum will get you that in the industry.Eiso Kant [00:48:51]: And I'll confirm that, when I talk to some researchers that this is a lot of the focus now is like how does training change and what is the next objective other than, next token prediction. I assume you don't have the answers, but you have some ideas.Vibhu [00:49:02]: We have some ideas. We're not ready to talk about it yet.Eiso Kant [00:49:05]: Yeah.Vibhu [00:49:05]: We've been working on them for years, and I think that's the one thing that's also like you asked earlier about, like what's not obvious about building a foundation model company is that you are constantly balancing the table stakes work, the recipe worksEiso Kant [00:49:19]: Yeah.Vibhu [00:49:19]: Versus like your, my crazyEiso Kant [00:49:22]: Pure researchVibhu [00:49:22]: Breakthrough.Eiso Kant [00:49:22]: Yeah.Vibhu [00:49:22]: Pure research and finding that balance and adjusting the percentage to it based on where you are in the race is really important.Eiso Kant [00:49:31]: I mean, so like, this is a nice way. I was gonna bring up auto research at some pointVibhu [00:49:35]: YesEiso Kant [00:49:35]: As another Andrej invention, or coinage, which is like, I honestly, like how many objective functions can there be, right? Like just try 1,000 of them, set it running, whatever.Vibhu [00:49:47]: Man, it's alsoEiso Kant [00:49:48]: Like what you're looking for. You're looking for loss curves like that, likeVibhu [00:49:51]: It's also a thing people take bets on, right? When you say more Neo labs, you're doing a version of we'll do foundation models, scale them up, next token predictors. A lot of other Neo labs that we see want to take a completely different approach, right? At some level, you're right. It's all, compute efficiency, and that's the net objective. But some are okay, different architecture, like vastly different amounts of compute spend. So some are different. They're not justEiso Kant [00:50:19]: YeahVibhu [00:50:19]: They're like, 99% not balancing, here's the vanilla and scale up. They're 99% on, here's novel research that'll change everything.Eiso Kant [00:50:27]: And I think, Luke, I think you. It depends when you started as well, right?Pure Research vs. Table StakesVibhu [00:50:30]: Yeah.Eiso Kant [00:50:30]: When we started, like the novel thing we did was reinforcement learning on code. No long- that's no longer novel by far, but we were like, - that's where we obsessed over when no one believed in RL. So you have to when you start the company, you have to have your own idea. You have to have something that's different that allows you to speed up, right? For us, it was RL to LLMs that later became common, like, Knowledge. But in the beginning, it wasn'tVibhu [00:50:53]: It's cool. this was like your original 2023 blogEiso Kant [00:50:57]: YeahVibhu [00:50:57]: Of purpose.Eiso Kant [00:50:58]: Yeah.Vibhu [00:50:59]: And like you do lay it all out here.Eiso Kant [00:51:01]: We laidVibhu [00:51:01]: The blog is pretty underrated, right? The whole RL on code was very early on.Eiso Kant [00:51:06]: Very early. And even we had to argue with people, like we say here things like to push beyond current capability, to train your own foundation model. We had to argue with people that it mattered that you had your own like, base model. you can fine-tune your way to success, right? major capabilities emerge from training a base model made accurate and useful during fine-tuning.Vibhu [00:51:23]: Which like, for perspective at the time, we knew closed models, OpenAI, Anthropic were huge. The open models we had were like Mistral 7B, a 30B, a 70B.Eiso Kant [00:51:35]: When weVibhu [00:51:35]: YeahEiso Kant [00:51:36]: The date on this thing is wrong. When we published this, it was April 2023. I think this was justVibhu [00:51:42]: YeahEiso Kant [00:51:42]: Happened on a migration, probably found it on archive.org.Vibhu [00:51:45]: Mistral.Eiso Kant [00:51:46]: Mistral had started, we started on the same month, right?Vibhu [00:51:49]: Yeah.Eiso Kant [00:51:49]: So this wasn't even, there was only, I think, Llama out at the timeVibhu [00:51:52]: SnellEiso Kant [00:51:52]: And that's it, right? And so, but I agree. I think we wan

Die Krypto Show - Blockchain, Bitcoin und Kryptowährungen klar und einfach erklärt
#1175 Kimi K3: Neues DeepSeek oder Börsenpanik bei KI?

Die Krypto Show - Blockchain, Bitcoin und Kryptowährungen klar und einfach erklärt

Play Episode Listen Later Jul 20, 2026 7:30


Daily Snippet vom 20.07.2026 Kimi K3 aus China verschärft den Druck auf Nvidia und andere KI Aktien. Ist das der nächste DeepSeek Moment oder verkauft der Markt gerade die falschen Gewinner? Moonshot AI meldet 2,8 Billionen Gesamtparameter, bis zu eine Million Token Kontext und deutlich niedrigere API Preise als westliche Spitzenmodelle. Gleichzeitig war der Andrang so groß, dass neue Privatkundenabos pausiert wurden. In dieser Episode erfährst du: • warum Kimi K3 den KI Abverkauf verstärkte, aber nicht allein verursachte • was die aktuellen Benchmarks wirklich zeigen und wo der Hype beginnt • warum günstigere KI die Margen der Modellanbieter drücken und gleichzeitig den Bedarf an Chips, Speicher und Strom erhöhen kann • welche Entwicklungen langfristige Anleger jetzt beobachten sollten Meine CIO Einschätzung: Kimi K3 ist ein ernstes Signal für mehr Wettbewerb. Der pauschale Schluss, dass dadurch weniger Rechenleistung gebraucht wird, greift zu kurz. Welche konkreten Positionen und Einstiegsszenarien ich daraus ableite, zeige ich im Daily:

The Compliance 911 Show
Community Development Benchmarks

The Compliance 911 Show

Play Episode Listen Later Jul 15, 2026 12:15 Transcription Available


In this episode of Compliance 911, Dean Stockford and Len Suzio discuss the OCC's proposed Community Development benchmarks for CRA performance, focusing on why community development has historically been difficult for banks to measure and plan. Len explains how the proposed benchmarks are organized by activity type—CD lending, qualified investments, services, and combined lending/investment activity—as well as by performance rating, bank size, and annual measures such as Tier 1 capital, assets, and volunteer service hours. The episode highlights how these benchmarks may give community banks a clearer framework for setting CRA goals and evaluating satisfactory or outstanding community development performance. Listeners can also download GeoDataVision's special PDF recap containing all 67 proposed Community Development lending, investing, and service benchmarks from the GeoDataVision website. https://geodatavision.com/content/occ-proposed-elective-goals-for-cra-strategic-planning/ Brought to you by GeoDataVision and M&M Consulting

Ransquawk Rundown, Daily Podcast
EU Market Open: Europe primed for a weaker open as energy benchmarks rally

Ransquawk Rundown, Daily Podcast

Play Episode Listen Later Jul 14, 2026 2:49


The US continued to launch strikes for a third night, after President Trump stated that he would hit Iran “very hard” on Monday and Tuesday. Separately, Trump threatened to hit Iran's “Pickaxe Mountain”, an underground nuclear facility; Brent Aug'26 +1.9%.US President Trump also announced a naval blockade, which is set to begin at 21:00 BST / 16:00 EDT.APAC stocks were mostly in the red given the geopolitical environment and tech sell-off; European equity futures are indicative of a weaker open.DXY takes a breather; Kiwi outperforms following regional data and hawkish comments from RBNZ's Conway.USTs and Bunds remain pressured amidst the elevated energy prices and after Fed's Waller delivered hawkish remarks.Looking ahead, Highlights include German Wholesale Prices (Jun), Chinese M2 Money Supply (Jun), US NFIB Business Optimism Index (Jun), US CPI (Jun), Fed Discount Rate Minutes (Jul), ECB President Lagarde-US Treasury Secretary Bessent meeting.Speakers including Fed Chair Warsh, Goolsbee, Barr, Cook & Bowman, BoE Governor Bailey, Supply from the Netherlands & Germany.Earnings from Citi, Goldman Sachs, JPMorgan, Bank of America, Wells Fargo.Read the full report covering Equities, Forex, Fixed Income, Commodites and more on Newsquawk

Sounds Profitable: Adtech Applied
Signal Hill's 2026 Podcast Ad Benchmarks, AI-Legible Audio's Competitive Edge, & More

Sounds Profitable: Adtech Applied

Play Episode Listen Later Jul 8, 2026 4:51


Today in the business of podcasting:Signal Hill Insights' 2026 benchmark report, "How Podcasts Impact Brand Perceptions," finds video podcast ads now make up 54% of aggregated brand lift data, up from zero in 2023, with video edging audio by one to two points on mid funnel and purchase intent metrics.Forbes contributor Damion Taylor argues podcasters need to make their audio "AI legible" with transcripts, speaker IDs, and timestamps, warning that platforms prioritizing video visuals over audio quality risk making hosts sound less intelligent and likable to listeners.Audacy research covered by Inside Audio Marketing finds news/talk radio listeners are more loyal and trusting than other formats and skew affluent, making the audience a strong fit for integrated sponsorships over traditional spot buys.To find links to these, and every article covered in today's episode, click here. You can also subscribe to The Download's newsletter to receive the full issue straight to your email inbox every day.

I Hear Things
Signal Hill's 2026 Podcast Ad Benchmarks, AI-Legible Audio's Competitive Edge, & More

I Hear Things

Play Episode Listen Later Jul 8, 2026 4:51


Today in the business of podcasting:Signal Hill Insights' 2026 benchmark report, "How Podcasts Impact Brand Perceptions," finds video podcast ads now make up 54% of aggregated brand lift data, up from zero in 2023, with video edging audio by one to two points on mid funnel and purchase intent metrics.Forbes contributor Damion Taylor argues podcasters need to make their audio "AI legible" with transcripts, speaker IDs, and timestamps, warning that platforms prioritizing video visuals over audio quality risk making hosts sound less intelligent and likable to listeners.Audacy research covered by Inside Audio Marketing finds news/talk radio listeners are more loyal and trusting than other formats and skew affluent, making the audience a strong fit for integrated sponsorships over traditional spot buys.To find links to these, and every article covered in today's episode, click here. You can also subscribe to The Download's newsletter to receive the full issue straight to your email inbox every day.

Ransquawk Rundown, Daily Podcast
US Market Open: Energy benchmarks firm as IRGC strikes vessels in Strait of Hormuz; NQ underperforms with semis hit post-Samsung

Ransquawk Rundown, Daily Podcast

Play Episode Listen Later Jul 7, 2026 2:17


Iran's IRGC fired at least two missiles at ships in the Strait of Hormuz, Axios reported, citing a US official. It was separately reported that one was a Qatari LNG tanker. Brent Sept'26 +1.2%.Fars reported that the Qatari tanker attempted to pass through the Omani route and ignored repeated warnings.Iranian Foreign Minister Araghchi said negotiations on a final deal will not commence if threats continue.Global tech stocks fall after Samsung Electronics (-6.9%) shares plummeted post-earnings; NQ -1%.DXY is incrementally firmer, JPY marginally outperforms after a Japanese official pushed back on reports that the government is urging the BoJ to lower interest rates.Looking ahead, highlights include US ADP Employment Change Weekly, Trade Balance (May), New York Fed SCE (Jun), Atlanta Fed GDP (Q2), Canadian Trade Balance (May), Ivey PMI (Jun), EIA STEO (Jul), NATO Ankara Summit. Supply from the US. Speakers include BoE's Mann.Read the full report covering Equities, Forex, Fixed Income, Commodites and more on Newsquawk

SaaS Metrics School
The Latest GRR Benchmarks

SaaS Metrics School

Play Episode Listen Later Jul 2, 2026 5:26


Is gross revenue retention under attack at your SaaS company? The latest benchmark data says the ground has shifted under everyone. In episode #380, Ben Murray breaks down the latest SaaS gross revenue retention benchmarks from Ray Rike's Benchmarkit report, the same data set Ben uses to benchmark his own client base. GRR is one of the power three metrics, and it is hard to scale without it. Pricing models are changing; seat-based pricing is under pressure, and AI is reshaping how revenue holds. If your board still treats 88% median GRR as the baseline, you are benchmarking against last year's reality. Why median GRR fell from 88% to 84% year over year, and why the top quartile slipped from 95% to 91% Whether the 95% GRR "elite" rule of thumb still holds, backed by three years of top-quartile benchmark data Which pricing model retains revenue best, comparing pure subscription against usage-based and subscription-plus-usage Why vertical SaaS is outperforming horizontal SaaS on retention, with a 90% median GRR versus 84% How to benchmark GRR the right way by ACV segment instead of relying on dangerous aggregate numbers Tune in to see where your gross revenue retention really stands, before your next board meeting or investor update. Resources Mentioned Benchmarkit SaaS Metrics Benchmark Report, Ray Rike: https://www.benchmarkit.ai/2026-saas-ai-native-metrics Ben's KPI app: https://softwaremetrics.ai/ Ben's blog post on 2026 GRR benchmarks: https://www.thesaascfo.com/saas-grr-benchmark-2026/

MoneywebNOW
[TOP STORY] Benchmarks demystified: From market indices to personal portfolio goals

MoneywebNOW

Play Episode Listen Later Jun 30, 2026 6:04


‘One of the biggest risks in choosing benchmarks is that you don't match what you actually want as an objective' – Satrix Quantitative Portfolio Manager Siyabulela Nomoyi.

The Six Five with Patrick Moorhead and Daniel Newman
Qualcomm's Data Center Debut, OpenAI's Jalapeño, and the Memory-as-Strategic Infrastructure Debate | The Six Five Pod Ep. 310

The Six Five with Patrick Moorhead and Daniel Newman

Play Episode Listen Later Jun 29, 2026 62:08


On Episode 310 of The Six Five Pod, Patrick Moorhead and Daniel Newman unpack the biggest stories from the week, including insights from Qualcomm Investor Day 2026, OpenAI and Broadcom's Jalapeño AI chip, Anthropic's Micron partnership, SpaceX's massive Reflection AI compute deal, Sakana AI's new Fugu orchestrator, and why memory is emerging as a critical layer of AI infrastructure. Plus, Bulls & Bears covers NVIDIA's $25B bond offering, Apple's MacBook price increases, Micron's record quarter, and Cerebras' first earnings as a public company. The handpicked topics for this week are: Qualcomm Investor Day 2026 — The Data Center Debut: Pat and Dan break down Qualcomm's push into the data center after the company took the stage with Microsoft's Satya Nadella and Meta's Mark Zuckerberg as named customers. They unpack the new Dragonfly platform, including the C1000 250-core data center CPU with PCIe Gen 7 and CXL, the AI200 and AI250 inference accelerators, and a novel High Bandwidth Compute (HBC) architecture that stacks compute under LPDDR memory at dramatically lower cost than HBM. They highlight Qualcomm's ambitious growth targets: $15B data center revenue target for FY 2029, an increased total non-handset revenue goal from $22B to  $40B, and a shortened timeline for automotive revenue by two years. They also debate the identity of Qualcomm's unnamed hyperscaler customer and why its robotics opportunity may be flying under the radar. (The Decode) OpenAI and Broadcom Unveil Jalapeño, OpenAI's First Custom Chip: A photo of Sam Altman and Hock Tan holding a wafer and packaged die kicked off OpenAI's reveal of Jalapeño, a custom inference chip built with Broadcom and slated for late-2026 deployment. The chip reached tape-out in roughly nine months, which is an aggressive cycle for an ASIC of this size, and uses HBM3E memory. Pat takes a victory lap on his long-standing heterogeneous compute thesis: every hyperscaler and now every model lab is building accelerators, and the XPU efficiency argument has played out as predicted. Dan frames OpenAI's broader move as existential: they cannot serve frontier models at premium margins if compute remains constrained. He flags that OpenAI is trying to do everything from chips and fabs to social networks and browsers, and that its IPO is now delayed. (The Decode)   Anthropic and Micron Sign a Strategic Multi-Year Memory Agreement: Anthropic and Micron announced a multi-year supply agreement for HBM, DRAM, and SSDs, including co-designed next-generation memory for AI workloads, along with a strategic investment by Anthropic in Micron. The pattern mirrors Samsung and SK Hynix's pre-funding Anthropic in May, and follows OpenAI's Jalapeño as another frontier lab moving to lock in supply chain control. Dan frames it as the same circular financing playbook NVIDIA ran two to three years ago, but with the ball now in the memory triopoly's court. Pricing-floor agreements with no ceilings, customized rather than commoditized memory architecture, and demand running well past the previously assumed 2027-2028 horizon. Pat notes that the rumored 14% free cash flow margin at Anthropic makes the strategic investment math work cleanly for both sides. (The Decode)   SpaceX Signs $6.3B Compute Deal with Reflection AI: SpaceX inked a $6.3B compute lease with open-source AI lab Reflection AI, at $150M per month from July 2026 through 2029, giving Reflection access to NVIDIA GB300 chips inside the Colossus infrastructure. Combined with the $920M-per-month Google compute contract and existing xAI commitments, SpaceX now has a contracted backlog larger than most public AI startups' entire revenue base, with some calling it the largest commercial AI infrastructure provider at $80B in contracted revenue. Pat reads it as XAI failing to land with developers, consumers, or enterprises, leaving SpaceX with a pot of gold worth far more as wholesale capacity than as XAI's own training compute. Dan flags that Google owning 7% of SpaceX ahead of an IPO is not accidental, and the open question is whether this becomes a Nebius-style infrastructure trade or a full-stack Google-equivalent platform. (The Decode)   Japan's Agentic Orchestrator Sakana AI Ships Fugu Plus and Fugu Ultra: Japan's Sakana AI released Fugu Plus and Fugu Ultra, an agentic orchestrator built on a multi-agent MOE approach that routes workloads across multiple underlying models rather than training a new frontier base model. Sakana claims agentic capabilities on par with or better than top frontier models at significantly lower input/output token costs, similar to the DeepSeek and GLM cost-undercut narrative. Pat compares the architecture to OpenRouter and notes the developer-facing parallel to Perplexity Computer's model-routing approach. Both agree that models themselves are no longer moats, and suggests the real moat is the harness, tooling, connectivity, looping, agentic stack, and total compute availability. Expect more sovereign agentic plays from Japan, the Middle East, and elsewhere on the same template. (The Decode)   The Flip — Is the Era of Memory as a Commodity Over? Daniel takes the FOR side: memory has moved from commodity to strategic AI infrastructure, citing 16 multi-year agreements covering $22B in committed volume booked through 2027, 84.9% gross margins higher than NVIDIA's, the technology barriers of HBM yield/stacking/packaging that only three companies can clear, and demand drivers tied to HBM as the binding constraint on every AI accelerator rather than to elastic consumer cycles. Patrick takes the AGAINST side: long-term agreements and SCAs signal a commodity in a strong cycle, not a structural rerating; nearly every relevant memory standard — DDR5, MRDIMM, HBM3/3E/4, LPDDR5X/6, GDDR6/7, LPCAM2 — is JEDEC-standard and therefore commodity at the pin; and CXMT's China DDR5 production ramps in 2H 2026 with Lenovo already shipping and HP and Dell qualifying. Custom HBM4 and Qualcomm-style HBC are where strategic memory genuinely lives. (The Flip)   NVIDIA's $25B Investment-Grade Bond Offering: NVIDIA priced a $25B multi-tranche bond offering on June 15, its first investment-grade debt sale since 2021, with seven tranches maturing between 2028 and 2056 and $85B in orders against an initial $20B target. Dan reads it as raising when capital is cheap, and oversubscription is real. NVIDIA doesn't need the money, it has a gold balance sheet, and is establishing a credit benchmark rather than funding CapEx. Pat agrees the optics are clean, but flags the irony of NVIDIA, with negative debt, borrowing while the stock trades like dead money at a sub-20x forward P/E. Both note that NVIDIA's underperformance reflects the market's skepticism on memory-as-strategic and on NVIDIA's own capex pace relative to the buildout opportunity ahead. (Bulls & Bears)   Tim Cook Calls Apple's Memory Crunch Price Raises on MacBook and iPad "Unsustainable": Apple announced MacBook and iPad price increases of up to $300, with Tim Cook telling the WSJ the memory cost environment is unsustainable. AAPL fell ~5% on the news, the broader rally was momentarily wiped out before Micron held the gains by close. Dan frames it as a moment when the market saw who is going to pay for the AI buildout: the consumer. He notes Apple's pricing power and inelasticity test is now live. Pat traces the backstory to Apple's negative-margin pricing pressure on Micron during the 2022-2023 memory downturn. The question is whether consumer-price blowback will eventually flow back to the memory vendors. (Bulls & Bears)   Micron Blows the Doors Off Fiscal Q3 — $41.46B Revenue, 84.9% Gross Margin: The memory story continues as Micron reported its largest beat in company history with fiscal Q3 revenue of $41.46B versus a $35.69B consensus, EPS of $25.11, year-over-year growth of more than 340%, and a record 84.9% gross margin that is roughly 10 points above NVIDIA's. Q4 guidance came in at a $50B midpoint against a $43B consensus. The 16 multi-year strategic customer agreements add up to $22B in committed volume, with most contracts containing pricing floors but no ceilings on most of the volume — a structurally asymmetric setup. Pat notes 95% of the beat came from price, not units, which reinforces his commodity argument; Dan flips it as the early innings of an NVIDIA-style run that puts Micron's 2027 profit on par with Google. (Bulls & Bears)   Cerebras' First Earnings Report Since IPO — Revenue Doubles, Margins Compress: Cerebras (CBRS) reported its first earnings as a public company, doubling year-over-year revenue and beating the top line while missing EPS, but the stock sold off hard amid gross margin deterioration. Core revenue came in at $191M, up 12% sequentially, with a $194M Q2 guide that is essentially flat, core gross margins at 47% guiding to 36-38% and 38-41% for the year, and operating margins flipping from positive 2% to a guided -30% to -32%. Customer concentration is shifting from Core42 and G42 (86% of FY25 revenue) to OpenAI, which loaned Cerebras $1B and gets paid quarterly in warrants. Pat flags that Cerebras' uncontested speed claim is no longer uncontested with Groq, TPU v8i, and Tenstorrent putting up real numbers. Cathie Wood is down 52% on her position. (Bulls & Bears) Watch the full video at sixfivemedia.com, and be sure to subscribe to our YouTube channel so you never miss an episode. The Decode Qualcomm Investor Day Lands the Data Center Pivot — Microsoft Deploying Qualcomm HBC XPUs in Azure (Per Satya Nadella) + Meta MOU on Three New Qualcomm Datacenter CPUs (Per Zuckerberg); $3.9B Modular Acquisition; Dragonfly Brand + AI200/AI250 Roadmap; HUMAIN 200MW Ramp; Qualcomm to Become Largest Automotive Silicon Company; Targets $3B Datacenter Revenue FY27, $35B by FY31 https://finance.yahoo.com/markets/stocks/articles/qualcomm-investor-day-detail-data-163247063.html  OpenAI Begins Vertical Integration — First Custom Inference Chip "Jalapeño" Unveiled With Broadcom June 24 (Hock Tan: As Good as Blackwell + TPU; ~50% Cost Savings; Late-2026 Microsoft Deployment, 10GW Multi-Gen Roadmap); Daybreak Cyber Stack (June 22) Confirms the Platform Shift https://x.com/OpenAI/status/2069770172802773292  Frontier AI Labs Are Now Financing Their Own Supply Chains — Anthropic Locks In Multi-Year Micron HBM/DRAM/SSD Supply + Micron Becomes Series H Investor; Same Pattern as Samsung + SK hynix Pre-Funded Anthropic in May; $965B Post-Money, $47B Revenue Run-Rate, October IPO Target https://investors.micron.com/news-releases/news-release-details/micron-and-anthropic-announce-strategic-agreement-scale-next  SpaceX Signs $6.3B Compute Deal With Reflection AI — $150M/Month July 2026 → End of 2029; NVIDIA GB300 + Colossus 2 Capacity; SpaceX Now Largest Commercial AI Infrastructure Provider With $80B+ Committed Compute Revenue Through 2029 https://finance.yahoo.com/technology/ai/articles/spacex-reportedly-grant-reflection-ai-162749237.html  The Sovereign AI Stack Lands — Japan's Sakana Ships Fugu + Fugu Ultra Multi-Agent System (June 22) That Beats Opus 4.8, GPT-5.5, and Gemini 3.1 Pro on 10 of 11 Benchmarks; Designed Around US Export-Control Risk; Completes the Three-Bloc Sovereign-AI Map With Mistral Compute (Europe) + DeepSeek $7.4B (China) https://www.datacamp.com/blog/sakana-fugu  The Flip Is the Era of Memory as a Commodity Over? FOR: Memory is now strategic AI infrastructure with multi-year supply lock-ins. The cycle dynamics that defined the last 30 years no longer apply. https://www.benzinga.com/markets/tech/26/06/60062500/micron-earnings-could-echo-nvidias-2023-moment-says-futurum-ceo  AGAINST: Memory is cyclical and priced for perfection. This print is either step change or top of the cycle, and the second one is more likely. https://www.cnbc.com/2026/06/25/apple-macbook-ipad-price-hike-memory.html Bulls & Bears NVIDIA (NVDA) $25B Bond Sale Anchors the AI Debt-Finance Boom — First Bond Offering Since 2021; Joins Alphabet $80B, Amazon $27.5B, Meta $30B, Oracle Stack; Dan: "Locking In Cheap Capital While It Can" https://finance.yahoo.com/technology/ai/articles/nvidia-record-us-25-billion-131039687.html  Apple (AAPL) Falls −5%+ Thursday June 25 on Confirmed MacBook + iPad Price Hikes — Tim Cook RAM "Unsustainable" Comment Lands as Real Price Action; Apple Hikes Erase Micron-Driven Tech Rally Mid-Session; Memory Beneficiaries (SanDisk, Micron) Surge; Analysts "Mostly Nonplussed" https://tickerspark.ai/market/apple-inc-aapl-drops-5-3-as-price-hikes-spook-investors-1782399950638  Micron (MU) Q3 FY26 ACTUALS — Largest Beat in Company History; Revenue $41.46B (+346% YoY) Crushes $35.69B Consensus; Non-GAAP EPS $25.11 (+1,215% YoY) Beats $20.49; Record 84.9% Gross Margin (Higher Than NVIDIA); Q4 Guide $50B Midpoint vs $43B Consensus; Stock +18-19% Overnight to $1,242 https://www.nasdaq.com/articles/nvda-who-micron-blows-doors-q3-earnings-revs  Cerebras Systems (CBRS) Q1 ACTUALS — First Earnings Post-IPO; Revenue $193.4M Nearly Doubled YoY; 2026 Guide $855-$865M Beats $824M; BUT Gross Margins Forecast 38-41% (Down From 45% Q1, Half of NVIDIA + Micron); Stock −20% AH on Margin Compression; Sets Up Inference-Tier Margin Debate https://investors.cerebras.ai/news-releases/news-release-details/cerebras-systems-announces-strong-first-quarter-2026-results  

UBC News World
Remote Contrast Supervision Benchmarks: What The Top Imaging Centers Track

UBC News World

Play Episode Listen Later Jun 29, 2026 11:02


Discover the four critical benchmarks top imaging centers track for remote contrast coverage—response time, reaction management, documentation, and cancellation rates—and why CMS's permanent 2026 virtual supervision rules demand better measurement. Learn more at https://www.contrast-connect.com/blog-post/remote-contrast-coverage-performance-benchmarks-for-imaging-networks ContrastConnect City: Las Vegas Address: Las vegas Website: https://www.contrast-connect.com/

No Priors: Artificial Intelligence | Machine Learning | Technology | Startups
Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI Research Scientist Noam Brown

No Priors: Artificial Intelligence | Machine Learning | Technology | Startups

Play Episode Listen Later Jun 26, 2026 36:18


When a new AI model drops, it's judged based on a static benchmark grid that doesn't account for how long the model is allowed to think. How then should we measure a model's true capability? OpenAI research scientist Noam Brown returns to talk with Sarah Guo about his latest essay on why the AI industry's traditional benchmark grids are broken, and how large-scale test-time compute is fundamentally changing how models are evaluated. Noam explains how, if properly scaffolded, today's models can reason for weeks or even months on complex tasks. He also discusses real-world implications of test-time compute, from building poker solver bots to disproving legendary math conjectures. Together, they also unpack the large gaps in current AI safety frameworks, explore the bottlenecks for recursive self-improvement, and look ahead at the future of multi-agent collaboration and global knowledge sharing. Read more: Implications of Large-Scale Test-Time Compute Sign up for new podcasts every week. Email feedback to show@no-priors.com Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @polynoamial | @OpenAI Chapters: 00:00 – Cold Open 00:43 – Noam Brown Introduction 01:23 – Why Benchmarks Are Broken 04:19 – Compute Budgets and Projections 05:34 – How Long Should Models Think? 06:47 – Benchmark-Maxxing 08:34 – Using Poker Bots as Evals 11:26 – Safety Evals When Model Capability Scales With Budget  14:41 – Release Cycle vs. Agent Runtime  17:06 – Latent Model Capability  20:59 – Limits on Recursive Self-Improvement 27:09 – Large-Scale Multi-Agent Coordination  29:11 – Competition at the Frontier  31:51 – Breaking the Benchmark Grid Equilibrium  33:29 – Why Benchmarks Should be Evaluated by Cost 36:18 – Conclusion

The Full Nerd
Episode 404: The Steam Machine Curse, Intel Arc G3 Extreme Benchmarks & More

The Full Nerd

Play Episode Listen Later Jun 23, 2026 136:10


Join The Full Nerd gang as they offer level-headed takes about the latest PC building news. In this episode the gang covers Adam's testing of the Intel Arc G3 Extreme processor inside the MSI Claw 8 EX AI+, all the reviews of the $1000 Steam Machine, and more. And of course we answer questions live! Timecodes: (00:00:00) Intro (00:06:41) Arc G3 Extreme benchmarks (01:01:36) Steam Machine reviews (01:50:13) Q&A (02:06:53) Outro Links: - MSI Claw 8 EX AI+ teardown: https://youtu.be/7-XY3S15IBg?si=RJsxgKoJixRVMO4i - Intel G3 Extreme Computex discussion: https://youtu.be/dmbzW9uUs_g?si=wbqKRFl6nxDDfGZ9 - @DigitalFoundry 's Steam Machine review: https://youtu.be/WhWtLi_FqLo?si=ujM1hnHkVPzaRQjk - Build your own Steam Machine: https://www.theverge.com/games/953411/valve-steamos-desktop-nvidia - DIY Steam Machine equivalent: https://www.pcworld.com/article/3173750/build-your-own-steam-machine-for-under-900.html Join the PC related discussions and ask us questions on Discord: https://discord.gg/UWhjwg778a Follow the crew on X and Bluesky: @AdamPMurray @BradChacos @MorphingBall Some links may contain affiliate links, which means if you buy something PCWorld may receive a small commission. ============= Read PCWorld! Website: http://www.pcworld.com Newsletter: http://www.pcworld.com/newsletters/signup ============= Learn more about your ad choices. Visit megaphone.fm/adchoices

The HERD FIT
Why do we need Benchmark workouts?

The HERD FIT

Play Episode Listen Later Jun 21, 2026 32:34 Transcription Available


We break down Bison Benchmark 28 and why two scored runs with short rests turn a “simple” workout into a real test of pacing, recovery, and grit. We also explain how to pick the right level, track your data, and train smart all summer so the retest proves you actually improved.• The format of Benchmark 28 and why two run times matter• How to choose RPE for the first run without ruining the second• Why levels exist and how scaling protects the intended stimulus• Benchmarks as a gauge for progress instead of a leaderboard obsession• How winners are determined by biggest percentage improvement• Intentional warm-ups and skill focus as the fastest path to better scores• Practical pacing using reps per minute and tighter transitions• Barbell complex work and pull-up volume ideas for the retest• A caution on adding extra running and the overtraining trapBe on the lookout for next week's episode.

Everyday AI Podcast – An AI and ChatGPT Podcast
Ep 802: ChatGPT's Task Comeback, Claude's Design upgrade, Codex Copies your workflow and 7 other Fresh AI features you'll Want to use Today

Everyday AI Podcast – An AI and ChatGPT Podcast

Play Episode Listen Later Jun 19, 2026 36:47


ChatGPT tasks are back, Jack. ✅While we were collectively ping-ponging the Anthropic vs. U.S. government saga, the big tech AI players rolled out a TON of fresh AI features that are available today. ↳ Claude Design got a big upgrade↳ Google Vids got some serious AI sparkle↳ And there's a new Open Weights model king We'll break it all down. ChatGPT's Task Comeback, Claude's Design upgrade, Codex Copies your workflow and 7 other Fresh AI features you'll Want to use Today -- An Everyday AI Chat with Jordan WilsonNewsletter: Sign up for our free daily newsletterMore on this Episode: Episode PageToday's Episode on LinkedIn: Thoughts on this? Join the convo on LinkedIn and connect with other AI leaders.Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineupWebsite: YourEverydayAI.comEmail The Show: info@youreverydayai.comConnect with Jordan on LinkedInTopics Covered in This Episode:ChatGPT Pulse Sunsetting and Tasks Comeback2. ChatGPT Scheduled Tasks Features and Access Tiers3. Claude Design June Update Overview4. WYSIWYG Editing and Design System Imports in Claude Design5. Claude Design Export Options and Third-Party Integrations6. Google Vids AI Avatars Upgrade with Veo 3.17. OpenRouter Fusion Multi-Model Synthesis Feature8. Claude Code Artifacts for Team and Enterprise Plans9. GLM 5.2 from ZAI Open Weights Model Overview10. GLM 5.2 Benchmarks and Enterprise Use Cases11. OpenAI Codex Record and Replay Feature Explained12. Codex Record and Replay vs. Traditional RPA ToolsTimestamps:00:00 Intro: 7 new AI features you can use today02:35 ChatGPT Tasks: Pulse is gone, Tasks are back04:31 Who has access to ChatGPT Tasks08:18 Claude Design June update overview09:14 WYSIWYG editing and Claude Code integration12:07 Claude Design export options and third-party integrations15:21 Google Vids AI Avatars upgrade18:15 OpenRouter Fusion multi-model synthesis21:54 Claude Code Artifacts for teams25:36 GLM 5.2 from ZAI open weights model29:02 OpenAI Codex Record and ReplayKeywords: ChatGPT Tasks, ChatGPT Pulse, OpenAI, scheduled tasks, proactive AI agent, Claude Design, WYSIWYG editor, Claude Code, design system import, PowerPoint export, Google Vids, AI avatars, Veo 3.1, Gemini 3.1 Flash, OpenRouter Fusion, model fusion, multi-model synthesis, Claude Code Artifacts, Claude Team plan, GLM 5.2, ZAI, open weights, MIT license, mixture of experts, Codex Record and Replay, RPA, workflow automation, Artificial Analysis, Hugging Face, Canva integrationSend Everyday AI and Jordan a text message. (We can't reply back unless you leave contact info) Start Here ▶️Not sure where to start when it comes to AI? Start with our Start Here Series. You can listen to the first drop -- Episode 691 -- or get free access to our Inner Cricle community and all episodes: StartHereSeries.com Also, here's a link to the entire series on a Spotify playlist. 

Be More Than A Fiduciary
FF5 #103 - Meaningful Benchmarks for Fiduciary Excellence

Be More Than A Fiduciary

Play Episode Listen Later Jun 19, 2026 11:51


Meaningful benchmarks can make or break your fiduciary process—and even land you in litigation if you get them wrong. In this Friday Fiduciary Five, Eric breaks down how to choose benchmarks that truly align with your investment policy, target date funds, and fiduciary duty.Connect with Eric Dyson: Website: https://90northllc.com/Phone: 940-248-4800Email: contact@90northllc.com LinkedIn: https://www.linkedin.com/in/401kguy/ The information contained herein is general in nature and is provided solely for educational and informational purposes.It is not intended to provide a specific recommendation of any type of product or service discussed in this presentation or to provide any warranties, financial advice, or legal advice.The specific facts and circumstances of all qualified plans can vary, and the information contained in this podcast may or may not apply to your individual circumstances or to your plan or client plan specific circumstances.The opinions expressed by guests on the Be More Than a Fiduciary podcast are not necessarily the same as the opinions held by 90 North Consulting, or of Executive Director Eric Dyson.

Be More Than A Fiduciary
Bonnie Treichel: Meaningful Benchmarks

Be More Than A Fiduciary

Play Episode Listen Later Jun 17, 2026 44:39


When does a benchmark actually become “meaningful” — and what does that have to do with your retirement committee meetings? In this episode, Eric and ERISA attorney Bonnie Treichel unpack retirement sketchbooks, DOL proposed regs, and how fiduciaries can align process, purpose, and benchmarks without getting lost in the legal weeds.In this episode, Eric and Bonnie Treichel discuss:Purpose and design of Your Retirement SketchbookMaking money conversations a “dinner table” topicBenchmarks and “meaningful benchmarks” in retirement plans3(21) vs. 3(38) fiduciary roles and investment policy statementsDOL proposed regulations, litigation trends, and action items for committeesKey Takeaways:Retirement conversations don't have to be intimidating; using accessible, bite-sized topics can turn money into a normal “dinner table” discussion across generations.An investment policy statement is only useful if it reflects reality; committees must periodically review it and ensure their actual practices match the documented process.Benchmarks are not just numbers on a report; selecting and understanding the right benchmark is central to evaluating performance and defending fiduciary decisions.Delegating to a discretionary investment manager does not eliminate responsibility; plan sponsors still “own” the policy and must prudently select, monitor, and understand their 3(38) relationship.Prudence is about process, and loyalty is about purpose; without both, even technically sound procedures can fail participants if they aren't anchored to what's right for that specific plan and its people.“The big action item is to look at your investment policy statement and see if it says anything about what benchmark is being used. Number two, look at your actual investment report and see, okay, what are the benchmarks being used?” - Bonnie TreichelBonnie's passion is sharing her knowledge with financial advisors. When she founded Endeavor Retirement, her goal was to make retirement legislation easy to understand. She keeps advisors up to date on the rules and regulations through her webinars, presentations, and consultations. The result — advisors and consultants help more people access their retirement savings.Connect with Bonnie Treichel:Website: https://endeavor-retirement.com/ LinkedIn: https://www.linkedin.com/in/bonnietreichel/ Connect with Eric Dyson: Website: https://90northllc.com/Phone: 940-248-4800Email: contact@90northllc.com LinkedIn: https://www.linkedin.com/in/401kguy/ The information and content of this podcast are general in nature and are provided solely for educational and informational purposes. It is believed to be accurate and reliable as of the posting date, but may be subject to change.It is not intended to provide a specific recommendation for any type of product or service discussed in this presentation or to provide any warranties, investment advice, financial advice, tax, plan design, or legal advice (unless otherwise specifically indicated). Please consult your own independent advisor as to any investment, tax, or legal statements made.The specific facts and circumstances of all qualified plans can vary, and the information contained in this podcast may or may not apply to your individual circumstances or to your plan or client plan-specific circumstances.The opinions expressed by guests on the Be More Than a Fiduciary podcast are not necessarily the same as the opinions held by 90 North Consulting, or of Executive Director Eric Dyson.

Ransquawk Rundown, Daily Podcast
EU Market Open: Energy benchmarks towards lows, Europe set for muted open with markets tentative into Fed announcement

Ransquawk Rundown, Daily Podcast

Play Episode Listen Later Jun 17, 2026 2:10


The US will allow Iran to immediately begin selling oil and fuel under the deal to end the war, offering Tehran an early financial incentive to wind down the conflict, WSJ reports, citing sources.Iran and Oman are already talking about how they will manage the Strait of Hormuz, as laid out in one of the reported points. They want to charge a "fee" for the management "services”, NYP reports, citing sources.IDF is prepared to stay in southern Lebanon for a significant period, Kann News reports. Reports of artillery shelling have continued in southern Lebanon.The US delayed the blacklisting of China's DeepSeek and over 100 Chinese firms deemed national security risks, to avoid escalating tensions with Beijing, Reuters reports, citing sources.APAC stocks were mixed, whilst US equity futures are indicative of a slightly weaker open.DXY trades tentatively as markets await today's Fed policy decision, and the debut of Chair Warsh; CHF incrementally leads, whilst the Kiwi lags.Looking ahead, highlights include UK Inflation Report (May), ECB Wage Tracker (Jun), EU Inflation Final (May), US Retail Sales (May), Atlanta Fed GDP (Q2), New Zealand GDP (Q1), Riksbank Policy Announcement, Fed Policy Announcement, BCB Policy Announcement, IEA OMR. Speakers include ECB's Cipollone & Lagarde, Riksbank's Thedeen & Fed's Warsh. Supply from Australia & Germany.Read the full report covering Equities, Forex, Fixed Income, Commodites and more on Newsquawk

The Wellness Mama Podcast
Resilience and Adaptability: Real Benchmarks of Health (Solo Episode)

The Wellness Mama Podcast

Play Episode Listen Later Jun 15, 2026 24:06 Transcription Available


Episode Highlights With KatieWhy resilience and adaptability...not restriction...are the true markers of vibrant health.How rigid diets and “perfect routines” often reflect a dysregulated nervous systemThe mindset and language shifts that changed your health from the inside out.The nervous system foundations that created real healing capacity.How gradually expanding inputs taught your body it was safe again.Why metabolic flexibility is impossible without nervous system flexibility.The identity-level transformation required to step into freedom.Practical steps you can use to build resilience and adaptability starting today.Resources MentionedLMNT mineralsSaunaBioptimizersI love and use so many products from them, but I especially love the magnesium (Magnesium Breakthrough) and digestive enzymes (Masszymes). Visit bioptimizers.com/wellnessmama to get the best deal!

Ransquawk Rundown, Daily Podcast
US Market Open: Risk assets rally as energy benchmarks soften on US-Iran framework peace agreement

Ransquawk Rundown, Daily Podcast

Play Episode Listen Later Jun 15, 2026 2:10


The US and Iran have reached a framework peace agreement; the US will lift its naval blockade, whilst the Iranians will reopen the Strait of Hormuz. The Pakistani PM suggested it would be signed in person on Friday, 19th June; Brent Aug'26 -4.5%.The deal includes the termination of military operations on all fronts, including in Lebanon. Israel's Katz said that they would not withdraw from Lebanon. US equity futures bid amid the constructive risk tone; NQ +1.9% DXY pressured as markets pare hawkish Fed pricing, ahead of Fed Chair Warsh's first meeting.Fixed income benchmarks firmer but off best levels, as yield curves bull steepen.Looking ahead, highlights include US Industrial/Manufacturing Production (May).Read the full report covering Equities, Forex, Fixed Income, Commodites and more on Newsquawk

Ransquawk Rundown, Daily Podcast
US Market Open: Crude benchmarks hit on Mehr MoU reporting, equities bid into SPCX debut

Ransquawk Rundown, Daily Podcast

Play Episode Listen Later Jun 12, 2026 2:50


Iran's Mehr News reports that the US-Iran MoU includes the reopening of the Strait of Hormuz, lifting oil sanctions, and releasing frozen Iranian funds. The draft is still being reviewed. Brent -4.1%.The US and Iran deal signing could occur around the June G7 meeting in Geneva (June 15th - 17th).Global equities gain on the constructive risk tone, SPCX set to debut today.DXY rangebound, EUR holds above 1.1580 despite somewhat conflicting ECB reports. Fixed income benchmarks benefit from the softer energy prices.Looking ahead, highlights include Canadian Wholesale Sales (Apr), US UoM Prelim. (Jun) & SpaceX Debut.Read the full report covering Equities, Forex, Fixed Income, Commodites and more on Newsquawk

Business By The Numbers
Auto Shop Benchmarks: Biggest Changes, Key KPI's and Trends [E226]

Business By The Numbers

Play Episode Listen Later Jun 11, 2026 44:49


Thanks to our partners Promotive, WickedFile, Maverick Shop Owners, and OverdryveYour technicians got more productive in 2025. Your labor rate went up. Your parts margins improved. So why is the average shop owner keeping almost exactly the same percentage of every dollar as they did the year before?The answer is hiding in your effective labor rate — and most shop owners haven't looked at it once.In this episode, Hunt Demarest breaks down the headline findings from Paar Melis & Associates' 2026 Auto Shop Benchmark Report — the largest study of its kind, built from more than 200 real shop locations across the country. From average repair order trends to technician productivity, overhead creep, benefits adoption, pay structure breakdowns, shop management software rankings, and the labor rate shame list nobody wants to be on, this is the financial state-of-the-industry episode you didn't know you needed. And it's only Part One.What You'll Learn(00:00) Intro — the 2026 Benchmark Report is live and how to get your free copy(03:16) Who made this report possible — methodology, participation, and what Paar Melis clients get that nobody else does(05:30) How to read benchmark numbers without misleading yourself — context, outliers, and the ARO trap(13:15) Sales are up 10-11% — but how much of that is real production vs. a labor rate increase you already gave yourself?(16:20) Productivity jumped from 47% to 55% — so why didn't net profit follow?(18:30) Effective labor rate: the silent margin killer hiding in plain sight in 2025(27:30) Benefits adoption hits an all-time high: 73% of shops now offer health insurance(29:30) Retirement plans, tool reimbursement, Trump Accounts, and the fringe benefits arms race(30:45) Four-day work weeks and non-cash comp — how shops are winning the talent war without raising base pay(34:30) Good management makes money, not good pay plans(36:30) 83% of shops are doing digital vehicle inspections — Hunt thought it would be closer to 100%(37:30) Shop management software rankings: Tekmetric at 56%, Mitchell likely on the way out, Shopware and Protractor tied for third(40:30) The labor rate shame list — 11% of shops haven't raised their rate in 18 months or moreIf you're ready to stop guessing where your numbers stand, start benchmarking against 200+ real shops across the country, and finally understand why doing more work doesn't always mean making more money — this episode is essential listening.Get the FREE 2026 Auto Shop Benchmark Report: https://hubs.ly/Q04j-grh0Thanks to our partner, PromotivePromotive has over 40 years of recruiting and automotive experience. If you need qualified technicians and service advisors and want to offload the heavy lifting, visit https://gopromotive.com/Thanks to our partner, WickedFileTurn chaos into clarity with WickedFile, the AI for auto repair shops. Transform invoices into insights, protect cash flow, and stop losing parts, cores, or credits to maximize your bottom line. visit https://info.wickedfile.com/Thanks to our partner, Maverick Shop OwnersYou're working on growing a more profitable shop - that's critical. That's exactly what the 24-video Blueprint course by Maverick Shop Owners addresses - customers, sales, profit, people, systems, and freedom. Get free access for our listeners only at https://maverickshopowners.com/blueprintThanks to our partner, OverdryveOverdryve is your AI-powered marketing operating system. It predicts slow weeks before they happen, automatically launches revenue-driving campaigns, tracks ROI down to the dollar, and optimizes performance in real time. Visit https://overdryvemarketing.com/Paar Melis and Associates – Accountants Specializing in Automotive RepairVisit us Online: www.paarmelis.comEmail Hunt: podcast@paarmelis.comText Paar Melis @ 301-307-5413Download a Copy of My Books Here:Beyond the Bays: A Financial Playbook for Auto Repair Shop OwnersWrenches to Write-OffsYour Perfect Shop The Automotive Repair Podcast Network: https://automotiverepairpodcastnetwork.com/Remarkable Results Radio Podcast with Carm Capriotto: Advancing the Aftermarket by Facilitating Wisdom Through Story Telling and Open DiscussionDiagnosing the Aftermarket A to Z with Matt Fanslow: From Diagnostics to Metallica and Mental Health, Matt Fanslow is Lifting the Hood on Life.The Weekly Blitz with Chris Cotton: Weekly Inspiration with Business Coach Chris Cotton from AutoFix - Auto Shop Coaching.Speak Up! Effective Communication with Craig O'Neill: Develop Interpersonal and Professional Communication Skills when Speaking to Audiences of Any Size.Business by the Numbers with Hunt Demarest: Understand the Numbers of Your Business with CPA Hunt Demarest.The Auto Repair Marketing Podcast with Kim and Brian Walker: Marketing Experts Brian & Kim Walker Work with Shop Owners to Take it to the Next Level.

Ransquawk Rundown, Daily Podcast
US Market Open: Energy benchmarks weaker as US-Iran diplomacy continues, ECB and US PPI due

Ransquawk Rundown, Daily Podcast

Play Episode Listen Later Jun 11, 2026 2:14


The US and Iran exchanged another round of strikes overnight, resulting in Iran announcing the complete closure of the Strait of Hormuz, effective immediately, and threatening to hit any vessel crossing the Hormuz.However, an Iranian source told Reuters that Iran and the US are still in negotiations over a preliminary deal, which includes a mechanism for unfreezing funds. US equity futures pare Wednesday's losses ahead of SPCX IPO pricing.DXY flips across the 100.00 handle; EUR muted ahead of ECB policy announcement.Fixed income muted, US 10yr remains above 4.50% with PPI ahead. Crude futures reverse earlier gains amid positive reports of continued US-Iran negotiations.Looking ahead, highlights include US PPI (May), Jobless Claims (May/30), ECB Policy Announcement (Jun), CBRT Policy Announcement (Jun), OPEC MOMR (Jun), Comments from ECB President Lagarde, Supply from the US and Earnings from Adobe.Read the full report covering Equities, Forex, Fixed Income, Commodites and more on Newsquawk

Behind the Steel Curtain: for Pittsburgh Steelers fans
BAD Language: Receiving Benchmarks for the Steelers in 2026

Behind the Steel Curtain: for Pittsburgh Steelers fans

Play Episode Listen Later Jun 8, 2026 21:29


Bryan Anthony Davis discusses his hope for pass catchers in 2026. Check out this and more on his solo show, BAD Language. Steel Curtain Network is courtesy of the Fans First Sports Network. Check out Meinelschmidt Distillery at meineldistillery.com and use the code SCNJUN to save 10% at checkout! Learn more about your ad choices. Visit megaphone.fm/adchoices

The Generative AI Meetup Podcast
The Best Open Source US Model (Right behind China)

The Generative AI Meetup Podcast

Play Episode Listen Later Jun 7, 2026 114:55 Transcription Available


https://novacut.ai/  https://genaimeetup.com/  Anthropic has officially closed a $65 billion Series H at a $965 billion valuation, nearly 2.5x its valuation from just 100 days ago. Meanwhile, funding is flowing across the ecosystem: Frameworks AI at $15B, Baseten at $11B, OpenRouter's $113M Series B, and Cognition AI's $1B Series D. NVIDIA went on an open-source super week with Nemotron 3 Ultra, Cosmos 3, and Nemotron 3.5 ASR. Microsoft dropped 5 new MAI models. Google released Gemma 4 12B, and Anthropic shipped Opus 4.8. On the benchmarks front, DeepSWE crowns GPT-5.5 as the leader in long-horizon coding tasks, while ITBench shows even frontier models struggle with real-world SRE incidents — Claude Opus 4.7 tops out at just 47%. Plus: Cloudflare acquires VoidZero to build the future of AI-native edge development, and Google is paying SpaceX $920M/month for compute. Topics covered: • Anthropic's $65B Series H and path to $1T • Fireworks AI, Baseten, OpenRouter & Cognition funding rounds • Microsoft's 5 new MAI models • NVIDIA's open-source super week (Nemotron, Cosmos 3) • MiniMax M3, Gemma 4 12B, JetBrains Mellum2, Opus 4.8 • DeepSWE benchmark: GPT-5.5 leads long-horizon coding • ITBench: Frontier models under 50% on real SRE tasks • Cloudflare + VoidZero for AI-native edge dev • Google's $920M/month SpaceX compute deal #AI #Anthropic #NVIDIA #OpenAI #AInews #TechNews #LLM     Funding rounds Anthropic formally confirmed the closure of its $65 billion Series H funding round at a post-money valuation of $965 billion. This represents a 2.5-fold increase over its $380 billion Series G valuation from February 2026, adding $585 billion in value in approximately 100 days https://www.anthropic.com/news/series-h  Frameworks AI raising at 15B valuation representing a near fourfold increase from its $4 billion Series C valuation recorded in October 2025 processing 15 trillion tokens daily for major production clients including Cursor, Notion, and Perplexity https://finance.yahoo.com/sectors/technology/articles/fireworks-ai-eyes-15-billion-174609357.html Baseten is raising 1B at 11B valuation annualized revenue, which skyrocketed from $200 million to $600 million over a single quarter https://techstartups.com/2026/05/26/ai-inference-startup-baseten-in-talks-to-raise-1-billion-at-11-billion-valuation/  OpenRouter has secured a $113 million Series B funding OpenRouter has experienced exponential traffic growth, with weekly production throughput expanding fivefold from 5 trillion to 25 trillion tokens over a six-month horizon https://www.businesswire.com/news/home/20260526953416/en/OpenRouter-Raises-%24113-Million-CapitalG-led-Series-B-as-Weekly-Volume-Explodes-to-25T-Tokens  Further up the stack: Cognition AI secured a $1 billion Series D round led by Lux Capital and 8VC https://cognition.ai/blog/series-d   Model Releases MAI models: MAI-Code-1-Flash: A 5-billion active parameter model optimized for ultra-low latency within GitHub Copilot and VS Code. MAI-Image-2.5: A high-fidelity image generation model ranking third on global image evaluation arenas, outperforming competing architectures like Nano Banana Pro. MAI-Transcribe-1.5: A multi-lingual speech processing engine offering fivefold speed improvements across 43 languages. MAI-Voice-2: Natural audio and voice generation across 15 languages, available at a highly competitive price point. Web IQ: A search-grounding API engineered to directly compete with Perplexity. https://microsoft.ai/models/    https://www.peoplematters.in/news/ai-and-emerging-tech/uber-imposes-dollar1500-monthly-ai-spending-limit-on-employees-amid-rising-costs-50073    Nvidia has executed an "Open-Source Super Week," positioning itself as a dominant software and model publisher: Nemotron 3 Ultra (best US open source open weights model but behind china): A massive 550-billion parameter MoE (55 billion active) designed with a 1-million token context window, optimized specifically for high-throughput, cyclical agent loops. It achieved peak throughput rates of 400 tokens per second on day-zero optimized clusters. Cosmos 3: A physical AI world-modeling framework comprising 16-billion Nano and 64-billion Super variants. Built on a Mixture-of-Transformers (MoT) architecture, Cosmos 3 natively binds textual, visual, auditory, and physical kinetic vectors. Nemotron 3.5 ASR: A highly compact 0.6-billion parameter streaming speech recognition model pushing sub-100 millisecond latencies across 40 language locales.   https://www.minimax.io/models/text/m3  MiniMax M3: A 1-million token context model hitting 59.0% on SWE-Bench Pro and 74.2% on MCP Atlas, though noted for high token consumption due to intensive internal self-validation loops.   https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12b/  Gemma 4 12B: Google's Apache 2.0 on-device model, which utilizes an encoder-free architecture that projects vision and audio vectors directly into the text-token space, bypassing separate CLIP-style encoders to minimize local memory footprints. https://www.jetbrains.com/mellum/  JetBrains Mellum2: A compact 12-billion parameter MoE (2.5 billion active) engineered for ultra-low latency routing and retrieval-augmented generation (RAG) sub-agents within developer IDEs. Opus 4.8 https://www.anthropic.com/news/claude-opus-4-8    https://www.cnbc.com/2026/06/05/google-to-pay-spacex-920-million-a-month-for-xai-compute-capacity.html      Benchmarks: https://deepswe.d atacurve.ai/blog https://venturebeat.com/technology/deepswe-blows-up-the-ai-coding-leaderboard-crowns-gpt-5-5-and-finds-claude-opus-exploiting-a-benchmark-loophole (GPT 5.5 the winner in long horizon tasks) a highly complex software engineering benchmark focused on original, long-horizon tasks across five distinct programming languages. Comprising 113 chaotic tasks across 91 live, production-grade repositories, DeepSWE forces agents to generate 5.5 times more code and modify an average of 7 separate files per task compared to standard evaluations. On this challenging leaderboard, GPT-5.5 leads with a score of 70%, establishing a significant 16-percentage-point lead over contemporary alternatives I think older benchmarks where models reach ~90% accuracy can be considered saturated. Few percentage points don't give us any good signal.  https://research.ibm.com/publications/developing-ai-agents-for-it-automation-tasks-with-itbench  ITBench-AA, an evaluation framework focusing on live Kubernetes incident response and Site Reliability Engineering (SRE) operations. Comprising 59 live, containerized SRE incident snapshots, the results are remarkably sobering: every frontier model scored under 50% on successful incident resolution, with Claude Opus 4.7 leading at 47% and GPT-5.5 following closely at 46%.   Edge AI announcements: https://www.cloudflare.com/press/press-releases/2026/cloudflare-acquires-voidzero-to-build-the-future-of-the-ai-native-web/  The consolidation of the AI-native developer stack has reached the runtime virtualization layer. Cloudflare recently completed the acquisition of VoidZero, the development group responsible for Vite, Vitest, Rolldown, and Oxc, backing the transaction with a $1 million open-source ecosystem fund. This acquisition is highly strategic; as autonomous agents write an increasing proportion of production software, local development environments, compilation pipelines, and bundlers must be optimized for execution speeds that match agent speeds. Cloudflare's goal is to construct a localized, full-stack edge playground. In this sandbox, AI agents can generate, test, bundle (utilizing the highly parallelized, Rust-based Oxc and Rolldown engines), and deploy entire web applications end-to-end within milliseconds. This architecture completely bypasses traditional local machine container bottlenecks, enabling high-velocity agent loops to execute in a fully sandboxed, web-scale edge runtime.

Cougar Sports with Ben Criddle (BYU)
6-4-26 - Hour 3 - Which benchmarks did A-Rod set for Bear Bachmeier?

Cougar Sports with Ben Criddle (BYU)

Play Episode Listen Later Jun 4, 2026 45:49 Transcription Available


Brett Hammer fills in for Ben Criddle and discusses LJ Martin's chances of winning another Big 12 Player of the Year award, the latest AJ Dybantsa news, all of the news and notes out of Cougar Country, and more! The Deseret News' Jay Drew and BYUtv's Jarom Jordan to the program.

Investing with IBD
Ep. 375 Momentum, Meet Strength: Inside The Forces Moving Today's Economic Benchmarks

Investing with IBD

Play Episode Listen Later Jun 3, 2026 54:07


Family matters here. Mish Schneider, chief strategist at MarketGauge.com, discusses market momentum and its effects on commodities, regional banks and more. She also gives an update on the modern “economic family,” a benchmark that helps investors assess market direction. Tune in to learn about the semiconductor industry's continued strength, AI's longevity and why small caps still matter. Learn more about your ad choices. Visit megaphone.fm/adchoices

The Morning Agenda
PA Headlines | June 1 | Gov. Shapiro's proposed benchmarks for data center developers receive mixed reviews.

The Morning Agenda

Play Episode Listen Later Jun 1, 2026 10:35


Governor Josh Shapiro is pitching details of his plan for managing data center growth,  months after broadly sketching out a strategy in his budget address. Shapiro, who is running for reelection, is calling on state lawmakers to work with his administration to make his proposal into law. It includes a series of benchmarks data center owners would need to meet in order to get tax benefits from the commonwealth.  And a deep dive:Staying with the topic of development – but with a twist...Think about the shingles on your home - are they made of asphalt? Aluminum? Wood? Imagine if those shingles were made of food waste - pineapple peels, egg shells, and shrimp shells. A group of researchers at the University of Pennsylvania are designing building materials that could be healthier for us and the planet.