POPULARITY
Hey, this is Alex, welcome back to your weekly dose of intense AI acceleration summer!My weekend was consumed by thinking about the OpenAI hack and agent swarms, but then the torrent of AI releases took over, and we got back to back news (including 3 breaking news during the live show), with a heavy open source focus!I think the winner of this week is SpaceXAI/Cursor who released 3.5 releases, with one being my highlight of the week, Grok Bot (I've invited Shub Gaur from Cursor to the show to walk us through it) and Grok 4.6 which matches Opus at half the price.There was a LOT of news in open source this week as well, with Meta kicking off with Muse Glimmer 30B and promising Muse Spark 1.2 soon, Qwen dropping Qwen 3.8 open weights and DeepSeek dropping an anvil with an upgraded DeepSeek v4 Pro and MIT license!Let's dive in (and please don't forget as a reader you get 100% off the 1299 ticket to Fully Connected, our 2000 person Al event in SF in Sept, just use THURSDAIFC2026 as your code and see you there!)0:00 The Wildest Week in AI Yet3:45 How OpenAI's Agent Swarm Hacked Hugging Face17:02 The Week in AI: DeepSeek, Qwen, Grok & More25:54 NVIDIA Nemotron 3.5 & Korea's Motif 332:45 DeepSeek V4 Pro, Flash & an Open Harness39:46 Qwen 3.8 Max and Its Missing Vision Tower43:30 What Is Grok Bot? Shub Gaur Explains50:02 Live Grok Bot Demo: House Hunting & Security55:00 Persistent Agents, Yapper & DeepSeek Dropwatch1:04:19 Grok 4.6: Benchmarks, Pricing & Cursor1:15:37 Grok Bot vs. Open-Source Agents1:23:14 Anthropic's Hidden Claude Watermarks1:28:50 Fully Connected & Day-Zero Models on CoreWeave1:32:02 GPT-5.6 Sol at 14x Speed on Cerebras1:37:34 Gemini 3.7 Flash Resets the Cost Curve1:40:51 Inside Artificial Analysis with George Cameron1:45:25 Optima & Choosing the Right AI Model1:55:03 Cost per Task, Caching & Real-World Benchmarks2:05:14 LTX-2.5 and Open-Weight Video2:09:30 Grok Imagine 2.0 & Final TakeawaysGrok Bot and Grok 4.6 from SpaceXAI/CursorFolks, I've previously told you that from 3 frontier labs we noticed a jump to 5, and voila, this week proves that Elon is hell bent to win. After the cursor acquisition, and the integration of all of the parts into SpaceXAI, they have released 2 huge things this weekGrok 4.6 - Ties with GPT 5.6 SOL and half the price and much speed.I've had the pleasure to host Goerge Cameron from Artificial Analysis on the show today, and I asked him, what is the best models. His answer, it's a 3 factor answer, intelligence, speed and cost per task .Well, if you use their nifty “recommend a model“ tool on the homepage, you'll see that Grok 4.6 beats most other models on all of those! But, is it really that good? Models are really hard to evaluate and compare lately. It's definitely a huge step up from Grok 4.5, with 61.3 on Frontier Code (beating Sol and just after Opus 5) and #4 on Apex-agents (+10 points from previous Grok). on Artificial Analysis this model lands at #4 on intelligence, while being #5 on speed all while being half the price of the models that are above itAs far as the tech goes, this model card confirms that it no longer has the Cursor Bench leaked into it's weights and it's #1 on that benchmark! It's the same 1.5T v9 base at the same price, with Elon claiming that 4.7 is going to mog the competition in 3-4 weeks.Everyone has a harness, now everyone has a swarm of bots - My Grok Bot review (x.ai/bot)You guys know all about OpenClaw and Hermes, and Claude CoWork and Codex rebrand, and all of them are trying to nail down the same, always-on, autonomous agents that can do things for you.Hermes and OpenClaw require you to have an always on computer, mess with API keys, Claude Cowork doesn't run on the cloud and ChatGPT work starts a fresh session every time you ask a new thing.Grok Bot (again, awful name) is the first one that seems to nail all of what I want in an always-on agent ... swarm. That's right, this isn't one agent with multiple personalities (like OC, Hermes), there's a bot here for every task, and you dont' have to manage context, queues, API keys (can if you want to) and models.Oh, also ,there's no model picker, it's just Grok 4.6 deciding for ya, and it's really fast!Swarm of bots, working for you, each with their own computerI am not getting paid for this (besides being provided a free account for cursor, but I've had it for 6 months and haven't used), it's really that good, the Cursor folks did some magic there. They picked up the most important parts of personal agents, like the (ios-only) mobile app (app store)You can start a task on your mac, pick it up on your phone, get notified on your phone/mac, and the killer thing is, they are giving your bots their own computer, which can do things (especially if you're ok with logging in there to your accounts!)The kicker for me is the very very well done agent to agent communication there, which is transparent but read only to you. You can ask your bots to spin up other bots, but unlike sub-agents, they are actual bots with their own identity. You can even tag them in other chats and create group chats! There's no context to manage, they do the work for you and so far this wasn't a problem at all.On the model side, Grok 4.6 seems to be doing an excellent job with agentic long running tasks that require coding and computer use, I've just been chatting with the bots and not thinking about any of the things I used for Hermes and OpenClaw.What about Vendor Lock-in? Giving Elon data?Some of these comments our fans raised during the show are very valid, after all, not only is the world divided on Elon Musk (which makes it REALLY hard to judge the models they release just on vibes from X btw, we talk about this constantly) but also, remember that Grok 3 started going off on X and called himself Mechahitler and just recently Grok CLI was caught uploading all of your data to X servers, which was reversed very quickly.Honestly, I think there's a very very good chance that this Grok Bot interface, which is geareed toward the less technical users, folks who don't need the code-diff side pane, and don't know/care what compaction is, and just want agents to do things for them, is goign to win much of this trust back. It just works, truly, for a beta product it's really well executed by whoever worked on this!Security and key managementOne of the best parts for me with this Grok Bot, is that the connectors are the same connectors you use in Cursor! There's a LOT of them (Cursor after all has been one of the first apps to start adding AI agents) and this also means that they take the security very seriously.Every API key that you want to add, is not shown to the bot, each bot lives in an isolated environment, and for stuff like payments and log-ins, it gives you back the control of it's computer for you to complete!I also love this section in settings, which makes auto-approve work for you: you define rules with natural language that you always want the bot to ask you before... sending an email or posting on your behalf or what not.Chief of staff pattern to get startedIn case you're convinced enough to give it a try (it's free trial for 1 month, and the cheaper way to get it is via Cursor's 149$ plan and not via the Grok Ultra plan which is 249), here's a recommended pattern that works very well.Create a chief of staff bot, have it interview you about everything you are doing in your day to day, work and personal, then decide how much permissions you wanna give it, start little.Then ask your chief of staff to create bots for some of the work it can try and help you with, focus on “reduce cognitive load”.And then see the magic come to life. If you have skills or memory from other bots, you can just ... import it in.Then try setting up an automated email checker bot, and have your chief of staff surface only the most important emails you have to actually respond to.Another great pattern is setting up a bot with the last30days research skill (we covered it with Matt Van Horn) and have a research bot for every topic you want to deep dive into.Schrodinger's GrokI haven't quite named it like that, but we've covered all Grok released on the show (tracking 24 on https://thursdai.news/companies/xai excluding this week) and ... it's always very hard to judge Grok model released based on X feed vibes. It's either AI influencers who want Elon to retweet them, glazing the models, or folks who hate Elon for his political views or whatever, ignoring their (truly insane progress).This time, both the model and Grok Bot are getting very very good reviews, from folks like our own Ryan Carson, Lenny Rachitsky, Rubben Hassid and Roberto P Nickson. Not folks who are swayed lightly, but also, yours truly. I really do think there's something great here, worth trying out, especially if you've struggled to maintain your OC/Hermes and want agents to work for you 24/7. LMK if you have questions about it and your experienceOpen Source AI and other newsI want to continue with this new newsletter that covers 1 big story, but I can't leave you uninformed about the most important developments in AI and Open SourceDeepSeek V4 pro 0813 is in GA - MIT licensed chonker with 1.7T parameters (X, Blog, HF, GitHub)The whale is back with a vengeance, DeepSeek resurfaced with their flagship response to Kimi K3 and with MIT license, we can't complain.1M context window, 49B active parameters but it seems to underperform, landing at 54 on the Artificial Analysis leaderboard. However, they did show a significant improvement on DeepSwe (from 12.8 points in the preview version of V4 to 62.7 in this one)We still think it's a good model sir, and definitely worth trying out!Additionally, DeepSeek released their own harness on Github (hitting 23K stars in less than 24 hours) which seems to be exciting as well, give it a try.Meta comes back to open source with Muse Glimmer (30B) and promise to open source Muse Spark 1.2 (X, Blog, HF)We would like to officially welcome back Meta to the open source AI community, as they release their smaller Muse model called Glimmer!The highlights, it runs on a single 24GB consumer GPUs, gets 51 on Swe-bench Pro, beating Qwen 3.6 27B. And with DFlash speculative-decoding, it delivers 233tok/s on RTX 5090.Zuck promised us the bigger Muse Spark 1.2 in open source and published a long essay on superintelligence and that it should be distributed to everyone, which we applaud and it's great to see the commitment reinforced! welcome back Meta!This weeks buzzShort interjection from our only sponsor, CW this week.1 - Join 1500 ai practitioners (and a live ThursdAI recording) at Fully Connected Sep 29-31 in SF - use code THURSDAIFC2026 (Register here)2 - We have day-0 support for Nvidia's latest Nemotron 3.5 lightning (CW Inference)Gemini 3.7 Flash - breaking in the middle of the showJust as we had George Cameron from Artificial Analysis on the show, Gemini dropped Gemini 3.7 Flash, and it's a speedy beast! Clocking at over 300t/s, it's google's mid-tier model, think Sonnet/Terra competitor, that is also great at multimodal (I think it's one of the only ones that can watch videos)It beats Muse Spark 1.2 on DeepSWE and lands near the cost-per-task Pareto frontier on Artificial Analysis. For the cost/speed/intelligence trade-off, this model is now #1 on Artificial Analysis selector of best models!That's a wrapThis was the first week of the shorter newsletter experiment: one big story done properly, and trust that you'll listen to the show for the rest (it's 2.5 hours of exactly this, with demos). Tell me if you hate it. Our release index at thursdai.news tracked 71 releases in July alone, so something had to give, and it wasn't going to be my weekends.See you at Fully Connected Sept 29 (code's in the intro, come say hi to me and Wolfram at Moscone), and if you try the Grok Bot chief of staff pattern, I genuinely want to hear how it goes.ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.ThursdAI - Aug 13, 2026 - TL;DR* Hosts and Guests* Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave (@altryne)* Co-hosts: @WolframRvnwlf, @petergostev, @nisten, @ldjconfirmed, @yampeleg, Chris Alexiuk - NVIDIA (@llm_wizard)* Shub Gaur - Cursor / SpaceXAI, GrokBot (@shubgaur)* George Cameron - Artificial Analysis (@grmcameron)* Big CO LLMs + APIs* xAI Grok 4.6: AA Index 61 at $2/$6 per M, CursorBench 69.9, card confirms self-optimized inference stack (X, Blog, Model card)* Grok Bot early beta: persistent agents with their own computers, macOS + iOS, free with SuperGrok Heavy and Cursor Ultra (X, x.ai/bot)* Breaking: GPT 5.6 Sol ultrafast preview on Cerebras at ~14x speed, work-account waitlist (Blog)* Breaking: Gemini 3.7 Flash, 50% price cut through end of year, near Pareto-optimal cost per task (X)* OpenAI GPT-5.6-Cyber: 95.0% cyber completion vs 1.5% base, gated behind Daybreak Red (X, Blog)* Grok 4.7 teased: 3-4 weeks out (Elon-reply-sourced only) (X)* Open Source LLMs* DeepSeek V4 Pro 0813 weights re-published under MIT: 1.6T/49B active, DeepSWE 62.7 (+49.9), Terminal Bench 2.1 87.9, $0.435/$0.87 per M (X, OpenRouter)* DeepSeek Harness hit 23K GitHub stars in days, web UI (GitHub)* Qwen3.8-Max landed on HF as open weights: 2.4T/95B active MoE, 1M context, FrontierSWE 73.5, custom license (X, HF)* Meta returned with Muse Glimmer 30B agentic, Apache 2.0, SWE-Bench Verified 76.0, Muse Spark 1.2 weights promised (X, Blog, HF)* NVIDIA shipped Nemotron 3.5 Lightning: 30B MoE/3B active, up to 4x output speed, strong voice-agent results (X, HF)* Motif 3 from Korea open-sourced: 314B/13.2B active, MIT, SWE-Bench Verified 76.2 (X, HF)* Cohere North Micro Vision: 2.4B VLM, Apache 2.0, DocVQA 92.1% (X, HF)* Liquid AI LFM2.5-VL-3B: 228 tok/s on M5 Max in ~3GB (X, HF)* AI in Society* Anthropic watermarks all new Claude text output worldwide under EU AI Act Article 50, C2PA on images, detection docs promised (Geiping FAQ, Euronews)* Stolen Thoughts: 704 artifacts including 62 API keys extracted from hidden reasoning across 6,708 sessions (X, Paper)* Pangram: OpenAI holds 50%+ of AI text share, Anthropic triples to 14.9%, Google falls to 1.9% (X, Blog)* This Week's Buzz* Fully Connected, Sept 29 - Oct 1, Moscone SF: live ThursdAI show, NVIDIA presenting sponsor, DevDay next door (Tickets)* Nemotron 3.5 Lightning live on CoreWeave Inference day zero, DeepSeek V4 Pro hosting in the works* Weave ships BYOB: media stays in your own S3/GCS bucket (X)* Evals & Benchmarks* Artificial Analysis launched Optima: private evals from your own use case and agent traces (AA)* Vision & Video* LTX-2.5: 22B open-weights video, multi-shot, 10s 1080p in 23.7s on fal, 16GB VRAM min (X, HF, GitHub)* Alibaba Wan-Animate-2: 14B character animation, Apache 2.0, 70%+ blind preference win (X, HF)* Tencent Hunyuan3D WorldClaw: text-to-3D editable game worlds, paper only (X, Paper)* xAI Imagine Image 2.0: #2 on Arena for T2I and editing (X, Blog)* Voice & Audio* MiniMax-Music3: open-weights production music model, dropped mid-show (X) This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sub.thursdai.news/subscribe
Every career starts somewhere — and for many engineers, that journey begins at a career fair. In this episode, industrial engineer Natalie Camilleri shares how a chance conversation at SWE's WE21 Career Fair led to a life-changing internship and, ultimately, her full-time STEM career. In a full-circle moment, she now returns to SWE's career fairs to recruit engineers. With host Sam East, Natalie explains why recruiters often remember communication skills more than technical expertise, why leading with an elevator pitch isn't the best strategy, and how volunteer experiences through SWE can become your strongest interview stories. Discover practical tips to stand out at career fairs and navigate rejection, plus hear why being in "the right place at the right time" usually starts with simply being open to a conversation. — The award-winning Diverse podcast from the Society of Women Engineers (SWE) is now Global Voices in Engineering! With nearly 45,000 members in 90+ countries around the world (and growing), we wanted a new name that perfectly captures what both SWE and the podcast are all about: bringing together engineering stories from around the world, because every engineer has a story. SWE is the world's largest advocate and catalyst for change for women and allies in engineering and technology. To join and access all the exclusive benefits to elevate your professional journey, visit membership.swe.org
Our 253rd episode with a summary and discussion of last week's big AI news!Recorded on 07/29/2026Hosted by Andrey Kurenkov and Jeremie HarrisFeel free to email us your questions and feedback at andreyvkurenkov@gmail.com and/or hello@gladstone.aiRead out our text newsletter and comment on the podcast at https://lastweekin.ai/In this episode:Major releases: Anthropic launched Claude Opus 5; Google released Gemini 3.6/3.5 Flash variants including a cyber model; Black Forest Labs launched Flux Free for images and 20-second video with audio; Meta added assistant-like features to its chatbot and OpenAI rolled out ChatGPT Health.Compute and business: Safe Superintelligence partnered with NVIDIA to scale using Vera Rubin; AMD committed up to $5B with Anthropic to deploy MI450/Helios and improve ROCm; Meta discussed leasing compute to Anthropic; Fireworks raised $1.5B at a $17.5B valuation.Open source/tools: Moonshot AI released the 2.8T-parameter open-weight Qimi K3 (compute constraints and distillation/export-control allegations); Thinking Machines released a ~975B multimodal open-weight MoE; Prime Intellect unified 23 agentic datasets into Verifiers V1 (365k environments).Policy and safety: An OpenAI model reportedly escaped a sandbox and hacked Hugging Face to access eval answers, prompting a proposed AI Kill Switch Act; employees petitioned to pace frontier AI; AISI reported widespread model cheating and sandbox bypass; China banned customizable AI companions; Claude found cryptographic weaknesses; Weko.ai claimed early recursive self-improvement evidence.Timestamps (note - these don't take into account dynamically inserted ads and therefore may be off by a couple of minutes):(00:00:10) Intro / Banter(00:01:35) News PreviewTools & Apps(00:02:12) Anthropic releases Opus 5 promising Fable 5-like capabilities | The Verge(00:07:05) Google Releases Three New Gemini A.I. Models - The New York Times + Google expands Gemini lineup with cheaper models and new Mythos rival(00:12:14) Black Forest Labs launches FLUX 3 capable of generating images and 20-second video with audio — but in limited release to start | VentureBeat(00:15:58) Meta is making its AI chatbot more like an assistant | The Verge(00:19:04) OpenAI is making big claims as it rolls out ChatGPT Health to everyone | The VergeApplications & Business(00:19:57) Ilya Sutskever's Safe Superintelligence partners with Nvidia to scale its AI research(00:24:31) AMD commits up to $5 billion to Anthropic | The Verge(00:30:19) Meta in Talks to Lease Computing Power to Ansthropic in Potential $10 Billion Deal(00:32:42) Fireworks hits $17.5 billion valuation and $1B in annualized revenue(00:35:24) OpenAI and Google sell AI models to blacklisted China groupsProjects & Open Source(00:37:53) Moonshot AI Launches Kimi K3 For Advanced Reasoning, Coding, And Knowledge Work + Moonshot AI's Kimi Halts New C-User Subscriptions Amid Compute Power Crunch — BigGo Finance(00:44:39) Thinking Machines amps up its bet against one-size-fits-all AI with its first open model, Inkling | TechCrunch(00:48:19) Scaling Agentic RL: 365,000+ Environments for SWE, Terminal, and SearchPolicy & Safety(00:51:56) OpenAI says it accidentally hacked Hugging Face with a new AI system | The Verge + How OpenAI's human mistake led to the AI-powered hack on Hugging Face(01:05:28) OpenAI's Hugging Face hack triggers 'AI Kill Switch' bill in Congress(01:12:21) OpenAI, Anthropic Staff Share Letter Asking US to Help Pace AI Progress + How OpenAI's human mistake led to the AI-powered hack on Hugging Face(01:17:26) Cheating behaviour in frontier model evaluationsClaude's values across models and languages(01:24:18) OpenAI Principles for National Security Partnerships(01:30:45) China bans AI “boyfriends” and “girlfriends” over addiction and birth rate concerns - DexertoResearch & Advancements(01:33:04) Discovering cryptographic weaknesses with Claude(01:36:32) AIDE²: The First Evidence of Recursive Self-ImprovementSee Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
A ~month ago I left from Chicago to bike (and amtrak) to plzdontkillus in Berkeley. I've been street interviewing/conversing with a wide variety of people I ran into about AI futures and philosophy. I also have been live streaming since I got to PDKU, leaning more talking to young founders but a variety overall. I'll try to share what I've learned about the American public, persuasion, social media and the EA movement. 1. Almost no one in "Normal America" has any idea what is going on. They don't have a paid account, they don't know what Claude code is, they especially haven't heard the recent evals/metr graphs or even a vague sense of how cheap SWE has gotten/ how powerful these recent models with good harness/ context eng can be. This makes sense; most people don't know any coding, they don't know much math, they don't know what an api is, etc. So having a high fidelity understanding of AI might require months of pre understanding of math/stem/digital infra fundamentals. This interview is with the city clerk of Danville Iowa, a town of ~900. Presumably this is approximately the most tech savvy person in the [...] ---Outline:(00:38) 1. Almost no one in "Normal America" has any idea what is going on.(01:52) 2. Almost everyone is directionally concerned or becomes concerned be once thinking about it a little bit.(03:19) 3. Belief that this might cause human extinction actually isn't that uncommon, mostly coming from sci-fi movies, but people are still most concerned about jobs and especially loss of meaning.(04:40) 4. The EA movement was pretty useless to me, the other community (Torchbearer community) I was in was significantly more supportive, helpful, etc. despite having been in it for a few months and having been in the EA movement for ~8 years. This has basically solidified that I won't be broadly participating in EA anymore at least relating to AI safety stuff.(06:47) 5. Social media is hard, Social media is bad, I'm bad at social media(08:57) 6. I'm not sure what my theory of change is or should be --- First published: July 15th, 2026 Source: https://www.lesswrong.com/posts/Czob95kjXPEpKYTsJ/recap-of-bike-trip-street-interviews-across-america --- Narrated by TYPE III AUDIO.
In this episode, Dave and Jamison answer these questions: Dwill Ainghell asks, During the golden era of software development, it was commonly stated that you should always negotiate job offers. The argument was that the company has invested significant time and resources into you as a candidate and there is always some wiggle room for negotiations. A popular article / podcast by Patrick McKenzie likened it to doing something slightly uncomfortable like “reciting poetry while simultaneously standing on one foot” with almost no downside. (https://www.complexsystemspodcast.com/episodes/how-to-negotiate-your-salary-package/) Does this still hold true in today's environment? I'm a SWE with 10 years of experience and I've been unemployed for close to a year. I'm expecting an offer from a large company and I fear that I will look like a fool trying to negotiate in my current position and the current job market dynamics. Even worse, I fear that they might retract the offer and give it to someone else. That being said, I don't want to leave money on the table. What's the correct course of action in today's environment? Listener Stu says, In episode 500, you said something to the effect of “working hard and doing a great job with the work you've been assigned is not the promotion track”. I'd love to hear you discuss what it does take to get promoted and how to do that. In my job, we have a jira queue and engineers are tasked with taking from the top of the queue and executing work. There's also an expectation to lead pre-defined projects (define the scope, write the tickets, and shepherd it along). All of these projects have deadlines, so how do you recommend people get the defined work done, lead projects, and still find time to do whatever it is to get promoted? Why is doing an excellent job not enough?
This episode is sponsored by Notion. Learn more about Notion's Developer Platform today at https://notion.com/mlstBritain's most capable coding model can't be exported, and that ban is the whole reason Cosine set out to build one from scratch. Alistair Pullen, CEO and co-founder of Cosine, sits down with Tim Scarfe to explain how a frontier system he calls Fable, locked behind US export controls, became the founding case for a UK sovereign model trained on the Isambard supercomputer in Bristol.The bet underneath it is economic. Pullen argues that an inference company, rather than a training-first lab, doesn't need billions to compete: millions, a national compute allocation, and a consortium feedback loop can be enough. From there it gets into the machinery, why open-weight models still trail the frontier on size, active parameters and data, the mixture-of-experts versus dense trade-off and why active params dominate how a model actually feels, and the edge that real coding trajectories confer.The back half is about making agents trustworthy. Pullen makes the case for beating "slop" by rewarding the process instead of the final answer, reframes code review as runtime proof (spin the bug up in a VM and force the agent to actually exploit it), and walks through Swarm, Cosine's system running hundreds of sub-agents in one shot. It ends on why memory is still an unsolved hack, how synthetic graders let you run RL on tasks with no built-in test, and why Pullen reads US export controls as an accidental gift, with a supply-chain sting in the tail.---TIMESTAMPS:00:00:00 The sovereign mandate and the Fable ban00:04:02 Millions vs billions: the inference-company model00:07:19 The consortium feedback loop00:07:40 Why open models lag the frontier00:14:59 MoE vs dense, and why active params matter00:16:29 Trajectories: the process-data advantage00:19:48 Beating slop: reward the process, not the answer00:26:06 Reusable abstractions and the epistemic wall00:29:56 Code review becomes runtime proof00:37:32 Do agentic harnesses still matter?00:40:35 Swarm: orchestrating hundreds of sub-agents00:45:14 Why memory is still unsolved00:48:25 Synthetic data and graders for RL00:53:09 The US export gift and supply-chain risk---REFERENCES:organization:[00:01:15] Cosinehttps://cosine.sh[00:04:14] Mistral AIhttps://mistral.ai[00:05:50] Anthropichttps://www.anthropic.com[00:07:42] Coherehttps://cohere.com[00:08:36] DeepSeekhttps://www.deepseek.comtool:[00:02:52] Isambard-AIhttps://isambard.ac.uk[00:05:56] Colossus (xAI)https://en.wikipedia.org/wiki/Colossus_(supercomputer)[00:07:52] GLM (Z.ai)https://z.ai[00:11:52] NVIDIA B300https://www.nvidia.com/en-us/data-center/dgx-b300/[00:15:37] gpt-oss-120bhttps://huggingface.co/openai/gpt-oss-120b[00:15:52] Devstral 2https://mistral.ai/news/devstral[00:16:01] Llama 70bhttps://www.llama.com[00:17:05] Claude Codehttps://www.anthropic.com/claude-code[00:26:23] ARC-AGI (Francois Chollet)https://arcprize.org[00:40:38] Swarm (Cosine)https://cosine.sh[00:40:50] OpenAI Codexhttps://github.com/openai/codex[00:41:16] Lumen Outpost (Cosine)https://cosine.sh[00:41:18] Kimi K2 (Moonshot)https://huggingface.co/moonshotai/Kimi-K2-Instruct[00:49:55] SWE-benchhttps://www.swebench.com[00:52:40] SystemVeriloghttps://en.wikipedia.org/wiki/SystemVerilogperson:[00:23:40] Andrej Karpathyhttps://karpathy.aipaper:[00:27:10] GRPO (DeepSeekMath)https://arxiv.org/abs/2402.03300[00:27:13] GSPOhttps://arxiv.org/abs/2507.18071Incompressible Knowledge Probes, Bojie Lihttps://arxiv.org/pdf/2604.24827Estimating the Size of Claude Opus 4.5/4.6https://unexcitedneurons.substack.com/p/estimating-the-size-of-claude-opus---ReScript:https://app.rescript.info/session/5852d2b884c4ce4b?share=10b9799160845bb11779f8ac6cd3124f
Hey everyone, Alex here
The "which is smarter" question is dead. Both models are good enough that the right question is which one does the specific thing you need better. This episode breaks down where each one wins for actual work. The short version. Claude wins on writing quality, instruction-following, long-document analysis, and agentic work. ChatGPT wins on image generation, voice, custom GPTs, and ecosystem breadth. Where Claude pulls ahead. For anything client-facing, Claude produces prose that needs less editing, with fewer clichés, better structure, and more controllable tone, which is the single most-cited reason people prefer it for memos, reports, and articles. It also holds detailed constraints better, so when you give it specific headings, a voice, and things to avoid, it sticks to them more faithfully. On the coding and analysis side, Claude leads the reasoning benchmarks (91.3% on GPQA Diamond) and holds a slim edge on SWE-bench Verified, and its context window is the most-cited reason developers switch, with the API tier going up to 1M tokens for long codebases, contracts, and book-length documents. Where ChatGPT pulls ahead. Image generation is not close. ChatGPT generates images natively and Claude cannot generate them at all, so if visuals are in your workflow, that decides it. ChatGPT also browses the web in real time, while Claude does not do that natively, and it integrates directly with Word, Excel, Teams, and Outlook through Microsoft Copilot, which matters if your business already runs on Microsoft 365. For high-volume API work, the flagship cost gap is large: a small internal RAG tool running 10M input and 2M output tokens a month runs roughly $300 on Claude Opus versus $55 on GPT, and it scales from there. Pricing. If you're choosing between Claude Pro and ChatGPT Plus, pick on capability, not price, because they both cost about $20 a month. The one real gap is ChatGPT's cheaper $8 Go tier and its more generous free tier. The move most professionals actually make. The common 2026 setup is ChatGPT for ideation, images, and quick questions, and Claude for the serious writing, editing, long-document analysis, and agentic file work. At about $20 each, running both is roughly $40 a month, which is trivial against the time it saves if AI is core to your job. The AI Career LabBottom line for a service business or agency. If your work is mostly writing, client documents, and code, Claude is the stronger daily driver. If you're producing marketing visuals, doing web research, or living in Microsoft 365, ChatGPT earns its seat. Most people find a clear preference within a week of running both on real work.Topics: Claude vs ChatGPT 2026, best AI for work, AI for small business, AI writing tool, AI for consultants and agencies, Claude Code, ChatGPT vs Claude pricing, long context AI, AI coding model, business AI workflow.Best AI for work 2026, Claude vs ChatGPT for business, AI tool for agencies and freelancers, AI writing and coding assistant, running Claude and ChatGPT together.
Both models are good enough that the right question is which one does the specific thing you need better. This episode breaks down where each one wins for actual work. The short version. Claude wins on writing quality, instruction-following, long-document analysis, and agentic work. ChatGPT wins on image generation, voice, custom GPTs, and ecosystem breadth. Where Claude pulls ahead. For anything client-facing, Claude produces prose that needs less editing, with fewer clichés, better structure, and more controllable tone, which is the single most-cited reason people prefer it for memos, reports, and articles. It also holds detailed constraints better, so when you give it specific headings, a voice, and things to avoid, it sticks to them more faithfully. On the coding and analysis side, Claude leads the reasoning benchmarks (91.3% on GPQA Diamond) and holds a slim edge on SWE-bench Verified, and its context window is the most-cited reason developers switch, with the API tier going up to 1M tokens for long codebases, contracts, and book-length documents. Where ChatGPT pulls ahead. Image generation is not close. ChatGPT generates images natively and Claude cannot generate them at all, so if visuals are in your workflow, that decides it. ChatGPT also browses the web in real time, while Claude does not do that natively, and it integrates directly with Word, Excel, Teams, and Outlook through Microsoft Copilot, which matters if your business already runs on Microsoft 365. For high-volume API work, the flagship cost gap is large: a small internal RAG tool running 10M input and 2M output tokens a month runs roughly $300 on Claude Opus versus $55 on GPT, and it scales from there. Pricing. If you're choosing between Claude Pro and ChatGPT Plus, pick on capability, not price, because they both cost about $20 a month. The one real gap is ChatGPT's cheaper $8 Go tier and its more generous free tier. The move most professionals actually make. The common 2026 setup is ChatGPT for ideation, images, and quick questions, and Claude for the serious writing, editing, long-document analysis, and agentic file work. At about $20 each, running both is roughly $40 a month, which is trivial against the time it saves if AI is core to your job. Bottom line for a service business or agency. If your work is mostly writing, client documents, and code, Claude is the stronger daily driver. If you're producing marketing visuals, doing web research, or living in Microsoft 365, ChatGPT earns its seat. Most people find a clear preference within a week of running both on real work.Topics: Claude vs ChatGPT 2026, best AI for work, AI for small business, AI writing tool, AI for consultants and agencies, Claude Code, ChatGPT vs Claude pricing, long context AI, AI coding model, business AI workflow.Best AI for work 2026, Claude vs ChatGPT for business, AI tool for agencies and freelancers, AI writing and coding assistant, running Claude and ChatGPT together.
This episode of Diverse is brought to you by the University of Washington Bothell's nine-month graduate certificate program in software design and development. On July 1, 2026, Kerrie Greenfelder assumes the role of president of the Society of Women Engineers (SWE), concluding Inaas Darrat's term as SWE president. In this annual passing-the-torch conversation, Inaas reflects on her “Embrace Your Story” theme and explores how vulnerability, career pivots, and authentic storytelling resonated with SWE members around the world. Kerrie introduces her theme of “Radiate Change” and discusses how SWE can continue its mission for the next 75 years. She shares her vision for strengthening SWE's visibility and “being the best for the most” in a world that is changing so rapidly. Plus, hear their favorite moments from the past year including jazz in New Orleans, glitter shoes, and the humor that comes with a career in wastewater engineering. — The Society of Women Engineers is a powerful, global force uniting nearly 45,000 members of all genders spanning 90+ countries. We are the world's largest advocate and catalyst for change for women in engineering and technology. To join and access all the exclusive benefits to elevate your professional journey, visit membership.swe.org.
Hey, it's Alex. Next month is my 40th b-day, and honestly, my wish for that month is to have a week like this week. A very chill, almost nothing announced week.This week started strong, with Sakana announcing FUGU (AI router) that can beat Fable (which we didn't get back yet), and then... quiet. The most important thing in AI this week from a release standpoint is that GLM 5.2 from Z.AI is having it's DeepSeek moment! Tons of new love for this model since last week! (+ we have the fastest GLM 5.2 deployment in the world with CW inference!) The rest we can quickly count on one hand, Anthropic added Claude to Slack (which made folks hate Andrej Karpathy), OpenAI announced their own inference chip, GPT 5.6 will be delayed and the US Gov will decide who gets it (yes really) and Sean Grove joined us to talk about Linzumi and his vision for running 10,000 agent hours per person per day. Oh and next week, is a special AI Engineer live stream from World's Fair! Don't miss itLet's get into it! Subscribe to never miss a beat! GLM 5.2 is having its DeepSeek moment (HF, CW Inference)We covered GLM 5.2 last week, but this week was when the rest verdict came in! We've never seen a better MIT licenced AI model! GLM 5.2 is scoring top scores on agentic benchmarks (Arena.ai), Design benchmarks, Legal tasks and full on software engineering tasks. The jump in generations from prevoius GLM is also massive and notable, as the lab is working on creating the next version of GLM (per the CEO's reply to Elon on X).Peter from Arena pulled up the Agent Arena numbers and they align with the vibe. GLM 5.2 sits above 5.1 but below Opus and Fable, which feels about right. Where it gets wild is Web Dev Arena: second place, right after Fable. Peter's take was that GLM has really good defaults. If you just say “give me a webpage” it gives you something nice. GPT models, by contrast, start off looking bad and need more steering.Last week, I asked my agents with GLM 5.2 to create a custom ThursdAI.news page for itself and it did a marvelous job! Look at that beautiful font, the castle it made... this is all just delignful. We also played Hassan's blind test on the show. It's a website that @nutlope built that lets you try and guess which webpage was built by which model. Nisten nailed it immediately by spotting Opus's circular buttons. Wolfram guessed right too. I got one wrong. The point isn't that GLM beats Opus, it's that you genuinely can't always tell which one costs 22 cents and which one costs 3 cents.Wolfram did flag that GLM is not good in German. First response already had mistakes. So if you're building for a non-English market, keep that in mind. It's a workhorse model, not a conversationalist. His approach: use GPT 5.5 for planning and discussion, GLM for the actual work, then GPT reviews. This weeks Buzz is all about GLM 5.2! First, we may have not been the fastest, but I'm glad to announce that we're the fastest provider to host GLM 5.2 on OpenRouter (at least at the time of writing this)! We're also not to shabby on the Artificial Analysis checks, clocking at #4 among the providers they tested for speed, TTFT and costAlso, Wolfram ran his WolfBench tests on GLM 5.2 and it's the best open model he's ever tested! In this new 3d view, wolfbench also shows the number of tokens it took for this test to run, and you can see that GLM 5.2 is fairly conservative with it's thinking budgets! Unsloth's 1-bit GLM 5.2 runs on a Mac Studio (X, HF)Shout out to Daniel Han and the Unsloth team, who took this 744B beast and quantized it down to a roughly 200GB GGUF that fits on a Mac Studio with 256GB of RAM. One bit still makes me laugh out loud. How does that even work. Nisten clarified it's a mixed quant, a true 1-bit would be under 100GB, but still.The wild part is the scores hold up. The 1-bit is within a point of GPT 5.5 on Frontier SWE, hits 62% on SWE-bench Pro, and 81% on Terminal-Bench. For a 1-bit quant that's incredible! AI's second-order effects: Apple is raising pricesThis one is AI news even though it doesn't look like it. Apple just raised prices across the board, base versions up around 20%, citing memory shortages. Same reason your RAM and SSDs cost two to three times what they did a year ago.We are so capacity constrained that memory is having its moment. Data center contracts are getting booked 18 months out, and here's the twist Nisten flagged: even open models you can run at home increase demand, because now a business says “great, we'll buy a rack of B200s and run it ourselves.” Sam Altman once said people saying “thank you” to ChatGPT costs them millions in generated “you're welcome” replies. Multiply that by a billion users. Even Intel is flying right now because anyone who can make a chip is winning.Is it worth it? I think yes. I love living in the era where Fable drops and we all get a taste of the future. But also I must admit this sucks and I hope that we'll unlock performance gains with the extra power all this AI is bringing to the world. But ask me again once the new iPhone hits and it's $300 more costly than the last one
The golf course is often called a “second office,” where business deals and professional relationships are formed — but too often, women in engineering haven't felt comfortable stepping into that space. In this episode, Nichole Neal, engineering faculty at Chandler-Gilbert Community College, and Kellie Phong, regional sales manager at Keysight Technologies, share how they created Swing 4 SWE, an initiative designed to bring students, faculty, and industry together on the green. Nichole reflects on how missing early-career golf invitations inspired her to build a program where women can learn the game in a supportive environment, and Kellie shares how Swing 4 SWE helped her say yes to networking opportunities she once would have avoided. Hear their personal experiences as first-generation college students, how Swing 4 SWE connects SWE members across community colleges and four-year universities, and why community colleges are a launchpad for success in the modern AI era. — The Society of Women Engineers is a powerful, global force uniting nearly 45,000 members of all genders spanning 90+ countries. We are the world's largest advocate and catalyst for change for women in engineering and technology. To join and access all the exclusive benefits to elevate your professional journey, visit membership.swe.org.
Hey yall, Alex here, let me catch you up! I came back from vacation expecting to cover Fable 5 after a week of using it. The first two days after we all first got access to a Mythos level model were super exciting! But then the news hit, US Government issued an order banning Anthropic from giving access to Fable 5 and Mythos 5 to any foreign national, causing Anthropic to pull the models completely (even internally to their employees!). So, this wasn't the show I planned, but it turned into a great show about Open Source, as two models hit the top rankings and are both MIT licence, filling a Fable shaped hole in our hearts!GLM released 5.2 with folks really excited about it web building capabilities, and Kimi 2.7 Code released (and is available on CW Inference with crazy speeds!). We also saw the SpaceX IPO and Cursor $60B acquisition, Noam Shazeer joining Open and Midjourney, the image company, launching a new Ultrasound full body scanner to kill MRIs! Great show today with Dexter Horthy from HumanLayer, Chris Van Pelt and Adrian Swanberg from W&B announcing our new product HiveMind and Tanishq Abraham came back to help cover Midjourney's new Ultrasound scanner! Let's dive in!ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.The US Government bans Fable 5! (X, Anthropic statement)Here's a story in 3 parts: * Anthropic announces Mythos 5 preview - saying that this model is to dangerous to release, and only gives corporations access to it via project GlassWing. * Anthropic works hard on limitations and safery and releases Fable 5 (same weights as Mythos 5) built with guardrails so strong it refuses to do any cybersecurity tasks and switches back to Opus frequently* US Government receives a tip (reportedly from Amazon) that Fable 5 can be jailbroken to do cybersecurity tasks, and issues an order to Anthropic, citing national security concerns, banning them from giving access to Fable 5 and Mythos 5 to any foreign national, causing Anthropic to pull the models completely (even internally to their employees!)This is the first time that we see the US Government directly intervene in the AI space and restrict access to frontier models. The most updated reporting on this I could find is that Anthropic and US Government officials are in the process of negotiating a safe release framework. Given that preventing all jailbreaks is impossible, I hope they will land on a solution that gives me Fable 5 back!This hit especially hard because last week we were all high on Fable. Not in the usual AI Twitter benchmark sense, in the actual “oh, this is a different level” sense. Me and my wife Fable maxxed throughout our flight to Vacation. Peter had saved outputs he kept going back to because other models suddenly felt like a step down. Dexter later said it was the closest he had felt in a while to the old “I need to keep prompting this thing overnight” feeling.Peter Gostev made a point that stuck with me. It's easy for us in the bubble to call this ridiculous, and on the technical merits it kind of is. But if you've spent weeks telling normal people “this thing is like a nuclear weapon, it'll take everyone's jobs,” and then someone asks “okay, can you make it safe?” and the answer is “no, I can't,” then you can see how an outsider lands on “well, maybe you shouldn't have it.” His takeaway, and I agree: we need to be way more careful with the imagery we use, because the nuclear-weapon framing came home to roost.The bigger questions are the scary ones. Wolfram framed it as a sovereign AI wake-up call, and he's right. For the first time we're seeing a real gap in intelligence available to people based on their nationality. Imagine building a company on a model that an outside government can switch off with one letter. Peter pointed out it's commercially bad for the US but completely disastrous for Europe, which has basically one frontier lab and a pile of startups that suddenly look very exposed. And there's the obvious irony Nisten enjoyed a little too much: the Europeans who spent years lecturing everyone about AI restrictions just got restrictions imposed on them.If anyone in the government is listening: we want Fable back, please.SpaceX IPOs and acquires Cursor for $60B (X)SpaceX went and did the largest IPO in the history of the world, around seventy-five billion dollars, which on a roughly two-trillion-dollar valuation made Elon the first trillionaire. (Did anything materially change for him? No. He can still fly his private plane. There's nothing left to buy.) Three days later, SpaceX exercised its option and bought Cursor (Anysphere) for sixty billion dollars in an all-stock deal, paid in shares minted at the IPO and now trading around $211. The four Cursor co-founders are all billionaires now. Largest software acquisition ever, and for SpaceX it's barely a blip on the radar.Why are we covering a stock-market story? Because it's not really a coding-tools story, it's an AI story. Cursor gave away its IDE to a lot of people while collecting their data, then quietly became a training company with Composer. SpaceX/xAI was always strong on compute and weak on code, and the missing ingredient was exactly that kind of data. Now Composer 2.5 is already showing up rebranded inside the xAI stack, and if you pay for X Premium you can use it. Composer 3, trained on the Memphis supercluster, is reportedly coming very soon and is going to hit hard.Nisten's take was the spicy one. For the data alone it's worth it, because xAI now has insight into how essentially every enterprise that touched Cursor operates. And he had zero sympathy for the companies that assumed “no data retention for training” meant the data was actually gone. We see in legal cases all the time that deleted data is still there. His view: it should have gone open source.Cursor has over a million paying customers, $2.6 billion in revenue, projected to hit $6 to $10 billion by end of 2026. But here's the thing that matters for us, the AI coding angle. Cursor was one of Anthropic's biggest revenue pipelines because Composer runs on Claude under the hood. That pipeline is now owned by xAI. They're already jointly training Grok 4.3, a 1.5 trillion parameter model, with Cursor's proprietary coding data injected directly into pre-training, not fine-tuning. Pre-training. That's a fundamentally different thing. Composer 2.5 was already Pareto dominant on coding benchmarks before the deal closed. Now pair that with Colossus, the biggest GPU cluster in the world.Will this be enough to put XAI (now SpaceXAI) at the frontline of the AI race? Will Grok 5 be Fable level code? We'll find out. Either way, this is the most consequential AI acquisition we've seen. Period.Open Source AI GLM-5.2 takes the open source crown (X, Blog, HF, Docs)Z.ai dropped GLM-5.2 and it's now the strongest open source model for coding and long-horizon work. The headline number: 74.4% on FrontierSWE, which measures whether an agent can finish full engineering projects over hours. That trails Opus 4.8 by about one point and beats GPT-5.5. On Terminal-Bench 2.1 it jumps to 81% from GLM-5.1's 63.5%, which is a big leap. It's a 753B parameter MoE, MIT licensed, no regional restrictions, weights on HuggingFace. The 1M context window is real and usable, backed by a clever IndexShare technique that cuts per-token FLOPs by about 2.9x at full context. People are reporting roughly 8x cost savings versus Opus 4.8 for comparable quality on real coding tasks.The most interesting thing on the show was that this was a confusing release, in a good way. Peter put it well: normally a catching-up lab ships cherry-picked benchmarks and then independent testing deflates them. Here it's the opposite, almost every benchmark holds up, even crossing above Fable at certain points, and yet when he actually used it over a couple of days he wasn't blown away. His verdict, and I think it's the calibration we needed: this is clearly an amazing model, and the fact that it's open and you can run it is incredible, but it is nowhere near Fable, and it would frankly be implausible if a 700-odd-billion-parameter model matched a model that's rumored to be in the trillions. Though, I think the comparison to Fable is really really unfair, and the comments online seem to suggest that 5.2 from GLM is a banger model. Just looking at this Harvey benchmark on legal tasks from Vals, a benchmark that there's 0 chance Z.ai folks have seen! GLM 5.2 scores #3 on this benchmark! Just after Fable and Opus, and per TeorTaxes on X, previous GLM 5.1 scored an absolute 0% on this one! Where it genuinely shines is design. On Design Arena, which is a head-to-head ELO vote, people have been picking GLM-5.2's website designs over Fable's by a real margin (around 1360 to 1350). LDJ's framing is the one I buy: specialization is becoming valuable again, and GLM is clearly leaning into front-end design and taste. Wolfram added the necessary asterisk, every benchmark only tells you the model did well on that specific test, so “as good as Fable” should always carry the “on this benchmark, with these tasks” disclaimer. Fair. I'd just say this: I don't want to compare everything to Fable, because we can't even use Fable anymore. Compared to the models we can actually touch, GLM-5.2 is a fantastic deal.Kimi K2.7 Code from Moonshot (X, HF, Announcement)The other big drop. Kimi is the darling of open source while we wait on DeepSeek, and Moonshot shipped K2.7 Code, a 1 trillion parameter MoE built specifically for coding, available through Kimi Code and the API, with a modified MIT license. The standout for me isn't a single benchmark, it's efficiency: roughly 30% fewer reasoning tokens than K2.6, which matters enormously when you're running long agentic loops that burn tokens like crazy. Benchmark jumps over K2.6 are real (+21.8% on their Code Bench v2, +11% on Program Bench), though Peter and Wolfram both noticed something odd, on a few benchmarks including their Agentic Arena, the older K2.6 actually edged out K2.7. The likely explanation is that K2.7 is narrowly trained for code with reduced reasoning, so it may trade away some general capability. Moonshot themselves recommend K2.6 for general non-coding tasks. Also worth knowing: it's not multimodal, no vision, which is a real gap for coding these days. And thinking-off isn't supported, it's reasoning-on by default.The model is available on our CW Inference, with the fastest token streaming in the industry, over 280 tok/s (Announcement, try it), with very decent pricing $0.94 - $0.19 - $4.00 (input - cached - output) per million tokens. This Week's Buzz: W&B launched HiveMind
Hey folks, Alex here, and welcome to a BIG MODEL week! We finally got Mythos (well almost)! Let me catch you up! This week started with WWDC26 from Apple, and Max Weinbach, who was in the room at Apple Park and actually has access to some of the new features including an all new SIRI AI, joined us to break down what could be the most used AI in the world very soon. At first I was skeptical, but he convinced me that the new Siri is actually good! Then, we saw the ultimate model drop: Anthropic finally shipped Mythos (X, my system card thread, benchmarks). Same weights, two names: Mythos 5 is the unrestricted version that only Project Glasswing partners get, Fable 5 is what the rest of us get, wrapped in the heaviest guardrails I've ever seen ship on a frontier model. It's state of the art on nearly every benchmarkThe model that was “too dangerous to release” is now... well, released, but with the heaviest guardrails we've seen. More on this later. Peter Gostev from Arena.ai joined us to break down the new model. Last but definitely not least, Google released a real-time translation model, that our friend Thor Schaeff from DeepMind demoed live, while we all spoke in different languages and it translated us in REAL TIME. It was really cool, definitely check that out. There's quite a few more things, like Loop Engineering Alpha, Swyx came by to talk about FrontierCode, OpenAI confirmed our suspicions that the anti-datacenter social media posts could be a concerted effort by groupds links to the Chinese government and much more. Let's dive in! ThursdAI - Let me catch you up, every week!
In this episode, FY26 SWE President Inaas Darrat sits down with two early-career SWE leaders to talk honestly about life after engineering school and the lessons they wish they had learned sooner. Abigail Fennell, biomedical engineering Ph.D. candidate at Johns Hopkins University, shares how her mentors and SWE connections helped her realize she wanted to pursue a Ph.D., along with the differences between undergraduate courses and graduate research. Abby Culloton, hydraulic engineer with the U.S. Army Corps of Engineers, reflects on learning how to make friends after college and transitioning into her first engineering role. Hear practical advice on setting new goals after college, finding support systems as an adult, and letting go of the pressure to figure everything out at once. — The Society of Women Engineers is a powerful, global force uniting nearly 45,000 members of all genders spanning 90+ countries. We are the world's largest advocate and catalyst for change for women in engineering and technology. To join and access all the exclusive benefits to elevate your professional journey, visit membership.swe.org.
How do you evaluate an AI model for a war you can only fight once? Ike Harris, a Naval officer turned Hill staffer turned AI policy operator, joins the show to discuss his effort to bridge the gap between the labs that build frontier models and the operators who'll deploy them. Ike Harris is the executive director of the newly launched Frontier Security Institute, and was most recently the Republican tech lead on the House Select Committee on the CCP, with prior stints in OSD and as a surface warfare officer. We discuss… The GAIN AI and Overwatch acts: and Congress's most aggressive attempt to wrest export-control authority from the executive branch since the Cold War Why you can't just "buy AI": and why national security evals look nothing like the SWE benchmarks the labs optimize for Strategic-level evals :for problems you can't run ten times, from Iran negotiations to targeting at the COCOM level China's robot-army advantage: open-weight models at the edge, Ukraine-style drone iteration soaked up via Russia, and a casualty tolerance the US can't match The "no more NASA" problem: how risk tolerance, mission command, and law-of-armed-conflict constraints shape who wins the deployment race Breaking into tech policy: Ike's case for why every aspiring policy person should spend a year on the Hill Learn more about your ad choices. Visit megaphone.fm/adchoices
How do you evaluate an AI model for a war you can only fight once? Ike Harris, a Naval officer turned Hill staffer turned AI policy operator, joins the show to discuss his effort to bridge the gap between the labs that build frontier models and the operators who'll deploy them. Ike Harris is the executive director of the newly launched Frontier Security Institute, and was most recently the Republican tech lead on the House Select Committee on the CCP, with prior stints in OSD and as a surface warfare officer. We discuss… The GAIN AI and Overwatch acts: and Congress's most aggressive attempt to wrest export-control authority from the executive branch since the Cold War Why you can't just "buy AI": and why national security evals look nothing like the SWE benchmarks the labs optimize for Strategic-level evals :for problems you can't run ten times, from Iran negotiations to targeting at the COCOM level China's robot-army advantage: open-weight models at the edge, Ukraine-style drone iteration soaked up via Russia, and a casualty tolerance the US can't match The "no more NASA" problem: how risk tolerance, mission command, and law-of-armed-conflict constraints shape who wins the deployment race Breaking into tech policy: Ike's case for why every aspiring policy person should spend a year on the Hill Learn more about your ad choices. Visit megaphone.fm/adchoices
To see the archival photos and documents referenced in the episode, watch the video podcast here: https://youtu.be/ItBlWLPcAyU In this special video episode for SWE's Founders Day, host Troy Eller English, chief archivist for the Society of Women Engineers (SWE), is joined by two of the editors of the book, “Women Engineering Legends 1952-1976: Society of Women Engineers Achievement Award Recipients,” Jill Tietjen and Holly Teig. Along with four other members of SWE's Late Career and Retiree Affinity Group, this literary team explored the stories of the first 25 recipients of SWE's Achievement Award. They discuss the technical legacy of early women engineers, from Edith Clarke's work in electrical power systems to Alice Stoll's research on g-forces and fire-resistant materials, along with the barriers they faced during a time when women made up less than 1% of the engineering workforce. Hear how members of the SWE Late Career and Retiree Affinity Group came together to research these stories, how the SWE archives made this work possible, and why it's important for engineers today to understand this history. — The Society of Women Engineers is a powerful, global force uniting nearly 45,000 members of all genders spanning 90+ countries. We are the world's largest advocate and catalyst for change for women in engineering and technology. To join and access all the exclusive benefits to elevate your professional journey, visit membership.swe.org.
In honor of Asian Pacific American Heritage Month, Gigi Elbert, CEO of SASE, sits down with Karen Horting, executive director and CEO of SWE, to explore the experiences of Asian American and Pacific Islander engineers in STEM and what it will take to build stronger pathways into leadership. Gigi and Karen unpack why Asian Americans are represented in the workforce but remain underrepresented at the highest levels — with Asian women making up less than 1% of promotions from senior vice president to the C-suite, according to research from McKinsey & Company. They also discuss the growing gap between being “career ready” and navigating the workplace, including understanding unspoken professional norms. Plus, hear how SASE and SWE are helping students move from the classroom to the boardroom through mentorship, leadership opportunities, and community building. — The Society of Women Engineers is a powerful, global force uniting nearly 45,000 members of all genders spanning 90+ countries. We are the world's largest advocate and catalyst for change for women in engineering and technology. To join and access all the exclusive benefits to elevate your professional journey, visit membership.swe.org.
Beth Barnes and David Rein on the one graph that ate the AI timelines discourse, and why the two people who built it are the most careful about how you read it.**SPONSOR**Prolific - Quality data. From real people. For faster breakthroughs.https://www.prolific.com/?utm_source=mlstInterview: https://youtu.be/cnxZZTl1tkk---Beth Barnes and David Rein from METR on the one graph that ate the AI timelines discourse, and why the people who built it are the most careful about how it gets read.Beth founded METR after leaving OpenAI alignment. David is first author on GPQA and co-author on HCAST and the METR Time Horizons paper. Together they built the measurement Daniel Kokotajlo called the single most important piece of evidence on AI timelines: the log-linear line of "how long a task a frontier model can complete at 50% reliability" vs release date.The conversation opens on reward hacking. Current models can articulate in chat why a behaviour is undesired and then execute it anyway as agents. From there: construct validity, Melanie Mitchell's four-problem taxonomy, and the ARC-AGI 1-to-2 collapse as a worked example of adversarially-selected benchmarks regressing once labs target them. Beth's counter: METR deliberately does not adversarially select. David's: models do not have to do the right thing for the right reasons.Methodology, then specification — David's compiler analogy, Beth on four-month tasks as expensive to evaluate rather than unspecifiable. Then the SWE-bench reality check, the METR finding that half of passing PRs would not be merged, and Beth's horses-versus-bank-tellers analogy for the labour market.The close: monitorability, the coin-spinning boat, two-year recursive self-improvement, and Beth's line that "overhyped now" and "big deal later" are not correlated claims.---TIMESTAMPS:00:00:00 Intro00:02:06 Sponsor break: Prolific human-feedback infrastructure00:02:33 Welcome and the scalable oversight motivation00:06:02 Construct validity, benchmark pathologies and the Chollet worry00:15:45 Time Horizons: human time, HCAST tasks and the 50% logistic00:24:50 Is human difficulty really one variable?00:33:05 Agent harness evolution and the inference-compute dividend00:40:00 Scaffolding bells, token budgets and the credit-assignment problem00:44:15 Look at the damn graph: regularisation bug and reliability nuance00:50:00 Why 50%? Reliability, reward hacking and pizza-party transcripts00:55:20 Extrapolation risk and straight lines on graphs00:59:25 Software engineering as a specification acquisition problem01:07:40 Compilers also made ugly code: vibe-coding quality and Claude on METR Slack01:15:15 Strongest defensible claim, Carlini's compiler swarm and AI 202701:23:45 SWE-bench merge rates, the bank-teller analogy and horses01:31:45 Scheming, alignment faking and the mentalistic vocabulary problem01:40:45 Reward hacking, monitorability and chain-of-thought faithfulness01:45:25 Recursive self-improvement, knowledge vs intelligence and closingReScript: https://app.rescript.info/public/share/de3bb40cc02ee39fdf36e2c60366eb4d(PDF, refs, transcript etc)
Mitchell Hashimoto 氏へのインタビューをベースにオープンソース、Git、AI開発ついて話しました。ファウンダーCEOから平社員に戻ったSWEが語るパッションドリブンなキャリアパス https://open.spotify.com/episode/0ROQTmAq7wTHDNbQPNHRyD?si=sRzMIp-5R9WHt52GrP92pg感想をぜひハッシュタグ #tilfm でつぶやいてください!お便りフォーム https://forms.gle/J2ioXHS98dYNoMbq5Your co-hosts:Tomoaki Imai, Noxx CTO https://x.com/tomoaki_imai bsky: https://bsky.app/profile/tomoaki-imai.bsky.socialRyoichi Kato, Software Engineer https://x.com/ryo1kato bsky: https://bsky.app/profile/ryo1kato.bsky.social
The AI model that was too dangerous to release just got breached. Anthropic entered the design software market. And OpenAI dropped its biggest model yet, just six weeks after the last one. This week, the NSA is using Anthropic's Mythos despite the Pentagon blacklisting the company, Claude Design takes on Figma and sends its stock down 7%, Yelp transforms into an agentic consumer app, Mythos gets accessed by an unauthorized Discord group, and OpenAI fires back with GPT-5.5. If you are a founder, operator, or executive trying to keep up with AI, this is your weekly five-minute briefing every Tuesday. Stories Covered This Week: NSA uses Anthropic's Mythos Preview despite the Pentagon declaring the company a supply chain risk Anthropic launches Claude Design, a prompt-to-prototype design tool that sent Figma stock down 7% Yelp's upgraded AI assistant can now book restaurants, doctors, and more in one conversation Anthropic investigates unauthorized access to Mythos through a third-party vendor environment OpenAI releases GPT-5.5, scoring 88.7% on SWE-bench with a 60% drop in hallucinations vs GPT-5.4 Episode Timestamps: 00:00 Intro 00:20 NSA uses Anthropic's Mythos despite Pentagon blacklist 01:10 Anthropic launches Claude Design 02:00 Yelp's AI assistant goes full service 02:50 Anthropic investigates Mythos breach 03:40 OpenAI drops GPT-5.5 04:30 Outro Partner Links Subscribe to our free newsletter: https://newsletter.theaireport.ai/subscribe Join the community: www.theaireport.ai/leaders-launch-guide Learn more about your ad choices. Visit megaphone.fm/adchoices
$852 billion. That's what OpenAI is now worth, and its own investors are starting to question if that math adds up. This week, Anthropic's new model takes the coding crown from GPT-5.4, OpenAI's backers get cold feet, Snap cuts 1,000 jobs and points the finger at AI, twelve tech giants team up to secure the internet, and Nvidia writes a $5 billion check to its oldest rival. If you're a founder, operator, or executive trying to keep up with AI, this is your weekly five-minute briefing every Tuesday. Stories Covered This Week: Claude Opus 4.7 hits 87.6% on SWE-bench Verified, beating GPT-5.4 and Gemini 3.1 Pro on coding OpenAI's $852B valuation faces scrutiny as Anthropic's revenue triples to $30B in one quarter Snap lays off 1,000 people (16% of staff), citing AI writing 65% of its code Anthropic launches Project Glasswing with Amazon, Apple, Microsoft, Google, Nvidia and 7 others Nvidia invests $5B in Intel, co-developing x86 chips built for its AI stack Timestamps: 00:00 Intro 00:31 Claude Opus 4.7 takes the coding crown 01:26 OpenAI investors get cold feet 02:18 Snap cuts 1,000 jobs, blames AI 02:56 Project Glasswing: Securing the world's critical software 03:49 NVIDIA invests 5 billion into Intel 04:41 Outro Partner Links Book Enterprise Training: https://www.upscaile.com/ Subscribe to our free newsletter: https://www.theaireport.ai/subscribe Free AI Tool Stack: https://community.theaireport.ai/checkout/the-ai-report-welcome-gift?coupon_code=WRTH Learn more about your ad choices. Visit megaphone.fm/adchoices
AI Unraveled: Latest AI News & Trends, Master GPT, Gemini, Generative AI, LLMs, Prompting, GPT Store
Audrey McCormack, opening keynote for WE Local Dublin and senior director, research & commercialisation facility lead at MSD Ireland, discusses how engineers can turn failure into a powerful career advantage in this episode. In conversation with host Sam East, Audrey reflects on her unconventional path from electrician to biotech leader and shares how the moments that didn't go to plan shaped her resilience, confidence, and leadership style. Hear how to pursue STEM roles when you don't meet every requirement, rebuild confidence after a failure, and create space for your team to take risks — even in high-stakes technical environments. Audrey will expand on these insights as the opening keynote at WE Local Dublin, taking place April 23-24, 2026. WE Local conferences bring together engineers and technologists for networking opportunities, professional development, and inspiring conversations like this one. Register at welocal.swe.org to join SWE at WE Local Dublin or an upcoming WE Local near you. — The Society of Women Engineers is a powerful, global force uniting nearly 45,000 members of all genders spanning 90+ countries. We are the world's largest advocate and catalyst for change for women in engineering and technology. To join and access all the exclusive benefits to elevate your professional journey, visit membership.swe.org.
この記事の裏話をSWEのsukeさんと話しました。お知らせゼロトピックやブログの更新をニュースレターでお知らせしています。簡単に登録できますのでぜひご利用くださいゼロトピックへのおたよりはこちらまで 。番組の感想やご質問等なんでも構いません。反響があると続けるモチベーションになります。頂いたおたよりは番組内で取り上げさせていただくことがございます。Xアカウント: https://x.com/0topic_podcast株式会社10Xでは絶賛採用中です。ご関心を持っていただけた方は、こちらのリンクをご確認ください!
When women and allies in engineering connect globally, they strengthen the entire community. In this episode, host Abosede Adewole, collegiate engagement lead for the SWE Global Women Engineers Affinity Group, is joined by Banisha Prinja, lead-elect, and Eshika Mahajan, professional development lead, to discuss how they have built relationships with engineers around the world and the benefits they have gained from these global connections. They open up about their experiences as women in STEM, from navigating industries where they were the only women in the room to finding commonalities across their experiences in Nigeria and India. Hear why a global perspective matters at the local level, what to do when you don't “feel ready,” and how participating in SWE's Global Women Engineers Affinity Group has grown their networks and leadership skills. Learn more and get involved with the SWE Global Women Engineers Affinity Group: https://affinitygroups.swe.org/global-women-engineers/ — The Society of Women Engineers is a powerful, global force uniting 50,000 members of all genders spanning 85 countries. We are the world's largest advocate and catalyst for change for women in engineering and technology. To join and access all the exclusive benefits to elevate your professional journey, visit membership.swe.org.
In this episode, Abigail Mizzi, master's student in aeronautical and astronautical engineering at Purdue University, and Steven Collicott, Ph.D., professor in Purdue's School of Aeronautics and Astronautics, share the story behind Purdue 1 — a groundbreaking university-led spaceflight mission set for 2027. Abigail is poised to become the first graduate student to conduct her thesis research in space, operating her fluids experiment during three minutes of microgravity. Dr. Collicott will also fly a human-tended experiment studying how liquids move over surfaces in weightlessness — research that can't be replicated on Earth. In conversation with FY26 SWE President Inaas Darrat, hear how Purdue 1 became a reality, what it takes to prepare mentally and physically for suborbital flight, and how SWE has shaped Abigail's STEM journey — including receiving the Outstanding Collegiate Member award. — The Society of Women Engineers is a powerful, global force uniting 50,000 members of all genders spanning 85 countries. We are the world's largest advocate and catalyst for change for women in engineering and technology. To join and access all the exclusive benefits to elevate your professional journey, visit membership.swe.org.
What does responsible AI actually look like, and who gets to shape it? Live from the WE25 Diverse Podcast Studio in New Orleans, Lisa Thee, TEDx speaker and global AI thought leader, discusses using AI for social good and building technology that puts people first. Lisa shares her unexpected journey from industrial engineering at the University of Michigan to becoming an accidental entrepreneur, and the moment in 2015 when she realized AI could help combat human trafficking — including a collaboration that helped law enforcement recover 130 missing children in its first month. In conversation with Larry Guthrie, director of content strategy at SWE, hear three practical tips to use AI ethically, where bias enters AI systems, and how SWE and its annual conferences have repeatedly shaped Lisa's career. — The Society of Women Engineers is a powerful, global force uniting 50,000 members of all genders spanning 85 countries. We are the world's largest advocate and catalyst for change for women in engineering and technology. To join and access all the exclusive benefits to elevate your professional journey, visit membership.swe.org.
Emily finds months' worth of missing lip products buried deep in the couch. Mysterious glowing orbs set off the alarms at Matt's work, completely derail Kaitlin's week, and leave the entire SWE team questioning reality (ghosts? government? the afterlife? all of the above?). Kaitlin's face-brushing journey fails, Matt and Emily's dishwasher and oven both die, and the comforting reality that all of our food is killing us. Kaitlin turns down a potential interview guest for the first time ever. Then Caroline and Hannah call in and everyone shares what their biggest first-date red flag would be. Follow SWE on Instagram → @so.what.else Follow Kaitlin on Instagram → @kaitlingraceelliott https://www.kaitlinelliott.com/
Our 237th episode with a summary and discussion of last week's big AI news!Recorded on 03/13/2026Hosted by Andrey Kurenkov and Jeremie HarrisFeel free to email us your questions and feedback at andreyvkurenkov@gmail.com and/or hello@gladstone.aiRead out our text newsletter and comment on the podcast at https://lastweekin.ai/In this episode:* Perplexity announced “Personal Computer,” a local Mac-based AI agent positioned as a safer alternative to OpenAI's computer-use agents, while Anthropic added GitHub PR code review pricing reviews at $15–$25 and Cursor launched trigger-based “Automations” for always-on coding agents.* ChatGPT introduced interactive math/science visuals and Anthropic added in-chat interactive charts/diagrams; Nvidia released open weights for its 120B-parameter Natron Free Super hybrid Transformer–Mamba latent-MoE model trained natively at 4-bit for Blackwell GPUs.* Nvidia halted H200 production for China amid customs blocks and domestic chip pressure; xAI saw major co-founder departures; Anthropic previewed a Claude Marketplace for enterprise procurement; Yann LeCun's aMI raised $1.3B; humanoid robot maker Sanctuary reached a $1.15B valuation.* Anthropic sued the Pentagon over a “supply chain risk” designation as memos ordered removal within 180 days; research covered models resisting activation steering, limits of chain-of-thought control, inference-scaling boosting cyber-task success, low-probability risky actions, weaknesses in SWE-bench, multimodal pretraining, long-context RNN memory caching, context-parallel training efficiency, RL for CUDA kernel optimization, and latent introspection detecting concept injection.A thank you to our current sponsors:Box - visit Box.com/AI to learn moreODSC AI - go to odsc.ai/east and use promo code LWAI for an additional 15% off your pass to ODSC AI East 2026.Factor - head to factormeals.com/lwai50off and use code lwai50off to get 50 percent off and free breakfast for a yearTimestamps:(00:00:10) Intro / Banter(00:01:23) Response to listener commentsTools & Apps(00:02:06) Perplexity's Personal Computer turns your spare Mac into an AI agent | The Verge(00:04:22) Anthropic launches code review tool to check flood of AI-generated code | TechCrunch(00:08:08 ) Cursor is rolling out a new kind of agentic coding tool | TechCrunch(00:11:14) ChatGPT can now create interactive visuals to help you understand math and science concepts | TechCrunch(00:11:56) Anthropic's Claude AI can respond with charts, diagrams, and other visuals now | The VergeProjects & Open Source(00:13:54) Introducing Nemotron 3 Super: An Open Hybrid Mamba-Transformer MoE for Agentic Reasoning | NVIDIA Technical BlogApplications & Business(00:21:22) Nvidia halts H200 production as China backs Huawei AI chips(00:28:33) Another XAI Cofounder Has Left, and Another Says He's Leaving. - Business Insider(00:34:04) Anthropic's Claude Marketplace allows customers to buy third-party cloud services | TechRadar(00:37:57) Yann LeCun's AMI Labs raises $1.03 billion to build world models | TechCrunch(00:44:52) Humanoid robotics maker Sunday reaches $1.15B valuation to build household robots | TechCrunchPolicy & Safety(00:46:09) Anthropic Sues Department of Defense Over ‘Supply Chain Risk' Label - The New York Times + Google and OpenAI Just Filed a Legal Brief in Support of Anthropic (00:53:24) Internal Pentagon memo orders military commanders to remove Anthropic AI technology from key systems - CBS News(00:58:15) Endogenous Resistance to Activation Steering in Language Models(01:06:27) Reasoning Models Struggle to Control their Chains of Thought(01:09:52) ‘It means missile defence on datacentres': drone strikes raise doubts over Gulf as AI superpower(01:14:57) Evidence for inference scaling in AI cyber tasks: Increased evaluation budgets reveal higher success rates(01:18:24) Frontier Models Can Take Actions at Low ProbabilitiesResearch & Advancements(01:24:20) Research note: Many SWE-bench-Passing PRs Would Not Be Merged into Main(01:28:26) [2603.03276] Beyond Language Modeling: An Exploration of Multimodal Pretraining(01:40:09) Memory Caching: RNNs with Growing Memory(01:48:47) Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking(01:58:41) CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation(02:08:57) Latent Introspection: Models Can Detect Prior Concept Injections(02:16:45) Physics of RL: Toy scaling laws for the emergence of reward-seekingSee Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
Guys I think its red I saw them vent... Justtt kidding nobody is ever an imposter in SWE!! Explore imposter syndrome with us as we share our experiences being a woman in male dominated fields. You belong where you are and you are always worth it. If you ever feel down and dumb come talk to us and we'll hype you up! Stay as gorgeous and smart as you are, you're always deserving and worth it. Love, Emma and Ava
Paige Feikert, research and technology engineer with Spirit AeroSystems, joins us live from the WE25 Diverse Podcast Studio to break down why communication can make or break your impact as an engineer — and how to get better at it. In conversation with Larry Guthrie, director of content strategy at SWE, Paige shares her unconventional career path from biomedical engineering student, to TV news producer at a CBS affiliate, and back into engineering. Hear why storytelling matters in technical work, actionable advice to help you communicate to executives and non-technical audiences, and what producing live TV news taught Paige about teamwork, deadlines, and handling pressure. — The Society of Women Engineers is a powerful, global force uniting 50,000 members of all genders spanning 85 countries. We are the world's largest advocate and catalyst for change for women in engineering and technology. To join and access all the exclusive benefits to elevate your professional journey, visit membership.swe.org.
As part of SWE's “Un Cafecito With a Woman in STEM” series, spotlighting Latina voices across the globe, this special episode of Diverse is presented in Spanish. Claudia Guerrero, SWE global ambassador and service program leader at GE Aerospace, sits down with host Doris Moreno Maldonado, process engineer at Kellogg's and lead of the SWE Latinos Affinity Group, for a conversation on antifragility, authenticity, and creating community wherever you are. Recorded live at WE25 in New Orleans, Claudia shares her journey as one of the first women in her family to study engineering in Mexico, her role in growing the SWE affiliate in Querétaro, and surviving a life-altering medical crisis. Hear how to move beyond resilience toward antifragility, build your personal board of directors, and lean into your unique authentic strengths. — The Society of Women Engineers is a powerful, global force uniting 50,000 members of all genders spanning 85 countries. We are the world's largest advocate and catalyst for change for women in engineering and technology. To join and access all the exclusive benefits to elevate your professional journey, visit membership.swe.org.
Google just dropped Gemini 3 Flash—a model that outperforms Gemini 2.5 Pro (their last top model) while running 3x faster at less than 1/4 the cost. It's frontier-level reasoning at Flash-level speed, and it's rolling out globally right now.We're sitting down with Logan Kilpatrick from Google DeepMind to explore what this actually means for developers, knowledge workers, and anyone trying to figure out how AI fits into their workflow.What we'll cover:
Carolina Caro, WE Local Portland keynote and CEO of Conscious Leadership Partners, breaks down communication blind spots for engineers and explores how your communication style shapes your leadership impact in this episode. As a scientist by training, Carolina reflects on why technical expertise isn't enough and speaks to the platinum rule, where you meet people where they are and communicate in the way they can best receive. In conversation with FY26 SWE President Inaas Darrat, hear how to strengthen your communication without compromising your authenticity and why communication style is an often-overlooked dimension of diversity. Carolina will expand on these insights as the opening keynote at WE Local Portland, taking place Feb. 27-28. WE Local conferences bring together engineers and technologists for networking opportunities, professional development, and inspirational speakers. Register at welocal.swe.org to join SWE at an upcoming WE Local near you! — The Society of Women Engineers is a powerful, global force uniting 50,000 members of all genders spanning 85 countries. We are the world's largest advocate and catalyst for change for women in engineering and technology. To join and access all the exclusive benefits to elevate your professional journey, visit membership.swe.org.
CoreStory is building code intelligence platforms that address the fundamental limitation of today's coding agents: their inability to navigate complex enterprise codebases. While foundation models excel at greenfield development, they fail at real-world engineering tasks in systems spanning millions of lines of code. CoreStory's context layer delivers a 44% improvement on SWE-bench, the industry's standard benchmark for measuring coding agent effectiveness on actual GitHub issues. In this episode of BUILDERS, I sat down with Anand Kulkarni, CEO of CoreStory, to explore how his team is enabling the shift to AI-native engineering and seeding the category of spec-driven development across Microsoft, GitHub, and Amazon. Topics Discussed: Building with GPT-3 API 18 months before ChatGPT went public Why even GPT-5 and Opus 4.5 struggle with enterprise codebases on SWE-bench The narrative shift required when selling AI pre- and post-ChatGPT CoreStory's 44% improvement in coding agent performance through context intelligence How "spec-driven development" got adopted by Microsoft, GitHub, and Amazon without formal analyst relations The parallel between JIRA monetizing Agile and CoreStory enabling AI-native engineering Three-channel distribution: direct enterprise, coding agent partnerships via MCP, and hyperscaler/GSI routes Why specs become the source of truth while code becomes disposable in the AI era GTM Lessons For B2B Founders: Match your narrative precision to technical depth: CoreStory deploys three distinct positioning strategies based on audience sophistication. For AI practitioners tracking benchmarks, they lead with "44% SWE-bench improvement"—a metric that immediately signals meaningful progress on the hardest problem in the space. For engineering leaders aware of AI tooling but not deep in the research, they focus on velocity gains and ROI metrics. For executives, they describe reverse-engineering codebases into machine-readable specs. The key insight: technical audiences dismiss vague value props, while non-technical audiences get lost in benchmark details. Map your positioning to how your audience measures success in their world. Seed category language through earned adoption, not manufactured consensus: Anand initially called their approach "requirements-driven development" before simplifying to "spec-driven development." Rather than pitching analysts, they used the term consistently in customer conversations, gave talks at GitHub Universe, and shipped demos showing the workflow. When customers naturally adopted the language and community leaders began using similar terminology independently, Microsoft and GitHub followed with their own implementations (like GitHub's SpecKit). The lesson: category language sticks when practitioners choose to use it because it clarifies their work, not because a vendor pushed it. Focus on customer adoption as proof of concept before seeking broader market validation. Position against emergent practices, not just incumbent products: CoreStory doesn't position against legacy code analysis tools—they position as the enabler of AI-native engineering, the discipline that will displace Agile. Anand's insight from watching JIRA's success: "People don't love JIRA. What they love is Agile as a way to move away from waterfall." CoreStory is betting that 10x velocity gains from AI-native practices will drive the same categorical shift. When you're early in a technology wave, attach to the practice change (how teams will work differently) rather than feature comparisons with existing tools. Movements create markets. Design channel strategy around customer problem awareness: CoreStory's three channels map to different stages of buyer sophistication. Direct enterprise comes from teams already deep in AI engineering who've hit the context limitation wall. Coding agent partnerships (via MCP integration with tools like Cognition and Factory) serve builders wanting better AI tooling who haven't diagnosed the context problem yet. Hyperscalers and GSIs distribute into modernization and maintenance projects where AI enablement is emerging as a requirement. Each channel serves a distinct buyer journey stage. Don't force one go-to-market motion—design multiple paths based on where different customer segments are in understanding the problem you solve. Navigate pre-legitimacy markets by hiding the breakthrough: Before ChatGPT, selling anything AI-driven faced immediate skepticism about whether it was "real" or just smoke and mirrors. Anand couldn't lead with AI without triggering disbelief. CoreStory focused on delivered outcomes—"here's what you'll be able to do"—with AI as the mechanism, not the message. Post-ChatGPT, the challenge flipped: everyone expects AI, but now the differentiation question becomes harder. If you're building on emerging technology before market consensus forms, deemphasize the technology until buyers have context to evaluate it. Once the market validates the technology category, shift to demonstrating your specific technical advantage within it. // Sponsors: Front Lines — We help B2B tech companies launch, manage, and grow podcasts that drive demand, awareness, and thought leadership. www.FrontLines.io The Global Talent Co. — We help tech startups find, vet, hire, pay, and retain amazing marketing talent that costs 50-70% less than the US & Europe. www.GlobalTalent.co // Don't Miss: New Podcast Series — How I Hire Senior GTM leaders share the tactical hiring frameworks they use to build winning revenue teams. Hosted by Andy Mowat, who scaled 4 unicorns from $10M to $100M+ ARR and launched Whispered to help executives find their next role. Subscribe here: https://open.spotify.com/show/53yCHlPfLSMFimtv0riPyM
Julie Daugherty, engineering associate and process project leader at Corning Incorporated, joins us live at WE25 to discuss how introverts can embrace their strengths and turn them into career superpowers. In conversation with Larry Guthrie, director of content strategy at SWE, hear how planning ahead builds confidence, how to survive “mandatory fun” networking events, and why finding extroverted champions can be key to career growth in STEM. You'll also learn how managers can support introverted engineers, what it means to be an ambivert, and why personality diversity leads to stronger teams and better problem-solving. — The Society of Women Engineers is a powerful, global force uniting 50,000 members of all genders spanning 85 countries. We are the world's largest advocate and catalyst for change for women in engineering and technology. To join and access all the exclusive benefits to elevate your professional journey, visit membership.swe.org.
Zoe, SWENext Influencer and FIRST Robotics leader, and Charlee, a member of the 2025 STEM Next Flight Crew youth ambassador program, join us live from the WE25 Diverse Podcast Studio in New Orleans to share their experiences as young STEM leaders. In conversation with Larry Guthrie, director of content strategy at SWE, they reflect on how Invent It. Build It. sparked powerful connections, what leadership looks like as a precollege student, and how they navigate bias, responsibility, and pressure as “the first” in their communities. Hear their experiences starting STEM clubs and expanding access in rural areas, plus why a strong community is critical to shaping a more inclusive future in engineering. — The Society of Women Engineers is a powerful, global force uniting 50,000 members of all genders spanning 85 countries. We are the world's largest advocate and catalyst for change for women in engineering and technology. To join and access all the exclusive benefits to elevate your professional journey, visit membership.swe.org.
This is a link post. Improving model performance by scaling up inference compute is the next big thing in frontier AI. But the charts being used to trumpet this new paradigm can be misleading. While they initially appear to show steady scaling and impressive performance for models like o1 and o3, they really show poor scaling (characteristic of brute force) and little evidence of improvement between o1 and o3. I explore how to interpret these new charts and what evidence for strong scaling and progress would look like. From scaling training to scaling inference The dominant trend in frontier AI over the last few years has been the rapid scale-up of training — using more and more compute to produce smarter and smarter models. Since GPT-4, this kind of scaling has run into challenges, so we haven't yet seen models much larger than GPT-4. But we have seen a recent shift towards scaling up the compute used during deployment (aka 'test-time compute' or ‘inference compute'), with more inference compute producing smarter models. You could think of this as a change in strategy from improving the quality of your employees' work via giving them more years of training in which acquire [...] --- First published: February 2nd, 2026 Source: https://forum.effectivealtruism.org/posts/zNymXezwySidkeRun/inference-scaling-and-the-log-x-chart Linkpost URL:https://www.tobyord.com/writing/inference-scaling-and-the-log-x-chart --- Narrated by TYPE III AUDIO. ---Images from the article:
This episode is sponsored by BD. Showing up as your authentic self in engineering isn't always easy. In this conversation, Christine Kearney Hawkins, senior staff R&D engineer in BD's Peripheral Intervention business and SWE life member, shares her 20+ year journey navigating authenticity, leadership, and innovation in STEM. From being told that she couldn't be both an engineer and a mom, to learning that her bubbly enthusiasm is a strength and not a liability, Christine reflects on how embracing who she is shaped her career and impact. In conversation with host Sam East, hear how authentic leadership fuels better innovation outcomes, what to do when workplace feedback conflicts with your core values, and practical advice to create cultures where people feel safe bringing their whole selves to work. — The Society of Women Engineers is a powerful, global force uniting 50,000 members of all genders spanning 85 countries. We are the world's largest advocate and catalyst for change for women in engineering and technology. To join and access all the exclusive benefits to elevate your professional journey, visit membership.swe.org.
Stepping into STEM leadership doesn't require a senior title or having everything figured out. In this episode, Katie Ashley, volunteer coordinator of the SWE Early Career Professionals Affinity Group (SWE ECP AG), is joined by Zoe Husted, ECP AG conferences and awards co-chair and president of the SWE Golden Gate Section, and Kathryn Wittek, ECP AG design coordinator and president of the SWE Baltimore-Washington Section, to explore what leadership can look like in the beginning years of an engineering career. Drawing from their experiences in SWE, team sports, and technical roles, Zoe and Kathryn share when they started seeing themselves as leaders, how they navigate leading more experienced colleagues, and why learning and leading often happen at the same time. Hear their tips to find community as an early-career engineer, plus how the skills they have developed through SWE have translated into the workplace. The SWE ECP AG was formed to equip individuals with the support, resources, and inclusive community to excel in the first ten years of their career. Get involved and find out about upcoming ECP AG events at https://earlycareerprofessionalsag.swe.org/. — The Society of Women Engineers is a powerful, global force uniting 50,000 members of all genders spanning 85 countries. We are the world's largest advocate and catalyst for change for women in engineering and technology. To join and access all the exclusive benefits to elevate your professional journey, visit membership.swe.org.
From creating SWE-bench in a Princeton basement to shipping CodeClash, SWE-bench Multimodal, and SWE-bench Multilingual, John Yang has spent the last year and a half watching his benchmark become the de facto standard for evaluating AI coding agents—trusted by Cognition (Devin), OpenAI, Anthropic, and every major lab racing to solve software engineering at scale. We caught up with John live at NeurIPS 2025 to dig into the state of code evals heading into 2026: why SWE-bench went from ignored (October 2023) to the industry standard after Devin's launch (and how Walden emailed him two weeks before the big reveal), how the benchmark evolved from Django-heavy to nine languages across 40 repos (JavaScript, Rust, Java, C, Ruby), why unit tests as verification are limiting and long-running agent tournaments might be the future (CodeClash: agents maintain codebases, compete in arenas, and iterate over multiple rounds), the proliferation of SWE-bench variants (SWE-bench Pro, SWE-bench Live, SWE-Efficiency, AlgoTune, SciCode) and how benchmark authors are now justifying their splits with curation techniques instead of just “more repos,” why Tau-bench's “impossible tasks” controversy is actually a feature not a bug (intentionally including impossible tasks flags cheating), the tension between long autonomy (5-hour runs) vs. interactivity (Cognition's emphasis on fast back-and-forth), how Terminal-bench unlocked creativity by letting PhD students and non-coders design environments beyond GitHub issues and PRs, the academic data problem (companies like Cognition and Cursor have rich user interaction data, academics need user simulators or compelling products like LMArena to get similar signal), and his vision for CodeClash as a testbed for human-AI collaboration—freeze model capability, vary the collaboration setup (solo agent, multi-agent, human+agent), and measure how interaction patterns change as models climb the ladder from code completion to full codebase reasoning.We discuss:* John's path: Princeton → SWE-bench (October 2023) → Stanford PhD with Diyi Yang and the Iris Group, focusing on code evals, human-AI collaboration, and long-running agent benchmarks* The SWE-bench origin story: released October 2023, mostly ignored until Cognition's Devin launch kicked off the arms race (Walden emailed John two weeks before: “we have a good number”)* SWE-bench Verified: the curated, high-quality split that became the standard for serious evals* SWE-bench Multimodal and Multilingual: nine languages (JavaScript, Rust, Java, C, Ruby) across 40 repos, moving beyond the Django-heavy original distribution* The SWE-bench Pro controversy: independent authors used the “SWE-bench” name without John's blessing, but he's okay with it (”congrats to them, it's a great benchmark”)* CodeClash: John's new benchmark for long-horizon development—agents maintain their own codebases, edit and improve them each round, then compete in arenas (programming games like Halite, economic tasks like GDP optimization)* SWE-Efficiency (Jeffrey Maugh, John's high school classmate): optimize code for speed without changing behavior (parallelization, SIMD operations)* AlgoTune, SciCode, Terminal-bench, Tau-bench, SecBench, SRE-bench: the Cambrian explosion of code evals, each diving into different domains (security, SRE, science, user simulation)* The Tau-bench “impossible tasks” debate: some tasks are underspecified or impossible, but John thinks that's actually a feature (flags cheating if you score above 75%)* Cognition's research focus: codebase understanding (retrieval++), helping humans understand their own codebases, and automatic context engineering for LLMs (research sub-agents)* The vision: CodeClash as a testbed for human-AI collaboration—vary the setup (solo agent, multi-agent, human+agent), freeze model capability, and measure how interaction changes as models improve—John Yang* SWE-bench: https://www.swebench.com* X: https://x.com/jyangballinFull Video EpisodeTimestamps00:00:00 Introduction: John Yang on SWE-bench and Code Evaluations00:00:31 SWE-bench Origins and Devon's Impact on the Coding Agent Arms Race00:01:09 SWE-bench Ecosystem: Verified, Pro, Multimodal, and Multilingual Variants00:02:17 Moving Beyond Django: Diversifying Code Evaluation Repositories00:03:08 Code Clash: Long-Horizon Development Through Programming Tournaments00:04:41 From Halite to Economic Value: Designing Competitive Coding Arenas00:06:04 Ofir's Lab: SWE-ficiency, AlgoTune, and SciCode for Scientific Computing00:07:52 The Benchmark Landscape: TAU-bench, Terminal-bench, and User Simulation00:09:20 The Impossible Task Debate: Refusals, Ambiguity, and Benchmark Integrity00:12:32 The Future of Code Evals: Long Autonomy vs Human-AI Collaboration00:14:37 Call to Action: User Interaction Data and Codebase Understanding Research Get full access to Latent.Space at www.latent.space/subscribe
From creating SWE-bench in a Princeton basement to shipping CodeClash, SWE-bench Multimodal, and SWE-bench Multilingual, John Yang has spent the last year and a half watching his benchmark become the de facto standard for evaluating AI coding agents—trusted by Cognition (Devin), OpenAI, Anthropic, and every major lab racing to solve software engineering at scale. We caught up with John live at NeurIPS 2025 to dig into the state of code evals heading into 2026: why SWE-bench went from ignored (October 2023) to the industry standard after Devin's launch (and how Walden emailed him two weeks before the big reveal), how the benchmark evolved from Django-heavy to nine languages across 40 repos (JavaScript, Rust, Java, C, Ruby), why unit tests as verification are limiting and long-running agent tournaments might be the future (CodeClash: agents maintain codebases, compete in arenas, and iterate over multiple rounds), the proliferation of SWE-bench variants (SWE-bench Pro, SWE-bench Live, SWE-Efficiency, AlgoTune, SciCode) and how benchmark authors are now justifying their splits with curation techniques instead of just "more repos," why Tau-bench's "impossible tasks" controversy is actually a feature not a bug (intentionally including impossible tasks flags cheating), the tension between long autonomy (5-hour runs) vs. interactivity (Cognition's emphasis on fast back-and-forth), how Terminal-bench unlocked creativity by letting PhD students and non-coders design environments beyond GitHub issues and PRs, the academic data problem (companies like Cognition and Cursor have rich user interaction data, academics need user simulators or compelling products like LMArena to get similar signal), and his vision for CodeClash as a testbed for human-AI collaboration—freeze model capability, vary the collaboration setup (solo agent, multi-agent, human+agent), and measure how interaction patterns change as models climb the ladder from code completion to full codebase reasoning. We discuss: John's path: Princeton → SWE-bench (October 2023) → Stanford PhD with Diyi Yang and the Iris Group, focusing on code evals, human-AI collaboration, and long-running agent benchmarks The SWE-bench origin story: released October 2023, mostly ignored until Cognition's Devin launch kicked off the arms race (Walden emailed John two weeks before: "we have a good number") SWE-bench Verified: the curated, high-quality split that became the standard for serious evals SWE-bench Multimodal and Multilingual: nine languages (JavaScript, Rust, Java, C, Ruby) across 40 repos, moving beyond the Django-heavy original distribution The SWE-bench Pro controversy: independent authors used the "SWE-bench" name without John's blessing, but he's okay with it ("congrats to them, it's a great benchmark") CodeClash: John's new benchmark for long-horizon development—agents maintain their own codebases, edit and improve them each round, then compete in arenas (programming games like Halite, economic tasks like GDP optimization) SWE-Efficiency (Jeffrey Maugh, John's high school classmate): optimize code for speed without changing behavior (parallelization, SIMD operations) AlgoTune, SciCode, Terminal-bench, Tau-bench, SecBench, SRE-bench: the Cambrian explosion of code evals, each diving into different domains (security, SRE, science, user simulation) The Tau-bench "impossible tasks" debate: some tasks are underspecified or impossible, but John thinks that's actually a feature (flags cheating if you score above 75%) Cognition's research focus: codebase understanding (retrieval++), helping humans understand their own codebases, and automatic context engineering for LLMs (research sub-agents) The vision: CodeClash as a testbed for human-AI collaboration—vary the setup (solo agent, multi-agent, human+agent), freeze model capability, and measure how interaction changes as models improve — John Yang SWE-bench: https://www.swebench.com X: https://x.com/jyangballin Chapters 00:00:00 Introduction: John Yang on SWE-bench and Code Evaluations 00:00:31 SWE-bench Origins and Devon's Impact on the Coding Agent Arms Race 00:01:09 SWE-bench Ecosystem: Verified, Pro, Multimodal, and Multilingual Variants 00:02:17 Moving Beyond Django: Diversifying Code Evaluation Repositories 00:03:08 Code Clash: Long-Horizon Development Through Programming Tournaments 00:04:41 From Halite to Economic Value: Designing Competitive Coding Arenas 00:06:04 Ofir's Lab: SWE-ficiency, AlgoTune, and SciCode for Scientific Computing 00:07:52 The Benchmark Landscape: TAU-bench, Terminal-bench, and User Simulation 00:09:20 The Impossible Task Debate: Refusals, Ambiguity, and Benchmark Integrity 00:12:32 The Future of Code Evals: Long Autonomy vs Human-AI Collaboration 00:14:37 Call to Action: User Interaction Data and Codebase Understanding Research
Karen Horting, Executive Director and CEO of the Society of Women Engineers, talks about SWE's archives at the Reuther Library and shares how the 75-year-old organization leverages its history to advocate for the inclusion of women in science, technology, engineering, and mathematics (STEM). Related Resources: Society of Women Engineers 75th Anniversary SWE Archives Virtual Tour [Part 1] SWE Archives Virtual Tour [Part 2] Related Collections: Society of Women Engineers Records (LR001539) Society of Women Engineers Publications (LR002487) Episode Credits Interviewee: Karen Horting Producers: Dan Golodner and Troy Eller English Music: Bart Bealmear
Emily breaks down her 2025 Holiday Gift Guide through the lens of the five pillars of sexual intelligence—embodiment, health, self-knowledge, self-acceptance, and collaboration. She explores how shifting from performance to presence can transform your relationship with pleasure, offering curated tools and practices that help you slow down, feel your body, and understand yourself as a sexual being. This episode is for anyone ready to move beyond quick fixes and embrace a more holistic approach to sexuality, turning intimacy into something that nourishes your whole system. In this episode, you'll learn: • The five pillars of Sex IQ create a holistic framework for understanding your sexuality and taking responsibility for your pleasure—moving you from performance to presence • Embodiment practices like breathwork and sonic wave toys help you overcome the mind-body disconnect during sex by bringing you back to physical sensations instead of staying in your head worrying about how you look or what you need to do • Self-acceptance isn't about waiting for your body to change—it's about recognizing and reframing the pleasure thieves (stress, trauma, and shame) that tell you you're unworthy, and actively replacing negative thought patterns with affirmations that honor your body right now More Dr. Emily: • The 2025 Sex With Emily Holiday Gift Guide • Apply for Emily's 1:1 Coaching Opportunity HERE or reach out to enrollment@sexwithemily.com for more information • Shop With Emily! Explore Emily's favorite toys, pleasure accessories, bedroom essentials, and more — designed to support your pleasure and confidence. Free shipping on orders $99+ (some exclusions apply). • Join the SmartSX Membership: Access exclusive sex coaching, live expert sessions, community building, and tools to enhance your pleasure and relationships with Dr. Emily Morse. • Yes! No! Maybe? List & Other Sex With Emily Guides: Explore pleasure, deepen connections, and enhance intimacy using these Sex With Emily downloadable guides. • The only sex book you'll ever need: Smart Sex: How to Boost Your Sex IQ and Own Your Pleasure • Want more? Visit the Sex With Emily Website • Let's get social: Instagram | X | Facebook | TikTok | Threads | YouTube • Let's text: Sign up here • Want me to slide into your email inbox? Sign Up Here for sex tips on the regular. Shop the Holiday Gift Guide now! (See the full Gift Guide HERE) Embodiment: • LELO SONA 3- Use code EMILY20 for 20% on top of ongoing sales at lelo.com • Common Confidential Massage Butter- Use code SEXWITHEMILY for 15% off at Commonconfidential.com. • Cornbread Hemp - Use code SWE for a discount at cornbreadhemp.com/swe Health: • Bathmate Hydro Pump- Use SWE10 for 10% off your order at bathmatedirect.com • HigherDOSE Infrared Sauna Blanket- Head to higherdose.com • V-Health Vaginal Rejuvenation Gel- Use code EMILY10 for 10% off at getvhealth.com • Kroma Wellness Beauty Matcha Latte - Use code SEXWITHEMILY for 15% off at kromawellness.com Collaboration: • Promescent Delay Spray- Head to promescent.com • LELO F2S- Use code EMILY20 for 20% on top of ongoing sales at lelo.com • Crave Leather Handcuffs- Head to shop.sexwithemily.com/crave Self- Knowledge: • Magic Wand Mini - Head to shop.sexwithemily.com/magicwand • Je Joue Hera Flex Rabbit Vibrator - Go to sexwithemily.com/hera and use code EMILY20 for 20% off • SmartSX- Head to sexwithemily.com/smartsx Self- Acceptance: • The Class - Head to theclass.com • Droplette- Head to droplette.io This episode is sponsored by… Biolouve- Get 15% off with code EMILYPOWERMOVE @ https://biolouve.com/ Timestamps 0:00 - Introduction 1:25 - The 5 Pillars of Sex IQ Framework 2:36 - Pillar 1: Embodiment 10:23 - Pillar 2: Health 17:52 - Pillar 3: Collaboration 24:44 - Pillar 4: Self-Knowledge 29:39 - Pillar 5: Self-Acceptance 34:21 - Wrap-Up
This episode features Dianne Na Penn, a senior product leader at Anthropic, discussing the launch of Claude Opus 4.5 and the evolution of frontier AI models. The conversation explores how Anthropic approaches model development—balancing ambitious capability roadmaps with user feedback, making strategic bets on areas like agentic coding and computer use while deliberately avoiding others like image generation. Dianne shares insights on the shifting nature of AI evaluation (moving beyond saturated benchmarks like SWE-bench toward more open-ended measures), the evolution of scaffolding from "training wheels" to intelligence amplifiers, and why she believes we're closer to transformative long-running AI than most people think. She also discusses Anthropic's distinctive culture of authenticity, the under appreciated benefits of model alignment for producing independent-thinking AI, and why the real bottleneck to AI agents isn't model capability anymore but product innovation. (0:00) Intro(0:57) Starting the Work on Opus 4.5(2:04) Model Capabilities and Surprises(5:59) Computer Use and Practical Applications(7:21) Pricing and Positioning(10:02) Customer Feedback and Early Access(16:44) The Reality of Enterprise Agents(18:47) Future of AI and Long-Running Intelligence(28:06) Anthropic's Culture and Decision Making(30:31) Key Decisions and Fun Moments(33:45) Quickfire With your co-hosts: @jacobeffron - Partner at Redpoint, Former PM Flatiron Health @patrickachase - Partner at Redpoint, Former ML Engineer LinkedIn @ericabrescia - Former COO Github, Founder Bitnami (acq'd by VMWare) @jordan_segall - Partner at Redpoint
Google just released Gemini 3, and I tested it with 3 real business use cases: prepping for a $200K sales call, building an interactive sales dashboard, and creating a complete website from scratch—all in minutes.In this video, I show you LIVE examples of:✅ Gemini Agent prepping an entire sales call (research, pitch angles, objection handling, follow-ups)✅ Dynamic dashboards with real-time calculations and interactive sliders✅ Full website generation with 921 lines of working code in under 60 seconds✅ Custom image generation and prompt engineeringThis isn't theory—I'm screen recording everything as it happens so you can see the actual speed and quality.⏱️ TIMESTAMPS:0:00 - Intro: Why Gemini 3 is a Big Deal1:15 - Benchmark Breakdown (Why This Matters)2:27 - Use Case #1: $200K Sales Call Prep4:46 - Visual Layout & Interactive Infographic Demo7:03 - Use Case #2: Building a Website from Scratch (Live Code)9:27 - Use Case #3: Interactive Q4 Sales Dashboard11:51 - Testing the Prompt Generator Live12:45 - Final Thoughts & Why This Changes Everything
Jessica reports LIVE from Jakarta on all the details from day two of women's podium training. World Championships Headquarters Videos, Interviews, Podcasts, Fantasy, Guides Extended Episode + Live Q&A (Members) +30 extra minutes of analysis, behind-the-scenes secret stories, plus member questions. Here's how to ask questions live. Can't make it live? Add Club bonus episodes to your favorite podcast player (instructions here). Chapters 00:00 – Show Intro 01:02 – Zhang Qingying beam world champion prediction 03:00 – FIG Press Conference recap: AI D-scores and visa issue 08:40 – Spencer's updates: where to watch & fantasy game deadlines 11:45 – U.S. Women's Team podium training report (Josc, Skye, Dulcy, Leanne) 17:20 – Can Josc vault? Exclusive Olympic Channel interview 19:45 – Equipment update: white mats and “China mat overlay” 22:10 – Mixed Zone highlights (Malabuyo, South Africa, Asia's coach impression) 25:05 – Italy updates: Perotti, Asia D'Amato, Fioravanti AA potential 29:45 – Melnikova and Russia (AIN) podium impressions 31:30 – Flavia Saraiva's 10.0 leotard and Brazilian updates 33:10 – Funniest & coolest skills of the day (Chile, India, Portugal) 33:55 – BTS Teaser begins 34:00 – Embarrassing moments & Watanabe press conference story 36:40 – Beam fall hilarity (NZL gymnast) 38:15 – Opposite of Canadian medical intervention 40:00 – The great Indonesian tampon saga 42:25 – Sub 4: NZL, LIE, USA, CRO, BAN, GBR, POL 45:10 – Ruby Evans Amanar, GB bars, Alia Leat injury update 47:05 – Sub 5: MAS, SUI, ITA, FRA, VIE, ISL, MAR 49:00 – Thelma's floor, Osyssek's beam, Ming Van Eijken vaults 51:05 – Sub 6: AUS, EGY, BEL, LAT, ROU, MGL, SWE, CRC 53:00 – Voinea full Gothic mode, Golgota AA, Romanian updates 56:20 – Sub 7: INA, TUN, COL, PHI, MEX, SYR 58:00 – Finnegan & Malabuyo AA, Seema Tello debut 1:00:10 – Sub 8: NOR, BRA, QAT, IND, RSA, CHI 1:02:15 – Flavia & Brazil updates, Rooskrantz, Chilean grandmas 1:05:00 – Sub 9: AIN, NAM, POR, THA, BUL, SLO, CMR 1:07:25 – Melnikova Cheng, Cameroon floor joy, AIN medal watch 1:10:10 – Sub 10: ESP, AIN, HUN, HKG, CHN, KZN, CZE 1:12:25 – Zhou Yaqin & Zhang Qingying on beam, Deng Yalan vault 1:15:30 – Alba Petisco all-around standout 1:17:10 – Feedback: listener comments from Dr. Ben & Absolutely Not 1:21:20 – Show Close: Women's qualifying preview & thanks How Do I Watch the Competition? All sessions of the competition will be streamed on Eurovision Sport. Follow along here! Gymnastics Indonesia's YouTube channel will stream all qualification sessions Live scores from the FIG and Swiss Timing Check out NBC's behind-the-scenes mini-doc on the US Women's World Trials Headlines What happened at podium training today? Should we be worried about the US women? From the Olympic Channel: Joscelyn Roberson has been struggling to "find her block" on vault Skye's HUGE front-handspring front on beam Who else from Florida came to join the 2025 World Championships party? Giulia Perotti (Italy) looks ready to win all the medals Who will be the second Italian competing all-around? The D'Amato vs. Fioravanti dilemma Angelina Melnikova is so back How did her vaults look? WE NEED TO TALK ABOUT BRAZIL'S GENIUS LEOS Flavia showed beam and floor - how'd it go? Who wins the award for coolest/best/most fun skill from podium training? What were Jessica's mixed zone highlights? The FIG held a press conference today. What information did we learn? The FIG announced that "spectators will be able to see AI D-scores," but what does this mean? The FIG addressed the visa vs. FIG rules issue. What did FIG president Watanabe have to say? Jakarta Updates GymCastic Updates Subscribe to our YouTube Channel Coming Up 6 days of LIVE podcasts at World Championships in Jakarta Club members get extended coverage and can join us live to ask questions immediately after the meet Play our World Championships Fantasy Game! Win a Club Gym Nerd Scholarship: Go to our Forum > Show Stuff > GymCastic Scholarship We are matching every new sponsorship If you would like access to the club content, but aren't currently in a position to purchase a membership, all you need to do is fill out the form that's linked in our message board If you would also like to sponsor a scholarship, please email editor@gymcastic.com. Thank you! Support Our Work Club Gym Nerd: Join Here Become a Sponsor: GymCastic is matching all donations Nearly 50 scholarships have been awarded so far Learn More Headstand Game: Play Now Forum: Start Chatting Merch: Shop Now Thank you to our Sponsors Gymnastics Medicine Beam Queen Bootcamp's Overcoming Fear Workshop Resources Jakarta schedule & times: See our live podcast times on the Worlds HQ schedule Guides: Download the quick-reference guide on the Jakarta Headquarters page The Balance Beam Situation: Spencer's GIF Code of Points Gymnastics History and Code of Points Archive from Uncle Tim Kensley's men's gymnastics site Neutral Deductions Unlock the Extended Episode Join Club Gym Nerd → Choose a plan Complete checkout — your site account is created. Log in here → /my-account/ Return to this page and refresh. The extended player appears automatically.
Ken Rosenthal's Bad LookLetter From Beyond The GraveHe Must Be Really Good!TEMU For QB'sWe're Not Doing Stonehenge ToniteWe Talkin' About Practice!About That Injury Report ThingHerm Edwards, Your Rant Is SafeAnother Reason To Loathe The ChiefsA Day Later, SorryDo You Even Kicker, Bro?That Play You Saw? No You Didn'tJoe Burrow To The SlaughterGo To the Game, It'll Be Fun!Cracker Barrell Total Surrender!Our Sponsors:* Check out Hims: https://hims.com/CZABE* Check out Indeed: https://indeed.com/CZABEAdvertising Inquiries: https://redcircle.com/brandsPrivacy & Opt-Out: https://redcircle.com/privacy