Podcasts about cpu

Central component of any computer system which executes input/output, arithmetical, and logical operations

  • 2,121PODCASTS
  • 6,376EPISODES
  • 53mAVG DURATION
  • 1DAILY NEW EPISODE
  • Sep 17, 2026LATEST
cpu

POPULARITY

20192020202120222023202420252026

Categories



Best podcasts about cpu

Show all podcasts related to cpu

Latest podcast episodes about cpu

Atareao con Linux
ATA 832 Watchbit, el monitor de uptime ultraligero hecho en Rust

Atareao con Linux

Play Episode Listen Later Sep 17, 2026 20:50


¿Cansado de que tu monitor de uptime consuma más recursos que los propios servicios que monitoriza? Llevaba años usando Uptime Kuma, una herramienta fantástica, potente y muy fácil de configurar. Pero cuando miras el docker stats y ves 150 MB de RAM, 900 MB de disco y un 2% de CPU para monitorizar solo seis páginas, algo no cuadra. Sobre todo cuando lo único que quieres es que te llegue una notificación si algo falla, no tener un mini-Netflix corriendo en tu servidor.Por eso creé Watchbit: un monitor de uptime escrito en Rust con frontend en React, base de datos SQLite embebida y autenticación OIDC con Pocket ID. El resultado: 7 MB de RAM, 20 MB de disco y 0,2% de CPU. Sí, has leído bien. Estamos hablando de reducir el consumo de memoria a una vigésima parte, el de disco a una cuadragésima parte y el de CPU a una décima parte. Y todo esto sin perder funcionalidad: monitores HTTP, TCP, Ping, heartbeats, notificadores vía Telegram, Matrix, NTFY, Gotify, Discord, Email, y páginas de estado públicas.En este episodio te cuento cómo migré de Uptime Kuma a Watchbit, te enseño el dashboard por tarjetas, los heartbeats, los monitores, los notificadores y las 7 razones por las que no pienso volver atrás. También te explico el stack técnico: Rust con Axum y Tokio en el backend, React con TypeScript en el frontend, SQLite como base de datos embebida (adiós a PostgreSQL y MariaDB), y OIDC con Pocket ID para olvidarte de gestionar usuarios y contraseñas. Todo en un solo binario, sin dependencias externas, con una imagen Docker que apenas ocupa 20 MB.Además te muestro cómo configurar los monitores con intervalos personalizados, las plantillas de notificación para eventos de down, up, latencia y expiración de certificados, y el sistema de backup integrado que te permite exportar e importar toda la configuración en un solo clic. También te cuento el proceso de optimización que seguí: inicialmente hacía un bucle que recorría todos los monitores, pero después los separé en tareas independientes con Tokio para maximizar la eficiencia.Si tienes un VPS ajustado, una Raspberry Pi Zero, o simplemente quieres ser racional con los recursos de tu servidor, este episodio te interesa. Porque al final, de esto va el self-hosting: de tener herramientas que hagan su trabajo sin que el servidor se resienta. Como siempre, te dejo el docker-compose, las variables de entorno y las instrucciones en las notas del episodio para que puedas probarlo tú mismo en cinco minutos.Capítulos:0:00 - Introducción: la necesidad de monitorizar páginas web1:54 - Uptime Kuma: características y panel de control4:45 - Ventajas e inconvenientes de Uptime Kuma5:36 - El problema del consumo: CPU, RAM y disco7:05 - Watchbit: OIDC, dashboard y diseño por tarjetas8:43 - Comparativa de consumo: Watchbit vs Uptime Kuma10:53 - Heartbeats y monitores en Watchbit12:37 - Configuración de monitores, notificadores y plantillas14:31 - Ajustes, backup y stack técnico (Rust + React + SQLite)17:26 - 7 razones para migrar e instalaciónMás información y enlaces en las notas del episodio

TechLinked
DLSS 5 Mods Go Wild, AI Safety Researchers Quit Anthropic and Google, Altman Open to Slowing OpenAI Down + more!

TechLinked

Play Episode Listen Later Sep 12, 2026 9:01


News Sources: https://lmg.gg/YZuHq Timestamps: 0:00 DLSS 5 mods run wild 1:53 AI researchers sound the alarm 4:50 QUICK BITS INTRO 4:58 Meta Project Phoenix leak 5:26 Meta AI prompts get invasive 5:55 AMD budget CPU lifeline 6:27 Steam age checks and California rules 7:14 Nitter is back 7:33 Fly brain plays Doom 8:05 Credits Learn more about your ad choices. Visit megaphone.fm/adchoices

Automation Ladies
Automate mini series: Whitney Sales

Automation Ladies

Play Episode Listen Later Sep 10, 2026 54:28 Transcription Available


“We do it with AI” is easy to say. Making robots work through contact, uncertainty, and variation on a real factory floor is the hard part. We sit down with Whitney Dills from Thought Forge AI to unpack an approach to industrial robotics control that flips a common assumption: more data does not automatically mean better performance.Whitney explains active inference, a neuroscience rooted framework that treats robot control as continuous prediction and correction using real time sensory feedback. Instead of bloated backpropagation models that demand huge datasets and retraining for every edge case, Thought Forge trains quickly on small amounts of data and adapts after deployment. We talk about why that matters for insertion and complex manipulation, where torque, joint state, and haptics decide success or failure. If you have ever watched a slick picking demo fall apart the moment parts vary, lighting shifts, or fixtures drift, this conversation connects the dots.We also get practical about adoption: edge AI that can run on a CPU, the risks of GPU dependence, and how data ownership and IP concerns shape buyer trust. Whitney shares why their models stay on premise, why off the shelf robots and sensors often win over novel vertically integrated stacks, and how to think about cobots versus humanoids when reliability and cycle time are non negotiable.Subscribe for more conversations that cut through robotics hype, share this with someone evaluating “physical AI” for manufacturing, and leave a review with the one automation task you most want AI to finally handle.Support the show_________________________________________________________________

In Numbers We Trust - Der Data Science Podcast
#102: [PAIQ5] Predictive AI Quarterly

In Numbers We Trust - Der Data Science Podcast

Play Episode Listen Later Sep 10, 2026 37:08


Die letzten Monate waren geprägt von neuen Foundation Models für tabellarische und sequentielle Daten: TabPFN 3 skaliert auf eine Million Beobachtungen, Google stellt mit TabFM ein eigenes Modell samt BigQuery-Integration vor, NXAI veröffentlicht TiRex-2 mit Kovariaten-Unterstützung, und an der Spitze von GIFT-Eval steht mit STRIDE eine Kombination aus LLM-Reasoning und Time Series Foundation Model. Dazu kommen der ClickHouse-MCP-Server und der Stand der Umsetzung des EU AI Act nach dem Digital Omnibus. Im Praxisteil vergleichen wir TabICL v2 mit einem getunten XGBoost, Meta Prophet und naiven Baselines auf stündlichen NO2-Messwerten von fünf Messstationen, ausgewertet über ein Jahr rollierender Kreuzvalidierung mit dem Mean Absolute Scaled Error. Wir zeigen, welches Feature-Engineering nötig ist, wie sich der Vorteil von TabICL mit der Länge der Historie verändert und was das an Rechenzeit kostet. Zum Schluss ordnen wir ein, wann sich ein Foundation Model für Zeitreihen anbietet und wann XGBoost die pragmatischere Wahl bleibt.   **Zusammenfassung** TabPFN 3 (Mai 2026) skaliert auf einer H100 auf bis zu 1 Mio. Beobachtungen; verbessertes KV-Caching senkt die Prognosezeit auf 0,1–3 ms pro Testbeobachtung und macht das Modell für schnelle Batch-Prognosen nutzbar. Googles TabFM ist von TabPFN und TabICL inspiriert, liegt im TabArena-Benchmark vor TabPFN 3 und ist direkt in BigQuery integriert. TiRex-2 von NXAI setzt auf eine xLSTM- statt Transformer-Architektur und kann jetzt zusätzliche Kovariate einbeziehen – ein Test steht bei uns noch aus. STRIDE führt den GIFT-Eval-Benchmark an: Das LLM prognostiziert nicht selbst, sondern steuert über destillierte Embeddings ein Time Series Foundation Model. Kurz notiert: Der ClickHouse-MCP-Server (v0.4.1) erlaubt LLM-Abfragen ohne SQL, etwa zur Log-Diagnose; beim EU AI Act gelten die Transparenzpflichten seit August, die Kennzeichnung von Bestandssystemen greift ab dem 2.12.2026. Praxis-Setup: TabICL v2 mit Kalender-, Fourier- und Lag-Features gegen getuntes XGBoost, Prophet, TabICL out-of-the-box sowie Naive und Seasonal-Naive; stündliche NO2-Daten, 24-Stunden-Horizont, Metrik MASE. Ergebnisse: Bei zwei Jahren Historie liegt TabICL klar vorn, bei rund 8.000–9.000 Trainingsbeobachtungen ist XGBoost praktisch gleichauf, bei drei Monaten Historie noch etwa 4 % schlechter; ohne jedes Feature-Engineering schlägt TabICL Prophet und die naiven Baselines deutlich. Kosten: Die Kreuzvalidierung mit TabICL auf einer L40S-GPU dauert etwa 17-mal länger als mit XGBoost, auf CPU ist das Modell nicht praktikabel – Caching dürfte diesen Nachteil künftig verkleinern.   **Links** Link zum begeleitenden Blogartikel "TabICL v2 für Zeitreihen: Das In-Context-Learning-Modell im Vergleich mit XGBoost und Meta's Prophet" https://www.inwt-statistics.de/blog/tabicl_v2_fuer_zeitreihen #72: TabPFN: Die KI-Revolution für tabulare Daten mit Noah Hollmann https://www.podbean.com/ew/pb-94ri2-18aca83 #57: Mehr als heiße Luft: unsere Berliner Luftschadstoffprognose mit Dr. Andreas Kerschbaumer https://www.podbean.com/ew/pb-u6xwt-16ff139 TabPFN-3 Technical Report: https://priorlabs.ai/technical-reports/tabpfn-3 TabPFN auf GitHub: https://github.com/PriorLabs/TabPFN Google Research zu TabFM: https://research.google/blog/introducing-tabfm-a-zero-shot-foundation-model-for-tabular-data/ TabFM in BigQuery: https://cloud.google.com/blog/products/data-analytics/tabfm-adds-predictive-ml-to-bigquery TiRex-2 (NXAI): https://www.nx-ai.com/en/tirex-2 | Code: https://github.com/NX-AI/tirex-2 | Paper: https://arxiv.org/abs/2607.01204 STRIDE – Reasoning-Aware Training for Time Series Forecasting: https://arxiv.org/abs/2605.08625 Time Series LLMs am Beispiel t0-alpha: https://towardsdatascience.com/time-series-llms-explained-with-t0-alpha/ ClickHouse MCP Server: https://github.com/ClickHouse/mcp-clickhouse EU AI Act nach dem Digital Omnibus (Überblick): https://www.deloitte.com/de/de/issues/innovation-ai/eu-ai-act-digital-omnibus.html TabICL v2 auf GitHub: https://github.com/soda-inria/tabicl TabICL-Dokumentation: https://tabicl.readthedocs.io/en/latest/ Tutorial zum TabICLForecaster: https://tabicl.readthedocs.io/en/latest/tutorials/time_series_forecasting.html Meta Prophet: https://github.com/facebook/prophet GIFT-Eval Leaderboard: https://huggingface.co/spaces/Salesforce/GIFT-Eval TabArena Leaderboard: https://huggingface.co/spaces/TabArena/leaderboard  

The Neuron: AI Explained
OpenAI Astra, Local AI, and the Hardware Race

The Neuron: AI Explained

Play Episode Listen Later Sep 9, 2026 94:40


What happens when AI models get powerful enough that the bottleneck stops being the model, and starts becoming the computer, the power grid, or the safeguards around it?Corey Noles and Grant Harvey break down a packed week in AI, starting with OpenAI's expected Astra model and the controversy around recurrent-depth reasoning, chain-of-thought monitoring, and critical cybersecurity capabilities. They also dig into Google's Gemini 3.8 Flash and Flash Cyber, and why benchmark charts increasingly matter less than what a model can actually do in real work.Then the conversation moves from models to machines: a $399 open-source robot duck, Dyson's wildly overengineered CameraJet toothbrush, the return of the CPU as AI agents increasingly operate computers directly, and Apple's increasingly data-center-like Mac Studio hardware for local AI workloads.Finally, Corey and Grant unpack the fight over AI data centers in Texas, what communities should demand in exchange for hosting them, and why ChatGPT's advertising business may become a major new piece of OpenAI's economics.Subscribe to The Neuron at theneuron.ai for a daily briefing on the AI stories that actually matter.The Neuron: https://www.theneuron.ai/The Neuron Academy: https://www.neuronacademy.com/

SQL Server Radio
Episode 191 - High CPU in SQL Server 2025

SQL Server Radio

Play Episode Listen Later Sep 8, 2026 29:04


Guy and Eitan discuss an interesting case study where crazy high CPU was detected after an upgrade to SQL Server 2025. Relevant links: SQL 2025 showing crazy-high CPU | LobsterPot Solutions Enable Automatic Tuning - Azure SQL Database & Azure SQL Managed Instance | Microsoft Learn Plan Forcing in SQL Server - Erin Stellato

David Bombal
#602: How Compilers Turn Secure C Code Into Vulnerable Binaries

David Bombal

Play Episode Listen Later Sep 8, 2026 30:39


Big thanks to ‪@ThreatLocker‬ for sponsoring my trip to Black Hat USA 2026 and also for sponsoring this video. To start your free trial with ThreatLocker please use the following link: https://www.threatlocker.com/davidbombal You can write secure C code, follow accepted best practices and still end up with a vulnerable binary. The reason is simple: the CPU does not run your source code. It runs whatever the compiler produces. David sits down with security researcher Chris Domas at Black Hat to examine how legal compiler optimizations can remove security protections, delete memory-clearing operations and introduce time-of-check to time-of-use vulnerabilities into code that appeared secure. Chris explains the C abstract machine, why compilers are allowed to transform code so dramatically and how register pressure, structure layout and even data size can affect whether a binary is vulnerable. In one striking example, 17 or 33 bytes can be safe while nearby sizes produce vulnerable code. They also discuss whether Rust solves the problem, why switching between GCC and Clang is not the answer and how AI helped analyse 500 million lines of open-source code to identify 300 potentially dangerous patterns. Most importantly, Chris explains what developers can do now, including enabling compiler warnings, using sanitizers, analysing optimized builds and testing the exact binary that will be shipped. // Christopher Domas' SOCIAL // LinkedIn: / christopher-domas GitHub: https://github.com/xoreaxeaxeax X: https://x.com/xoreaxeaxeax // David's SOCIAL // Discord: discord.com/invite/usKSyzb Twitter: www.twitter.com/davidbombal Instagram: www.instagram.com/davidbombal LinkedIn: www.linkedin.com/in/davidbombal Facebook: www.facebook.com/davidbombal.co TikTok: tiktok.com/@davidbombal YouTube: / @davidbombal Spotify: open.spotify.com/show/3f6k6gE... SoundCloud: / davidbombal Apple Podcast: podcasts.apple.com/us/podcast... // MY STUFF // https://www.amazon.com/shop/davidbombal // SPONSORS // Interested in sponsoring my videos? Reach out to my team here: sponsors@davidbombal.com // MENU // 0:00 - Coming Up 0:48 - Intro 02:05 - Different Ways of Exploiting CPU's 04:10 - The C Specifications 06:17 - The Compiler Deleting Nemsec 08:40 - Do we need to use a new Compiler ? 10:09 - Compiler Inventing Vulnerabilities 12:13 - Don't Give up Writing Secure Code 12:44 - Sponsored Section 14:25 - Any Easy Options To Create A New Compiler ? 15:09 - Chris's Presentation at Black Hat 20:00 - Weird Situations with Size of Data 21:22 - What Can Developers Do ? 23:32 - Who Can Leverage this Vulnerability ? 25:02 - Could AI Make it Easy For Attackers To Leverage This? 28:27 - Recommendations For Developers 29:48 - Advice To Be Like Chris 30:36 - Conclusion & Outro Please note that links listed may be affiliate links and provide me with a small percentage/kickback should you use them to purchase any of the items listed or recommended. Thank you for supporting me and this channel! Disclaimer: This video is for educational purposes only. #bhusa2026 #securecoding #compiler

Mac Geek Gab (Enhanced AAC)
Maybe It's "S-Bot"

Mac Geek Gab (Enhanced AAC)

Play Episode Listen Later Sep 7, 2026 86:35 Transcription Available


Start Mac Geek Gab 1158 with Quick Tips you’ll actually use: trigger your iPhone’s camera from your Apple Watch, pull a full-quality photo out of any video in iOS 27, and find out why hybridsearchd is quietly eating your CPU in macOS Golden Gate. You’ll also get the right way to set up Homebrew on a brand-new Mac, meet mackup for keeping your settings in sync across every Mac you own, and watch Siri AI turn a business card into a contact without you typing a thing. Then your two favorite geeks and pilot dig into your questions: how much RAM is actually enough, how iOS call screening and spam filtering really behave, which outdoor cameras earn their keep, and how to reconnect to that remote NAS that keeps ghosting you. You’ll finally learn whether you can delete someone from a group text, get the early word on Sonos 27, and settle the great “how do we even say Siri AI?” debate in a lightning round that lands on S.A.I., S-Lady, Papaya, and — maybe — S-Bot. Press play, get the answers before your gear gets you, and whatever you do, Don’t Get Caught! 00:00:00 Mac Geek Gab 1158 for Monday, September 7th, 2026 – https://mgg.fm/1158 September 7th: National Beer Lovers Day MGG Monthly Giveaway – Win a license to SuperDuper! 4 The MGG Merch Store is Live! Quick Tips 00:00:01 Todd-QT-Use your Apple Watch Camera Remote 00:03:40 Pilot Pete-QT-Save Video Still as Full Quality Photo iOS 27 00:06:10 hybridsearchd in macOS Golden Gate using CPU 00:10:15 Pensacola Craig-Nuke & Dave: using Homebrew on a new Mac 00:11:40 mackup to copy your settings and keep them in sync between your Macs 00:12:40 Dom Bettinelli-Create a new contact from business card with Siri AI MGG Reviews at https://macgeekgab.com/review 00:15:07 iGuy1984-MGG Review-Fantastic 00:18:22 TaiPai59-MGG Review-Great show!! Sponsors 00:19:28 SPONSOR: Gusto. Get three months free when you run your first payroll when you start at https://gusto.com/MGG 00:20:43 SPONSOR: OneSkin. Born from over a decade of longevity research, OneSkin's OS-01 Peptide is proven to target the visible signs of aging, helping you unlock your healthiest skin now and as you age. Get 15% off OneSkin with the code MGG at https://www.oneskin.co/MGG  #oneskinpod #ad Your Questions Answered – feedback@macgeekgab.com 00:22:00 Dan-How much RAM is enough? 00:38:10 Mark-Call Screening options and behavior in iOS Settings > Apps > Phone > Call Filtering > Spam 00:45:50 Mike-Which outdoor cameras do you use? Eufy Security Cameras Thermacell LIV 00:56:29 Peer Timo-How can I re-connect to my remote NAS? 01:04:15 Ericka-Can you delete someone from a group text? 01:11:57 Sonos 27mcp in Early Access How to Say Siri AI – The Lightning Round 01:16:33 Bill-1157-Maybe it’s just SAI 01:17:25 Dan-I think S-Lady is the right answer! 01:19:00 Simon-1157–Maybe Just Notch out The Siri Frequencies? 01:20:26 Ben-1157-Papaya! 01:21:09 Adam-1157-Another way to say Siri 01:22:47 Shree AI? …or maybe it’s S-Bot! 01:23:44 MGG 1158 Outtro MGG Monthly Giveaway Bandwidth Provided by CacheFly Pilot Pete's Aviation Podcast: So There I Was (for Aviation Enthusiasts) The Debut Film Podcast – Adam's new podcast! Dave's Business Brain (for Entrepreneurs) and Gig Gab (for Working Musicians) Podcasts MGG Merch is Available! Mac Geek Gab iOS app Mac Geek Gab YouTube Page Mac Geek Gab Live Calendar This Week's MGG Premium Contributors MGG Apple Podcasts Reviews feedback@macgeekgab.com 224-888-GEEK Active MGG Sponsors and Coupon Codes List BackBeat Media Podcast Network

The Six Five with Patrick Moorhead and Daniel Newman
NVIDIA's $12.93B Hugging Face Deal, Anthropic's Compute Buildout, and a Blowout AI Earnings Week

The Six Five with Patrick Moorhead and Daniel Newman

Play Episode Listen Later Sep 7, 2026 60:52


NVIDIA closes its $12.93 billion acquisition of Hugging Face, Anthropic signs its third $30 billion-plus compute deal in three weeks, and Signal65 launches its new PINNACLE agentic AI benchmark just in time to catch Claude Fable 5.1 topping the charts, all while Dell, HPE, Broadcom, and Snowflake post a blowout week of AI-driven earnings. Patrick Moorhead and Daniel Newman have all the details on Ep. 318 of The Six Five Pod. The handpicked topics for this week are: NVIDIA's $12.93 Billion Acquisition of Hugging Face: NVIDIA closed its purchase of Hugging Face this week after days of rumors, structured as a full acquisition rather than a licensing deal like NVIDIA's earlier Grok arrangement. Moorhead frames the deal as NVIDIA's bid to own the entire developer pipeline, with GitHub as the first stop and Hugging Face as the second, and the platform's Spaces service giving NVIDIA a way to run workloads once developers arrive. (The Decode) Anthropic's $35 Billion Compute Deal With Lambda: Anthropic signed its third $30 billion-plus compute commitment in three weeks, this time with NVIDIA-backed Lambda managing infrastructure inside a data center owned by crypto infrastructure company Hut 8. The deal covers 350 megawatts over six years, and Newman ties it directly to Anthropic's revenue run rate climbing from about $10 billion to $65 billion since the start of 2025. (The Decode) A 72-Hour Wave of Frontier Model Launches: Claude Fable 5.1, OpenAI's GPT-6 Astra, Google's Gemini 3.8 Flash, Meta's Muse Spark, and six new models from the UAE's MBZUAI all shipped within days of each other. Moorhead and Newman frame the pace as evidence that model rankings now shift day to day, a much faster cycle than the weeks-long stretches leaderboards used to hold. (The Decode) John Ternus Takes Over as Apple CEO: Moorhead calls it a continuity pick that promotes the engineer who built Apple's hardware moat. Newman notes Tim Cook's buyback record, $867 billion in total repurchases, exceeded the combined total of the rest of the Magnificent Seven during his tenure. (The Decode) Signal65 Launches the PINNACLE Agentic AI Benchmark: Signal65 President Ryan Shrout joined Pat and Dan to detail PINNACLE, a new benchmark built to measure real, enterprise-relevant agentic work across model intelligence, infrastructure performance, and full-stack cost per correct task. Shrout says the benchmark already caught Fable 5.1 jumping ahead of Opus 5 and GPT-5.6 Sol on intelligence within days of the model's release. That same result also came in as the most expensive model per correct answer. (Off The Record) Dell Technologies (DELL): Dell delivered $47 billion in revenue, up 58%, with $16.4 billion in AI servers and a $95 billion AI backlog. Newman calls out strong storage and CPU server performance, along with the backlog. He flags financing risk tied to Dell's Neo Cloud customers as the one weak spot in an otherwise dominant quarter. (Bulls and Bears) Hewlett Packard Enterprise (HPE): HPE beat on revenue and non-GAAP EPS with a double beat on guidance, and Moorhead points to networking orders compounding faster than revenue can be recognized as the standout signal. The stock sold off anyway on questions about margin mix and supply. CEO Antonio Neri's focus on selective, less price-sensitive deals is already showing up in a repaired balance sheet. (Bulls and Bears) Broadcom (AVGO): Broadcom tripled its AI business to $16.7 billion, up 221%, and raised its fiscal 2028 AI revenue outlook to $230 billion. Newman says the market wanted the number closer to $300 billion and sold off on the guide. Broadcom still posted record $29.6 billion total revenue and continued strength across networking and storage controllers. (Bulls and Bears) Snowflake (SNOW): Snowflake beat on revenue, non-GAAP EPS, and forward guidance, with AI products across Cortex and its new coding agent driving roughly half the beat. Sell-side price targets moved sharply higher across more than a dozen firms, and Newman ties the 111 percent stock gain since the software sell-off to the same pattern he's tracked through the DeepSeek moment and the "software is dead" narrative: markets consistently overreact to those stories before reversing. (Bulls and Bears) Thanks for tuning in to the pod. Hit that subscribe button, and check out the new Signal65 PINNACLE benchmark at pinnacle.signal65.com. The Decode NVIDIA's $12.93 Billion Acquisition of Hugging Face https://techcrunch.com/2026/09/03/nvidia-confirms-it-will-buy-hugging-face-for-12-9-billion/ Anthropic's $35 Billion Compute Deal With Lambda https://www.reuters.com/technology/anthropic-signs-35-billion-cloud-deal-with-nvidia-backed-lambda-source-says-2026-08-31/ A 72-Hour Wave of Frontier Model Launches https://www.anthropic.com/claude-fable-and-mythos-5-1 https://openai.com/index/path-to-astra/ John Ternus Takes Over as Apple CEO https://www.apple.com/newsroom/2026/04/tim-cook-to-become-apple-executive-chairman-john-ternus-to-become-apple-ceo/ Off The Record Signal65 Launches the Pinnacle Agentic AI Benchmark https://pinnacle.signal65.com/ Bulls and Bears Dell Technologies (DELL) https://www.barrons.com/articles/dell-stock-soars-q2-earnings-guidance-beat-d07d0ee7 Hewlett Packard Enterprise (HPE) https://www.hpe.com/us/en/newsroom/pressrelease/2026/09/hpe-reports-fiscal-2026-third-quarter-results.html Broadcom (AVGO) https://investors.broadcom.com/news-releases/news-release-details/broadcom-inc-announces-third-quarter-fiscal-year-2026-financial Snowflake (SNOW) https://investors.snowflake.com/news/news-details/2026/Snowflake-Reports-Financial-Results-for-the-Second-Quarter-of-Fiscal-2027/default.aspx  

The React Native Show Podcast
Browser AI, WebGPU, and the Road to On-Device Agents With Nico Martin | React Universe On Air

The React Native Show Podcast

Play Episode Listen Later Sep 7, 2026 55:33


Many AI features in web apps send the user's audio, images, or text to a server and wait for the result. That adds network latency, moves data off the device, and ties the experience to a connection. Transformers.js gives developers another option: run the model in the browser. In this episode, Mike Grabowski speaks with Nico Martin, Open Source ML Engineer at Hugging Face, about what browser AI can do today. They break down the Transformers.js pipeline API, ONNX Runtime, WebGPU acceleration, model downloads, browser caching, CPU fallbacks, and performance across devices. Nico explains when local inference can beat a server round trip, why model size and hardware variation shape the user experience, and how browser code can handle parts of an agent workflow. He also shares early work on structured output and a custom WebGPU inference engine, where current experiments point to 5x to 10x speedups. For developers deciding which AI tasks belong in the browser and which need the cloud, this episode maps the tradeoffs through speech recognition, background removal, embeddings, local tool calling, and on-device agents. Check out episode resources on our website ➡️ https://clstk.com/4zMwzq9 Catch more React Universe On Air episodes

Geek News Central
The Mainframe Learned to Speak Arm #1875

Geek News Central

Play Episode Listen Later Sep 5, 2026 44:51 Transcription Available


In this episode, Ray Cochrane breaks down Hot Chips 2026, the engineering conference where IBM, NVIDIA, Intel, AMD, Arm, and Fujitsu all showed how their next processors actually work. The headline disclosure is a mainframe core that runs Arm natively. Ray also covers Apple’s odd M6 Mac mini naming, London’s first autonomous Uber rides, Amazon’s purchase of the company behind DuckDB, GitHub’s HydraFusion, the best of IFA 2026, and new USDA research on farmed salmon. – Want to start a podcast? It’s easy to get started! Sign-up at Blubrry – Thinking of buying a Starlink? Use my link to support the show. Subscribe to the Newsletter. Email Ray if you want to get in touch! Like and Follow Geek News Central’s Facebook Page. Support my Show Sponsor: Best Godaddy Promo Codes Get 1Password Full Summary Cochrane opens the show with a personal update. He apologizes for the late rollout and the missed Monday episode, having spent the week fighting a cold. He’s also heading to Michigan to spend time with family and visit his father’s gravesite. He hopes to record a couple of shows from his dad’s old studio while he is there, including a special episode planned for Tuesday. He also points listeners to a fresh site redesign that trades the old techie look for something cleaner and friendlier. The featured segment starts from a wrap-up post on Arm’s newsroom. However, Cochrane broadens it to cover the whole conference rather than a single article. Hot Chips has run every August since 1989, and this year’s event was the 38th, held August 23rd through the 25th at Stanford’s Memorial Auditorium. What Hot Chips Actually Is Cochrane draws a line between Hot Chips and the big consumer trade shows. CES and Computex exist for product announcements and marketing. Meanwhile, Hot Chips is an IEEE engineering conference where chip architects present block diagrams and die photos for thirty minutes at a stretch. The audience matters as much as the content. Roughly five hundred people who design chips for a living fill the room, and they would spot a fudged number immediately. No written paper is required, just the talk and the slides. In-person tickets sold out this year, as did Stanford’s dorm housing. Cochrane says he plans to cover the conference annually going forward. IBM Built a Mainframe Core That Speaks Arm The disclosure that stopped Cochrane cold came from IBM. Its next processor for IBM Z and LinuxONE runs two completely different instruction sets natively, on every one of its eleven cores. Those are z/Architecture, IBM’s own mainframe language, and AArch64, which is 64-bit Arm. Crucially, this is not emulation. Nor is it Arm cores glued onto the die beside the mainframe cores. IBM built 2,792 Arm instructions directly into the hardware, which it says is more than double the mainframe instruction count. Each core carries two separate decoders while sharing the caches, branch predictor and register files downstream. It switches between the two in nanoseconds. IBM even added dedicated hardware to flip byte order, because Arm and the mainframe store numbers in opposite directions. The payoff is Arm SystemReady compliance, meaning off-the-shelf Arm Linux runs on a mainframe unmodified. Patrick Kennedy of ServeTheHome, who was in the room, wrote: “I am sitting here still in awe of what IBM is doing here; this is not Z plus Arm cores, this is Z and Arm in one core.” Cochrane flags one precision point that is easy to get backward. IBM did not license Arm’s core designs and drop them in. Instead, it took its own mainframe core and taught it AArch64 under an architecture license, which is considerably harder engineering. The specifications are striking. The chip uses a 2nm process, with eleven cores running above 5.7GHz sustained and no turbo mode at all. Each core gets 36MB of L2 cache, backed by a 3.5GB virtual L4 pool. Furthermore, the reliability target is eight nines, which works out to roughly three tenths of one second of unplanned downtime per year. IBM gave it no name and no ship date, though the press expects “Telum III” around 2028. Arm, Fujitsu and NVIDIA Show Their Hands Arm itself had plenty to discuss, starting with an unfortunate name. Its first chip in thirty-five years is called the AGI CPU, which is a product name rather than any claim about artificial general intelligence. For three and a half decades, Arm designed processor blueprints and licensed them out, collecting royalties without competing. That era is now over. The AGI CPU is Arm’s own silicon, co-designed with lead customer Meta, running up to 136 cores on TSMC’s 3nm process at 300 watts. Arm’s CEO says the company has more than $2 billion in customer demand across the next two fiscal years. Fujitsu brought the detail Cochrane called the coolest of the conference. Its MONAKA chip packs 144 Arm-based cores, but the trick is the cache. Rather than sitting alongside the cores and eating die area, the entire last-level cache lives on a separate 5nm die with the 2nm compute die stacked directly on top. It ships in 2027 in 350-watt and 500-watt versions. NVIDIA had more stage time than anyone, with six sessions. Its new Vera CPU carries 88 cores of NVIDIA’s own Olympus design, which marks a change: the previous Grace CPU used Arm’s off-the-shelf cores. Consequently, Arm’s win here is the instruction set, not the blueprint. The memory disclosure drew the most attention, with a fully loaded system reaching 1.5TB at 1.2TB/s while the whole memory subsystem draws just 30 to 40 watts. The Caveat on NVIDIA’s Benchmark Slides Cochrane pushes back on how NVIDIA presented its numbers. On the standard SPEC integer benchmark, Vera scored 925 against AMD’s 128-core EPYC score of 898, about three percent ahead. However, the slide NVIDIA showed normalizes that same result per physical core, which makes a three percent gap look enormous. NVIDIA defends the choice, arguing that per-core throughput matters when thousands of AI agents run at once. Cochrane grants that it is a fair argument to make. Even so, his verdict is blunt: it is a different number from the headline one, and presenting it that way is not the best look. Sponsor: GoDaddy Economy hosting $6.99/month, WordPress hosting $12.99/month, domains $11.99. Website builder trial available. Use codes at geeknewscentral.com/godaddy to support the show. Apple’s Newest Chip Landed in Its Cheapest Mac Last week’s episode covered the Mac Studio half of Apple’s August announcement. Tonight Cochrane takes the other half, the new Mac mini, and finds the numbering genuinely strange. The $899 base Mac mini gets the M6, which is Apple’s first 2nm chip and the newest process the company has shipped. It brings twelve CPU cores, twelve GPU cores, and 170GB/s of memory bandwidth. Apple also introduced a third CPU core class called super cores. So Apple’s most advanced chip sits in its cheapest desktop, while the M5 Pro, M5 Max and M5 Ultra above it all carry a lower number. It gets stranger. The M6 mini has Thunderbolt 4 while the pricier M5 Pro mini has Thunderbolt 5, and the memory ceiling runs backward too. Apple explains none of it across three announcement pages. Reading between the lines, Cochrane figures the M6 is the entry point of a new generation that shipped ahead of its larger siblings. London’s Robotaxis Started Carrying Passengers Arm’s monthly roundup covers everything outside the data center, and Wayve stood out. Transport for London granted the British self-driving company private hire vehicle licenses on August 5th, the same category a minicab needs. Then on September 3rd the service launched with Uber, marking the first autonomous rides ever offered to UK passengers. A small fleet of Ford Mustang Mach-Es covers anywhere in London except the airports, with a TfL-licensed safety driver still aboard. Over 140,000 Londoners signed up. The compute runs on NVIDIA’s Arm-based automotive platform, which is why Arm claims the win. Cochrane notes the pattern: once you set a standard nobody can move off, these wins keep arriving. Two more items round it out, both about squeezing AI onto phones. Google’s Pixel 11 shipped in August with the Tensor G6, and Google claims on-device AI runs up to 3.5x faster using 3.5x less energy. Those are Google’s own unbenchmarked figures. Separately, Graphcore’s research team worked with Arm to run an 11-billion-parameter vision model on phone-class processors by squeezing each parameter to 2.7 bits, taking the model from roughly 22GB down to 3.7GB. Intel Is Pitching AI Infrastructure From a Long Way Back Intel’s newsroom post previews the AI Infra Summit, running September 15th through 17th in Santa Clara. CEO Lip-Bu Tan takes a fireside chat on Tuesday morning, with three Intel sessions across the show. Cochrane unpacks two terms first. Physical AI means AI that acts in the real world through sensors and motors rather than living on a screen, and Intel takes it seriously enough to have renamed its PC division the Client Computing and Physical AI Group in May. Disaggregated inference splits the two phases of running a model: reading your prompt is compute-hungry, while writing the answer back is memory-hungry. The context is where this gets interesting. Intel’s revenue rose 25 percent last quarter, its fastest growth since 2011. Nevertheless, in AI accelerators the company barely registers next to NVIDIA. Gaudi has effectively been abandoned, to the point that Intel stopped maintaining its open-source driver, and AMD passed Intel in data center revenue last quarter. Tellingly, when Intel demoed disaggregated inference at Computex, NVIDIA GPUs handled one phase and SambaNova chips the other while Intel supplied the coordinating CPU. Crescent Island, Intel’s actual inference chip, does not sample until later this year, and Intel declined to publish its memory bandwidth. Amazon Bought the Company Behind DuckDB On August 26th, Amazon signed a deal to acquire DuckLabs, the Amsterdam company behind DuckDB. Cochrane spends time explaining what DuckDB is, since listeners outside the data world may never have encountered it. The problem it solves is familiar. Querying a large pile of data files traditionally meant either running a database server or spinning up a data warehouse with a cluster, a bill, and a loading pipeline. Both are heavy machinery for a question you wanted answered in ten seconds. DuckDB instead ships as a library rather than a server. You add it to your program like any other package, point it at your files, and write ordinary SQL directly against them. The data never moves. The usual comparison is SQLite, which is embedded in nearly every phone and browser on earth. Where SQLite excels at looking up one record, DuckDB rebuilds that embedded idea for chewing through millions of rows. It now sees roughly 62 million monthly downloads on Python’s package index alone, up from about 25 million last October. AWS says it is buying the company, not the project. DuckDB stays free and open source under the MIT license, held by a Dutch nonprofit foundation, and the founders join AWS while continuing to run technical direction from Amsterdam. Andy Warfield, a VP and distinguished engineer at AWS, described DuckDB as “the glibc of structured data: a lean, unglamorous, ubiquitous dependency that a great deal of software links against and almost nobody has to think about.” Cochrane sits with what that ownership means. He reaches for an analogy: imagine Daniel Stenberg selling curl. He doubts it would ever happen, but the concept alone is startling given how much infrastructure depends on it. His read is that AWS is betting DuckDB becomes as foundational as curl and SQLite already are. GitHub Has One AI Model Grade Another One’s Homework GitHub shipped Project HydraFusion into Copilot as a research preview. Instead of routing your request to a single model, it picks one of three approaches per request. Sometimes one model simply answers. Alternatively, a cheaper model drafts, and a quality gate decides whether to escalate. The interesting one is Critique. One model writes the code, a separate model from a different family reviews it read-only, and the original gets one pass to revise. Despite the name, nothing is fused here. There is no voting and no merging, just one model at a time with a gate deciding whether to spend more. GitHub explained the reasoning in an earlier post: “a model reviewing its own work is still bounded by its own training biases: the same training data and techniques, the same blind spots.” Research supports it. A team at NeurIPS in 2024 showed that models recognize their own writing and score it higher than human graders do. GitHub’s earlier number had a Claude Sonnet and GPT critic pairing closing about three-quarters of the gap between Sonnet and the larger Opus model. This mirrors a workflow Cochrane uses constantly and has described on a previous episode. He runs a cross-check review with a second model from a different company, and it routinely surfaces issues the first model missed. He explains that different training data, different engineers, and different reinforcement approaches build different internal biases about what counts as correct. Looking ahead, he expects more of these “Frankenstein patterns” where models from different training families work together. The Best of IFA 2026, and What You Can Actually Buy IFA opened to the public in Berlin for its 102nd year, with about 1,900 brands. Cochrane splits The Verge’s roundup in two, since much of what generates headlines at these shows never ships. Starting with real products, iRobot’s flagship Roomba Max 875 Combo runs $1,199 and ships in about two weeks. Its SealForce feature drops a hidden skirt from the chassis when it detects carpet, sealing against the fibers so suction concentrates instead of leaking out the sides. That reaches 35,000 pascals, iRobot’s strongest yet. A step-down model at $899 carries the same trick, though Cochrane balks at both prices. Anker’s Soundcore Sleep 4 Pro earbuds arrive in November at $349.99. The charging case carries its own round touchscreen, so you pick soundscapes, set alarms, and read sleep stats without your phone. Optical sensors read heart rate and variability from the ear canal, which beats the wrist for accuracy, and the case masks a snoring partner. Cochrane remains unconvinced about sleeping with earbuds in. Philips also has smart rope lights, the Hue Liane 360, on sale now. They glow evenly around the tube rather than showing individual LEDs. They also run $400 for three meters, which works out to about $130 per meter of rope light. As for concepts nobody can buy, iRobot showed a robot vacuum that carries a smaller robot vacuum on its back in a garage and lowers it to deploy. Lenovo brought a 14-inch laptop whose screen rolls out to 17 inches at the press of a button, which reviewers call the first rollable that feels close to shippable. Tecno showed a phone with essentially no border around the screen, and Acer had a Windows gaming handheld that swivels its screen up over a keyboard. Two themes ran through the show. Humanoid robots were the loudest thing on the floor, and IFA’s own CEO framed the event as being about robots that work rather than robots that demo. Meanwhile, AI stopped being its own product category and became an ingredient, showing up in refrigerators, treadmills, dishwashers, and motorized TV mounts. Farmed Salmon Isn’t the Omega-3 Machine It Used to Be USDA scientists measured farmed salmon and found considerably less of the good fat than the government’s own database claims. EPA and DHA are two fatty acids you get almost entirely from fish. Your body can build them from the plant version, but only in tiny amounts, so the NIH’s position is that eating them is the only practical way to raise your levels. Those fatty acids are structural pieces of every cell, with DHA concentrating in the brain and retina. That is why the federal dietary guidelines, issued jointly by USDA and Health and Human Services, recommend at least eight ounces of fish a week and steer you toward salmon. That amount is calibrated to deliver about 250 milligrams a day. Published in Frontiers in Nutrition last month, the study found EPA and DHA in farmed Atlantic salmon came in 54.7 percent lower than USDA’s own reference values, last updated in 2018. A three-ounce serving fell from roughly 1,670 milligrams to about 756. Consequently, two servings a week now fall about 14 percent short of the target. Plant-derived fats meanwhile rose two to three times over. The likely cause is feed. Salmon are carnivores, and farms once fed them oily little fish. There was never going to be enough of those as the industry scaled, so crops filled the gap: soy, canola, sunflower and linseed. Importantly, the study does not claim to have proven this and calls the feed shift a plausible explanation. Independent corroboration lends it credibility. Researchers at Stirling measured a similar halving in Scottish salmon between 2006 and 2015, and Norway’s marine institute saw it across thousands of samples. There is a land dimension too. Roughly half the world’s soy grows in South America, where rainforest gets cleared for feed. Matthew Hayek, who studies the environmental cost of protein at NYU, told Inside Climate News that “soy is a major, important protein and oil ingredient in fish farming.” Adding up two decades of soy across all fish farming, he puts the extra forest clearing at around the area of Nicaragua or Bangladesh. That figure covers all fish farming rather than salmon alone, and Hayek notes it is hard to attribute soy use to any single species. Cochrane closes with two caveats. First, the study measured only eight fish, bought around Maryland, DC and Virginia over six weeks in 2023, and nearly all sourced from Chile. That is not a national survey, and the authors say plainly the sample was not large enough to change government advice. Rather, it flags that a federal database value needs rechecking, which is what the paper set out to do. Second, on whether you should care, farmed salmon still beats beef, chicken and eggs by a mile, since those carry essentially zero EPA and DHA. What it loses is its crown among fatty fish, dropping to mid-pack behind herring, sardines and mackerel and roughly level with trout. The broader health case is also softer than the 2000s suggested. A review of 86 trials covering 162,000 people found supplements barely moved heart attacks or deaths, so eating fish and swallowing fish oil are not the same claim. No producer has responded to the findings, and USDA, whose own scientists ran the study, declined an interview and did not answer emailed questions. Cochrane wraps up with housekeeping and a note that he will be back on Labor Day. The post The Mainframe Learned to Speak Arm #1875 appeared first on Geek News Central.

IOSYS / haitenai.com
NLP ぬるぽ放送局 第1095回 本名 miko田miko三郎 #nurupo

IOSYS / haitenai.com

Play Episode Listen Later Sep 4, 2026 82:08


ぬるぽ放送局おたより投稿フォーム https://forms.gle/6tbmBzK6wbyavJG47 2026年9月パワープレイ 臨界モスキー党 - どすこいファンク (千秋楽バージョン) 2026・8・26 Digital Release https://notebookrecords.net/discographyportal.php?cdno=NBDL-1006 力士は黙ってちゃんこに廻しに懸賞金! 2013年リリースの3rdアルバム『青空家族』にて初入幕を果たした原曲『おすもうさん!』、その後は鳴かず飛ばずで遂に廃業を決意…と思いきや!2024年に装い新たにファンキーな化粧廻しを着けて某公募に挑戦するも…見事に落選!今度こそ廃業か!?…と思いきや!フル尺でサブスク界の横綱を目指してしれ〜っと三度目の土俵入りでごっつぁんです! 番組時間:82分8秒 出演者:夕野ヨシミ、たくや VOICEVOX:ずんだもん VOICEVOX:四国めたん ---- 2026/9/3に公開録音したものを配信いたします。 ラジオ記事はリスナーのEEチャンピオンさんが書いてくれているので楽してます。 <オープニング> ・よいお年を! ・まだ稲刈ってないか ・来週は9年ぶりの広島です ・この100円引きは何でしょうか? ・10年前はバニーガールのお店に行ってなかったよ ・活動くんのイオシスコーナーです ・steamの売上1位のゲームをやるのはめずらしい ・ちょっとえっちなゲームは再生数が伸びるんだよな ・じゃあ、前しっぽのおかげか ・毎週水曜深夜24時放送 InterFM897『TRIBALCON. presents NIGHT HIKE Radio』にて  「D.watt」がスペシャルマンスリーDJを担当します! ・radiko等でご視聴ください ・提供楽曲にMVがつきました!  「クラブこわい/ななひら」  Music: D.watt(IOSYS)  Lyrics: 七条レタス(IOSYS) ・YouTube 400万再生!  【東方MV】行列のできるえーりん診療所【IOSYS】 ・YouTube 300万再生!  【東方MV】お空のニュークリアフュージョン道場【IOSYS】 ・食べ美ちゃん今年も出てましたね ・イオシスは何も関わっていません ・本名 miko田miko三郎 ・画数多美 ・この放送はじめて聞きましたね ・車は早く動かしてくださいね ・今すぐ降り美 ・謎の台風は沖縄にいます ・台湾の球場で楽曲が流れるよ ・イベントはいっぱいありまーす ・最近の夕野ヨシミは暇なのです ・ゆっくりしていってねだった ・暇な2人でお送りします <Aパート> ・ふつおたです ・地球大変だっぴ ・歳はほどよく取りました ・治水は大事ってKOEIの三国志で学びましたからね ・リーチの時は脱がなくていいんですか? ・靴下をリー棒代わりに ・スーパー稲刈りシーズン ・ちゃんとした農家はちゃんとかまえてる ・収入保険とか概算金とか ・農業関係者が集まるぬるぽ放送局 ・実家で農業を ・ミントとシソの二毛作 ・チェンソーマン レゼ編 ・オデュッセイアは見に行きます ・AIでエロチャットのつづき ・突然出てくる田中さん ・AIってなんなんでしょうか? ・連休が潰れました ・最近使ってないチャットgpt ・おじさん心配になっちゃうよー ・手で前しっぽを動かしているのかい? <Bパート> ・これ、誰が作ってるんですか? ・どすこいアニマルプレゼント ・MV付けますか? ・みつをたです ・タイパ至上主義麻雀 ・麻雀を覚えようとする海外ニキ ・ドンドン脱げちゃうね ・愚民みつを ・オーロラソースはうまい ・オロールソースだね ・芸人ピックアップニュース ・カタカナのヨネスケに改名 ・クラブ以外全部怖い ・創作昔話 ごんぎつね ・5時に起こして みつを ・この名前怒られる ・助かるなー ・チャットが審議待ちだらけ ・にじさんじピックアップニュース ・すけべ怪談 ・7時間もやってたのか ・また瞬間停電 ・イースⅠ.Ⅱはいつでもいけます ・どんどんデビューしてるにじさんじさん ・るんちょまさん復帰 ・ハマコーみたいに言いました? ・ホロピックアップニュース ・あんなに早く負けることあるんだ ・法定速度が30kmに ・Vピックアップニュース ・しぐれういがアタック25の問題に ・ホロ甲子園ニュース ・ラミ谷ダメだったか ・あのゲーム2人は効率悪すぎな <エンディング> ・ロストテクノロジーでしたねー ・1999年のデジカメの写真の話をファンボックスでやってます ・CPUが早すぎる ・電子ポローン ・バニーガーデンでもいいですよ ・みなさまからのお便りお待ちしてます

Atareao con Linux
ATA 828 De Docker a tu Cerebro Digital, el roadmap de IA para Linuxeros

Atareao con Linux

Play Episode Listen Later Sep 3, 2026 31:36


Este episodio 828 es la carta de presentación de la Temporada 9 de atareao con Linux. Treinta y cuatro episodios ya guionizados, siete etapas, y un objetivo claro: construir tu cerebro digital sobre Linux con herramientas locales, sin depender de nubes ni suscripciones.Pero antes de mirar adelante, toca hacer balance. La T08 empezó prometiendo Docker, selfhosting y Android, y sí, hablé de todo eso. Pero en abril de 2025 la IA local irrumpió con fuerza y la temporada viró hacia Ollama, modelos locales, RAG, MCP. Fue un giro desordenado, lo reconozco. Pero también fue el germen de todo lo que viene ahora.De eso va esta T09: de poner orden al caos. Siete etapas, de menos a más, para que sigas el hilo hagas el nivel que hagas.Etapa 1 — Recursos básicos: los cimientos de tu laboratorio de IA. Skills para tu agente, herramientas de publicación, el botiquín del explorador.Etapa 2 — Skills y MCPs: el pegamento. El Model Context Protocol ha madurado hasta ser un estándar abierto — lo soportan Claude, ChatGPT, VS Code, Cursor. Ya no es un experimento, es el USB-C de la IA. Y de paso, herramientas del ecosistema atareao como watchbeat (monitor de uptime en Rust) y alloy (dashboard Docker con OIDC).Etapa 3 — GraphRAG, el gran hito: de RAG vectorial a grafos de conocimiento. Mientras el RAG clásico devuelve fragmentos sueltos y tú unes los puntos, GraphRAG construye un grafo con entidades y relaciones. Preguntas como "qué contenedores están detrás de Traefik" pasan a ser una consulta directa a tu mapa de conocimiento. Usaremos LightRAG, que con 39.000 estrellas ya superó al Microsoft GraphRAG original. Esto ocupa tres episodios.Etapa 4 — Multimedia: Whisper para speech-to-text, TTS local, ffmpeg, visión artificial, y el pipeline de YouTube a conocimiento con yt-dlp. Rematamos con RAG multimodal.Etapa 5 — Orquestación: systemd timers, asyncio, just, y CrewAI para montar equipos de agentes.Etapa 6 — Proyecto final: dos episodios para construir El Asistente que te Conoce y ponerlo en producción con Quadlets.Etapa 7 — El futuro: mantenimiento de tu cerebro digital y hacia dónde va todo esto.Entre medias, herramientas Linux: shuul, sqlite-utils, yq + jq, Rust en el kernel, Wayland vs X11, la guerra de los filesystems.No necesitas una GPU de 3000 euros ni un doctorado. Con 16 GB de RAM y un CPU decente ejecutas modelos de 7B a 14B. Esto es IA local, en tu máquina, con tus datos.Capítulos del episodio:00:00 — Introducción y bienvenida a la Temporada 901:47 — Balance T08: de Docker y Selfhosting al boom de la IA04:37 — El momento adecuado para cada tecnología06:46 — El gran objetivo: tu cerebro digital08:44 — Roadmap T09: 30 episodios ya guionizados11:28 — Skills y MCPs imprescindibles12:43 — GraphRAG: de RAG a grafos de conocimiento14:01 — RAG vs GraphRAG: el mapa de tu conocimiento17:28 — Herramientas del ecosistema: alloy, populater, watchbeat20:08 — ¿Para quién es esto? De veteranos a escépticos22:08 — No es hype: es un cambio de paradigma24:30 — El momento perfecto para el linuxeroToda la info y el roadmap completo en atareao.es/828.Más información y enlaces en las notas del episodio

Hacker Public Radio
HPR4718: Programmable Logic Controls - Episode 2

Hacker Public Radio

Play Episode Listen Later Sep 2, 2026


This show has been flagged as Clean by the host. -------------------- 01 Introduction This is the second episode in an 8 part series. 02 In the previous episode we discussed the predecessors of PLCs, in particular relay logic. In this episode we will discuss how relay logic came to be expressed in software rather than in actual hardware components. 03 The topics to be covered include * Early computers in industry. * The first PLC. * PC versus PLC - what's in a name. * Who the major brands are. * What does a PLC actually look like. * Machine architecture. * PLC programs. * The scan. * PLC programming languages. * Relative popularity of PLC programming languages versus more conventional computer programming languages. -------------------- 04 Early Computers in Industry Computers came to be used in industry fairly early on. Minicomputers were used in some large industries to control things like electric power plants. For example, the DEC PDP8 was used to control the refuelling machines in CANDU nuclear power plants. These would load and unload fuel from the reactor, which happens continually on a daily basis while the reactor is running. 05 The first microprocessor based computer sold on the commercial market was the MICRAL from France, based on the Intel 8008. It was sold as a cost effective replacement for minicomputers used in industry. It preceded what are generally considered to be the first Personal Computers. So you can see that industry were not reluctant to adopt new technology. 06 However, mini computers were large complex systems that did not fit well into a factory floor. They had particular niches in very large complex integrated systems, but were not suited to controlling many individual machines in a factory that produced things like automobiles or appliances. 07 What was needed was something that would fit into a standard electrical enclosure, withstand the temperatures found in a factory, could stand up to vibration and noise, was tolerant of voltage fluctuations, interfaced directly to sensors and actuators, and could be readily programmed by engineers, technicians, and tradesmen who were familiar with the processes to be controlled, but had little or no experience with computers. 08 This required a complete integrated package covering hardware, software, and product distribution through industrial supply retailers. Something that could meet these criteria is what was needed to become the PLC. -------------------- 09 History Many histories point to a US developer that became Modicon. However, their first hardware was more a proof of concept than a viable commercial product. In fact multiple companies in the US, Europe, and Japan were working on the problem and released comparable products all at around the same time in the early 1970s. 10 The idea preceded the implementation by a number of years. For example, General Motors had asked industrial control equipment suppliers in the mid 1960s for some sort of programmable control device to replace relay logic. This showed potential suppliers that there was a market for this sort of thing, and gave them a clearer idea of what customers were looking for. 11 What made this possible was the development of the first 8 bit microprocessor in that time period, combined with readily available bit slice processors from the minicomputer industry and other components. 12 People knew what to do, the problem was waiting for the development of suitable component hardware. Minicomputers were already being used in industry, but were normally located in control rooms. PLCs were an effort to take that technology out of control rooms and put it on the shop floor. 13 PC Versus PLC - What's in a Name In the early days, these systems were known as either PLC, which means "Programmable Logic Controller", or PC, which means "Programmable Controller". Different vendors favoured different terminology, but both were common. I have a book which was at one time one of the standard reference handbooks for this industry which was published in 1989, and refers to them as Programmable Controllers, or PCs. 14 The term PC was also used to mean "Personal Computer", but that wasn't really a problem in the early days. However, a certain large company decided to call their entry into the personal computer market the "IBM PC", and suddenly "PC" became a generic term for desktop computers. 15 After a long struggle to keep referring to their products as "PCs", even the biggest vendors caved in and gave up the fight and switched to using the "PLC" term. I will therefore use the term "PLC" in this podcast series even when talking about products which were originally called "PCs" at the time they were introduced. 16 Who the Major Brands Are The companies that came to dominate the industry fairly early on were generally companies that already made industrial electrical hardware. These were Siemens Allen Bradley, later known as Rockwell Schneider Mitsubishi Omron 17 These are still the dominant companies in the business, although there are many small brands, particularly at the cheaper end of the market. All of them were suppliers of a wide range of industrial control hardware, such that you could build all or nearly all of the parts of your control system using only their products. 18 Each of these sells an integrated hardware and software package, including development software, that is completely proprietary. If you thought that the mainframe business had a lot of vendor lock-in, you haven't seen the industrial market. -------------------- 19 What Does One of These Things Actually Look Like? At this point you are probably still very confused as to what a PLC actually is. I will attempt to describe one in basic terms by giving an example of one. The smallest and simplest ones are what are termed "shoebox" PLCs. These are basically a rectangular box with a plastic case. 20 Here are the dimensions for a typical low end model, a Mitsubishi FX3U-32M. According to specs in the manual, it is 150mm wide, 90mm high, and 86mm deep. 21 It can be mounted to the panel in an electrical enclosure by snapping it onto what is termed a DIN rail. DIN rails are standard mounting systems which use a strip of metal which is U shaped with a lip at the top of each end of the U. You screw the DIN rail to the electrical panel in the enclosure, and most industrial components such as PLCs, terminal blocks, circuit breakers, and all sorts of other things will simply snap onto it and be ready to wire together. 22 Along the top and bottom of the PLC are terminals to which you can attach wires from the things you want to sense or control in the rest of the machine, such as push buttons, pilot lights, proximity sensors, and valves. 23 Inputs are along the top, and outputs are along the bottom. The FX3U-32M has 16 inputs and 16 outputs. These I/O can be 24 volts DC, or 100 to 120 volts AC, depending on the model. Alternatively it may have small relays for outputs. While AC I/O once predominated, 24VDC became the standard in most industries several decades ago as it interfaced with electronic devices more readily and also allowed for cheaper and more compact wiring. 24 More inputs and outputs can be added by connecting additional I/O modules to the main PLC next to it on the same DIN rail, up to a total of 256 I/O The connection is typically via short ribbon cables which plug into the adjacent module. These I/O modules look like the main PLC, but just add I/O. 25 Inside the PLC are a CPU, memory, and firmware. This model has 64K "steps" of memory, which can be thought of as how many instructions you can have. The program runs in RAM, but you can add a small flash memory device to save the program to. There is a run/stop switch that allows you to start or stop the user program. There is a port that you can use to connect the PLC to a laptop computer with a cable so you can download your program to it. 26 The biggest part of the market for PLCs is for these "shoebox" style. However, there are bigger ones as well which have more I/O, more memory, faster CPUs, etc. These are meant for tasks such as coordinating large assembly lines and things like that. These are what are termed rack systems, where the rack is an empty box which is open at the front and has a backplane bus running along the back. You plug the main CPU into the bus, normally in the leftmost position, and then plug I/O modules into the rack. The I/O modules are tall, thin, fully enclosed boxes with the I/O terminals along the front. 27 You should have a rough idea of their appearance at this stage, so I'll leave the physical description aside for now and go on to the concepts behind how they actually work and what makes them different from something like say a Raspberry Pi. -------------------- 28 Machine Architecture It is not really the hardware which defines the PLC. Rather, it is the software system architecture. Every PLC that I am aware of has a common set of features. 29 Data Table A data table is simply a block of memory which may or may not be subdivided into different parts. All logic, and all, or nearly all, I/O works by writing to and reading from addresses in the data table. 30 I/O Every input and output point maps to a fixed memory address in the PLC. To read the state of an input, you read its memory address. To change the state of an output, you write to its memory address. 31 These digital I/O appear as single bit boolean values. The PLC programming language has instructions which allow single bits to be read or written to directly without have to mask off other bits as you would have to do if you were using a conventional programming language. I have so far only mentioned digital I/O, which are single bit on/off values. 32 Other types of I/O Many PLCs also have other types of I/O such as Analogue I/O, which represent variable voltages rather than just on/off values. For example you may wish to read a temperature from a thermocouple where the voltage level varies with the temperature. Often these are mapped into the data table as byte or word values. 33 Internal Bit or Boolean Memory There will be a range of single bit or boolean memory which is used for internal logic state. Your program may need to come up with a series of intermediate logical values, and you store them in these internal bits or flags. 34 Timers and Counters Control of machinery often involves timing and counting. Timer and counters are mapped into the data table. Timer and counter presets and values can be read as words, and there will be bit addresses which indicate the status of the timer or counter, such as whether it has reached its preset and is "done". 35 Integer and Floating Point Values There will be word addresses where integer and floating point words can be stored. If you need to do some mathematical calculations and store the results, you would save them in an integer or floating point word. 36 Typed Memory Unlike when programming something like a PC where you simply have a range of undifferentiated bytes and it is up to your software to impose meaning on it, PLC data table memory has defined types and meanings and the system firmware enforces the correct memory access methods. 37 Size of Data Table Data tables vary greatly in size. A bigger data table allows for bigger and more complex program. Generally, new PLCs will have bigger data tables than older ones, and more expensive PLCs will have bigger data tables than cheaper ones of the same generation. 38 Example Having picked a current PLC as an example for physical dimensions, I will pick an old and obsolete one for an example of a data table. The Siemens S5-100 series was a small PLC that came in several sizes. The basic S5-100 had the following data table size. 39 Digital I/O had a maximum of 128 inputs and outputs taken together. Analogue inputs and outputs had a maximum of 8 taken together. Flag, or bit, memory was 1024 Timers - 16 Counters - 16 40 The S5-103 offered more of everything in the same basic package, but of course at a higher cost. Digital I/O had a maximum of 256 inputs and outputs taken together. Analogue inputs and outputs had a maximum of 32 taken together. Flag, or bit, memory was 2048 Timers - 128 Counters - 128 -------------------- 41 The PLC Program A PLC will come with what amounts to an operating system and virtual machine built into it. You as the programmer use the vendor's proprietary development software, what is generally called the "PLC programming software" to write an application for it. 42 You then connect your PC to the PLC using a cable of some sort, and download your program into it. The program gets stored in the PLC's RAM. There may be flash memory which serves to back up the RAM when the power is off so you don't lose the program. Before flash was available, a battery was typically required to hold the static RAM memory. 43 There is just one program in the PLC. There is no file system, just the program memory and the data table. When the PLC starts up, it runs the program that you wrote. It continues running the program until you turn off the power or you set the run/stop switch to the stop position. 44 I will come back to programming later, but we needed to cover these basic points before I could describe some further concepts we need to cover. -------------------- 45 The Scan A key thing to understand about PLCs is the "scan" concept. This works as follows. There is a repeated cycle called the scan. 46 First, the PLC reads all the physical inputs and updates the corresponding input addresses in the data table. Next it runs the user program once. Finally it reads the output addresses in the data table and writes them to the corresponding physical outputs. Then it repeats the scan from step 1. 47 A scan can take anywhere from a few milliseconds to hundreds of milliseconds, depending on the model of PLC. Older PLCs and cheaper PLCs tended to be slower. However, as integrated circuit technology advanced, programs got faster. 48 The PLC has a watchdog timer. This is a timer which runs in the background and resets at the start of a scan. 49 If a scan takes too long, the watchdog will trip and stop the program and set the outputs to some defined state, either turning them all off or holding them at the last value. This is known as "faulting" the processor. 50 A watchdog fault may be caused by a program that is too long, or uses a lot of very slow instructions, or in some cases it may be due to a bug in your program. However, this last cause is not as common or as easy to cause as you may think when it comes to faults. -------------------- 51 PLC Programming Languages The key to all of this is the unique way in which the programming languages used in PLCs work. I will first briefly describe the two main programming languages and then explain how they work in a PLC as part of a complete system. 52 Ladder Logic There are several programming languages, but the main one is called ladder logic. If you listened to the previous episode, the word "ladder" will ring a bell. When machines were controlled by electromechanical relays, the electrical drawings which specified and documented the types of relays and the connections between them were called "ladder diagrams". 53 PLCs simply take these ladder diagrams and reproduce them on your computer screen in the programming software using the same standard schematic symbols. 54 You do not need to draw the symbols like it was CAD. Instead you simply use your keyboard or mouse to say that you want this symbol to be entered in the current cursor location or you want this wire to be connected from here to there. When your rung is complete the software will check at some point if it is syntactically correct. You then go on to entering the next rungs one after another until you have your complete program. 55 Instruction List An alternative representation is something called "instruction list", although various vendors may use other terms. This is a text representation that looks like a sort of assembly language. However, it's actually really ladder, just shown in a different way as text and some types of programming software will allow you to toggle between ladder and instruction list mode, translating between them automatically. 56 This should give you a clue as to how a PLC can execute a schematic diagram. Behind the scenes the programming software will translate the diagram into instruction list, and then the PLC will either execute those instructions in an interpreter, or compile them to machine code and execute those. 57 Other Programming Languages There are other programming languages, but they are rarely seen. I won't go into detail in terms of describing them, just briefly mention them. 58 One is called Sequential Function Chart. This is a type of flow chart which is designed to show complex sequences, including ones with alternate or parallel paths. The original name for this is Grafcet, spelled G R A F C E T. This stands for "Graphe Fonctionnel de Commande Étape Transition". As well as being a programming language, it is also a very good design and analysis tool even if the PLC being used doesn't support it. 59 Another is Function Block Diagram. This is another graphical language which resembles flow diagrams used in process industries. This is supposedly used mainly in industries such as chemical manufacturing. However it appears to be very niche and most people who use PLCs will never have seen it, and many will not have even heard of it. 60 Yet another language is Structured Text. This is very obviously derived from Modula-2 and strongly resembles it. If you are not familiar with Modula-2, it was created by Nikolas Wirth and intended as the successor to Pascal, which it was derived from. Structured Text seems to be much loved by a small number of academics, but seems to have little actual use in the field. You lose pretty much all of the monitoring and debugging facilities built into a PLC if you use it, so it's kind of pointless except perhaps for a few niche applications. There are a few other languages as well, mainly proprietary ones, often using flow charts or the like. 61 PLC Languages are Non-Blocking The important thing about a PLC program is that it executes from top to bottom, left to right, without stopping or blocking. There is no waiting on input or output or for a system call to return. Each rung immediately yields a result which is written to a memory address or which enables a timer or counter. Every rung is executed every scan. 62 You can reproduce this in a traditional computer programming language. In ladder logic it is inherent to the syntax and it is either very difficult or impossible to do it any other way. I'll give an example in Python of what I mean. 63 a = (b or c) and not d 64 If b or c are true and d is false, then a is set to true. If b and c are false or d is true, then a is false. 65 Now imagine that we are not executing this statement once, but rather are executing it repeatedly. This statement will never block execution. It will always immediately yield a result. Now let's modify that a bit. 66 a = (b or a) and not d 67 Note that we have replaced c with a. Now the value of a depends the previous value of a as well as b and d. However, once a becomes true, the value of b no longer matters because b is in an or condition with a. Now if we execute it over and over again, only d becoming true will make a go false. 68 This is a standard push button circuit, also known as a "seal in circuit", because it "seals" around the start condition. 69 b is the start push button with a normally open contact. d is the stop push button with a normally closed contact. a is a relay with one of its contacts being used to hold it on. 70 If you wired this up with actual push buttons and relays or if you programmed it into a PLC with ladder logic it would work the same way. A PLC program will consists of rung after rung of logic like this and executes it repeatedly scan after scan, only pausing to update its I/O. 71 Async Programming Some of you may be thinking that this cyclical scan sounds like the async programming that is all the rage with web server applications these days. Essentially it is exactly the same principle. 72 With async programming, your program must never use blocking instructions and execution proceeds on a repeated cycle. PLCs solve this by not having any instructions which block, and many PLCs do not allow backwards jumps. Tradesmen in coveralls in factories were doing async programming decades before the cool kids heard about it. 73 Most PLC Programming Languages are Visual or Graphical Languages If ladder, Sequential Function Chart, or Grafcet and Function Block Diagram sound like visual programming languages, they are. In fact ladder is probably one of the earliest visual languages in commercial use. Wikipedia defines a visual programming language as: 74 In computing, a visual programming language (visual programming system, VPL, or, VPS), also known as diagrammatic programming, graphical programming or block coding, is a programming language that lets users create programs by manipulating program elements graphically rather than by specifying them textually. A VPL allows programming with visual expressions, spatial arrangements of text and graphic symbols, used either as elements of syntax or secondary notation. For example, many VPLs are based on the idea of "boxes and arrows", where boxes or other screen objects are treated as entities, connected by arrows, lines or arcs which represent relations. VPLs are generally the basis of low-code development platforms. Scratch is an example of a VPL 75 End of quote. As you can see, what is cool today was on the factory floor more than 40 years ago. -------------------- 76 Popularity of PLC Programming Languages PLCs are about as proprietary as you can get. Pretty much every vendor has his own proprietary take on each language. How you can program any particular PLC is determined by that vendor's programming software. If the vendor doesn't support it, you can't use it. An individual vendor may even have multiple incompatible product lines which have to be programmed in different ways with different software, although that is not as common these days as it once was. 77 Nearly all PLCs support programming in Ladder. Some support programming in instruction list as well as ladder. Anything else is much less common. 78 Ladder logic happens to be on the Tiobe Index by the way. For those who have not heard of it, the Tiobe Index, that is T I O B E, is a web site that lists the popularity of numerous programming languages based on various criteria. 79 Number one on their list happens to be Python, currently at 19.98%. Number two is C, at 11.55%. 80 At the time of writing this script, Ladder Logic was at number 46 with a 0.28% rating, which put it just below Erlang and just above Haskell. I'm not sure whether that means that Ladder Logic is not as obscure as you thought it was, or whether it means that Erlang and Haskell are in fact more obscure than you thought they were. -------------------- 81 Episode Summary In this episode we covered we covered The early history of computers in industrial control The early history of PLCs, including how they got their name Who the major brands are What they look like physically 82 A basic description of the abstract machine architecture A very brief look at what a PLC program is like The scan concept The main PLC programming languages The minor PLC programming languages The relative popularity of each of the programming languages 83 In the next episode we will take a look at one of the early PLCs from the era when they began seeing widespread use. This PLC was hugely successful and was for many companies the first PLC they used. This is the Allen Bradley PLC2. 84 This has been the second episode in an 8 part series. -------------------- Provide feedback on this episode.

David Bombal
#600: This Free Tool Decodes Game Boy ROM From a Photograph

David Bombal

Play Episode Listen Later Sep 2, 2026 41:33


Can you recover working software from a photograph of a microchip? Embedded systems reverse engineer Travis Goodspeed demonstrates how to extract the original Game Boy's 256-byte boot ROM from a microscopic photograph of the mask ROM inside its CPU. The source image was created by combining 22 microscope photographs captured at 50x magnification. Using his free, open-source Mask ROM Tool, Travis marks 2,048 ROM bit locations, separates the ones from the zeros and checks for possible recognition errors. He then determines the logical order of the bits, disassembles the program and shows how it can be exported as a ROM file for use in an emulator. Travis also explains the Game Boy's unusual copy-protection system. During startup, its boot ROM displays logo data supplied by the cartridge and compares it with Nintendo's internal copy. If the logos do not match, the game does not boot. Requiring cartridges to contain Nintendo's trademark gave the company legal leverage against unlicensed publishers. The video also explores techniques from Travis's book, Microcontroller Exploits. These include extracting protected firmware from an access-control reader, chemically decapsulating chips while preserving their operation and using ultraviolet light with a nail-polish mask to remove memory protection without erasing the program. The Mask ROM Tool, Game Boy chip photograph and step-by-step tutorial are publicly available, allowing you to reproduce the ROM-decoding demonstration without owning a microscope or chemistry lab. // Sponsored SEGMENT // Big thank you to Proton Pass for sponsoring this video. Take your security to the next level by getting Proton Pass using the following be www.proton.me/davidbombal // Link to No Starch Website for Travis' Book // Order Microcontroller Exploits and get the eBook free. https://nostarch.com/microcontroller-... Use Coupon Code GOODSPEED25 for 25% off Microcontroller Exploits at NoStarch.com // Travis Goodspeed SOCIAL // GitHub: https://github.com/travisgoodspeed // GitHub link to GameBoy ROM Tutorial // https://github.com/travisgoodspeed/gb... // David's SOCIAL // Discord: discord.com/invite/usKSyzb Twitter: www.twitter.com/davidbombal Instagram: www.instagram.com/davidbombal LinkedIn: www.linkedin.com/in/davidbombal Facebook: www.facebook.com/davidbombal.co TikTok: tiktok.com/@davidbombal YouTube: / @davidbombal Spotify: open.spotify.com/show/3f6k6gE... SoundCloud: / davidbombal Apple Podcast: podcasts.apple.com/us/podcast... // MY STUFF // https://www.amazon.com/shop/davidbombal // SPONSORS // Interested in sponsoring my videos? Reach out to my team here: sponsors@davidbombal.com // MENU// 0:00 - Coming Up 0:40 - Intro 03:28 - Book Overview 05:27 - Proton Pass Ad 07:29 - Demonstration Context 09:56 - Demo Begins 12:48 - Marking the Chip Rows 18:27 - How Travis Wrote his Book 19:55 - Design Rule Check 23:11 - How to Decode the Binary 28:36 - Can the Binary Numbers Change? 30:30 - Why is this Method Useful? 34:24 - More book Overviews 40:48 - Conclusion Please note that links listed may be affiliate links and provide me with a small percentage/kickback should you use them to purchase any of the items listed or recommended. Thank you for supporting me and this channel! Disclaimer: This video is for educational purposes only. #gameboy #reverseengineering #rom

Python Bytes
#494 Python Wrapture

Python Bytes

Play Episode Listen Later Sep 1, 2026 28:37 Transcription Available


Topics covered in this episode: OpenAI's Python SDK has migrated to HTTPX2 TMOG - Native Task Manager for macOS, Windows, and Linux wrapture - one wrapper for mocking, tracing, and observability linkedin2md: turn your LinkedIn export into 40+ Markdown files Extras Joke Watch on YouTube About the show Sponsored by us! Support our work through: Our courses at Talk Python Consulting from Six Feet Up Connect with the hosts Michael: Mastodon / BlueSky / X / LinkedIn Calvin: Mastodon / BlueSky / X / LinkedIn Show: Mastodon / BlueSky / X Join us on YouTube at pythonbytes.fm/live to be part of the audience. Usually Tuesday at 7am PT. Older video versions available there too. Finally, if you want an artisanal digest of every week of the show notes in email form? Add your name and email to our friends of the show list, we'll never share it. Calvin #1: OpenAI's Python SDK has migrated to HTTPX2 The OpenAI Python SDK has migrated to HTTPX2, the Pydantic-stewarded fork of httpx. Pydantic picked it up citing "limited activity recently" in the original project, promising "a reliably maintained path forward." If you just use the default client, nothing to do. No code changes. The catch is TLS. Quoting the guide: HTTPX "previously verified certificates against the CA bundle provided by certifi. HTTPX2 instead uses the operating-system trust store, and the SDK no longer installs certifi." That "can break certificate verification in minimal container images without system CA certificates, environments using corporate TLS-inspecting proxies, and deployments that relied on a custom or modified certifi bundle." The fix is SSL_CERT_FILE or SSL_CERT_DIR, or pass your own ssl.SSLContext via verify. Deeper integrations need real edits: custom clients, auth handlers, hooks, and request mocking all take HTTPX2 objects now, and plain httpx is no longer pulled in transitively. So import httpx in your own code means declaring it yourself or moving over. Temporary escape hatch: a legacy HTTPX client Michael #2: TMOG - Native Task Manager for macOS, Windows, and Linux A native, deeply instrumented system monitor for macOS, Windows, and Linux, now in public beta - from Plummers' Software, i.e. Dave Plummer, who wrote the original Windows Task Manager and donated it to Microsoft in 1995. Wikipedia Three real native apps: Swift/AppKit on macOS, Win32 on Windows, C++/Qt 6 on Linux, with a shared C++ core keeping metric semantics aligned - no browser shell anywhere. One dense summary: CPU, clocks, thermals, GPU, memory, storage, network, energy, and the processes responsible for the load, all click-through. Per-core honesty: logical processor and NUMA views, P and E cores color-coded, optional kernel time, 60 FPS live meters. Memory with context: pressure, wired, compressed, cached, committed, available, and swap, plus configurable scrolling history. Processes that act like processes: tree view, filtering, sorting, follow mode, and native verbs including service and launchd control. Phosphor themes: light, dark, green, amber, blue, or mono, with color and saturation you tune yourself. Calvin #3: wrapture - one wrapper for mocking, tracing, and observability Graham Dumpleton, author of wrapt and the original New Relic Python agent, has released wrapture. The name is wrapt plus capture. The core idea: wrap real code instead of replacing it, so the real code still runs while you watch every call. Name a method with wrapture.binding(Class, "method"), open a timeline(), and you get a tape of what actually happened. Real return values, real nesting, arguments normalised against real signatures. tape.tree() prints the call graph as it ran. One mechanism, three jobs: monkey patching with a real lifecycle (apply, remove, suspend, plus returns, raises, transforms_args), unit testing that asserts on real call flow instead of a flat MagicMock call list, and ad-hoc tracing of a running app. The testing pitch is error paths. Inject TimeoutError at the payment gateway, then assert the ledger was never written. Stubs and mocks are strict and spec-required, and there is deliberately no bare Mock(). Tracing needs no code at all. A wrapture.toml naming targets and a sink, run with python -m wrapture main.py, and you get a live call tree with timings. It captures ordinary logging calls as nested events, and with the otel extra it exports spans, metrics and correlated logs with W3C trace ids that join across services. Every line of code and docs was AI-written under their direction, and they say so up front. Two weeks from first commit, eleventh alpha, over 1000 tests, 150+ pages of docs. Alpha on PyPI, needs Python 3.12+ and wrapt 2.4.0+. Michael #4: linkedin2md: turn your LinkedIn export into 40+ Markdown files Via Juan Manuel Daza - a Python CLI that unpacks LinkedIn's data-export ZIP into clean, per-category Markdown you can drop straight into an LLM. One command: linkedin2md Complete_LinkedInDataExport.zip, plus o for output dir, -lang en|es, and -pdf. 40+ output files: profile, experience, education, skills, connections, posts, comments, reactions, recommendations, endorsements, job applications, even ad targeting and LinkedIn's inferences about you. Built for LLM analysis: the README pitches NotebookLM, Claude Projects, Obsidian, and Ollama, with example prompts like "what patterns do you see in my career transitions?" PDF resume mode: -pdf renders an A4 CV via weasyprint, and degrades gracefully to Markdown-only if it isn't installed. Dependency note: "pure Python / zero-dep" holds for the Markdown path only - the PDF path needs weasyprint and markdown installed. Install: pipx install linkedin2md recommended, pip in a venv otherwise - 86% Python, 10 releases, v0.3.1 in May. Agentic dev angle: repo ships opencode config and an N3RV subagent pipeline, including a "judgment day" dual-model adversarial PR review. Extras Calvin: EVE Online Migrates to Python 3 Michael: Dinkus by Will McGugan Joke: Tao of Programming: Book 5 Maintenance

Antreprenori care Inspira cu Florin Rosoga
Cum este construită tehnologia pe care o folosim în fiecare zi? - cu Darian Domocoș

Antreprenori care Inspira cu Florin Rosoga

Play Episode Listen Later Sep 1, 2026 15:00


Inteligența artificială a adus din nou hardware-ul în centrul atenției. În spatele fiecărui model pe care îl folosim astăzi stau decizii luate cu ani înainte, în laboratoare, în echipe de ingineri și în companii care proiectează infrastructura pe care rulează software-ul. Despre această lume vorbește episodul de astăzi, alături de Darian Domocoș, cofondator și CEO al Fiora5, un start-up românesc care dezvoltă procesoare și proprietate intelectuală pentru industria semiconductorilor.Discuția pornește dintr-un loc personal. Darian povestește cum a intrat în acest domeniu încă din adolescență și cum o experiență dintr-o companie de inginerie i-a schimbat complet direcția profesională. Câțiva ani mai târziu, împreună cu fostul său coleg, avea să pună bazele Fiora5. Compania dezvoltă partea de CPU dintr-un procesor și se adresează unei piețe aflate într-o schimbare accelerată, în care cererea pentru arhitecturi noi și soluții eficiente crește odată cu dezvoltarea inteligenței artificiale.Pentru mai multe resurse despre episodul de astăzi, notițe, ideile sumarizate - click aici pentru pagina episodului***Găsești notițe, ideile principale, insight-uri și cărțile menționate în toate episoadele pe florinrosoga.ro. Aici te poți înscrie și la un newsletter.Dacă îți plac aceste podcasturi, ajută-ne cu o recenzie pe Spotify sau Apple Podcasts. Este un gest simplu care ne ajută să abordăm subiecte și invitați interesanți.***Podcasturile noastre sunt aici:

Badlands Media
The Narrative Ep. 80: The Strait Truth

Badlands Media

Play Episode Listen Later Aug 31, 2026 147:53


A rare solo outing. Burning Bright kicks off Episode 80 by taking full credit for its title, one of the cleverest plays on words you have probably heard in the last few minutes, if not your entire life. No guest was lost. He simply wanted to get the ideas out of his head before they pile up like bad cache in a CPU. Only his second solo stream ever, this one doubles as practice and as a table setter. Starting into September he plans to alternate solo weeks with guest weeks, pending audience reaction. The substance is a State of the War address. Expect a run through his eight pieces of the Iran War thesis, the Strait of Hormuz as "Schrodinger's Strait," global energy and the oil cartel angle, and how the same template maps onto the Taiwan situation so you can keep a level head when the narrative arrives. Pure pattern recognition, delivered with his usual dry confidence.

Atareao con Linux
ATA 827 No Necesitas una GPU de 3000€ para IA Local

Atareao con Linux

Play Episode Listen Later Aug 31, 2026 22:36


Cerramos la octava temporada con un episodio que me apetecía grabar desde hace meses. Igual te ha pasado como a mí: empecé hablando de un laboratorio de IA para cualquiera, y terminé recomendando GPUs de 3000 euros. Me fui creciendo, pero no hace falta. Te cuento cómo montar un laboratorio de IA local con el equipo que ya tienes. Da igual si tienes 8 GB de RAM o 16, CPU modesta o sin GPU. La clave está en elegir los modelos adecuados. Muchas veces nos perdemos buscando el modelo más grande, cuando con uno pequeño y bien cuantizado tenemos de sobra para el 80% de las tareas.Te hablo de Ollama, el gestor de modelos estándar para ejecutar modelos locales. Más de 180.000 estrellas en GitHub, API compatible con OpenAI, modelos para todos los presupuestos: desde Phi 3.5 con 3.8B parámetros hasta Qwen 1.5B que ocupa 1 GB. También la cuantización: reduces la precisión numérica de los pesos para que ocupen menos y vayan más rápido. El punto dulce es Q4_K_M, que reduce el tamaño a menos de un tercio. Para 8 GB de RAM, Q3_K_S puede ser tu salvación.También te hablo de Open WebUI, la interfaz que le da mil vueltas a ChatGPT. No solo chateas: tiene RAG local, Whisper integrado para transcribir voz (75 MB en CPU), TTS con Kokoro-82M para que el modelo te hable en tiempo real, búsqueda web, plugins y memoria persistente. Todo en un contenedor Docker que levantas con un solo comando.Y de SQLite Vec, extensión de SQLite sponsorizada por Mozilla para búsqueda semántica sin servidores vectoriales. Ni ChromaDB, ni Qdrant, ni Milvus. C puro que funciona hasta en Raspberry Pi. Creas tablas virtuales para vectores de 768 dimensiones, generas embeddings con nomic-embed-text, y buscas por similitud coseno en milisegundos. RAG local sin complicaciones.Y te explico cómo organizarlo todo con Docker o Podman. Un docker-compose.yml que levanta Ollama y Open WebUI en segundos, con healthchecks, redes separadas y volúmenes persistentes. También a limitar recursos con --memory y --cpus. He preparado scripts: inicialización que comprueba requisitos, crea directorios y descarga modelos; otro para descargar por niveles según tu hardware (nivel 1 para 8 GB, nivel 2 para 16 GB, nivel 3 para 32 GB); y uno de respaldo.Y la estrategia híbrida local + nube, que es lo que realmente tiene sentido. El enfoque Minions del Stanford Hazy Research Lab: el modelo local hace el trabajo pesado, y solo consulta al grande en la nube para tareas complejas. El 90% de las consultas se resuelven localmente. Ahorras dinero, mantienes privacidad de tus datos, y cuando necesitas potencia, la tienes.Con 16 GB de RAM y un SSD te sobra para el 80% de las tareas: traducciones, resúmenes, código, asistentes, RAG, transcripción de audio, texto a voz... Todo en tu máquina, sin enviar datos a servidores, sin suscripciones, sin depender de internet. Con 8 GB también puedes, con modelos más pequeños. Cerramos temporada, la novena arranca en el episodio 828. Capítulos del episodio:0:00 - Introducción — cierre de temporada 8 y replanteamiento2:30 - Hardware mínimo: 8-16 GB RAM + SSD obligatorio5:00 - Software base: instalar Ollama en tu distribución7:30 - Contenedores: Docker vs Podman para el laboratorio10:00 - Modelos pequeños: Phi 3.5, Qwen 1.5B y cuantización13:00 - Herramientas complementarias: SQLite Vec, Whisper, TTS16:00 - Organización del laboratorio: script y estructura de directorios19:00 - Demo: probando Ollama en local con modelos ligeros22:00 - Combinación local + nube: lo mejor de ambos mundos24:30 - Cierre, avance temporada 9 y despedidaMás información y enlaces en las notas del episodio

The MacRumors Show
208: Apple's Surprisingly HUGE Week of Announcements

The MacRumors Show

Play Episode Listen Later Aug 28, 2026 46:47


It's been an unusually big week for Apple announcements. On this week's episode of The MacRumors Show, we discuss the new Mac mini and Mac Studio, the M6 and M5 Ultra chips, and invites to Apple's September event. Apple confirmed on Wednesday that its annual iPhone event will take place on Wednesday 9 September at Apple Park, starting at 10:00 a.m. Pacific Time, with select members of the media invited to attend. The event's tagline is "Surprise and Shine." Apple is expected to introduce at least six new products, including the the Apple Watch Series 12, Apple Watch Ultra 4, iPhone 18 Pro, ‌iPhone 18 Pro‌ Max, and the first foldable iPhone. Pre-orders should follow shortly after, and release dates for iOS 27, iPadOS 27 and macOS 27 Golden Gate should also be revealed at the event.Apple announced the new Mac mini on Tuesday, offering M6 and M5 Pro chip options. The machine gains an N1 networking chip, bringing Wi-Fi 7 and Bluetooth 6, and moves to 2.5Gb Ethernet as standard, up from 1Gb, with a 10Gb option. Thunderbolt 5 is restricted to the M5 Pro model, which gets three Thunderbolt 5 ports while the M6 model keeps Thunderbolt 4. Genlock supportover USB-C is also new, syncing a display's refresh timing to a camera such as the iPhone 17 Pro. Pricing starts at $899 with 16GB and 256GB, rising to $1,699 for the M5 Pro at 24GB and 512GB. Pre-orders are now open, with launch on 22 September. M6 is Apple's first 2nm chip. Its 12-core CPU is split into 2 super cores, 4 performance cores and 6 efficiency cores, making it the first Apple chip to combine all three core types in one design. Super cores are Apple's top tier, optimised for single-threaded speed. The 12-core GPU features a Neural Accelerator in each core for the first time in a ‌Mac mini‌. The GPU also adds an updated shader architecture, Dynamic Caching, and hardware-accelerated ray tracing, while an all-new Dual 16-core Neural Engine doubles peak compute over the previous generation, with system frameworks able to drive both engines at once. Memory starts at 16GB and tops out at 32GB, running at up to 170GB/s. M5 Pro scales further, to an 18-core CPU, a 20-core GPU and up to 64GB at 307GB/s.The new Mac Studio arrived alongside the ‌Mac mini‌, with M5 Max and M5 Ultra chip options. The M5 Max offers an 18-core CPU with 6 super cores and 12 performance cores, an up-to-40-core GPU with Neural Accelerators in every core, and up to 128GB of memory at 614GB/s, with the GPU running up to 50% faster than the previous generation. Both configurations move to a next-generation SSD architecture on PCIe Gen 6 for up to twice the storage performance, and gain up to six Thunderbolt 5 ports at 120Gb/s, and the same genlock support as the ‌Mac mini‌. Apple's N1 chip also brings Wi-Fi 7 and Bluetooth 6 to the ‌Mac Studio‌. Pricing starts at $2,499 for the M5 Max at 36GB and 512GB, and $5,499 for the M5 Ultra at 96GB and 1TB, with the 512GB memory option not arriving until late October.M5 Ultra is Apple's first quad-die chip, using a next-generation version of UltraFusion to join two dual-die M5 Max chips into a single processor. The result is a CPU with up to 36 cores, and 12 super and 24 performance, delivering 1.25x the single-threaded and 1.3x the multithreaded performance of the M3 Ultra. Its up-to-80-core GPU is the most powerful Apple silicon GPU ever made and the first on an Ultra chip to feature Neural Accelerators. Second-generation Dynamic Caching, hardware-accelerated mesh shading and third-generation ray tracing lift graphics up to 40% over M3 Ultra, and the chip features a 32-core Neural Engine. Memory reaches 512GB at 1.2TB/s, which is a 50% bandwidth increase.Other announcements this week included a new, softer Apple Polishing Cloth at $9, down from $19, and refreshed U.S. Magic Keyboards for the Mac, now featuring glyphs instead of edge text. Apple Support's 1-800-APL-CARE line is also now answered by a generative AI assistant in the U.S. and Canada.Ready to tackle bigger problems? Get started with Claude today at — https://www.Claude.ai/mac

Geek News Central
Eyes, Hands, and a Sense of Timing #1874

Geek News Central

Play Episode Listen Later Aug 28, 2026 51:40 Transcription Available


In this episode, Ray Cochrane digs into Anthropic’s Model Hardware Standard. It is a shared driver that lets an AI agent run real lab equipment, from pipetting robots to the lasers inside a quantum computer. He also covers OpenAI’s builder’s guide to GPT-5.6, Google’s new Expert Intelligence book feature, Apple’s M5 Ultra Mac Studio, and a judge’s order forcing Google to stop hiding rival app stores. Finally, he weighs in on Apple’s proposed 15 percent link-out fee, Meta’s Australia numbers, the White House deputizing private hackers, and why rivers obey a 1957 math rule. – Want to start a podcast? Its easy to get started! Sign-up at Blubrry – Thinking of buying a Starlink? Use my link to support the show. Subscribe to the Newsletter. Email Ray if you want to get in touch! Like and Follow Geek News Central’s Facebook Page. Support my Show Sponsor: Best Godaddy Promo Codes Get 1Password Full Summary Cochrane opens with a quick personal update. He is hunting for tickets to Michigan for his dad’s anniversary, and he has been learning Blender and Godot on the side, mostly modeling and blocking out levels. Consequently, he asks listeners for advice on starting a big game project, and he plans to record his progress, maybe as a time lapse. Then it is straight into the featured story. Anthropic’s Model Hardware Standard: A Driver for the Physical World The featured story comes from Anthropic, which opened a research preview of the Model Hardware Standard, or MHS. Cochrane frames it as the other side of the question NVIDIA’s world models raised two weeks ago: when do AI agents start touching actual machines? A typical lab runs a microscope, a liquid handler, a robotic arm, and a plate reader, each from a different vendor with its own control software. One Janelia researcher in the post launches seven programs in three languages just to start an experiment. Anthropic says wiring a setup like that takes weeks or months of specialist work. MHS is a driver, the same kind of translation layer a printer uses, except every device gets described with a tiny set of commands like read and write. Devices announce themselves on the network. A plain-English reference file then records what each machine measures, what can be adjusted, and which safety limits get enforced no matter what the agent asks. Agents then reach the hardware through the Model Context Protocol, the command line, or plain code. Cochrane sees the same move the industry keeps making, from coding harnesses to RSS and JSON: agree on a standard and let everyone build against it. In fact, he calls MHS the hardware version of MCP. The partner results carry the segment. QuEra builds quantum computers from individual atoms held by lasers that must hold their frequency to about one part in a trillion. A four-person team spent months on a relock script that worked 58 percent of the time. However, four copies of Claude iterating overnight through MHS produced a decision-tree script that recovers the laser in about six seconds, and it passed 99.3 percent of 700 blind trials. Carnegie Mellon wrote MHS drivers for four instruments across three incompatible computers in about eight hours, then ran dose-response experiments three times faster and blocked all six deliberately induced faults. Genentech, meanwhile, showed the limits. Claude used the same pump speed for water, a foamy protein solution, and a human had to explain that the bubbles were a physics problem. That gap in physical intuition is what sticks with Cochrane. He doubts it will change soon, and he suspects the fix will arrive as sub-agents or sub-models that judge a request against an expected outcome. He also connects MHS to a video of racing robots that never learned to stop at the finish line. What happens, he wonders, once they can read a distance sensor through a shared standard? Still, he calls the announcement a fantastic read and points listeners to the full article. Sponsor: GoDaddy Economy hosting $6.99/month, WordPress hosting $12.99/month, domains $11.99. Website builder trial available. Use codes at geeknewscentral.com/godaddy to support the show. GPT-5.6 Does the Same Work for a Fraction of the Cost OpenAI’s builder’s guide to GPT-5.6 leads the headlines. Cochrane recaps the three tiers from episode 1870, Sol, Terra, and Luna, plus the separate dial for reasoning effort. On BrowseComp, a benchmark for digging up obscure facts on the web, the old GPT-5.5 flagship scored about 84 percent on a run that cost 33 dollars three months ago. Luna now matches that score for a dollar thirty-three, and OpenAI has since cut Luna’s price another 80 percent. Browser Use reports Luna finishing 78 percent of its hardest browser tasks for about 14 dollars, against 80 percent for roughly 235 dollars from the best available model. The guide’s other big addition is a multi-agent beta flag. It lets the model handling a request spawn parallel helper agents that report back to a root agent inside a single API call. However, Cochrane is unimpressed by the timing. He has been running that pattern in Claude Code for months, so he sees OpenAI copying a workflow other companies already ship rather than inventing its own. Along the way, he plugs Claude Code’s remote-control sessions, which let him send prompts from his phone to a terminal session at home. Google Lets Gemini Read the Books You Actually Bought Google launched Expert Intelligence, a name Cochrane calls quite the reach. The feature lets you drop a book you bought on Google Play Books into Gemini Notebook, formerly NotebookLM, and ask questions answered only from that book, with citations. Cochrane sees real power here for students, since he once used NotebookLM to organize scattered course PDFs. Additionally, publishers get a cut, which he calls a far better deal than the wholesale scraping of books that trained earlier models. Nevertheless, he asks who loses out, because a paid publisher does not automatically mean a paid author. He floats the same idea for artists, even a penny per use, then admits that may be too idealistic. Apple’s M5 Ultra Mac Studio Is Built to Run Big Models at Home Back in episode 1861, when Apple killed the Mac Pro, an M5 Ultra Mac Studio was expected later this year. Now it is here. The M5 Ultra brings up to a 36-core CPU, an 80-core GPU, and 512GB of unified memory moving 1.2 terabytes per second. Apple claims up to 4.3 times the AI performance of the M3 Ultra. Thunderbolt 5 can also cluster four machines into one memory pool for up to three times faster inference. The M5 Max model starts at $2,499 and the Ultra at $5,499, with shipping on September 22 and the 512GB configuration arriving in late October. Cochrane finds the clustering pitch ridiculous at that price, but he invites anyone who spends the money to report back. Apple Opens a Manufacturing School in Houston Apple also opened a 20,000-square-foot Advanced Manufacturing Center in Houston. It offers free classes for small and midsize manufacturers, from circuit board design to hands-on time on a scaled-down production line, with college students joining later. Cochrane calls it a solid step in the bring-manufacturing-home movement. The bigger story is the campus itself, which builds Apple’s AI servers and will add the first US-assembled Mac mini line later this year. That ties back to the Mac mini shortage that followed the OpenClaw rush, when Tim Cook warned of months-long waits. Cult of Mac was still reporting four-month waits in late July. However, Cook blamed chip supply rather than assembly, so Cochrane is not counting on relief just yet. Amazon EC2 Turns Twenty Amazon EC2 turned twenty this week, which Cochrane admits makes him feel old. The 2006 beta offered one server size in one region for ten cents an hour. Each came with a 1.7 gigahertz Xeon and under two gigabytes of memory, and accounts were capped at twenty servers. Today AWS offers more than 1,200 instance types across 39 regions. Consequently, Cochrane credits the company with turning that tiny product into the backbone of cloud and AI computing. Intel Gamer Days: Two Free Games, With Fine Print Intel Gamer Days runs through September 13. Buy a qualifying Core Ultra Series 2 or 14th Gen desktop chip, a Core Ultra Series 3 laptop, or an Arc graphics card. In return you get Star Wars: Galactic Racer plus the Tomb Raider: Legacy of Atlantis remake. GamesRadar values the pair at about 120 dollars. However, neither game is out yet, and codes must be redeemed by October 31 even though the Tomb Raider remake ships in February. Cochrane calls that awful, but he still tells qualifying buyers to claim the deal early. Note that 13th Gen chips do not qualify. Judge Orders Google to Stop Hiding Rival App Stores A jury found Google’s Android app monopoly illegal in late 2023, and Judge James Donato ordered rival stores into the Play Store in 2024. On August 13, Epic’s lawyer demonstrated that searching Play for “store for apps” returned Walmart instead of any app store. Donato called that “not acceptable” and ordered three fixes within a week. Searches must surface third-party stores, listings need a plain install button, and the “are you looking for” interstitial has to go. Cochrane welcomes the monopoly being chipped away, but he notes that a controlling entity still sits atop every app store. In his view, community hubs like app stores and social media need a public infrastructure layer. He suspects governments skip that investment because companies already run the services, while selling your data. Apple Wants 15 Percent of Purchases Outside Its Store The other half of the Epic saga is Apple’s proposed link-out commission. After the 2021 anti-steering injunction, Apple charged 27 percent on purchases made through external links. A judge held it in contempt last year, and the Ninth Circuit then allowed a fee limited to the cost of running the system. Judge Yvonne Gonzalez Rogers refused to wait for the Supreme Court, writing that “further delay is unwarranted.” Apple filed 15 percent for standard apps, 10 percent for subscription renewals and partner programs, and 5 percent for small businesses. It also conceded the rate would be “essentially zero” under the appeals court’s cost yardstick. Since Apple has charged nothing on link-outs since the contempt ruling, Cochrane sees this as a raise. He calls a cut on purchases made on a developer’s own website disturbing. He also recalls reading about the size of Uber’s payments to Apple, and he questions whether that kind of percentage is sustainable for companies without funding. Meta Says It Has Cut Off 750,000 Australian Kids Meta reported locking out more than 750,000 Facebook and Instagram accounts in Australia by the end of June under the country’s under-16 social media law. Over 500,000 of those were removed before the law even took effect. Detection relies mostly on AI scanning posts and bios for tells like birthday messages, plus user reports and blocks on re-registration. However, the post gives no count of mistaken removals or appeals, and the regulator’s early data shows under-16 usage falling only from about 86 to 81 percent. Meta wants a single age signal at the operating system or app store level, and Cochrane agrees completely. He connects it to the MHS idea from the top of the show: platforms need a standard flag to reference instead of guessing. The White House Deputizes Private Hackers Earlier this month the White House signed a National Security Presidential Memorandum that lets vetted private security firms run surveillance and disruption operations against overseas criminal groups. The Justice Department and Homeland Security hold the contracts and oversee the work. Firms need a proven track record, vetted staff, and a bond of at least $1 million, and must submit operating procedures within 60 days. Cochrane finds the measure aggressive in a good way and hopes it deters attacks on innocents. Still, he takes Kevin Beaumont’s warning seriously that the private security industry profits from ransomware existing. He compares it to the old Head and Shoulders myth: why solve the problem that drives your revenue? A Weather Satellite Watched the Eclipse Shadow Cross Europe Cochrane skips the readout on this one and simply sends listeners to ESA’s site. The MTG-I1 weather satellite captured the Moon’s shadow sweeping across Europe during the August 12 eclipse. Watching a shadow cross an entire continent, he says, was a first for him. Additionally, it leaves him excited about the research happening beyond the planet. Rivers, Deltas, and the Number 0.6 Quanta Magazine explains Hack’s law, which John Hack discovered in 1957 while measuring streams in Virginia and Maryland. A stream’s length tracks its drainage area raised to the power of 0.6, regardless of the rock underneath, and satellite data later confirmed it worldwide. Computer models in the 1990s showed why. Channels that capture extra runoff cut deeper and steal from their neighbors until the network settles into the arrangement that wastes the least energy. Now a University of Texas Rio Grande Valley team has found the same 0.6 exponent in river deltas, which spread water out rather than gathering it. Nobody knows why yet, and Cochrane calls it a really cool read. Sugar Helped Grow the Human Brain, Too A new paper in Science, co-authored by Jennie Brand-Miller at the University of Sydney, adds a third ingredient to the story of early human brain growth. Alongside meat and cooking, natural sugars from ripe fruit and honey may have fueled it too. The brain is about two percent of body weight but burns twenty percent of resting energy. It runs on glucose, which meat and marrow barely supply and raw starch cannot release without fire. The team modeled ancestral diets from a chimp-like baseline through Homo erectus and concluded that the earliest hominins may have drawn over 65 percent of their energy from natural sugars. Cochrane stresses that it is a model, not fossils, and notes that paleoanthropologist Marina Lozano thinks the authors place widespread cooking too early. Still, he loves this kind of deep research. Retracing the steps to our own intelligence, he suggests, could hint at what it takes for intelligent life to develop at all. A Brain Rhythm That Tells Doctors Where to Aim Finally, Science Daily covered a University of Cologne study on deep brain stimulation. That is the implanted-electrode treatment that eases Parkinson’s tremors for some patients but not others. Andreas Horn’s team recorded from 50 patients using both the implanted electrodes and an external magnetic scanner. They identified a circuit between the electrode’s target and the frontal cortex that oscillates at 20 to 35 cycles per second. Stronger coupling there predicted bigger improvement after surgery, though the study, published in Brain, shows correlation rather than cause. First author Bahne Bahners hopes the finding helps tune DBS more precisely, especially for patients who have not responded well. Cochrane half-jokingly asks whether MHS might one day drive those electrodes, and he calls brain disorders the hardest thing in the body to treat. Cochrane wraps with housekeeping: become a GNC Insider at geeknewscentral.com/insider, email geeknews@gmail.com with questions or comments, subscribe to the newsletter, and grab a modern podcast app at podcastapps.com. He thanks GoDaddy for over twenty years of keeping the show on the air, promises to catch everyone next Monday, and wishes listeners a great night. The post Eyes, Hands, and a Sense of Timing #1874 appeared first on Geek News Central.

Aeon Byte Gnostic Radio
Jason Reza Jorjani on Gnostic Glitches: Escaping the Demiurge's Labyrinth

Aeon Byte Gnostic Radio

Play Episode Listen Later Aug 27, 2026 149:52


Prometheus is unbound as Jason Reza Jorjani arrives at the Virtual Alexandria to discuss his latest book, Occult Horizons. Of course, we'll take the Gnostic angle. Jason will grant us a mind-bending exploration of existence where consensus reality is unmasked as a plastic, simulated matrix governed by a machinic demiurge. We'll investigate how corporate archons and panoptic wardens deploy hyperreal signals to police our perception and trap us in scripted, bureaucratic loops. By tracking the cracks, glitches, and ontological shocks that fracture this counterfeit system, we map the initiatory process where trauma becomes an aperture to higher sight. Join us as we discard spiritual submission and reclaim the sovereign, un-spooked ego capable of hacking the cosmic CPU and authorizing its own future. Get the book: https://amzn.to/4qFrMmw More on Jason: https://jasonrezajorjani.com/ Stargazer returns: a Gnostic vampire apocalypse of false paradise, forbidden memory, the Moon Queen, and the monster who remembers. Join the Resurrection List before the gates open this Halloween: https://thegodabovegod.com/stargazer Get the fall of Sophia: https://www.patreon.com/aeonbyte/posts/fall-of-sophia-167697521?utm_medium=clipboard_copy&utm_source=copyLink&utm_campaign=postshare_creator&utm_content=join_link Get The Occult Elvis: https://amzn.to/4jnTjE4 Virtual Alexandria Academy: https://thegodabovegod.com/virtual-alexandria-academy/ Gnostic Tarot Readings: https://thegodabovegod.com/gnostic-tarot-reading/ The Gnostic Tarot: https://www.makeplayingcards.com/sell/synkrasis Homepage: https://thegodabovegod.com/ Patreon: https://www.patreon.com/aeonbyte AB Prime: https://thegodabovegod.com/members/subscription-levels/ Voice Over services: https://thegodabovegod.com/voice-talent/ Support with donation: https://buy.stripe.com/00g16Q8RK8D93mw288 Merch store: https://aeonbyte.creator-spring.com/ Equipment Wishlist: https://www.amazon.com/hz/wishlist/ls/2WEJ2CCWHALZB?&sort=default Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

airhacks.fm podcast with adam bien
Turning Back Time: Strings, Locks and Garbage Collectors in Java

airhacks.fm podcast with adam bien

Play Episode Listen Later Aug 27, 2026 60:57


An airhacks.fm conversation with Ian Rogers about: Java's early API design strengths versus verbose C++/STL manuals, the Java 7 String.substring behavior change from an O(1) view over the character array to an O(n) copy, substring performance regressions at Google, implementing a custom substring as a workaround, designing APIs without knowing future customers, the SPEC JVM98 jack parser generator throwing an exception per matched token, exceptions used for control flow violating the exceptions-are-exceptional principle, JVM benchmarks skewed toward fast exception throwing, ANTLR and JavaCC as later parser tools, class loaders pairing a type with a runtime notion, ClassNotFoundException thrown from nested jar files, the DaCapo benchmark as a test of exception-throwing speed, Effective Java advice to return interfaces such as Map instead of HashMap, why substring should have returned CharSequence, the argument that String should be an interface and CharSequence the concrete type, CharSequence length limited to int and capped at 2GB, Guava Rope as a collection of CharSequences unable to implement CharSequence because of the int length limit, signed versus unsigned sizing of arrays and strings, byte signed and char unsigned inconsistencies, StringBuilder versus StringBuffer, the race condition avoided by StringBuilder cloning its array in toString, StringBuffer passing array ownership under a lock, biased locking making uncontended locks cheap, the Attack of the Clones problem of defensive cloning, optimizing clones away in the JVM, transactional memory as an alternative to locking, copy-on-write and CopyOnWriteArrayList, Swift copy-on-write, Linux kernel read-copy-update and epochs, writing the Azul C4 garbage collector, the C4 read barrier now used in ZGC and Shenandoah, moving Azul from custom hardware read-barrier instructions to x86, the two-space invariant in concurrent copying collectors, IBM Metronome fixups versus two-space invariants, detecting same-page references by XOR of two pointers, trading computation for memory accesses in read barriers, the Transitive Corporation binary translator, Rosetta for Apple and a Sparc-to-Power translator behind IBM's planned Sun acquisition, writing Android Runtime ART, ahead-of-time compilation of Dex replacing Dalvik, disk-size constraints of AOT compilation, Hans Boehm and sticky mark bits for generational marking without moving objects, ART generational concurrent garbage collection in Android KitKat, HashMap interface dispatch replacing LinkedList iteration from JikesRVM, alphabetically sorted interfaces putting AbstractCollection first, Google Maps interface dispatch consuming half the frame time, frame-rate and pause-time gains from ART, WhatsApp broken by an unbalanced-lock bytecode obfuscator, balanced locks required for biased locking, thread safety annotations in Clang/LLVM guarding the Java heap, the mutator lock and stale pointer risks, a reader-writer lock model of the Java heap, managing risk when replacing an operating system runtime, CyanogenMod as a delivery path for ART, Project Valhalla value types, current work on the Linux perf tool and observability, OProfile origins, GPU versus CPU visibility in system tools Ian Rogers on linkedin: irogers

Techmeme Ride Home
New Macs

Techmeme Ride Home

Play Episode Listen Later Aug 25, 2026 21:21


Apple refreshed the Mac Studio with M5 Max and M5 Ultra and gave the Mac mini M6 silicon, at higher prices. OpenAI's Jalapeño chip beat Nvidia on efficiency, Perplexity went fully local, WhatsApp toughened logins, and Uber livestreamed teen rides. Links Apple updates the Mac Studio with M5 Max and M5 Ultra, with up to 4.3x faster AI performance, faster graphics, and up to 512GB of unified memory for $2,499+ (Apple Newsroom) Apple unveils a Mac mini with M6 and M5 Pro, with up to 4x faster AI performance and 2x faster graphics, for $899+ and $1,699+, with preorders today and shipping September 22 (The Verge) Apple's M6 is its first 2nm chip with a 12-core CPU and GPU, while the M5 Ultra fuses two dual-die M5 Max chips into a 36-core CPU, 80-core GPU "most powerful chip ever" (The Verge) OpenAI says its Jalapeño chip delivered 1.5x-1.9x more AI work per watt and 1.7x-3.6x lower latency than Nvidia chips across GPT-OSS, DeepSeek R1, Kimi K2.5 1T (The Verge) WhatsApp upgrades its two-step verification, letting users replace the six-digit PIN with a longer alphanumeric password, and adds support for multiple passkeys (TechCrunch) Perplexity launches Portable Computer, a local AI agent platform running fully on-device with zero token costs, starting with Nvidia DGX Spark and RTX Linux PCs (VentureBeat) Uber launches an optional safety feature allowing parents or guardians to watch a livestream of their teen's ride via the driver's front-facing phone camera (Bloomberg) Subscribe to the ad-free feed.

The Cloud Pod
369: Thirteen Billion Reasons to Hug This Face

The Cloud Pod

Play Episode Listen Later Aug 25, 2026 63:17


Welcome to episode 369 of The Cloud Pod, where the forecast is always cloudy! Justin, Ryan, and (eventually) Matt are in the studio this week to bring you all the latest news in AI and Cloud, including a new local zone in Vegas, a 20th birthday, and some OAuth news thanks to Cloudflare. There's a lot to cover, so let's get into it!  Titles we almost went with this week What Happens In Local Zones Stays Low-Latency When Git Push Comes to Scaling Shove Twenty Policies Walk Into a Role AWS Bets Big on Latency in Vegas Local Zone AWS Hits the Jackpot with New Local Zone Two Decades of Instances, Zero Midlife Crisis EC2 Turns 20, Still Refuses to Retire Happy Birthday EC2, Now With 1,200 Candles Lambda Finally Lets IAM Policies Multitask Like Adults Cloudflare’s OAuth Diet: Trimming the Permission Fat Hugging Face Squeezes Out a 13 Billion Dollar Valuation Bedrock Slashes GPT-5.6 Sol Prices, Wallets Rejoice GitHub’s Capacity Crisis Sparks Retry Storm Reckoning A big thanks to this week's sponsors: We're sponsorless! Want to get your brand, company, or service in front of a very enthusiastic group of cloud news seekers? You've come to the right place! Send us an email or hit us up on our Slack channel for more info. Follow Up 01:45 The August 17 outage, and the work ahead Update on GitHub’s August outages: root cause analysis published for the August 17 incident, which lasted nearly 8 hours and followed an earlier August 6 Actions failure. Root cause identified as a capacity failure, not a code or configuration change: a critical infrastructure component in the Central US data center failed to scale at a new traffic peak, triggering authentication failures and cascading disruption across services including Copilot, which was prolonged by a client-side retry loop. Since April, GitHub has added over 3 million CPU cores and 120 petabytes of storage, and accelerated Azure migration; Azure now handles approximately 58 percent of platform load and half of Git operations, up from 12 percent in May. Monthly commit volume has roughly doubled since April, from 1.4 billion to 2.9 billion, underscoring the scaling pressure behind both incidents and explaining, though not excusing, per GitHub, the repeated failures. Concrete remediation steps include consistent retry limits and budgets across service-to-service calls to prevent retry storms, a review of lower-priority CPU and memory alerts, and continued work isolating critical systems to reduce shared dependencies and blast radius. 03:07 Justin – “It felt a little ‘woe is me, capacity is a problem,' but it feels like more of the same lip service from them… maybe we need to rethink some core fundamentals of how Git works. Git was designed for humans… around human speed and human scale. ”  General News 14:03

What's Next Wall Street?
Apple M6 Mac Mini Local AI Bet, Trade Wars, Iran Sanctions vs Embargos

What's Next Wall Street?

Play Episode Listen Later Aug 25, 2026 39:49


Desktop AI is becoming a reality, and Apple wants the Mac sitting on your desk to replace at least some of the AI compute happening in massive data centers today. On this episode of What's Next Wall Street, we break down Apple's new Mac Mini and Mac Studio, escalating US-Canada tariffs, the campaign against Iran's oil exports, Callaway Golf's controversial Good Good Golf ad, and a potential multibillion-dollar sleep-tech IPO. First, the US-Canada trade fight is escalating. Canada is imposing 50% tariffs on more than 700 US goods as well as steel and aluminum, while President Trump is planning 50% tariffs on Canadian cars, trucks, and auto parts by January. Greg Krause looks at who actually gets hurt, including Ford, GM, Stellantis, Canadian companies, and railroads, and why the broader impact on US GDP and inflation may be relatively small. His view: tariffs work best as protection, not manipulation. Next, Washington's campaign against Iran's oil exports is being described as an "economic D-Day." The Treasury wants to isolate Iran financially and target countries helping it sell oil, including China and Russia. Greg explains the important difference between sanctions and a blockade and puts the US military presence in the region into perspective. We also look at whether disrupting Iranian oil really threatens global markets or whether buyers and producers will simply adapt. Then comes Apple. The new Mac Mini starts at $899 with Apple's M6 chip, built on a 2nm process with a 12-core CPU, 12-core GPU, and 16 dual-core Neural Engine. Apple says it can deliver up to 4x the performance of the M4 and process AI prompts up to 4.8x faster. More importantly, Apple is positioning the Mac Mini as an "always-on machine for AI agents." Instead of sending every prompt, document, or inference job to a cloud data center, developers can run LLMs locally using tools such as LM Studio and Krystallize.ai. The Mac Studio takes that idea much further. Starting at $2,499 with M5 Max and around $5,000 with M5 Ultra, configurations reach up to a 36-core CPU, 80-core GPU, 512GB of unified memory, and 1.2TB/s of memory bandwidth. Apple says M5 Ultra delivers 4.3x the AI compute performance of M3 Ultra and 10x that of M1 Ultra. Thunderbolt 5 and RDMA also make it possible to cluster Macs together. Apple says a four-system configuration can deliver roughly 3x the inference performance of a single Mac Studio. That's potentially a major shift: powerful AI models running locally, privately, and without a recurring cloud GPU bill. We also examine Callaway Golf's apology over a Good Good Golf ad in which an influencer knocks a woman down while going after a driver. Does controversy like this actually affect a public company's stock, or does Wall Street ignore it unless the company's core customers care? Finally, Aura Health, maker of smart sleep-tracking rings, is reportedly considering a roughly $3 billion IPO at a $16 billion valuation after generating about $1 billion in revenue last year. Greg looks at the opportunity and the challenge of competing with Apple Watch, Withings, and an increasingly crowded health-tracking market. Chapters:0:00 Welcome to What's Next Wall Street0:43 US-Canada Trade War Escalation3:01 Impact of Tariffs on Markets and Companies7:01 US Economic Pressure on Iran9:37 Iran Sanctions & Blockades17:09 Pro Tip: Diversify Gains17:53 Apple's New Mac Mini for Local AI19:31 Mac Studio: Desktop AI Supercomputer22:50 The Future of Local AI23:13 Callaway Golf's Controversial Ad Download your local AI for any M series Mac at Krystalize.AI now! Learn more about your ad choices. Visit megaphone.fm/adchoices

10 minutos con Sami
Nvidia acelera agentes, IBM mezcla Arm y mainframe, y AMD supera el 30%

10 minutos con Sami

Play Episode Listen Later Aug 25, 2026 5:38


Nvidia estrena Groq 3 LPX para acelerar la inferencia de agentes; DeepSeek aparece en operaciones vinculadas a ciberataques chinos; IBM prepara un procesador que combinará IBM Z y Arm; AMD supera por primera vez el 30% de los envíos de CPU x86 para PC; y la FDA autoriza un análisis de sangre para ayudar a evaluar patología asociada al alzhéimer.Puedes seguirnos en YouTube en https://youtube.com/olivernabani y puedes unirte al Discord Mashain en https://olivernabani.com/discord

Hackaday Podcast
Ep 383: QR Codes, Caving Gear, and the Old School Way to Learn Electronics

Hackaday Podcast

Play Episode Listen Later Aug 21, 2026 87:44


In this week's episode, Hackaday Editors Elliot Williams and Tom Nardi start things off by getting excited about the recently announced 2026 Retrocomputing Challenge. From there the conversation will cover efforts to improve desktop 3D printing with lasers, an expensive grill with an ESP32 controller and an open source firmware, open source tools underground, and some impressive techniques to squeeze a bit more utility out of the common QR code. You'll also hear about turning PVC pipes into flat stock, old school Radio Shack electronic kits, and VR soldering demos. Stick around to the end of the episode learn about the latest developments in over-the-counter hearing aids and the 1-bit CPU that's enjoying an unexpected fandom nearly 50 years after its release. Check out the links if you want to follow along, and as always, tell us what you think about this episode in the comments!

Microsoft Mechanics Podcast
Windows 365: Five Years In, See What's New

Microsoft Mechanics Podcast

Play Episode Listen Later Aug 21, 2026 15:04


An AI agent works inside your Windows 365 Cloud PC from a Teams chat on your phone — finding downloaded files, editing documents, and drafting emails with your laptop lid closed. Microsoft Scout, a new Autopilot agent powered by the open source OpenClaw project, runs on your own Cloud PC using your identity, your files, and your apps. Developers get a Developer Optimized Cloud PC preloaded with GitHub, VS Code, and Node.js that runs Foundry Local for token-free AI coding, and input and output protection capabilities black out the screen when anyone tries to capture a sensitive file. Bhavya Chopra, Windows 365 Partner Director, joins Jeremy Chapman, Microsoft 365 Director, to show how Windows 365 now covers every seat — end users, developers, admins, and the AI agents working alongside them. 

PC Perspective Podcast
Podcast #881 - Intel CPU Leak, Arc Linux VRAM Zip, Twitch AI Training, MOZA Racing Sim Gear Review, Ryzen Delidding + MORE!

PC Perspective Podcast

Play Episode Listen Later Aug 20, 2026 77:44


Get the fanciest Ryzen 9 without a lid, Intel appears to be looping back on their naming CPUs and Arc squeaks out more performance (sometimes), CoPilot assists in "hacking" itself, your Twitch streams might be training AI, and Epic appears to be coming to Linux.  It's the year of Linux, Am I Right(tm)? or what?   All that and way to much for this 5Lb bag again.Timestamps:00:00 Intro02:29 Patreon03:17 Food with Josh05:02 News begins - Intel might Haswell name their new CPU the 4950K06:34 You can now buy a de-lidded 9950X3D2 with a warranty07:53 VRAM compression with Arc on Linux?11:22 Microsoft's rebrand registry14:28 Microsoft working on Windows 11 context menu15:50 Intel, money motivated, may make more memory memories21:44 Twitch ai training is opt-out, because no one would opt-in26:47 Fairphone Gen 6 Plus arrives in USA28:51 (In)Security Corner40:48 Gaming Quick Hits50:10 Josh reviews MOZA CS Pro and R9 V31:03:06 Picks of the Week1:16:19 Outro ★ Support this podcast on Patreon ★

php[podcast] episodes from php[architect]
The PHP Podcast 2026.08.20

php[podcast] episodes from php[architect]

Play Episode Listen Later Aug 20, 2026 58:01


PHP Podcast – August 20, 2026 Hosts: Joe Ferguson, Sara Golemon & Holly Schilling Shirley MacLaine is the answer — but what’s the question? RAM prices went up 500%, GitHub fell over for eight hours, and the crew figured out how you can actually contribute to PHP. Glasses optional. Shirley MacLaine, Green Room Shade, and the Show’s New Normal The episode opens with Joe flustered — not because it’s a “flustered day,” but because there’s shade being thrown in the green room backstage chat. Rather than spoil the drama, the crew turns a mysterious green-room answer into a running bit: “Shirley MacLaine” is the answer, and listeners are invited to submit what the question was. It’s declared the first round of PHP Architect’s Jeopardy, complete with its own Easter egg sound cue. Joe also lays out where the show is heading. Eric and John took over the podcast the previous week to relive their glory days, and the plan is for the old guys to step back to roughly one episode a month. Joe and Sara are working on something to fill one slot, another mystery show is in the works pending signed contracts, and Alive and Kicking is confirmed still alive and kicking after a great recent episode with Derek. Joe also plugs the PHP 8.6 Beta 1 tag and Scott Keck-Warren’s PHP Community Podcast interview with release managers Daniel Scherzer and Matteo Bacotti. The RAM Apocalypse: 500% in Twelve Months The big topic of the week is the ongoing memory and chip crisis. Memory prices have climbed 500% in twelve months, and the crew commiserates about the parts they wish they’d bought bigger. Sara explains the brutal math: you can’t build a semiconductor fab fast enough — two or three years minimum — and by the time one comes online nobody knows if we’ll have overproduction or a burst bubble. Worse, Sara notes that essentially every RAM stick through the end of 2027 is already accounted for and sold to vendors. The ripple effects are everywhere. Console prices are going up instead of down late in their cycle, with next-gen consoles projected to start at $1,000 or more. The hobbyist single-board computer market is getting crushed — Pine64 announced they’re stepping back from making hardware, and a Raspberry Pi 5 that debuted cheap now runs over 100 pounds. Holly’s new PC, bought at the start of 2025, ended up shipped 3,000 kilometers the wrong way to California and is stuck awaiting import paperwork, leaving her leaning on a laptop and some now-precious Raspberry Pi zeros. The nostalgia gets thick as the crew reminisces about DIP chips, EDO RAM, Pentiums, Epson 486s, Tandy 1000s, and the TRS-80. Special venom is reserved for Packard Bell (“utter freaking trash”) and Gateway 2000’s cow-pattern branding — which Sara confirms sold better in Wisconsin than anywhere, though hay bales in a Berkeley storefront still boggle the mind. A chat comment about technicians de-soldering and reballing BGA RAM chips becoming economically viable gets a hearty “absolutely.” Memory-Aware Development and Why PHP 7 Doubled Down Bringing the RAM crisis back to PHP, Joe wishes more developers were aware of how their applications consume and release memory — a lesson he credits to learning enough C back in the day, where you have to manage memory yourself. He connects sloppy memory awareness to the N+1 query problems web developers keep tripping over. Sara drops a great deep dive: a significant reason PHP 7 was roughly twice as fast as PHP 5 was changes in the memory layout. Every variable became referenced by one fewer pointer, and while eight bytes sounds trivial, every level of indirection adds time across every single instruction and access. Sara adds the CPU-level detail — one layer of indirection can be a single instruction on most architectures, but adding a second layer can push a lookup from one instruction to three. That leads into a warm tangent about learning C to become game developers. Sara’s evergreen joke: “I’m going to be a game developer” is the programmer’s version of “I’m going to buy a bar.” Great people, brutal hours, endless competition, and the reality of hitting spacebar 400 times to figure out why you can phase through a wall. The cat-reading-the-paper “I should buy a boat” meme makes an appearance to seal it. The GitHub Outage and the Monoculture Problem Monday’s eight-hour GitHub outage hit the crew directly. Holly couldn’t use a site that only offered “log in with GitHub,” and Joe got kicked out of his CLI auth session mid-PR with no way to re-authenticate. To GitHub’s credit, they published an incident update and a follow-up blog post: a service auto-scaled so aggressively to handle network traffic that the sidecar and supporting services couldn’t keep up, bringing the whole thing down. The conversation turns to whether this is self-inflicted. Joe recalls GitHub’s pre-Microsoft, gold-standard reliability and wonders aloud how much the decline lines up with Copilot’s arrival and internal AI adoption, with uptime reportedly slipping below a single nine at points. Sara defends them somewhat — the number of actions, CPU cores, and pull requests has genuinely hockey-sticked, partly because AI has emboldened people who previously wouldn’t have opened a PR. But as Joe puts it, the call is coming from inside the house, since GitHub itself has been pushing AI. On alternatives, Joe says the least-jarring migration for PHP Architect’s clients would be self-hosting GitLab, since GitHub Actions and GitLab runners are nearly identical in syntax — Atlassian’s Bitbucket, by contrast, is a bridge too far, mostly because the entire ecosystem assumes you’re on GitHub. Sara names this the core problem: monoculture. The crew discusses package mirrors, local caches, 12-factor thinking, and Composer’s support for custom mirrors, all while remembering the PHP repo intrusion years ago that came from an unmaintained self-hosted Git server. The takeaway: owning your pipeline end-to-end is the only way an outage can’t stop you — and Joe teases spinning up a self-hosted GitLab now that “the boss” (Sara) has signed off. How to Contribute to PHP (and Handling Security Reports) The crew highlights two PHP Foundation blog posts. First, Matt Stauffer’s “How to Contribute to PHP,” adapted from a talk he gave at Atlanta PHP. It goes well beyond “learn C,” clearly separating the PHP project, the PHP ecosystem, and the Foundation, and lays out approachable on-ramps: testing pre-releases (PHP 8.6 Beta 1 is out, Beta 2 lands next week), improving documentation, and writing tests — which, spoiler, are written in PHP, not C, using PHP’s own test format that’s simple enough to learn from any single example. Other contribution paths include triaging and reviewing issues across PHP repositories — invaluable work that frees core developers from wading through AI-generated slop bug reports — and participating in internals via the well-documented mailing list process, up to and including running for release manager (8.7 managers will be needed before you know it). Sara points folks to discord.phpc.chat for the PHP Discord, with dedicated Internals and Foundation channels for anyone the mailing list intimidates. Second, Sebastian Bergmann’s “So you received a security report. Now what?” is a jump-around reference for application developers rather than a front-to-back read, walking through roughly ten steps to triage, validate, and resolve reported issues the right way. Sara shares a real-world example from mobile: a flagged package that was only exploitable on a rooted device with an actively hostile package installed alongside it — a very different risk profile than a SQL injection on an API endpoint. Cue reminiscing about writing SQL against Access databases over ODBC from PHP (and Perl) back in the 90s, and Joe’s advice for handling any security report: don’t panic, and always know where your towel is. Links from the show: PHP Tek 2027 — April 27–29, 2027 in Chicago; early bird tickets & hotel available now PHP Tek 2027 CFP Audio versions of the podcast at phparch.com Join us live on Discord at discord.phparch.com PHP Discord — discord.phpc.chat Community Corner Podcast: PHP 8.5 + 8.6 Release Manager Daniel Scherzer Memory prices climb 500% in 12 months So You Received a Security Report. Now What? How to Contribute to PHP Shirley MacLaine Host: Joe Ferguson Mastodon: @joepferguson@phpc.social PHPArch.me: @svpernova09 Sara Golemon Mastodon: @pollita@phpc.social Holly Schilling Mastodon: @TheCodeLorax@tech.lgbt Streams: Youtube Channel Twitch Connect & Hire PHP Architect Website Twitter/X Mastodon Hire PHP Developers Looking to hire PHP developers? Email support@phparch.com – Joe and the team are available for consulting, infrastructure work, Ansible playbooks, and code review. Partner This podcast is made a little better thanks to our partners Displace Infrastructure Management, Simplified Automate Kubernetes deployments across any cloud provider or bare metal with a single command. Deploy, manage, and scale your infrastructure with ease. https://displace.tech/ OurCVEs Your security posture, on autopilot with OurCVEs CodeRabbit Cut code review time & bugs in half instantly with CodeRabbit. PHP Architect Consulting Your PHP codebase deserves a partner, not a contractor PHP Architect provides long-term technical partnerships for organizations that need senior-level PHP expertise that you can depend on https://www.phparch.com/consulting/ Music Provided by Epidemic Sound https://www.epidemicsound.com/ Join Us Live Next Week Youtube Channel Got feedback? Join us on Discord at discord.phparch.com The post The PHP Podcast 2026.08.20 appeared first on PHP Architect.

Castle Super Beast
CSB385: No One's Ready For a Battle Poop

Castle Super Beast

Play Episode Listen Later Aug 19, 2026 157:29


Download MP3 | Watch Video Episode Full Timestamps: https://docs.google.com/document/d/e/2PACX-1vQXzUWeL7-kPfvD63XJEQDoLYOUTIG0ZABPETWD7Czvy9RPxhs_Jny13GyOTCka4M3iXHq_WrvW9iXl/pub Watch full episodes: https://www.youtube.com/@CastleSuperBeastArchive New Steam Controller Review Durability is Just Ammo for your Sword XenoLIES How Asterdam 1666 Just Lost Everybody Rooting For It Super Bosses: I Love This, Now Get It Away From Me NEW CASTLE SUPER BEAST "LEGACY" SHIRT & DESKMAT AVAILABLE NOW: https://www.orchideight.com/collections/castle-super-beast Visit http://drinkag1.com/SUPERBEAST to get a FREE AG1 Pro Yeti Shaker in your AG1 Pro Welcome Kit. Head to http://factormeals.com/castle50off and use code castle50off to get 50% off and 1 free breakfast item per box for a year! Exclusive $35-off Carver Mat, Aspen, and Walden frames at https://on.auraframes.com/SUPERBEAST. Promo Code SUPERBEAST Docket: (Aftermath) 1666: Amsterdam uses so much GenAI they've essentially forgot which are AI assets Twitch now uses your channel to train generative AI by default. You can opt out of some training It's on by default because if it was opt In nobody would opt in Twitch isn't sure whether Amazon has used Twitch creator VODs prior to today to train its AI models. ︀︀The opt-out will apply to your previous and future content, so it will not be used to train future models. Saber Interactive has denied renewed claims that it replaced a former Lead Writer so it could "use ChatGPT" in the development of the recently announced title, Rideshare "Stimulator." In a statement to PC Gamer, Saber also acknowledged previously undisclosed genAI use elsewhere in the title. this CEO comment is how you know with 1000% certainty Stella is right  Hey everybody. I've seen the statement by Matt Karch, the CEO of Saber Entertainment, and all of the truly horrendous things he said about me. I think it should be obvious to anyone who reads that statement that he's lashing out because he did not want people to know there was AI use in the game. Invincible VS Developer Reportedly Laid Off "Something Like 75%" of Staff Tokon got cracked in less than 2 weeks and not only does it run better than the retail version, it also works on Linux. What was even the point of Sony shooting itself in the foot with all of that DRM? Not entirely so it turns out you can terminate the Playstation SDK in task manager and the game and online will still work you just won't be able to matchmake against PS players. So the PSN login BS is just for crossplay nothing else Retail version of the game is encrypted and decrypts the game in real time in 64kb chunks on 1 CPU core. Kingdom Hearts The Series announced for Disney+ Kingdom Hearts 4 - Official Coco Showcase Trailer | D23 2026 https://www.reddit.com/r/GamingLeaksAndRumours/comments/1vrklpn/comment/p4dyhik/  

PING
DNS Cold Start

PING

Play Episode Listen Later Aug 19, 2026 42:04


In this episode of PING, APNIC Chief Scientist Geoff Huston and I discuss the behaviour of the DNS system when you have to come up from nothing: the “cold start” where no data exists from prior queries in your “cache”. This stems from a talk given by Ondřej Surý, that Geoff saw at the recent RIPE-92 Meeting held in Edinburgh, in the DNS working group. Cache, is the things you hold on to, from prior work done. It's a fundamental technique in computer science to avoid the massive disparity in speed between parts of the systems, keeping things you use a lot in “fast” memory close to the CPU, and avoiding having to go to “slow” memory or even worse disk or tape files, to find the data. The size of your cache and it's speed can have a huge influence on the speed of your program. Comparing CPU performance with and without prior cached data can be very instructive to it's benefits! -In the DNS, Cache is how you avoid having to wait for a remote system to answer over the network (with all the round-trip delay of question and answer) because you know parts of your answer from the prior queries and answers you hung onto. “Cold Start” is a well known problem in large distributed systems. The problem is not just that industrial systems like coal plants and gas turbines need time to warm up and get sufficient energy to drive the generator, in the so-called “black start” there is the problem of energising the magnetic coils associated with generation of electricity. It's distinct DC energy, used to make the generator enter the state where the rotational force from the turbine actually makes power. Without a source of power, you cannot make the coils “excite” and so the spinning generator won't actually make any electrical power. Typically, this DC voltage comes from an independent source like a small diesel generator, or a source of electrical power which hasn't been affected by the outage. If you contract to supply this contingency power, failure to be available is a major issue. The DNS isn't an electrical generation network, but it does have this kind of complex dependency, mapping from DNS fully qualified name, to the specific IP addresses bound to the name. The “D” part of DNS stands for “Domain” and the elements of the name in a sequence form “domains” of control (or at least potentially do) separated by the dots in the name which may have a distinct “name server” detailing how things are named under-neath that domain boundary. So a single question “what is www.potaroo.net” has the potential to invoke at least 3 if not more questions, for each “domain” in the name. When you have a cache. The cache of known name-to-address pairs includes the data you need to know the addresses of the nameservers for each domain you have seen, which can tell you the name-to-address pairs for the other things you are asking: either directly or indirectly by passing you on to another system. But if you have no cache, when you ask (for example) the nameservers for .NET and are told they are A.NET and B.NET you are no wiser: how do you find these hosts, when you don't know how to find .NET? Geoff has been chasing down what he sees exploring this “cold start” behaviour, and what it tells us about how people chose to name their hosts and services, how they provide the “name servers” behind these hosts and services, and how intruding intermediaries, or even trying to work around the risks of a cold start can increase the query burden on the clients worldwide.

The Full Nerd
Episode 412: Intel Reveals Upcoming CPU Plans, 16GB GPU Market Sales Data & More

The Full Nerd

Play Episode Listen Later Aug 18, 2026 111:21


Join The Full Nerd gang as they offer level-headed takes about the latest PC building news. In this episode the gang is joined by Jake Roach from Tom's Hardware to chat about his recent interview with Intel which reveals the companies plans for future CPU launches, as well as looking at current market share numbers for GPUs with 16GB of VRAM, and more. And of course we answer questions live! Timecodes: (00:00:00) - Intro (00:05:42) - Intel CPU plans (01:00:01) - GPU market share (01:19:39) - Q&A Links: - Nova Lake on desktop: https://www.tomshardware.com/pc-components/cpus/intel-says-it-will-launch-new-core-with-nova-lake-on-desktop-first-not-in-data-center-vp-robert-hallock-hopes-enthusiasts-do-the-math-compared-to-amd - DDR4 Raptor Lake: https://www.tomshardware.com/pc-components/cpus/raptor-lake-is-a-core-part-of-the-portfolio-for-years-to-come-says-intel-theres-been-a-sudden-inrush-of-demand-for-lga-1700-chips-due-to-ddr5-prices - GPU sales data: https://wccftech.com/gpu-sales-data-by-german-retailer-shows-that-16-gb-gpus-still-lead-the-market-despite-being-way-more-expensive-than-ever/ Join the PC related discussions and ask us questions on Discord: https://discord.gg/UWhjwg778a Follow the crew on X and Bluesky: @AdamPMurray @BradChacos @MorphingBall Music by Our Ghosts: https://ourghosts.bandcamp.com/ Some links may contain affiliate links, which means if you buy something PCWorld may receive a small commission. ============= Follow PCWorld: Website: http://www.pcworld.com Newsletter: http://www.pcworld.com/newsletters ============= Learn more about your ad choices. Visit megaphone.fm/adchoices

Telecom Reseller
Stealthium and Tenstorrent Target the AI Infrastructure Security Blind Spot, Podcast

Telecom Reseller

Play Episode Listen Later Aug 18, 2026 16:34


By Doug Green “If you can't see something, you can't secure it.” AI security discussions often focus on models, applications and data. But Chris Hosking, GTM Advisor at Stealthium, says organizations may be overlooking a critical layer: the accelerated computing infrastructure where AI actually runs. In this Telecom Reseller podcast, Hosking discusses Stealthium's partnership with Tenstorrent and a broader emerging challenge for enterprises, cloud providers and operators — securing GPUs and other AI accelerators that traditional security tools often cannot fully observe. Stealthium describes itself as a runtime observability and security company for AI infrastructure. Its goal is to give organizations visibility directly into accelerated compute environments, below the point where many conventional security tools stop. “There's an architectural gap here when it comes to controls,” Hosking says. Existing security instrumentation may provide extensive visibility at the endpoint, CPU and driver layers, but organizations increasingly need to understand what is happening inside the accelerator runtime itself. That becomes especially important as GPUs move from specialized computing resources to critical enterprise infrastructure. Hosking describes GPUs as environments that increasingly contain an organization's “crown jewels” — including sensitive model weights, proprietary AI workloads and potentially valuable customer or corporate data. And GPUs are not simply black boxes performing calculations. They are programmable compute environments with memory, network access and the potential for persistence. That means they also create new attack surfaces. Stealthium's recently announced partnership with Tenstorrent is designed to address that challenge by bringing runtime observability and security controls deeper into Tenstorrent's AI infrastructure. Hosking points to Tenstorrent's open, full-stack architecture as particularly important because it enables security and observability to be incorporated into the infrastructure rather than added later. The discussion also explores the security risks associated with shared GPU infrastructure. Hosking cites vulnerabilities such as Januscape as an example of how weaknesses involving virtualization and isolation can become particularly serious when multiple customers share physical accelerated-compute infrastructure. Customers need more than an architecture diagram saying workloads are isolated, he argues — they increasingly need evidence that the isolation is actually working. That leads to several basic questions AI infrastructure operators should be able to answer: What actually ran on the infrastructure? Was my workload isolated? Am I receiving all the compute capacity I paid for? Is another workload consuming resources? Is something behaving abnormally? Can I prove that the infrastructure is operating securely? Runtime visibility can therefore become more than a cybersecurity tool. It can also help organizations identify resource hijacking, inefficient GPU utilization and expensive compute capacity being consumed without operators realizing it. The issue takes on additional importance as enterprises pursue private and sovereign AI strategies. Organizations may own or control their infrastructure, but Hosking argues that sovereignty ultimately requires the ability to verify what is happening inside that infrastructure. Over the next six to twelve months, he expects security attention to move increasingly toward AI hardware and accelerated computing. As GPU adoption grows, vulnerabilities and attacks targeting these environments are likely to receive far greater scrutiny. For enterprises, service providers and cloud operators, the takeaway is straightforward: AI infrastructure is becoming too important — and too privileged — to remain a security blind spot. “If you can't see something, you can't secure it,” Hosking says. Learn more at Stealthium.io.

Additive Snack
AI Is Getting Hot: Can Additive Manufacturing Help Cool It Down?

Additive Snack

Play Episode Listen Later Aug 18, 2026 49:46


Host Fabian Alefeld interviews Professor PS Lee, Head of National University of Singapore Mechanical Engineering & Program Director of STDCT, and founder of CoolestDC, about heat as a key constraint in AI infrastructure as data centers shift from CPU-heavy to GPU-dense systems with higher power density, hotspots, and bursty workloads. Lee explains why cooling must be treated as an integrated “chip to grid” system spanning cold plates, racks, CDUs, piping, controls, and operations, and discusses single-phase versus two-phase direct-to-chip liquid cooling and the added need for condensation in closed loops.  Professor Lee describes how additive manufacturing (AM) enables complex fin structures and unibody cold plates that reduce leakage risk, lower junction temperatures (~10°C), cut pumping power (60–70%+), and reduce material use, while requiring TCO evaluation. At NUS's live testbed (22 liquid-cooled racks, ~500 kW), AM cold plates are being deployed and show promising thermal, power, and compute-performance improvements; future work includes two-phase cooling and broader system-level heat rejection, warm-water operation, and waste-heat reuse, plus a discussion of challenges and potential AM roles in space-based data centers. 02:13 Meet Professor Lee 05:26 Chip To Grid Cooling 10:22 Liquid Cooling Maturity 14:15 Two Phase Explained 18:12 Cold Plate Tradeoffs 20:40 AM Unibody Cold Plates 24:41 Benchmarking AM Gains 30:18 AM For CDUs 34:13 Inside STDC Testbed 38:09 Reliability And Leakage 39:00 Next Phase Roadmap 42:31 Orbital Data Centers

Unplugged: An IIoT Podcast
Native PLC Communication & Why OPC-UA Is Wasting Your CPU

Unplugged: An IIoT Podcast

Play Episode Listen Later Aug 18, 2026 61:45


Every IIoT project starts the same way: you need data out of a PLC. The standard answer is OPC-UA — but what if that protocol is overloading your PLCs, eating bandwidth, and costing you thousands in upsized hardware?In this episode, Christofer Dutz — creator of Apache PLC4X and CEO of ToddySoft GmbH — joins Phil Seboa and Ed Fuentes to explain why speaking the native tongue of your PLCs is faster, cheaper, and more efficient than wrapping everything in OPC-UA. Chris shares how PLC4X pulled 2,600 data points every 200 milliseconds at just 2–3% CPU overhead — compared to OPC-UA choking at 200 data points every two seconds on the same hardware.The conversation covers the origin of PLC4X (a "JDBC for industrial automation"), the not-invented-here syndrome that created today's protocol mess, why open-source adoption in OT is still painfully slow, the role of Apache IoTDB as a time-series storage engine handling 1.5 billion insertions per second, and Chris's new commercial venture ToddySoft — bringing production-grade Rust-based drivers to platforms like Inductive Automation's Ignition.Whether you're an automation engineer drowning in protocol adapters or an IT architect trying to bridge the OT gap, this episode is packed with practical insight from someone who has spent a decade making PLCs talk.Connect with Christofer on LinkedIn: linkedin.com/in/christoferdutzConnect with Phil on LinkedIn: linkedin.com/in/phil-seboaConnect with Ed on LinkedIn: linkedin.com/in/ed-fuentesSign up for the Unplugged newsletter: unpluggediiot.com/newsletterLearn more about ToddySoft: toddysoft.comLearn more about Apache PLC4X: plc4x.apache.org

The CyberWire
Please hold while we decide.

The CyberWire

Play Episode Listen Later Aug 17, 2026 28:10


Internal policy conflicts hamper U.S. military AI leadership. Clop claims GE, Philips and Shell. Attackers actively probe internet-facing GeoServer instances. “The Hatman” offers millions of alleged employee records for sale. ETSI begins the approval process for European cyber standards. Microsoft is still working on a patch for the ShieldBreak vulnerability. Autonomous AI systems create CPU bottlenecks. Monday business briefing. Our guest is Nick Warner, CEO at Neo.ai, on the shifting landscape around AI and agentic security. AI agents kneecap each other with self-replicating malware. Remember to leave us a 5-star rating and review in your favorite podcast app. Miss an episode? Sign-up for our daily intelligence roundup, Daily Briefing, and you'll never miss a beat. And be sure to follow CyberWire Daily on LinkedIn. CyberWire Guest On today's Industry Voices segment, we are joined by Nick Warner, Neo.ai's CEO, discussing the shifting landscape around AI and agentic security. If you enjoyed this conversation, be sure to check out the full interview here. Selected Reading The U.S. Military Wants A.I. Dominance. Feuds and China May Thwart It. (The New York Times) Philips and GE investigating Clop ransomware data theft claims (Bleeping Computer) Attackers Probe Critical GeoServer SQL Injection Vulnerability (Hack Read) Crook hawks millions of records allegedly plundered from corporate Azure tenants (The Register) ETSI Proposes 17 Cybersecurity Standards to Support EU CRA (Infosecurity Magazine) Microsoft working on Defender patch for ShieldBreak zero-day (Bleeping Computer) Agentic AI Crunch Creates CPU Comeback (IEEE Spectrum) Corma raises $60 million in seed funding. (N2K) Conflicting Test Goals Pushed Claude Agents to Deploy Self-Replicating Malware (SecurityWeek) Share your feedback. What do you think about CyberWire Daily? Please take a few minutes to share your thoughts with us by completing our brief listener survey. Thank you for helping us continue to improve our show. Want to hear your company in the show? N2K CyberWire helps you reach the industry's most influential leaders and operators, while building visibility, authority, and connectivity across the cybersecurity community. Learn more at sponsor.thecyberwire.com. The CyberWire is a production of N2K Networks, your source for strategic workforce intelligence. © N2K Networks, Inc.

SANS Internet Stormcenter Daily Network/Cyber Security and Information Security Stormcast
SANS Stormcast Friday, August 14th, 2026: AI vs. Honeypot Data; CPU Bugs; GeoServer 0-Day; Windows USB Driver Confusion

SANS Internet Stormcenter Daily Network/Cyber Security and Information Security Stormcast

Play Episode Listen Later Aug 14, 2026 7:23


Using Gemma4 with Ollama - Testing File Hash Analysis and Recommendations with AI https://isc.sans.edu/diary/Using%20Gemma4%20with%20Ollama%20-%20Testing%20File%20Hash%20Analysis%20and%20Recommendations%20with%20AI/33242 CPU Privilege Escalation https://github.com/xoreaxeaxeax/smiiiiiiiiiiiiiiii https://github.com/xoreaxeaxeax/skitter-creek-bath-salts GeoServer Vulnerability https://x.com/q1uf3ng/status/2087490992723407096 Windows USB Driver Vulnerability https://x.com/0xedh/status/2085842285481062887 My Upcoming Classes https://www.sans.org/profiles/dr-johannes-ullrich

Geekoholics Anonymous Video Game Podcast
Another Halo on the Battlefield 6 while Raiding with the Splatoon – Geekoholics Anonymous Video Game Podcast 543

Geekoholics Anonymous Video Game Podcast

Play Episode Listen Later Aug 14, 2026 146:10


BS Section and House Keeping Discord Server geekoholics.com/discord/ Hosts - Dayne, Joe Whatcha Been Playing? Halo Campaign Evolved - Completed Battlefield 6 More Splatoon Raiders News: Cross Platform / PC / Misc Supermassive Games is laying off up to 75 staff 'to ensure sustainability' Fortnite confirms that Sonic is coming to its next season, Override Take-Two CEO Strauss Zelnick says there's currently no interest in selling the company to Netflix or anyone else. The $100 Ultimate Edition of GTA 6 is selling more pre-orders than the $80 Standard Edition, Take-Two CEO says Halloween The Game has been banned in Australia and New Zealand because players can smoke weed in it Hello Games CEO Sean Murray thanks players on No Man's Sky's 10th anniversary New free Quake campaign from MachineGames available now for the 30th Anniversary Smilegate will rescue Lost Ark's western servers, as Amazon ditches support while retreating from MMO development Marvel Tōkon's dire PC performance improved massively by new patch, drastically reducing CPU use and stutters Nintendo  Aliens: Fireteam Elite's cloud version on Switch has been shut down, making it unplayable for those who bought it Crimson Desert getting a Switch 2 release, and Pearl Abyss sees "significant potential" for multiplayer, with "live service operation models" mentioned PlayStation PlayStation reveals limited edition Wolverine PS5 console, controllers and console covers Big Walk has sold more than one million copies, not counting PlayStation Plus downloads Xbox Xbox promises that the causes of the recent outage, which rendered even game discs unplayable, have been fixed Limited Edition XBOX Wireless Controllers PSA's: Epic Games Store Freebies: Caravan Sandwitch Free 4 All The Odyssey IMAX 70mm Ghost In The Shell on prime Intro/Outro Music: Geeky Beat – @johnsbernardo Help support the show: - Subscribe to our Twitch channel http://twitch.tv/geekoholics - Please review the show (bit.ly/geekoholics) on Apple Music, Apple Podcasts and to share with your friends. Reviews help us reach more listeners, and the feedback helps us to produce a better show. Join our Discord server: CLICK HERE

Ethereum Cat Herders Podcast
Consensus Layer Meeting 184 [2026-08-06] | ACDC 184

Ethereum Cat Herders Podcast

Play Episode Listen Later Aug 14, 2026 83:52


The conversation covers various topics related to DevNet updates, beacon API PR, and distributed blob space reconstruction. It also discusses the implementation of quick slots and the proposal for tapered issuance bond. The conversation covers the presentation of EIP 8363 called Tappered Issuance Burn, which aims to remove the issuance mechanisms incentive for stake growth once the staking ratio exceeds 50%. The proposal's primary goal is to prevent 50% of the supply from being at stake with the current trend.TakeawaysDevNet updates and discussions on beacon API PR are crucial for the Ethereum network's development.The implementation of distributed blob space reconstruction aims to distribute the reconstruction process and reduce CPU load on supernodes.The proposal for tapered issuance bond is a significant topic that requires further discussion and consideration. Tappered Issuance Burn aims to prevent 50% of the supply from being at stake.The proposal seeks to remove the issuance mechanisms incentive for stake growth once the staking ratio exceeds 50%.Chapters00:00 Introduction and DevNet Updates09:12 Beacon API PR and DevNet Discussion23:35 Quick Slots Implementation43:31 Distributed Blob Space Reconstruction52:33 Tapered Issuance Bond57:59 Introduction to EIP 836301:03:38 Community Feedback and Consultation01:13:26 Discussion on EIP Retirement and Balance Sunset

Startup Project
Bringing Robotics for Electronics Manufacturing & AI Infrastructure | Bright Machines Founder

Startup Project

Play Episode Listen Later Aug 13, 2026 48:07


Startup Project sits down with Sviat, CEO of Bright Machines, to unpack how the company is using software-first robotics to manufacture complex electronics closer to where they're deployed. The conversation focuses on why AI infrastructure is a strategic category, how Bright Machines differs from traditional contract manufacturing, and what onshoring really means for speed, quality, and security.Key Topics:In this episode, Sviat explains that Bright Machines is focused on AI infrastructure, specifically the electronics that go inside modern data centers, including compute nodes, storage, and racks.He traces the company's thesis back to a broader idea: use software and robotics to manufacture electronics anywhere, then narrow that focus to the data center market as demand became clearer.The discussion breaks down the market stack, from chip designers like NVIDIA and AMD, to ODMs, OEMs, hyperscalers, and contract manufacturers.Sviat shares why data center hardware became the right bet before ChatGPT accelerated the market: the products are expensive, strategically important, and driven by quality and throughput more than labor cost alone.The show compares traditional assembly lines with Bright Machines' approach, which uses more robotics, sensors, cameras, traceability, and humans in the loop where automation does not make sense.Sviat explains how Bright Machines starts with design, using Bright Designer to simulate and improve manufacturability before lines are built, which helps reduce bottlenecks and improve automation over time.He says the company's main differentiator is its software platform, which orchestrates the line, powers smart skills for navigation and inspection, collects data, and feeds insights back into design.The conversation covers line flexibility, including how much can be reused when switching between CPU, GPU, or different accelerator-based server designs, and when end-of-arm tooling must change.Sviat says Bright Machines is growing rapidly, expects more than 3x growth this year, and can produce high volumes from a small number of sites because of robotics efficiency.The episode closes on the broader case for onshoring AI infrastructure manufacturing in the US: security, time to market, quality, and a labor shortage that makes robotics necessary.Timestamps:06:39 - The market stack: chip designers, ODMs, OEMs, hyperscalers, and CMs09:07 - Why Foxconn, Jabil, and similar contract manufacturers matter10:04 - Why large factories still rely on massive manual labor12:20 - Why data centers are different from cheap consumer electronics13:49 - Security, strategic sectors, and why AI infrastructure belongs onshore16:26 - The first Bright Machines product: CPU compute servers for a hyperscaler17:58 - How the line works: modular stations, yields, and automation levels19:26 - Bright Designer and design-for-manufacturing feedback loops21:20 - Robots, sensors, traceability, and humans in the loop22:19 - Why time to market matters as much as cost23:31 - Yield and throughput: 98% line-level yields and up to 2x throughput25:25 - The Bright Robotic Cell and how the assembly line is structured27:35 - Reusability across products and when tooling changes are needed30:31 - Manufacturing as a service, not repair or field service31:24 - Growth, gigawatt-scale capacity, and output from a single site33:00 - Why current hyperscaler capex is not expected to slow near term34:45 - The bottlenecks before deployment: chips, components, power, permits36:54 - Bright Machines' three pillars: platform, data layer, and Bright Designer39:15 - Why humanoid robotics is exciting but not ready for industrial use41:16 - Where LLMs and newer AI tools can help the robotics workflow43:57 - The overlooked advantages of onshoring manufacturing in the US45:59 - What Bright Machines could build next: more complex electronics and future AI devices

The New Stack Podcast
Why CPUs still matter in the age of AI agents

The New Stack Podcast

Play Episode Listen Later Aug 11, 2026 26:18


As AI evolves from conversational chatbots to autonomous agents, CPUs are becoming an increasingly important part of the infrastructure equation. In this episode, The New Stack speaks with Bhumik Patel of Arm and Mo Farhat of Google about how CPUs act as an “air traffic controller” for agentic workloads, handling orchestration, data preparation, semantic search, vector databases, code execution and API calls alongside GPUs and TPUs. Smaller AI models, including summarizers and evaluators, can also run effectively on CPUs for specialized tasks.  As agents increasingly generate and execute code, secure sandboxing becomes critical. Google's gVisor and GKE Agent Sandbox provide isolation and scalable environments, with the latter supporting up to 300 sandboxes per second per cluster. The discussion also explores efficiency and cost, with Google highlighting Axion's price-performance and energy-efficiency advantages across different workload types. Ultimately, the shift toward agentic AI is creating a more diverse compute environment where CPUs, GPUs and TPUs each play complementary roles in delivering scalable, efficient AI applications. Learn more from The New Stack around the latest in CPUs in the world of AI agents:  AI Agents Will Eat Enterprise Software, Just Not in One Bite  How to ground AI agents in accurate, context-rich data  Join our community of newsletter subscribers to stay on top of the news and at the top of your game. 

The Information's 411
SpaceX Nears $60B Cursor Acquisition, Meta's Open Source Model Family, Microsoft's AI Chip Ramp Up

The Information's 411

Play Episode Listen Later Aug 10, 2026 38:04


Amp Co-Founder and CEO Quinn Slack talks with guest TITV Host Stephanie Palazzolo about Meta's return to open-source AI models. We also talk with The Information's Grace Kay about SpaceX closing its $60B acquisition of Cursor, Aaron Holmes about Microsoft ramping production of homegrown AI chips, and Catherine Perloff about AWS telling engineers to cut CPU waste.Articles discussed on this episode: https://www.theinformation.com/articles/microsofts-homegrown-ai-chip-effort-shows-signs-life-slow-starthttps://www.theinformation.com/articles/cursor-maps-branding-changes-spacex-acquisition-nearshttps://www.theinformation.com/articles/aws-tells-engineers-cut-cpu-waste-amid-crunchSubscribe: YouTube: https://www.youtube.com/@theinformation The Information: https://www.theinformation.com/subscribe_hSign up for the AI Agenda newsletter: https://www.theinformation.com/features/ai-agendaTITV airs weekdays on YouTube, X and LinkedIn at 10AM PT / 1PM ET. Or check us out wherever you get your podcasts.Follow us:X: https://x.com/theinformationIG: https://www.instagram.com/theinformation/TikTok: https://www.tiktok.com/@titv.theinformationLinkedIn: https://www.linkedin.com/company/theinformation/Chapters:00:00 - Introduction02:08 - SpaceX Nears $60B Acquisition of Cursor06:55 - Meta Bets Big on Open-Source Models20:25 - Microsoft Ramps Up Homegrown AI Chips28:45 - AWS Tells Engineers to Cut CPU Waste

The Circuit
EP 187: Inside the Future of Memory Summit & AMD's "Boring" Bullishness

The Circuit

Play Episode Listen Later Aug 9, 2026 40:43


n this episode of The Circuit, Ben and Jay unpack a busy week of tech earnings and the buzzing Future of Memory Summit. The hosts kick things off by analyzing AMD's latest earnings, discussing why the street had a mixed reaction to a quarter that demonstrated solid, reliable execution and surprisingly bullish signals for data center CPU demand.Ben then reports back from the Future of Memory Summit, where the convention center was completely packed—a clear sign of the times for the booming memory industry. He explains why the ecosystem around CXL (Compute Express Link) is finally maturing, driven by the rise of rack-scale AI architectures and the urgent need for hyperscalers to reuse vast amounts of legacy DDR4 memory for inference workloads. Finally, the duo debates the trajectory of the current memory supercycle—touching on the durability of long-term agreements (LTAs)—and highlights why SiTime is perfectly positioned to benefit from the growing need for precision timing and synchronization in AI data centers

The Investing Podcast
SpaceX Picks Nvidia Over AMD: SPCX -11%, AMD -9%, ANET +14% | August 5, 2026 – Morning Market Briefing

The Investing Podcast

Play Episode Listen Later Aug 5, 2026 17:43


Andrew, Ben, and Tom discuss SpaceX falling 11% despite laying out an ambitious AI and Starlink roadmap including Grok 5 trained on internal SpaceX data by year-end, 1.4GW of current compute growing toward 10GW by end of 2027, deploying NVL72-designed data centers on the ground rather than in space, boots on the moon by 2028, and the goal of scaling launch cadence to one per day in 2027, SpaceX's decision to go exclusively with Nvidia sending AMD down 9% despite good results and CPU acceleration while Arista Networks jumped 14% on accelerating networking demand tied to Nvidia, CVS raising guidance and re-adding Zepbound to its formulary while publicly backing Eli Lilly with expanded GLP-1 support, Lilly beating on stronger-than-expected GLP-1 pricing, Disney rising 4% on parks strength and a TikTok short-form video partnership, Uber slipping despite record first-time user additions, and the JOLTS report showing softer job openings but likely not a signal Warsh will weigh.Join our live YouTube stream Monday through Friday at 8:30 AM EST:http://www.youtube.com/@TheMorningMarketBriefingPlease see disclosures:https://www.narwhal.com/disclosure

Latent Space: The AI Engineer Podcast — CodeGen, Agents, Computer Vision, Data Science, AI UX and all things Software 3.0

Watch the full episode on YouTube:We first covered Baseten last year when DeepSeek mania was at peak hype. Now they have raised a monster $13B round and become one of the new cohort of AI Infra decacorns that are (with Nvidia, Intel, and the semis complex) chief beneficiaries of the Inference Inflection. We return to Baseten at the peak of the 2026 edition of Open Weights debate. Ali has published a viral breakdown of Kimi K3:And since you last saw him, Philip has spoken at AI Engineer and written the definitive book on Inference Engineering spotted all over SF:Three years ago, inference engineering barely existed as a category.Today, it is one of the most critical disciplines in AI. Inference engineering inherently tackles a different question than standard model training: “How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?” Focusing on these creates an entirely new optimization problem.In one recent GLM-5.2 experiment, quantizing more of the model actually preserved its benchmark quality while increasing throughput by 20%, because the errors introduced in different layers could cancel each other out.Inference is no longer just the final step after training. It is becoming its own engineering discipline, with its own research problems, infrastructure, and increasingly specialized roles.In this episode, Baseten's Philip Kiely and Ali Taha join swyx and Vibhu to explain what actually happens after a new open model is released and what it takes to turn “we generated a token” into a fast, reliable, production-ready API.We go deep on cache-aware routing, disaggregated prefill and decode, quantization, speculative decoding, KV-cache movement, model parallelism, GPU kernels, and the race to make frontier models up to 10× faster. Philip and Ali explain why inference optimizations can still produce gains of 20%, 100%, or even 200%; how quantization errors can cancel one another out; why identical weights can behave differently across clusters; and how Baseten grafted a Kimi vision encoder onto GLM-5.2 without changing the underlying language model.The conversation then expands beyond LLMs into NVIDIA Dynamo, mega kernels, Rubin, AI-specific chips, local inference, video generation, diffusion versus autoregressive models, and the enormous compute barrier to generating coherent long-form video. Finally, we explore the convergence of training and inference, continual learning through persistent KV cache, and the emerging loop where models help optimize the infrastructure that runs them.We discuss:* What happens when a 200,000-token request enters an inference system* Cache-aware routing and reusing previously computed KV cache* Why prefill and decode are increasingly handled by different GPUs* When dedicated deployments become cheaper and more reliable than shared APIs* How speculative decoding uses a smaller model to accelerate a larger one* Tool calling, structured outputs, and what LLMs actually do* What it takes to support a new open model on day zero* Grafting Kimi's vision encoder onto GLM-5.2* Retrofitting inefficient model layers with components from other architectures* Why models sometimes collapse into repeating the same token* How hardware, kernels, and race conditions create nondeterministic failures* Preserving model fidelity while making inference faster* How quantization errors can cancel each other out* Why inference optimizations still deliver gains of 20%, 100%, and 200%* How optimized serving can make a model up to 10× faster* NVIDIA Dynamo, KV-aware routing, and distributed model serving* Speculative decoding the speculative decoder* Why local AI is about making models less dumb while data-center AI is about making them less slow* Tensor, expert, and pipeline parallelism across GPUs* Hardware-aware model design, auto-tuning, and the case against mega kernels* Rubin and why inference is becoming a systems problem* Whether modern GPUs are evolving into programmable AI ASICs* Why enormous models like Kimi K3 require GB300-class hardware* Why open-source video generation still trails Veo, Kling, and other closed models* The quadratic attention bottleneck behind long-form AI video* Autoregressive video, real-time generation, and compounding quality drift* Why future video systems may combine autoregressive and diffusion architectures* Training for inference and inference for training* Continuous post-training, deployment, evaluation, and improvement loops* How GLM-5.2 helped optimize the kernels serving GLM-5.2 itself* Why faster networking could unlock dramatically faster decoding* Continual learning, KV-cache compaction, and persistent model memoryShow Notes* How to build a day-0 API for Kimi K3* 22580: From GPT2 to Kimi3, ExplainedPhilip Kiely* LinkedIn: https://www.linkedin.com/in/philipkiely* X: https://x.com/philipkiely* Inference Engineering: https://www.baseten.co/inference-engineering/Ali Taha* LinkedIn: https://www.linkedin.com/in/aliestaha/* X: https://x.com/waterloointernTimestamps00:00:00 Introduction and the 200K-Token Prompt00:03:18 Dedicated Deployments, Speculative Decoding, and Tool Calling00:11:26 Launching Production-Ready Open Models00:19:06 Model Retrofits, Failure Modes, and Nondeterminism00:28:22 Quantization and Canceling Errors00:32:15 The Race to 10× Faster Inference00:40:48 Dynamo, Speculation, and Local vs. Data-Center AI00:50:18 Model Parallelism, Auto-Tuning, and Mega Kernels01:00:55 Rubin, GPUs vs. ASICs, and Custom AI Chips01:10:03 Giant Models and the Limits of GPU Memory01:12:42 AI Video, Quadratic Attention, and Autoregressive Generation01:21:47 Audio, Images, and Diffusion Models01:27:32 Training, Self-Optimizing Models, and Continual Learning01:40:06 Closing ThoughtsTranscriptIntroduction: Baseten, Waterloo Intern, and Inference EngineeringSwyx [00:00:00]: Okay, we're here in the studio with Philip, old friend from Inference Engineering, the book, as well as Baseten and everything that you've done, you and I have done before, as well as Ali. Welcome.Ali [00:00:15]: Pleasure to meet you.Swyx [00:00:15]: Waterloo intern.Ali [00:00:16]: Waterloo intern, always.Swyx [00:00:17]: When did you get “Waterloo intern” as a handle?Ali [00:00:19]: As a handle? Oh.Ali [00:00:20]: I think the rebranding happened mid-March. When I saw it was open, I was like, “I have to take it. Up for grabs.”Philip [00:00:26]: The problem is that Ali is really good at his job and is not gonna be an intern much longer.Philip [00:00:30]: So we have to figure out who's gonna get the handle.Ali [00:00:33]: Well, I'll pass the torch over to the next intern.Swyx [00:00:34]: Oh, okay. It can be, like, you just pass it to another Waterloo grad.Ali [00:00:37]: To another Waterloo intern. No, bruh.Philip [00:00:39]: Yeah.Ali [00:00:39]: Intern.Swyx [00:00:40]: Intern, yeah.Ali [00:00:40]: And no.Philip [00:00:41]: You gotta get an intern from Waterloo.Ali [00:00:42]: Yeah, I've gotta get an intern from Waterloo.Swyx [00:00:44]: Right.Ali [00:00:44]: But they have to follow the path.Swyx [00:00:45]: Oh, it could, but it could come from Baseten, so it's like whoever Baseten gets from Waterloo.Ali [00:00:48]: Right.Swyx [00:00:49]: Has the title of Waterloo.Ali [00:00:50]: It stays in the ecosystem.Philip [00:00:51]: Exactly.Ali [00:00:52]: Halfway through the internship, you either get it or you're out.Philip [00:00:55]: You should also do, like, a big graduation ceremony where you change the handle.Ali [00:00:59]: Just say it.Philip [00:00:59]: For everybody.Swyx [00:01:00]: You guys are good at ceremonies, clearly. We had a nice launch of the book, very successful. But before we get into all that, I wanna start off with a fun question for you. Okay, you're an expert inference engineer. What happens when I send a long query, say two hundred thousand tokens into Baseten's inference? What's the process of query through GPU model routing, balancing, all that? What is all the stuff that we don't think about?Long Context Requests, KV Cache, and Cache-Aware RoutingPhilip [00:01:26]: With a long query specifically, the first thing that I'm gonna ask is, “Have you sent me this query before, or at least part of it?” and I really hope you have, because it's gonna be a lot easier for me and a lot cheaper for you. So the first thing that we're gonna look at is some cache-aware routing, where we're going to see, we probably have a number of instances, a number of replicas up serving whatever model you're hitting. We want to send this one to something with, number one, available prefill workers, and number two, ideally some cached input already there so that we can skip prefill on at least part of these two hundred thousand tokens. If you're doing two hundred thousand tokens, it's probably coding or a multi-turn agent or something where you would expect to have that cached. If you don't, we're gonna have to send it to a prefill worker. We've at least on certain models disaggregated prefill and decode, so you're going to have one set of GPUs that's solely going to process the input, create the KV cache, and get you your first token, and then that's going to be passed over to a separate set of GPUs, which is going to run decode. We're going to iteratively make those tokens. We're probably going to have some speculator model in front of that. I'm going to assume that you're doing coding, and because of that, our speculator model, which assumes you're doing coding, is gonna have a high draft token acceptance rate. If I'm wrong and you're asking me to summarize every Harry Potter book, it's gonna be slower. And then we stream that output to you and account for it, charge you, a couple of pennies and say, “Hey, would you like to send another one?”Swyx [00:03:04]: Except Baseten doesn't charge by pennies.Philip [00:03:07]: Well, yeah, we charge. I'm assuming that we're talking about the public model APIs. If you are setting up a dedicated deployment, then yeah, it's not pennies.Public APIs vs. Dedicated DeploymentsSwyx [00:03:18]: Yeah, one of the key differentiators when I was talking with Baseten initially was that people who want very high volume just need to rent by the box, ‘cause then it's up to you to figure out how to saturate the box.Ali [00:03:31]: And more often than not, it's, like, way cheaper if you're pushing, like, millions of tokens per hour, if you just pay per hour instead of pay per token.Philip [00:03:37]: Yeah, they do. I think that we've increasingly seen a lot of demand for the pay per token APIs, just because everyone wants to try open models, and then once they find a use case that's really sticky, then they move over to dedicated.Swyx [00:03:51]: Is there a best practice on when it's time to swap over?Philip [00:03:54]: Couple reasons. Yeah, reliability, that's a big one, right?Ali [00:03:57]: Like, if they have a very specific use case, they want you to train something specifically for them, like they want their own spec dec, for instance, for their own traffic.Swyx [00:04:04]: Spec dec is speculative decoding.Speculative Decoding and Custom SpeculatorsAli [00:04:05]: Speculative decoding, yeah.Swyx [00:04:07]: You have to explain.Ali [00:04:07]: Sorry. Like, speculative decoding is like, if you have a huge model, right? And so the model is going to be generating one token at a time every single turn, every single forward pass. So we attach, like, this little, like, parasite, like this layer that goes on top of the model, and this model just has to predict. It does three very fast autoregressive forward passes, and it will predict, like, three certain tokens, and then you do one forward stage over the entire original model in order to see if those predictions were correct or not, and then you accept them or you reject them. Now, this draft model is traffic specific, so if you, like, Philip said, if you're summarizing Harry Potter books, I can train exclusively that draft model on Harry Potter books, and I can guarantee you that I'm gonna accept the three tokens every single time. And so with that case, I increase your decode speed. I wouldn't be able to provide this to you if you're a shared endpointSwyx [00:04:53]: YeahAli [00:04:53]: ‘cause I have no idea if you're doing Harry Potter, if you're doing coding, if you're doing English. We don't know. Also, there was a thing in the book that mentioned that if they really cared about a specific threshold, chapter four, I think. Do you remember that?Philip [00:05:06]: Yeah. The things that you can do is you can set a specific, like, batch sizing, a specific, like, parallelism strategy if you're trying to optimize for, like, throughput versus latency. You can. Maybe a NVFP4 quant doesn't pass your benchmarks and you wanna run a model at higher precision, you could do that. There's just a bunch of reasons why you might wanna have your own endpoint and the biggest one, of course, just being, like, you don't have to deal with someone else doing a hundred million tokens of benchmarking traffic at the endpoint when you happen to be trying to serve your users.Swyx [00:05:40]: Yeah. I think one thing that is. That is a classic journey. Like, it's people is asking the, what happens when you type Google into the browser. Tool calling, is that just, you're generating JSON or is there more complication beyond that?Tool Calling, JSON, and Structured OutputsAli [00:05:58]: Certain customers that we have, they have their own post-trained models, and so they demand a tool calling that's not just, like parse a file or go find the weather. It's something that's very specific and you have to do post-training on this. And if the post-training on the model is not good or if the quantization after the post-training to get the inference to be fast, the model will struggle reading the JSON file and reading the tool calling. But it doesn't require its own like sandbox. It's not like it's going to use that tool calling to like escape a sandbox or like it doesn't have to be contained. It can just be a normal dedicated deployment. The challenge with tool calling more and more seems to be that the companies want certain tool calling which is a very sensitive thing to train. And because you're dealing with all of the JSON outputs, if it doesn't like close the end of the request in a very certain manner, you end up with a model that did the tool calling and like the thinking and so as a result of that, it didn't see the result and just hallucinated the result as it decoded. That seems to be the most challenging thing with tool calling, not really the sandboxes model.Philip [00:06:56]: Yeah, that's a challenge on the training side and then on the inference side, there's work that you can do to scope the possible output. So we published this at this point close to two years ago, the solution to this problem which is you make a state machine and you use that to constrain the output to a specific format. So this is the structured output problem. If you remember backSwyx [00:07:27]: Yeah, the specific grammar is,Philip [00:07:29]: Yeah, exactlySwyx [00:07:30]: GML had this thing.Philip [00:07:31]: Yeah. So it's like the old-school “make sure this is only JSON”, return only JSON orSwyx [00:07:38]: YeahPhilip [00:07:38]: Grandma's gonna die type of prompts.Swyx [00:07:39]: Is it BNF grammar? At some point OpenAI had released a thing that was like, yeah, if you want to constrain your output, write BNF grammar, back as NOR.Philip [00:07:47]: In our inference system, it's just a specified output format. And you get the guarantee that your output's gonna be structured along that format. And so applying that to tool calls can like help cut down on. You can still call the wrong tool or call no tool. It doesn't solve the certainty problem but it at least solves the output structuring problemSwyx [00:08:10]: YeahPhilip [00:08:10]: Within tool calls.Swyx [00:08:12]: And MCP is just another form of tool, right.Philip [00:08:14]: Yeah, exactly.Swyx [00:08:15]: As far as there's no special thing there.Philip [00:08:16]: The thing I'm always like explaining to people is the LLM is not capable of doing anything. It's only capable of making suggestions of what to do and then if those suggestions are formatted in a certain way and applied to a system that knows what to do with them, then an action occurs.Swyx [00:08:32]: Yeah. Part of the fun stuff is, this is solved outside of tool calling too. Like in an agent loop if the output is not correct or you're right, like reasoning, tool calling was done in the reasoning trace, just be like, “Oh, I don't know what to do. Let me just try again.” And it might get there after a few tries. And on your point of training, sometimes this is harder in smaller models, so you don't have the same exact quality outputAli [00:08:56]: Right.Swyx [00:08:57]: When you just swap from a big model, right?Ali [00:08:59]: Yeah. I will say that, before, I think we need to go back to inference engineering proper.Ali [00:09:04]: But, I had expected that something would replace JSON because it's hard to stream JSON ‘cause JSON must be complete and you must have open and close brackets and everything. So it's hard to parse something or validate something while it's being streamed. So people invented all sorts of things that are like, I forget the name of some of these alternatives, but it's something like TOML, something like YAML. But JSON seems to be dominant still.Philip [00:09:30]: The JSON outputs aren't that long, right? Like you could have a long-- ‘cause tool calls also contain the arguments in them and perhaps for a certain tool you might pass like a very long argument. But my impression of the median tool call is that it's a relatively small number of tokens, right? So I would expect that speculators are generally fairly good at something as formatted as JSON. And so you would have like a pretty fast decode step there and that the streaming wouldn't be as valuable, but maybe I'm wrong about that.Ali [00:10:02]: I think you're also bounded by the software or that the model is gonna integrate with if the software is built with JSON for the tool calls or if the company that you'- if your customer says that this is how our software works and our tools are interfaced with JSON, you can ask them to like, change their software and say like, “Yeah, this is gonna be better for the model.” but like with the right training shouldn't be that much of a difference. Also more profitable if it outputs more tokens probably.Swyx [00:10:25]: Depends on your business model.Swyx [00:10:27]: It really depends. But I will say that, as a writer with like experience a lot with generated output, I do try to move from text to JSON text which is very long JSON, right? Like there's paragraphs in every field because I'm trying to structure it, right?Philip [00:10:44]: Right.Swyx [00:10:44]: I want you to first make factual statements, then make opinions then make bullet point summaries, have dates, have entity references have your sources for references, all these things. Anyway, so these are things that like I think people who really experiment with structural output have to really care about. But, let's, let's recurse up the stack a little bit. Before we started recording, you mentioned something really cool, which is that there's a lot of engineering that-- inference engineering that goes on when a new model provider releases a new model, right? So let's call it GLM-5.2, Kimi K3. I had previously assumed, especially if it's like, well, GLM 5 to 5.1 to GLM-5.2, like that you've supported them before. Is it that much work?What It Takes to Support a New Open ModelAli [00:11:26]: It's a lot of work.Swyx [00:11:28]: Yeah. Okay. So like, a lot of people, all you guys, right whenever a new model launch like, people rush to say like, “Oh, Hugging Face supports this, Fireworks supports this, Spacetime supports this,” and I'm like, “Yeah, of course we support it.” But what goes into that? What goes intoPhilip [00:11:40]: I think it's more than just support it too, right? It benefits the consumer a lot. Like I think it was with Kimi K2.5 or GLM-5.2 the latest, there was an inference war, right? X provider is at 90 tokens a second. The next day we're at 150. The nextSwyx [00:11:55]: I kinda kicked that off with the GLM-5.2.Swyx [00:11:58]: I wrote a Twitter article about. It got like half a million views,Ali [00:12:02]: Based on being numberSwyx [00:12:03]: YeahAli [00:12:04]: Or it's for something else.Swyx [00:12:05]: Yeah. Which,Ali [00:12:06]: Oh my GodSwyx [00:12:07]: Which then got everyone really excited about, hey, how can we, bend tracks a little bit further and,Philip [00:12:14]: There's a difference between support the model, as in I can make a token out of this model, and support a model, as in I have a production-ready API from this model.Philip [00:12:26]: Getting to the point of I can make a token out of this model is not that hard because generally the, open source inference engines, vLLM, SGLang of the world oftentimes even receive weights ahead of time, maintainers do, or the people making the model merge PRs to ensure support. So you generally can, just get it working on the standard open source stack without too much pain in most cases. The challenge is, every inference company is gonna have own proprietary stack. Some open source components, some in-house stuff. And for any arbitrary model, there's going to be some new stuff. Sometimes you get lucky, like K, two five to two six was, like, pretty similar.Quantization, Speculators, and Production ReadinessAli [00:13:16]: Yeah. It was pure continued post-trainingPhilip [00:13:18]: YeahAli [00:13:18]: If I remember correctly.Philip [00:13:19]: Even in those cases, there's still stuff you have to do. You have to redo the quantization work. You're taking the model from. Generally, these models are not released in NVFP4, and we want them to be in NVFP4 for maximum Blackwell compatibility. So we have to perform that quantization, and, calibrate the quantization to make sure that we're not causing any regression in the model's intelligence. And then we also have to train the speculator, as we've talked about. Generally, we have. We have ZDR, zero data retention on our model APIs, so we don't know exactly the traffic that people are sending us, but we know what's popular. We know that coding use cases are popular. We know that agents, agentic use cases are popular. So we can get public data sets that are representative of that traffic and train general speculators. Now, with speculators today, you need to train the speculator using the base model itself because you're getting hidden states out of the model from running inference on these specific prompts, and that is the training data you use to create the speculator. So there's that process which you need the real model weights for. And then there's of course just the process of, standing up all the infrastructure behind it, loading all this stuff, testing it. And then when there's a new model with a newer architecture, I think that, like, the DeepSeek models tend to be the most challenging as they have, like, the most novel architectural stuff going on, model after model. But every new model has something. Kimi K2 had. Oh, sorry, GLM-5.2 hadAli [00:14:53]: Sparse attention.Philip [00:14:54]: Yeah,Ali [00:14:54]: YeahPhilip [00:14:54]: the DSA.Ali [00:14:55]: Right. Which is brought from DeepSeek.Philip [00:14:57]: Yeah. AndAli [00:14:59]: So you can copy-paste then?Philip [00:15:01]: It kindAli [00:15:01]: I don't know how this works.Philip [00:15:02]: So, like we had to, like, build support for that into our runtime. And you're right, like it is really interesting the way that all of these open source labs borrow from each other. For example, like GLM-5.2 doesn't have vision. So something that, Haley, a guy on our team, if we could take a look at this, he, like, grafted the Kimi vision encoder onto GLM-5.2.Retrofitting Vision into GLM-5.2Ali [00:15:27]: We'll be training the projector.Philip [00:15:28]: Exactly. So if you think about, like, the encoder, there's the encoder, which is the part that looks at the image and turns it into latent information, and then there's the projector which likeAli [00:15:38]: You can say latent space. It's okay.Philip [00:15:41]: And then there's the projector that maps it onto, the model itself, and then there's the model weights. You don't wanna mess with the model weights because you run a chance of making the model dumber at something else for the purpose of giving it vision. So instead, Haley started with just a projector, which is only a handful of millions of parameters.Ali [00:16:02]: That would be, yeah.Philip [00:16:02]: Yeah.Ali [00:16:03]: Can you show the training one?Ali [00:16:04]: Like the way it groksPhilip [00:16:05]: YeahAli [00:16:06]: Very interesting.Philip [00:16:06]: And maybeAli [00:16:07]: That right therePhilip [00:16:07]: Maybe Ali, you should take it from here. You've got a betterAli [00:16:10]: Ooh, double the sandPhilip [00:16:11]: Understanding of this than I do.Ali [00:16:11]: Yeah. You can see, like, he. The way he trained this is really cool. At the beginning, he was training it using just like, “Here's a picture of a mountain. Can you describe what's in this mountain?” And that caused it just like the first, learning walls. Like here you can see this all we're trying to teach it is to translate the encoded. Like it's already taken the encoder from Kimi K. It's taken the image. It'Philip [00:16:31]: Yeah. FrozenAli [00:16:31]: FrozenPhilip [00:16:32]: With adapter.Ali [00:16:32]: Exactly.Philip [00:16:33]: Yeah.Ali [00:16:33]: So the brain is frozen and the eyes are frozen. It's just we're tryingPhilip [00:16:37]: AlignAli [00:16:38]: Interconnect between the eye and the brain, right? So the projector. And so you take the tokens and then he's like, “Oh, can you describe what's in this image?” And he's like, “Oh, it's a mountain,” or it's a person or it's a human, whatever the case is. But that didn't cause complete understanding. So he changed it such that every image was associated with a data set of questions. Like, does this image have a white male? Does this image have birds in the top corner? Does this image have a scientist in it? All of that stuff. And it would have to answer questions correctly. And using not just training on describing an image, but being able to answer question, another question, answer over time. Like you can see the grokking, which is like genuinely insane, that retrofitting vision into a large LLM can learn to that extent. And even for images that it doesn't perform well on, for instance, if you ask it a picture of like Stephen Hawking, “Who is this?” Maybe it doesn't get it, but it will say something like, “This is Albert Einstein.” Like it still understandsPhilip [00:17:25]: Close enoughAli [00:17:26]: That this is a scientist who is a man who has, some significant achievements, all that stuff. So that's like really cool.Philip [00:17:32]: Yeah. So, we've covered Hao Tian before, who the author of the LLaVA paper that did this, a while ago. And I think that's very foundational work for anyone who hasn't done vision work before.Ali [00:17:41]: Same with the CLIP and MetaCLIP, where you go from just captioning to building out questionsPhilip [00:17:47]: RightAli [00:17:47]: Off the image and how much better you can get performance.Philip [00:17:50]: Right. Right. Right. Yeah. But what's, what's so exciting about this is if you look at a model like this. Now, this is a little bit more of a research project. It's not. It got to 56% on MMLU Pro, I think. So not quite frontier. But if you're running this model, you haven't suffered any loss on your GLM-5.2 quality. If you don't have an image, it'll just behave exactly the way it used to. And ultimatelyAli [00:18:14]: Which in the inference code you literally do not include the other part, right?Philip [00:18:18]: Yeah. You would just skip the encoder if you don't have an image input.Ali [00:18:22]: Okay.Philip [00:18:22]: Just confirming.Philip [00:18:23]: YeahAli [00:18:23]: Does it affect a lot on the overall inference side? Like you're not adding much, you're adding a very small vision encoder. These are typically likePhilip [00:18:30]: They're super fineAli [00:18:31]: Less than a billion parameters, right?Philip [00:18:32]: Yeah. It's, - There's a little bit less standardization among vision encodersSwyx [00:18:37]: YeahPhilip [00:18:37]: So the support matrix can be a little bit, sparser. But overall, yeah, it's a pretty, it's a pretty minor component of the overall system. And ultimately what you get out of the system is all of a sudden you have Kimi Vision, GLM weights, and DeepSeek attention all in one model.Open Source Model Grafting and Franken-MergesPhilip [00:18:56]: And that's, I think, a lot of the power and beauty of open source, is that you can take all of these different components and combine them together into a system that's better than anyoneSwyx [00:19:05]: YeahPhilip [00:19:05]: Can be individually.Swyx [00:19:06]: People used to say that you would also do Franken-merges where you would take likePhilip [00:19:10]: YeahSwyx [00:19:10]: Layers from each model.Swyx [00:19:11]: Does anyone do that anymore?Ali [00:19:13]: Well, to your point previously when you were mentioning like, the work that goes into supporting a model when it first comes out, like GLM-5.2 or MiniMax M3 or whatever the case is. Sometimes you do have to like, you do have to switch out some things. Like, for instance, the MiniMax M3 head uses full attention, and with full attention you end up with this like insane bottleneck in spec dec ‘cause you're doing auto-regressive token generation for three tokens, and you're doing this like N squared over all of the tokens that are in your sequence. Your KV cache is like very large because it's not sparse, it's not top K. So we find it better to like, okay, we're gonna replace this, we're gonna replace this layer with a layer from another model that's using like GQA, for instance. And then just with the right training, you can get it to have the same acceptance rate. So it is very possible to retrofit layers from other models and very much needed. If a layer is like inefficient, the training just becomes the challenge, like how do you ensure that you train it properly? Which again to your earlier point is like the mesh between training and inference. As in like you need very good training in order to do fast inference. That's like, I feel like more and more becoming true.Swyx [00:20:21]: Yeah. Anything else on the support side when you say like get it to fully production ready?Loop Detection, Race Conditions, and Non-DeterminismPhilip [00:20:26]: Yeah. I think that there's also a question of just, we can test a model to a pretty extensive degree, but we're trying to get it out quickly and then you see a bunch of other people test it and you get interesting results. There was an issue with, GLM briefly where we had some like mode collapses where it would just output the same token over and over again for certain prompts on certain temperatures. Like once you expose an endpoint to the real world, there's going to be, so many more varieties of things given to it that you're able to, discover and patch things. So it's not just a, day zero process, it's then like for the first week, for the first month, if a model remains popular, like how do you both fix bugs and then continue to push the envelope on performance?Ali [00:21:21]: What do you mean you don't want your model outputting S?Swyx [00:21:24]: Is there loop detection on that stuff, by the way? It still happens like quite a lot, which is surprising.Ali [00:21:30]: We have like we, in our endpoint, like if a model was to output the same token like four plus times, we just cut the generation. We say like, “Oh, sorry, this-- Like try again,” or like we will reprocess the request. ‘Cause we know then, like if it, like if, yeah, it's four times the same token, it's probably collapsed.Swyx [00:21:45]: Yeah. Is there a way to opt out in case I really want that?Ali [00:21:48]: You want that?Ali [00:21:50]: I think there's a way that we have to handle it. I'm not exactly certain, but I feel like in certain models, like when they output something like you can imagine, like a table for instance, and so they want, they wanna draw like 12 dashes and 12 dashes. Yeah, I think there's a way for that to happen. I think we only do it on certain tokens. Like we exclude certain special characters.Swyx [00:22:07]: Yeah.Ali [00:22:07]: So we only do it on like certain like S is the most common almost. GLM-5.2Swyx [00:22:11]: OhAli [00:22:11]: And I think it was DSV 4 as well. Like you'd just have like looping issues where like you literallySwyx [00:22:17]: ItAli [00:22:17]: Just have like S.Swyx [00:22:18]: Yeah. Is there a special, something special about S? No, just randomlyAli [00:22:21]: It just seems to be the one token involved.Swyx [00:22:23]: Yeah. And it'Philip [00:22:24]: Is thereSwyx [00:22:24]: And it's only temperature 0Ali [00:22:27]: NoSwyx [00:22:27]: Even at other temperaturesAli [00:22:27]: Even at like 0.9 or whatever, it will still, it will still collapse.Swyx [00:22:30]: That's weird, right?Ali [00:22:30]: It's, it is an inference problem to be honest, like a software problem. Like oftentimes, the image you run will-- like NVIDIA will release an image for instance, and if we will upstream the changes from their latest TensorRT-LLM image into our stack, we'll find that it fixes it. Or oftentimes this will only happen in an inference engine that you're using like SGLang. But if you were to switch to vLLM, that isn't the case. So it seems to be like an extremely like deterministic software issue and not really a model issue. It's not like a weights problem. Like I'- we'll say like, “Oh, it's a problem with the quant. We did PTQ wrong,” right? But that isn't, that doesn't make sense because the same weights used with a different inference engine does not repeat the problem. And sometimes it's, the kernels that are being used in the backend have like these very subtle sometimes race conditions, where if you were to use this model hosted on one cluster, you will never get this problem.Swyx [00:23:19]: Oh my God.Ali [00:23:19]: But if you host it on a different cluster, you will. And the reason is the KV cache transfer from a node to node in that one cluster is using a slower interconnect than the node to node in another cluster. So that exposes the race, whereas in another cluster it doesn't. So then you end up just like, okay, this model is not gonna be hosted on this cluster. We're gonna host it on, another cluster because that cluster exposed that problem. But then it ends up with like, okay, is it the software? Is it the model weights or is it the hardware?Swyx [00:23:42]: There is a thing about this with temperature 0 still not being deterministic, right?Ali [00:23:46]: Right.Swyx [00:23:46]: Mostly because of hardware. Even at temperature 0 same model, you won't always get the same output.Swyx [00:23:52]: Even-- But I'm surprised by the race condition one because, I thought PyTorch was a graph that like guarantees that you at least, execute things in the right order.Ali [00:24:02]: Well, yeah, true. Like I'm not, I'm not saying that there is. Like well, you have things like PTL optimizations where like you can start a kernel before the end of the previous kernel, and that's like ‘cause you want to do that because there'sSwyx [00:24:12]: It's like pipeliningAli [00:24:12]: Expense. Exactly.Swyx [00:24:13]: Yeah.Ali [00:24:13]: But it'- But you don't do it cleanly. Like you overlap a little bit of the execution. No, it is very possible that the kernel itself, like that one block that is supposed to be running in this instance of time, that kernel itself has a race condition. For instance, like a missing barrier. Like often if you're designing a kernel and you want it to make it to be very fast, if you don't test it extensively, you'll, you'll have certain threads access data points from registers before they've been written to by other threadsSwyx [00:24:36]: YeahAli [00:24:36]: For example, because like your barrier is wrong or your synchronization was wrong. But yeah, like the testing itself is very difficult in those like, andSwyx [00:24:42]: And there's no like borrow checkerAli [00:24:45]: What does that mean?Swyx [00:24:46]: Like Rust. Like the. If you're trying to have like memory safety It sounds like a comparable problem.Ali [00:24:52]: Well, yes, but you're working in CUDA, right, NVIDIA GPUs. Like- You just need a higher level language like modular Maybe that's what modular is supposed to do. I don't know.Quantization Quality and Vendor FidelityVibhu [00:25:00]: How do you see keeping quality of the model? So you talked about all these steps of, okay, you gotta do quantization, train your own speculative decoderAli [00:25:07]: RightVibhu [00:25:07]: Run on different hardware. Looking at other model providers, okay, you kicked off a inference speed race on the consumer end. What goes into keeping quality the same across them, right? Sure, you can run benchmarksAli [00:25:22]: YeahVibhu [00:25:22]: But, like, how do you determine how much quantization are there standards? What goes intoPhilip [00:25:27]: There's a few things on quality. Most inference optimizations are lossless. KV caching, for example. You are just recomputing or preventing recomputing the same values. Speculation, of course, if a draft token is wrong, it gets rejected. The main lossy optimization is quantization. And that really comes down to, number one, data format, number two, which parts of the model you choose to quantize, which layers, and number three, like doing a lot of calibration on the quantized weights, to ensure that you're preserving all the outliers. There's other tricks that you can do, though. A big one is long context, ‘cause one thing you asked at, right at the beginning is, “Oh, what's gonna happen if I send a 200,000 token request in?” So with a long input sequence, you need to, store a lot more information. You need to process a lot more tokens. And so even if a model has a context of a certain length, you might, as an inference provider, choose to build an API with a shorter context length, and of course a full length one as well. Because if someone doesn't need the full million token context, for example, you can get them better performance. I don't know if that's exactly like quality of the model. The way that I think about quality is to what degree are we faithfully serving the original model? If you think of a golden implementation of a model that performs exactly the way the model is designed to perform, I think of quality as how close are we getting to that, 100% fidelity of the model.Philip [00:27:13]: You can also, of course, think about quality from the training side and how do you push yourself past 100%. But when I think about purely inference optimizations, it's getting faster while staying as close to that 100% fidelity mark as possible. And certainly our standard internally is that, like you should not be able to tell the difference between our API and a, official API. I think Kimi in particular does a good job of vendor benchmarking hereAli [00:27:41]: YesPhilip [00:27:41]: Where they haveAli [00:27:42]: They released an actual vendor benchmark.Philip [00:27:43]: Exactly, yeah.Ali [00:27:44]: ‘Cause they accused, some people, Amazon? There was some provider that was not doing very well on Kimi's benchmark.Philip [00:27:50]: Yeah.Philip [00:27:51]: So, with Reflect we probablyVibhu [00:27:52]: This was a long time ago, right?Philip [00:27:54]: No.Ali [00:27:54]: Yeah, like threeVibhu [00:27:55]: They alsoAli [00:27:55]: Four, five months agoVibhu [00:27:57]: This also happened with, I don't remember which model, but they pulled out quite a few, and then they started a whole chart about this. It might have beenPhilip [00:28:03]: Kimi Vendor Verifier.Ali [00:28:04]: Yeah.Philip [00:28:05]: Yeah.Ali [00:28:05]: Yeah, ‘cause you, ‘cause you'd be pissed, right? Like if you'Philip [00:28:07]: Yeah.Ali [00:28:07]: If like if I'm a consumer and I'm using like Amazon's endpoint for instance, and I've used Kimi and I'm like, “Oh my God, like this is bad,” I'm not gonna say, “Oh, Amazon quantized the model in a bad way.” I'm gonna say, “Oh, Kimi sucks.” Right?Philip [00:28:17]: Yeah.Ali [00:28:17]: So it seems like that makes sense.Philip [00:28:19]: Yeah, they care. They care.Vibhu [00:28:21]: Justifiably.Ali [00:28:21]: Yeah, justifiably.Vibhu [00:28:22]: This is probably a stupid question, but just checking, has anything improved from main quantization?Philip [00:28:28]: Yeah.Vibhu [00:28:28]: Like, is quantization always strictly worse?Ali [00:28:30]: Well technicallyVibhu [00:28:32]: NoAli [00:28:32]: It's a lossy. QuantizationPhilip [00:28:33]: YeahAli [00:28:33]: Is a lossy, it's a lossy implementation.Philip [00:28:36]: Speed improvesVibhu [00:28:36]: Speed improves.Ali [00:28:37]: It the number, likeVibhu [00:28:38]: No, I' always look for inverse scaling laws.Philip [00:28:40]: Yeah.Ali [00:28:40]: Yeah.Vibhu [00:28:40]: This is something I learned from Noam Brown, where like things that normally act in one direction sometimes do.Philip [00:28:45]: Well, technically when you run a benchmark, because these models are deterministic, sometimes your,Ali [00:28:52]: YeahPhilip [00:28:52]: NVFP4 quant is like, two basis points higher than yourAli [00:28:56]: No, it's noise. It's noise.Philip [00:28:57]: Yeah, exactly. I'm like, yeah, it's, it's within. That's why I always say within margin of error.Philip [00:29:01]: And I stopped saying that because everyone assumes that what is, well, within some margin of error, we're barely inside of that to the worst, so we're saying. But yeah, sometimes it's just like, gives you a higher output score. But like Ali said, that's noise. To my knowledge, you're not necessarily making the results better. You're just trying to, again, like keep your fidelity as close to 100% to the original model.Layer Selection, KL Divergence, and Better QuantizationAli [00:29:27]: There is, to your point, research that we did on MP. I don't know if you are able to pullPhilip [00:29:31]: YeahAli [00:29:32]: A tweet we did. One of our research interns, Joshua, I think it's a tweet on how we have 20% better quantized GLM-5.2 than NVIDIA. Essentially what we found throughout like this month research is, okay, quantization is a lossy. It's. You're compressing the data from, occupying 16 bits to occupying, four bits, for instance. And so you're losing some information, and you're trying to minimize that. And so when I say that I'm gonna quantize the model, my job becomes how do I find the layers that I can quantize, and how to find the layers to not. For instance, with image models, I don't quantize modulation layers, and I don't quantize out projections because those two are. Like out projection is what you see as the user. Modulation is what the model sees or understands. Right, exactly. And so to his paper, do you have the. It doesn't have the. Yeah. It's a long paper. I don't know if I can findVibhu [00:30:25]: If there's a part to search or it's probably in the thread.Ali [00:30:28]: It's probably in the thread.Vibhu [00:30:29]: Yeah.Ali [00:30:29]: But the long and the short is it is very possible that quantizing more of the model makes the results. Like if I have a model that I quantize layers one, five, and 10, and another model where I only quantize layers one and It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in, is that you can predict which layers are going to have quantization errors that will cancel out with each other, and you choose to quantize those layers. And so the result of doing this mathematical quantization is you end up with a model that's 20% more quantized than another provider, so you get 20% more throughput of it because there's more layers than running an NVFP4, and your quality is better than that other quant because the layers that you chose to quantize have their errors cancel out, like one layer skewed to the right one layer skewed to the left, one layer skewed to the right. Your final logits distribution is more similar to the original distribution of the model, so you have better fidelity. And so the way we proved this was with KL divergence. So instead of just scoring on the benchmarks, we scored the KL divergence between the logit distribution of the quantized model and the logit distribution of the original full precision model, and we showed that with this technique we get. If your probability distribution on the logits which token it wants to select is more of the same as the original model, you're probably gonna end up staying true to the original model. So yeah, so it seems like previously before this, it seemed like the industry was, well, the more you quantize, the worse it's gonna be, ‘cause the more loss you introduce. That's not exactly, not necessarily true. So yeah, doesn't improve it, but can cancel out.Philip [00:31:57]: I think it might be this, but reminds me a good bit about pruning where you can prune off certain layers.Philip [00:32:03]: But very interesting. Didn't know this was a whole paper you guys put out.Ali [00:32:06]: It's. Fun fact, it was originally 72 pages, this paper, and then we decidedPhilip [00:32:11]: WowAli [00:32:11]: We can't tell. We couldn't release it. So it's now 45.Swyx [00:32:15]: Still 39 pages, so very substantive. We talked about evals and all these things and, like what's possible in terms of speedup? Like it's like probably like the numberInference Speedups and BenchmarkingSwyx [00:32:25]: Thing that people do wanna care about, and it's something that you wrote about in your post. Like official API is 70 tokens per second, and you push it up to 90. Is that like a normal thing?Philip [00:32:36]: So what's cool about working in inference, the reason that I think inference is going to be a useful place to do engineering for a long time, is that if you look at highly optimized domains like, say, finance, if you're in finance, you measure how much better you got in basis points. It's like, “Oh, I got five basis points better, like twentieth of 1% better,” that's huge news because everything is so optimized. When we publish optimizations, it's 20%, it's 100% it's 200%. So there's still probably like a lot further to go, honestly. Like you'll, you'll know that inference is pretty much solved when researchers start publishing about how they got 1% faster at something.Swyx [00:33:19]: Which by the way, because I am from the finance background, in the ‘70s, that was the margin at the time. When you did quantitative finance research, you would findAli [00:33:27]: And like 20%, tens of percent.Swyx [00:33:29]: That's. Yes.Philip [00:33:29]: Yeah.Swyx [00:33:30]: And now it'Philip [00:33:31]: Tiny fractionsSwyx [00:33:32]: For those people interested, look up Andrew Lo's paper. He had a really interesting illustration of quant, stat arb, distribution, narrowing down from like those kinds of 20% differences in the ‘70s, down to nothing today, which is very cool.Philip [00:33:48]: Exactly, and we're at the beginning of the same type of thing. Now benchmarking is hard. I think anyone will tell you that, and benchmarking provider speeds is hard because there's so many variables that go into it. What hardware are you using? How much load do you have on the system? What's the exact nature of the prompts and input and output sequence lengths? All that stuff. But overall, when you start stacking these improvements, you're looking at multiples. You can look at it. The most common form, of course, is TPS, tokens per second, which is bad naming by us in the industry, ‘cause there's two tokens per second. There's tokens per second, the throughput number, and the latency number.Ali [00:34:31]: TTMT, yeah.Philip [00:34:32]: Like total tokens per second out of the, out of the GPU as a throughput number. Most people only care about tokens per second as the latency number, which we should call ITL, intertoken latency, but we don't.Philip [00:34:44]: Anyway, so you can imagine a standard API without many optimizations for a 1 trillion parameter model operating somewhere in the 30 to 50 tokens per second range for reasonable traffic profile. And we generally see the goal of, pushing to 10X that. But, not necessarily day zero, but by stacking enough optimizations, if you have, say like four optimizations, each of which doubles performance. Or sorry, three optimizations, each of which doubles performance, then you stack that up, that's an 8X gain. That's the order of magnitude that we're working with in this space. We're trying to make things substantially faster, not just go from like 70 to 90.Swyx [00:35:38]: Are you saying you've. You have done that?Philip [00:35:40]: So let's say you have as a reasonable baseline, 30 or 40 tokens per second. You can achieve 10X that. So like on GLM-5.2, if you run it unquantized, perhaps on H100s even, and you're just using an off-the-shelf inference engine with no particular optimizations, no speculator, nothing extra around like KV routing, no disaggregation, you're, you're probably, yeah, looking at that like 30 to 40. You think that's like a reasonable baseline?Swyx [00:36:12]: Right. Right.Philip [00:36:12]: To get to something like 10X, there's a lot of trade-offs that you're making. If we're running at more like a 300, 400 tokens per second range, you are using the best hardware possible. You have a optimized speculator. You have done all of your quantization work. You are Seeing a pretty high cache hit rate. You are running with a reasonably small batch size and a parallelism configuration that is tuned for latency versus throughput, but it is possible. So the spreads that you see if you, like, go on artificial analysis or you go on OpenRouter and you look at, the worst provider to the best provider, oftentimes can hit that range. 10X is of course very aggressive. It's oftentimes maybe more of a four to six times improvement. But that's the performance that makes us really excited, is when we can get these huge gains, not just go from 70 to 90 tokens.Stacking Optimizations: NVFP4, Speculation, and DisaggregationAli [00:37:19]: It's also, like, hardware dependent. Like, ifPhilip [00:37:20]: YeahAli [00:37:20]: If you have a thing where you're serving it on just, like, a node of H100s and then you throw, like, you shard the model across, like, four nodes of B200s. Like, you can definitely increase the speed with just throwing more hardware at it. Like, normalizing for the same exact hardware and the same number of GPUs.Philip [00:37:35]: Yeah. Then you're looking at, like, a two to 4X improvementAli [00:37:38]: Right. RightPhilip [00:37:38]: Depending on the inference optimizations. So yeah, it's. Some of it's, what's the call, and some of it's who's the driver.Vibhu [00:37:46]: If you break down the two to 4X, say the example is run GLM-5.2Ali [00:37:51]: YeahVibhu [00:37:51]: On B200sAli [00:37:53]: YeahVibhu [00:37:53]: Single node, right? What's, like, the cost trade-off for effort to get, like, the last bit of juice out versus what should people just think of, right?Ali [00:38:01]: Spectre quantization. Yeah.Vibhu [00:38:03]: Spectre quantization.Ali [00:38:04]: That's, that's, that's like 95%. LikeVibhu [00:38:06]: And how far does that get you? And how easy is that for the average person to do? So say right I wanna throw the weights of GLM-5.2 on a node of B200s, how easy is it to find speculative decoder- decoder model or already quantized model? How much work goes into it?Philip [00:38:23]: If you're doing it up front, it's quite a lot of work. If you're doing it today, there's going to be people who have published things that you can just, you can just grab some NVFP4 weights. You can grab a speculator. Yeah, if we're thinking about, like, what are the 2Xs we're stacking, going from, BF16 to NVFP4 is, it's not quite a 2X, right? It's like. I think it's about, like, 30 to 40%, from 16 to 8, and then another 30 to 40% multiplied from, 8 to 4. So that doesn't quite get you a 2X, but, like, roughly a 2X. Speculator, roughly a 2X. Disagg on top of that if you're able to get enough hardware and put enough traffic through it, another roughly a 2X. And then you add in some, double-digit percent increase from having just a better runtime with, the latest kernels and stuff behind it. And that's how it stacks up.Ali [00:39:21]: YeahPhilip [00:39:21]: So building each of those, like, building the, quantized weights is, for someone who really knows what they're doing, hours to days of work. Building the speculator, again, like, hours to days of work. And the, disagg setup, hours to days. Well okay, but like once you haveAli [00:39:39]: Once set up. Once set up. YeahPhilip [00:39:40]: Yeah, getting disagg working for the first time, I'm saying, of course, is very difficult.Philip [00:39:44]: The marginal implementationAli [00:39:48]: Like, if you're just grabbing, like if you are a person, like just a normal consumer who has access to, like, a node of B200s and you're wondering, “How can I just host it myself?” You don't need to quantize the model yourself. There's always gonna be, like, an open source quantized checkpoint. NVIDIA's gonna push one out if no one else does. You. Usually, the providers will have their own spec dec that they've trained as well. You don't need to train your own spec dec. You can just use that as well.Philip [00:40:09]: Yeah. Like, GLM-5.2 has its own MTP.Ali [00:40:13]: Right. Right.Vibhu [00:40:14]: What's multi token prediction?Philip [00:40:15]: Yes.Ali [00:40:16]: I'm justVibhu [00:40:16]: Can you explain that?Ali [00:40:16]: I'm just an expert.Ali [00:40:18]: I can do it for you in case I get it wrong?Vibhu [00:40:20]: No.Vibhu [00:40:21]: Yeah, you should correct if we're wrong, but their multi-token prediction can be used for self-speculative decoding.Ali [00:40:27]: I'm not sure. I'm not gonna correct that.Vibhu [00:40:28]: Okay. I'm semi-confident in thatAli [00:40:30]: Okay. YeahVibhu [00:40:30]: But someone can check. But it's useful to paint the story of, okay, not just the average person, but say a company wants to switch from serverless inference I wanna throw this up on. I wanna rent some GPUs, throw it up. These are the steps you take to do significantly faster than just put it behind vLLM.Ali [00:40:48]: Right.Vibhu [00:40:49]: I was waiting for a mention of Dynamo.Vibhu [00:40:51]: I feel like, that's supposed to be the baseline that you measure against.Dynamo, KV Routing, and Disaggregation ToolkitsPhilip [00:40:55]: I would think of Dynamo as less of a box system and more of a toolkit for building with. So when we talk about doing aware routing, when we talk about doing KV offloading, when we talk about doing, PD disaggregation, Dynamo fundamentally is. By the way, Dynamo is an open source library from NVIDIA.Ali [00:41:17]: We've done a pod with KylePhilip [00:41:18]: OkayAli [00:41:19]: Kyle Cranin.Philip [00:41:19]: Cool. So then your listeners know then that it supports all the different inference frameworks. And it is multi hardware, which is interesting.Ali [00:41:28]: But it's just a router, it's not like an optimizer layer.Philip [00:41:30]: Yeah. All it does, like, what Dynamo is good at, it is a library for moving information around your cluster, around your hardware. So if you have, KV cache on one place and you need it to be somewhere else, Dynamo coordinates NIXL for you to move that around.Philip [00:41:49]: That doesn't mean that, like, out of the box, you just say, “Pip install Dynamo,” and then you get, like, a massive performance speed up. It's more of a developer toolkit.Ali [00:42:01]: Yeah. I would have said it would. It comes with a set of defaults that you can then swap out.Philip [00:42:06]: It does. If the industry at large, I think, was, like, rolling out all of these deployments, standard, then I think it would be, like, a credible baseline. But, we've got to, we've got to benchmark against, like, what we're seeing in the wild.Speculative Decoding Methods: Medusa, EAGLE, n-Gram, and Spec-SpecVibhu [00:42:23]: I did wanna talk a little bit more about PD disagg, because that is probably, like, number three after quantized and speculative decoding. In your book though, I was just gonna pull out the book.Philip [00:42:31]: Yeah.Vibhu [00:42:32]: Like section 522 on Medusa, 523 on EAGLEPhilip [00:42:35]: YeahVibhu [00:42:36]: 524 on gram.Philip [00:42:37]: It's 55, would be disaggregationAli [00:42:42]: Yeah. Well, no, I just wanted to dwell a little bitPhilip [00:42:44]: YeahAli [00:42:44]: The other. Like, so what do you choose to include? What do you choose to not to include? Because there was all these other techniques.Philip [00:42:51]: Yeah.Ali [00:42:51]: Are these still relevant? Because I think they came out, like, a year and a half ago maybe.Vibhu [00:42:55]: Medusa is quite old.Philip [00:42:56]: Yeah, Medusa's old.Ali [00:42:58]: It was old.Vibhu [00:42:58]: But is it in the book as a good, here'sPhilip [00:43:01]: BaselineVibhu [00:43:01]: Baseline vanilla understand it?Philip [00:43:02]: Like you should know this.Vibhu [00:43:03]: Like I read the paper, I'm like, “ it makes so much sense.”Philip [00:43:05]: Yeah.Philip [00:43:05]: So with the book, I had a couple goals. One was to give people just a working vocabulary for the space as a whole, and the other was to give them some intuition about how each of these techniques works. As I mentioned in my AI Engineer talk, which is the first public addendum to this, the speculation space has moved much faster than everything else. So yeah, even at the time that I wrote the book Medusa, I very much included as a way for people to understand how the space evolved rather than what the most modern technique is. And now of course, there's DFlash, dSpark. There's, there's newer techniques even than EAGLE, although EAGLE is still very commonly used.Ali [00:43:51]: SpecSpecta.Philip [00:43:52]: Yes. Speculative decoding.Vibhu [00:43:54]: What canAli [00:43:56]: Oh, it's a paper by Tri Dao and it's like, it's doing speculative decodingVibhu [00:44:00]: HuhAli [00:44:01]: For the speculative decoder.Philip [00:44:02]: Oh, in spec- oh my God.Ali [00:44:02]: It's literally just an another. It's like, yeah, that's the most simple way to explain it, and it seems like he got trivial speed ups there. But it seems that the complexity with training, it's almost like in our mind at least, it's almost as complex as training GANs. Like it's like a very delicate balance and oftentimes you, it's just but yeah, it's literally speculative decoding on speculative decoding.Vibhu [00:44:21]: Speculative.Ali [00:44:22]: Yeah. We saw this paper.Vibhu [00:44:24]: It's interesting, right?Ali [00:44:24]: Yeah.Vibhu [00:44:24]: I wouldn't even expect it to be very particular to train, I wouldAli [00:44:29]: Right.Vibhu [00:44:29]: The naive part of me is like, okay, train speculative decoder.Ali [00:44:32]: But like, and it makes sense, like the whole idea of speculative decoding is you. It's like, it's like almost like the iPhone auto predict version but for a normal model, right? Like you're just, you're just, generating three tokens and you're like, okay, I'll do prefill on them. And so you save those three turns for your original model. Now your speculative decoder is doing three turns of auto regression, so why not just have an even smaller model?Ali [00:44:53]: The other question there is what are the size of speculators? So say forPhilip [00:44:58]: Right. It's like a billion parameters.Ali [00:45:01]: Like for MiniMax, it's. Yeah. It's like one layer. It's like one 60th of the original model usually.Philip [00:45:06]: Yeah. I think we should do a paper when we get back to the office.Philip [00:45:10]: SpeculativeAli [00:45:11]: SpeculativePhilip [00:45:11]: Decoding.Ali [00:45:13]: No, it's, it does seem like how, when do you stop? But then it also seems like if you're able to train spec-spec decode for instance, right? Like if you're able to have a small model that is accurately predicts what the intermediate speculator is gonna predict, that is able to predict what the original target model's gonna predict, then why not just use that smallest model directly, right?Vibhu [00:45:34]: Yeah. This isAli [00:45:35]: Like it seems likeVibhu [00:45:35]: Adjacent to the routing problem.Ali [00:45:36]: Right.Vibhu [00:45:36]: Yeah.Ali [00:45:36]: Right.Philip [00:45:37]: The thing with speculators is one of the practical constraints on using them is that you do have to run a small model on the same hardware that you're running the big model on. There is a orchestration and resource competition problem inherent in that, and that is one of the constraints on speculation in general, is that draft tokens cost resources to create and cost software complexity to manage. And so if you have like infinitely recursive speculators, you add in quite a bit of that complexity on the actual implementation within the inference engine as well, not just in the training process.Vibhu [00:46:17]: I was gonna say, I would wonder if you could do similar, like distillation and pruning of, it's the same thing, it's just a model. Can we not just distill a lot of the weights, quantize the speculator, out of my domain? The question that also comes up is, this is all for big server workloads, right? How much of this applies to, say I have this MacBook, I wanna run Gemma really efficiently. Similar problems, not the same?Local AI vs. Data Center InferencePhilip [00:46:45]: Pretty different. I talked to Selo, about this on his podcast a couple weeks ago. The difference between inference engineering for the data center and for production workloads versus inference engineering for local AI, is that we start with fundamentally like different constraints and different goals. With local AI, it's how do I fit this model onto my hardware and then make it less dumb? And with data center influence, it's how do I load this model and then make it less slow? And we care about less dumb, and they care about less slow. But the local AI inference engineering ecosystem, I think has a lot for us to learn from in the data center space. They are experts in various forms of quantization, including dynamic quantization that we just don't touch, in the pruning, in the distillation, in the, layer removal. There'Ali [00:47:42]: Layer removal matters less.Philip [00:47:43]: Yeah. There'Ali [00:47:44]: No one loves pruning really.Philip [00:47:45]: Yeah. Well, but the, but they doVibhu [00:47:46]: Which is surprising, right? But that's, that's a whole different thingPhilip [00:47:48]: Just to fit something on the laptop.Ali [00:47:50]: Right.Philip [00:47:50]: So yeah, it's a, it's an interesting, it's an interesting space. Not necessarily that like their techniques make sense for us to do in the data center, because we have different resources and different goals, but more that the process as well as the openness of that field is something to, admire.Ali [00:48:12]: Yeah. Like to your point, like, certain optimizations that would. Like for instance, Turbo Quantum Sharper, like it made such huge hype on that and we did like a whole deep dive on Twitter and like said, what is it? How does it work? Why is it good or not? And it took off and it was implemented on local devices because your memory bandwidth is so slow on like a MacBook, for instance. But try putting the same thing on like an NVIDIA GPU on a B200 Turbo quant would not be. Like, it would not be used. Like, NVIDIA - Like, NVIDIA made it clear that this is not a good optimization, and we've seen it firsthand where the overhead of doing dequantization, quantization of, in the kernel itself with turbo quant kernel, each end is much slower than the time that you save from doing the bandwidth. ‘Cause on the B200s, you have like 3.5 terabytes per second. You don't need decrease the storage that much. You don't need to do, FP4 KV cache. You don't need to use a requant. There's, there's, there's better optimizations to be made. But on Edge devices, it's extremely important, it's extremely useful. So, seems to be, like, different optimizations there, but then they're all uniquely combined with like all you wanna quantize the model, you wanna do speculative decoding, like certain common prefixes with bothPhilip [00:49:18]: Principles.Ali [00:49:19]: Yeah, exactly. Exactly. Exactly.Philip [00:49:20]: They also do a lot of work on, model parallelism, especially over, heterogeneous topology, where you have, some sparks and they are wired together with, Ethernet, DGX sparks.Ali [00:49:35]: Yeah, this is the Exo Labs guys.Philip [00:49:36]: Yeah. You have, a nu

The Tech Blog Writer Podcast
Running Enterprise Computer Vision on CPUs With Ultralytics YOLO26

The Tech Blog Writer Podcast

Play Episode Listen Later Jul 29, 2026 25:22


What becomes possible when enterprise computer vision no longer depends on expensive GPU infrastructure? In this episode of Tech Talks Daily, I speak with Glenn Jocher, founder and CEO of Ultralytics, about YOLO26, CPU inference, edge AI, open vocabulary vision, deployment economics, and the practical work required to move computer vision from a promising pilot into production. Glenn's route into AI began inside the U.S. intelligence community. He worked with the National Geospatial Intelligence Agency and Defense Intelligence Agency on particle physics applications, attempting to detect and track antineutrinos. Antineutrinos are extraordinarily difficult to detect because they pass through almost everything. Glenn describes them as the perfect spy. While searching for better detection methods, he discovered that computer vision researchers were solving similar problems with images. His original attempt to transfer those techniques into particle physics did not succeed. However, the work introduced him to a field where the technology could create a visible effect on everyday life. That led him toward open source development and eventually the YOLO models for object detection, classification, segmentation, and tracking. Glenn believes computer vision research has historically placed too much attention on small gains in accuracy while overlooking deployment economics. A model can perform impressively inside a laboratory and still remain unsuitable for a factory, warehouse, store, vehicle, drone, or medical environment. Price, latency, power consumption, data privacy, and deployment speed can determine whether the technology is commercially useful. This led Glenn and Ultralytics toward smaller models capable of running close to where images and video are generated. YOLO26 continues that approach with architectural changes designed specifically for CPU inference. Glenn says the model can process camera streams in real time at 30 frames per second and run across Intel CPUs, AMD CPUs, and lower power devices such as Raspberry Pi computers. This matters because specialist GPUs can increase the equipment cost and power requirements of a computer vision project. Running inference on existing CPUs or edge hardware can make deployment economically possible across larger numbers of cameras and locations. The scale already involved is difficult to comprehend. Glenn says Ultralytics models now process approximately three billion inference jobs each day, equivalent to around 30,000 every second. These jobs include images, videos, and collections of images being analyzed to detect, segment, or track objects. He attributes the platform's maturity to thousands of mistakes and bugs corrected through a rapid feedback cycle. New models are released, users report problems and request features, and the team incorporates that information into later versions. We also discuss the respective roles of cloud and edge infrastructure. Glenn sees cloud platforms continuing to provide the computing power required for training, while computer vision inference often belongs at the edge. Local processing can reduce latency, control operating costs, and keep sensitive video or medical information closer to where it was created. The smallest YOLO model is approximately three megabytes, according to Glenn. That allows it to reach mobile phones, vehicles, drones, battery powered devices, and other environments where a large language model would be impractical. Open vocabulary vision provides another development. Traditional object detection models are trained to recognize a fixed collection of objects. If a model learns to detect dogs and the user later wants it to detect cats, retraining can cause it to forget earlier knowledge unless both categories appear in the new training data. Glenn explains how promptable models can identify common everyday objects from text or visual instructions without additional training. A user could request a person wearing a blue shirt and white shoes, for example, and the system could search an image for that description. That flexibility could benefit businesses whose requirements change regularly. It reduces the need to create and label a new data set every time the company wants the model to recognize another common object. The range of current applications is already extensive. Glenn describes YOLO being used across robotics, parking, industrial safety, PPE detection, warehouses, aviation, security, traffic management, food quality, and manufacturing. Some of his favorite examples involve environmental problems. One company uses YOLO with underwater vehicles to identify and recover plastic from the ocean. Other applications detect smoke and fire early enough to support forest fire response. For leaders considering computer vision, Glenn recommends beginning with a defined problem and measurable outcome. A manufacturing company may want to reduce defects, but it still needs labeled examples showing the model what acceptable and defective products look like. He advises testing the idea through a limited pilot, measuring the return, and expanding only when the evidence supports further investment. Computer vision has become easier to deploy, but practical problems involving data, cameras, integration, reliability, and operating conditions still separate a demonstration from a production system. Could CPU inference and open vocabulary models make computer vision practical for processes your organization previously considered too expensive? Listen to the episode and share your thoughts with me. Useful Links   Ultralytics website Ultralytics Platform      

Python Bytes
#490 It's a vibe coding party

Python Bytes

Play Episode Listen Later Jul 28, 2026 37:14 Transcription Available


Topics covered in this episode: Some more things about Django I've been enjoying Who cleans up after the vibe-coding party? Where Did All Your AI Tokens Go? AgentsView to the rescue! Careful with phishing all Extras Joke Watch on YouTube About the show Sponsored by us! Support our work through: Our courses at Talk Python Consulting from Six Feet Up Connect with the hosts Michael: Mastodon / BlueSky / X / LinkedIn Calvin: Mastodon / BlueSky / X / LinkedIn Show: Mastodon / BlueSky / X Join us on YouTube at pythonbytes.fm/live to be part of the audience. Usually Tuesday at 7am PT. Older video versions available there too. Finally, if you want an artisanal, hand-crafted digest of every week of the show notes in email form? Add your name and email to our friends of the show list, we'll never share it. Calvin #1: Some more things about Django I've been enjoying Julia Evans is learning "2010-style" web dev (Django + SQL + server-rendered HTML) after years of Go backends and JS-heavy frontends Query builders: likes defining custom QuerySet classes with chainable filter methods (.approved().future().with_tags()) — more readable than raw SQL Template filters: highlights urlize, linebreaksbr, json_script, and especially querystring for building/modifying query-string links in templates Migrations: still loves Django's auto-generated migrations — 19 and counting on her project Skips inheritance for class-based views; prefers function-based views for sharing code, though fine using Django's own mixins/interfaces Performance surprise: CPU profiling (via py-spy) — not slow DB queries — revealed the culprit; she'd accidentally disabled the cached template loader, and re-enabling it took throughput from ~2-3 req/s to ~12 req/s on a $10/mo VM Michael #2: Who cleans up after the vibe-coding party? FT Magazine piece by Sam Learner (July 11) on AI coding tools overwhelming open source maintainers - sent in by listener Dylan McConnell, whose main point was that this ran in the Financial Times, not a dev blog. cURL as the case study - Daniel Stenberg has been the only full-time person on it for years; libcurl has been installed an estimated 20+ billion times with 3,000+ listed contributors. Bug bounty killed - cURL ended its paid security bounty program in January, citing an "explosion of AI slop reports" that take real time to debunk and drain morale. Extractive contributions - authoring a PR is now nearly free, reviewing one still costs a human; tldraw's Steve Ruiz closed outside contributions entirely, asking why he'd want someone else writing the easy part. Guido weighs in - van Rossum says projects are holding emergency meetings over the slop flow, and notes LLM patches tend to touch unrelated parts of a file, making review more tedious. "Vibe Coding Kills Open Source" - paper from Miklós Koren's group: packages frequently recommended by coding models saw big download jumps with no matching engagement, breaking the reputation loop that sustains maintainers. Stack Overflow flatlined - over 100,000 questions a month before ChatGPT, under 1,500 last month, with the response rate cut roughly in half; the public archive is now stale training data. The course-creator angle - Josh Comeau's newest web dev course launched at about a third of prior enrollment, and he worries about devs who never learn which questions to ask. But the most interesting portion is what was omitted. Focused on: The end of the curl bug-bounty Omitted: High-Quality Chaos Why the omission is interesting It fits a narrative. The FT piece is a maintenance-and-decline story, and January-Stenberg is a perfect witness for it. April-Stenberg complicates it - same person, same project, better data, opposite direction on the specific claim being used. The tell is already in the article. Learner quotes Stenberg saying AI tools are much better at finding problems than fixing them. That's the April thesis in one line, and it goes undeveloped. Reason for the shift is process, not vibes. Killing the bounty removed the cash incentive and the venue change filtered the rest. Worth saying out loud, because "AI reports got better" isn't quite it - "no bounty plus a real triage platform" is closer. Joke too: Sarah O'Connor wrote a related piece (is this just before skynet launches?) Calvin #3: Where Did All Your AI Tokens Go? AgentsView to the rescue! Local-first desktop/web app for browsing, searching, and analyzing your past AI coding agent sessions (Claude Code, Codex, Copilot, Cursor, Gemini, Aider, and dozens more) Auto-discovers session files on your machine — no config needed; everything stored locally in SQLite, no cloud/accounts agentsview usage is a drop-in ccusage alternative — reads from pre-indexed SQLite, reports run 80–220× faster on large histories New Activity dashboard shows peak concurrency, active vs. idle time, agent-minutes, and cost — filterable by project/agent/machine, with a -json CLI report too Full-text + optional semantic search across every session; also imports Claude.ai/ChatGPT chat exports Install via pip install agentsview, uvx agentsview, brew install --cask agentsview, or download desktop binaries from GitHub Releases Michael #4: Careful with phishing all The situation I pass this along because it was a pretty sneaky bit of targeted phishing, and happened to play off an old interaction in bandit's repo. As usual with phishing scams there are a bunch of tells that this isn't legitimate, but just enough plausibility that I could see falling for it in a weak moment. Relative nobodies like me haven't historically been worth the effort to hit with scams this specific. Agents change the game though :-/. Be careful out there folks! Original message From: "Patrick (Blacktrace)" [HTML_REMOVED] To: LISTENER EMAIL Subject: Your Bandit #1350 (B105 NextToken false positive) -- just fixed that exact case Date: Wednesday, July 15, 2026 12:02 AM Hi AJ, Saw your Bandit issue #1350 -- the B105 hardcoded-password false positive on the string NextToken. I build a deterministic gate that filters that class of Bandit noise, and #1350 was literally the case I just fixed: NextToken / next_token / page_token / nextPageToken now stay quiet, while a genuine hardcoded token like api_token="sk-live-..." still fires. Verified against your exact case. 30-second paste: https://blacktrace.co/noise-eraser Where it still trips, published: https://blacktrace.co/kruc Curious whether it clears what you hit -- and if it trips on something of yours, that's the more useful reply. Patrick, Blacktrace I asked Claude for some analysis too. It was pretty good at finding them. The message name-drops enough real detail to feel legit, but the structure is pure phishing - everything in it exists to get AJ onto blacktrace.co. The strongest ones: Freemail sender, corporate signoff. Signs as "Patrick, Blacktrace" but sends from emailpjv@gmail.com. Real company outreach comes from the company domain, not a personal Gmail - and there's no last name. Over-specific targeting. It mirrors AJ's exact public activity - issue #1350, the B105 rule, the NextToken false positive, even the token variants. That's the "just enough plausibility" AJ flagged, and it's exactly what agents make cheap: scrape a GitHub issue, auto-generate tailored bait. Legit cold outreach rarely reads your history back to you this precisely. The entire payload is two links. Strip the technical flattery and the message is just "paste here" plus "see results here." When the whole point of an email is the click, that's the tell. "30-second paste." Low-friction urgency, and "paste" most likely means paste your source into their tool - handing your code to a stranger's site. Exfiltration dressed as convenience. Brand-new, no-reputation domain. blacktrace.co has no track record, and the name is doing some ominous work. The /kruc slug is random noise, not how real product pages get named. Precise-sounding jargon that's actually vague. "Deterministic gate," "noise-eraser" - impressive, empty. Bolted onto correct real details (B105 is the Bandit hardcoded-password test, sk-live- is a Stripe live-key prefix) to borrow credibility. The disarming close. "if it trips on something of yours, that's the more useful reply" - engineered humility that flatters your expertise and baits a response. Makes engaging feel like you're doing them a favor, which drops your guard. Extras Calvin: DjangoCon US 2026 is rapidly approaching, August 24-28, Chicago Ruff v0.16.0 massively expands its default rule set Ruff now enables 413 rules by default, up from 59 https://astral.sh/blog/ruff-v0.16.0 Michael: Completely redesigned the home page. Try /insights in Claude Code (terminal) Joke: We're Safe

AWS Morning Brief
New Tools, Ancient Failures, Same Old AWS

AWS Morning Brief

Play Episode Listen Later Jul 27, 2026 7:51


AWS Morning Brief for the week of July 26th, with Corey Quinn. Links:Amazon ECS now provides Action Logs for deployment and orchestration visibilityAmazon Managed Service for Prometheus supports 1.5B active metrics and 200K rules per workspaceAmazon SES introduces pricing plansAWS Network Load Balancer now supports Listener Rules for custom traffic routingAWS Organizations increases RCP quota to 2,000 per organizationAWS now supports automatic credit memo application preferencesAWS Secrets Manager now publishes secret update notifications to Amazon EventBridgeJVM memory, CPU, and classpath best practices for Java containers on AWSConnection pooling strategies in Amazon Aurora DSQLCustom OS installation now available on AWS DeepRacer devicesThermodynamic sampling of disordered materials with an analog Hamiltonian Rydberg simulatorUnlocking data residency use cases with Amazon S3 in AWS Local ZonesAWS Security Bulletins: New Tools, Ancient Failures