POPULARITY
Categories
As enterprise environments become increasingly distributed and complex, the volume of telemetry data has outpaced human capacity to analyze it. This has been future compounded by the challenges of tokenomics. Jeetu Patel sits down with the founders of Galileo to discuss how we are moving beyond passive dashboards toward a future of autonomous, intelligent agents that can reason through system issues in real time.
As enterprises integrate greater levels of AI into their operational workflows, many are concerned about the effectiveness and accuracy of the decisions with which it's being entrusted. Mike Fratto joins host Eric Hanselman to talk about his recent research into the ways in which observability can provide greater context to AI systems and administrative teams. Data is the raw material from which AI value is created and a lack of solid operational data will mean that management systems won't have a complete perspective. Observability platforms have been expanding to collect and correlate a broader range of data and they can be particularly useful, if they're integrated well. Observability data can also assist in building confidence in automation and support automation efforts by creating scripts and validating control decisions. It can be a huge help to staff to digest and make sense of existing run books and processes and guide decisions in the event of failures. When novel incidents occur, being able to rule out certain areas can be just as powerful as finding the root cause. No, the network isn't always the cause of the outage… More S&P Global Content: Next in Tech | Ep. 271: AI Networking Next in Tech | Ep. 247: Security and Observability For S&P Global subscribers: Updating processes and building trust are key to full operational autonomy – Highlights from VotE: DevOps Observability Market Monitor & Forecast Agent orchestration and observability go hand in hand: The case for governance in motion Host/Author: Eric Hanselman Guest: Mike Fratto Producer/Editor: Dylan Scheible Published With Assistance From: Feranmi Adeoshun and Sophie Carr
SaaS Scaled - Interviews about SaaS Startups, Analytics, & Operations
Today, we're joined by Ariel Assaraf, Co-Founder and CEO of Coralogix, the data and AI platform for observability. We talk about:Why observability was an underutilized data resource for many yearsHow engineering skills will be found in more areas of the organizationFour phases of creating distinct user interfaces for AI agentsThe impacts of widespread improved observabilityPersonalizing insights per person, not persona
David Soria Parra is an Engineering Lead at Anthropic and one of the core maintainers of the Model Context Protocol (MCP). We explore the biggest evolution of the protocol since its launch, and why MCP is becoming the foundation for the next generation of AI agents.We discuss why MCP is moving toward stateless communication, what developers misunderstand about state, sessions, and transport layers, and how lessons from real-world deployments at massive scale have shaped the protocol's future. We also dive into MCP v2, SDK migrations, protocol design, extension architecture, governance, developer experience, and how Anthropic thinks about balancing simplicity with long-term flexibility.Along the way, we explore progressive disclosure, tool search, programmatic tool calling, context bloat, forward compatibility, long-running AI tasks, protocol evolution, open-source governance, observability, and why the future of AI infrastructure will depend on designing protocols that can evolve without breaking the ecosystem.Timestamps:[00:00] Introduction[01:59] Why MCP Had to Become Stateless[04:28] The Tradeoffs of Stateless Design[06:13] What We Learned About Agent State[08:04] Sessions, Models & Implicit State[09:33] Migrating to MCP v2[12:19] Lessons from HTTP & Open Source Standards[18:16] Shipping Fast Without Breaking Everything[20:35] The Future Complexity of MCP[22:44] Core Features vs Extensions[26:47] Progressive Disclosure Explained[28:16] Solving Context Bloat[30:50] Why Tool Search Beats Progressive Disclosure[32:10] The Biggest MCP Anti-Pattern[34:25] Designing for Forward Compatibility[38:41] Why "Tasks" Matter[40:53] JSON, Tokens & Better Tool Calling[44:44] Observability & Tracing AI Agents[47:34] Will MCP Ever Be Finished?[50:22] What's Next for MCP
From the archive: This episode was originally recorded and published in 2022. Our interviews on Entrepreneurs On Fire are meant to be evergreen, and we do our best to confirm that all offers and URL's in these archive episodes are still relevant. Martin Mao is the co-founder and CEO of Chronosphere, provider of the leading observability platform for cloud-native. He previously held roles at leading tech companies including Uber, Microsoft, Google and Amazon. Top 3 Value Bombs 1. Luck and timing play a big role in success. 2. More industries are adopting new technology, creating demand for better solutions. 3. Be ready for opportunities, study trends and position yourself to act. Observability for what's next - Chronosphere.io Sponsors HighLevel - The ultimate all-in-one platform for entrepreneurs, marketers, coaches, and agencies. Learn more at HighLevelFire.com. ThriveTime Show - Is your business stuck? Join Eric Trump and Clay Clark at America's highest rated business conference November 5th and 6th in Tulsa, Oklahoma. Request life-changing tickets today at ThriveTimeShow.com/eofire.
Send us Fan MailAriel Assaraf is the CEO and co-founder of Coralogix, a leading observability platform that most recently raised $115 million at a unicorn valuation. He started the company in 2014, and today Coralogix serves more than 4,000 customers, monitors more than 500,000 applications, and processes over 3 million events per second.Before co-founding Coralogix, Ariel served in Israel's Elite Intelligence Unit 8200, where the sheer scale and complexity of data he worked with made one thing clear: existing architectures weren't built for what was coming. He later joined Varent Systems, a homeland security company, leading automation, integration, and QA.In this episode, Ariel draws on more than a decade of building at the frontier of data infrastructure to argue that observability is no longer just a tool for preventing downtime. It is becoming the most truthful, real-time source of intelligence any company owns.In this conversation, we discuss:The evolution from the data collection era ("oil phase") to a landfill crisis, and now the AI-driven "brain phase" where telemetry has become the most valuable raw material for business decision makingHow Coralogix separates the data plane from the control plane, storing data in open format on the customer's own infrastructure to enable data ownership, infinite retention, and freedom from vendor lock-inWhy Ariel invested over $100 million in R&D to build query engines that always return answers, not constrained by predefined schemas that agents will quickly exceed How SREs evolve from reactive incident responders into autonomous operators as agents like Ollie, Coralogix's AI agent, take over incident triage, root cause analysis, and narrative generationThe emerging role of the AI-forward product manager who sits between customer needs and autonomous agents, reshaping how software gets built, priced, and sold in real timeHow Ariel thinks about linear versus exponential impact as a leadership principle, and why the intuition to prioritize exponential value is something agents will never replicateExplore more in the conversation:00:00 Welcome & AI's societal impact paradox01:53 AI Fun Fact: New data on productivity and employment04:15 Introducing Ariel Assaraf and the Coralogix origin story05:15 Evolving data architecture: from oil to landfill to brain phase07:28 Coralogix's data ownership and open format advantages13:27 From telemetry data lake to autonomous, agent-driven analytics20:00 Building scalable, answer-guaranteeing query engines for complex data25:44 The future of natural language interfaces like Olly for SREs and DevOps28:15 Orchestrating multiple AI agents for better decision-making30:20 Responsibility, autonomy, and the evolving role of customer success35:37 How Coralogix turns telemetry into strategic business decisions38:18 Linear vs. exponential value in your careerResources:Subscribe to the AI & The Future of Work NewsletterConnect with Ariel on LinkedInAI fun fact articleOn How Personalized Healthcare Is Being Transformed Through AI and the Human Microbiome
"There is no single way to deploy OpenTelemetry at scale—and that's exactly the challenge."As organizations adopt OTel across teams and environments, they face tough questions around standardization, configuration, and operating resilient observability pipelines.To address these challenges, the OpenTelemetry community has introduced Blueprints and Reference Implementations—practical guidance on topics like data standards, consistent agent and collector configuration, pipeline resilience, and intelligent sampling.In this episode, we're joined by Dan Gomez Blanco, maintainer of the OpenTelemetry End-User SIG, to explore real-world reference architectures from organizations like Skyscanner, Adobe, and Mastodon.Tune in to learn how the community is turning OTel complexity into shared best practices—and how you can contribute your own blueprint
Ameet Talwalkar, Carnegie Mellon ML professor and Chief Scientist at Datadog, joins Ben Lorica to trace time series foundation models from skepticism to Datadog's Toto V1 and V2. Subscribe to the Gradient Flow Newsletter
Take a Network Break! We start with a critical vulnerability in Adobe Coldfusion. On the news front, Infoblox acquires Kentik to add network observability to its portfolio, data center electricity consumption jumps worldwide, and Exabeam rolls out AI-agent focused detection in its Agent Behavior Analytics platform. DriveNets and WhiteFiber connect two AI data centers over... Read more »
Take a Network Break! We start with a critical vulnerability in Adobe Coldfusion. On the news front, Infoblox acquires Kentik to add network observability to its portfolio, data center electricity consumption jumps worldwide, and Exabeam rolls out AI-agent focused detection in its Agent Behavior Analytics platform. DriveNets and WhiteFiber connect two AI data centers over... Read more »
Take a Network Break! We start with a critical vulnerability in Adobe Coldfusion. On the news front, Infoblox acquires Kentik to add network observability to its portfolio, data center electricity consumption jumps worldwide, and Exabeam rolls out AI-agent focused detection in its Agent Behavior Analytics platform. DriveNets and WhiteFiber connect two AI data centers over... Read more »
Join us as John Mark Troyer and Rakesh Gupta break down what AI observability actually means once agents leave the demo and hit production - and why the old playbook for monitoring doesn't cut it anymore. John Mark and Rakesh walk through why errors and latency are just the starting point for agents, how quality became a much harder thing to measure once bots went from answering questions to taking autonomous action, and why token-based costs are creating a confusing new economics problem for engineering teams. You'll learn the difference between online and offline evals, why a new engineering role has emerged just to build testing harnesses for agents, how trace data works differently when every prompt is its own trace, and what teams are doing to catch prompt injection and other AI-specific failure modes before they become expensive mistakes. Timestamps 0:00 Welcome & Introduction 3:20 Full Disclosure - Observe, Snowflake, and How This Conversation Started 7:07 From Developer Concerns to Boss's Boss's Boss - Spending Out of Control 8:29 What Actually Gets Measured - Errors, Latency, Quality, and Cost 10:30 The Casino Chip Problem - Confusing Token Pricing Models 13:47 Defining Quality When the Task Itself Is Nebulous 18:41 The New Role - Engineers Who Just Build Testing Harnesses 22:00 Non-Determinism and Why Testing Agents Is Expensive 32:10 Trace Data, Tool Calls, and What Observability Tools Actually See 55:08 Prompt Injection, Zero-Width Characters, and Real World Failures How to find John Mark: https://www.linkedin.com/in/johnmarktroyer/ How to find Rakesh: https://www.linkedin.com/in/rg0/ Links from the show:
In this episode of The IT Experts Podcast, I hosted an MSP Insights Roundtable on AI, automation, and observability at scale, bringing together three brilliant guests, Joe Burns, Fiona Challis, and Nick Horner. What struck me straight away was how all three agreed on one thing before we even got into the detail. AI, automation, and observability only work when you understand your own processes first. Joe walked us through how his MSP, Reformed, built its operational maturity by identifying repetitive tasks, spotting where human error crept in, and asking his team what they actually disliked doing. Only once that groundwork was done did he bring in automation, including an early AI triage system on the service desk, and it paid off. Nick added a perspective I loved, describing how starting small with clients avoids the scope creep that can derail an automation project before it even gets going. He shared a story about a modest HR automation that grew organically once the client saw the value for themselves, and how bringing end users into the process from day one builds the kind of trust that makes AI, automation, and observability actually stick. Getting genuine buy in, as Nick put it, turns a nervous stakeholder into a project sponsor rather than a blocker. Fiona introduced an idea I keep coming back to, becoming your own customer zero, assessing your own readiness before you ever take an AI conversation to a client. She told us most MSPs score only two or three out of five on her readiness assessment, which shows how much foundational work is still undocumented across our sector. We spent time discussing how observability has changed, moving away from juggling dashboards across Microsoft 365, PSA, and RMM tools towards a single intelligent layer that pulls everything together securely and quickly. Nick made the point that speed and accuracy no longer need a dedicated Power BI specialist, and Fiona reminded us that the ROI conversation always starts with measuring a baseline before you change anything. One of my favourite moments came from Joe, describing a law firm that spent three to four hours every week cross checking court lists against their case management system, a task his team solved with an agent in fifteen minutes. Fiona echoed this, encouraging MSPs to lead with one practical win rather than an overwhelming pitch, because solving a single small problem tends to open the door to many more conversations. We also got into the tension between compliance and outcome. Fiona argued that clients buy the outcome AI delivers rather than the technology itself, and Joe pushed back with honest feedback from his law firm clients, who want compliance answered first given how sensitive that sector is to reputational risk. Both agreed that governance needs to be built into the foundations of any deployment rather than bolted on afterwards. Drawing on Daniel Priestley's thinking around demand and supply tension, Joe warned us against pouring all our energy into operational capacity while neglecting sales and marketing, a gap that can quietly erode margins even as efficiency improves. Fiona picked this up with real enthusiasm, describing how AI and automation can make selling feel far more natural, framing every client conversation as a business problem to solve rather than a service to pitch. She introduced the three pillars she coaches MSPs towards, capacity, experience, and revenue, and stressed the importance of owning your intellectual property rather than giving away your hard built agents for free. We closed by sharing how each of us measures success, from outcome-based tracking to client and employee satisfaction scores, before final takeaways. Nick urged everyone to embrace the shift and stay ahead of the curve, Joe reminded us that capacity means nothing without the ability to sell it, and Fiona encouraged listeners to stop overthinking and take the first step. I came away from this session convinced, more than ever, that AI, automation, and observability can genuinely transform an MSP, provided the fundamentals are respected along the way. Connect with Fiona Challis through LinkedIn and website. Connect with Joe Burns through LinkedIn and website. Connect with Nick Horner through LinkedIn. Make sure to check out our Ultimate MSP Growth Guide, a free guide that walks you through a proven process to take your MSP from stuck to scalable, without working even more hours. It's 44 pages rammed with advice, insights and inspiration to help you decide what support is available to you now if you want to grow and scale your business. Click HERE to get your copy. Connect on LinkedIn HERE with Ian and also with Stuart by clicking this LINK And when you're ready to take the next step in growing your MSP, come and take the Scale with Confidence MSP Mastery Quiz. In just three minutes, you'll get a 360-degree scan of your MSP and identify the one or two tactics that could help you find more time, engage & align your people and generate more leads. If you're serious about growth and want to explore what this could look like for your MSP, you can book a Right Fit Clarity Call with us HERE. OR To join our amazing Facebook Group of over 400 MSPs where we are helping you Scale Up with Confidence, then click HERE Until next time, look after yourself and I'll catch up with you soon!
Join us as Du'An digs into the real mechanics of running AI locally and in production - from GPU memory math to multi-agent architectures, observability, and the economics of self-hosted inference. Du'An walks through how model weights and KV cache compete for GPU memory, why continuous batching matters when you have more than a handful of users, and how agent architectures like single-agent, workflow, graph, swarm, and supervisor patterns each solve different problems. You will learn how to instrument your agents with Langfuse for observability and cost tracking, when to use Ollama versus vLLM, how prompt caching can cut provider costs by up to 75%, and why GPUs should never sit idle. Episode two of three - the next episode covers deploying at scale. Timestamps 0:00 Welcome & Introduction 1:47 Du'An's New Role at Akamai Cloud 3:10 Data Privacy and the Case for Self-Hosted AI 7:21 Anthropic and OpenAI as the New Cloud Layer 12:48 Local Models for Specific Use Cases - Cancer Detection Example 15:02 GPU Memory Math - Weights, KV Cache, and Context Windows 19:32 Continuous Batching and GPU Time Slicing 20:03 Observability with Langfuse - Live Demo 27:44 Agent Architectures - Single Agent, Workflow, Graph, Swarm, Supervisor 36:36 Token Economics, Prompt Caching, and GPU Cost Planning 45:32 Ollama vs vLLM - Prototyping vs Production How to find Du'An: https://duanlightfoot.com https://www.linkedin.com/in/duanlightfoot/ Links from the show: https://langfuse.com/ https://github.com/akamai-developers/akamai-workshop-solution-architect-agent https://amzn.to/4bvHn1p https://vllm.ai/
I still hear people say, “OpenTelemetry is vendor-neutral, so you can switch any time!”In this episode, Adriana Villela and Josh Lee (both active OpenTelemetry contributors) help bust that myth.While OTel standardizes instrumentation and signal transport—and unlocks a rich ecosystem of tools—switching vendors isn't as simple as it sounds. There's real cost in retraining engineers, migrating dashboards, SLOs, and alerts, and reworking deep integrations across your delivery pipeline.We also dive into a key challenge the community is tackling: helping engineers instrument by value, not by default—making it easier to capture the right signals with high quality instead of just collecting everything.Here the links we discussed:Adriana's LinkedIn: https://www.linkedin.com/in/adrianavillela/Josh's LinkedIn: https://www.linkedin.com/in/joshuamlee/The blog article: https://thenewstack.io/opentelemetry-vendor-neutrality-guide/CND Austria Talk: https://www.youtube.com/watch?v=1gxLseuaTdMKCD Prague Talk: https://www.youtube.com/watch?v=pPXG20CXKxQOpenTelemetry Project Website: https://opentelemetry.io/
Advanced software systems have long been more complex than any single engineer can fully understand. Observability is the established solution to this problem, but with AI agents now generating code, deploying changes, and operating autonomously, the challenge of understanding large software systems is entering a new dimension. Grafana is an open source observability platform, and The post Grafana's Approach to AI-Native Observability appeared first on Software Engineering Daily.
Advanced software systems have long been more complex than any single engineer can fully understand. Observability is the established solution to this problem, but with AI agents now generating code, deploying changes, and operating autonomously, the challenge of understanding large software systems is entering a new dimension. Grafana is an open source observability platform, and one of the most widely used in the world. The company builds tools that help teams collect, visualize, and act on telemetry data across logs, metrics, and traces. They are now extending that capability into the agentic era with AI-powered investigation and monitoring tools. Anthony Woods is a co-founder of Grafana Labs. In this episode, he joins Matt Merrill to discuss how AI-generated code is straining software operations, why telemetry data volume has become as much a problem as a solution, how Grafana is adapting to a world where agents are the primary consumers of observability data, and what keeps him up at night about where the industry is headed. Sponsorship inquiries: sponsor@softwareengineeringdaily.com The post Grafana’s Approach to AI-Native Observability appeared first on Software Engineering Daily.
Advanced software systems have long been more complex than any single engineer can fully understand. Observability is the established solution to this problem, but with AI agents now generating code, deploying changes, and operating autonomously, the challenge of understanding large software systems is entering a new dimension. Grafana is an open source observability platform, and one of the most widely used in the world. The company builds tools that help teams collect, visualize, and act on telemetry data across logs, metrics, and traces. They are now extending that capability into the agentic era with AI-powered investigation and monitoring tools. Anthony Woods is a co-founder of Grafana Labs. In this episode, he joins Matt Merrill to discuss how AI-generated code is straining software operations, why telemetry data volume has become as much a problem as a solution, how Grafana is adapting to a world where agents are the primary consumers of observability data, and what keeps him up at night about where the industry is headed. Sponsorship inquiries: sponsor@softwareengineeringdaily.com The post Grafana’s Approach to AI-Native Observability appeared first on Software Engineering Daily.
Advanced software systems have long been more complex than any single engineer can fully understand. Observability is the established solution to this problem, but with AI agents now generating code, deploying changes, and operating autonomously, the challenge of understanding large software systems is entering a new dimension. Grafana is an open source observability platform, and one of the most widely used in the world. The company builds tools that help teams collect, visualize, and act on telemetry data across logs, metrics, and traces. They are now extending that capability into the agentic era with AI-powered investigation and monitoring tools. Anthony Woods is a co-founder of Grafana Labs. In this episode, he joins Matt Merrill to discuss how AI-generated code is straining software operations, why telemetry data volume has become as much a problem as a solution, how Grafana is adapting to a world where agents are the primary consumers of observability data, and what keeps him up at night about where the industry is headed. Sponsorship inquiries: sponsor@softwareengineeringdaily.com The post Grafana’s Approach to AI-Native Observability appeared first on Software Engineering Daily.
Advanced software systems have long been more complex than any single engineer can fully understand. Observability is the established solution to this problem, but with AI agents now generating code, deploying changes, and operating autonomously, the challenge of understanding large software systems is entering a new dimension. Grafana is an open source observability platform, and The post Grafana’s Approach to AI-Native Observability appeared first on Software Engineering Daily.
Federal Tech Podcast: Listen and learn how successful companies get federal contracts
In this episode of the Federal Tech Podcast, John Gilroy interviews Justin Fessler, Vice President of Public Sector at LogicMonitor, about the growing role of autonomous AI and observability in federal government IT operations. Fessler explains that autonomous AI is not about replacing people but about automating repetitive operational tasks, correlating complex system events, and helping IT teams make faster, better-informed decisions. Rather than allowing AI to operate without oversight, LogicMonitor focuses on keeping humans "in the loop" while AI handles time-consuming analysis and routine remediation. A central theme is the importance of complete visibility across increasingly complex federal environments. LogicMonitor's agentless monitoring technology discovers devices, cloud resources, applications, and shadow IT without requiring software agents on every endpoint. This broad visibility enables agencies to identify unmanaged assets, reduce blind spots, optimize cloud costs, and strengthen security. The discussion also highlights observability's critical role in Zero Trust. Fessler notes that agencies cannot secure or verify assets they cannot see. By discovering everything connected to the network—including servers, cloud services, IoT devices, cameras, badge readers, and physical infrastructure—LogicMonitor helps agencies build a stronger Zero Trust foundation. Gilroy and Fessler examine the challenges of managing hybrid and multi-cloud environments, emphasizing that agencies require a single operational view regardless of where workloads reside. LogicMonitor integrates information across cloud providers and third-party platforms, including ServiceNow, Splunk, Dynatrace, Datadog, and IBM Watsonx, enabling AI-driven event correlation and faster incident response. The conversation concludes with the future of autonomous IT. Fessler predicts increased automation, AI-assisted self-healing infrastructure, and significantly reduced mean time to identify and resolve incidents. Rather than replacing IT professionals, autonomous AI will eliminate repetitive work, allowing skilled personnel to focus on higher-value mission objectives while improving operational resilience, reducing alert fatigue, and delivering better digital services to citizens. For more information, visit www.logicmonitor.com/solutions/federal-government.
Когда у тебя 3 миллиарда семплов в секунду, 60 гигабайт логов и 44 миллиона спанов — ты уже не «настраиваешь мониторинг», ты его пишешь с нуля. В гостях Владимир Гордийчук, CTO Yandex Monium — системы наблюдаемости, которая выросла внутри Яндекса, а теперь доступна как отдельный продукт. 9 лет разработки, большая часть кодовой базы написана лично, и чёткое понимание, почему Prometheus + Grafana — это не всегда ответ. Кто: Владимир Гордийчук — CTO Yandex Monium Что обсудим: Зачем писать свой мониторинг, когда есть Prometheus, Grafana и ELK 3 млрд семплов/сек — какие архитектурные подходы позволяют держать такие нагрузки Сколько стоит мониторить всё — и как это обосновать перед менеджментом Alert fatigue: когда алертов столько, что на них перестают смотреть, и что с этим делать Кто мониторит мониторинг, когда мониторинг падает Самые запоминающиеся инциденты — и почему после них мониторинг стал другим Рекомендации, которые можно внедрить уже завтра Для тех, кто хочет мониторить, а не тонуть в дашбордах и ложных срабатываниях. Оставайтесь на связи Пишите нам: info@linkmeup.ru Канал в телеграме: t.me/linkmeup_podcast Канал на youtube: youtube.com/c/linkmeup-podcast Подкаст доступен в iTunes, Google Подкастах, Яндекс Музыке, Castbox Сообщество в вк: vk.com/linkmeup Группа в фб: www.facebook.com/linkmeup.sdsm Добавить RSS в подкаст-плеер. Пообщаться в общем чате в тг: https://t.me/linkmeup_chat Поддержите проект:
Cut through alert noise and move from detection to root cause using the Azure Copilot Observability Agent. It autonomously investigates incidents, correlates signals across logs, metrics, alerts, application health, and ML anomalies, then surfaces root cause with charts and recommended next steps. Extend coverage to your AI agents in Microsoft Foundry, track Gen AI errors and token consumption with trace-level detail, and write plain-language instructions to tune autonomous behavior to match your team's workflow. Matt McSpirit, Microsoft Azure expert, shares how to take full control of incident response at scale. ► QUICK LINKS: 00:00 - Azure Copilot Observability Agent 00:43 - How to use it as you work 01:33 - Unified Full-Stack Telemetry 02:39 - Root Cause Investigation 04:12 - Investigate further 04:55 - Re-run the investigation 05:36 - Autonomous Alert Correlation & Triage 07:13 - Natural Language Agent Customization 07:34 - Wrap up ► Link References Get started at https://aka.ms/ObservabilityAgent ► Unfamiliar with Microsoft Mechanics? As Microsoft's official video series for IT, you can watch and share valuable content and demos of current and upcoming tech from the people who build it at Microsoft. • Subscribe to our YouTube: https://www.youtube.com/c/MicrosoftMechanicsSeries • Talk with other IT Pros, join us on the Microsoft Tech Community: https://techcommunity.microsoft.com/t5/microsoft-mechanics-blog/bg-p/MicrosoftMechanicsBlog • Watch or listen from anywhere, subscribe to our podcast: https://microsoftmechanics.libsyn.com/podcast ► Keep getting this insider knowledge, join us on social: • Follow us on Twitter: https://twitter.com/MSFTMechanics • Share knowledge on LinkedIn: https://www.linkedin.com/company/microsoft-mechanics/ • Enjoy us on Instagram: https://www.instagram.com/msftmechanics/ • Loosen up with us on TikTok: https://www.tiktok.com/@msftmechanics
With enterprises now rushing to integrate AI agents into their operations and security, the most imperative focus now becomes the AI model itself. However, Eric Tschetter, Chief Architect at Imply, believes the real challenge is within the data infrastructure that supports these systems.In the recent episode of the Tech Transformed podcast, Kevin Petrie, BARC Vice President of Research, sat down with Tschetter to talk about how AI is actually increasing the current needs around scale, performance, and data access.“Agents are always running queries. They're always doing stuff,” Tschetter stated.Unlike human analysts, AI systems work continuously, producing much higher query volumes and putting more pressure on the data platforms underneath. This leads to a greater demand for observability architectures that can manage more data, more users, and more machine-to-machine interactions without losing speed.For Tschetter, the solution is not to create new observability tools, but to rethink the data layer that supports them.Key TakeawaysAI is transforming observability and security disciplines.The observability warehouse concept is gaining traction.AI agents increase the volume of queries significantly.Data silos remain a major challenge for enterprises.Collaboration between IT and security teams is essential.Observability and security teams often consume the same data.A decoupled architecture can enhance data accessibility.The semantic layer must support multiple query languages.Effective data management is crucial for AI-driven workloads.Data should be stored once and accessed from multiple platforms.Chapters00:00 Introduction to AI and Observability02:08 Challenges in Observability with AI06:44 Modernising Architecture for Observability10:49 Decoupled Observability and Semantic Layers16:31 Collaboration Between IT and Security Teams22:23 Imply's Observability Warehouse and Data LakesFor more information on AI, observability and Imply's observability warehouse and data lakes, please visit imply.io.For further information on all things B2B Tech, please visit em360tech.comImply LinkedIn: @Imply Imply X: @implydataImply YouTube: @ImplydataEM360Tech YouTube: @enterprisemanagement360EM360Tech LinkedIn: @EM360TechEM360Tech X: @EM360TechFollow: @EM360Tech on YouTube, LinkedIn and XStay connected for more expert insights, podcast episodes, and enterprise data strategy discussions
Foundry Observability empowers developers to take agents from prototype to production with end-to-end observability. We demonstrate how evaluations, tracing, monitoring, and optimization work together to identify issues, measure quality, and improve outcomes over time. Through a live demo spanning the Foundry portal and VS Code, we showcase a practical workflow for building more reliable, production-ready agents. Chapters 00:00 - Introduction 01:33 - Foundry portal demo 06:58 - VS Code demo 17:43 - Wrap up Recommended resources Observability in Generative AI - Microsoft Foundry Build 2026 Resources Connect Scott Hanselman | Twitter/X: @SHanselman Azure Friday | Twitter/X: @AzureFriday
Foundry Observability empowers developers to take agents from prototype to production with end-to-end observability. We demonstrate how evaluations, tracing, monitoring, and optimization work together to identify issues, measure quality, and improve outcomes over time. Through a live demo spanning the Foundry portal and VS Code, we showcase a practical workflow for building more reliable, production-ready agents. Chapters 00:00 - Introduction 01:33 - Foundry portal demo 06:58 - VS Code demo 17:43 - Wrap up Recommended resources Observability in Generative AI - Microsoft Foundry Build 2026 Resources Connect Scott Hanselman | Twitter/X: @SHanselman Azure Friday | Twitter/X: @AzureFriday
This presentation was recorded at GOTO Copenhagen 2025.https://gotocph.comAbby Bangser - Principal Engineer at Syntasso & Team Topologies AdvocateDave Farley - Bestselling Author, Founder & Director of Continuous Delivery Ltd.RESOURCESAbbyhttps://bsky.app/profile/abangser.bsky.socialhttps://twitter.com/a_bangserhttps://github.com/abangserhttps://www.linkedin.com/in/abbybangserhttps://www.syntasso.io/members-area/abby/profileDavehttps://bsky.app/profile/davefarley77.bsky.socialhttps://www.continuous-delivery.co.ukhttps://linkedin.com/in/dave-farley-a67927https://twitter.com/davefarley77http://www.davefarley.netDESCRIPTIONDave Farley and Abby Bangser open with a clear statement: Continuous Delivery isn't a relic of the pre-AI era — it's the foundation that makes the AI era survivable. Dave's definition is simple but consequential: software should always be in a releasable state, verified after every small change. That's not just a workflow preference; it's the same incremental, hypothesis-driven approach that underpins science and engineering. In an AI-assisted world where code can be generated far faster than humans can reason about it, the discipline of small, safe, verifiable steps becomes more critical, not less. The danger isn't AI writing bad code — it's AI writing a lot of code very fast that nobody is properly checking.The conversation turns to a genuinely alarming DORA report statistic: 70% of developers using AI tools don't distrust the output. Abby draws a parallel to the long-running debate over whether developers can be trusted to test their own code — they usually can't, without a deliberate change in perspective. The same challenge applies to AI-generated code: you need to consciously shift from "prompter" mode to "verifier" mode, and most developers aren't making that switch. Dave closes with a surprising note of optimism: AI may be the industry's best-ever opportunity to finally get XP practices — small increments, automated tests, continuous feedback — embedded into how teams actually work. Not because anyone chose to adopt them ideologically, but because working without them while using AI is visibly, measurably risky.Read the full abstract here:https://gotocph.com/2025/sessions/3779RECOMMENDED BOOKSKief Morris • Infrastructure as Code • https://amzn.to/4e6EBQcMatthew Skelton & Manuel Pais • Team Topologies • http://amzn.to/3sVLyLQDave Thomas • simplicity • https://amzn.to/43FghBJDave Farley & Jez Humble • Continuous Delivery • https://amzn.to/3ocIHwdDavid Farley • Modern Software Engineering • https://amzn.to/3GI468MDave Farley • Continuous Delivery Pipelines • https://amzn.to/3rjetdiBlueskyInstagramLinkedInFacebookCHANNEL MEMBERSHIP BONUSJoin this channel to get early access to videos & other perks:https://www.youtube.com/channel/UCs_tLP3AiwYKwdUHpltJPuA/joinLooking for a unique learning experience?Attend the next GOTO conference near you! Get your ticket: gotopia.techSUBSCRIBE TO OUR YOUTUBE CHANNEL - new videos posted daily!
Before he founded Render, Anurag Goel was the fifth engineer at Stripe, where he watched roughly a fifth of the engineering team disappear into managing AWS, writing brittle, repetitive, error-prone infrastructure scripts that had nothing to do with the actual product. That experience became the seed for Render: a platform that automates away the undifferentiated DevOps work and lets application teams ship without standing up their own cloud team. Today, millions of developers build on it, and Render has raised over $260M from Bessemer and General Catalyst. In this episode, Tobi and Anurag get into what's actually changing as AI moves from hype to production. Anurag makes the case that agents are simply a new kind of application, long-running, stateful, tool-heavy, and a new kind of end user you have to design for. He explains why Render deliberately refuses the "AI cloud" label, what he's building with Workflows and sandboxes, and why the hardest part of shipping agents isn't building them but seeing inside them. The conversation also goes wide: how to hire executives when interviews lie, why short-lived keys and blast-radius thinking matter more than container escapes, how distribution is shifting from SEO to getting ChatGPT and Claude to recommend you, and why, despite all the "SaaS is dead" noise, specialization isn't going anywhere. Topics covered: Why ~20% of Stripe's engineers were stuck managing AWS and how that became Render "We're not the AI cloud, we're the application cloud," and why the distinction matters Agents, as a new type of application (and a new end user), you have to build for Render Workflows and sandboxes: the consolidated AI runtime Hiring executives when interviews are an imperfect signal Security as blast-radius management: short-lived keys over "admin forever" The shift from SEO to GEO, getting chatbots to recommend your product Why SaaS isn't dying, and specialization still wins
OpenChoreo is an opinionated, “batteries included”, AI-native Kubernetes platform stack for Platform Engineers that combines GitOps, Observability, AI Agents, and Workflows into a custom K8s distribution “super pack” that is managed via Backstage, CLI, API, or MCP. Now a CNCF project.Check out the video podcast version here:
Brandon talks with OpenObserve's Prabhat Sharma and Shani Shoham: why observability is still broken, how they fixed it, and where AI takes it next. Watch the YouTube Live Recording of Episode 576 Show Links OpenObserve OpenObserve on GitHub Series A and Observability 3.0 announcement blog post Launching OpenObserve OpenObserve 2-Minute Demo Download OpenObserve Contact Prabhat Sharma LinkedIn: hiprabhat Twitter/X: @hiprabhat Contact Shani Shoham LinkedIn: shanishoham Twitter: @shohams SDT News & Hype Join us in Slack. Get a SDT Sticker! Send your postal address to stickers@softwaredefinedtalk.com and we will send you free laptop stickers! Follow us: Twitch, Twitter, Instagram, Mastodon, BlueSky, LinkedIn, TikTok, Threads and YouTube. Use the code SDT to get $20 off Coté's book, Digital WTF, so $5 total. Become a sponsor of Software Defined Talk! Special Guests: Prabhat Sharma and Shani Shoham.
At Infosecurity Europe 2026 in London, Bill Peterson, Senior Director of Product Marketing at Sumo Logic, joins us to unpack a tension every regulated security team knows well. When an incident hits, the business has to keep running. At the same time, regulators expect sensitive data to stay in region. For a long time, those two demands have pulled in opposite directions. Sumo Logic has spent 15 years as a SaaS platform on AWS, processing roughly four exabytes of data a day for around 2,000 customers. The core promise is speed, driving mean time to resolve as low as possible. Peterson frames it in business terms, because the person signing the check wants to know the return, not the bits and bytes. The news from the show is Sumo Logic availability on the AWS European Sovereign Cloud. EU organizations can keep their data in region, handled by EU staff, while still running the full platform for incident response. That turns a painful either/or into a checklist a regulated buyer can complete. Genesys is the first customer live in the sovereign cloud, with payment processor OpenPay preparing to follow. How does this play out for highly regulated industries? Sumo Logic is focused on finance, healthcare, telco, and government, the verticals feeling the most pressure. The path Peterson describes is simple: let Sumo Logic handle incident management, let AWS move and grow the data in region, and check the sovereignty box without giving up operational readiness. Underneath sits a full-featured SIEM and Dojo AI, the agentic approach Sumo Logic launched earlier this year. The goal is not to replace analysts but to keep a human in the loop while handing proven, repetitive work to an agent. Fix one server, confirm the solution, then let an agent patch the other 599 under oversight. A SOC Analyst Agent reaches general availability at Black Hat later this year, alongside an MCP server. On observability, the differentiator is reading both structured and unstructured data without normalizing it first. A zip code is structured; a cryptic web hook error is not. Sumo Logic reads both, which feeds directly into faster time to identify and faster time to resolve. For any leader weighing sovereignty against uptime, Bill Peterson makes a clear case that they can finally live in the same plan. This is a Brand Spotlight. A Brand Spotlight is a ~15 minute conversation designed to explore the guest, their company, and what makes their approach unique. Learn more: https://www.studioc60.com/creation#spotlight GUEST Bill Peterson, Senior Director of Product Marketing, Sumo Logic LinkedIn: https://www.linkedin.com/in/williampetersonjr/ RESOURCES Learn more about Sumo Logic: https://www.sumologic.com/ Sumo Logic on the AWS European Sovereign Cloud (announced at Infosecurity Europe 2026): https://www.sumologic.com/newsroom Infosecurity Europe 2026 event coverage: https://www.itspmagazine.com/infosecurity-europe-2026-infosec-london-cybersecurity-event-coverage Are you interested in telling your story? ▶︎ Full Length Brand Story: https://www.studioc60.com/content-creation#full ▶︎ Brand Spotlight Story: https://www.studioc60.com/content-creation#spotlight ▶︎ Brand Highlight Story: https://www.studioc60.com/content-creation#highlight ▶︎ Get your own Brand Briefing at an upcoming event: https://www.studioc60.com/buy-brand-briefings KEYWORDS Bill Peterson, Sumo Logic, Sean Martin, brand story, brand marketing, marketing podcast, brand spotlight, AWS European Sovereign Cloud, data sovereignty, incident response, mean time to resolve, SIEM, security operations, Dojo AI, agentic AI, SOC analyst agent, observability, log analytics, Infosecurity Europe 2026 Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
As AI matures, it becomes increasingly important to know how it's performing and what it actually costs. Ned and Kyler are joined by Anuj Tyagi, Senior Site Reliability Engineer for RingCentral, to discuss the critical shift toward AI observability. AI observability is not just about costs; Anuj breaks down why observability has to include agent... Read more »
As AI matures, it becomes increasingly important to know how it's performing and what it actually costs. Ned and Kyler are joined by Anuj Tyagi, Senior Site Reliability Engineer for RingCentral, to discuss the critical shift toward AI observability. AI observability is not just about costs; Anuj breaks down why observability has to include agent... Read more »
As AI matures, it becomes increasingly important to know how it's performing and what it actually costs. Ned and Kyler are joined by Anuj Tyagi, Senior Site Reliability Engineer for RingCentral, to discuss the critical shift toward AI observability. AI observability is not just about costs; Anuj breaks down why observability has to include agent... Read more »
In this episode of Data Driven, we're diving into the rapidly evolving world of agentic AI—where autonomous AI agents collaborate, communicate, and occasionally collide. Our guest, Vlad Luzin, co-founder and CTO of Band, joins us to explore the technical challenges and real-world implications of building collaboration layers for agents that act like distributed, non-deterministic microservices. We'll unpack the myths and realities surrounding orchestration, governance, and security, and discuss how enterprises can operationalize these agent ecosystems safely. Tune in as we share lessons learned, amusing engineering mishaps, and get a glimpse of what the future holds as agents become everyday colleagues in the digital enterprise.LinksVlad's LinkedIn Profile -https://www.linkedin.com/in/luzin/Watch this episode on YouTube -https://youtu.be/MZztFagEX_EBand Website -https://www.band.ai/Band Docs -https://docs.band.ai/Time Stamps00:00 Explaining orchestration in tech03:42 Understanding models and harnesses09:38 Misconceptions about A2A communication10:41 Understanding multi-agent systems16:18 Observability for distributed systems18:54 Agent communication and collaboration24:28 Unauthorized agent interactions25:49 Remote agent collaboration ideas28:54 How foundational AI models communicate33:20 Agent communication protocols overview35:39 Discussing tech standards and AI velocity40:53 Learning to Work with AI Agents42:41 Using Band AI tools
Its rare - but it happens: A guest-free episode of PurePerformance, allowing Andi Grabner and Brian Wilson reconnect to share real-world insights from recent months in the cloud-native and observability space. From KubeCon Amsterdam experiences and the strength of open-source collaboration to emerging challenges like AI-generated contributions, they explore how the industry is evolving beyond the hype.Your co-hosts of PurePerformance discuss the changing role of observability in the AI-native era—both as a foundation for understanding complex systems and as a tool to monitor AI itself. Brian shares his personal shift from AI skepticism to practical adoption, highlighting how AI can significantly improve productivity when used thoughtfully.Hope you all enjoy this episode!
The Pure Report welcomes Mark Wilkinson, a Consulting Field Solutions Architect at Everpure and a former Database Administrator (DBA) and manager. Mark shares his unique perspective on the changes reshaping the Database Administrator role from the perspective of a DBA practitioner. Drawing on his experience as a 10-year Everpure customer who was freed from storage concerns, Mark highlights that the DBA function has not been eliminated but rather has been elevated and broadened in scope. Mark explains how the role continues to shift from routine, fire-fighting tasks to high-value, strategic contributions. The modern DBA role is expanding beyond traditional relational databases and SQL Server dominance, now intersecting with big data, AI, and unstructured data. We discuss how adopting technologies like cloud for data mobility, containers (which force teams to prioritize resilience), and automation (leading to self-service workflows) creates more time for the DBA team to grow their expertise. Automation, often driven initially by laziness, is seen as the key force multiplier, enabling DBAs to stop asking "Am I adding any value right now?" and start using their knowledge to benefit the business. Crucially, the entire evolution points to the necessity of building stronger relationships throughout the organization—with developers, finance, and leadership. This shift allows DBAs to move from a stereotypical gatekeeper role to a business partner, gaining a seat at the table and increasing their visibility and impact. While new challenges like AI accuracy (especially for new DBAs) and compliance (GDPR) exist, the expansion of the role makes it a cool time to be a DBA, with many options to specialize, build skills (e.g., via open source), and drive corporate success. To learn more, visit: https://www.everpuredata.com/solutions/databases.html Check out the new Everpure digital customer community to join the conversation with peers and Everpure experts: https://purecommunity.purestorage.com/ 00:00 Intro and Welcome 05:05 Career Journey 09:55 Everpure Benefit for App Environments 15:01 Stat of the Episode 20:05 Slow Storage Impact on DBAs 25:05 Key Changes to DBA role 30:15 Containers and DBAs 35:15 Automation and Workflows 41:10 Observability and Telemetry 43:43: AI and DBAs 55:08 Hot Takes
Cisco Splunk: Agentic Observability, Token Economics and the Smaller War Room, Podcast, Cisco and Splunk are focused: helping customers bring the right information together, with the right context, so AI can be useful rather than overwhelming By Doug Green “The real opportunity is helping customers pull together all the different sources of data into an environment where they can understand when they need to pay attention, how to find and fix problems, and how to layer AI on top of that.” In this Technology Reseller News podcast, recorded at Cisco Live, I spoke with Patrick Lin of Cisco Splunk about the changing role of observability in a hybrid, AI-driven IT environment — and why the conversation now also includes token economics. As AI becomes part of everyday IT operations, enterprises are beginning to ask a new economic question: how much does it cost to reason over all this data? In an AI-native environment, every log, metric, trace, network signal and security event may become part of a larger decision-making process. That creates value, but it also creates cost. Token economics becomes part of the observability discussion because customers need to know what data matters, when to use AI, and how to get better answers without flooding systems with unnecessary context. That is where Cisco and Splunk are focused: helping customers bring the right information together, with the right context, so AI can be useful rather than overwhelming. Lin described how Cisco and Splunk are connecting observability, networking intelligence and AI-native workflows to help teams see across complex environments. A key example is the integration between ThousandEyes and Splunk Observability Cloud, giving teams the ability to understand whether a problem is happening in the application or in the network — and, if it is in the network, whether the issue is in the part of the network they own or the part they do not. That distinction matters. In hybrid environments, responsibility is often shared across enterprise infrastructure, cloud platforms, service providers, SaaS applications and third-party systems. Knowing where the problem lives can dramatically reduce the time teams spend in war rooms trying to determine what went wrong. Lin also pointed to Cisco Cloud Control and AI Canvas as part of a broader AI-native approach. Rather than forcing users to jump across separate tools and interfaces, Cisco is working toward a model where information from Splunk, Cisco platforms and the wider ecosystem can be brought into a collaborative environment. That includes human teammates as well as agentic assistants that can help teams reason across data, identify patterns and accelerate troubleshooting. For channel partners, Lin said the opportunity is significant. Customers need help bringing together data sources, building the right observability foundation and applying AI in practical ways. Partners can play a key role in making agentic observability real for customers by helping them move from disconnected monitoring tools to a more unified, intelligent operating model. The goal, Lin said, is not just more data. It is a “much, much smaller war room” when incidents happen. For Cisco Partners, that message is timely. As customers modernize applications, adopt AI, expand hybrid environments and depend on increasingly distributed infrastructure, observability becomes more than an IT operations tool. It becomes a business resilience capability. Learn more about Cisco Splunk at: https://www.splunk.com/ Learn more about Cisco at: https://www.cisco.com/
Sherwood Callaway is the founder of Sazabi (YC P26), the AI-native observability platform built for engineering teams who ship fast. He previously founded and exited a YC company — now he's back, betting that logs are all you need to replace Datadog.Logs Are All You Need: Rethinking Observability with AI Agents // MLOps Podcast #381 with Sherwood Callaway, the Founder of Sazabi
BONUS: How AI Is Reshaping Software Teams From the Inside — Lessons From Google, Meta, and Snowflake In this episode, Dwarak Rajagopal — VP of AI Engineering and Research at Snowflake — shares what he's seeing firsthand as AI agents become part of the software development process. From compressed sprint cycles to automated standups across time zones, Dwarak draws on two decades of building AI infrastructure at Google, Meta, Uber, and Apple to show what's actually changing inside engineering organizations today. From Compiler Engineer to AI Leader — The Thread That Connects Two Decades "In AI, the hardest part isn't just the models itself, it's making them work in real environments where data is messy, fragmented, and governed." Dwarak started his career as an open-source GCC compiler engineer over two decades ago, optimizing hardware performance. He moved into graphics at Apple, then pivoted to AI when AlexNet started running on GPUs around 2011-2012. From there, he built autonomous driving software at Uber, led Meta's PyTorch core framework team bridging research and production, and at Google led AI Frameworks including getting Gemini training on TPUs. The common thread: always working at the intersection of research and production, making powerful technology work in the real world. That focus on real-world application is what drew him to Snowflake — where enterprise data meets AI at scale. AI Is Changing What Engineers Actually Do All Day "Engineers are spending more time on system design, validation, production reliability — and less time doing the implementation itself, because AI is helping that." The shift Dwarak sees is concrete: AI is accelerating development, but the real value comes when it's grounded in enterprise data and context. At Snowflake, teams use tools like Cortex Code, Snowflake Intelligence, and other LLMs to generate code and tests faster — because the friction cost of development has dropped dramatically. Customer example: Whoop, the fitness band company, used Cortex Code with conversational data assistance and agents to reduce development cycles from weeks to hours, freeing teams to focus on high-value work. The End of "This or That" — Try Both, Kill Fast "There's a lot more choices now. You don't have to think about this versus that. Do both and then figure out what is the best." One of the most practical shifts Dwarak describes: teams no longer need to commit to one architectural approach upfront. Because AI reduces the cost of building, teams can pursue two designs in parallel and evaluate both. A concrete example: instead of choosing a cross-platform framework like Flutter or React Native for a mobile app, Snowflake's teams now build native iOS and Android apps simultaneously — one human-led, the other agent-built — at roughly the same speed. But this creates a new challenge: teams have to learn to kill projects faster. When you can build more, you also discard more — and engineers need to detach from "their baby." Smaller Teams, Bigger Output — The Cross-Functional Shift "You could build multiple products now faster with different smaller teams. One back-end person, one front-end person — build vertically end-to-end." Dwarak's teams moved from functional structures (separate backend, frontend, and feature teams) to project-based teams that own the full vertical stack. This isn't theoretical — Snowflake Intelligence was built this way. The result: fewer dependencies, faster delivery, more products in parallel. The tradeoff is coordination cost — more things running in parallel means more decisions to synchronize. Recruiting Has Fundamentally Changed — Systems Thinking Over Syntax "We used to ask an engineer to code a specific search algorithm. Now we ask them to build a whole search system within an hour." Dwarak is clear: fundamentals matter more than ever. Systems thinking, judgment, the ability to work with complex data and production systems — these are what hiring evaluates now. AI handles execution; humans need to define problems clearly and ensure systems behave at scale. For junior engineers, the news is encouraging: onboarding is faster because team-specific skills are codified and shared, and the barrier to building end-to-end systems has dropped. "Learning by building is more true than ever now." Monday Planning, Friday Demos — The Compressed Sprint "You basically decide what to do on Monday, and you're testing together as a team on Friday and getting the feedback for the next week." Daily work has transformed at Snowflake. The traditional multi-week sprint has compressed to a single week: Monday planning, Friday team demos and testing. Standups still happen — but faster, sometimes multiple times per day. For distributed teams across Bay Area, Seattle, and Poland, an automated skill scans each day's code changes and posts a summary in a shared Slack channel — so the next timezone knows exactly what happened without waiting for a meeting. This solves one of the oldest problems in distributed development. The Road to Lights-Out Codebases — Governance, Observability, Reversibility "Can agents take actions? Which of these actions cannot be taken back? You need the concept of committing actions or rolling back." Building on the "lights-out codebases" concept from Philip Su's episode, Dwarak agrees the direction is clear — agents are already writing more code than humans in some contexts. But enterprise adoption requires governance, observability, traceability, and reversibility of agent actions. The shift from "AI as a tool" to "AI as part of the system" is happening now, with the focus moving from getting answers to enabling actions at scale. What Most People Get Wrong About AI in Software "It's very easy to build prototypes, even end-to-end systems. But it's very hard to get it working in enterprises where the data is so messy." The gap between demo and production is where most organizations hit the wall. Enterprise data is scattered across invoices, factory outputs, and dozens of systems — combining it meaningfully for AI to generate insights and actions is the real challenge. This is different from the "AI will replace developers" narrative. The bottleneck isn't code generation; it's data integration, governance, and controlled execution at scale. About Dwarak Rajagopal Dwarak Rajagopal is VP of AI Engineering at Snowflake, where he leads the Cortex AI and AI Research teams. Before Snowflake, he led Google's AI Frameworks and On-Device ML teams (including Gemini), ran Meta's PyTorch Core Frameworks team, and built autonomous driving software at Uber. Two decades of shipping AI at the companies that define the field. You can link with Dwarak Rajagopal on LinkedIn.
GoConf, Sept 11 & Moscow, RussiaCFPProposalsAccepted: Formal GODEBUG removal policyNew: Allow explicit conversion from function to 1-method interfaceBlog: The 10 Go Error Handling Commandments by Preslav RachevLearn Logging & Observability in Go @ boot.dev, use code CUPOGO to save 25%Video: Practical Go Development with AI Agents by Miki TebekaBlog series: Understanding the Go Runtime by Jesús Espino ★ Support this podcast on Patreon ★
This is the Everyday AI episode we probably shoulda done a while ago....
The new AIEWF website is live! CFPs close in 2 days and we will run our first New Engineer Orientation this weekend, get your tickets booked ASAP as they -will- sell out. Take the AI Engineering Survey and get >$2k in credits and free AIE WF tickets!One of the central tensions in the agents industry is that even while there are major decacorn agent labs like Sierra, Decagon, Notion and Cursor being built up, it is also true that it has never been easier to DIY agents, with a plethora of agent frameworks like LangGraph and Pydantic and Flue, and managed agents from Anthropic and Gemini and Amazon. There has been a wave of companies building their own background agents from Shopify to Stripe to Paradigm to Razorpay, and even Cognition's friends Ramp have built their own coding agent with other friend Modal.You'd think Cognition might feel a bit threatened, but they're not - even after all this, they were way oversubscribed for the $1B Series D they just announced:Walden Yan, coiner of context engineering and Chief Product Officer/Cofounder of Cognition, invited OpenInspect's Cole Murray to talk about why the Devin is in the Details.Full conversation live on the pod today: In retrospect, async agents were the most AGI pilled bet you could make in 2024 - the models weren't good enough yet to vibecode, and people didn't trust AI enough to let it rip, nobody (including early Cognition) was sure about the form factors. Now it is obvious:* The first wave of AI coding tools made the developer faster but remain heavily in the loop. Copilor and Cursor's tab autocomplete are prime examples However, the workflow was still heavily centered around and bottlenecked by the developer's local workflow: a developer in an IDE, watching the model, accepting or rejecting changes, and pushing code one interaction at a time.* The second wave was local agents: Claude Code, Windsurf, Cursor's agents pane: first one and increasingly many terminals all running concurrently.* The current Age of Async Agents points to a different future focused more on agent orchestration which drives end-to-end development.According to previous guest Steve Yegge, there are finer-grained 8 levels to agent adoption, but we have collapsed it into three.As Cursor's Michael Truell put it in The third era of AI software development:Cursor is no longer primarily about writing code. It is about helping developers build the factory that creates their software. This factory is made up of fleets of agents that they interact with as teammates: providing initial direction, equipping them with the tools to work independently, and reviewing their work.The agent should not sit solely inside the developer's flow. It should be setup to work in the background so that you can give it a task, a repo, a machine, a shell, a browser, tests, memory, and review loops to go do the work somewhere else.In less than a year, the sentiment has shifted from avoiding multi-agent systems:to suggesting approaches that actually work:From coining “context engineering” to building the infrastructure behind Devin's 7x PR growth and jump from 16% to 80% of commits across Cognition repos, Walden Yan has had a front-row seat to the background-agent shift. In this episode, Cognition co-founder and CPO Walden Yan joins swyx alongside Cole Murray, creator of OpenInspect, to unpack why everyone is building their own Devin, what changed after the December 2025 model inflection, and why “spec to pull request” is now becoming a real production workflow.We go deep on the architecture of background agents: harness-in-the-box vs out-of-the-box, why Devin separates the “brain” from the machine, why repo setup is still one of the hardest problems, why Docker is not always enough, and how full VMs, snapshots, scoped secrets, GitHub bots, Slack integrations, and video-based testing all fit together. Walden and Cole also dig into memory, MCP limitations, multi-agent orchestration, AI code review, SRE auto-triage, PMs shipping code from Slack, Windsurf 2.0, hybrid frontier/sub-frontier systems, and the real failure mode of uncontrolled vibe coding: your codebase regressing to your worst engineer.And as agents eat software… and software eats the world… you can draw the conclusion on what is next:We discuss:* Why the engineering world is waking up to background agents and cloud agents* The December 2025 model inflection that made spec-to-PR workflows practical* Devin's 7x merged PR growth and rise from 16% to 80% of commits* Why Cole built OpenInspect as an open-source background-agent system* The economics of $20/seat agent products and why monetization is tricky* What Cognition actually sells beyond Devin: infra, onboarding, integrations, and adoption* Harness in the box vs out of the box, and why architecture matters* Why Devin separates the brain from the machine for security and permissions* Repo setup, scoped secrets, Docker Compose, and agent-ready dev environments* Why full VMs matter when agents need to run real applications and test them* Android, macOS, Windows, nested virtualization, and machine-specific agent work* Why testing is much harder than “computer use”* Screenshots, video verification, and the “I know it works” merge moment* GitHub UX, Devin Review, AI reviewers, and agents responding to PR comments* Why MCP alone is not enough for first-class Slack and enterprise integrations* Memory, Knowledge, skills, Claude.md, and why retrieval is still unsolved* Devin's auto-generated memories and the challenge of memory pruning* Always-on agents as permanent PMs for issues, tickets, and product areas* Sub-agents, meta-Devin management, and what multi-agent systems actually add* Why pure auto-merge vibe coding breaks down after about two weeks* AI code smells, lint rules, reward hacking, and Semgrep for agent-written code* GitAI, inline context, and preserving the “why” behind code changes* Local testing, mock servers, older codebases, and preparing companies for agents* Windsurf 2.0 and the handoff between local foreground agents and cloud background agents* SRE auto-triage, support workflows, and agents as first responders* PMs, marketing, and non-engineers creating pull requests from Slack* AI agent budgets, $1k-$5k per engineer spend, and hybrid frontier/sub-frontier systems* The rise of autonomous coding factories and who Cognition is hiringWalden Yan* X: https://x.com/walden_yan* LinkedIn: https://www.linkedin.com/in/waldenyan/Cole Murray* X: https://x.com/_colemurray* LinkedIn: https://www.linkedin.com/in/colemurray/* OpenInspect / Background Agents: https://github.com/ColeMurray/background-agentsTimestamps00:00:00 Introduction00:00:43 Why Everyone Is Building Their Own Devin00:01:57 Devin's 2025 Ramp: 7x PR Growth and 80% of Commits00:03:49 OpenInspect and the Rise of Open-Source Background Agents00:07:59 What Cognition Actually Sells Beyond Devin00:09:56 Background Agent Architecture: Harness In vs Out of the Box00:12:08 Separating the Brain from the Machine00:14:07 Repo Setup, Secrets, Docker, and Full VMs00:19:13 Why Testing Is Harder Than Computer Use00:22:40 Video Verification and the “I Know It Works” Merge Moment00:23:19 GitHub UX, Devin Review, and AI Code Review00:25:42 MCP, Slack, and Enterprise Agent Integrations00:28:59 Memory, Knowledge, and Always-On Agents00:36:16 Sub-Agents, Multi-Agent Orchestration, and Meta-Devin00:43:55 Vibe Coding, Auto-Merge, and Codebase Decay00:48:38 Agent Infra, VPCs, Cloud Providers, and Fast VM Restore00:52:25 AI Code Smells, Reward Hacking, and Code Review Systems00:56:10 Making Codebases Agent-Ready00:58:30 Windsurf 2.0 and the Local-to-Cloud Agent Handoff01:01:15 SRE Auto-Triage, PMs Shipping Code, and Agent Use Cases01:04:32 Agent Budgets, Hybrid Models, and Autonomous Coding Factories01:06:51 Hiring at Cognition and OpenInspect Consulting01:07:45 OutroTranscriptIntroduction: Walden Yan, Cole Murray, and Context EngineeringSwyx [00:00:00]: All right, we're in the studio with Walden Yan, co-founder of Cognition, CPO.Walden [00:00:08]: Happy to be here.Swyx [00:00:09]: Which is a cool title. And coiner of context engineering.Walden [00:00:15]: Although I think there are many people who'd used the terms in various ways beforehand, but I did find that people, both internally and externally, enjoyed the upgrade from prompt engineering or model wrapping into maybe a more thoughtful way to build agents.Swyx [00:00:33]: For those who haven't caught up on that, I have on screen the Don't Build Multi-Agents post, which you should go read on and we might refer to, and Cole Murray, who created OpenInspect.Cole [00:00:43]: Great to be here.Swyx [00:00:43]: So let's talk about it. Everyone is building their own Devins. What's going on?The December Shift: From Handholding Models to Autonomous PRsCole [00:00:51]: So I think the engineering world is waking up to this idea of background agents, cloud agents, whatever you'd like to call it. And I think we saw a shift around the December timeframe of 2025, where the models Opus 4.5 and GPT 5.2, they reached a capability where we moved away from handholding the model and being able to actually more or less autonomously drive the model. And what I mean by that is that we could pretty much go from a specification to a completed pull request, assuming the spec was good enough, with very little friction. And that paradigm alone, I think, changed a lot of how we interact with agents, and opened this world where background agents became more practical.Swyx [00:01:41]: I think for Cole, everyone experienced this in December, but I feel like there was just this increasing ramp, right? There was this moment which was, I think, Sonnet 3.7, where, You guys rewrote Devin in one night or something. So describe 2025 or how it felt from your side.Walden [00:02:01]: In retrospect, we always thought it was ramping up, but then even now, over the last three, four months from today, it's been ramping up even faster. So it's almost funny to be talking about how, big of a leap Sonnet 3.7 was, and honestly, a lot of it was stripping out parts of Devin that were no longer needed with that jump in of intelligence. But I also just think that a lot of the recent leaps, especially, you look at, models like Opus and the latest GPT models, they are reaching levels of autonomy where people are actually finding that they actually can just be hands-off. And people who were once debating, “Oh, do I need to be in the weeds with my model in the IDE? Can I just completely move it off into the cloud?” That's a more serious conversation, and we've seen that in all of our growth charts. Internally there's this funny graph where our usage has, of PRs, our merged PRs, has grown 7X since I forget what it was called.Swyx [00:02:57]: I think Dev, maybe tweeted that. Yes.Walden [00:03:01]: it grew like 7X over, the last, I think it was, two months, three months, something like that. And then you see our engineering headcount growth. It's, gone up by, 10% or something.Swyx [00:03:11]: We were, we were afraid To release this. So this is Devin commit percentages on all Devin repos, was 16% in January and now 80% in March.Walden [00:03:25]: It's a big shift right now. And so it makes sense that a lot of people are now thinking about, buying Devin, but also maybe, trying to build their own and there's Lots of I have a lot of fun building Devin, so I can see why other people would want to build their own cloud agents as well. Matt, well, maybe it's good to hear, what initially inspired you to try to build OpenInspect?OpenInspect: Ramp, Cloud Agents, and Open SourceCole [00:03:49]: OpenInspect came about, through primarily my clients observing how they were using tools like Claude, OpenAI's Codex at the time, and seeing some of the friction that they were having with it. Primarily the Claude was being used through Slack, and a big issue they ran into was that the sessions that were launched were specific to whoever called it via Slack. And so if a PM was the one who invoked the session and they would then go to pass context to engineering can't see the session. And that in itself was a deal breaker because the PM, “Hey, engineering, can you jump in?” But there's nothing to jump in on unless they're copy-pasting out or the single response that came back. And so seeing some of these problems, I had built a similar architecture internally, just to experiment with, test out different ideas as this trend of moving off of localhost was starting to become, And as Ramp released their blog post, I had a lot of the pieces for this already in place, and just thought it would be funny to, see what Claude could do just purely from the blog post. And on my X account, there's actually a thread of where I live tweeted, going through thisCole [00:05:14]: comparing GPT and Claude as both of them are going through it.Swyx [00:05:17]: On the announcement thing or something else?Cole [00:05:19]: right after it got released. We can put it in the show notes. Yeah, it was helpful that I had already knew how to verify the system. I knew what I was looking for. I think Ramp did a great job of really illustrating, the technical aspects of how to build something. It was much more than just like, “Hey, we built a great system.” It was, “And here's how you can build it too.” And so, I resonated a lot with that, just with the problems that I was already seeing, and I thought that, looking around, I didn't really see anything in the open source community that, met this type of system. I think there's a lot that run, in localhost like Superset, Conductor, and many others.But nothing that was actually running in the cloud. And so, I built it, and I thought it was interesting to just open source it and allow anyone to then have a foundation that they can mix and match on top of.The Business of Background Agents: Open Source vs. DevinSwyx [00:06:16]: So literally after Devin was launched was, there was OpenDevin Which became All Hands. I don't know if you tried that orWalden [00:06:22]: I was going to say, one of the things that interested me a lot with OpenInspect was, you didn't try to go make it then something you monetize. There are a lot of, I think, these open source projects would then go and really try to, raise VSwyx [00:06:36]: That's why no OpenDevin. Yeah.Walden [00:06:38]: yeah, and how did you think about that? I thought that was very interesting.Cole [00:06:44]: I thought, and just what I had seen across my clients, was that having a background agent system is going to become a critical infrastructure within their company. And so because of that, I think that I wanted to open source it so that they could fork it and put in whatever customization they wanted. To that question though, I get asked all, “Oh, are you going to raise? Are you going to turn this into a service?”Walden [00:07:08]: I'm sure you've gotten offers.Cole [00:07:09]: but primarily I don't want to do that for a few reasons. One, I think that I don't want to compete for, $20 a seat. I think that is just a really difficult business. I think it's very easy to copy the main pieces of it. Again, I built this fairly quickly. And I think because you are not owning, I guess, the entire stack, it's hard to monetize. You have money being made at the sandbox layer with Daytona, E2b, many other players. You have money being made at the model layer. And you sit in this weird in-between gray area where what are you actually selling? You're selling, I guess, the infrastructure. You're selling, the integrations maybe.Swyx [00:07:55]: let's ask the guy. What are you What are you selling?Walden [00:07:59]: Well, yeah, there's multiple layers to this in practice, and actually it's funny you mentioned the infrastructure, ‘cause when we got started building Devin as well, we had to go figure out how to make the infrastructure as well because,Swyx [00:08:10]: You had to build this two years before everyone else,?Swyx [00:08:15]: Including, the model sideWalden [00:08:17]: It was not, it was not very polished at the start, when we just built it off of raw VMs from cloud providers like EC2, the boot up time was so slow, I think, And especially then, turning off the machines, saving them, and then to be able to bring them back up again when the, when you want Devin to wake up again later. It would just be out cold for like 10 minutes because that's just how long these systems took. They were not built for this repeated down and up usage. And so we actually had to go do all of that. And as a result now, one thing we offer when we go and sell Devin to people is, you don't have to worry about all the compute side of things. We'll make it work. We'll make it work in your cloud if you want it to. But aside from the product, and I want to go into the agents and the tuning of the intelligence part later, but I think a big part of what we do at Cognition as well is to just make sure that your company learns and uses and adopts these coding agents. ‘Cause I think for especially the largest enterprises in the world, you find that there is a lot of people who want to move over to using AI for their day-to-day workloads. But because of the way projects are planned, because, not everyone is literate in using AI in these ways, having a team of engineers who can actually go in and onboard you, set up all the integrations you need, the automations you need to really get to that level of, leverage with AI, is super helpful. And so We do that. We show thought partners to the customers that we work with as well.Swyx [00:09:56]: So let's talk about, architectural stuff. I think that's always, that is something that was the topic of conversation between the two of you. Is this, the mental model that you want to start with or something else? I'll just leave the floor open to you guys.Agent Architecture: Harness in the Box vs. Out of the BoxCole [00:10:11]: I think, maybe we can start here as just a general what are the pieces of a background agent system. And then maybe we can go into some of the nuances of, Decisions that you can make.Swyx [00:10:22]: But I guess I also Like, what, maybe what Walden is saying is the agent is like in this open code box, I guess. Right? This is infra, and then there's, that's the agent. And you had this discussion about whether you put the agent in here or in Out externally. Can you tease that out?Cole [00:10:39]: In a background agent systems, you have a decision to make of where the agent is actually going to run. This is typically described as the harness in the box or out of the box. With running the agent in the box, you're making some trade-offs by doing that. The negative trade-off you're making is primarily security. Because the agent is running in that box, unless you otherwise design it, all of your secrets need to go into that box as well. And given the nature of AI, it can be unpredictable, and you could very easily end up accidentally exfilling your secrets, or other unintended behavior. Now, the out of the box is the idea that we are going to have the actual agent running not directly in the sandbox, and we will have, quote-unquote, the brain of the agent running in some type of worker, control plane. That sandbox then is going to serve as the hands where the brain is basically operating and making tool calls into that environment to manipulate it. I guess other trade-off that you're making between the two systems is that, in my opinion, running it out of the box is much more complex because, you have state that has to be managed, whereas if you're running it in the box, all of the state of that agent is actually in the box, and yes, it's you could persist it elsewhere, but it's all localized and you have less concerns to worry about.Walden [00:12:08]: I think a lot of that, what you mentioned, is why we actually from the start built Devin to what we called separate the brain from the machine. The other thing that this allows you to do is reuse any existing infrastructure you have for dev boxes Perhaps. And so you don't have to worry as much about making a new type of dev box that has all the dependencies the brain needs, as you mentioned, the secrets the brain needs as well. One thing that we've seen some customers run into is, you have a GitHub app and you want Devin, your agent, whatever, be able to interact with GitHub through this application, but then you have different users with different actual permissions. If they are all interacting through the same GitHub app and there's no actual, separation between the system that decides, what it does and the actual secrets on the machine, then you run into an issue where, okay, it's hard to do the separation. But in practice, with Devin, it's much easier because we just say whatever you put on the machine, that is, the scope of basically what the user is free to do, what the agent is free to do. So only put the most scoped secrets on that machine, and then the brain is fully not accessible from the machine. So you don't have to worry about messing with the, any of the most secure parts of the brain if the user is free to do whatever they want with the machine.Swyx [00:13:31]: I was going to just bring, I have this, chart from OpenAI, where I don't know if this is, in the box, out of the box. That is something that they do use to describe it. And then also recently Anthropic did, managed agentsSwyx [00:13:44]: Which is, this is their thing. I don't know. It's all, it's all variations of the same pattern, right?Cole [00:13:49]: So this would be out of the box.Swyx [00:13:51]: Which, is preferable for them because it's less work?Cole [00:13:56]: I would say it's more work.Swyx [00:13:58]: It's more work?Cole [00:13:58]: But it, in my opinion, it is the better architecture of the two. It's just, you're taking on a bit of complexity by doing that.Repo Setup, Docker, and VM-Based Development EnvironmentsWalden [00:14:07]: One thing I've not seen a lot of other players do well is how do you manage what's actually on the box? And this can be complex for many reasons. Let's say you have a big repository that's changing and updating a lot with changing dependencies. How do you make sure that the working environment of the agent actually stays up to date, has all the credentials it needs to, let's say, run the app and test it, and all the things you want your autonomousSwyx [00:14:34]: So a repo setup.Walden [00:14:35]: Exactly. So in, internally At Cognition, we call this repo setup.Cole [00:14:39]: The hardest part ofWalden [00:14:40]: It's been a perennial problem since the start of the company, of how do we help people get this set up? Because not everyone just has, working cloud environments working out of the box. And do you find this to be a common problem withSwyx [00:14:53]: How do you solve it?Walden [00:14:53]: Your clients?Cole [00:14:54]: This is a very common problem, and through my consulting, this is a lot of what I help teams do. A lot of teams don't really have great developer environment setups, if any. A lot of the times it's, “Go talk to Bob and get the secrets,” and that obviously doesn't work when the agent needs to actually set this up. And so a lot of that, most teams are using Docker Compose or some type of microservices. And so for theSwyx [00:15:19]: Even in prod?Cole [00:15:20]: Not in prod. With the OpenInspect, you are using this primarily to interact, and make code changes. There is other use cases, but you can hook, whether through CLI, MCPs, other tools, you can then hook that into your production systems primarily for, SRE type use cases. But you are not, necessarily, trying to test your prod internal microservice through the system.Walden [00:15:48]: And you mentioned Docker Compose. I think one direction we saw some of our friends take early on was, using Docker containers as the level of abstraction for their models. There's lots of reasons, I think, why Docker containers are not great. One thing is, Docker container's not really a true security boundary, for one. But the other is, if you are running real applications, a lot of times those applications use Docker, and then you have to think about Docker in Docker, which is, really weird. And so I think part of, the really hard challenge of getting VMs to work, why did we do that? Well, it was because we realized that you actually needed, full VMs to be able to do these types of things. And especially nowadays where there's actually value in running the application and clicking around and sending you screen recordings of these things. The value just, keeps adding on top of that. But it is a decision I see people run into when they try to build their own systems, is, “Oh, do we, in addition to this, do we put the agent in the machine or out of the machine? Do we use Docker? Do we use something else?” What do you recommend people nowadays?Cole [00:16:57]: I think Docker is a good solution for maybe not running the agent, but running your infrastructure, because that is more or less the same setup your engineers are probably already using. If they're not, then I don't know what they're using. But they're probably already using Docker Compose.Swyx [00:17:14]: I've always had a small candle for web containers. I don't know if you guys have tried them before.Swyx [00:17:19]: To me, they were, supposed to be like Docker Light.Cole [00:17:22]: Is it?Swyx [00:17:22]: I don't know.Cole [00:17:22]: No, I haven't tried it. But yeah, I think any environment that you've set up that is a good experience for your developer naturally lends itself to being easy to set up for the agent. And once you figure out that local developer story, you've more or less solved the agent in a sandbox, environment setup. OpenInspect does have hooks as well, where you can, run a setup SH script that will pre-install everything. You can then pre-snapshot that build so it starts instantly, and then there is a second hook to actually then, restore the state of the sandbox when it comes back. And so you can already have all of those microservices running and basically get the same experience that you would on your machine within the sandbox.Testing Agents: Computer Use, Screenshots, and Real App WorkflowsWalden [00:18:08]: Another thing that we've been thinking a lot about is like Different VM service offerings. Have you had customers where they needed like macOS specific VMs or like Windows specificWalden [00:18:20]: VMs?Walden [00:18:22]: There are like many technologies in the world that only work on specific types of machines, right? If you're building a.NET application that has to run on Windows or like, maybe more commonly if you want to build iOS or macOS Does that workSwyx [00:18:32]: Does Commission supportSwyx [00:18:33]: Choices like that?Walden [00:18:35]: The fundamental architecture we do, because we do the separation, it does support, but the actual work in progress is happening right now on these. Another thing that we've actually recently added support now for, it's in beta, is doing Android development. To do that, we needed to support, I think, nested virtualization within our machines because the VM itself is like a, is a virtualized Firecracker instance, and then you had to then run another Android emulator inside. And there's like weird performance issues that like, it, which is why it's like still in beta. We have to think through these problems, but it unlocks a lot for anyone who wants to do Android development.Swyx [00:19:13]: I was trying to find like a reference video for the testing thing. I couldn't find it, but I think you worked on the testing, capability. Why call it testing and not like computer use or I don't know, it's, what's the general Category of problem?Walden [00:19:26]: I think that when people think about the ability of an AI to run your app and test it, I think they actually over-index on the computer use part of it because computer use in my mind is the literal, okay, you want what button you want to click. Can you emit the right coordinates to go click that button? I think testing is actually a really interesting likeWalden [00:19:48]: Problem-solving, challenge for these AIs because if you wanted to do arbitrary testing, imagine you make a change that spans the frontend and the backend, maybe, even some other like even more deeply nested service. To actually test that change, we have to reason through what-- how do you first run these applications to orchestrate with each other with the right version of the code? Then, okay, how do I trigger the feature or how do I make the thing actually happen? And this can get arbitrarily hard, maybe you have to be an admin. Maybe a certain thing has to be feature flagged on. Maybe, you have to like run two sessions and then send us a very specific word into one of them to trigger a specific behavior. And figuring out how do you do that requires a lot of code base context, requires, a lot of orchestration that we've specifically done. And in some cases, we found that you actually, no one frontier model can actually do this full end-to-end task itself.Walden [00:20:42]: We've seen cases where we actually had to orchestrate different frontier models together to solve this problem together. That is where we spend most of our time when we think about this testing problem, not so much the computer use part. Computer use for what it's worth has gotten a lot better with recent models and it's made that part of the job certainly easier.Swyx [00:20:58]: Especially with like even 4.7, that they released yesterday, apparently like way better in terms of the vision stuff, which is going to be encompassing computer use.Walden [00:21:08]: Having evals for all these as well is something that like takes a while to build up. And having the evals be right is tricky as well. Do you ever see like, clients who are building their own agents have to start standing up evals to make sure things don't regress?Swyx [00:21:25]: Not so much evals in the traditional sense, but specific to the testing part that has just gone in. I just added support for screenshots And in theory you can also do video. I need to put in a plugin to do that. But they do show up natively, and it was a very heavily requested feature, especially after Cursor's recording came out. I think that was very enlightening for everyone of like, “Oh, this is a very good feature to actually have.”, I think with Devin you guys have had this for a while.Swyx [00:21:57]: Oh, yeah. See how screenshots work. Yeah, I don't know if there's anything, super and not obvious. It's like once what feature to build, you can just prompt it and it Will mostly work.Walden [00:22:09]: I think to Walden's point, though, the computer use is a subset of the larger testing problem, and I think that's very specific to the code base that you're working and it's not something that, out of the box that you could just solve it. The-- you do need the code base context to actually know how to test it. And I think in the case of a background agent system, you fortunately do have that code base locally that what is changing and could then inspect it and use that to drive the model.Swyx [00:22:40]: For those who haven't seen it before, this is an example of how it works. You, after the PR is done, you click testing approved, and then it sends you back a video. What I really like is that it labels, It's very small here, but it actually labels what it's testing. And then it-- and then you actually see the cursor and everything. So I don't know, yeah, the engineering in this, just Whatever you want to show. ‘cause this is like, this is one of those like, oh, few of the AGI moments, right? ‘cause Once I look at this, I actually don't I wish I can just merge inside Of Slack instead of going to GitHub ‘cause I don't need to see the code. I know it works.Walden [00:23:19]: Maybe a new feature in Cursor. Yeah, the annotations at the bottom was also a big difference for me when I, when I added those.Swyx [00:23:27]: It's just like, what am I looking at? What are you trying to demonstrate?Walden [00:23:30]: Exactly. There's a surprisingly long tail of small details that ends up making a big difference for this end metric of like how fast do you actually merge the code in. One experience that we spent a lot of time tuning early on was what is the right experience on GitHub for these tools. Because I think, most tools out there when you build the agent, you'll think about, oh, it'll create the PR for you. We try to take that a step further and say, “Oh, what if we actually made sure you could interact Devin, with direct Devin directly on GitHub?” And so we made sure that you can comment on GitHub, and Devin would actually receive those comments and address them back. But there's actually quite a bit of tuning you have to do here because you can imagine that actually like-We recently have Devin Review, for example. Devin Review will post comments on his own PR And then Devin has to then goGitHub Workflows: Devin Review, Comments, and PR AutomationSwyx [00:24:23]: He answers his own comments, which is Really loopy. So like, yeah, I like that it just updates here that it's, that I have commented But usually it's just me saying like, “Hey, merged, fix any merge conflicts.”Walden [00:24:37]: The, so when Devin fixes his own comments, you might be scared that, oh, maybe I'll infinite loop. But we've put a lot of work into making sure it doesn't, both by making sure that the comments are high signal, but also that the agent is thoughtful about what comments it immediately goes and tries to fix, and what comments it's like, “Wait a second, I think you're wrong.” Actually, that's one of my favorite moments is when Devin tells me that I'm wrong, when I try to get it to do something different. But tuning that behavior, actually makes a big difference in terms of how useful the actual GitHub experience is.Cole [00:25:06]: I think to touch on that as well, I think having the AI reviewer integrated into the system is a critical part of this background system. OpenInspect does have that. It has a GitHub code reviewer that you can control the prompt. It does do comments as well. It doesn't do them automatically yet. The capability is there, but it's not fully used.Swyx [00:25:27]: So you have to ask for it?Cole [00:25:28]: you do, yeah. You can tag it on GitHub, and then whatever you named your, GitHub bot, it will then follow up on it. It will then, if you have merge conflicts or whatever you have asked it to resolve, it will then resolve it, but it doesn't do it automatically yet.Integrations: Slack, MCP, and First-Party Agent InterfacesWalden [00:25:42]: Well, I'm curious, what is, the most common thing that people end up requesting, that they still need on top of OpenInspect when you help them go implement it?Cole [00:25:52]: I think a lot of it comes down to actually integrating it into the company. It's one thing to have the background agent system set up, but if it isn't actually integrated into your larger ecosystem, it isn't that useful. It is useful to be able to kick off sessions, but what we really want to be able to do is hook it into all of our other systems, whether that is the production database with read-only credentials, the logs, a Confluence or internal knowledge-based system. I think that is where I see the huge leap for companies, and that can be a challenge for companies as well who are maybe not familiar with exactly how to approach it, especially if they're in environments that have more compliance type things where, access control can be pretty big and how do you deliberately think about these problems, I find to be, one of the problems that comes with a system like this.Walden [00:26:46]: The thing we found is So, MCPs, obviously it has been like this, really big explosion of, oh, you can go, integrate it with all these different things. But to actually get the integration right and the and get the right experience, oftentimes we found that we had to go build our own ad hoc things. I think Slack is a great example of this. You could give your agent a Slack MCP and okay, it can post messages back to you on Slack. But we actually use Devin like a coworker in Slack, and that's how it's been built from the ground up. But to do that, you actually need to, support webhooks that come back, right? And then Devin has to respond in a natural way and then hopefully don't spam your threads too much and annoy the people in your company. So you got to tune that experience just right. Especially when there's a lot of back and forths, we find that we actually have to go beyond the simple MCP integrations in these places.Swyx [00:27:39]: I just pulled up the MCP marketplace. I know this is a Fair amount of work. Is the answer to eventually take first party control of all the top MCPs? Is that theWalden [00:27:48]: I would love a world where you could have something that's more expressive than MCP. That, goes both ways, not just a set of tools, but a proper system that interacts back and lets it Have the right experience with all these interfaces.Swyx [00:28:03]: So there actually is sampling in the MCP spec, but nobody Uses it, right?Walden [00:28:07]: And so I think that's the other part is, actually we found that when the MCP spec starts to get too complicated, it starts to lose its original promise of Being like a simple one-step connect. Now then we have to go figure out how to support all these different variations of things and It starts to look a lot like just building the first party integrations in a lot of these cases now.Cole [00:28:29]: I think it matters, too, how critical it is to your company, right? If this is something that nearly every session is going through, it probably makes sense to own it so that you can make optimizations on top of it Versus just whatever is off the shelf.Swyx [00:28:43]: Awesome. Other than MCPs, what else, sorry, well, I don't know if that's Narrowing in too much on, integrations. But what else? What other elements of building OpenInspect or Devin that you guys really sink on?Memory and Knowledge: What Agents Should RememberCole [00:28:59]: I think, a problem that comes up very frequently is this idea of memories or knowledge base.Swyx [00:29:05]: Oh, boy. How do you solve it?Cole [00:29:08]: so not solved yet, is the short answer.Cole [00:29:11]: it's something, there's a open issue for it, someone asking about it.Swyx [00:29:16]: There's, I, D Wiki hasn't indexed anything about memory yet.Cole [00:29:20]: how I'm seeing it solved across my clients is primarily through skills. I find that skills can be a good gap within that or updating Claude MD, but I think memory as a whole is a pretty unsolved problem, and it is why I've been hesitant to add it. I think there is parts of memory and that can be addressed, but I think as a whole it's a very difficult retrieval problem.Swyx [00:29:44]: Oh my God. RAMP didn't write anything about memory? I see zero search results.Walden [00:29:50]: No. Memory can be quite tricky to get right because it's the retrieval, but also the generation of the memories that can be really tricky. You don't want it to just like Remember very specific details.Swyx [00:29:59]: Walk us through the Devin memory journey because I know there's been a journey.Walden [00:30:03]: the first version of memory that like stuck around for a while was A system we have called Knowledge. And the idea was we wanted it to pick up things over time and not need the user to be proactive about teaching Devin things. So, okay, any time you remind Devin, “Wait, no, that's not quite the way you're supposed to use Git”Like, we actually want Devin to say, “Hey, do you want me to actually just remember this for the future?” And for you to just basically quickly approve or reject and for it to build up over time. ‘Cause I find that, 95%, I think, or some crazy stat like that of the memories that Devin has are all through these auto-generated things. Very few people actually just want to sit down and write big docs on Here's how you're supposed to work with the technology, et cetera. The generation and the retrieval has been something that we've been trying to tune a lot over the years. Generation, you don't want it to remember something like, if you asked one time to like, “Oh, please open as a draft PR,” you don't want to be like, “Oh, everyone forever now should get their PRs as draft PRs.” But you do want some, conveyor. Maybe you want to say like, “Oh, Cole generally likes, things to be created as draft PRs.” Same with retrieval, if you have thousands of these memories, how do you actually make sure they're retrieved at the right time? And that can be quite tricky to do right without exploding the context with a bunch of useful yeah, useless information. Surprising amount of just, eval work to just make sure that, memory is, remains a reliable system as new models come and go.Cole [00:31:31]: Do you have anything that you could share on, memory pruning? And like the temporal aspect of memory?Swyx [00:31:36]: Deleting and forgetting?Walden [00:31:39]: The, today, the, So the things they could do is it could edit memories. And so if your memory used to say like, “Oh, Cole likes to open everything as like a draft PR,” then you can imagine, “No, don't do that.” And then it'll say, “Oh, do you want me to update the memory to be Cole now want everything as, open PRs?” I think that at the same time we don't know if this is going to be the final version of the system. Whatever we have here will probably, translate into the new system that we'll be coming up with. But I think one big difference between two years ago and today is these agents are really good at using anything that resembles a file system natively. And so part of us are, is thinking, “Oh, should we rebuild memories to feel more like a file system that we let the agent navigate on its own?” That's been an interesting exploration. Also similar ideas in the scale space.Swyx [00:32:35]: I am pulling up OpenClaude's memory thing right now. So memory, OpenClaude has like this like daily memory journal thing, right? And you can I mean, that is a file system you can grep through and is a source of truth. I don't know if it's the best. It's probably super noisy, but at least, if you lose something you can discover it or you can apply some, forgetting algorithm to, more ancient memories that don't get recalled again or something. I don't know.Walden [00:33:01]: One thing we've been trying to do to push the boundaries of how you use agents at your company is letting an agent basically have a very similar file, a memory.md or something, and just like be your permanent PM for a specific set of issues maybe. So we have like some Slack channels internally, maybe a Slack channel dedicated to, a specific product like DeepWiki maybe. And you can imagine that, or you want a Devin that never stops, it's just always awake, but it has this like memory dock that it can just maintain for itself about, okay, what are like the number one priorities of what we have to fix and prioritize? Who is responsible for some upcoming work? Maybe they'll even Devin will even tag you on some recurring basis. And so it's been an interesting move to see, okay, how can we actually use Devin for more than just engineering? Can we actually upstream above the engineering process and maybe it's just Devin creating tickets, which then maybe some humans do, but then maybe other Devins do.Swyx [00:34:00]: One of my more fun automations is go research competitors and just suggest stuff to me on a weekly basis. That's the automation. I can't find it right now, but basically it just like, “Look at competitors and suggest things.” “And here are three things that you've suggested that I don't want any more of,” and you just stick that in the prompts. But like I wish actually So for like when I, for example, when I reject a PR, I wish that it updated memory so that I can then just not have to go up, go back and update the scheduled, sync, but anyway, feature request.Walden [00:34:31]: what? We might change it soon. I guess OpenInspect, in the time you've been around, has there been anything you tried to implement but then you had to like undo and like do a different way?OpenInspect Architecture: Webhooks, Control Planes, and Agent StateCole [00:34:41]: Nothing yet, but something that is on my mind. The initial way that I built it was that each of the integrations lives as its own package. And so you have The Slack bot, which is what's handling the webhooks, and then is basically interacting with the control plane. As I'm seeing the system starting to be more integrated, specifically with the GitHub bot integration, I'm considering bringing that all into the central control plane because especially now I want to start, And a request that I'm getting is the ability to monitor, the actual, pull requests being merged, as well as just tracking ofSwyx [00:35:19]: What do I have open?Cole [00:35:21]: What do I have open? How many of these are getting merged? How many comments are showing up? To just understand the health of the system. And so in the case of a GitHub app, you only have one webhook. And so then it's a question of do I put that webhook in that GitHub bot package? That's weird. It doesn't really make sense to live there because that package is more for like the code reviewer. Or do I like centralize it? So that's something that's on my mind of, making that decision. I think the other one we touched on earlier is the harness in the box versus out of the box. I think long term the architecture will eventually come back out of the box. Some of the newer tools that I've added are calling back into the control plane so that you don't have the secrets in the sandbox. And so I think long term I probably will pull the actual, agent out of the box, but I think for now it's fine.Subagents and Multi-Agent Systems: When Parallelism Helps or HurtsSwyx [00:36:16]: Just, a quick question on pulling the agent out of the box. I'm One thing I'm very bullish on this year is agents calling other agents or spawning sub-agents or Whatever you want to call it. Does that make it harder or easier? I can't tell. Because if the harness is in the box, you can just spin up more boxes. If the harness is outside the box, then you're, it's less easy because you are, you have a unicorn pet of a, of a harness that's, living outside the box.Cole [00:36:45]: In theory it would be the same way, right? Whether, one agent has launched many, sub-sessions within it, OpenInspect, for example, can launch sub-sessions and actually create other environments and then monitor them. In the case where it is out of the box, that would basically just be an additional session that's running. And so that session is also running outside of the box. It's running in your worker plane, wherever you're running this. And then you really just have to think about how does your top level agent then interact with it. I do think it can be more complex, just ‘cause again, you have now a more difficult architecture. But I think if you figured it out once, it's probably fine.Swyx [00:37:26]: Well, then I'm just, throwing it open to you in terms of, I call this like meta Devin management. Which is like the, Devin's calling Devins or Devin scheduling Devins or querying trajectories or anything like that. What have you built or unshipped, anything?Cole [00:37:46]: I think one of the surprising things we've seen is that a lot of the ways that, these, separate agents work with each other, and you want them to, parallelize their work, has still mostly followed the same manager sub-agents regime. And a lot of people I think are excited about this world where you have swarms of agents that, talk with each other all over the place. We've actually given Devin an MCP so they can just go arbitrarily message other Devins And create new Devins, et cetera. But I guess, it somehow creates, a really chaotic world in that sense. And so we've still found that most practical use on a day-to-day basis has been one single Devin.Cole [00:38:33]: Figuring out how to segregate the work and get, have other Devins work on it in, a relatively isolated sense, each with their own boxes Not sharing machines, so there's, a very little room for conflict is the regime that you have to create today.Swyx [00:38:50]: I'll call out, the experiments from Cursor, right? This is Wilson Lin's work on Single agent to multi-agent, and you're obviously famously on the side of don't build multi-agent. But they went through the whole thing, only to arrive at, this Which is exactly what Devin has, I think.Cole [00:39:08]: I think there will be a revision to that post at some point AboutSwyx [00:39:12]: Tell us about itCole [00:39:12]: I think multi-agents were very much not at all possible a year ago. You do see more multi-agent experiments today, but you can argue, are they really multi-agents, or are they just just, tool calls,? There are people who, will create sub-agents to go look for XYZ file, XYZ implementation. Has really nice context management benefits because all of the tool calls and tokens that it spends then get collapsed back to just the answer for the main agent. There's a lot of benefits to doing this. We basically have Devin do this with Deep Bookie, make a call out to Deep Bookie, give you back the results, but that feels like a tool call,? It's not like these, two collaborators actually talking back with each, back and forth with each other. But I think the thing that gives me the most bullishness that multi-agents might actually be possible is actually what I said earlier about Devin will actually sometimes tell me I'm wrong and push back, and I think that demonstrates a level of maturity and communication today that makes a multi-agent world possible. One, can two agents who have seen different information come back to each other and actually figure out who is right, what is the correct implementation? They're not just, yes men. Claude, I guess is like, used to just say, what is it? “You're right,” or,Swyx [00:40:25]: “You're absolutely right.”Cole [00:40:26]: “You're absolutely right.” Yeah.Swyx [00:40:28]: The Have you seen, did you seeCole [00:40:29]: The age is overSwyx [00:40:30]: The Codex app troll in Topic? This is the Codex app. Inside of Settings, there's a little, there's a little Easter egg, right? So if you go to, the Themes or Appearance, right? There's all these, color codes, and the top is absolutely, and it's the Topic's colors. Which is such a troll. Anyway.Model Behavior: Pushback, Adversarial Prompts, and Agent SkepticismCole [00:40:53]: I love that Easter egg. Did you discover that yourself?Swyx [00:40:54]: No, it was, someone was, tweeting about it And I was like, I was like, “Is this true?” Because, sometimes people just tweet stuff to, get a rise out of you. But yeah, there you go, in Topic colors.Cole [00:41:06]: Yeah. So yeah, we're out of this regime where, it just says you're absolutely right, and they can have real conversations and real back and forths.Swyx [00:41:13]: You can prompt it as well to be more adversarial or whatever. Yeah. Okay. Yeah, that, I mean, to me, that is more intelligence, right? That is not just something that's, a dumb tool, it's actually pushing back on you I think. Yeah.Cole [00:41:24]: when you mentioned, of course, the blog posts. There was one blog they had where they fed a swarm of agents together and built a browser.Swyx [00:41:34]: That was I think that was the one.Cole [00:41:36]: You can have, likeSwyx [00:41:37]: I think it's the same oneCole [00:41:37]: Creation of it. We found a surprising success of, don't do a swarm or anything, just have one Devin, it does its own context management. Just let it keep running for a while and give it some crazy tasks. I think we asked it to, rebuild, a Windows OS system. And it managed to do it just like, going on for long enough. It'sSwyx [00:41:55]: Was this Andrew's thing?Cole [00:41:58]: there were lots of demos that we ended up not posting, ‘cause at some point we'd just be posting way too much a bunch of, Demos. But I love that because it shows that I think the multi-agent thing still has, a bit of exciting sexiness to it, which is maybe still beyond still, the actual delta it adds to the capabilities of these systems. But it's absolutely the future. I think we're heading in that direction and we can see the progress being made there already.Swyx [00:42:25]: If I were to, make one super minor pushback because I don't feel that confident about it yetCole [00:42:33]: Go for itSwyx [00:42:33]: But I've had Ryan Lopopolo from OpenAI on the pod And he's a super slop cannon, right? Oh my God, that's my coding agent being done. I downloaded this, Peon Ping. I don't know if you guys have heard this. It takes like-, sound packs from popular games like, Command and Conquer and Warcraft, and then it plays it whenever it's done. And so it's like, “Work,” or whatever, “At your command,” or something. Anyway, what I got from the Cursor code base and from Ryan's thing was that there's a slop cannon approach where you try to loosen the single agent's, bottleneck, and I feel like that is, probably an, a very important thing to try to figure out. I don't think anyone's, really solved it. Because then you just have more reviewer slop on top of the agent slop To try to wrangle it all. Ryan will probably very strongly object that I say that he hasn't solved it, but he thinks he's He thinks he's completely solved it. But I think it's still I think it's, very important, ‘cause, that is a bottleneck, right? I feel Devin is slow sometimes Because I'm like, well, yeah, this is very readable and very sensible, but also it is slower than it could be if I just, I want a button to just say, “Just ramp this up 1,000 next parallel, in parallel and just, see what happens,”? And I don't know if that's, feasible at some point in the future.Code Review, Entropy, and AI SlopWalden [00:43:55]: I And we've also run experiments internally where we've basically tried to build entire products, true products that we knew we would eventually ship, but for now, let's try to see if we can do it just by purely, vibe coding on top of each other, auto merge, no code review at all. And then there's this benchmark of how many weeks can you go onto this for Before you say, “We have the trashiest code base.”Walden [00:44:18]: “Let's actually rewrite it from scratch.”Swyx [00:44:19]: Start a new factory, yeah. What'd you find?Walden [00:44:21]: I think we found that the state-of-the-art in December was you can probably, run this for about two weeks. By the end of those two weeks, you'd find that, hey, you want to, change the color of a button. Well, it turns out this button is implemented in, 10 different places, and they, have All these different variations, and oh, you forgot one of them, and actually it's a slightly different color in one spot. And you're like, “Okay, this is too much to work with. Let's actually try to do code review at the same time.” And make sure that we're on top of our software, actually cleaning it up a bit And making sure it's done in a scalable way.Cole [00:44:54]: I think building on that, the idea of, you don't have to look at code, I think is generally a bad idea. And the meme that I have for thatWalden [00:45:03]: What timeline, all right, is Do you think that statement will be true on?Cole [00:45:06]: I think probably for a while it'll be true that you should continue to look at your code. A problem that I see a lot of teams run into that I work with who are embracing AI native, AI first coding, is The meme that I have is that your code base regresses to your worst engineer, because that engineer who is, very gung-ho about AI and is not auditing their code, their pattern starts cementing into the code, and now the AI is referencing their patterns. And so now their if/else block that, is 20 if/elses back and forth, the AI is seeing that as the pattern of how things are done and starts to then exponentially grow this slop. And I find to your point, a pretty good approach to that is having scheduled cleanup, whether by humans or through systems, that are looking for duplication. They then address that. You'll end up with like 12 helpers for how to format a date. And you need to address that, because otherwise it will continue to sprawl.Swyx [00:46:09]: Within balance, I think it's fine to have some duplication, and then sometimes To have garbage collection, right? Yeah. The What I've been, talking about with a lot of engineering leaders is that you want to be very strict about the boundaries between modules, and it's your job as an architect, as a CTO, whatever, to say like, “Okay, here's the hard contract between you guys and you guys. Whatever you do inside this black box is your business. You do whatever. But between these guys, let's be, really damn clear, and any movement must be signed off by a human or me,” or. Then, and like that's that. I don't know if you have any other modifications or advice.Walden [00:46:44]: Well, I guess generally on the topic of, where humans can be useful, I found that ‘cause, some of these, really deep infra problems, sometimes just having a human that just has, really deep expertise can make a big difference. I've actually seen this come into play when actually building agents. So we've had a few friends now, try building their own coding agents, and I think one same problem that I recurringly heard a lot of them run into was this problem of like, “Oh, Grep is really slow on our agents' machines.” And so a lot of them, I assume because they're using AI and they themselves don't have, super deep infra background knowledge, say, “Okay, we're going to go build our own custom Grep index. It's going to be really fast,” and use that as a way around this problem. When we ran into this problem About like, maybe like a year and a half ago when we were, in the early days of building Devin, we obviously didn't have AI then. We just asked our, how to, how to do this. You can just swap out a new Grep index, so.Infrastructure Details: Grep, File Systems, and SandboxesSwyx [00:47:45]: What do you mean you hand-coded Devin? What?Walden [00:47:48]: It's like, can you believe we hand-wrote this code? And we had, our infra people who are really amazing, they were looking into it and they're like, “Oh, what? We realized that actually the root cause of this problem is actually super simple, but like fine-grain detail,” which is that a lot of these virtual machines actually underlying them don't use real file systems. They use these, network file systems where things are actually cached over the network actually in S3. So when you're Grepping, you're actually making network calls Every time you're doing these things, and that's why Grep is extremely slow on these machines. And so again, goes back to, what is all of the crazy infra work that we had to do to actually get these machines working. If you try to do this yourself, there are tons of small details like this, and so we had to eventually go swap out that network file system. ButSwyx [00:48:35]: I think there's a write-up about it, right? Silas did one about the virtual file system.Walden [00:48:38]: Oh, that was a whole other thing. TheSwyx [00:48:39]: Oh, that's a different thingWalden [00:48:40]: The BlockDev file storage formatSwyx [00:48:42]: I'll bring it upWalden [00:48:42]: Which is, a file system format that we built so that the VMs could be spun up and down very quickly. Basically, the intuition behind this is-Imagine you have, a terabyte of disk, and your agent only, wrote, a hundred lines of code on top of that disk. How long does it, say, take to, save and re-bring up that disk? And most systems, because you're not optimizing for this case, it's just, on the order of a terabyte of work because you have to Save all of that and bring it back up. In our system, we try to build a file system that incrementally builds on top of each other. So every time you save and bring the machine back up, you're only doing work that is proportional to effectively the diff in the file system. And so this, shaves off a lot of time in the boot-up process of Devin. I think we This is actually now outdated. We have a newer system inside of Devin. But yeah, there's a lot of tiny details you have to get right here to actually get the day-to-day experience of Devin to be good.Swyx [00:49:39]: It's, not technically agents, but it is agent infra, and when you sell an agent as a company, you sell agent plus agent infra.Walden [00:49:46]: At least the way we do it be And the other The nice thing about having the agent infra being done together is, you We get to deploy Devin in whatever environment we want now. We don't need to wait for some underlying infra provider to also go and support VPC or on-prem or FedGovCloud, for instance. So we can actually go and figure out, okay, since we own the infrastructure, how can we get that set up for you?Cloud Providers: Modal, Daytona, and Enterprise SandboxesSwyx [00:50:12]: Whereas you're Cloudflare dependent.Cole [00:50:15]: so Cloudflare runs the control plane. The sandboxes, Modal is supported. A contributor just added Daytona. E2B is on the roadmap, and I think there's an abstraction in place that if any contributor wants to add a new provider, they can add that in.Walden [00:50:32]: Well, what are, How are the customers you work with Do they generally try to then go set up a contract with another one of these third-party providers? Do they try to do the VMs in-house?Cole [00:50:44]: most of them I see using Modal. I think Modal has a greatWalden [00:50:48]: Shout out Modal.Swyx [00:50:48]: Shout out Modal.Cole [00:50:50]: I think Modal has a great offering. It captures all of the sandbox pieces you need, snapshots being a pretty big piece of that, and given that they also offer GPUs, I think it's a pretty nice offering as a whole.Swyx [00:51:04]: no debate there.Walden [00:51:07]: Modal is great, especially, I think their container offering is, the most natural, and so especially if you are willing to, forego, the full VM requirements Modal is, a really vast place you can spin something up on.Swyx [00:51:20]: Is there a point So Modal's very Python, and I feel like most workload, has really shifted to JavaScript. I don't know if you guys Get the same feeling. So, okay, when I started Landspace and IE and all these things, I was like 50/50 Python and JS, right? That's roughly. I think that's wrong now. I think JS has won. I don't know if you guys Like, I Maybe I'm overstating it, and maybe for cognition, there's, C# and Java and what have you. But for, new greenfield apps, do you feel that Do you get that sense? Does it matter?Cole [00:51:52]: I think that most of the libraries that I see in this space are Python native first, especially in theCole [00:51:58]: Observability space. That said, I think that there is a pretty big appeal of having your entire system in one language. Especially when you have both your frontend and backend communicating, you can have one central type Which is very nice.Swyx [00:52:11]: That's my case against Modal, which is Then you have to run JS. You can run JS inside Modal. It's just, one extra step That, isn't native to the runtime. I don't know ifWalden [00:52:22]: I don't knowSwyx [00:52:23]: Reviews. Do you have numbers? I don't know.Walden [00:52:25]: the one thing I don't like about Python is whenever AI, whenever it writes Python, it always does, the weirdest patterns, andSwyx [00:52:32]: Oh, because it's, mixing two and three or what?Walden [00:52:34]: I think it's something mixing two and three, yeah. The I don't know if you see this. It always tries to do, has attribute on objects as likeCole [00:52:41]: Oh, my God.Walden [00:52:41]: But it's like But that you shouldn't be doing that. It should error if there wasSwyx [00:52:45]: Because it's training on library code?Cole [00:52:47]: I think it's more of, likeCole [00:52:48]: From what I've seen, it's more of, a reward hacking mechanism where it doesn't want to basicallyWalden [00:52:54]: It'll never error.Cole [00:52:54]: It doesn't want the code to fail. And so it Even when it knows it has the attribute, it'll call getattr on a, and for a lot of my clients who have moved towards more autonomous coding, we've put that in as a lint rule That if you do getattr, your pull request is going to fail.Slop Signatures: Comments, Backwards Compatibility, and TypesSwyx [00:53:12]: Ooh, this is a fun topic. Can you tell me more about this? What else is a sign of AI coding that you have to put guards in?Walden [00:53:21]: So we were talking just before this about Opus 4.7. One of the things this new model likes to do is it writes lots of comments. Not like, it'll, comment every line, but it'll write, paragraph, PRDs, on top of every function. But I will say, to its credit, these aren't slop, descriptions like they were before. “Oh, here's what this function does.” It's like, “Oh, here's actually the r
The Bike Shed celebrates its 500th episode with hosts new and old as they reflect on the show's history and ask, what's new in your world? Our past hosts look back at their time on the show, their favourite moments while hosting, what they took away from producing the Bike Shed, and what they might do today if they were still in the hosting chair. — Your hosts for this special episode of The Bike Shed have been Joël Quenneville, Sally Hall and Aji Slater. Joining them have been our returning hosts Derek Prior, Sage Griffin, Stephanie Viccari, Chris Toomey and Stephanie Minn. Listen back to some of our guest's highlighted episodes Bike Shed 14: An Acceptable Level of Hassle with David Heinemeier Hanson Bike Shed 172: What I Believe About Software Bike Shed 180: A Citizen of the Internet with John Resig Bike Shed 302: Observability with Charity Majors Bike Shed 325: Pranting Bike Shed 404: Estimation If you would like to support the show, head over to our GitHub page, or check out our website. Got a question or comment about the show? Why not write to our hosts: hosts@bikeshed.fm This has been a thoughtbot podcast. Stay up to date by following us on social media - YouTube - LinkedIn - Mastodon - BlueSky © 2026 thoughtbot, inc.
What if you stopped treating observability as a simple insurance policy and started viewing it as a profit center? This week, Andrew sits down with Honeycomb CEO Christine Yen to explore how observability, data science, and product development are colliding in the agentic era. Christine explains why production signals must become compiler inputs for autonomous agents and how MCP tools are democratizing telemetry for entire organizations. Finally, the two discuss Honeycomb's latest Innovation Week announcements and the exact strategy for reframing observability from basic risk mitigation into a clear revenue accelerant.Learn why: LinearB is a Leader in the 2026 Gartner® Magic Quadrant™ for Developer Productivity Insight PlatformsFollow the show:Subscribe to our Substack Follow us on LinkedInSubscribe to our YouTube ChannelLeave us a ReviewFollow the hosts:Follow AndrewFollow BenFollow DanFollow today's stories:Honeycomb Blog: Read deep dives on SLOs at honeycomb.io/blogHoneycomb Innovation Week: Explore the latest announcementsRequired Reading: Check out the book Observability Engineering by Charity Majors, Liz Fong-Jones, and George Miranda.HumanX Interview: A Codebase Is No Longer the Source of Truth"Production is a Compiler Input": Chad Fowler's take on the future of code generation. Follow Christine on LinkedInOFFERSStart Free Trial: Get started with LinearB's AI productivity platform for free.Book a Demo: Learn how you can ship faster, improve DevEx, and lead with confidence in the AI era.LEARN ABOUT LINEARBAI Code Reviews: Automate reviews to catch bugs, security risks, and performance issues before they hit production.AI & Productivity Insights: Go beyond DORA with AI-powered recommendations and dashboards to measure and improve performance.AI-Powered Workflow Automations: Use AI-generated PR descriptions, smart routing, and other automations to reduce developer toil.MCP Server: Interact with your engineering data using natural language to build custom reports and get answers on the fly.
Selector is extending its AI-driven network observability capabilities into public clouds. On today’s sponsored episode, we dig into how Selector gathers and analyzes public cloud network telemetry, how it integrates cloud and on-prem network data to provide end-to-end visibility, how it integrates with third-party Application Performance Monitoring (APM) systems to correlate network and application performance,... Read more »
Selector is extending its AI-driven network observability capabilities into public clouds. On today’s sponsored episode, we dig into how Selector gathers and analyzes public cloud network telemetry, how it integrates cloud and on-prem network data to provide end-to-end visibility, how it integrates with third-party Application Performance Monitoring (APM) systems to correlate network and application performance,... Read more »
Selector is extending its AI-driven network observability capabilities into public clouds. On today’s sponsored episode, we dig into how Selector gathers and analyzes public cloud network telemetry, how it integrates cloud and on-prem network data to provide end-to-end visibility, how it integrates with third-party Application Performance Monitoring (APM) systems to correlate network and application performance,... Read more »
Did you ACTUALLY MISS OpenAI's big Workspace Agents announcement?
The race for AI dominance has created a dangerous imbalance between business velocity and cyber resilience. In this episode, host Caleb Tolin is joined by Joe Hladik, Head of Rubrik Zero Labs, and Staff Security Researcher Amit Malik to break down the findings of their latest report on agentic adoption. The discussion centers on the Agentic Paradox. This is the technical reality that tools designed to automate high-level tasks are inherently built to find the most efficient path around obstacles, including existing security policies. A primary focus is implementing a three-layer framework for AI Operations. This model targets the Tool Layer, where agents interact with databases; the Cognitive Layer, which serves as the LLM brain; and the critical Identity Layer. The conversation explores stories in which agents, without malicious intent, have caused catastrophic data loss simply by following an optimized logic path. These instances prove that agents need not be sentient to be destructive when they lack proper human-in-the-loop checkpoints. Technical hurdles of Identity Resilience are also addressed, specifically the explosion of non-human identities that spin up and down like elastic cloud infrastructure. The episode examines the fear index regarding job security, noting that 92% of leaders fear for their roles post-breach. Joe and Amit join Caleb to explore the evolution of personal liability for CISOs and the urgent need to move from basic visibility to deep observability. This is a forward-looking briefing for leaders who recognize that, in an era of autonomous routines, the human must remain the ultimate command-and-control center. What You'll Learn Define the agentic paradox to understand why AI efficiency naturally compromises traditional security guardrails. Implement a three-layer framework to secure the tool, cognitive, and identity components of AI. Transition from basic visibility to deep observability to track autonomous decision-making in real time. Mitigate prompt injection risks by auditing the input and output flows of the cognitive layer. Utilize ephemeral containers to sandbox agentic tools and prevent unauthorized database alterations. Manage the elasticity of non-human identities to maintain control over rapidly spinning AI agents. Anchor AI operations with human-in-the-loop checkpoints to ensure integrity during high-stakes executions. Episode Highlights Defining the Agentic Identity and Autonomous Routines Revenue vs. Resilience: The Drivers of AI Urgency The Three-Layer Framework for Agentic Defense Shadow AI and the Rise of Invisible Insider Threats The Context Gap: Why Rolling Back AI Actions is Hard The CISO Fear Index and Personal Liability Post-Breach Visibility vs. Observability in Elastic Identity Environments Learn more about your ad choices. Visit megaphone.fm/adchoices