POPULARITY
By Doug Green “AI goes blind at exactly the moment when you need it most.” In this Technology Reseller News podcast, Vishal Gupta, Director of Product Management at ZPE Systems, explains why AI-driven infrastructure management needs an independent path to the devices it is expected to monitor, troubleshoot and recover. AIOps platforms have become increasingly effective at detecting problems, correlating events and automating routine infrastructure operations. The problem, Gupta says, is that these systems often run on the same production infrastructure they manage. When a network outage or hardware failure occurs, the AI platform can lose both its connection to the affected equipment and access to the telemetry it needs to diagnose the problem. “That is the gap out-of-band fills,” says Gupta. Out-of-band management provides an independent management plane that remains separate from the production network. Even when the primary infrastructure is unavailable, IT teams—and increasingly AI agents—can still reach devices through console access, examine system logs and kernel messages, and take corrective action. Gupta compares the architecture to an airport. Aircraft use the runway for normal operations, while emergency and service vehicles have separate roads and infrastructure. If the runway becomes unavailable, the service infrastructure can still reach the aircraft. The same principle applies to resilient IT operations. An isolated management environment should have its own connectivity, security, routing, switching, storage and compute capabilities. It may also include failover connectivity through 4G, 5G or satellite services such as Starlink. Building Out-of-Band for a Larger Edge ZPE Systems developed its Nodegrid Net Services Router 2U, or NSR 2U, in response to customers operating increasingly large and complex edge environments. These environments can include branch offices, remote facilities, ships, oil rigs, cell sites and other locations outside traditional data centers. They frequently contain more devices, require greater bandwidth and have fewer trained personnel available on-site. The NSR 2U was designed around three priorities: greater capacity, increased resiliency and support for AI workloads. The modular platform offers 10 expansion-card slots, allowing customers to configure the system around their particular deployment. It also includes redundant, field-serviceable power supplies and fans, two NVMe storage slots with RAID support, four native 10-gigabit SFP+ ports and an increased Power over Ethernet budget. ZPE has even addressed the possibility that the out-of-band device itself could fail. Two NSR 2U systems can be interconnected so that one system can provide remote console, power and reset control for the other—effectively providing out-of-band management for the out-of-band infrastructure. Taking NVIDIA Jetson AI to Remote Locations ZPE Systems has also developed an NVIDIA Jetson AI Expansion Card for the Nodegrid NSR family. The card supports NVIDIA Jetson Orin Nano and Orin NX modules, providing local AI processing within the isolated management environment. This allows organizations to deploy AI agents close to the infrastructure and data they manage, without relying entirely on a remote cloud connection. A key capability is remote lifecycle management. IT teams can remotely flash the Jetson operating system, deploy or update AI agents and models, access the console, and power the device on, off or into recovery mode. Ordinarily, updating or recovering an edge AI device may require someone to travel to the location and connect directly to the hardware. ZPE's approach is intended to reduce those truck rolls while allowing organizations to manage distributed AI infrastructure centrally. Potential applications extend beyond AIOps. The platform can support real-time video analytics, object detection, smart recording, manufacturing quality control, sensor-data aggregation and local automation. GPIO and I2C interfaces also allow sensors measuring conditions such as temperature, vibration or voltage to feed information directly into locally running AI models. Asking the Hard Infrastructure Questions Gupta says much of the AI conversation remains focused on models, software and the token economy. Those areas are important, but they can obscure fundamental infrastructure questions. Where will an AIOps platform run? Can it survive the outage it is expected to resolve? Will it still have a path to the affected equipment? Can it access sufficiently accurate data to diagnose the problem and select the right recovery action? “If you can't answer these questions, then there's a gap in your AIOps strategy,” Gupta says. “No software and no model will fix it for you.” As AI becomes more autonomous, infrastructure resilience will determine whether AI agents can move beyond identifying failures to actually recovering from them. ZPE Systems is positioning isolated out-of-band infrastructure, the NSR 2U and edge-based Jetson AI processing as the foundation for making that transition possible. More at Enterprise Network Management Solution | ZPE Systems
What if your MCP server shipped with its own manual? Angie Jones, VP of Developer Experience at the Agentic AI Foundation, joins William and Eyvonne to break down the Skills Over MCP working group effort, which delivers Agent Skills through MCP’s existing resources primitive (think voice over IP, not skills versus MCP). Angie shares her... Read more »
What if your MCP server shipped with its own manual? Angie Jones, VP of Developer Experience at the Agentic AI Foundation, joins William and Eyvonne to break down the Skills Over MCP working group effort, which delivers Agent Skills through MCP’s existing resources primitive (think voice over IP, not skills versus MCP). Angie shares her... Read more »
In this episode, Mark Russinovich, CTO of Microsoft Azure revealed Brain, the AI-powered AIOps system that continuously monitors Azure's health, detects incidents, identifies root causes, and increasingly automates responses such as pausing problematic deployments and notifying affected customers. Built on Azure Resource Graph, Brain creates a real-time digital twin of Azure, mapping dependencies across hundreds of services, data centers, and regions. Although Brain predates the generative AI boom, years of data engineering, standardized service-level indicators (SLIs), and machine learning laid the foundation for today's capabilities. Brain combines standardized SLIs, service-specific monitoring, and third-party signals to detect anomalies, while ML models dynamically establish service baselines and correlate outages with software rollouts. Microsoft says automated notifications have reduced customer support tickets by four to six times, with 80–90% of Brain-covered services receiving notifications within 15 minutes, often in under five. The company is also layering LLM-powered agents, called Triangle, on top of Brain to streamline incident routing and eventually enable AI agents to autonomously troubleshoot and remediate outages. Learn more from The New Stack around the latest in Microsoft Azure: Meet Brain, the AI that decides when Azure is officially down Microsoft's pitch to enterprises: Ditch Azure Repos for GitHub, despite its rocky reliability record Join our community of newsletter subscribers to stay on top of the news and at the top of your game.
Malcolm Matalka joins William and Eyvonne to challenge the narrative that Infrastructure as Code (IaC) is dead. Malcolm argues that the real value of IaC was never the syntax, but state and governance. Together they examine whether the state was a file problem at all, or a distributed systems problem in a JSON costume. Episode... Read more »
Malcolm Matalka joins William and Eyvonne to challenge the narrative that Infrastructure as Code (IaC) is dead. Malcolm argues that the real value of IaC was never the syntax, but state and governance. Together they examine whether the state was a file problem at all, or a distributed systems problem in a JSON costume. Episode... Read more »
The Pope issued a recent encyclical on AI, urging developers to safeguard human agency in the age of artificial intelligence. Eyvonne and William explore this encyclical, moving beyond the headlines to the core message regarding human dignity. They examine how the document provides a values-based framework for evaluating technology and the need for a balanced... Read more »
The Pope issued a recent encyclical on AI, urging developers to safeguard human agency in the age of artificial intelligence. Eyvonne and William explore this encyclical, moving beyond the headlines to the core message regarding human dignity. They examine how the document provides a values-based framework for evaluating technology and the need for a balanced... Read more »
en Fallon, vice president of worldwide channel and partner ecosystem networking sales At HPE Discover Las Vegas this week, HPE pushed its networking story to the centre of the event – from autonomous AIOps capabilities to a unified SASE platform – and the channel is central to how it plans to execute on some ambitious market share targets. ChannelBuzz.ca sat down on-site with Ben Fallon, vice president of worldwide channel and partner ecosystem networking sales, to talk about what the announcements mean in practice for Canadian partners. On the self-driving network vision – a major theme in the general sessions this week – Fallon pointed to HPE Aruba Mist as the concrete proof point: autonomous remediation that partners can toggle on in the dashboard for known network problems, no human click required. “Autonomous networking, with that human deciding where they want that to take place, is already real,” he said. On the Aruba and Juniper Networks platform integration – a frequent question from partners navigating two management platforms – Fallon described a “build once, deploy twice” philosophy built on microservices architecture, keeping both platforms differentiated by use case while accelerating innovation through cross-pollination rather than forced convergence. The SASE and security opportunity produced one of the clearest channel statements of the conversation: “Pretty much 100% of our security sales go through partners. There is no other path.” With HPE publicly targeting a $1 billion security business, Fallon said the partner base is nowhere near saturated – and that competency-based incentives within the Partner Ready Vantage program are in place to bring more networking-pedigreed partners into that conversation. A formal partner program unification is on track for November, with a stated focus on simplifying certification, deal registration, and rebates – and new incentives aimed squarely at winning net-new networking customers away from competing vendors. Read Full Transcript Robert Dutt: Today’s episode of In The Channel is brought to you by HPE Discover 2026. Discover runs June 15-18 at the Venetian in Las Vegas. Discover what’s next at hpe.com/discover. Hello and welcome to In The Channel from ChannelBuzz.ca, bringing news and information to the Canadian IT channel community for the last 16 years. I’m Robert Dutt, editor of ChannelBuzz.ca, and your host for the show. We’re coming to you this week from HPE Discover Las Vegas, where HPE has been rolling out a significant set of announcements across networking, cloud, and AI infrastructure. The embargoes are lifted, and the Partner Growth Summit is in the books, so we can actually get into the substance of things. My guest is Ben Fallon, vice president of worldwide channel and partner ecosystem networking sales at HPE. Ben came to this role via the Juniper side of the house. He was running global partner and commercial sales for Juniper Networks when the acquisition closed, and moved into leading the combined networking channel earlier this year. His session at Discover this week was called “Betting on HPE Networking,” which turned out to be a pretty useful frame for a conversation. We got into what self-driving networks actually mean for a partner having a Monday morning conversation with a customer, the Aruba and Mist integration story, the SASE and security opportunity, and what partners can expect when the unified program formally launches in November. Let’s get right into it. My chat with Ben Fallon. Ben, thanks for taking the time. I appreciate it. I know it’s a busy week on site here, I’m sure. Ben Fallon: It is. It’s a fun week. We’ve got thousands of partners here, but it’s great to be here with you. Robert Dutt: For listeners who don’t know you or your role, can you give me a quick rundown on what you do here and how you came to be leading networking channels for HPE? Ben Fallon: Yeah, so like you said, I lead the global networking channel for HPE. I’ve spent the last 25-odd years in the industry, have led channels for a number of the significant vendors in the market. I was part of the Juniper acquisition, most recently running one of the global sales segments, and in January moved over to lead the channel. We’ve got a fantastic opportunity in front of us. Robert Dutt: I like that you frame it as you’re part of the Juniper acquisition. You’re not taking entire credit for them acquiring Juniper to get your talent. Ben Fallon: Absolutely not, no. It was a bonus. Robert Dutt: Absolutely. Your session this week is called “Betting on HPE Networking.” It’s a pretty confident way of looking at it, and obvious given the milieu. Walk me through what the bet looks like from where you sit. What are you asking partners to bet on, and why now? Ben Fallon: Yeah, so for me, it’s like when you look at a bet, you’ve got to make sure it’s a good one. No one wants to be playing the lottery. That’s got the worst chance of winning. The more strategy that you actually bring into a game, along with some execution, increases your chance of winning. So for us, what increases the chance of winning with HPE Networking is cross-selling. The more you’re selling across the portfolio, the more you’re going to engage with our account teams, the more problems you’re going to solve for our customers. And also, that’s where you can earn the most amount of rebates, and where the program is really geared towards. So if you make a bet on us, we’re making a bet on you, and you’ll get that back in profitability and customer satisfaction. Robert Dutt: Cross-selling within networking, across the HPE portfolio, or… Ben Fallon: All of the above. So you can absolutely cross-sell within the portfolio, whether you’re selling campus and branch, or you want to move into selling more security solutions. Or if you’re selling the hybrid cloud solution portfolio from HPE, you need to start getting involved in networking, because it’s going to expand your opportunity, and we know the network is at the heart of all of these AI workloads. Robert Dutt: One of the big presentations here is about taking the idea of self-driving networks from vision to reality. For a lot of partners, though, the question is always, “What do I take to my customer?” On Monday morning, how do partners translate that message around self-driving networks to a concrete conversation with staff at a customer, and make it map with their care-abouts? Ben Fallon: Yeah, sure. Well, look, complexity is only increasing. We know there are talent shortages. We know that it’s almost an impossible task to keep up with all the vulnerabilities that are created through AI. And so you have to have AI as part of your defense. So what’s real? Let’s take something like HPE Mist, where that has autonomous actions now built into the dashboard. So we know for certain problems that come up on the network, we know how to remediate them. We don’t need a person to go and click a button. You can literally switch on a toggle, and off it goes. So autonomous networking, with that human deciding where they want that to take place, is already real. Robert Dutt: You touch on Mist. One thing I do hear from partners sometimes is with the Aruba and Juniper integration, the two platforms you’ve got with Aruba Central and Mist, moving toward common capabilities, but it sounds like the vision is not to merge. What do you tell the partner who’s been selling one side of that equation or the other? And now that we’ve kind of got one HPE networking, what does it mean in practice, basically? Ben Fallon: Yeah, well, you touched on self-driving. That’s a unified vision across the entire portfolio. And then we’ve got this strategy of cross-pollination. I think if you look at a lot of acquisitions over the years, they’ve spent so long arguing over maybe not a feature, but how do you actually get to that feature to be capable? And innovation dies when that happens. If you want innovation to actually accelerate, which is what we’re seeing, you take the best from each platform, and because they’re built with a microservices architecture, you can build once, deploy twice, and it becomes this incredible boon of innovation on the platform. So I’d say that is real, because customers are voting with their wallet. So there’s a decent amount of cross-pollination, but each kind of remains aimed towards its focus. Robert Dutt: That’s it. Ben Fallon: And really what I see with partners is they see this as a growth play in the same way that we do. This is about finding new opportunity. So they may have served some SMB customers with some on-prem part of the Aruba portfolio. Now they’re wanting to get into some mid-sized lower enterprise, and they’re seeing that Mist has some capability that helps get them there. So it’s a growth play for us, and it’s a growth play for the partner. Robert Dutt: One of the things that caught my attention in the announcements this week was the unified SASE story – bringing SD-WAN and SSE under one management pane. You guys have talked about a billion-dollar security ambition. Pretty big number. What’s the channel’s role in getting to that? And for a partner who hasn’t historically led with networking security, what’s kind of the on-ramp or the easiest first step? Ben Fallon: Yeah. So first of all, obviously, we’ve got this universal zero-trust network architecture, which we’re really leaning into. And it’s about bringing together the different parts of the security portfolios from across HPE. And obviously with the Juniper acquisition, that brought an even richer portfolio. For partners, pretty much 100% of our security sales go through partners, so there is no other path. And what we’re really looking for is – we have some very, very capable, specialized partners on security – I think there’s a bigger opportunity for more partners to be selling HPE networking and security solutions. We’re just getting started. We’re already posting some great numbers. We had some incredible growth just last quarter, and there’s still more partners can do. We are not saturated from the partner landscape selling our security portfolio, so lots of opportunity there. Robert Dutt: Those additional partners in that space – do you see them being primarily folks who come in from other parts of the HPE network, existing specialists in security who maybe haven’t worked with you in the past, a little bit of both? What’s kind of the… Ben Fallon: It’s a bit of a combination, but you always have to focus. You can’t go everywhere. And where we’re focusing is on partners that have a pedigree in networking with us, because we’re increasingly seeing that there’s a great attach opportunity, and the convergence of the network and security we think is only going to accelerate. Robert Dutt: Are we at the point of having a formal program, that kind of thing, to bring those partners on board, or to enable and encourage the partners who are in the HPE sphere, but not yet? Ben Fallon: Yeah, we do. We have, as part of our Partner Ready Vantage program, our broad certifications that are part of that, and that’s how you get to platinum, gold, silver, etc. But then we have competencies, and we have a number of security competencies that partners can build up that capability. They can pick different parts of the portfolio. They could be brand new to networking, but build up competency in security, and that will bring technical competence and capability, but also incremental profitability for them as well. Robert Dutt: A lot of talk this week, obviously, about the disruption around VMware – customers reconsidering virtualization strategies and how that drives the refresh cycles within the data center on some of the compute and storage hardware, all that kind of good stuff. Does that also create a network refresh opportunity? Ben Fallon: So there can be opportunities that do arise. I don’t know if that’s the biggest piece that’s driving growth in data center networking right now. I think the AI boom is doing a significant job there, and probably dwarfs anything else. But what you’ll see is announcements this week around how we’re, really from a technology perspective, bringing more parts of the portfolio together from across the hybrid cloud portfolio and networking. Because really, that’s what customers want. They want integrated technology that solves their problems, and that’s what we’re focused on. Robert Dutt: From a Canadian channel perspective, where do you see the biggest networking opportunities today? I’m going to guess your answer to the last question strongly informs the answer to this one. But what are the biggest opportunities in the back half of the year? And what’s your ask of Canadian partners who are listening to this? Ben Fallon: Yeah. Well, there are two things I think are the biggest opportunity. One is cross-selling. If you’re selling part of the HPE portfolio today, look at how you can integrate across the stack – whether that’s the full HPE stack, or whether it’s specific to networking. There’s a huge opportunity there, and we’re seeing that partners that have adopted that are growing faster than anyone else. Second, new logos – going after new customers. We’re here to win. We’re here to be number one, and we’ll do that first in wireless networking. And to do that, we need new customers. And you’ll see new incentives and new programs come out in November that will put even more wood behind the arrow – that’s going to make it an incredible opportunity for partners to go and solve the networking crimes of other vendors and bring them into the light of a self-driving network. Robert Dutt: You guys are obviously deep into the process of integrating programs between legacy HPE and legacy Juniper. We have the November 1 date, I believe, as the formalized launch date for that becoming one. What can partners expect coming out of that at a programmatic level on the networking side? Ben Fallon: Yeah. So what we’re doing is, first of all, looking at the experience partners have – everything from how they get certified, trying to simplify that and make sure that they’re not having to do multiple layers and duplicative actions. We’re working on the experience when it comes to things like registering a deal, getting rebates, keeping it simple. I think other vendors I’ve seen, you need a bit of a rocket science degree to figure out how all of these different programs and rebates come together. We’re focusing on keeping it simple, we’re focused on driving action, and most of all – which I think is often missed – we’re making sure that our sales teams know how to engage with partners really well and go and win deals together. Robert Dutt: Good luck on a big week here at Discover, and thanks for taking the time once again. Ben Fallon: Appreciated. And we love working with our Canadian partners, and just a big thank you to all of them that are on board already. Robert Dutt: There you have it, Ben Fallon from HPE. I’d like to thank Ben for his time. We were literally recording between sessions at Discover, and I appreciate him making it work. And thank you for listening as well. A few things that stuck with me from this one. The self-driving network story has been fairly abstract for a while, but his Mist example – autonomous remediation actions you can toggle on in the dashboard, no human in the loop for known problem types – it’s the most concrete I’ve heard it get. That’s actually something you can put in front of a customer. The other thing worth sitting with: “pretty much 100% of our security sales go through partners. There is no other path.” That’s what Ben said. If you’re an HPE networking partner who hasn’t yet built a security practice, and HPE is out there talking about a billion-dollar security ambition, someone is going to capture that opportunity. Make sure it’s you. And for partners who may have walked away from the Juniper side of the portfolio at acquisition time and have been watching from the sidelines, November is shaping up to be the moment to take another look. Simplified programs, new incentives, a unified experience. It’s worth paying attention to. If you found the episode useful, we’d love to have you subscribe to the podcast. You’ll find us on Apple Podcasts, Spotify, YouTube, and most of the major podcast directories. If you have a moment to leave a rating or a review, it always helps. Until next time, I’m Robert Dutt for ChannelBuzz.ca, and I’ll see you in the channel.
By Doug Green “We're absolutely on the path, and we're not talking five, six, seven years. We're talking in the next 18 to 24 months.” In this episode of the Technology Reseller News podcast, Doug Green speaks with Josh Kindiger, COO and co-founder of Grokstream, about the company's new L1 Agent and what it means for the future of AI-driven network and IT operations. Grokstream is the company behind Grok, an AI-powered predictive agent platform for network and IT operations. The platform comes out of the event intelligence and AIOps space and is designed to help operations teams identify, triage, and resolve recurring issues more efficiently. Kindiger says Grokstream recently released its first role-based agent, the L1 Agent, in beta. The full production release is expected in Q2. The agent is already being used with customers to prove out real-world capabilities. Because many organizations remain cautious about AI-driven automation, Grokstream is starting with low-risk, repeatable use cases. In many operations centers, Kindiger notes, the same incidents occur repeatedly, sometimes accounting for as much as 70% of activity. The L1 Agent is designed to recognize those patterns and guide operators through triage and resolution. For example, if a recurring issue requires a service restart, the system can recommend or automate that step. If a pattern points to a commercial power outage at a site, the agent can help avoid unnecessary dispatches while monitoring backup power systems. Kindiger says the goal is not to remove human oversight immediately, but to build trust through guardrails, staged automation, and operator control. Low-risk automations can be handled end to end, while higher-risk actions may require human approval. The podcast also explores the broader opportunity for enterprises, MSPs, and CSPs. Kindiger says service providers and managed service providers face growing pressure to improve efficiency, reduce costs, and differentiate in competitive markets. AI-driven operations can help them respond faster, lower manual workload, and deliver better service outcomes. The long-term direction is clear: autonomous network operations are coming. Kindiger says companies should begin now because foundational work is needed before they can fully benefit from automation. For MSPs and CSPs, he says the urgency is even greater. Cost pressure is shaping renewals and new customer wins, and AI-powered operations may become a competitive advantage. Learn more at www.grokstream.com
William and Eyvonne discuss recent tech news, including the growing political and community opposition to AI data centers driven by fears over power and water usage. They also analyze the “AI Chip War” as hyperscalers such as AWS and Google invest in specialized silicon for training and inference. Episode Links: Amid backlash, O'Leary Digital CEO... Read more »
William and Eyvonne discuss recent tech news, including the growing political and community opposition to AI data centers driven by fears over power and water usage. They also analyze the “AI Chip War” as hyperscalers such as AWS and Google invest in specialized silicon for training and inference. Episode Links: Amid backlash, O'Leary Digital CEO... Read more »
Grokstream: Predictive and Agentic AI Moves IT Operations Toward Self-Healing, Podcast, Grokstream's platform is designed to operate from signals, not noise. The system fuses telemetry across domains, learns continuously from operational data and human feedback, and creates a unified source of truth for IT operations. That allows teams to move beyond correlation and toward understanding what is happening, why it is happening and what should be done next. By Doug Green Grokstream says the next generation of IT operations will not be built around more dashboards, more rules, or faster alert routing. It will be built around AI that can learn, reason, remember, recommend and eventually act with governed autonomy. “Agentic AI must be governed by design,” said Josh Kindiger, CEO of Grokstream. “Predictive intelligence is powerful, but safe, explainable autonomy is what drives real adoption.” In this Technology Reseller News podcast, Doug Green speaks with Josh Kindiger, Co-Founder and COO of Grokstream, about how the company is helping MSPs, CSPs and enterprise IT organizations move from reactive operations toward predictive, self-healing IT environments. The conversation comes as Grokstream advances its Grok L1 Agent, a new role-based agent designed for frontline IT operations teams. The L1 Agent is intended to reduce alert noise before incidents reach the queue, provide intelligent summaries, identify likely root causes, recommend next-best actions and trigger approved remediations inside tools such as Slack, Microsoft Teams and existing IT workflows. For service providers and enterprise operations teams, the problem is familiar. More tools often mean more alerts, but not necessarily more clarity. Traditional rules-based AIOps platforms can help with deduplication and routing, but they often stop short of true incident compression, causal reasoning and prevention. Grokstream is taking a different approach by combining classical machine learning, causal intelligence and generative AI into a single cognitive AI layer. Kindiger explains that Grokstream's platform is designed to operate from signals, not noise. The system fuses telemetry across domains, learns continuously from operational data and human feedback, and creates a unified source of truth for IT operations. That allows teams to move beyond correlation and toward understanding what is happening, why it is happening and what should be done next. A central theme of the podcast is the difference between AI that summarizes and AI that reasons. Grokstream argues that true agentic AI is not simply an LLM attached to a workflow. It requires memory, context, policy guardrails, procedural intelligence and the ability to improve over time. In Grokstream's model, agents begin as assisted tools, then move toward trusted operators and eventually toward predictive autonomous systems. The first practical on-ramp is the L1/NOC environment, where many organizations see the fastest measurable impact. Grokstream says its approach can deliver 2–3x more incident compression beyond traditional deduplication and rules-based correlation, while reducing L1 workload by more than 50% through noise compression, guided resolution and fewer unnecessary escalations. The timing is significant. Grokstream recently announced that Cirion Technologies selected the Cognitive Grok AI platform to support AI-driven predictive operations across Latin America's digital infrastructure. That deployment highlights the growing demand for systems that can detect emerging issues across network, transport and infrastructure layers before customer-facing impact occurs. For MSPs, CSPs and enterprise IT leaders, the message is clear: operational scale cannot be achieved simply by adding more people or more monitoring tools. The next step is an intelligence layer that can unify data, predict impact, explain cause and support governed automation. Grokstream is positioning Grok as that layer: a predictive and agentic AI platform that helps operations teams reduce noise, prevent incidents, improve engineer experience and move toward self-healing IT operations. Learn more at https://grokstream.com/ Related Grokstream Stories on Telecom Reseller Grokstream's Cognitive Grok® AI Platform Selected by Cirion Technologies to Power AI-Driven, Predictive Operations Across Latin America's Digital Infrastructure https://telecomreseller.com/2026/05/20/grokstreams-cognitive-grok-ai-platform-selected-by-cirion-technologies-to-power-ai-driven-predictive-operations-across-latin-americas-digital-infrastructure/ Grokstream Announces Grok® L1 Agent to Advance Predictive and Agentic AI for IT Operations https://telecomreseller.com/2026/04/06/grokstream-announces-grok-l1-agent-to-advance-predictive-and-agentic-ai-for-it-operations/ More Grokstream coverage on Telecom Reseller https://telecomreseller.com/?s=grokstream/
Today our Packet Pushers team assembles to discuss whether the grass is greener on the NetOps or DevOps side of the telemetry fence. William of The Cloud Gambit, Scott of Total Network Operations, and Ned and Kyler of Day Two DevOps discuss the difficulties and differences of getting telemetry and state from devices across different... Read more »
Today our Packet Pushers team assembles to discuss whether the grass is greener on the NetOps or DevOps side of the telemetry fence. William of The Cloud Gambit, Scott of Total Network Operations, and Ned and Kyler of Day Two DevOps discuss the difficulties and differences of getting telemetry and state from devices across different... Read more »
As the software world is transforming from cloud native to AI-native, observability must transform with it. But how exactly? How do we apply this in an existing enterprise with established processes and practices?In this PurePerformance episode, Andi Grabner hosts Hilliary Lipsig and Rob Rati to discuss their new book, Observability in the AI‑Native Era. The conversation explores how AIOps, automation, and modern observability must evolve as systems become cloud‑native, data‑heavy, and AI‑driven.We talk about why old alerting and SLO models no longer scale, how to balance AI with automation and human judgment, and why trust, security, and compliance matter more than ever when machines start making operational decisions. A must‑listen for SREs, platform engineers, and engineering leaders navigating the AI‑native future.Links we discussedBook on Amazon: https://www.amazon.com/Observability-AI-Native-Era-Artificial-Intelligence-ebook/dp/B0GHZH1YFLHilliary LinkedIn: https://www.linkedin.com/in/hilliary-lipsig-a5935245/Rob LinkedIn: https://www.linkedin.com/in/roberthrati/Andi LinkedIn: https://www.linkedin.com/in/grabnerandi/
Eyvonne and William sit down with Joseph Nicholson, a Network Operations Engineer with NTT DATA, to share how public speaking transformed his career and technical experience. Joseph went from a terrifying ten minute lightning talk at AutoCon 2 to presenting 45-minute sessions at conferences like NANOG. Together they discuss how conversations in conference halls influenced... Read more »
Eyvonne and William sit down with Joseph Nicholson, a Network Operations Engineer with NTT DATA, to share how public speaking transformed his career and technical experience. Joseph went from a terrifying ten minute lightning talk at AutoCon 2 to presenting 45-minute sessions at conferences like NANOG. Together they discuss how conversations in conference halls influenced... Read more »
Eyvonne Sharp and William Collins speak with Sif Baksh, Principal Solutions Architect at Tines, to discuss the power of automation. Sif shares some personal stories of how he has been able to use automation to innovate and modernize networking operations. They also discuss the importance of learning AI and using it as a tool, how... Read more »
Eyvonne Sharp and William Collins speak with Sif Baksh, Principal Solutions Architect at Tines, to discuss the power of automation. Sif shares some personal stories of how he has been able to use automation to innovate and modernize networking operations. They also discuss the importance of learning AI and using it as a tool, how... Read more »
The episode identifies a structural shift in the integration of generative AI within organizational workflows: variable cost models, unpredictable output quality, and heightened accountability requirements are converging to reshape managed services operations. This shift is exemplified by Anthropic's move toward usage-based pricing for Claude Enterprise, combining compute consumption with per-user fees, and by reports of major enterprises and intelligence agencies piloting dedicated cybersecurity-focused generative AI models. These trends expose IT service providers, especially MSPs, to cost volatility, operational risk, and new governance challenges as generative AI transitions from experimental implementation to core workflow tooling. Primary evidence includes Anthropic's revised pricing strategy, which replaces predictable licensing with usage-based billing, introducing financial unpredictability for heavy users. The episode cites reporting from The Verge and The Guardian, noting that AI-generated outputs can create hidden labor through the need for manual review and corrections, while undetected errors escalate into operational disputes and rework. The implementation of generative AI in security-sensitive environments underscores the need to scrutinize how AI-driven processes are metered and governed. Supporting developments reinforce this shift: MSP platform providers such as Enable are embedding generative AI directly into operational workflows, connecting third-party tools to live data. This creates the need for controls over what AI systems can access, approve, and log, particularly in multi-tenant environments. Meanwhile, outcome-based service agreements—such as fixed response-time SLAs—set new client expectations for measurable performance and accountability in AI operations. The market is also rewarding those who wrap unmanaged technology surfaces, like BYOD or AI tooling, with enforceable policies and auditable evidence trails. Operational implications for MSPs include increased pressure on margins due to AI's variable usage costs colliding with fixed-fee contracts, the challenge of capturing and reporting hidden labor from AI output review, and the necessity for evidence-based governance. Service providers unable to implement and sell AI operations management (“AIOps”) as a billable, controlled service risk becoming de facto shock absorbers for unpriced spend, rework, and disputes. Those who standardize on enforceable budgets, approval gates, audit trails, and compliance-ready reporting stand to protect service margins and reduce liability exposure. 00:00 AI Cost Reckoning 02:39 AI Governance Gap 04:44 Govern or Lose 07:12 Why Do We Care? Supported by: TimeZest Zero Networks
Vibe coding: give AI a description of what you want, the model writes the code, you ship it, and then you hope for the best. It works great for side projects, but it can fall apart the moment you point an AI agent at production infrastructure. Today, William and Eyvonne sit down with John Capobianco,... Read more »
Vibe coding: give AI a description of what you want, the model writes the code, you ship it, and then you hope for the best. It works great for side projects, but it can fall apart the moment you point an AI agent at production infrastructure. Today, William and Eyvonne sit down with John Capobianco,... Read more »
Web and Mobile App Development (Language Agnostic, and Based on Real-life experience!)
In this episode, Michael Nappi, Chief Product and Engineering Officer at ScienceLogic, shares insights into AI Ops, its role in modern IT management, and how it helps large enterprises and MSPs streamline their infrastructure monitoring and management. Discover how AI-driven automation and observability are transforming IT operations.
Databricks Roundtable episode: Operationalizing AI Agents: From Experimentation to Production. Join the Community: https://go.mlops.community/YTJoinInGet the newsletter: https://go.mlops.community/YTNewsletterMLOps GPU Guide: https://go.mlops.community/gpuguideBig shout-out to Databricks for the collaboration!// AbstractThis panel discusses the real-world challenges of deploying AI agents at scale. The conversation explores technical and operational barriers that slow production adoption, including reliability, cost, governance, and security.The panelists also examine how LLMOps, AIOps, and AgentOps differ from traditional MLOps, and why new approaches are required for generative and agent-based systems. Finally, experts define success criteria for GenAI frameworks, with a focus on robust evaluation, observability, and continuous monitoring across development and staging environments.// BioSamraj MoorjaniSamraj is a software engineer working on the Agent Quality team. Previously, Samraj worked at Meta on ads/product classification research and AppLovin on MLOps. Samraj graduated with a BS+MS in Computer Science from UIUC, advised by Professor Hari Sundaram, where he worked on controllable natural language generation to produce appealing, interpretable science to combat the spread of misinformation. He also worked with Professor Wen-mei Hwu on accelerating LLM inference through extreme sparsification.Apurva MisraApurva is an AI Consultant at Sentick, focusing on assisting startups with their AI strategy and building solutions. She leverages her extensive experience in machine learning and a Master's degree from the University of Waterloo, where her research bridged driving and machine learning, to offer valuable insights. Apurva's keen interest in the startup world fuels her passion for helping emerging companies incorporate AI effectively. In her free time, she is learning Spanish, and she also enjoys exploring hidden gem eateries, always eager to hear about new favourite spots!Ben EpsteinBen was the machine learning lead for Splice Machine, leading the development of their MLOps platform and Feature Store. He is now the Co-founder and CTO at GrottoAI, focused on supercharging multifamily teams and reducing vacancy loss with AI-powered guidance for leasing and renewals. Ben also works as an adjunct professor at Washington University in St. Louis, teaching concepts in cloud computing and big data analytics.Hosted by Adam Becker// Related LinksWebsite: https://www.databricks.com/https://mlflow.org/~~~~~~~~ ✌️Connect With Us ✌️ ~~~~~~~Catch all episodes, blogs, newsletters, and more: https://go.mlops.community/TYExploreJoin our Slack community [https://go.mlops.community/slack]Follow us on X/Twitter [@mlopscommunity](https://x.com/mlopscommunity) or [LinkedIn](https://go.mlops.community/linkedin)] Sign up for the next meetup: [https://go.mlops.community/register]MLOps Swag/Merch: [https://shop.mlops.community/]Connect with Demetrios on LinkedIn: /dpbrinkmConnect with Samraj on LinkedIn: /samrajmoorjani/Connect with Apurva on LinkedIn: /apurva-misra/Connect with Ben on LinkedIn: /ben-epstein/Connect with Adam on LinkedIn: /adamissimo/Timestamps:[00:00] Introduction[02:30] AI Agents in Operations[04:36] AI Strategy Consulting[05:30] Agent Quality Focus[06:17] AI Agent Expectations[11:44] AI Use Cases Evolution[15:25] Agent Expectations Adjustment[17:41] Agent Quality Monitoring[23:22] Trust in GenAI Systems[33:33] Data Prep vs Product Thinking[40:27] Quality Systems Distinction[44:54] Q & A[1:00:57] Wrap up
William Collins and Eyvonne Sharp invite Skylar Sands, Senior Automation Engineer at World Wide Technology, to discuss what it means to integrate AI into the daily workflow in a meaningful way. Together they break down the shift in the automation engineer's role now that AI can instantly generate the “toolkit” of Python, Ansible, and Bash,... Read more »
William Collins and Eyvonne Sharp invite Skylar Sands, Senior Automation Engineer at World Wide Technology, to discuss what it means to integrate AI into the daily workflow in a meaningful way. Together they break down the shift in the automation engineer's role now that AI can instantly generate the “toolkit” of Python, Ansible, and Bash,... Read more »
In this sponsored episode, FluidCloud co-founders Sharad Kumar and Harshit Omar sit down with William and Eyvonne to discuss how FluidCloud tackles multi-cloud portability. They detail how FluidCloud acts as a cloning platform that scans an existing cloud or VMware environment, extracts complex infrastructure configurations (including compute and storage, as well as firewall rules and... Read more »
In this sponsored episode, FluidCloud co-founders Sharad Kumar and Harshit Omar sit down with William and Eyvonne to discuss how FluidCloud tackles multi-cloud portability. They detail how FluidCloud acts as a cloning platform that scans an existing cloud or VMware environment, extracts complex infrastructure configurations (including compute and storage, as well as firewall rules and... Read more »
The tech industry is split between two fantasies – that AI writes production software while you get coffee, and that everything AI touches is slop. The reality is messier and more interesting: AI tools are force multipliers for people who already know what good looks like, and an expertise amplifier disguised as an easy button. ... Read more »
The tech industry is split between two fantasies – that AI writes production software while you get coffee, and that everything AI touches is slop. The reality is messier and more interesting: AI tools are force multipliers for people who already know what good looks like, and an expertise amplifier disguised as an easy button. ... Read more »
William and Eyvonne tackle the biggest AI stories of early 2026. They dissect Matt Schumer’s viral “Something Big is Happening” essay – agreeing professionals need to skill up now while pushing back on the doomsday framing with real-world examples from engineering disciplines. The conversation takes a fascinating turn as Eyvonne draws a parallel between AI-assisted... Read more »
William and Eyvonne tackle the biggest AI stories of early 2026. They dissect Matt Schumer’s viral “Something Big is Happening” essay – agreeing professionals need to skill up now while pushing back on the doomsday framing with real-world examples from engineering disciplines. The conversation takes a fascinating turn as Eyvonne draws a parallel between AI-assisted... Read more »
We’ve spent a decade figuring out how to (more or less) securely authenticate humans. Now AI agents are crashing the party, and identity just got a whole lot more complicated. Today we sit down with Dan Moore, Senior Director of CIAM Strategy and Identity Standards at FusionAuth, to explore the collision course between artificial intelligence... Read more »
We’ve spent a decade figuring out how to (more or less) securely authenticate humans. Now AI agents are crashing the party, and identity just got a whole lot more complicated. Today we sit down with Dan Moore, Senior Director of CIAM Strategy and Identity Standards at FusionAuth, to explore the collision course between artificial intelligence... Read more »
Cloud bills are climbing, AI pipelines are exploding, and storage is quietly becoming the bottleneck nobody wants to own. Ugur Tigli, CTO at MinIO, breaks down what actually changes when AI workloads hit your infrastructure, and how teams can keep performance high without letting costs spiral. In this conversation, we get practical about object storage, S3 as the modern standard, what open source really means for security and speed, and why “cloud” is more of an operating model than a place. Key takeaways• AI multiplies data, not just compute, training and inference create more checkpoints, more versions, more storage pressure • Object storage and S3 are simplifying the persistence layer, even as the layers above it get more complex • Open source can improve security feedback loops because the community surfaces regressions fast, the real risk is running unsupported, outdated versions • Public cloud costs are often less about storage and more about variable charges like egress, many teams move data on prem to regain predictability • The bar for infrastructure teams is rising, Kubernetes, modern storage, and AI workflow literacy are becoming table stakes Timestamped highlights00:00 Why cloud and AI workloads force a fresh look at storage, operating models, and cost control 00:00 What MinIO is, and why high performance object storage sits at the center of modern data platforms 01:23 Why MinIO chose open source, and how they balance freedom with commercial reality 04:08 Open source and security, why faster feedback beats the closed source perception, plus the real risk factor 09:44 Cloud cost realities, egress, replication, and why “fixed costs” drive many teams back inside their own walls 15:04 The persistence layer is getting simpler, S3 becomes the standard, while the upper stack gets messier 18:00 Skills gap, why teams need DevOps plus AIOps thinking to run modern storage at scale 20:22 What happens to AI costs next, competition, software ecosystem maturity, and why data growth still wins A line worth keeping“Cloud is not a destination for us, it's more of an operating model.” Pro tips for builders and tech leaders• If your AI initiative is still a pilot, track egress and data movement early, that is where “surprise” costs tend to show up • Standardize around containerized deployment where possible, it reduces the gap between public and private environments, but plan for integration friction like identity and key management • Treat storage as a performance system, not a procurement line item, the right persistence layer can unblock training, inference, and downstream pipelines What's next:If you're building with AI, running data platforms, or trying to get your cloud costs under control, follow the show and subscribe so you do not miss upcoming episodes. Share this one with a teammate who owns infrastructure, data, or platform engineering.
In this episode, we sit down with Adam Zimman, author and VC advisor, to explore the world of progressive delivery and why shipping software is only the beginning. Adam shares his fascinating journey through tech—from his early days as a fire juggler to leadership roles at EMC, VMware, GitHub, and LaunchDarkly – and how those... Read more »
In this episode, we sit down with Adam Zimman, author and VC advisor, to explore the world of progressive delivery and why shipping software is only the beginning. Adam shares his fascinating journey through tech—from his early days as a fire juggler to leadership roles at EMC, VMware, GitHub, and LaunchDarkly – and how those... Read more »
The industry has pivoted from scripting to automation to orchestration – and now to systems that can reason. Today we explore what AI agents mean for infrastructure with Chris Wade, Co-Founder and CTO of Itential. We also dive into the brownfield reality, the potential for vendor-specific LLMs trained on proprietary knowledge, and advice for the... Read more »
The industry has pivoted from scripting to automation to orchestration – and now to systems that can reason. Today we explore what AI agents mean for infrastructure with Chris Wade, Co-Founder and CTO of Itential. We also dive into the brownfield reality, the potential for vendor-specific LLMs trained on proprietary knowledge, and advice for the... Read more »
Dr. Chris Marshall analyzes AI from all angles including market dynamics, geopolitical concerns, workforce impacts, and what staying the course with agentic AI requires.Chris and Kimberly discuss his journey from theoretical physics to analytic philosophy, AI as an economic and geopolitical concern, the rise of sovereign AI, scale economies, market bubbles and expectation gaps, the AI value horizon, why agentic AI is harder than GenAI, calibrating risk and justifying trust, expertise and the workforce, not overlooking Rodney Dangerfield, foundational elements for success, betting on AIOps, and acting in teams. Dr. Chris L Marshall is a Vice President at IDC Asia/Pacific with responsibility for industry insights, data, analytics and AI. A former partner and executive at companies such as IBM, KPMG, Oracle, FIS, and UBS, Chris's mission is to translate innovative technologies into industry insights and business value for the digital economy.Related ResourcesData and AI Impact Report: The Trust Imperative (IDC Research)A transcript of this episode is here.
In this year-end episode, William and Eyvonne recap their experiences at AutoCon 4 in Austin, Texas. They discuss the conference’s new multi-track format, including Eyvonne’s presentation in the leadership track on why technical projects fail. The conversation dives into how AI tools like Google Gemini can augment – not replace – human creativity, from research... Read more »
In this year-end episode, William and Eyvonne recap their experiences at AutoCon 4 in Austin, Texas. They discuss the conference’s new multi-track format, including Eyvonne’s presentation in the leadership track on why technical projects fail. The conversation dives into how AI tools like Google Gemini can augment – not replace – human creativity, from research... Read more »
In this sponsored episode recorded live at AutoCon 4 in Austin, we sit down with Peter Sprygada, Chief Architect at Itential, to discuss Itential’s on-stage announcement of FlowAI. Peter shares his journey from network engineering skeptic to AI advocate, explaining how Itential securely connects AI agents to infrastructure with enterprise-grade governance and traceability. We dive... Read more »
In this sponsored episode recorded live at AutoCon 4 in Austin, we sit down with Peter Sprygada, Chief Architect at Itential, to discuss Itential’s on-stage announcement of FlowAI. Peter shares his journey from network engineering skeptic to AI advocate, explaining how Itential securely connects AI agents to infrastructure with enterprise-grade governance and traceability. We dive... Read more »
Recorded live at AutoCon4, William Collins and Eyvonne Sharp join forces with John Capobianco for some in the moment thoughts and reflections on the AutoCon experience – from the in-person connections to the workshops to the stage presentations. John gives us the inside story on his very own workshop and the latest version releases in... Read more »
Recorded live at AutoCon4, William Collins and Eyvonne Sharp join forces with John Capobianco for some in the moment thoughts and reflections on the AutoCon experience – from the in-person connections to the workshops to the stage presentations. John gives us the inside story on his very own workshop and the latest version releases in... Read more »
Today we delve into the tech expertise deficit and why technical depth and decades of doing the work matter more than social media followers and content creation hype. Our guest is Russ White, engineer, author, teacher, and certification developer. We begin with current events in AI, and then investigate the differences between career and influence... Read more »
When Cloudbeds faced a post-sales organization at 120% capacity, no budget, and declining efficiency, Colin Slade chose to rebuild the operation through AI. Within nine months, his four-person AIOps team deployed more than 150 workflows and agents, automating 75% of repetitive work and reclaiming 7,000 hours every month.This episode details how Colin turned a resource-starved customer success organization into an AI-driven engine. It explores the early missteps, the shift from overengineering to small, quick wins, and how incremental adoption evolved into company-wide transformation.A practical study in applied AI, organizational change, and measurable outcomes—showing how constraint, not abundance, can drive real innovation.Timestamps0:00 – Preview & Introduction0:57 – Meet Colin Slade and the Situation at Cloudbeds9:25 – Mitigating Team Fears Around AI Replacing Jobs13:13 – The Stepwise Approach to Implementing AI19:50 – Scaling Securely: Working with IT, Risk-Taking, and Adoption24:00 – Roles and Team Structure for Effective AI Operations33:10 – Documentation as a Hidden Bottleneck39:45 – Build vs. Buy: Why Cloudbeds Built In-House42:20 – The Impact and a Culture of Fearless ExperimentationWhat You'll Learn* How to rebuild a post-sales org around AI without additional headcount* The step-by-step approach to deploying 150+ workflows in under a year* How to identify and structure AI roles: visionary, operators, knowledge masters, and project leads* The cultural and psychological levers for AI adoption* How to optimize documentation for AI readability (and boost SEO at the same time)* The measurable impact of AI on cost savings, efficiency, and morale---Check out the Key Takeaways & Transcripts: https://www.gainsight.com/presents/series/unchurned/---Where to Find Colin:LinkedIn: https://www.linkedin.com/in/colinslade/Where to Find Josh: LinkedIn: https://www.linkedin.com/in/jschachter/---Resources: n8n – https://n8n.io/Forethought – https://forethought.ai/Google AI Studio – https://aistudio.google.com/Anthropic Claude – https://claude.ai/Gemini – https://gemini.google.com/appLovable – https://lovable.dev/Pinecone – https://www.pinecone.io/Snowflake – https://www.snowflake.com/en/Zendesk – https://www.zendesk.nl/Salesforce – https://www.salesforce.com/Slack – https://slack.com/
Today we are joined by Dario Pasquini, Principal Researcher at RSAC, sharing the team's work on WhenAIOpsBecome “AI Oops”: Subverting LLM-driven IT Operations via Telemetry Manipulation. A first-of-its-kind security analysis showing that LLM-driven AIOps agents can be tricked by manipulated telemetry, turning automation itself into a new attack vector. The researchers introduce AIOpsDoom, an automated reconnaissance + fuzzing + LLM-driven telemetry-injection attack that performs “adversarial reward-hacking” to coerce agents into harmful remediations—even without prior knowledge of the target and even against some prompt-defense tools. They also present AIOpsShield, a telemetry-sanitization defense that reliably blocks these attacks without harming normal agent performance, underscoring the urgent need for security-aware AIOps design. The research can be found here: When AIOps Become “AI Oops”: Subverting LLM-driven IT Operations via Telemetry Manipulation Learn more about your ad choices. Visit megaphone.fm/adchoices