Platform Engineering Playbook Podcast

The Platform Engineering Playbook Podcast is where AI meets open-source infrastructure knowledge—and you're part of the editorial process. Every episode is researched, scripted, and produced with AI, then reviewed by the community and published on GitHub for anyone to improve. Facing tool sprawl across 130+ platforms? Justifying PaaS costs to your CFO? Navigating the Shadow AI crisis hitting 85% of organizations? We tackle the messy realities of platform engineering that most content avoids, delivering data-backed insights and decision frameworks you can use Monday morning. Built for senior engineers, SREs, and DevOps practitioners with 5+ years in production, we dissect cloud economics, AI governance, infrastructure trade-offs, and career strategy—with the receipts to back it up. Think we got something wrong? Have better data? Open a pull request at platformengineeringplaybook.com. This is infrastructure podcasting as a living document, where the community keeps us honest and the content gets better with every contribution.

Read the playbook at https://platformengineeringplaybook.com

Episodes

Mar 30, 2026

18 min

**Is your company bleeding $43,800 annually on hidden Kubernetes costs?** Most platform teams have no idea they're paying this "invisible tax" – but the smartest engineers are already eliminating it.
In today's Platform Engineering Playbook, we expose the shocking truth about Kubernetes cost isolation and dive deep into virtual clusters as the solution. Plus, we break down the biggest platform engineering news shaking up the industry right now.
**What You'll Learn:**• How to identify and eliminate the $43,800 hidden Kubernetes tax• Virtual clusters: Running Kubernetes inside Kubernetes (and why it works)• Practical implementation strategies for 47+ development clusters• Microsoft's new Azure Copilot Migration Agent breakdown• Kubescape 4.0's game-changing runtime security features• Why WebAssembly is crushing containers at the edge
**Timestamps:**0:00 Cold Open: The $43,800 Hidden Tax2:15 Today's Platform Engineering News8:30 Deep Dive: Virtual Clusters Explained15:45 Implementation Strategy & Analysis
Whether you're managing a handful of clusters or hundreds, this episode will save you serious money and headaches. Perfect for platform engineers, SREs, and anyone tired of Kubernetes cost surprises.
**Sources & References:**• Virtual Clusters Cost Analysis: https://thenewstack.io/virtual-clusters-kubernetes-cost-isolation/• Azure Copilot Migration Agent: https://www.infoq.com/news/2026/03/azure-copilot-migration-agent/• Kubescape 4.0 Release: https://www.infoq.com/news/2026/03/kubescape-40/• Teleport AI Security Report: https://www.infoq.com/news/2026/03/teleport-ai-report/• WebAssembly vs Containers: https://thenewstack.io/webassembly-component-model-future/• Platform Engineer AI Takes: https://alienchow.dev/post/ai_takeaways_mar_2026/
#PlatformEngineering #DevOps #CloudNative #Kubernetes

Mar 30, 2026

18 min

Mar 27, 2026

18 min

**Why do 87% of AI models never reach production? It's not the AI - it's the infrastructure underneath.**
In this deep dive episode of Platform Engineering Playbook, we tackle the critical challenge of building cloud-native platforms that can actually support AI workloads at scale. While everyone's talking about model performance, the real bottleneck is happening at the infrastructure layer.
**What You'll Learn:**• Why traditional Kubernetes setups fail with AI workloads• How to architect platforms for GPU-intensive applications• Real-world strategies from teams successfully running AI in production• Practical steps to prepare your platform for the AI transformation
**Episode Breakdown:**00:00 Cold Open: The AI production problem02:30 Today's platform engineering news roundup08:45 Deep Dive Act 1: AI workloads are breaking Kubernetes18:20 Deep Dive Act 2: How successful teams handle AI infrastructure
**Today's News:**• CNCF adds 21 new silver members focused on AI infrastructure• Critical Grafana security vulnerabilities (CVE-2026-27876 & CVE-2026-27880)• GitHub Actions 2026 security roadmap• AWS launches new AI development tools• Aurora PostgreSQL joins AWS Free Tier
Perfect for platform engineers, SREs, and infrastructure teams navigating the AI infrastructure challenge.
**Sources & References:**- https://www.cncf.io/blog/2026/03/26/the-platform-under-the-model-how-cloud-native-powers-ai-engineering-in-production/- https://www.cncf.io/announcements/2026/03/25/cncf-welcomes-21-new-silver-members-as-global-demand-surges-for-observability-ai-and-secure-cloud-native-infrastructure/- https://grafana.com/blog/grafana-security-release-critical-and-high-severity-security-fixes-for-cve-2026-27876-and-cve-2026-27880/- https://github.blog/news-insights/product-news/whats-coming-to-our-github-actions-2026-security-roadmap/- https://aws.amazon.com/about-aws/whats-new/2026/03/agent-plugin-aws-serverless/- https://aws.amazon.com/about-aws/whats-new/2026/03/amazon-aurora-postgresql-aws-free-tier/
#PlatformEngineering #DevOps #CloudNative #Kubernetes

Mar 27, 2026

18 min

Mar 26, 2026

19 min

**What happens when AI agents in your Kubernetes cluster start making their own scaling decisions without proper guardrails?**
In this episode of Platform Engineering Playbook, we dive deep into the emerging world of cloud native agentic standards and why they're becoming mission-critical for modern infrastructure. As AI agents become more autonomous in managing our clusters, the need for standardized communication protocols has never been more urgent.
**What You'll Learn:**• How cloud native agentic standards are reshaping Kubernetes operations• Real-world scenarios where unsupervised AI agents can wreak havoc on your infrastructure• Practical strategies for implementing these standards in existing systems• Latest CNCF developments including 21 new silver members and expanded AI inference capabilities• Istio's new AI-era features and AWS Load Balancer Controller GA release
**Episode Chapters:**0:00 - Cold Open: The AI Agent Standardization Crisis2:15 - Industry News Roundup8:30 - Deep Dive: Cloud Native Agentic Standards15:45 - The Problem: When AI Agents Go Rogue
Whether you're already running AI workloads on Kubernetes or planning your first implementation, this episode provides the frameworks and insights you need to avoid costly mistakes while building resilient, agent-aware infrastructure.
**Sources & References:**- Cloud native agentic standards: https://www.cncf.io/blog/2026/03/23/cloud-native-agentic-standards/- CNCF Celebrates Innovators: https://www.cncf.io/announcements/2026/03/25/cncf-celebrates-innovators-advancing-cloud-native-at-kubecon-cloudnativecon-europe/- CNCF AI Inference Expansion: https://cloudnativenow.com/features/cncf-expands-efforts-to-run-ai-inference-workloads-on-kubernetes-clusters/- Istio AI Era Features: https://www.cncf.io/announcements/2026/03/25/istio-brings-future-ready-service-mesh-to-the-ai-era-with-new-ambient-multicluster-gateway-api-inference-extension-and-more/- CNCF New Silver Members: https://www.cncf.io/announcements/2026/03/25/cncf-welcomes-21-new-silver-members-as-global-demand-surges-for-observability-ai-and-secure-cloud-native-infrastructure/- AWS Gateway API GA: https://www.infoq.com/news/2026/03/aws-gateway-api-ga/
#PlatformEngineering #DevOps #CloudNative #Kubernetes

Mar 26, 2026

19 min

Mar 25, 2026

19 min

What happens when hundreds of AI agents start running in your Kubernetes cluster but can't communicate with each other? By 2026, this isn't a hypothetical problem—it's the reality platform engineers are facing right now.
In this episode of Platform Engineering Playbook, we dive deep into the CNCF's new cloud-native agentic standards and what they mean for your infrastructure. We'll break down why these standards exist, how they solve critical interoperability challenges, and most importantly—what you need to implement today to stay ahead.
**What You'll Learn:**• How to prepare your platform for the AI agent explosion• CNCF's new agentic workflow validation requirements• Why IBM, Red Hat, and Google just donated their LLM inference blueprint• The latest enterprise networking developments with multi-cloud SD-WAN• How the cloud-native community reached 19.9 million developers
**Episode Timestamps:**0:00 Cold Open - The AI Agent Problem2:30 Platform Engineering News Roundup8:15 Deep Dive: Cloud Native Agentic Standards15:45 Analysis: What These Standards Actually Mean
Whether you're managing existing Kubernetes workloads or planning your AI strategy, this episode gives you the practical insights to build resilient, future-ready platforms.
**Sources & References:**• CNCF Cloud Native Agentic Standards: https://www.cncf.io/blog/2026/03/23/cloud-native-agentic-standards/• Colt Multi-Cloud SD-WAN Launch: https://totaltele.com/colt-targets-enterprise-digital-transformation-with-multi-cloud-sd-wan-launch/• CNCF Developer Community Report: https://cloudnativenow.com/kubecon-cloudnativecon-europe-2026/cncf-and-slashdata-report-finds-cloud-native-developer-community-has-reached-19-9-million/• Kubernetes LLM Inference Blueprint: https://thenewstack.io/llm-d-cncf-kubernetes-inference/• CNCF AI Platform Certifications: https://cloudnativenow.com/kubecon-cloudnativecon-europe-2026/cncf-nearly-doubles-certified-kubernetes-ai-platforms-adds-agentic-workflow-validation/
#PlatformEngineering #DevOps #CloudNative #Kubernetes

Mar 25, 2026

19 min

Mar 24, 2026

20 min

**Are you burning through your LLM budget with zero visibility into why?** You're not alone - 73% of production deployments are facing this exact problem right now.
In today's Platform Engineering Playbook, we tackle the monitoring crisis plaguing AI infrastructure and break down five game-changing developments reshaping how we deploy and secure production systems.
**🎯 What You'll Learn:**• How to implement proper LLM observability using Grafana Cloud, OpenLIT, and OpenTelemetry• Step-by-step rollout strategies that won't break your production environment• Why Teleport's new Beams could revolutionize AI agent security in your infrastructure
**📺 Chapters:**0:00 - Cold Open: The LLM Budget Crisis2:15 - Today's Platform Engineering News8:30 - Deep Dive: LLM Monitoring in Production15:45 - Implementation Walkthrough
**🔥 This Week's News:**• Teleport Beams: Trusted runtimes for AI agents• Metal3 joins CNCF incubation at KubeCon Europe 2026• Cloudflare's Gen 13 servers deliver 2x edge compute performance• New AI-compatible certification frameworks• Pi-Hole deployment strategies for network-wide ad blocking
Perfect for platform engineers, DevOps teams, and infrastructure leaders dealing with AI workloads in production.
**Sources & References:**• https://grafana.com/blog/ai-observability-llms-in-production/• https://cloudnativenow.com/kubecon-cloudnativecon-europe-2026/teleport-launches-beams-to-provide-trusted-runtimes-for-ai-agents-in-production-infrastructure/• https://www.cncf.io/blog/2026/03/23/metal3-at-kubecon-cloudnativecon-europe-2026-meet-the-cncfs-freshly-incubated-bare-metal-project/• https://blog.cloudflare.com/gen13-launch/• https://letsdatascience.com/news/made-in-usa-introduces-ai-compatible-certification-framework-b0b018ac• https://thenewstack.io/pihole-docker-network-adblocking/
#PlatformEngineering #DevOps #CloudNative #Kubernetes

Mar 24, 2026

20 min

Mar 23, 2026

14 min

**What if 94% of Helm chart vulnerabilities could be prevented with one unexpected technology?**
Today's Platform Engineering Playbook dives deep into the surprising intersection of WebAssembly and Kubernetes security, plus breaking news that every platform engineer needs to know.
**What You'll Learn:**• How WebAssembly is revolutionizing Helm chart security (spoiler: it's not replacing Kubernetes)• Why Trivy is under attack again and what it means for your CI/CD pipelines• Critical findings from a security audit of 22,511 AI coding skills• Whether AI will evolve code or make it extinct• GrapheneOS's bold stance against age verification laws
**Episode Breakdown:**0:00 - Cold Open: The 94% Helm vulnerability stat that will shock you2:30 - Today's platform engineering headlines5:15 - Deep Dive: WebAssembly + Kubernetes security analysis15:45 - Practical implementation strategies for platform teams
Perfect for platform engineers, DevOps professionals, and anyone building resilient cloud-native infrastructure. Get the technical depth you need with practical insights you can implement immediately.
**Sources & References:**• WebAssembly & Helm Security: https://thenewstack.io/helm-webassembly-kubernetes-security/• Trivy Attack Analysis: https://socket.dev/blog/trivy-under-attack-again-github-actions-compromise• AI Coding Skills Audit: https://thenewstack.io/ai-agent-skills-security/• AI Programming Future: https://thenewstack.io/ai-programming-languages-future/• GrapheneOS News: https://www.tomshardware.com/software/operating-systems/grapheneos-refuses-to-comply-with-age-verification-laws
#PlatformEngineering #DevOps #CloudNative #Kubernetes

Mar 23, 2026

14 min

Mar 20, 2026

23 min

**87% of AI workloads are sitting idle on GPUs right now** - yet companies keep buying more hardware. What if the problem isn't capacity, but how we're running AI on Kubernetes?
In today's Platform Engineering Playbook, we tackle the massive inefficiencies plaguing AI infrastructure at scale. You'll discover why traditional Kubernetes patterns break down with AI workloads, what's actually happening under the hood when you try to serve ML models in production, and concrete strategies to fix GPU utilization without throwing more money at the problem.
**What You'll Learn:**• Why current Kubernetes-native AI patterns are failing at scale• The hidden bottlenecks destroying your GPU efficiency • Runtime security developments from Grafana Labs and Miggo• Amazon ECR's new pull-through cache support for Chainguard• How to evolve from Kubernetes Gatekeeper to full-stack governance with OPA
**Timestamps:**0:00 Cold Open - The AI Infrastructure Crisis2:15 Today's Platform Engineering News8:30 Deep Dive: Kubernetes + AI at Scale15:45 Under the Hood Analysis22:10 Actionable Takeaways
Whether you're scaling AI workloads or just trying to understand why your GPU bills keep growing while performance stays flat, this episode gives you the platform engineering perspective you need.
**Sources & References:**• Building Kubernetes-native AI infrastructure: https://thenewstack.io/kubernetes-native-ai-infrastructure/• Grafana Cloud and Miggo runtime protection: https://grafana.com/blog/grafana-cloud-and-miggo-for-runtime-protection/• Amazon ECR Chainguard support: https://aws.amazon.com/about-aws/whats-new/2026/03/amazon-ecr-pull-through-cache-chainguard/• AWS Cloud 20 years retrospective: https://aws.amazon.com/blogs/aws/20-years-in-the-aws-cloud-how-time-flies/• LLM Compressor v0.10: https://developers.redhat.com/articles/2026/03/18/llm-compressor-010-faster-compression-distributed-gptq• Kubernetes Gatekeeper to OPA governance: https://www.pulumi.com/blog/kubernetes-gatekeeper-full-stack-governance-opa/
#PlatformEngineering #DevOps #CloudNative #Kubernetes

Mar 20, 2026

23 min

Mar 19, 2026

18 min

**Are 73% of Kubernetes clusters really flying blind?** According to recent industry reports, most K8s deployments are drowning in meaningless metrics while missing the signals that actually matter for performance and cost optimization.
In today's Platform Engineering Playbook, we tackle the Kubernetes observability crisis head-on. You'll discover why traditional monitoring approaches are failing platform teams and learn actionable strategies to build metrics that drive real business value.
**What You'll Learn:**• Why most K8s metrics collection strategies are fundamentally broken• How to identify and implement performance indicators that actually matter• Practical frameworks for establishing effective observability in your clusters• Real-world approaches to turning metrics into cost savings and performance gains
**Episode Breakdown:**00:00 - Cold Open: The K8s Observability Crisis02:30 - Industry News Roundup08:45 - Deep Dive: Fixing Kubernetes Metrics (Part 1)
**Today's News:** Container security innovations from Chainguard, Grafana's new cost optimization tools, custom metrics scaling strategies, and the latest observability trends including AI integration challenges.
Perfect for platform engineers, DevOps teams, and engineering leaders looking to move beyond vanity metrics to actionable observability.
**Sources & References:**- CNCF Kubernetes Metrics Best Practices: https://www.cncf.io/blog/2026/03/18/understanding-kubernetes-metrics-best-practices-for-effective-monitoring/- Grafana Cost Optimization Guide: https://grafana.com/blog/from-signals-to-savings-optimizing-cloud-costs-with-grafana-assistant-and-mcp-servers/- Chainguard Container Security Analysis: https://thenewstack.io/chainguard-os-packages-containers/- Datadog Custom Metrics Scaling: https://www.datadoghq.com/blog/autoscaling-custom-metrics/- Grafana Observability Standards Report: https://grafana.com/blog/observability-survey-OSS-open-standards-2026/- AI in Observability Survey: https://grafana.com/blog/observability-survey-AI-2026/
#PlatformEngineering #DevOps #CloudNative #Kubernetes

Mar 19, 2026

18 min

Mar 18, 2026

18 min

**87% of enterprise AI deployments have a critical security vulnerability that red teams aren't even testing for.** Are you one of them?
In today's Platform Engineering Playbook, we expose the massive security hole plaguing enterprise AI systems and dive deep into prompt injection attacks that are slipping past traditional security measures. Plus, we cover the latest platform engineering news that's reshaping how enterprises build and deploy.
**What You'll Learn:**• The hidden AI security vulnerability affecting 9 out of 10 enterprise deployments• Step-by-step breakdown of how prompt injection attacks work in production• Actionable security strategies for platform engineers deploying AI agents• Microsoft's aggressive PostgreSQL push and what it means for your data strategy• Cloudflare's evolution from legacy architecture to modern SASE solutions
**Timestamps:**0:00 Cold Open - The 87% Problem1:30 Introduction3:00 Deep Dive: The AI Security Crisis8:45 How Prompt Injection Attacks Actually Work15:20 Platform Engineer Action Items
Whether you're currently deploying AI systems or planning your enterprise AI strategy, this episode delivers the security insights and platform engineering intelligence you need to stay ahead of emerging threats.
**Sources & References:**• AI Security Research: https://thenewstack.io/red-teaming-enterprise-ai-agents/• PostgreSQL on Azure: https://azure.microsoft.com/en-us/blog/from-legacy-to-leadership-how-postgresql-on-azure-powers-enterprise-agility-and-innovation/• Cloudflare SASE Evolution: https://blog.cloudflare.com/legacy-to-agile-sase/• AI Tooling Survey: https://newsletter.pragmaticengineer.com/i/189777574/2-most-used-ai-tools• Azure DevOps MCP Server: https://devblogs.microsoft.com/devops/azure-devops-remote-mcp-server-public-preview/
#PlatformEngineering #DevOps #CloudNative #Kubernetes

Mar 18, 2026

18 min

Mar 17, 2026

17 min

**Is your Kubernetes cluster blind to AI model poisoning attacks?** 73% of companies running AI workloads can't detect when their models are compromised - and traditional monitoring tools are completely useless against these threats.
In today's Platform Engineering Playbook, we dive deep into why AI workloads are breaking traditional Kubernetes observability strategies and what platform teams need to do about it. Plus, we cover the latest developments shaking up the cloud native ecosystem.
**What You'll Learn:**✅ Why traditional Kubernetes monitoring fails with AI workloads✅ How to detect AI model poisoning in production environments✅ Critical AWS security vulnerabilities affecting managed services✅ New authentication strategies for Kubernetes registry mirrors✅ Latest developments from the cloud native community
**Timestamps:**0:00 Cold Open - The AI observability crisis1:30 Today's Platform Engineering News8:45 Deep Dive: AI Workloads vs Traditional Monitoring15:20 The Real-World Impact on Autoscaling
Whether you're running AI workloads today or planning for tomorrow, this episode gives you the strategies and tools to maintain visibility and security in your Kubernetes environments.
**Sources & References:**- Why AI workloads are breaking traditional Kubernetes observability strategies: https://thenewstack.io/ai-kubernetes-observability-practices/- AWS Launches Managed Openclaw on Lightsail Amid Critical Security Vulnerabilities: https://www.infoq.com/news/2026/03/aws-lightsail-openclaw-security/?utm_campaign=infoq_content&utm_source=infoq&utm_medium=feed&utm_term=global- LLM Architecture Gallery: https://sebastianraschka.com/llm-architecture-gallery/- Cursor built a fleet of security agents to solve a familiar frustration: https://thenewstack.io/cursor-open-sources-security-agents/- Registry Mirror Authentication with Kubernetes Secrets: https://www.cncf.io/blog/2026/03/16/registry-mirror-authentication-with-kubernetes-secrets-2/- KubeCon + CloudNativeCon Europe 2026 Co-located Event Deep Dive: Open Sovereign Cloud Day: https://www.cncf.io/blog/2026/03/16/kubecon-cloudnativecon-europe-2026-co-located-event-deep-dive-open-sovereign-cloud-day/
#PlatformEngineering #DevOps #CloudNative #Kubernetes

Mar 17, 2026

17 min

Copyright 2025 All rights reserved.

Podcast Powered By Podbean

Version: 20241125