Platform Engineering Playbook Podcast

The Platform Engineering Playbook Podcast is where AI meets open-source infrastructure knowledge—and you're part of the editorial process. Every episode is researched, scripted, and produced with AI, then reviewed by the community and published on GitHub for anyone to improve. Facing tool sprawl across 130+ platforms? Justifying PaaS costs to your CFO? Navigating the Shadow AI crisis hitting 85% of organizations? We tackle the messy realities of platform engineering that most content avoids, delivering data-backed insights and decision frameworks you can use Monday morning. Built for senior engineers, SREs, and DevOps practitioners with 5+ years in production, we dissect cloud economics, AI governance, infrastructure trade-offs, and career strategy—with the receipts to back it up. Think we got something wrong? Have better data? Open a pull request at platformengineeringplaybook.com. This is infrastructure podcasting as a living document, where the community keeps us honest and the content gets better with every contribution.

Read the playbook at https://platformengineeringplaybook.com

Episodes

Jan 23, 2026

16 min

**Will 70% of DevOps engineers disappear in the next 5 years?** That's the bold prediction kicking off today's deep dive into the massive career shift happening in tech right now.
In this episode of Platform Engineering Playbook, we explore the critical transition from DevOps to Platform Engineering and what it means for your career survival. You'll discover why traditional DevOps roles are evolving, how companies like Spotify are leading this transformation, and the concrete roadmap you need to navigate this shift successfully.
**What You'll Learn:**• Why the DevOps-to-Platform Engineering transition is inevitable• Real-world examples from industry leaders like Spotify's Backstage platform• A practical career roadmap for making the transition• Breaking news: Railway's $100M funding to challenge AWS with AI-native infrastructure• GitHub Actions' new 1 vCPU Linux runner and what it means for CI/CD• The AI slop problem plaguing Kubernetes communities
**Timestamps:**0:00 Cold Open - The 70% Prediction2:15 Industry News Roundup8:30 Deep Dive: DevOps Career Evolution15:45 Platform Engineering Success Stories22:10 Your Career Transition Roadmap28:30 Wrap-up & Key Takeaways
Whether you're a DevOps engineer feeling uncertain about the future or a platform engineering leader building your team, this episode provides the insights and actionable strategies you need.
**Sources & References:**- From DevOps to Platform Engineer: https://platformengineering.org/blog/from-devops-to-platform-engineering- Railway secures $100 million: https://venturebeat.com/infrastructure/railway-secures-usd100-million-to-challenge-aws-with-ai-native-cloud- GitHub Actions 1 vCPU runner: https://github.blog/changelog/2026-01-22-1-vcpu-linux-runner-now-generally-available-in-github-actions- r/kubernetes AI discussion: https://www.reddit.com/r/kubernetes/comments/1qiezxc/rkubernetes_over_taken_with_ai_slop_projects/- OpenStack & OpenShift monitoring: https://developers.redhat.com/articles/2026/01/22/monitoring-openstack-and-openshift-together
#PlatformEngineering #DevOps #CloudNative #Kubernetes

Jan 23, 2026

16 min

Jan 22, 2026

17 min

The monitoring tool everyone trusts is actually blind to 40% of your infrastructure failures—and the vendor knows it. Are you using an industry standard that misses almost half of all incidents? In this episode, we unravel the mystery of infrastructure monitoring tools and why your choice could be costing you dearly.
As platform engineering teams grapple with an overwhelming array of options—from battle-tested open source tools to shiny SaaS platforms—the stakes have never been higher. The shift in focus from simple server monitoring to comprehensive observability is crucial for modern development.
🔑 What you’ll learn in this episode:- The shocking truth about popular monitoring tools that leave critical gaps in your observability.- Key indicators that signal it’s time to consider paid solutions for your monitoring needs.- A strategic playbook for evaluating vendors without falling into the lock-in trap.- Real-world examples of how companies manage their monitoring expenses, including a mid-sized SaaS company facing an $85,000 monthly bill.
Don't miss out on understanding how to choose the right tools that catch what others miss. Tune in as we dive deep into the world of infrastructure monitoring and equip you with the insights you need to make informed decisions. 
[Timestamps below]
⏱️ TIMESTAMPS:00:00:03 - Cold Open00:00:30 - Intro00:00:49 - Deep Dive - Act 1: The Setup00:05:32 - Deep Dive - Act 2: The Analysis00:09:44 - Deep Dive - Act 3: Takeaways00:14:21 - News00:16:57 - Outro
📌 IN THIS EPISODE:
---🎙️ Platform Engineering Playbook🔗 https://platformengineeringplaybook.com📧 Subscribe for weekly platform engineering insights
#PlatformEngineering #DevOps #Kubernetes #CloudNative #SRE #Podcast

Jan 22, 2026

17 min

Jan 21, 2026

11 min

**73% of engineering teams are drowning in technical debt because of their CI/CD pipelines. Not despite them—because of them.**
Are your automation tools secretly sabotaging your codebase? Today's Platform Engineering Playbook dives deep into the hidden ways CI/CD pipelines create technical debt and reveals practical strategies to break the cycle.
**What You'll Learn:**• Why inheritance beats copying in platform design• Docker's new hardened images for bulletproof container security• How OpenTelemetry's log deduplication processor can slash your log volume• Critical vulnerabilities in Chainlit and Cloudflare you need to patch NOW• Actionable steps to audit and optimize your CI/CD debt
**Episode Chapters:**0:00 Cold Open - The 73% Problem1:30 Platform Engineering News Roundup8:45 Deep Dive: CI/CD Technical Debt Crisis15:20 The Theory vs Reality of Inheritance22:10 Practical Solutions for Your Organization28:30 Security Alert Roundup
Whether you're a platform engineer drowning in legacy pipelines or a team lead trying to prevent future debt, this episode gives you the frameworks and tools to build sustainable automation that actually reduces complexity instead of adding to it.
**Sources & References:**• CI/CD Technical Debt Analysis: https://thenewstack.io/are-your-ci-cd-pipelines-accidentally-increasing-technical-debt/• Docker Hardened Images: https://feeds.dzone.com/link/23568/17258698/docker-hardened-images-container-security• OpenTelemetry Log Deduplication: https://opentelemetry.io/blog/2026/log-deduplication-processor/• Chainlit Security Advisory: https://www.securityweek.com/chainlit-vulnerabilities-may-leak-sensitive-information/• Cloudflare Zero-Day Alert: https://cybersecuritynews.com/cloudflare-zero-day-vulnerability/
#PlatformEngineering #DevOps #CloudNative #Kubernetes

Jan 21, 2026

11 min

Jan 20, 2026

17 min

**Are major tech companies secretly abandoning Kubernetes certifications?** What we discovered about the future of K8s learning will change how you approach platform engineering in 2026.
In today's Platform Engineering Playbook, we uncover why traditional Kubernetes education is becoming obsolete and what platform teams are doing instead. Plus, breaking news that could revolutionize your infrastructure stack.
**What You'll Learn:**• Why the volume of Kubernetes resources reveals a hidden shift in the industry• Microsoft's game-changing Azure Functions announcement for Model Context Protocol servers• How Pinterest's Moka is rewriting big data processing rules with Kubernetes• Practical strategies for platform engineers navigating the evolving K8s landscape• Critical vulnerability insights from Cloudflare's ACME validation logic
Whether you're leading a platform team or building cloud-native infrastructure, this episode delivers actionable insights you can implement immediately.
**Sources & References:**- Microsoft Azure Functions Model Context Protocol announcement- Cloudflare ACME vulnerability mitigation report- Pinterest Moka Kubernetes big data processing case study
#PlatformEngineering #DevOps #CloudNative #Kubernetes #Azure #CloudSecurity #BigData #InfrastructureAsCode

Jan 20, 2026

17 min

Jan 19, 2026

16 min

"92% of European companies don’t trust US cloud providers with their data anymore. So, AWS just locked itself out of its own Euro Cloud! This shocking move raises critical questions about data sovereignty and compliance for businesses operating in Europe. 
In this episode, we dive deep into AWS's groundbreaking decision to create a completely isolated European cloud infrastructure, one that even Amazon employees can't access. Why would they cut off their own access, and what does this mean for your data strategy?
🔑 Learn about the implications of AWS's European Sovereign Cloud and how it represents a shift in data sovereignty.🔑 Discover the parent company structure AWS is using with local subsidiaries in Germany and what that means for compliance.🔑 Get actionable insights on navigating data classification and regulatory exposure in a post-CLOUD Act world.🔑 Understand how this decision impacts your compliance strategy if you store customer data in AWS.
This episode unpacks the complexities of data sovereignty and the geopolitical risks that come with using US cloud providers. Don't miss this crucial information that could change your approach to cloud infrastructure forever. 
---🎙️ Platform Engineering Playbook🔗 https://platformengineeringplaybook.com
---
 

Jan 19, 2026

16 min

Jan 18, 2026

15 min

Terraform’s biggest competitor just made a move that could redefine infrastructure-as-code in 2026.
Pulumi now runs Terraform and HCL natively—better than HashiCorp does. That’s not a migration tool, not a compatibility shim, but full native execution through the Pulumi engine, plus Terraform state hosted in Pulumi Cloud and financial credits to help teams exit existing HashiCorp contracts.
In this episode of the Platform Engineering Playbook Daily Podcast, we break down why this announcement is one of the most important platform engineering stories of the year—and what it actually means for SREs, platform teams, and infrastructure leaders.
We cover:- Why Pulumi supporting Terraform is not cooperation, but displacement  - How native HCL execution inside the Pulumi engine actually works  - What “polyglot infrastructure” means in practice for platform teams  - Pulumi Cloud as a Terraform Cloud replacement (state, RBAC, policy, AI)  - The real risks: compatibility gaps, lock-in concerns, and beta limitations  - A practical framework for deciding whether your organization should care  
We also cover today’s top platform engineering headlines:- How AI is shifting SRE from reactive firefighting to failure prevention  - Cloudflare’s observability redesign for large-scale configuration management  - RunPod’s unlikely rise to $120M ARR  - CoreWeave’s infrastructure challenges and why execution matters more than hype  
If you’re managing Terraform at scale, evaluating OpenTofu, building an internal developer platform, or navigating the post–HashiCorp license-change landscape, this episode will directly impact decisions you’ll be making over the next 6–12 months.
Subscribe for daily platform engineering analysis, deep dives, and practical insights you can actually apply.
Links and sources discussed are in the show notes.

Jan 18, 2026

15 min

Jan 17, 2026

13 min

Cloudflare acquires the Astro Technology Company, adding a 1M-downloads-per-week web framework to their edge platform. We analyze the strategic implications, what stays open source, and lessons about framework sustainability for platform engineering teams.
Key Topics:- Astro framework overview: islands architecture, framework-agnostic components, content-first approach- Why Cloudflare acquired Astro: Developer ecosystem capture, edge compute alignment, workerd integration- Open source sustainability: MIT license preserved, historical patterns (Gatsby, Remix)- What changes for platform teams: Framework evaluation criteria, portability concerns, exit strategies- News: AWS European Sovereign Cloud, Let's Encrypt 6-day certs, Datadog LLM Observability
Duration: 13 minutes
Subscribe and share with colleagues who'd find this valuable!
#PlatformEngineering #Astro #Cloudflare #WebFramework #EdgeComputing #OpenSource #CloudflareWorkers #DevOps #SRE

Jan 17, 2026

13 min

Jan 16, 2026

11 min

ScyllaDB just launched X Cloud with claims of double the performance at half the cost compared to DynamoDB. This episode breaks down the technical architecture behind their tablet-based approach, how they're achieving 80% data compression on ARM Graviton4 instances, and when this actually makes sense for platform engineering teams running high-throughput workloads.
Key Topics:- ScyllaDB X Cloud tablet-based architecture (5GB chunks) vs traditional consistent hashing- Claims of 6x performance improvement with 50% cost reduction vs DynamoDB- 80% compression on ARM Graviton4 instances, 25x faster data streaming- High-throughput workload targets: Discord, Disney, Starbucks use cases- News: TerraFormer AI IaC generation, AWS supply chain vulnerability, ML pipeline security
Duration: 11 minutes
Subscribe and share with colleagues who'd find this valuable!
#PlatformEngineering #ScyllaDB #DynamoDB #NoSQL #CloudNative #DatabaseArchitecture #AWS #DevOps #SRE

Jan 16, 2026

11 min

Jan 15, 2026

16 min

Your Linux servers aren't just running containers anymore—they're hosting invisible tenants that security teams can't even detect.
In this episode, we deep dive into VoidLink, the new cloud-native malware framework that Check Point Research just uncovered. This isn't your typical malware that got retrofitted for the cloud—this thing was born in the cloud, designed from the ground up to evade every detection tool in your security stack.
We explore:• How VoidLink achieves its terrifying persistence in cloud environments• Why every major cloud provider is vulnerable to this new threat class• eBPF-based rootkits and kernel-level persistence techniques• Why traditional security tools fail against cloud-native threats• How VoidLink learns and adapts to your environment over time• Defense-in-depth strategies for cloud-native infrastructure
Key takeaway: VoidLink represents a new generation of threats built specifically for the cloud. Platform teams must evolve their security posture to include runtime detection, eBPF observability, and defense-in-depth strategies.
---
Platform Engineering Podcast provides deep dives into infrastructure, DevOps, and cloud-native security. New episodes weekly.
Subscribe: https://platformengineering.org/podcast

Jan 15, 2026

16 min

Jan 14, 2026

14 min

By 2025, 90% of new enterprise applications will be AI-powered and cloud-native. This episode explores the symbiotic relationship between AI and Kubernetes - where AI isn't just another workload, but is fundamentally transforming how we build and operate cloud native platforms. We cover real-world examples like Netflix's predictive scaling achieving 92% accuracy, the emergence of AI-driven observability platforms, and why platform engineers need to evolve from infrastructure operators to AI-infrastructure orchestrators.
In this episode:- AI transforming the Kubernetes control plane with predictive scheduling- Netflix's AI-driven traffic management: 92% prediction accuracy, 35% resource reduction- AI-native observability: anomaly detection on metric relationships, not just metrics- GPU orchestration: NVIDIA GPU Operator achieving 80%+ utilization vs 30-40% baseline- Edge AI patterns: federated learning, model distillation, intermittent connectivity- Skills evolution: Understanding AI workload characteristics without becoming ML experts- News: Red Hat connects AI to Istio via Kiali MCP Server, AWS CloudWatch adds Apache Iceberg support
Perfect for senior platform engineers, SREs, DevOps engineers looking to understand the convergence of AI and cloud native technologies.
New episodes every week. Subscribe wherever you listen to stay current on platform engineering.
Episode URL: https://platformengineeringplaybook.com/podcasts/00090-ai-cloud-native-symbiosis
Duration: 15 minutes
Host: Alex and Jordan
Category: TechnologySubcategory: Software How-To
Keywords: AI, cloud native, Kubernetes, symbiosis, intelligent infrastructure, platform engineering, GPU orchestration, predictive scaling, observability, machine learning, Netflix, edge AI, federated learning

Jan 14, 2026

14 min

Copyright 2025 All rights reserved.

Podcast Powered By Podbean

Version: 20241125