103 afleveringen
- Karpenter can reduce Kubernetes infrastructure costs, but aggressive node consolidation can also expose workloads that lack disruption safeguards.
Ahmad Asmar explains how Zencity uses Kyverno to automatically generate Pod Disruption Budgets, while accounting for existing PDBs, percentage-based availability targets, single-replica workloads, and environment-specific policies.
In this interview:
How Karpenter consolidation changes the availability risks of cluster operations
Why Kyverno's generated policies can provide safer defaults than manual enforcement
How to handle duplicate PDBs, scaling workloads, and single-replica edge cases
How aggregated ClusterRoles keep custom permissions separate from Helm-managed resources
Sponsor
This episode is sponsored by LearnKube. Download the free book, The Technical Guide to Kubernetes Rightsizing, to understand what Prometheus and Grafana cannot tell you about safely reducing requests and limits.
More info
Find all the links and info for this episode here: https://ku.bz/xrlPJg54D
Interested in sponsoring an episode? Learn more. - An unmaintained identity component can remain invisible until a routine Kubernetes upgrade turns it into an incident.
Fabián Sellés Rosa, Platform Engineer and Runtime Tech Lead at Adevinta, explains how his team moved from KIAM to EKS Pod Identities without discarding the security boundaries and application interface that their internal platform depended on.
In this interview:
Why KIAM became urgent to replace after years of stable operation
How Crossplane, a custom controller, and KRO with ACK compared against the team's criteria
Why managed EKS Capabilities reduced toil but introduced observability and rollout trade-offs
How Kyverno preserved namespace-level authorization for IAM roles
Sponsor
This episode is sponsored by LearnKube. Download the free book, The Technical Guide to Kubernetes Rightsizing, to understand what Prometheus and Grafana cannot tell you about safely reducing requests and limits.
More info
Find all the links and info for this episode here: https://ku.bz/R_06hwnCn
Interested in sponsoring an episode? Learn more. - GPU inference throughput depends on more than accelerator generation or count.
Memory bandwidth, model parallelism, cache configuration, and the load generator itself all influence measured throughput.
Federico Iezzi, Customer Engineer at Google Cloud, explains how his team achieved 1 million output tokens per second using Qwen 3.5 27B, vLLM, GKE Autopilot, and NVIDIA B200 GPUs.
The discussion covers:
Why memory bandwidth limits decode performance
How Federico chose between tensor and data parallelism
What changed after enabling multi-token prediction and reducing the KV cache footprint with FP8 quantization.
Sponsor
This episode is sponsored by LearnKube. Download the free book, The Technical Guide to Kubernetes Rightsizing, to understand what Prometheus and Grafana cannot tell you about safely reducing requests and limits.
More info
Find all the links and info for this episode here: https://ku.bz/1xD9Md0mb
Interested in sponsoring an episode? Learn more. - Forced platform migrations are usually treated as something to survive. At Scout24, a mandatory OS migration became an opportunity to rethink Kubernetes autoscaling, node provisioning, and infrastructure efficiency.
John Ford explains how Scout24 moved its EKS-based Infinity platform from a polling autoscaler and over-provisioned capacity to Karpenter and Bottlerocket. The result was faster node startup, a safer migration path, and about a 30% infrastructure reduction without major downtime.
In this interview:
Why two-minute node provisioning forced a 25% capacity buffer
How Karpenter made the Bottlerocket migration safer
What broke around EC2 metadata, AWS SDKs, and cgroups
How the new foundation enables Spot, ARM, and GPU workloads
Sponsor
This episode is sponsored by LearnKube — get started on your Kubernetes journey through comprehensive online, in-person or remote training.
More info
Find all the links and info for this episode here: https://ku.bz/DdmVC2_7v
Interested in sponsoring an episode? Learn more. - Most teams scale Kubernetes by thinking about pods and nodes. At Render, Brian Stack ran into a different dimension: hundreds of thousands of namespaces per cluster, multiplied across DaemonSets that list-watch every namespace.
Brian explains how Render traced the issue through Calico and Vector, worked with upstream maintainers, and turned memory profiling into operational wins: lower node costs, lighter API-server load, and faster rollouts.
In this interview:
Why namespaces can become a hidden scaling bottleneck
How DaemonSets multiply memory and control-plane pressure
How profiling, staging clusters, and upstream collaboration freed 7 TiB
Why pushing from an 80% fix to a complete fix can make teams faster
Sponsor
This episode is sponsored by LearnKube — get started on your Kubernetes journey through comprehensive online, in-person or remote training.
More info
Find all the links and info for this episode here: https://ku.bz/0mrvCsXrV
Interested in sponsoring an episode? Learn more.
Meer Technologie podcasts
Trending Technologie -podcasts
Over KubeFM
Discover all the great things happening in the world of Kubernetes, learn (controversial) opinions from the experts and explore the successes (and failures) of running Kubernetes at scale.
Podcast websiteLuister naar KubeFM, De Technoloog | BNR en vele andere podcasts van over de hele wereld met de radio.net-app

Ontvang de gratis radio.net app
- Zenders en podcasts om te bookmarken
- Streamen via Wi-Fi of Bluetooth
- Ondersteunt Carplay & Android Auto
- Veel andere app-functies
Ontvang de gratis radio.net app
- Zenders en podcasts om te bookmarken
- Streamen via Wi-Fi of Bluetooth
- Ondersteunt Carplay & Android Auto
- Veel andere app-functies


KubeFM
Scan de code,
download de app,
luisteren.
download de app,
luisteren.
































