560 blogs tracked4,950 posts indexed

Kubernetes

We track 65 posts about Kubernetes from 22 engineering blogs. Most active: Learnk8s, Pulumi, William Kennedy. Latest post: Oct 6, 2026.

Raw tag behind this topic:kubernetes

Companies writing about Kubernetes

Recent posts

  • Rossoctl for agent discoverability, security, and observability on Kubernetes (opens on the source site)

    Red Hat ·

    Imagine deploying a pair of agents to Red Hat OpenShift and watching them thrive. But fast-forward 6 months and you find your environment cluttered with 40 of them, their purposes largely unknown. Redundant agents emerge across namespaces, duplicating work. Meanwhile, a security review uncovers agents utilizing static API keys over insecure HTTP. When an agent falters, the lack of tracing makes it impossible to distinguish between model hallucinations, tool errors, or downstream failures. The post Rossoctl for agent discoverability, security, and observability on Kubernetes appeared first on…

  • Go CPU and Memory Requests and Limits in Kubernetes (opens on the source site)

    Learnk8s ·

    Set Kubernetes CPU and memory requests and limits for Go services by checking GOMAXPROCS, GOMEMLIMIT, live heap, goroutine stacks, CPU throttling, and garbage collection cost.

  • How to Reduce Overprivileged Kubernetes Service Accounts (opens on the source site)

    Teleport ·

    IN THIS BLOG Overprivileged Kubernetes ServiceAccounts persist when broad RBAC, cloud IAM permissions, and long-lived credentials outlive their intended use. Reduce overprivileged access with least privileged RBAC, scoped cloud permissions, and short-lived certificates that eliminate static credentials. I spent two days last quarter tracking down why a developer could delete production secrets. The RBAC looked fine. The ClusterRoleBinding said edit, not cluster-admin. But someone had bound the default service account to cluster-admin permissions in a CI namespace three years ago, and that…

  • Kube AuthKit: Unified Kubernetes and OpenShift auth in Python (opens on the source site)

    Red Hat ·

    Kube AuthKit is a lightweight Python library that unifies Kubernetes and Red Hat OpenShift authentication behind a single, consistent API. Whether you're running locally with kubeconfig, inside a pod with a service account, or authenticating via OIDC or OpenShift OAuth—it's one line of code. The post Kube AuthKit: Unified Kubernetes and OpenShift auth in Python appeared first on Red Hat Developer.

  • Building Kubernetes PR Previews with Shared Pulumi Components (opens on the source site)

    Pulumi ·

    My team develops a microservices application on Kubernetes, with hundreds of PRs opened each day. To let engineers test and review those changes in isolation before they’re merged, we give every pull request its own ephemeral environment. We use Pulumi to define those short-lived PR environments from a component resource that’s shared with our long-lived Dev, Stage, Prod environments. Each PR gets its own Pulumi stack and Kubernetes namespace, which we tear down once the PR is merged or closed. In this post, I’ll walk through how we’ve implemented this pattern and what we’ve learned from…

  • How to Build a PostgreSQL CI/CD Pipeline with GitOps and Kubernetes (opens on the source site)

    SeveralNines ·

    Application teams have been living in Git-driven, immutable-deployment land for years. Databases lagged for good reason: That gap has closed considerably now that Kubernetes operators understand PostgreSQL well enough to manage replication, failover, and storage declaratively. What’s still missing in a lot of pipelines is the operational layer that sits on top of “it deployed”, […] The post How to Build a PostgreSQL CI/CD Pipeline with GitOps and Kubernetes appeared first on Severalnines.

  • Self-driving infrastructure with Pulumi and Jev (opens on the source site)

    Pulumi ·

    It’s been a weirdly great time to be building software. We’ve never had so many tools that help us get things done: endless cloud providers, regions, deployment frameworks, and now AI agents that can actually build and manage infrastructure for us. That’s part of what made this past week feel so big. TypeSafe AI opened early access to Jev, their SystemOne model, and we here at Pulumi were bitten by the excitement bug and got straight to work building. Almost immediately we were able to whip up some test agentic workflows that were making real, live decisions about infrastructure. If you’re…

  • Health for Laravel: Kubernetes Probes and Prometheus Metrics (opens on the source site)

    Laravel ·

    Health for Laravel adds Kubernetes liveness, readiness, and startup probes, a Prometheus metrics endpoint, and 10 built-in health checks to your app. The post Health for Laravel: Kubernetes Probes and Prometheus Metrics appeared first on Laravel News. Join the Laravel Newsletter to get Laravel articles like this directly in your inbox.

  • Red Hat OpenShift networking: Default OVN-Kubernetes vs. Cilium operator (opens on the source site)

    Red Hat ·

    As platform engineering teams scale their Red Hat OpenShift environments to support microservices, AI workloads, and declarative multi-cluster management with Cluster API (CAPI), decisions around Container Network Interface (CNI) architecture carry far-reaching consequences. In modern enterprise Kubernetes, networking is the foundation of your cluster's security posture, application performance, and operational visibility. The post Red Hat OpenShift networking: Default OVN-Kubernetes vs. Cilium operator appeared first on Red Hat Developer.

  • Beyond the Hyperscalers: What Actually Protects You (opens on the source site)

    Pulumi ·

    Recorded September 3, 2026. Quotes are lightly edited for clarity. Maybe this sounds familiar. You run infrastructure at a company that isn’t American. Your workloads are on AWS, Azure, or Google Cloud, probably more than one, because that is what everyone picked. Until recently nobody asked you where the data lives or who can reach it. Now you’re getting questions. Legal wants to know what NIS2 means for where your systems run. Someone on the leadership team read that the US government locked the cloud accounts of judges at the International Criminal Court and wants to know if that could…

  • Implementing GitOps from Infrastructure to DB Operators to Unify Ops for Kubernetes Databases (opens on the source site)

    SeveralNines ·

    Most platform teams already use GitOps for their Kubernetes apps. The config lives in Git, Argo CD applies it, and deployments are predictable. But look one layer down and things get messy. Infrastructure setup is often still done by hand, running Terraform locally or through scattered scripts. Database operations, especially for databases running through Kubernetes […] The post Implementing GitOps from Infrastructure to DB Operators to Unify Ops for Kubernetes Databases appeared first on Severalnines.

  • Reference architecture for HA scanning with Red Hat Advanced Cluster Security for Kubernetes (opens on the source site)

    Red Hat ·

    Red Hat Advanced Cluster Security for Kubernetes provides an image scanning API through its Central component that CI/CD pipelines depend on for vulnerability assessment. Red Hat Advanced Cluster Security upgrades and restarts introduce downtime windows that can block critical build paths. This post describes a reference architecture that eliminates scan API downtime by running two Central service instances with a client-side failover mechanism. The post Reference architecture for HA scanning with Red Hat Advanced Cluster Security for Kubernetes appeared first on Red Hat Developer.

  • Rerouting the Stream: How Lyft Moved to the Apache Flink Operator (opens on the source site)

    Lyft ·

    Written by Maheep Myneni, Arda Kuyumcu, and Prem Santosh Udaya Shankar at Lyft.Why We Migrated: Technical Debt Meets Modern Streaming DemandsOver the past several quarters, Lyft’s Streaming Compute team retired our internally developed Flink Kubernetes operator and moved our entire streaming fleet onto the open-source Apache Flink Kubernetes operator. This post is about why we made the switch, how we pulled it off incrementally without disrupting users, and the follow-on work it took to actually get the benefits we were after.Back in 2020, when we first architected the Lyft Flink Kubernetes…

  • Stop wasting GPU allocation in Kubernetes with GPU-pruner (opens on the source site)

    Red Hat ·

    In all Kubernetes platforms, idle GPU waste is one of the toughest capacity issues to resolve. You may have seen it yourself: GPUs allocated by other users apparently run for days, doing pretty much nothing. Kubernetes shows that the pods are still running, and the usage bill is still accumulating due to that allocation, but Data Center GPU Manager NVIDIA metrics reveal close to zero engine activity for hours. The post Stop wasting GPU allocation in Kubernetes with GPU-pruner appeared first on Red Hat Developer.

  • Java JVM CPU and Memory Requests and Limits in Kubernetes (opens on the source site)

    Learnk8s ·

    Setting Kubernetes CPU and memory requests and limits for a JVM service is four coupled decisions: container memory, heap size, GC selection, and CPU quota. Use the calculator to explore the space.

  • Pulumi Kubernetes v4.34.0: CRDs as provider extensions (opens on the source site)

    Pulumi ·

    We’re really excited to bring you v4.34.0, the newest version of the Pulumi Kubernetes provider, which includes improved support for Kubernetes Custom Resource Definitions (CRDs). As with any release, we’ve also shipped standard dependency updates and bug fixes. This provider release includes the newest resources for Kubernetes v1.37.0, which was recently cut. So that in itself is very exciting! But the feature we’re proudest of is that you can now extend the Kubernetes provider with any Kubernetes Custom Resource Definition of your choice by passing its manifest file to Pulumi, using the new…

  • Best Kubernetes Infrastructure as Code Tools in 2026 (opens on the source site)

    Pulumi ·

    There is no single best Kubernetes infrastructure as code tool, because “Kubernetes IaC” actually spans three different jobs. For provisioning the cluster and its cloud dependencies, Pulumi and Terraform (or OpenTofu) are the strongest general-purpose options. For templating and packaging workloads, Helm and Kustomize dominate. For continuous reconciliation once things are running, Argo CD and Flux lead the GitOps category. The right stack usually combines one tool from each layer, not a single tool that claims to do all three. What counts as infrastructure as code for Kubernetes? Kubernetes…

  • Setting the right requests and limits in Kubernetes (opens on the source site)

    Learnk8s ·

    Learn how Kubernetes CPU and memory requests and limits control scheduling, cgroup weights and quotas, eviction, throttling, and OOM kills.

  • We Gave Our Service More CPU — and It Took Down Production (opens on the source site)

    Housing.com ·

    Every engineer has a story about the “harmless” change that wasn’t. This is mine.Ours didn’t start with a bad deploy. No risky feature flag, no sketchy migration, no Friday-evening hotfix. It started with a change every ops playbook calls safe:We gave a struggling service more CPU.We went from 1.5 cores to 2.5 cores. That’s it. And within minutes, our lead-creation API — the thing that turns a visitor tapping “Contact” into an actual business lead — stopped creating leads. Entirely across web and app.The twist that cost us hours: the more CPU we threw at it, the worse it got. This is the…

  • Terraform and Kubernetes: A Practical Guide for 2026 (opens on the source site)

    Pulumi ·

    Yes, Terraform can manage Kubernetes: the official hashicorp/kubernetes provider lets you declare Deployments, Services, and other objects as HCL resources, and community providers like kubectl fill in the gaps. It works well for many teams. The friction shows up around two well-documented limits — provider ordering and plan-time API access — and around testing, where a general-purpose language changes what’s possible. That friction matters more in 2026 than it did a few years ago. Kubernetes infrastructure now sits next to AI-driven engineering workflows: agents that propose changes, run…

  • Your MVP doesn’t need a Kubernetes cluster (opens on the source site)

    Stack Overflow ·

    Ryan welcomes Anurag Goel, CEO and co-founder of Render, to discuss why most startups shouldn’t start by managing their Kubernetes and cloud infrastructure.

  • How to Run AI Agents on Kubernetes with Pulumi (opens on the source site)

    Pulumi ·

    Kubernetes has become the default place teams run agentic AI workloads: CNCF’s 2026 annual survey found that 66% of organizations hosting generative AI models use Kubernetes to manage some or all of their inference workloads.1 An entire ecosystem has grown up around that fact — agent runtimes, model servers, GPU schedulers — and most of it assumes the infrastructure underneath is already handled. It usually isn’t. An AI agent is not a stateless web service, and provisioning for one takes more than copying a Deployment YAML and swapping the image. I spend a lot of my time these days thinking…

  • Kubernetes Agent Sandbox: What It Is and How to Deploy It with Pulumi (opens on the source site)

    Pulumi ·

    When you use a coding agent, it can seem like there’s a trade-off between autonomy and permissions. If you approve every command, it’s safe but slow. Let it do whatever it likes and it works more autonomously, but as the nx supply-chain attack showed, that can go badly. The fix is to give the agent a sandbox: a box it’s allowed to wreck, with limited permissions and scoped network access. The only files are the checkout you handed it, the only credentials are the task’s own, and trashing the machine just means a disposable pod gets garbage-collected early. Pulumi Neo works this way, and if…

  • Announcing the public beta of Vault Kubernetes key management (opens on the source site)

    HashiCorp ·

    Announcing the public beta of Vault Kubernetes key management, enabling Kubernetes to use Vault Enterprise as a KMS provider for etcd encryption.

  • Kubernetes for Agentic AI: Best Practices for Identity and Access (opens on the source site)

    Teleport ·

    When agents act autonomously beyond a user's session, they need their own identity, least-privilege access, and full audit trails.

  • Migrating Counter Service storage: Design choices and learnings (opens on the source site)

    Grab ·

    Introduction Counter Service is used across Grab’s anti-fraud platform to answer time-windowed count questions, such as recent ride requests by a user or failed payment attempts on a card. The service handles tens of thousands of queries per second (QPS) with about a billion requests per day, while maintaining strict requirements around latency and reliability to support real-time fraud rule evaluation. For most of its life, Counter Service was backed by a wide-column database that served the workload reliably as the service scaled. As part of a broader infrastructure review mandated at an…

  • Upgrade Amazon EKS clusters with confidence using Kubernetes version rollbacks (opens on the source site)

    AWS ·

    Learn how Kubernetes version rollbacks for Amazon EKS let you reverse cluster upgrades within seven days. This new feature provides a safety net for upgrade failures—no cluster rebuilds required—turning Kubernetes version upgrades into a reversible, low-risk operation.

  • How Netflix Simplified Batch Compute with Kueue (opens on the source site)

    Netflix ·

    By Alvin Bao, Alex Petrov, Jennifer Lai, Aidan Sherr, and Samartha ChandrashekarAs a part of the journey to transition Netflix’s compute infrastructure to be more Kubernetes-native, we have leaned into incorporating components from the Kubernetes ecosystem into our container platform Titus. One example of this is our use of Kueue, a cloud-native job queueing system for batch workloads, which has largely replaced the custom queuing and scheduling logic in our homegrown managed batch solution Compute Managed Batch (CMB). In this post, we’ll give an overview of what motivated the migration, how…

  • Palana (Part 2): Architecting isolation, identity, and auditability for AI agents (opens on the source site)

    Grab ·

    Introduction In Part 1, we introduced Palana, Grab’s Kubernetes-native secure execution platform for autonomous AI agents. We discussed the underlying need for isolated environments and covered its core design principles: treating isolation as the unit of trust, keeping credentials out of agent hands, and mediating all network access. In this second part, we’ll dive under the hood into Palana’s architecture, look at the agent lifecycle, and share the key lessons we learned from putting this system into production. Architecture overview The core request path looks like this: Figure 1. Palana…

  • Authentication between microservices using Kubernetes identities (opens on the source site)

    Learnk8s ·

    Learn how you can secure communications between microservices to prevent unauthenticated requests using Kubernetes identities.

Related topics

This page is generated automatically from the engineering blogs we follow. Every post links to its source, where it was published. See all sources.

Privacy choices

Reading never requires analytics. These choices last 90 days on this browser.

Essential sign-in and security storage always stays on. Read the privacy notice.