My team develops a microservices application on Kubernetes, with hundreds of PRs opened each day. To let engineers test and review those changes in isolation before they’re merged, we give every pull request its own ephemeral environment. We use Pulumi to define those short-lived PR environments from a component resource that’s shared with our long-lived Dev, Stage, Prod environments. Each PR gets its own Pulumi stack and Kubernetes namespace, which we tear down once the PR is merged or closed. In this post, I’ll walk through how we’ve implemented this pattern and what we’ve learned from…
It’s been a weirdly great time to be building software. We’ve never had so many tools that help us get things done: endless cloud providers, regions, deployment frameworks, and now AI agents that can actually build and manage infrastructure for us. That’s part of what made this past week feel so big. TypeSafe AI opened early access to Jev, their SystemOne model, and we here at Pulumi were bitten by the excitement bug and got straight to work building. Almost immediately we were able to whip up some test agentic workflows that were making real, live decisions about infrastructure. If you’re…
Recorded September 3, 2026. Quotes are lightly edited for clarity. Maybe this sounds familiar. You run infrastructure at a company that isn’t American. Your workloads are on AWS, Azure, or Google Cloud, probably more than one, because that is what everyone picked. Until recently nobody asked you where the data lives or who can reach it. Now you’re getting questions. Legal wants to know what NIS2 means for where your systems run. Someone on the leadership team read that the US government locked the cloud accounts of judges at the International Criminal Court and wants to know if that could…
kubernetesawsexcerpt only · body stays at the source
Most platform teams already use GitOps for their Kubernetes apps. The config lives in Git, Argo CD applies it, and deployments are predictable. But look one layer down and things get messy. Infrastructure setup is often still done by hand, running Terraform locally or through scattered scripts. Database operations, especially for databases running through Kubernetes […] The post Implementing GitOps from Infrastructure to DB Operators to Unify Ops for Kubernetes Databases appeared first on Severalnines.
Written by Maheep Myneni, Arda Kuyumcu, and Prem Santosh Udaya Shankar at Lyft.Why We Migrated: Technical Debt Meets Modern Streaming DemandsOver the past several quarters, Lyft’s Streaming Compute team retired our internally developed Flink Kubernetes operator and moved our entire streaming fleet onto the open-source Apache Flink Kubernetes operator. This post is about why we made the switch, how we pulled it off incrementally without disrupting users, and the follow-on work it took to actually get the benefits we were after.Back in 2020, when we first architected the Lyft Flink Kubernetes…
We’re really excited to bring you v4.34.0, the newest version of the Pulumi Kubernetes provider, which includes improved support for Kubernetes Custom Resource Definitions (CRDs). As with any release, we’ve also shipped standard dependency updates and bug fixes. This provider release includes the newest resources for Kubernetes v1.37.0, which was recently cut. So that in itself is very exciting! But the feature we’re proudest of is that you can now extend the Kubernetes provider with any Kubernetes Custom Resource Definition of your choice by passing its manifest file to Pulumi, using the new…
There is no single best Kubernetes infrastructure as code tool, because “Kubernetes IaC” actually spans three different jobs. For provisioning the cluster and its cloud dependencies, Pulumi and Terraform (or OpenTofu) are the strongest general-purpose options. For templating and packaging workloads, Helm and Kustomize dominate. For continuous reconciliation once things are running, Argo CD and Flux lead the GitOps category. The right stack usually combines one tool from each layer, not a single tool that claims to do all three. What counts as infrastructure as code for Kubernetes? Kubernetes…
Every engineer has a story about the “harmless” change that wasn’t. This is mine.Ours didn’t start with a bad deploy. No risky feature flag, no sketchy migration, no Friday-evening hotfix. It started with a change every ops playbook calls safe:We gave a struggling service more CPU.We went from 1.5 cores to 2.5 cores. That’s it. And within minutes, our lead-creation API — the thing that turns a visitor tapping “Contact” into an actual business lead — stopped creating leads. Entirely across web and app.The twist that cost us hours: the more CPU we threw at it, the worse it got. This is the…
Yes, Terraform can manage Kubernetes: the official hashicorp/kubernetes provider lets you declare Deployments, Services, and other objects as HCL resources, and community providers like kubectl fill in the gaps. It works well for many teams. The friction shows up around two well-documented limits — provider ordering and plan-time API access — and around testing, where a general-purpose language changes what’s possible. That friction matters more in 2026 than it did a few years ago. Kubernetes infrastructure now sits next to AI-driven engineering workflows: agents that propose changes, run…
Kubernetes has become the default place teams run agentic AI workloads: CNCF’s 2026 annual survey found that 66% of organizations hosting generative AI models use Kubernetes to manage some or all of their inference workloads.1 An entire ecosystem has grown up around that fact — agent runtimes, model servers, GPU schedulers — and most of it assumes the infrastructure underneath is already handled. It usually isn’t. An AI agent is not a stateless web service, and provisioning for one takes more than copying a Deployment YAML and swapping the image. I spend a lot of my time these days thinking…
When you use a coding agent, it can seem like there’s a trade-off between autonomy and permissions. If you approve every command, it’s safe but slow. Let it do whatever it likes and it works more autonomously, but as the nx supply-chain attack showed, that can go badly. The fix is to give the agent a sandbox: a box it’s allowed to wreck, with limited permissions and scoped network access. The only files are the checkout you handed it, the only credentials are the task’s own, and trashing the machine just means a disposable pod gets garbage-collected early. Pulumi Neo works this way, and if…
kubernetesaiexcerpt only · body stays at the source
Introduction Counter Service is used across Grab’s anti-fraud platform to answer time-windowed count questions, such as recent ride requests by a user or failed payment attempts on a card. The service handles tens of thousands of queries per second (QPS) with about a billion requests per day, while maintaining strict requirements around latency and reliability to support real-time fraud rule evaluation. For most of its life, Counter Service was backed by a wide-column database that served the workload reliably as the service scaled. As part of a broader infrastructure review mandated at an…
By Alvin Bao, Alex Petrov, Jennifer Lai, Aidan Sherr, and Samartha ChandrashekarAs a part of the journey to transition Netflix’s compute infrastructure to be more Kubernetes-native, we have leaned into incorporating components from the Kubernetes ecosystem into our container platform Titus. One example of this is our use of Kueue, a cloud-native job queueing system for batch workloads, which has largely replaced the custom queuing and scheduling logic in our homegrown managed batch solution Compute Managed Batch (CMB). In this post, we’ll give an overview of what motivated the migration, how…
kueuenetflixexcerpt only · body stays at the source
Introduction In Part 1, we introduced Palana, Grab’s Kubernetes-native secure execution platform for autonomous AI agents. We discussed the underlying need for isolated environments and covered its core design principles: treating isolation as the unit of trust, keeping credentials out of agent hands, and mediating all network access. In this second part, we’ll dive under the hood into Palana’s architecture, look at the agent lifecycle, and share the key lessons we learned from putting this system into production. Architecture overview The core request path looks like this: Figure 1. Palana…
You are in the process of putting together your application. While designing your authorization solution, you realize you will need to send emails to potential clients.Using a third-party service (like SendGrid or Mailgun) to cover your needs for now looks pretty attractive. After all, you don’t have any users yet, they offer free tiers, and […] The post Postfix on Kubernetes: A Step-by-Step Email Guide appeared first on RisingStack Engineering.
Are you ready for more self-service datastore adventures? If you haven’t already, have a look at our previous entries in this series:Unlocking Efficiency: A New Era for Datastore ProvisioningSimplifying Datastore Provisioning with Kubernetes OperatorsResolving Incidents With The Remote Incident ConsoleThey’re a fun read.The story so farLast time, in Simplifying Datastore Provisioning with Kubernetes Operators, we talked about making datastores easy to provision by just writing a few lines of YAML in a file called service.yml, like this:version: 1.0name: "Terrific Tents"description:…
Scaling Kubernetes efficiently is an art that blends strategic planning with technical expertise. In Azure Kubernetes Service (AKS), mastering the allocation of IP addresses is crucial for ensuring seamless scalability and optimal performance. When configuring an AKS cluster with default Azure CNI networking, it’s essential to address the nature of scaling Pods, Nodes, and internal LoadBalancer-type services.This blog delves into the intricate process of IP address planning, explicitly focusing on scaling Pods, Nodes, and Services in an AKS cluster. Discover how thoughtful IP management can…
Optional Google Analytics helps us understand visits. Microsoft Clarity records masked interactions to improve the site. Optional tools stay off unless you choose them. Privacy details.