560 blogs tracked4,950 posts indexed

#kubernetes

17 posts · 9 companies · newest first

Curated view of this subject:Topic: Kubernetes65
1

Building Kubernetes PR Previews with Shared Pulumi Components (opens on the source site)

My team develops a microservices application on Kubernetes, with hundreds of PRs opened each day. To let engineers test and review those changes in isolation before they’re merged, we give every pull request its own ephemeral environment. We use Pulumi to define those short-lived PR environments from a component resource that’s shared with our long-lived Dev, Stage, Prod environments. Each PR gets its own Pulumi stack and Kubernetes namespace, which we tear down once the PR is merged or closed. In this post, I’ll walk through how we’ve implemented this pattern and what we’ve learned from…

pulumikubernetesexcerpt only · body stays at the source
From the web
2

Self-driving infrastructure with Pulumi and Jev (opens on the source site)

It’s been a weirdly great time to be building software. We’ve never had so many tools that help us get things done: endless cloud providers, regions, deployment frameworks, and now AI agents that can actually build and manage infrastructure for us. That’s part of what made this past week feel so big. TypeSafe AI opened early access to Jev, their SystemOne model, and we here at Pulumi were bitten by the excitement bug and got straight to work building. Almost immediately we were able to whip up some test agentic workflows that were making real, live decisions about infrastructure. If you’re…

ai-agentsautomation-apiexcerpt only · body stays at the source
From the web
3

Beyond the Hyperscalers: What Actually Protects You (opens on the source site)

Recorded September 3, 2026. Quotes are lightly edited for clarity. Maybe this sounds familiar. You run infrastructure at a company that isn’t American. Your workloads are on AWS, Azure, or Google Cloud, probably more than one, because that is what everyone picked. Until recently nobody asked you where the data lives or who can reach it. Now you’re getting questions. Legal wants to know what NIS2 means for where your systems run. Someone on the leadership team read that the US government locked the cloud accounts of judges at the International Criminal Court and wants to know if that could…

kubernetesawsexcerpt only · body stays at the source
From the web
4

Implementing GitOps from Infrastructure to DB Operators to Unify Ops for Kubernetes Databases (opens on the source site)

Most platform teams already use GitOps for their Kubernetes apps. The config lives in Git, Argo CD applies it, and deployments are predictable. But look one layer down and things get messy. Infrastructure setup is often still done by hand, running Terraform locally or through scattered scripts. Database operations, especially for databases running through Kubernetes […] The post Implementing GitOps from Infrastructure to DB Operators to Unify Ops for Kubernetes Databases appeared first on Severalnines.

hybrid-operationssovereign-dbaasexcerpt only · body stays at the source
From the web
5

Rerouting the Stream: How Lyft Moved to the Apache Flink Operator (opens on the source site)

Written by Maheep Myneni, Arda Kuyumcu, and Prem Santosh Udaya Shankar at Lyft.Why We Migrated: Technical Debt Meets Modern Streaming DemandsOver the past several quarters, Lyft’s Streaming Compute team retired our internally developed Flink Kubernetes operator and moved our entire streaming fleet onto the open-source Apache Flink Kubernetes operator. This post is about why we made the switch, how we pulled it off incrementally without disrupting users, and the follow-on work it took to actually get the benefits we were after.Back in 2020, when we first architected the Lyft Flink Kubernetes…

data-platformskubernetesexcerpt only · body stays at the source
From the web
6

Pulumi Kubernetes v4.34.0: CRDs as provider extensions (opens on the source site)

We’re really excited to bring you v4.34.0, the newest version of the Pulumi Kubernetes provider, which includes improved support for Kubernetes Custom Resource Definitions (CRDs). As with any release, we’ve also shipped standard dependency updates and bug fixes. This provider release includes the newest resources for Kubernetes v1.37.0, which was recently cut. So that in itself is very exciting! But the feature we’re proudest of is that you can now extend the Kubernetes provider with any Kubernetes Custom Resource Definition of your choice by passing its manifest file to Pulumi, using the new…

kubernetesreleasesexcerpt only · body stays at the source
From the web
7

Best Kubernetes Infrastructure as Code Tools in 2026 (opens on the source site)

There is no single best Kubernetes infrastructure as code tool, because “Kubernetes IaC” actually spans three different jobs. For provisioning the cluster and its cloud dependencies, Pulumi and Terraform (or OpenTofu) are the strongest general-purpose options. For templating and packaging workloads, Helm and Kustomize dominate. For continuous reconciliation once things are running, Argo CD and Flux lead the GitOps category. The right stack usually combines one tool from each layer, not a single tool that claims to do all three. What counts as infrastructure as code for Kubernetes? Kubernetes…

kubernetesinfrastructure-as-codeexcerpt only · body stays at the source
From the web
8

We Gave Our Service More CPU — and It Took Down Production (opens on the source site)

Every engineer has a story about the “harmless” change that wasn’t. This is mine.Ours didn’t start with a bad deploy. No risky feature flag, no sketchy migration, no Friday-evening hotfix. It started with a change every ops playbook calls safe:We gave a struggling service more CPU.We went from 1.5 cores to 2.5 cores. That’s it. And within minutes, our lead-creation API — the thing that turns a visitor tapping “Contact” into an actual business lead — stopped creating leads. Entirely across web and app.The twist that cost us hours: the more CPU we threw at it, the worse it got. This is the…

kubernetesproduction-outageexcerpt only · body stays at the source
From the web
9

Terraform and Kubernetes: A Practical Guide for 2026 (opens on the source site)

Yes, Terraform can manage Kubernetes: the official hashicorp/kubernetes provider lets you declare Deployments, Services, and other objects as HCL resources, and community providers like kubectl fill in the gaps. It works well for many teams. The friction shows up around two well-documented limits — provider ordering and plan-time API access — and around testing, where a general-purpose language changes what’s possible. That friction matters more in 2026 than it did a few years ago. Kubernetes infrastructure now sits next to AI-driven engineering workflows: agents that propose changes, run…

kubernetesterraformexcerpt only · body stays at the source
From the web
10

How to Run AI Agents on Kubernetes with Pulumi (opens on the source site)

Kubernetes has become the default place teams run agentic AI workloads: CNCF’s 2026 annual survey found that 66% of organizations hosting generative AI models use Kubernetes to manage some or all of their inference workloads.1 An entire ecosystem has grown up around that fact — agent runtimes, model servers, GPU schedulers — and most of it assumes the infrastructure underneath is already handled. It usually isn’t. An AI agent is not a stateless web service, and provisioning for one takes more than copying a Deployment YAML and swapping the image. I spend a lot of my time these days thinking…

kubernetesai-agentsexcerpt only · body stays at the source
From the web
11

Kubernetes Agent Sandbox: What It Is and How to Deploy It with Pulumi (opens on the source site)

When you use a coding agent, it can seem like there’s a trade-off between autonomy and permissions. If you approve every command, it’s safe but slow. Let it do whatever it likes and it works more autonomously, but as the nx supply-chain attack showed, that can go badly. The fix is to give the agent a sandbox: a box it’s allowed to wreck, with limited permissions and scoped network access. The only files are the checkout you handed it, the only credentials are the task’s own, and trashing the machine just means a disposable pod gets garbage-collected early. Pulumi Neo works this way, and if…

kubernetesaiexcerpt only · body stays at the source
From the web
12

Migrating Counter Service storage: Design choices and learnings (opens on the source site)

Introduction Counter Service is used across Grab’s anti-fraud platform to answer time-windowed count questions, such as recent ride requests by a user or failed payment attempts on a card. The service handles tens of thousands of queries per second (QPS) with about a billion requests per day, while maintaining strict requirements around latency and reliability to support real-time fraud rule evaluation. For most of its life, Counter Service was backed by a wide-column database that served the workload reliably as the service scaled. As part of a broader infrastructure review mandated at an…

securityartificial-intelligenceexcerpt only · body stays at the source
From the web
13

How Netflix Simplified Batch Compute with Kueue (opens on the source site)

By Alvin Bao, Alex Petrov, Jennifer Lai, Aidan Sherr, and Samartha ChandrashekarAs a part of the journey to transition Netflix’s compute infrastructure to be more Kubernetes-native, we have leaned into incorporating components from the Kubernetes ecosystem into our container platform Titus. One example of this is our use of Kueue, a cloud-native job queueing system for batch workloads, which has largely replaced the custom queuing and scheduling logic in our homegrown managed batch solution Compute Managed Batch (CMB). In this post, we’ll give an overview of what motivated the migration, how…

kueuenetflixexcerpt only · body stays at the source
From the web
14

Palana (Part 2): Architecting isolation, identity, and auditability for AI agents (opens on the source site)

Introduction In Part 1, we introduced Palana, Grab’s Kubernetes-native secure execution platform for autonomous AI agents. We discussed the underlying need for isolated environments and covered its core design principles: treating isolation as the unit of trust, keeping credentials out of agent hands, and mediating all network access. In this second part, we’ll dive under the hood into Palana’s architecture, look at the agent lifecycle, and share the key lessons we learned from putting this system into production. Architecture overview The core request path looks like this: Figure 1. Palana…

securityartificial-intelligenceexcerpt only · body stays at the source
From the web
15

Postfix on Kubernetes: A Step-by-Step Email Guide (opens on the source site)

You are in the process of putting together your application. While designing your authorization solution, you realize you will need to send emails to potential clients.Using a third-party service (like SendGrid or Mailgun) to cover your needs for now looks pretty attractive. After all, you don’t have any users yet, they offer free tiers, and […] The post Postfix on Kubernetes: A Step-by-Step Email Guide appeared first on RisingStack Engineering.

kubernetesexcerpt only · body stays at the source
From the web
16

Connecting Applications to Self-Service Datastores (opens on the source site)

Are you ready for more self-service datastore adventures? If you haven’t already, have a look at our previous entries in this series:Unlocking Efficiency: A New Era for Datastore ProvisioningSimplifying Datastore Provisioning with Kubernetes OperatorsResolving Incidents With The Remote Incident ConsoleThey’re a fun read.The story so farLast time, in Simplifying Datastore Provisioning with Kubernetes Operators, we talked about making datastores easy to provision by just writing a few lines of YAML in a file called service.yml, like this:version: 1.0name: "Terrific Tents"description:…

storagecredentialsexcerpt only · body stays at the source
From the web
17

Future-Proof Your AKS Cluster with Strategic IP Address Planning (opens on the source site)

Scaling Kubernetes efficiently is an art that blends strategic planning with technical expertise. In Azure Kubernetes Service (AKS), mastering the allocation of IP addresses is crucial for ensuring seamless scalability and optimal performance. When configuring an AKS cluster with default Azure CNI networking, it’s essential to address the nature of scaling Pods, Nodes, and internal LoadBalancer-type services.This blog delves into the intricate process of IP address planning, explicitly focusing on scaling Pods, Nodes, and Services in an AKS cluster. Discover how thoughtful IP management can…

container-networkingazure-kubernetes-serviceexcerpt only · body stays at the source
From the web
17 shown

Privacy choices

Reading never requires analytics. These choices last 90 days on this browser.

Essential sign-in and security storage always stays on. Read the privacy notice.