560 blogs tracked4,950 posts indexed

#devops

17 posts · 12 companies · newest first

Curated view of this subject:Topic: DevOps19
1

From 4 Months to 2 Years: Scaling Systems and Stepping into Seniority at Bazaarvoice (opens on the source site)

A follow-on from the smash hit that was “My First 4 Months at Bazaarvoice.” While my first post focused on onboarding and adapting to a global organization, this post shifts focus to developing at scale, navigating major organizational shifts, and my journey to becoming a Senior Engineer. What’s Been Going On? (The High-Level View) To […]

culturedevopsexcerpt only · body stays at the source
From the web
2

Important Update for GitHub Actions OIDC: Immutable Subject Claims (opens on the source site)

GitHub has introduced Immutable Subject Claims for OIDC (OpenID Connect) in GitHub Actions. This update fixes a security issue where the subject claim in the OIDC token used to be based on organisation and repository names instead of permanent identifiers. What was the problem? Organisation and repository names can be changed or reused. This means that if an organisation or repository is deleted or renamed, another actor could, in theory, create a new repository or organisation with the same name. This would let them land in the same namespace within the subject claim, and in the worst case…

github-actionsoidcexcerpt only · body stays at the source
From the web
3

Best Kubernetes Infrastructure as Code Tools in 2026 (opens on the source site)

There is no single best Kubernetes infrastructure as code tool, because “Kubernetes IaC” actually spans three different jobs. For provisioning the cluster and its cloud dependencies, Pulumi and Terraform (or OpenTofu) are the strongest general-purpose options. For templating and packaging workloads, Helm and Kustomize dominate. For continuous reconciliation once things are running, Argo CD and Flux lead the GitOps category. The right stack usually combines one tool from each layer, not a single tool that claims to do all three. What counts as infrastructure as code for Kubernetes? Kubernetes…

kubernetesinfrastructure-as-codeexcerpt only · body stays at the source
From the web
4

Test reporting in Microsoft.Testing.Platform: from red build to root cause (opens on the source site)

Microsoft.Testing.Platform brings failures into GitHub Actions and Azure DevOps, uses pipeline history to separate regressions from flakes, and preserves usable reports when a test host crashes. The post Test reporting in Microsoft.Testing.Platform: from red build to root cause appeared first on .NET Blog.

.netc#excerpt only · body stays at the source
From the web
5

Treating Pricing Changes Like Code Deploys (opens on the source site)

Making large pricing changes safe to run, safe to resume, and safe to undo.Most days, shipping code to production at Thumbtack is a non-event. You merge your change, a deployment pipeline takes it from there and it rolls out while you monitor the rollout. If a deploy stalls, the system already knows which step it stalled on. If it goes wrong, rolling back is a button away. It’s the return on years of platform work, so routine now that we mostly forget it’s there. Which is the point: Safety is built into the road, not into how carefully each person drives.Changing pricing data never had that…

marketplacesengineeringexcerpt only · body stays at the source
From the web
6

Incident Response as Code: Managing PagerDuty with Pulumi (opens on the source site)

You usually find out in the postmortem: the alarm fired, but it paged a schedule nobody was on anymore. Or the service had been running in production for three months before anyone created the matching PagerDuty service, so the first person to notice the outage was a customer. The infrastructure was code, reviewed and versioned. The incident response setup was forty clicks in a web UI, done once, by someone who has since changed teams. PagerDuty’s own engineering team has been making the case for managing PagerDuty as code for years, and the OneUptime folks recently published a hands-on guide…

pagerdutyawsexcerpt only · body stays at the source
From the web
7

Best Terraform Alternatives in 2026 (opens on the source site)

The strongest Terraform alternatives in 2026 split into three groups: general-purpose-language platforms like Pulumi and AWS CDK, HCL-compatible forks like OpenTofu, and cloud- or platform-specific tools like AWS CloudFormation, Azure Bicep, and Crossplane. (Pulumi also supports HCL directly and can serve as a Terraform-compatible state backend, so teams don’t have to choose between staying in HCL and gaining Pulumi Cloud’s platform capabilities.) Which one fits depends less on syntax preference than on how well it lets your team, and increasingly your AI coding agents, read, test, and safely…

terraforminfrastructure-as-codeexcerpt only · body stays at the source
From the web
9

The Most Expensive Milliseconds Are Unmeasured (opens on the source site)

Expedia Group Technology — EngineeringHow a screen-level performance metric reshaped platform decisions, engineering ownership, and release disciplinePhoto by Pietro De Grandi on UnsplashFor the last few years, my responsibility has been straightforward to state but hard to execute — owning the traveler login experience across mobile platforms.Not just whether a feature works, but whether it feels responsive, predictable, and trustworthy in the moments that matter most. During login, those moments are unforgiving: if a login screen hesitates travelers don’t interpret it as ‘a slow render’,…

time-to-interactivesoftware-engineeringexcerpt only · body stays at the source
From the web
10

Using the OSI Model for Effective Production Issue Debugging (opens on the source site)

In production environments, debugging alerts can sometimes feel like finding a needle in a haystack. Over the years, I’ve found the OSI (Open Systems Interconnection) model to be a reliable guide during Root Cause Analysis (RCA) of production issues.What is the OSI Model? The OSI model is a conceptual framework that standardizes the functions of a telecommunication or computing system into seven layers:Physical Layer — Hardware, cables, switchesData Link Layer — MAC addresses, switches, network topologyNetwork Layer — IP addressing, routingTransport Layer — TCP/UDP, ports, session…

production-debuggingsreexcerpt only · body stays at the source
From the web
11

How We Migrated Millions of UGC Records to Aurora MySQL (opens on the source site)

Discover how Bazaarvoice migrated millions of UGC records from RDS MySQL to AWS Aurora – at scale and with minimal user impact. Learn about the technical challenges, strategies, and outcomes that enabled this ambitious transformation in reliability, performance, and cost efficiency Bazaarvoice ingests and serves millions of user-generated content (UGC) items—reviews, ratings, questions, answers, and […]

databasedevopsexcerpt only · body stays at the source
From the web
12

Unleashing the Magic of Zendesk Datastore Management: Your One-Stop Self-Service Hub! (opens on the source site)

You may be new to this series; and if so welcome! If so, I encourage you to start at the beginning of our datastore journey and see the blog post “Unlocking Efficiency: A New Era for Datastore Provisioning”.Already up to date in our series? MAGICAL — then let’s continue with a quick re-cap.Where are we?We have introduced you to a multitude of aspects all pertaining to how we make the provisioning and utilisation of datastores quick, no-fuss and simple — as simple as clicking your fingers or making a wish.By continuing our theme of MAGIC, we enabled engineers at Zendesk to:Make a wish (detail…

storagedevopsexcerpt only · body stays at the source
From the web
13

Connecting Applications to Self-Service Datastores (opens on the source site)

Are you ready for more self-service datastore adventures? If you haven’t already, have a look at our previous entries in this series:Unlocking Efficiency: A New Era for Datastore ProvisioningSimplifying Datastore Provisioning with Kubernetes OperatorsResolving Incidents With The Remote Incident ConsoleThey’re a fun read.The story so farLast time, in Simplifying Datastore Provisioning with Kubernetes Operators, we talked about making datastores easy to provision by just writing a few lines of YAML in a file called service.yml, like this:version: 1.0name: "Terrific Tents"description:…

storagecredentialsexcerpt only · body stays at the source
From the web
14

Migrating to Confluent Kafka: A Comprehensive Guide (opens on the source site)

Confluent Kafka is an enterprise-ready distribution of Apache Kafka, providing a reliable and scalable platform for real-time data streaming. It includes additional features such as schema registry, REST proxy, connectors, and advanced monitoring tools.And migrating your existing Kafka cluster can be driven by several factors, including the need to upgrade hardware, move to a cloud-based infrastructure, enhance performance, improve scalability, or optimize resource utilization. This migration can help ensure your data infrastructure remains robust and capable of handling future…

confluentdevopsexcerpt only · body stays at the source
From the web
16

In Focus: Sumanth Reddy (opens on the source site)

A conversation with engineers who help run BlinkitChinthakunta Sumanth Kumar Reddy is an SDE 3 at Blinkit. He joined us in March 2021 and has since helped us build a resilient application platform at Blinkit. He currently works as a part of Software Resilience Engineering (SRE)–enabling scalable database migrations for Blinkit’s applications.Tell us about your background and your journey in Blinkit so far.I have always been curious to learn new things and don’t have barriers as long as I can engage with the task. Hailing from a small village in Andhra Pradesh, I have travelled across several…

quick-commercedevopsexcerpt only · body stays at the source
From the web
17

In Focus: Jay Dihenkar (opens on the source site)

A conversation with engineers who help run BlinkitJay Dihenkar is a Staff Engineer at Blinkit. He joined us in December 2020 and has helped different teams manage and streamline their build and release processes. He is currently working towards continuously improving the reliability, scalability, observability, developer productivity, and other such aspects of a software system critical for ensuring that the system can meet the needs of its users and stakeholders over time.Tell us something about yourself and your journey in Blinkit so far.I started out as an engineer on the Release…

people-at-blinkitcultureexcerpt only · body stays at the source
From the web
17 shown

Privacy choices

Reading never requires analytics. These choices last 90 days on this browser.

Essential sign-in and security storage always stays on. Read the privacy notice.