A follow-on from the smash hit that was “My First 4 Months at Bazaarvoice.” While my first post focused on onboarding and adapting to a global organization, this post shifts focus to developing at scale, navigating major organizational shifts, and my journey to becoming a Senior Engineer. What’s Been Going On? (The High-Level View) To […]
culturedevopsexcerpt only · body stays at the source
GitHub has introduced Immutable Subject Claims for OIDC (OpenID Connect) in GitHub Actions. This update fixes a security issue where the subject claim in the OIDC token used to be based on organisation and repository names instead of permanent identifiers. What was the problem? Organisation and repository names can be changed or reused. This means that if an organisation or repository is deleted or renamed, another actor could, in theory, create a new repository or organisation with the same name. This would let them land in the same namespace within the subject claim, and in the worst case…
There is no single best Kubernetes infrastructure as code tool, because “Kubernetes IaC” actually spans three different jobs. For provisioning the cluster and its cloud dependencies, Pulumi and Terraform (or OpenTofu) are the strongest general-purpose options. For templating and packaging workloads, Helm and Kustomize dominate. For continuous reconciliation once things are running, Argo CD and Flux lead the GitOps category. The right stack usually combines one tool from each layer, not a single tool that claims to do all three. What counts as infrastructure as code for Kubernetes? Kubernetes…
Microsoft.Testing.Platform brings failures into GitHub Actions and Azure DevOps, uses pipeline history to separate regressions from flakes, and preserves usable reports when a test host crashes. The post Test reporting in Microsoft.Testing.Platform: from red build to root cause appeared first on .NET Blog.
Making large pricing changes safe to run, safe to resume, and safe to undo.Most days, shipping code to production at Thumbtack is a non-event. You merge your change, a deployment pipeline takes it from there and it rolls out while you monitor the rollout. If a deploy stalls, the system already knows which step it stalled on. If it goes wrong, rolling back is a button away. It’s the return on years of platform work, so routine now that we mostly forget it’s there. Which is the point: Safety is built into the road, not into how carefully each person drives.Changing pricing data never had that…
You usually find out in the postmortem: the alarm fired, but it paged a schedule nobody was on anymore. Or the service had been running in production for three months before anyone created the matching PagerDuty service, so the first person to notice the outage was a customer. The infrastructure was code, reviewed and versioned. The incident response setup was forty clicks in a web UI, done once, by someone who has since changed teams. PagerDuty’s own engineering team has been making the case for managing PagerDuty as code for years, and the OneUptime folks recently published a hands-on guide…
pagerdutyawsexcerpt only · body stays at the source
The strongest Terraform alternatives in 2026 split into three groups: general-purpose-language platforms like Pulumi and AWS CDK, HCL-compatible forks like OpenTofu, and cloud- or platform-specific tools like AWS CloudFormation, Azure Bicep, and Crossplane. (Pulumi also supports HCL directly and can serve as a Terraform-compatible state backend, so teams don’t have to choose between staying in HCL and gaining Pulumi Cloud’s platform capabilities.) Which one fits depends less on syntax preference than on how well it lets your team, and increasingly your AI coding agents, read, test, and safely…
This article is inspired by our LinuxCommunity.io forum discussion thread (thanks to users @tmick and @shybry747 for the feedback). Let's walk through what Podman is and how to use it as a Docker alternative on Linux.Continue reading...
Expedia Group Technology — EngineeringHow a screen-level performance metric reshaped platform decisions, engineering ownership, and release disciplinePhoto by Pietro De Grandi on UnsplashFor the last few years, my responsibility has been straightforward to state but hard to execute — owning the traveler login experience across mobile platforms.Not just whether a feature works, but whether it feels responsive, predictable, and trustworthy in the moments that matter most. During login, those moments are unforgiving: if a login screen hesitates travelers don’t interpret it as ‘a slow render’,…
In production environments, debugging alerts can sometimes feel like finding a needle in a haystack. Over the years, I’ve found the OSI (Open Systems Interconnection) model to be a reliable guide during Root Cause Analysis (RCA) of production issues.What is the OSI Model? The OSI model is a conceptual framework that standardizes the functions of a telecommunication or computing system into seven layers:Physical Layer — Hardware, cables, switchesData Link Layer — MAC addresses, switches, network topologyNetwork Layer — IP addressing, routingTransport Layer — TCP/UDP, ports, session…
Discover how Bazaarvoice migrated millions of UGC records from RDS MySQL to AWS Aurora – at scale and with minimal user impact. Learn about the technical challenges, strategies, and outcomes that enabled this ambitious transformation in reliability, performance, and cost efficiency Bazaarvoice ingests and serves millions of user-generated content (UGC) items—reviews, ratings, questions, answers, and […]
You may be new to this series; and if so welcome! If so, I encourage you to start at the beginning of our datastore journey and see the blog post “Unlocking Efficiency: A New Era for Datastore Provisioning”.Already up to date in our series? MAGICAL — then let’s continue with a quick re-cap.Where are we?We have introduced you to a multitude of aspects all pertaining to how we make the provisioning and utilisation of datastores quick, no-fuss and simple — as simple as clicking your fingers or making a wish.By continuing our theme of MAGIC, we enabled engineers at Zendesk to:Make a wish (detail…
storagedevopsexcerpt only · body stays at the source
Are you ready for more self-service datastore adventures? If you haven’t already, have a look at our previous entries in this series:Unlocking Efficiency: A New Era for Datastore ProvisioningSimplifying Datastore Provisioning with Kubernetes OperatorsResolving Incidents With The Remote Incident ConsoleThey’re a fun read.The story so farLast time, in Simplifying Datastore Provisioning with Kubernetes Operators, we talked about making datastores easy to provision by just writing a few lines of YAML in a file called service.yml, like this:version: 1.0name: "Terrific Tents"description:…
Confluent Kafka is an enterprise-ready distribution of Apache Kafka, providing a reliable and scalable platform for real-time data streaming. It includes additional features such as schema registry, REST proxy, connectors, and advanced monitoring tools.And migrating your existing Kafka cluster can be driven by several factors, including the need to upgrade hardware, move to a cloud-based infrastructure, enhance performance, improve scalability, or optimize resource utilization. This migration can help ensure your data infrastructure remains robust and capable of handling future…
Software isn't what it used to be. That's not necessarily a bad thing, but it does come with its own set of challenges. In the past, if you wanted to build a feature, you'd have to build it from scratch, without AI 😱 Fast forward from the dark ages of just
A conversation with engineers who help run BlinkitChinthakunta Sumanth Kumar Reddy is an SDE 3 at Blinkit. He joined us in March 2021 and has since helped us build a resilient application platform at Blinkit. He currently works as a part of Software Resilience Engineering (SRE)–enabling scalable database migrations for Blinkit’s applications.Tell us about your background and your journey in Blinkit so far.I have always been curious to learn new things and don’t have barriers as long as I can engage with the task. Hailing from a small village in Andhra Pradesh, I have travelled across several…
A conversation with engineers who help run BlinkitJay Dihenkar is a Staff Engineer at Blinkit. He joined us in December 2020 and has helped different teams manage and streamline their build and release processes. He is currently working towards continuously improving the reliability, scalability, observability, developer productivity, and other such aspects of a software system critical for ensuring that the system can meet the needs of its users and stakeholders over time.Tell us something about yourself and your journey in Blinkit so far.I started out as an engineer on the Release…
Optional Google Analytics helps us understand visits. Microsoft Clarity records masked interactions to improve the site. Optional tools stay off unless you choose them. Privacy details.