How orphaned BigQuery storage, Delta retention, DMS right-sizing, and Databricks System Tables became our biggest cloud cost wins.The biggest cloud cost optimization we made wasn’t shrinking clusters.It was deleting data we’d forgotten we were paying for.Like most teams, our first instinct was to tune infrastructure first. Instead, we discovered a treasure trove of hidden costs: orphaned BigQuery datasets, 90-day Delta retention, 24-hour jobs no one monitored, and DMS infrastructure that no longer matched business needs.We stopped treating cloud bills as a finance problem and started treating…
by Aarti Laddha, Richard Diaz-Cool, Rishika Idnani, Venkatesh SelverajNetflix supports a vast and evolving set of features and content types, ranging from 4K streaming and immersive audio to live streaming and cloud gaming, across a diverse ecosystem of devices. However, not all devices are created equal. Hardware limitations such as available RAM, CPU cores, display capabilities, or platform support mean that some features cannot be supported on certain device models. To ensure the best possible user experience, we rely on a deep understanding of device capabilities. We have invested in…
“We have multiple dashboards showing different numbers. Which one is correct?”If you’ve worked in data long enough, you’ve probably heard this question more times than you’d like.Sales reports one revenue figure.Finance reports another.Product Analytics has a third.Executives spend more time debating whose dashboard is correct than discussing what action to take.The natural response is often:“Let’s build a Single Source of Truth.”Sounds simple.Build a few centralized tables.Move everyone onto the same dashboards.Problem solved.Except…it rarely is.After leading an enterprise-wide Single Source…
By Parth Jain, Rakesh Sukumar, Yingwu Zhao, Renzo Sanchez-Silva & Nathan FisherA deep dive into the engineering challenges of building a real-time service dependency map at Netflix scale: from streaming architectures and distributed aggregation pipelines to time-travel queries and the methodology that made it work.IntroductionIn our first post, we introduced the problem: engineers at Netflix needed a unified, real-time view of service dependencies to troubleshoot faster, understand blast radius, and navigate our distributed architecture. We described our multi-source approach, combining eBPF…
Expedia Group Technology — InnovationUsing large language models to reveal bottlenecks in Spark SQL execution plansPhoto by Luis del RíoIf you’ve ever stared at a 300-plus-node physical plan at 2 a.m. trying to spot a missing broadcast or one cursed skewed partition, this is for you.Spark makes it deceptively easy to write complex SQL that looks correct but quietly turns into a performance and cost problem at scale. A query that runs fine on day one can slow to a crawl as data grows, joins get wider, and aggregations become more nested. Suddenly, jobs take hours instead of minutes, clusters…
Before We Knew BetterAs our data platform grew, so did the number of pipelines, scheduled tasks, and data quality checks running every day.While Snowflake provided a reliable platform for storing and processing data, operational monitoring was fragmented across multiple systems. Data quality failures were often discovered only after downstream reports showed inconsistencies. Pipeline issues sometimes required engineers to manually inspect logs, query tables, and trace execution paths before identifying the root cause.The challenge wasn’t detecting failures — we already had mechanisms to…
How Airbnb’s data engineers and analytics engineers built a consistent and flexible data modeling framework to support the expansion into Homes, Experiences, and Services.By: Patrick Lam, Namrata Lamba, Jamie StoberWith the May 2025 Summer Release, Airbnb redesigned its app, relaunched Experiences, and debuted Services, pushing us beyond our traditional Homes focus. For the data teams, this meant rapidly evolving a decade-old infrastructure to integrate two brand-new product pillars. Our data engineers and analytics engineers rose to the challenge by building a consistent and flexible…
Photo by Corinne Kutz on UnsplashBefore we knew betterOur orchestration system started as a simple internal solution to manage event pipelines and trigger downstream jobs. Over time, as more workflows and dependencies were added, it gradually evolved into a tightly coupled monolithic scheduler that became increasingly difficult to understand and maintain.Understanding how a workflow executed often meant looking through multiple files, configurations and database tables.For newer team members, onboarding into the system took time because much of the workflow context was distributed across…
Excerpt By 2024, Slack’s data platform had accumulated 700+ SSH-based operators orchestrating critical data pipelines. We’re talking daily search indexing that processed terabytes of data, analytics jobs powering business intelligence, the whole shebang. Every single one of these jobs required direct SSH access to production AWS Elastic MapReduce (EMR) clusters. We had a massive security…
Hi my name is Bisman and I studied Computer Science at University of California, Santa Barbara. During summer of 2022, I had the most amazing experience working as a Software Engineer Intern on Strava’s Data Platform Team. In the first fews weeks, I learned the tools my team uses and then spent the rest of the time working on my project.TRACKING BAD EVENTSFor my major summer project, I created a data pipeline that pulls user behavior data out of external storage and persists it in our data warehouse. Strava uses a service called Snowplow to collect this user behavior data, like loading a club…
Optional Google Analytics helps us understand visits. Microsoft Clarity records masked interactions to improve the site. Optional tools stay off unless you choose them. Privacy details.