560 blogs tracked4,950 posts indexed

#data-engineering

10 posts · 7 companies · newest first

1

We Cut Cloud Waste Before Touching Cluster Sizes: Lessons from Running a Data Platform (opens on the source site)

How orphaned BigQuery storage, Delta retention, DMS right-sizing, and Databricks System Tables became our biggest cloud cost wins.The biggest cloud cost optimization we made wasn’t shrinking clusters.It was deleting data we’d forgotten we were paying for.Like most teams, our first instinct was to tune infrastructure first. Instead, we discovered a treasure trove of hidden costs: orphaned BigQuery datasets, 90-day Delta retention, 24-hour jobs no one monitored, and DMS infrastructure that no longer matched business needs.We stopped treating cloud bills as a finance problem and started treating…

bigquerycloud-computingexcerpt only · body stays at the source
From the web
2

Modeling Device Capabilities for Analytics (opens on the source site)

by Aarti Laddha, Richard Diaz-Cool, Rishika Idnani, Venkatesh SelverajNetflix supports a vast and evolving set of features and content types, ranging from 4K streaming and immersive audio to live streaming and cloud gaming, across a diverse ecosystem of devices. However, not all devices are created equal. Hardware limitations such as available RAM, CPU cores, display capabilities, or platform support mean that some features cannot be supported on certain device models. To ensure the best possible user experience, we rely on a deep understanding of device capabilities. We have invested in…

data-engineeringdevicesexcerpt only · body stays at the source
From the web
3

Why Most Single Source of Truth Initiatives Fail (And What Successful Teams Do Differently) (opens on the source site)

“We have multiple dashboards showing different numbers. Which one is correct?”If you’ve worked in data long enough, you’ve probably heard this question more times than you’d like.Sales reports one revenue figure.Finance reports another.Product Analytics has a third.Executives spend more time debating whose dashboard is correct than discussing what action to take.The natural response is often:“Let’s build a Single Source of Truth.”Sounds simple.Build a few centralized tables.Move everyone onto the same dashboards.Problem solved.Except…it rarely is.After leading an enterprise-wide Single Source…

leadershipdatabricksexcerpt only · body stays at the source
From the web
4

Building Service Topology at Scale: Architecture, Challenges, and Lessons Learned (opens on the source site)

By Parth Jain, Rakesh Sukumar, Yingwu Zhao, Renzo Sanchez-Silva & Nathan FisherA deep dive into the engineering challenges of building a real-time service dependency map at Netflix scale: from streaming architectures and distributed aggregation pipelines to time-travel queries and the methodology that made it work.IntroductionIn our first post, we introduced the problem: engineers at Netflix needed a unified, real-time view of service dependencies to troubleshoot faster, understand blast radius, and navigate our distributed architecture. We described our multi-source approach, combining eBPF…

backend-developmentdistributed-systemsexcerpt only · body stays at the source
From the web
5

Using LLMs to Analyze Spark SQL Plans: A Practical Approach to Debugging Long-Running Jobs (opens on the source site)

Expedia Group Technology — InnovationUsing large language models to reveal bottlenecks in Spark SQL execution plansPhoto by Luis del RíoIf you’ve ever stared at a 300-plus-node physical plan at 2 a.m. trying to spot a missing broadcast or one cursed skewed partition, this is for you.Spark makes it deceptively easy to write complex SQL that looks correct but quietly turns into a performance and cost problem at scale. A query that runs fine on day one can slow to a crawl as data grows, joins get wider, and aggregations become more nested. Suddenly, jobs take hours instead of minutes, clusters…

apache-sparkbig-dataexcerpt only · body stays at the source
From the web
6

Building a Centralized Alerting Framework for Data Quality Monitoring and Incident Management (opens on the source site)

Before We Knew BetterAs our data platform grew, so did the number of pipelines, scheduled tasks, and data quality checks running every day.While Snowflake provided a reliable platform for storing and processing data, operational monitoring was fragmented across multiple systems. Data quality failures were often discovered only after downstream reports showed inconsistencies. Pipeline issues sometimes required engineers to manually inspect logs, query tables, and trace execution paths before identifying the root cause.The challenge wasn’t detecting failures — we already had mechanisms to…

snowflakeincident-managementexcerpt only · body stays at the source
From the web
7

Scaling beyond one: How Airbnb evolved its data architecture for a multi-product world (opens on the source site)

How Airbnb’s data engineers and analytics engineers built a consistent and flexible data modeling framework to support the expansion into Homes, Experiences, and Services.By: Patrick Lam, Namrata Lamba, Jamie StoberWith the May 2025 Summer Release, Airbnb redesigned its app, relaunched Experiences, and debuted Services, pushing us beyond our traditional Homes focus. For the data teams, this meant rapidly evolving a decade-old infrastructure to integrate two brand-new product pillars. Our data engineers and analytics engineers rose to the challenge by building a consistent and flexible…

data-engineeringdataexcerpt only · body stays at the source
From the web
8

Migrating from a Monolithic Orchestrator to Apache Airflow (opens on the source site)

Photo by Corinne Kutz on UnsplashBefore we knew betterOur orchestration system started as a simple internal solution to manage event pipelines and trigger downstream jobs. Over time, as more workflows and dependencies were added, it gradually evolved into a tightly coupled monolithic scheduler that became increasingly difficult to understand and maintain.Understanding how a workflow executed often meant looking through multiple files, configurations and database tables.For newer team members, onboarding into the system took time because much of the workflow context was distributed across…

etlapache-airflowexcerpt only · body stays at the source
From the web
9

From SSH to REST: A Security-Driven Modernization of Slack’s EMR Data Pipelines (opens on the source site)

Excerpt By 2024, Slack’s data platform had accumulated 700+ SSH-based operators orchestrating critical data pipelines. We’re talking daily search indexing that processed terabytes of data, analytics jobs powering business intelligence, the whole shebang. Every single one of these jobs required direct SSH access to production AWS Elastic MapReduce (EMR) clusters. We had a massive security…

uncategorizedairflowexcerpt only · body stays at the source
From the web
10

AN EVENTFUL SUMMER AT STRAVA (opens on the source site)

Hi my name is Bisman and I studied Computer Science at University of California, Santa Barbara. During summer of 2022, I had the most amazing experience working as a Software Engineer Intern on Strava’s Data Platform Team. In the first fews weeks, I learned the tools my team uses and then spent the rest of the time working on my project.TRACKING BAD EVENTSFor my major summer project, I created a data pipeline that pulls user behavior data out of external storage and persists it in our data warehouse. Strava uses a service called Snowplow to collect this user behavior data, like loading a club…

software-engineeringdata-platformsexcerpt only · body stays at the source
From the web
10 shown

Privacy choices

Reading never requires analytics. These choices last 90 days on this browser.

Essential sign-in and security storage always stays on. Read the privacy notice.