560 blogs tracked4,950 posts indexed

#spark

6 posts · 4 companies · newest first

1

Partition Finalization in Pinterest’s Next-Generation DB Ingestion Framework (opens on the source site)

Qianrui Zhang | Sr Software Engineer, Logging PlatformKanchi Masalia | Software Engineer II, Stream Processing PlatformLiang Mou | Sr Staff Software Engineer, Logging PlatformYi Pan | Principal Engineer, Agent PlatformIntroductionThis is the third post in our series on Pinterest’s next-generation database ingestion framework. Part 1 introduced the DB ingestion framework built on Kafka, Flink, Spark, and Iceberg, and Part 2 covered automated schema evolution. This post tackles another challenge in migrating downstream customers to the new ingestion framework: knowing when data is complete…

pinteresticebergsexcerpt only · body stays at the source
From the web
2

Scaling Grab's Data Lake: Our journey to Apache Iceberg adoption (opens on the source site)

Introduction: The evolution of Grab’s Data Lake At Grab’s scale, managing petabytes of data across billions of S3 objects demands more than a storage layer. It demands a robust architectural primitive that supports the high-concurrency needs of a modern “Lakehouse.” Our goal is full storage-compute separation, leveraging S3 as an elastic foundation for both near-real-time metrics and large-scale batch transformations. For years, the vast majority of our tables were Hive Parquet, managed through the Hive Metastore with a directory-based layout. This model served us well, but as data volume…

datadatabaseexcerpt only · body stays at the source
From the web
3

Automated Schema Evolution in Pinterest’s Next-Generation DB Ingestion Framework (opens on the source site)

Yisheng Zhou | Software Engineer IILiang Mou | Sr Staff Software EngineerGabriel Raphael Garcia Montoya | Staff Software EngineerIstvan Podor | Staff Software EngineerIntroductionIn the first post of this series, we introduced Pinterest’s next-generation CDC-based ingestion platform built on Kafka, Flink, Spark, and Iceberg. In production, upstream schemas are constantly evolving, and in a distributed CDC pipeline, schema is not just metadata — it is a cross-system contract spanning ingestion, transformation, storage, and historical backfill. A schema change that is not handled carefully can…

schema-evolutionengineeringexcerpt only · body stays at the source
From the web
4

Rain: A key-value store for Strava’s scale (opens on the source site)

Much of our heatmaps are built on batch data outputs stored in RainAt Strava, we love maps — some of our most loved features are nestled on map surfaces. My team, the Geo team, is focused on building and improving these products. On the Geo and Metro teams, we tend to work with large datasets: aggregations of open source map data via OpenStreetMaps, GPS data points from uploaded activities, third-party datasets for properties like elevation, and beyond. This aggregated dataset eventually turns into Geo features we know and love, like the global heatmap, Strava Metro, the routing product,…

mapssparkexcerpt only · body stays at the source
From the web
6 shown

Privacy choices

Reading never requires analytics. These choices last 90 days on this browser.

Essential sign-in and security storage always stays on. Read the privacy notice.