560 blogs tracked4,950 posts indexed

#big-data

8 posts · 7 companies · newest first

2

How INTEGER and INT Produced Different Schemas in Debezium (opens on the source site)

A Debezium investigation: how MySQL synonyms INT and INTEGER were treated as different types.One of our CDC pipelines started failing with a ClassCastException on every batch.java.lang.ClassCastException: class java.lang.Integer cannot be cast to class java.lang.LongThe pipeline was producing an Integer, but the downstream consumer expected a Long. Every run failed in the same way, which pointed us toward a schema mismatch rather than an issue with individual records.Background: how CDC worksTo see why a mismatch like that can hide for years, it helps to know how CDC actually works. Most…

big-datadebeziumexcerpt only · body stays at the source
From the web
3

Using LLMs to Analyze Spark SQL Plans: A Practical Approach to Debugging Long-Running Jobs (opens on the source site)

Expedia Group Technology — InnovationUsing large language models to reveal bottlenecks in Spark SQL execution plansPhoto by Luis del RíoIf you’ve ever stared at a 300-plus-node physical plan at 2 a.m. trying to spot a missing broadcast or one cursed skewed partition, this is for you.Spark makes it deceptively easy to write complex SQL that looks correct but quietly turns into a performance and cost problem at scale. A query that runs fine on day one can slow to a crawl as data grows, joins get wider, and aggregations become more nested. Suddenly, jobs take hours instead of minutes, clusters…

apache-sparkbig-dataexcerpt only · body stays at the source
From the web
4

Distilling Long-Tail User Behavior into Scalable Embeddings for Job Search (opens on the source site)

Authors : Marsan Ma, Nikhil Lopes, Raj Amrit, Hong Lu, Dipankar Biswas, Trent KyonoLeadership: Iris Wang, Madhu Kurup Recommendation and ranking systems power many of the most important experiences on large internet platforms. Yet the models that run in production are rarely the largest models we can train. They are usually compact, latency-sensitive supervised models […]

big-datadata-scienceexcerpt only · body stays at the source
From the web
5

How a Deadlock Froze Blinkit’s Supply Chain (opens on the source site)

A silent deadlock in our query engine was stalling inventory replenishment jobs with no error, no crash — just infinite waiting. This is the story of how we found it, traced it to an open-source bug, and fixed it upstream.TL;DRTrino’s Hudi connector used a single thread pool for both producing file splits and signalling when there was room for more. Under load, every thread ended up waiting for a signal that had no thread left to run it. The fix was to switch the producer side to a cooperative scheduling pattern: yield the thread when the buffer is full, and resume when space opens.Our…

blinkittrinosexcerpt only · body stays at the source
From the web
6

From SSH to REST: A Security-Driven Modernization of Slack’s EMR Data Pipelines (opens on the source site)

Excerpt By 2024, Slack’s data platform had accumulated 700+ SSH-based operators orchestrating critical data pipelines. We’re talking daily search indexing that processed terabytes of data, analytics jobs powering business intelligence, the whole shebang. Every single one of these jobs required direct SSH access to production AWS Elastic MapReduce (EMR) clusters. We had a massive security…

uncategorizedairflowexcerpt only · body stays at the source
From the web
7

How Data Powers Agent Productivity (opens on the source site)

As a data engineer, I used to see metrics as just numbers on a dashboard — until I realized they’re the lens through which customers view and run their operations. In customer support, for example, agent productivity metrics aren’t just figures, they’re actionable insights that drive efficiency, shape staffing decisions, and directly impact customer satisfaction.These aren’t just charts — they help customers understand the value we provide, how well things are working, and what decisions to make next. Realizing this changed how I think about building analytics.➡️💡The Question That Shifted…

apache-sparkanalyticsexcerpt only · body stays at the source
From the web
8

Leveraging Spark 3 and NVIDIA’s GPUs to Reduce Cloud Cost by up to 70% for Big Data Pipelines (opens on the source site)

By Ilay Chen and Tomer AkiravAt PayPal, hundreds of thousands of Apache Spark jobs run on an hourly basis, processing petabytes of data and requiring a high volume of resources. To handle the growth of machine learning solutions, PayPal requires scalable environments, cost awareness and constant innovation. This blog explains how Apache Spark 3 and GPUs can help enterprises potentially reduce Apache Spark’s jobs cloud costs by up to 70% for big data processing and AI applications.Our journey will begin with a brief introduction of Spark RAPIDS — Apache Spark’s accelerator that leverages GPUs…

cloud-computinggpuexcerpt only · body stays at the source
From the web
8 shown

Privacy choices

Reading never requires analytics. These choices last 90 days on this browser.

Essential sign-in and security storage always stays on. Read the privacy notice.