560 blogs tracked4,950 posts indexed

#data

18 posts · 11 companies · newest first

Curated view of this subject:Topic: Data156
1

Beyond synthetic testing: Capturing and replaying real database workloads at Airbnb (opens on the source site)

How we capture real production database traffic at Airbnb and replay it offline to load-test, plan capacity, and de-risk upgrades.By: Zuofei Wang, Erluo LiIntroductionAt Airbnb, MySQL-compatible databases are a critical backbone of our online database infrastructure: a fleet of hundreds of clusters supporting thousands of use cases at millions of queries per second (QPS). Operating databases at scale brings hard problems, including sizing clusters for future growth, keeping behavior consistent across version upgrades and migrations, and reproducing production incidents well enough to debug…

datatechnologyexcerpt only · body stays at the source
From the web
2

Mixpanel Is Easy to Install. Trusting the Data Takes Work. (opens on the source site)

Mixpanel makes it easy to start collecting events and building dashboards. The harder part is knowing whether those dashboards represent real product usage. A report can look polished while counting employees, automated browser tests, web scrapers, duplicate events, or actions that users attempted but never completed. It can also be technically correct while answering a […] The post Mixpanel Is Easy to Install. Trusting the Data Takes Work. appeared first on Atomic Spin.

developer-toolsdataexcerpt only · body stays at the source
From the web
4

Data Mesh at Grab (Part III): Operationalizing data reliability with automated DPIs (opens on the source site)

Introduction In the first two parts of this series, we described how Grab approaches data mesh through the Signals Marketplace: a way for teams to publish, discover, and reuse trusted data products across domains. Part II introduced the foundational tools behind certification: Hubble for metadata and ownership, Genchi for data quality observability, and the Data Contract Registry for explicit producer-consumer guarantees. Certification is the starting point for a trusted data marketplace. It gives downstream consumers confidence in an asset’s ownership, documentation, lineage, and quality…

datadatabaseexcerpt only · body stays at the source
From the web
5

Project Lighthouse — Part 3: Introducing project-lighthouse-anonymize (opens on the source site)

Project Lighthouse — Part 3: Introducing project-lighthouse-anonymizeThe data in Project Lighthouse is powered by privacy-preserving anonymization code. We’ve put this code into open source, and published two new technical papers detailing the scalable algorithms and data quality frameworks behind it.By: Adam BloomstonIntroductionIn 2020, we launched Project Lighthouse, which we developed in partnership with leading civil rights and privacy organizations. As our 2020 announcement details, Project Lighthouse enables us to measure potential disparities in user experiences. This work uses…

engineeringopen-sourceexcerpt only · body stays at the source
From the web
6

How we knew COVID was over (and what our models had to unlearn) (opens on the source site)

When we retrain, when we rebuild, and when we leave a model alone.By: Harrison KatzA forecast that carries weightThe Forecasting Data Science team at Airbnb produces many of the forecasts the rest of the company plans around: demand, bookings, cancellations, and a range of finer cuts by market and segment, refreshed continuously across thousands of markets. The targets differ, and the models differ, but they have one thing in common: Other teams build on top of them.This means a forecast that is casually wrong is not a clean miss, as it might be in an academic setting. That’s because a small…

data-sciencedata-modelingexcerpt only · body stays at the source
From the web
7

AI adoption starts with truth (opens on the source site)

The semantic layer is the foundation AI adoption is limited by trust. A user who gets burned by a confidently wrong answer will double-check the next one, eventually routing consequential work around the system entirely. Once that happens, AI remains a tool at the edges rather than infrastructure at the center… useful, but never trusted with the workflows where its value compounds. Before a company can benefit from more capable agents, those agents need a reliable way to know what the company considers true. A semantic layer tells an agent which tables are sources of truth and how they…

aiengineeringexcerpt only · body stays at the source
From the web
9

Crowdsourced taxonomy verification: A feedback-driven framework for refining knowledge graph relationships via online search interactions (opens on the source site)

Introduction The efficacy of semantic search relies on the accuracy of the underlying Knowledge Graph (KG). In high-velocity domains like on-demand food delivery or e-commerce, the catalog of entities like dishes, products, and merchants changes rapidly. Current methods for KG construction and maintenance face three critical challenges: Inaccuracy and hallucination from Large Language Models (LLMs): Automated models often infer relationships based on statistical text co-occurrence rather than semantic reality. For instance, an LLM might incorrectly classify “Pho” as a child of “Italian Noodle…

engineeringdataexcerpt only · body stays at the source
From the web
10

Blinkit at OpenSearchCon India 2026 (opens on the source site)

A few weeks ago, we spoke at OpenSearchCon India about how we make search fast and reliable at scale at Blinkit.I’m Harshit, from the Search Engineering team at Blinkit, and this year we got to speak at OpenSearchCon India at the Jio World Convention Centre in Mumbai. In this post, I’ll walk you through our talk, the conversations that followed, and what our team took away from being part of the OpenSearch community in a more active way this year.First Day : Keynote by BlinkitThe highlight of this year’s conference for our team was the opportunity to deliver a keynote session at…

software-engineeringopensearchexcerpt only · body stays at the source
From the web
11

How Expedia Group Builds AI That Lasts at Scale (opens on the source site)

Expedia Group Technology — InnovationA framework for how we build, deploy, and evolve AI systems for impact and scalePhoto by Florian Wehde on UnsplashThere’s an important distinction between Artificial Intelligence (AI) that just works today and AI that lasts at scale. Many companies optimize hard for the first one without ever asking whether they’re building the second.Velocity without discipline and strategic direction is a liability, not an asset. The hardest part of building AI at scale isn’t getting a model to work once. It’s building systems that continue to work, scale beyond…

datainnovationexcerpt only · body stays at the source
From the web
12

Scaling Grab's Data Lake: Our journey to Apache Iceberg adoption (opens on the source site)

Introduction: The evolution of Grab’s Data Lake At Grab’s scale, managing petabytes of data across billions of S3 objects demands more than a storage layer. It demands a robust architectural primitive that supports the high-concurrency needs of a modern “Lakehouse.” Our goal is full storage-compute separation, leveraging S3 as an elastic foundation for both near-real-time metrics and large-scale batch transformations. For years, the vast majority of our tables were Hive Parquet, managed through the Hive Metastore with a directory-based layout. This model served us well, but as data volume…

datadatabaseexcerpt only · body stays at the source
From the web
14

Scaling beyond one: How Airbnb evolved its data architecture for a multi-product world (opens on the source site)

How Airbnb’s data engineers and analytics engineers built a consistent and flexible data modeling framework to support the expansion into Homes, Experiences, and Services.By: Patrick Lam, Namrata Lamba, Jamie StoberWith the May 2025 Summer Release, Airbnb redesigned its app, relaunched Experiences, and debuted Services, pushing us beyond our traditional Homes focus. For the data teams, this meant rapidly evolving a decade-old infrastructure to integrate two brand-new product pillars. Our data engineers and analytics engineers rose to the challenge by building a consistent and flexible…

data-engineeringdataexcerpt only · body stays at the source
From the web
15

Structured outputs with Pydantic AI (opens on the source site)

One of the challenges of working with LLMs is getting them to respond with a consistent format, such as a given JSON schema. Anyone who has tried to solve this issue with prompt engineering knows how frustrating it can be. You add a ‘MUST’ here and an ‘always return JSON’ there, but still the output […]

data---mlaiexcerpt only · body stays at the source
From the web
16

Working at the intersection of data and AI with Kenza Boulisfane (opens on the source site)

Kenza Boulisfane, Software Engineer at Thumbtack, works at the intersection of data, AI, and real business impact. In this Q&A, she shares how she’s building an AI-powered Marketing Analytics Agent designed to make complex marketing data accessible to everyone. She also reflects on team culture, technical challenges, and why diverse perspectives make engineering stronger.What are you currently working on?​​I’m working on building a Marketing Analytics Agent. It’s an AI-powered marketing expert that provides companywide support, regardless of technical background. The idea is simple: marketing…

careerssoftware-developmentexcerpt only · body stays at the source
From the web
17

Scaling Subscriptions at The New York Times with Real-Time Causal Machine Learning (opens on the source site)

How real-time algorithms and causal ML transformed our digital subscription funnel from static paywalls to dynamic, millisecond decision-makingIllustration by Mathieu LabrecqueThe New York Times became a subscription-first news and lifestyle service with the launch of its paywall in 2011. Since then, our subscription strategy has evolved substantially. Initially, users could access a limited number of free articles per month before they encountered the paywall. In 2019, we began personalizing this number using a Machine Learning (ML) model — The Dynamic Meter. In the past few years, we have…

data-sciencedataexcerpt only · body stays at the source
From the web
18

Make your reports faster : Beginner’s guide to Tableau Optimization (opens on the source site)

Make your reports faster: A beginner’s guide to Tableau OptimisationIn today’s world, given the pace at which data operates, we need a tool that can help us to generate reports faster and bring out insights within milliseconds. In order to solve this challenge, several companies have started utilising a few Business Intelligence (BI) tools such as Tableau/Power BI/Superset/Looker/Qlikview, etc. We at Blinkit have also moved away from the traditional way of reporting via spreadsheets to a more scalable and robust tool — Tableau.Earlier, we had no single source of truth for metrics; it took a…

tableaudataexcerpt only · body stays at the source
From the web
18 shown

Privacy choices

Reading never requires analytics. These choices last 90 days on this browser.

Essential sign-in and security storage always stays on. Read the privacy notice.