How we capture real production database traffic at Airbnb and replay it offline to load-test, plan capacity, and de-risk upgrades.By: Zuofei Wang, Erluo LiIntroductionAt Airbnb, MySQL-compatible databases are a critical backbone of our online database infrastructure: a fleet of hundreds of clusters supporting thousands of use cases at millions of queries per second (QPS). Operating databases at scale brings hard problems, including sizing clusters for future growth, keeping behavior consistent across version upgrades and migrations, and reproducing production incidents well enough to debug…
Mixpanel makes it easy to start collecting events and building dashboards. The harder part is knowing whether those dashboards represent real product usage. A report can look polished while counting employees, automated browser tests, web scrapers, duplicate events, or actions that users attempted but never completed. It can also be technically correct while answering a […] The post Mixpanel Is Easy to Install. Trusting the Data Takes Work. appeared first on Atomic Spin.
Clearing the confusion about what Bayesian A/B testing is. The post Why Spotify Is Not Using Bayesian A/B Testing appeared first on Spotify Engineering.
Introduction In the first two parts of this series, we described how Grab approaches data mesh through the Signals Marketplace: a way for teams to publish, discover, and reuse trusted data products across domains. Part II introduced the foundational tools behind certification: Hubble for metadata and ownership, Genchi for data quality observability, and the Data Contract Registry for explicit producer-consumer guarantees. Certification is the starting point for a trusted data marketplace. It gives downstream consumers confidence in an asset’s ownership, documentation, lineage, and quality…
datadatabaseexcerpt only · body stays at the source
Project Lighthouse — Part 3: Introducing project-lighthouse-anonymizeThe data in Project Lighthouse is powered by privacy-preserving anonymization code. We’ve put this code into open source, and published two new technical papers detailing the scalable algorithms and data quality frameworks behind it.By: Adam BloomstonIntroductionIn 2020, we launched Project Lighthouse, which we developed in partnership with leading civil rights and privacy organizations. As our 2020 announcement details, Project Lighthouse enables us to measure potential disparities in user experiences. This work uses…
When we retrain, when we rebuild, and when we leave a model alone.By: Harrison KatzA forecast that carries weightThe Forecasting Data Science team at Airbnb produces many of the forecasts the rest of the company plans around: demand, bookings, cancellations, and a range of finer cuts by market and segment, refreshed continuously across thousands of markets. The targets differ, and the models differ, but they have one thing in common: Other teams build on top of them.This means a forecast that is casually wrong is not a clean miss, as it might be in an academic setting. That’s because a small…
The semantic layer is the foundation AI adoption is limited by trust. A user who gets burned by a confidently wrong answer will double-check the next one, eventually routing consequential work around the system entirely. Once that happens, AI remains a tool at the edges rather than infrastructure at the center… useful, but never trusted with the workflows where its value compounds. Before a company can benefit from more capable agents, those agents need a reliable way to know what the company considers true. A semantic layer tells an agent which tables are sources of truth and how they…
aiengineeringexcerpt only · body stays at the source
Your semantic layer is a risk mitigation strategy. Not risk in the abstract, compliance-framework sense, but the practical, operational risk that quietly drains organizations every day.
oreillydataexcerpt only · body stays at the source
Introduction The efficacy of semantic search relies on the accuracy of the underlying Knowledge Graph (KG). In high-velocity domains like on-demand food delivery or e-commerce, the catalog of entities like dishes, products, and merchants changes rapidly. Current methods for KG construction and maintenance face three critical challenges: Inaccuracy and hallucination from Large Language Models (LLMs): Automated models often infer relationships based on statistical text co-occurrence rather than semantic reality. For instance, an LLM might incorrectly classify “Pho” as a child of “Italian Noodle…
A few weeks ago, we spoke at OpenSearchCon India about how we make search fast and reliable at scale at Blinkit.I’m Harshit, from the Search Engineering team at Blinkit, and this year we got to speak at OpenSearchCon India at the Jio World Convention Centre in Mumbai. In this post, I’ll walk you through our talk, the conversations that followed, and what our team took away from being part of the OpenSearch community in a more active way this year.First Day : Keynote by BlinkitThe highlight of this year’s conference for our team was the opportunity to deliver a keynote session at…
Expedia Group Technology — InnovationA framework for how we build, deploy, and evolve AI systems for impact and scalePhoto by Florian Wehde on UnsplashThere’s an important distinction between Artificial Intelligence (AI) that just works today and AI that lasts at scale. Many companies optimize hard for the first one without ever asking whether they’re building the second.Velocity without discipline and strategic direction is a liability, not an asset. The hardest part of building AI at scale isn’t getting a model to work once. It’s building systems that continue to work, scale beyond…
Introduction: The evolution of Grab’s Data Lake At Grab’s scale, managing petabytes of data across billions of S3 objects demands more than a storage layer. It demands a robust architectural primitive that supports the high-concurrency needs of a modern “Lakehouse.” Our goal is full storage-compute separation, leveraging S3 as an elastic foundation for both near-real-time metrics and large-scale batch transformations. For years, the vast majority of our tables were Hive Parquet, managed through the Hive Metastore with a directory-based layout. This model served us well, but as data volume…
datadatabaseexcerpt only · body stays at the source
At Spotify, data problems used to follow a specific pattern. You'd look for the relevant dashboard, there... The post Encoding Your Domain Expert: The Context Layer Behind Spotify's Data Assistant appeared first on Spotify Engineering.
dataplatformexcerpt only · body stays at the source
How Airbnb’s data engineers and analytics engineers built a consistent and flexible data modeling framework to support the expansion into Homes, Experiences, and Services.By: Patrick Lam, Namrata Lamba, Jamie StoberWith the May 2025 Summer Release, Airbnb redesigned its app, relaunched Experiences, and debuted Services, pushing us beyond our traditional Homes focus. For the data teams, this meant rapidly evolving a decade-old infrastructure to integrate two brand-new product pillars. Our data engineers and analytics engineers rose to the challenge by building a consistent and flexible…
One of the challenges of working with LLMs is getting them to respond with a consistent format, such as a given JSON schema. Anyone who has tried to solve this issue with prompt engineering knows how frustrating it can be. You add a ‘MUST’ here and an ‘always return JSON’ there, but still the output […]
data---mlaiexcerpt only · body stays at the source
Kenza Boulisfane, Software Engineer at Thumbtack, works at the intersection of data, AI, and real business impact. In this Q&A, she shares how she’s building an AI-powered Marketing Analytics Agent designed to make complex marketing data accessible to everyone. She also reflects on team culture, technical challenges, and why diverse perspectives make engineering stronger.What are you currently working on?I’m working on building a Marketing Analytics Agent. It’s an AI-powered marketing expert that provides companywide support, regardless of technical background. The idea is simple: marketing…
How real-time algorithms and causal ML transformed our digital subscription funnel from static paywalls to dynamic, millisecond decision-makingIllustration by Mathieu LabrecqueThe New York Times became a subscription-first news and lifestyle service with the launch of its paywall in 2011. Since then, our subscription strategy has evolved substantially. Initially, users could access a limited number of free articles per month before they encountered the paywall. In 2019, we began personalizing this number using a Machine Learning (ML) model — The Dynamic Meter. In the past few years, we have…
Make your reports faster: A beginner’s guide to Tableau OptimisationIn today’s world, given the pace at which data operates, we need a tool that can help us to generate reports faster and bring out insights within milliseconds. In order to solve this challenge, several companies have started utilising a few Business Intelligence (BI) tools such as Tableau/Power BI/Superset/Looker/Qlikview, etc. We at Blinkit have also moved away from the traditional way of reporting via spreadsheets to a more scalable and robust tool — Tableau.Earlier, we had no single source of truth for metrics; it took a…
tableaudataexcerpt only · body stays at the source
From the web
18 shown
Your visit, your choice.
Optional Google Analytics helps us understand visits. Microsoft Clarity records masked interactions to improve the site. Optional tools stay off unless you choose them. Privacy details.