How Thumbtack built an AI pipeline that generates, evaluates, and refines marketing content at scale, maintaining human-level quality programmatically.The ChallengeThumbtack connects customers with local service professionals across a wide range of categories and geographies. For many customers, the first entry point to Thumbtack is a landing page. When someone searches Google for “plumber near me” or “duct cleaning Raleigh NC”, these pages are where they land, and they are the user’s first impression of the marketplace for that query. The footprint targeted by this work is roughly 500K such…
How Airbnb’s agent harness transforms unstructured data exploration by encoding scientific methodology into scalable, reproducible, and audit-ready infrastructure.Wren DoughertyAsk a coding agent to analyze 100,000 customer support conversations and within minutes you’ll have a polished taxonomy, precise prevalence numbers, and an executive-ready summary. What you can’t see is the investigation that produced them: the methods it chose, the evidence it weighed, how much to trust it, or whether a second request would agree. All that reaches you is the polish. The model is undeniably…
Clearing the confusion about what Bayesian A/B testing is. The post Why Spotify Is Not Using Bayesian A/B Testing appeared first on Spotify Engineering.
When we retrain, when we rebuild, and when we leave a model alone.By: Harrison KatzA forecast that carries weightThe Forecasting Data Science team at Airbnb produces many of the forecasts the rest of the company plans around: demand, bookings, cancellations, and a range of finer cuts by market and segment, refreshed continuously across thousands of markets. The targets differ, and the models differ, but they have one thing in common: Other teams build on top of them.This means a forecast that is casually wrong is not a clean miss, as it might be in an academic setting. That’s because a small…
TL;DR: LLM predictions can stand in for human outcomes in A/B tests, but only by assumption, not by design.... The post When Can LLMs Replace Humans in A/B Tests? appeared first on Spotify Engineering.
data-scienceexcerpt only · body stays at the source
Key Idea Human validation is not only for evaluating an LLM. It can also calibrate how the LLM is used as a scalable measurement instrument for population estimation. An LLM can classify thousands of records at low cost, but the proportion it classifies as positive is not necessarily the true proportion in the population. By […]
Expedia Group Technology — DataWhat happened when we treated a framework migration as an architecture modernisation — and cut P99 inference latency by two-thirdsSt Paul’s and millennium bridge, LondonExpedia Group™ has always been a market leader in providing personalised search experiences for travellers. As our ranking models evolved, we saw an opportunity not just to migrate to Keras 3, but to modernise the broader stack around it so we can better serve travellers. This led us to rewrite key parts of our pipelines that made model training 30% faster and cut P99 inference latency by…
IntroductionAt OLX, professional sellers pay for higher-tier packages because they promise more exposure. More visibility, and, in theory, better results. But when we looked at the data, something unexpected happened. In some cases, ads published with premium packages appeared to perform worse than ads using cheaper packages.That raised an uncomfortable question: If higher-tier packages provide more exposure, shouldn’t they consistently perform better?At first glance, there were several possible explanations. Perhaps the extra visibility weren’t creating as much value as we expected. Perhaps…
From Day 1 to Production: Building Lyft’s Analytics & Rides Intelligence Assistant as Onboarding ProjectWritten by Sagar Baronia at Lyft.A Different Kind of Day OneMost onboarding journeys follow a familiar arc: orientation sessions, benefits enrollment, setting up your laptop, and gradually finding your footing over the first few weeks. Mine followed that arc too, but with an additional thread running alongside it from the very start.I joined Lyft in March 2026 as a Senior Data Scientist — Algorithm on the Marketing, Business & Ads team, bringing close to a decade of experience in data…
Introduction As Large Language Models (LLMs) move from research prototypes to production systems, the developers of these systems need rigorous performance evaluation. In particular, we need confidence intervals around estimates of system accuracy. However, LLMs introduce a challenge that is unusual for ML systems: they are (operationally) non-deterministic. Even with the temperature set to zero, […]
Benjamin S. KnightScaling Marketplace experiments requires specialized statistical techniques. We examine why standard ordinary least squares regression (OLS) becomes computationally intractable when controlling for high-cardinality categories. We then dive into the underlying math and demonstrate how modern packages — specifically Fixest and Pyfixest — bypass these limitations. We conclude by benchmarking these methods to show their real-world impact on processing speed, memory efficiency, and estimator precision.At Instacart we strive to give our customers access to all the fresh foods and…
Authors : Marsan Ma, Nikhil Lopes, Raj Amrit, Hong Lu, Dipankar Biswas, Trent KyonoLeadership: Iris Wang, Madhu Kurup Recommendation and ranking systems power many of the most important experiences on large internet platforms. Yet the models that run in production are rarely the largest models we can train. They are usually compact, latency-sensitive supervised models […]
How Airbnb used sequential geographic recovery signals and prior propagation to generate reliable corridor-level forecasts when local data was scarce.By: Harrison KatzThe problem with unprecedented shocksAlmost every forecasting system is built on the same implicit assumption: the future will resemble the past. You train on historical data, you validate on holdout periods, and you trust that past patterns will at least roughly indicate future performance. When this assumption breaks, the model does not gracefully degrade; it fails confidently. It produces precise, well-calibrated intervals…
aitechnologyexcerpt only · body stays at the source
A practical look at how Thumbtack navigates evaluation for emerging AI experiences and what we’ve learned along the way.By: Shishir Dash, Director of Applied Science & Teja Venkat Kolli, Senior Applied ScientistEvaluating AI at ScaleIntroductionAI is reshaping how people interact with products, and Thumbtack is no exception. We’re introducing AI into more aspects of our customer and local service professional (pro) experiences — from helping customers articulate what they need, to generating helpful summaries, to offering clearer explanations of how pros may fit those needs.But evaluating…
Text classification is often done through fine-tuning of a pretrained foundation model with domain-specific data. In FreeAgent we use transformer based models to automatically classify incoming bank transactions. Specifically we use a DistilBERT model that is fine-tuned on hundreds of millions of bank transactions with customer-labelled accounting categories. The model inputs are currently text-based, built from a combination of bank transaction descriptions and amounts. In this post we describe an approach to fine-tuning the DistilBERT model and training the classifier including the…
data---mlaiexcerpt only · body stays at the source
At Lyft, understanding how riders go through our user experience is fundamental to operating a healthy marketplace. Specifically, it is important to have a robust model determining if a rider will actually request a ride after entering a destination and viewing a price and ETA. Accurately predicting this decision, that we call conversion, informs countless decisions across our platform. Whether it is to better balance supply and demand, improve user experiences, optimize recommendations and advertisement, understand long-term engagement, decide how to distribute coupons… rider conversion…
Expedia Group Technology — DataWorkload‑aware routing for TrinoPhoto by Joseph Barrientos on UnsplashTrino — a fork of PrestoSQL — is a powerful tool in modern data analytics, enabling organizations to query large datasets quickly and efficiently. As a distributed SQL query engine, Trino provides fast, scalable insights without requiring data relocation. While Trino is robust on its own, its capabilities are further enhanced when paired with a Gateway, which introduces features such as query routing, strong security, and streamlined cluster management.A brief overviewThe Gateway project…
One of the challenges of working with LLMs is getting them to respond with a consistent format, such as a given JSON schema. Anyone who has tried to solve this issue with prompt engineering knows how frustrating it can be. You add a ‘MUST’ here and an ‘always return JSON’ there, but still the output […]
data---mlaiexcerpt only · body stays at the source
Written by Rohan Varshney, with support from Devon Mittow & Janice Lee.This article expands upon a presentation from the Feature Store Summit 2025, which can be viewed in full here. There is also another video available on the evolution of Lyft’s Feature Store from DE4AI 2024.Introduction and Core PurposeLyft’s Feature Store stands as a core infrastructural pillar within its Data Platform organization, designed to optimize the management and deployment of Machine Learning (ML) features at massive scale. Its primary objective is to centralize feature engineering efforts, guaranteeing…
Introduction At Indeed, our mission is to help people get jobs. We connect job seekers with their next career opportunities and assist employers in finding the ideal candidates. This makes matching a fundamental problem in the products we develop. The Ranking Models team is responsible for building Machine Learning models that drive matching between job […]
How real-time algorithms and causal ML transformed our digital subscription funnel from static paywalls to dynamic, millisecond decision-makingIllustration by Mathieu LabrecqueThe New York Times became a subscription-first news and lifestyle service with the launch of its paywall in 2011. Since then, our subscription strategy has evolved substantially. Initially, users could access a limited number of free articles per month before they encountered the paywall. In 2019, we began personalizing this number using a Machine Learning (ML) model — The Dynamic Meter. In the past few years, we have…
Data scientists use different Jupyter notebooks every day — ranging from disposable ones for quick tasks to those shareable with clients. Over time, more and more notebooks accumulate, making it increasingly difficult to reuse them in whole or in part. To mitigate this problem and make the most relevant pieces of code quickly accessible to every data scientist, we developed JupyterLab Snippets at Feedzai — our take on leveraging code snippets directly on JupyterLab.JupyterLab is a computational notebook platform that enables us to carry on data science work (and beyond) via notebooks. These…
In recent years, online shopping has surged, revolutionizing how people purchase products and services. E-commerce’s convenience has reshaped consumer behaviour and the retail landscape. Unlike traditional stores, online shoppers often face sizing challenges, leading to hesitancy and missed sales. Myntra has been a pioneer in addressing size and fit challenges in India, leading the way with innovative solutions that have significantly enhanced the shopping experience.Building on its leadership in this space, Myntra’s latest initiatives take these solutions to the next level, offering even…
A clustering-based approach to create deep learning datasets in a dayIntroductionUnderstanding what’s happening in an image is both an important task, as well as a costly one. In the last few years, the field of computer vision has greatly accelerated due to the advances in neural networks. At Bumble Inc., we see potential value in computer vision for a variety of use cases, such as improving the safety of our platform and providing our members with a better user experience.The most common way to train these neural networks is by showing it many images with the corresponding label.…
How we got started with Mesos: Pain-points in a hybrid environment and surviving the transition The post Mesos at Opentable appeared first on OpenTable Technology Blog.
We can build expert knowledge of cities with our corpus of unstructured reviews The post Using data science to create a dining expert appeared first on OpenTable Technology Blog.
Using Word Mover’s Distance to find similar reviews across our corpus — even without similar keywords The post Navigating themes in restaurant reviews with Word Mover’s Distance appeared first on OpenTable Technology Blog.
The promise of data science for the restaurant industry The post Mining data to power great dining experiences appeared first on OpenTable Technology Blog.
data-scienceexcerpt only · body stays at the source
Diner reviews are a powerful tool in making better recommendations The post Mining a treasure trove of textual data appeared first on OpenTable Technology Blog.
Optional Google Analytics helps us understand visits. Microsoft Clarity records masked interactions to improve the site. Optional tools stay off unless you choose them. Privacy details.