Machine learning
We track 61 posts about Machine learning from 27 engineering blogs. Most active: Instacart, Feedzai, Pinterest. Latest post: Oct 9, 2026.
Companies writing about Machine learning
Recent posts
From Activity to Intent: Generating User Journeys with LLMs (opens on the source site)
Lin Zhu | Sr. Staff Machine Learning Engineer; Manan Kalra | Machine Learning Engineer II; Logan Jeon | Sr. Machine Learning Engineer; Ye Liu | Staff Machine Learning Engineer; Xiangyi Chen | Sr. Machine Learning Engineer; Jaewon Yang | Principal Machine Learning Engineer; Jinwen Xu | Manager II, Machine Learning Engineering; Tingting Zhu | Sr. Manager, Engineering; Sudarshan Lamkhede | Director, Machine Learning EngineeringPinterest is built to get inspired and then turn the inspiration into realization — a dinner, a renovation, a wedding, a new skill. That only works if we understand more…
Explainable AI: If You Can’t Evaluate It, Can You Trust It? (opens on the source site)
Feedzai ·
Machine learning models increasingly support decisions that carry real consequences. For example, they help fraud analysts identify suspicious transactions, assist doctors in assessing patient risk, and inform hiring decisions that can shape people’s careers.As these models become more complex, explainability has become a cornerstone of trustworthy AI. The idea is straightforward: if people understand why a model reached a particular prediction, they should be able to make better, more informed decisions.But there is a fundamental question that often goes unasked: How do we know whether an…
Powering AI-led research through simulation (opens on the source site)
Grab ·
The short story Consider a Friday evening. A food order arrives from a mall in the city center. One driver is nearby; another is finishing a drop-off and will be available shortly; a second order from the same mall may or may not appear in the next two minutes. Dispatch the nearby driver now, or hold briefly for a batching opportunity? The decision window is short. A fulfillment marketplace makes these decisions continuously. Each one is small. Across a city, those decisions determine whether your dinner arrives hot and whether a driver’s hour is well spent. And that is one decision. There…
Personalization without user identity (opens on the source site)
Airbnb ·
How Airbnb uses proximity signals to personalize without relying on individual user history.By: Wei Jiang, Bin Xu, Bharathi Thangamani, Weiwei Guo, Sundar Srinivasavaradhan, Tracy Yu, Huiji Gao, Swapnil Ghike, Michael KinotiGreat personalization starts with knowing your user. But what happens when the user is a stranger?A significant share of Airbnb users arrive without a login, without a recent search history, or without any prior booking — especially those landing from paid advertising or organic search. For these users, the ML models that power search ranking, destination recommendations,…
Evolving our calendar assistant Reclaim to be AI-native without starting over (opens on the source site)
Dropbox ·
How we redesigned Reclaim for AI and natural-language requests while preserving the scheduling experience people rely on.
Quiz: Setting Up Python for Machine Learning on Windows (opens on the source site)
Test your understanding of setting up a Python machine learning environment on Windows with Miniconda, Conda environments, packages, and channels.
The guest journey, updated in real time: extending Airbnb’s sequence recommender with Chronon (opens on the source site)
Airbnb ·
How two new Chronon capabilities, Push Mode and NRT Model Transform, allows us to provide more relevant search results instantly as a guest explores, rather than waiting for the next batch run.By: Pengyu Hou, Yuli Han, Daochen Zha, Haozhen Ding, Xin Liu, Sophie Wang, Pallavi Adusumilli, Sherry Li, Henry Saputra, Chun How Tan, Huiji Gao, Yan Zhang, Stephanie Moyerman, Yi Li, and Sanjeev KatariyaA guest’s interaction with Airbnb doesn’t pause to wait for a nightly batch job. Someone might browse a dozen listings on a Tuesday afternoon, run a new search that evening, and expect the next search…
Beyond Two Towers: Launching the 3-Tower Engagement Co-Train Model (Part 2) (opens on the source site)
Authors: Longyu Zhao (Staff Machine Learning Engineer), Gwendolyn Zhao (Staff Machine Learning Engineer), Peng Yan (Senior Machine Learning Engineer), Yuanlu Bai (Senior Machine Learning Engineer), Yuan Wang (Senior Machine Learning Engineer), Yao Cheng (Staff Machine Learning Engineer), Ang Xu (Principal Machine Learning Engineer), Zhaohong Han (Manager II, Ads Lightweight Ranking)IntroductionPreviously¹, we launched the next-generation serving stack for standard ads, which we call Nexus. Nexus decoupled candidate generation from scoring and moved us beyond the classic two-tower-only world,…
Three principles for building a vector platform at Thumbtack (opens on the source site)
Reusing what we already had, treating embeddings as data, and lowering the next team’s costToday, an ML engineer at Thumbtack can stand up production vector search without negotiating database access, building a custom ETL, or writing a query service. The team brings their choice of embedding model, the data, and the query; the platform handles what connects them. It took several iterations to get to this point. In this post we’ll walk through how we got there and the three principles that shaped what we built.A vector database stores high-dimensional numeric arrays (embeddings) and serves…
Agentic Machine Learning Modeling at Instacart (opens on the source site)
Tilman Drerup, Moe Moazzami, Shih-Ting Lin, Greg Reda (and many more)IntroductionAt Instacart, artificial intelligence is fundamentally changing the way our machine learning engineers operate. In a prior blog post, we used one of our teams as a case study to illustrate how the emergence of agents has reshaped what machine learning engineers spend their time on. The post below goes a few levels deeper and zooms in on the machine learning modeling process itself, an area where recent developments in AI-assisted research have opened up exciting new frontiers that we are now actively exploring.…
MAPS: Netflix’s Multimodal Asset Personalization at Scale (opens on the source site)
Netflix ·
By Emma Yanyang Kong, Aditya Deshpande, Asad Abbasi, Bowei Yan, David Fagnan, Ashish Rastogi, Dhaval Patel, Ray ZhangIntroductionThe Netflix experience is a journey of discovery. Every visual cue, from the artwork on a title to the video previews that autoplay while you browse, is there to connect you with a story you will love. We call these visual cues assets, and choosing the right one for each member is a personalization problem of its own. But which image or video preview of Squid Game should we show you? And what do we do right after a title launches, when there’s far too little…
How we think about text classification in the LLM era (opens on the source site)
Medium ·
Why we think LLMs can be useful and why we will not replace all of our models with themContextAt Medium, we have many Machine Learning models that we use to label stories automatically. These affect what stories we recommend to readers.Here’s some examples:a few of our text classification models. All diagrams and charts made by the authorSome Clarifications on our Machine Learning policyBefore we go deep on this project, I just wanted to clarify a few things about how we stand regarding AI in general.Medium has been training internal models with user and post data for a long time now. We…
How we knew COVID was over (and what our models had to unlearn) (opens on the source site)
Airbnb ·
When we retrain, when we rebuild, and when we leave a model alone.By: Harrison KatzA forecast that carries weightThe Forecasting Data Science team at Airbnb produces many of the forecasts the rest of the company plans around: demand, bookings, cancellations, and a range of finer cuts by market and segment, refreshed continuously across thousands of markets. The targets differ, and the models differ, but they have one thing in common: Other teams build on top of them.This means a forecast that is casually wrong is not a clean miss, as it might be in an academic setting. That’s because a small…
Calibrating LLM-Based Population Estimates with Human Validation (opens on the source site)
Indeed ·
Key Idea Human validation is not only for evaluating an LLM. It can also calibrate how the LLM is used as a scalable measurement instrument for population estimation. An LLM can classify thousands of records at low cost, but the proportion it classifies as positive is not necessarily the true proportion in the population. By […]
Grab Bench: Evaluating AI on Grab-shaped production work (opens on the source site)
Grab ·
Introduction What worried us wasn’t the hallucination, it was the subtle plausibility. Answers an engineer could easily read past and accept: a right-looking Structured Query Language (SQL) query, a plausible tool call, an innocent profile update, or a patch that satisfied the surface tests. When we analyzed the row-level failures, a clear pattern emerged: SQL generation: kept the query shape but changed the underlying metric. Tool calling: selected the right tool family but drifted on parameters. Profile updates: cited every event instead of only the evidence that supported the claim. Coding…
How Keras 3 Helped Modernise Expedia Group’s Lodging Ranking Stack (opens on the source site)
Expedia ·
Expedia Group Technology — DataWhat happened when we treated a framework migration as an architecture modernisation — and cut P99 inference latency by two-thirdsSt Paul’s and millennium bridge, LondonExpedia Group™ has always been a market leader in providing personalised search experiences for travellers. As our ranking models evolved, we saw an opportunity not just to migrate to Keras 3, but to modernise the broader stack around it so we can better serve travellers. This led us to rewrite key parts of our pipelines that made model training 30% faster and cut P99 inference latency by…
Correlation Lied to Us: Rethinking Product Impact with Causal Inference (opens on the source site)
OLX ·
IntroductionAt OLX, professional sellers pay for higher-tier packages because they promise more exposure. More visibility, and, in theory, better results. But when we looked at the data, something unexpected happened. In some cases, ads published with premium packages appeared to perform worse than ads using cheaper packages.That raised an uncomfortable question: If higher-tier packages provide more exposure, shouldn’t they consistently perform better?At first glance, there were several possible explanations. Perhaps the extra visibility weren’t creating as much value as we expected. Perhaps…
Crowdsourced taxonomy verification: A feedback-driven framework for refining knowledge graph relationships via online search interactions (opens on the source site)
Grab ·
Introduction The efficacy of semantic search relies on the accuracy of the underlying Knowledge Graph (KG). In high-velocity domains like on-demand food delivery or e-commerce, the catalog of entities like dishes, products, and merchants changes rapidly. Current methods for KG construction and maintenance face three critical challenges: Inaccuracy and hallucination from Large Language Models (LLMs): Automated models often infer relationships based on statistical text co-occurrence rather than semantic reality. For instance, an LLM might incorrectly classify “Pho” as a child of “Italian Noodle…
Agent platform (Part 1): How we help Grab build and run AI agents at scale (opens on the source site)
Grab ·
Part 1: From one support bot to a framework At Grab, AI agents have evolved from interesting team prototypes into production services used every day by millions of merchants, drivers, and consumers. Today, more than 500 services run on our internal agent framework, over 50 Model Context Protocol (MCP) servers are registered on our remote MCP framework, and a single Large Language Model (LLM) gateway fronts every model call across the company, handling billions of tokens each month. None of this was designed up front. It began as the plumbing behind one internal support bot, which then…
Personalizing Airbnb search by learning from the guest journey (opens on the source site)
Airbnb ·
How we built a Transformer-based sequence model that encodes years of guest behavior to surface the right listings at the right time.By: Daochen Zha, Chun How Tan, Xin Liu, Bin Xu, Han Zhao, Xiaowei Liu, Jun Shi, Tracy Yu, Hui Gao, Huiji Gao, Liwei He, Michael Kinoti, Stephanie Moyerman, and Sanjeev KatariyaIntroductionPlanning a trip on Airbnb rarely happens in a single session. A guest searching for a place to stay in San Francisco might browse dozens of listings over several days, leaving behind a trail of views. Typically, over a period of years, that same guest will have accumulated many…
Building a Transformer-Based Category Recommender at Thumbtack (opens on the source site)
A look at compensating for position bias in recommender systems using negative sampling strategiesBy: Andrew Morss, Senior Applied ScientistIntroductionA recommender system is a machine learning model that, given a user and a catalog of items, predicts which items that user is most likely to want. Recommenders set your YouTube playlist, determine what items Amazon suggests for you, push you songs on Spotify and customize your Steam store. If you’re a homeowner, Thumbtack’s recommender systems can suggest home projects for you such as house cleaning or lawn mowing.Thumbtack connects users with…
How Expedia Group Builds AI That Lasts at Scale (opens on the source site)
Expedia ·
Expedia Group Technology — InnovationA framework for how we build, deploy, and evolve AI systems for impact and scalePhoto by Florian Wehde on UnsplashThere’s an important distinction between Artificial Intelligence (AI) that just works today and AI that lasts at scale. Many companies optimize hard for the first one without ever asking whether they’re building the second.Velocity without discipline and strategic direction is a liability, not an asset. The hardest part of building AI at scale isn’t getting a model to work once. It’s building systems that continue to work, scale beyond…
Bootstrap Confidence Intervals for LLM Evaluation (opens on the source site)
Indeed ·
Introduction As Large Language Models (LLMs) move from research prototypes to production systems, the developers of these systems need rigorous performance evaluation. In particular, we need confidence intervals around estimates of system accuracy. However, LLMs introduce a challenge that is unusual for ML systems: they are (operationally) non-deterministic. Even with the temperature set to zero, […]
Variance Reduction Below the Randomization Grain (opens on the source site)
Sergio Camelo, Caitlin Kearns, Matias Cersosimo, and Tilman DrerupAs artificial intelligence increases the velocity of engineering and science teams, experimental throughput is set to become a bottleneck for many product decisions. Many companies can now build faster than they can experiment, with queues of good ideas running the risk of not being tested because of lack of experimental capacity.This problem is particularly severe in marketplaces, where the presence of spillover and cannibalization effects between experimental units requires cluster-level randomization techniques. That…
GenPage: Towards End-to-End Generative Homepage Construction at Netflix (opens on the source site)
Netflix ·
Authors: Lequn Wang, Jiangwei Pan, and Linas BaltrunasFigure 1. Autoregressive homepage generation. GenPage builds a Netflix homepage one row or entity at a time, each one conditioned on what’s already on the page and the user’s context.IntroductionThe Netflix homepage is the first thing users see when they open the app and the primary way they discover content to enjoy. Almost every part of it is personalized, including which rows appear, which entities show up within those rows, and how everything is arranged on the page.Constructing that homepage is a genuinely hard problem. It is not…
Achieving Near-Linear Training Scalability for Pinterest’s Foundation Models (opens on the source site)
Sheng Huang | Software Engineer, AI Platform; Pong Eksombatchai | Machine Learning Engineer, Applied Sciences; Saurabh Vishwas Joshi | Software Engineer, AI Platform; Gaurav Arora | Software Engineer, AI Platform; Karthik Anantha Padmanabhan | Engineering Director, AI PlatformAt Pinterest, foundation models power recommendations for over 600 million monthly active users. Our latest Foundation Model (ACM RecSys 2025) pre-trains on two years of user activity data and is deployed into Home feed and Related Pins ranking, the platform’s two most important recommendation systems. Multi-node…
Distilling Long-Tail User Behavior into Scalable Embeddings for Job Search (opens on the source site)
Indeed ·
Authors : Marsan Ma, Nikhil Lopes, Raj Amrit, Hong Lu, Dipankar Biswas, Trent KyonoLeadership: Iris Wang, Madhu Kurup Recommendation and ranking systems power many of the most important experiences on large internet platforms. Yet the models that run in production are rarely the largest models we can train. They are usually compact, latency-sensitive supervised models […]
From Scoring to Spelling: Rebuilding Ads Retrieval at Instacart (opens on the source site)
Key Contributors: Karuna Ahuja, Marko Avdalovic, Soroush Sobhkhiz, Shrikar Archak, Xiyu Wang, Ji Chao Zhang, Hao YanIntroductionEvery time a user opens Instacart, they see product recommendations: on the retailer home page, in search results, and alongside their cart. Many of these recommendations are sponsored products surfaced by a retrieval model that decides which products to show from a vast ads product catalog. A relevant ad helps users discover products they didn’t know they needed; a less relevant one generates friction.Two years ago, we introduced Contextual Recommendations (CR), a…
When history fails you, borrow from geography (opens on the source site)
Airbnb ·
How Airbnb used sequential geographic recovery signals and prior propagation to generate reliable corridor-level forecasts when local data was scarce.By: Harrison KatzThe problem with unprecedented shocksAlmost every forecasting system is built on the same implicit assumption: the future will resemble the past. You train on historical data, you validate on holdout periods, and you trust that past patterns will at least roughly indicate future performance. When this assumption breaks, the model does not gracefully degrade; it fails confidently. It produces precise, well-calibrated intervals…
Semantic IDs: Product Understanding at Scale (opens on the source site)
Key Contributors: Shrikar Archak, Karuna Ahuja, Soroush Sobhkhiz, Marko Avdalovic, Xiyu Wang, JiChao Zhang, Hao Yan, Chris HartleyIntroductionOperating a grocery catalog at Instacart’s scale means managing millions of products across thousands of categories. Every product is assigned to a category in our hierarchical taxonomy like “Dairy > Cheese > Parmesan”. These categories provide broad classification, but they miss the connections that drive how customers actually shop.For example, a customer is building a cheese board. They’ve added Parmigiano Reggiano, and now they need accompaniments.…
Related topics
This page is generated automatically from the engineering blogs we follow. Every post links to its source, where it was published. See all sources.