Lin Zhu | Sr. Staff Machine Learning Engineer; Manan Kalra | Machine Learning Engineer II; Logan Jeon | Sr. Machine Learning Engineer; Ye Liu | Staff Machine Learning Engineer; Xiangyi Chen | Sr. Machine Learning Engineer; Jaewon Yang | Principal Machine Learning Engineer; Jinwen Xu | Manager II, Machine Learning Engineering; Tingting Zhu | Sr. Manager, Engineering; Sudarshan Lamkhede | Director, Machine Learning EngineeringPinterest is built to get inspired and then turn the inspiration into realization — a dinner, a renovation, a wedding, a new skill. That only works if we understand more…
Writes are accelerating, and this growth can add stress to systems all repos depend on. Here's our approach to building architecture that can scale. The post Building Git infrastructure for agent-scale development appeared first on The GitHub Blog.
How we capture real production database traffic at Airbnb and replay it offline to load-test, plan capacity, and de-risk upgrades.By: Zuofei Wang, Erluo LiIntroductionAt Airbnb, MySQL-compatible databases are a critical backbone of our online database infrastructure: a fleet of hundreds of clusters supporting thousands of use cases at millions of queries per second (QPS). Operating databases at scale brings hard problems, including sizing clusters for future growth, keeping behavior consistent across version upgrades and migrations, and reproducing production incidents well enough to debug…
How Airbnb uses proximity signals to personalize without relying on individual user history.By: Wei Jiang, Bin Xu, Bharathi Thangamani, Weiwei Guo, Sundar Srinivasavaradhan, Tracy Yu, Huiji Gao, Swapnil Ghike, Michael KinotiGreat personalization starts with knowing your user. But what happens when the user is a stranger?A significant share of Airbnb users arrive without a login, without a recent search history, or without any prior booking — especially those landing from paid advertising or organic search. For these users, the ML models that power search ranking, destination recommendations,…
Authors: Bowen Zhou | Staff Software Engineer; Shan Gao | Senior Software Engineer; Jingwen Hu | Software Engineer II; Wenjiang Chu | Staff Software EngineerThe Billion-Embedding ChallengeAt Pinterest, the “signal” is our lifeblood. Whether it’s a home decor enthusiast finding the perfect rug or a fashion seeker discovering a new aesthetic, our discovery engine relies on understanding deep semantic relationships to help our users find inspirations. Over the last few years, the explosive growth of embedding-based retrieval has fundamentally transformed how we surface these signals — and at the…
John Grass | Sr. Manager, EngineeringA Fundamental TransformationAn AI team is fundamentally more than just a group whose members incorporate AI tools into their existing workflows. The journey to becoming an AI team necessitates a fundamental and comprehensive paradigm shift in how the team defines ownership, engages in strategic planning, and, most critically, executes on its core goals and objectives. This transformation is not merely an addition of new technology; it is a restructuring of the team’s operating model, philosophy, and individual roles.Becoming an AI team requires a holistic…
Devin Kreuzer | Sr. Machine Learning Engineer; Yichi Wang | Machine Learning Engineer I; Sujan Reddy Ale | Machine Learning Engineer I; Zelun Wang | Sr. Machine Learning Engineer; Hongtao Lin | Sr. Machine Learning Engineer; Piyush Maheshwari | Staff Machine Learning EngineerPinterest home feed candidate generation is a large-scale User-to-Pin retrieval problem. A common approach is a two-tower model: a user tower encodes the user, an item tower encodes candidate Pins, and approximate nearest neighbor search retrieves Pins close to the user embedding. But Pinterest users often have multiple…
Rebuilding login and signup surfaced product insights, not just technical challenges. Here’s how we designed Flexible Authentication at the intersection of product intuition and technical architecture.By: Jose Santos, Mike BarryFor Airbnb, logins at irregular intervals are normal. A guest books a trip in January and may not open the app again until summer. A host checks back only when a reservation comes in, and may be busy with other activities when it does. For a two-sided marketplace where a failed login means a lost booking, and lost revenue for both the guest and the host, long gaps…
Companies like Spotify need vast quantities of data accessible at low latency for online services and,... The post Indexing the Data Lake for Online Point Queries appeared first on Spotify Engineering.
Training an LLM is the easy part. The hard part is designing experiments and evaluations that you can trust enough to know whether the new model is actually an improvement.By: Baharak SaberidokhtIntroductionShipping a production LLM system means iterating fast on improvements to something that is, by construction, non-deterministic. Models drift, judges disagree with themselves, references regenerate as different strings, and bugs may persist until the next release, because retraining takes weeks. Most of this friction comes from infrastructure challenges, not model quality, and the fixes…
technologyaiexcerpt only · body stays at the source
Introduction Our Android team uses Gradle as our build tool of choice. Gradle offers lots of options for tuning its resource consumption, giving engineers an opportunity to optimize performance for running tasks on well-known hardware. In this post, we’ll explore how the team tuned our Gradle setup for Amazon’s m7i.8xlarge EC2 instances. General concepts Before... Read more
How Airbnb built a Kubernetes sidecar to deliver dynamic configuration reliably at scale.By: Bo Teng, Cosmo Qiu, Siyuan Zhou, Ankur Soni, Xin Huang, Willis HarveyIntroductionIn our previous post, we explored Airbnb’s dynamic configuration system, Sitar, with a focus on service architecture and configuration change safety. Now for the harder question: once a config change is committed, which happens several times each minute, how does it actually reach the thousands of Airbnb’s service instances reliably, quickly, and without redeploying the services?This post describes sitar agent: a…
How Airbnb shifts from PaaS to an internal knowledge graph infrastructure at scale.By: Lucen Zhao, Shukun Yang, Ashish JainKnowledge graphs offer a natural and powerful way to represent relationships between entities. Many real-world systems are fundamentally about connections.Airbnb’s identity graph captures relationships between users in a graph database. The identity graph serves aggregated insights that enable user identity resolution and relationship understanding. These capabilities support a wide range of Trust and Safety use cases, from detecting suspicious activities to identifying…
Author: Stanislav GlukhovWhen you run a large production footprint in Google Cloud, changing a VM family is never just a hardware refresh. In our case, HAProxy sits on a critical path of the platform, serving as part of the traffic layer that hundreds of downstream systems quietly depend on every day. That means even a seemingly straightforward migration from one instance type to another has to be treated as a reliability exercise first and an infrastructure optimization second.We decided to migrate our HAProxy fleet from N2D to C4D, not because the old setup was failing, but because at our…
Guangtong Bai | Staff Software Engineer, Product ML Infrastructure*; Shantam Shorewala | Software Engineer II, Product ML Infrastructure*; Chi Zhang | Staff Software Engineer, AI Platform*; Neha Upadhyay | Software Engineer II, AI Platform*; Haoyang Li | Director, Product ML Infrastructure*These authors contributed equally to this article.BackgroundAt Pinterest, our online ML serving systems employ a root-leaf architecture. On a high level, the architecture looks as follows:Figure 1: Root-leaf Architecture of Online ML Serving Systems at PinterestIn the diagram, “Client Service” is…
At Wealthfront, the “tech” in financial technology isn’t just a buzzword—it’s the foundation of everything we build. Beneath the intuitive frontend our clients interact with lies a complex ecosystem of distributed systems. One of the most critical pieces of that backend architecture is the engine that enables us to route massive volumes of trades efficiently. ... Read more
Learn how Github uses eBPF to detect and prevent circular dependencies in its deployment tooling. The post How GitHub uses eBPF to improve deployment safety appeared first on The GitHub Blog.
By turning compaction into a layered, adaptive pipeline and strengthening our monitoring and controls, we made Magic Pocket more resilient to workload changes.
The Problem: Legacy Tooling and Its Limitations Currently, Slack utilizes a hybrid approach to network measurement, incorporating both internal (such as traffic between AWS Availability Zones) and external (monitoring traffic from the public internet into Slack’s infrastructure) solutions. These tools comprise a combination of commercial SaaS offerings and custom-built network testing solutions developed by our…
Engineering at Wealthfront is centered on the idea that code should be written to facilitate testing, not the other way around. Without a staging environment to fall back on, we maximize confidence through a sophisticated, multi-layered testing strategy. While unit tests provide our most rigorous line of defense, our Integration Server is the workhorse that... Read more
Every iOS application starts as a monolith. Xcode’s default project structure places all source files, resources, and build configuration into a single module (or target, for all the iOS devs reading this). For small apps, this works fine. For a 10+ year old financial application with roughly 2,000 Swift and a handful of Objective-C files,... Read more
In the fast-paced world of engineering, the dream of easy infrastructure management and provisioning is a common aspiration. At Zendesk, this sentiment resonates deeply among our engineers. When we talk about infrastructure, we refer to a wide range of tools such as MySQL, S3, DynamoDB, Kafka topics, compute resources, network and routing configurations, security groups, secrets, credentials, configuration settings, dashboards, monitors, and log management.Challenges with self-service provisioningIn our recent blog post, Unlocking Efficiency: A New Era for Datastore Provisioning, we…
Defra is a long-standing client for Capgemini, with both organisations sharing a commitment to sustainability. Recently we have been focussing on whether we can introduce “Green coding standards” to our teams. Whilst there can be benefits from developers focussing on making each line of code “greener”, much greater wins can be achieved from using best practices to minimise the infrastructure and resources used in the software development lifecycle. Listed below are the best practices that our teams follow at Defra to help make the application estate as sustainable as possible. Prerequisites…
On Sep 29, we had a period of about 2.5h from 14:00 to 16:30 (Pacific Time), in which an incomplete build was pushed to our infrastructure that handles storing Repl data. This caused Repls opened during that time window to become read-only or stop working. Any Repl not opened during that timeframe was not affected. We have addressed the root cause, recovered 98% of the affected Repls, and continue to work on recovering the remaining 2%. We understand that your data not being available is unacceptable for both you and your users. This post summarizes what happened and what we're doing to…
At Small Improvements, we are always keen to learn about our customers and how we can make the product better for them. Speaking to customers is great (and we do it all the time) but using the data we hold to find trends and usage patterns helps to find things that the customer won’t tell […]
Optional Google Analytics helps us understand visits. Microsoft Clarity records masked interactions to improve the site. Optional tools stay off unless you choose them. Privacy details.