560 blogs tracked4,950 posts indexed

#data-infrastructure

7 posts · 2 companies · newest first

1

MetaRoCE: A New RDMA Transport Built for AI-Scale Ethernet (opens on the source site)

Training and serving frontier AI models depends on fast, reliable networks that move data between GPUs without wasting compute cycles. To meet this challenge at scale, Meta designed MetaRoCE – a clean-sheet RDMA transport protocol purpose-built for AI workloads on commodity Ethernet. We’re releasing the MetaRoCE specification, a reference software implementation and a compliance test [...] Read More... The post MetaRoCE: A New RDMA Transport Built for AI-Scale Ethernet appeared first on Engineering at Meta.

data-center-engineeringdata-infrastructureexcerpt only · body stays at the source
From the web
2

MTIA 300: Meta’s First Training Chip with Built-in NICs and Communication-Offloading Engines (opens on the source site)

MTIA 300 is the first of Meta’s family of in-house training and inference accelerators optimized for training ranking and recommendation models. We’re sharing how MTIA 300’s built-in NIC chiplets allow it to meet the communication needs associated with training recommendation models with superior performance over general-purpose GPUs. By co-designing MTIA’s communication library, HCCL, alongside the [...] Read More... The post MTIA 300: Meta’s First Training Chip with Built-in NICs and Communication-Offloading Engines appeared first on Engineering at Meta.

data-infrastructuredevinfraexcerpt only · body stays at the source
From the web
3

From User Sequences to Scaling Laws: A Multi-Stage Architecture for Meta’s Ads Ranking (opens on the source site)

Every day, Meta’s recommendation platforms handle billions of user interactions, generating rich temporal signals that capture individual preferences and intent across products, ads, and content. In our 2024 post on sequence learning for ads recommendations, we showed how modeling the order and timing of user actions (rather than relying on static, manually engineered sparse features) [...] Read More... The post From User Sequences to Scaling Laws: A Multi-Stage Architecture for Meta’s Ads Ranking appeared first on Engineering at Meta.

data-infrastructureml-applicationsexcerpt only · body stays at the source
From the web
4

GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model (opens on the source site)

Meta’s Generative Ads Recommendation Model (GEM), the foundation model behind ads recommendations across Instagram and Facebook, now trains at LLM scale on several thousand of the latest-generation GPUs. This post goes into the details on how we achieved: doubling end-to-end (E2E) training efficiency to 20–25% Model FLOPs Utilization (MFU) while scaling training FLOPs 4x in [...] Read More... The post GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model appeared first on Engineering at Meta.

ai-researchdata-infrastructureexcerpt only · body stays at the source
From the web
5

Meta’s AI Storage Blueprint at Scale (opens on the source site)

Over the past several years, model capabilities and training dataset sizes have experienced exponential growth. During the past year or so, the time between new-frontier-model releases has gone down from months to weeks. Reliable and fast access to storage is important to both the speed and computational cost of this AI innovation. If AI is [...] Read More... The post Meta’s AI Storage Blueprint at Scale appeared first on Engineering at Meta.

data-center-engineeringdata-infrastructureexcerpt only · body stays at the source
From the web
6

10 Years of Meta’s Commitment to Python (opens on the source site)

This year marks Meta’s 10th consecutive year as a sponsor of the Python Software Foundation (PSF), the charitable organization dedicated to advancing, supporting, and protecting the open-source Python programming language and the community that sustains it. Python is one of the world’s most influential programming languages, and we use it across our engineering stack, from [...] Read More... The post 10 Years of Meta’s Commitment to Python appeared first on Engineering at Meta.

ai-researchcultureexcerpt only · body stays at the source
From the web
7

Making User-Sequence Data More Cost-Efficient, Faster, and Easier to Use (opens on the source site)

Authors (listed alphabetically)Ads Feature Engineering Infra team: Ajay Venkatakrishnan, Le ZhangCore ML Infra team: Eric Shang, Pihui WeiML Data team: Connor Votroubek, Yi HeUser Understanding team: Camilo Munoz, Simin LiIf you work on ranking, retrieval, or recommendation systems, you’ve probably asked for some version of the same thing: “Give me the last N meaningful actions this user took, with the right enrichments, in a format that’s easy to train and serve ML models.”On paper, that sounds simple. In practice, “user sequences” often become one of the most expensive and fragile parts of…

machine-learningrecommendation-systemexcerpt only · body stays at the source
From the web
7 shown

Privacy choices

Reading never requires analytics. These choices last 90 days on this browser.

Essential sign-in and security storage always stays on. Read the privacy notice.