560 blogs tracked4,950 posts indexed

#ai-infrastructure

17 posts · 1 company · newest first

1

How NVIDIA GPUs Help Accelerate OpenAI’s GPT-6 Astra Ultrafast (opens on the source site)

GPT-6 Astra Ultrafast, running on NVIDIA Blackwell GPUs, is available now in the OpenAI API and to eligible ChatGPT Work and Codex users. Accelerated by inference optimizations through OpenAI’s models that tap into the capabilities of the NVIDIA Blackwell architecture, Ultrafast offers up to 8x faster token generation than the Astra Standard mode. For developers, […]

aiai-infrastructureexcerpt only · body stays at the source
From the web
2

Productive, Durable, Fungible: How NVIDIA AI Factories Maximize Return on Investment (opens on the source site)

AI factories are built by the megawatt, even by the gigawatt. Each megawatt factory costs roughly $60 million, and AI factory operators will only commit capital on that scale with a clear view of the return on investment. Three key things shape AI factory returns: Earning capacity: What the factory could earn in a year […]

ai-infrastructurehardwareexcerpt only · body stays at the source
From the web
3

From Training to Production, NVIDIA and CoreWeave Close the Loop on Agentic AI (opens on the source site)

Building on nearly a decade of co-engineering, CoreWeave has built NVIDIA compute, networking and software into a cloud purpose-built for AI that’s still returning on investment across multiple generations of deployment. Now, CoreWeave is bringing the next generation of NVIDIA infrastructure to production. At CoreWeave Fully Connected, running this week in San Francisco, CoreWeave announced […]

ai-infrastructurecloudexcerpt only · body stays at the source
From the web
4

Sakeena Fiza Helps NVIDIA Hardware Succeed at Scale (opens on the source site)

When Sakeena Fiza describes her work as a validation engineer at NVIDIA, she does so in terms more befitting a detective story than a world-class engineering lab. “Validation engineers look in the shadows and shine a light into every corner,” Fiza said. “Every time we get a system, our first thought is: how can it […]

ai-infrastructurehardwareexcerpt only · body stays at the source
From the web
5

NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut (opens on the source site)

System performance, efficient infrastructure scaling and continuous software optimization are key levers that determine AI inference economics. Higher system performance means more tokens generated, resulting in higher revenue. Efficient scaling means throughput grows proportionally as hardware gets added, requiring fewer resources to serve users at scale. Continuous optimization means generating more value from infrastructure investments. […]

ai-infrastructurehardwareexcerpt only · body stays at the source
From the web
6

Emerald AI, Google and NVIDIA Launch Alliance to Advance Flexible AI Data Centers (opens on the source site)

AI factories are the infrastructure of the intelligence era. Scaling them responsibly will depend as much on innovation across the grid as inside the data center. Today, Emerald AI, Google and NVIDIA announced the launch of the AI Energy Management Alliance (AEMA), a first-of-its-kind coalition advancing data centers that can dynamically manage their electricity use […]

ai-infrastructurecorporateexcerpt only · body stays at the source
From the web
7

From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production (opens on the source site)

On a sweltering August evening in Silicon Valley, as the sun dropped and air conditioning loads spiked, Silicon Valley Power sent a signal to an AI factory to adjust its power consumption. Varun Sivaram was watching on Zoom with about forty others — his team at Emerald AI in their San Francisco conference room, engineers […]

ai-infrastructureexcerpt only · body stays at the source
From the web
8

AI Infra Summit: NVIDIA Vera Rubin and DSX Platform Advancements Showcase Energy Efficiencies of Optimizing Tokens Per Watt for AI Factories (opens on the source site)

Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA, Tuesday spoke on AI factory efficiency at the AI Infra Summit, the Santa Clara Convention Center event that has morphed into a Coachella of infrastructure tech. Before a packed audience — with more than 8,000 attendees this year, up from 3,500 last year — […]

ai-infrastructurenvidia-dsxexcerpt only · body stays at the source
From the web
10

NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memory (opens on the source site)

The next wave of AI is placing new demands on infrastructure. As AI agents and trillion-parameter workloads become mainstream, the performance of AI infrastructure depends not only on compute, but on how compute, memory, storage, networking and software are designed together as a unified system. To help hyperscalers and AI innovators build the next generation […]

ai-infrastructurehardwareexcerpt only · body stays at the source
From the web
11

How XPUs Meet a World-Class AI Factory (opens on the source site)

To generate intelligence at scale, AI factories run continuously, and their economics are defined by delivered output: tokens per second, tokens per watt, cost per token, utilization and uptime. That requires AI infrastructure designed and built as a full factory, not a collection of individual accelerators. Hyperscalers and AI-native companies building custom XPUs must consider […]

ai-infrastructurehardwareexcerpt only · body stays at the source
From the web
12

With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents (opens on the source site)

The next era of AI inference won’t be defined by a single breakthrough chip, network or system. It’ll be defined by how every layer of the AI factory works together. That’s why NVIDIA is extending Vera Rubin NVL72 with fast token generation for agentic systems. Announced today, the NVIDIA Vera Rubin rack-scale system NVIDIA Groq […]

ai-infrastructurebluefiexcerpt only · body stays at the source
From the web
13

Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents (opens on the source site)

According to OpenRouter data, agentic AI workloads consume 15x more tokens than a simple chat request. Why? Consider what happens when an AI agent researches a company for an investment decision. The agent queries financial databases, searches news and filings, invokes a sub-agent to run peer comparisons and model valuations, then synthesizes everything into a […]

ai-infrastructurehardwareexcerpt only · body stays at the source
From the web
14

Securing the Infrastructure of Intelligence (opens on the source site)

AI factories are the defining infrastructure of the AI era — where compute transforms energy and data into intelligence that powers every business, industry and country. In the AI economy, compute is revenue. AI factories require a full stack of critical resources: advanced chips, packaging, memory and networking — as well as land, power and […]

ai-infrastructurecorporateexcerpt only · body stays at the source
From the web
15

Universitas Gadjah Mada, Indosat and NVIDIA Open Indonesia’s First University AI Center to Develop Local AI Talent (opens on the source site)

Indonesia is taking charge of its AI future. This week, the Ministry of Communication and Digital Affairs (Komdigi), Indosat Ooredoo Hutchison (Indosat or IOH), NVIDIA and Universitas Gadjah Mada (UGM) launched the UGM Indosat NVIDIA AI Technology Center (NVAITC) in Yogyakarta — the country’s first university-based AI technology center. Established under Indonesia’s AI Center of […]

aiai-infrastructureexcerpt only · body stays at the source
From the web
16

NVIDIA AI Factory Compute Is Becoming an Investable Asset Class (opens on the source site)

We announced partnerships with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR to establish independent financing platforms designed to mobilize over $500 billion of third-party capital to support the buildout of AI infrastructure over time. This is a major milestone for NVIDIA and the AI industry. We have moved from an era in which companies […]

ai-infrastructurecorporateexcerpt only · body stays at the source
From the web
17

Why Scaling AI Compute Performance Requires a New Power Architecture (opens on the source site)

Every new generation of accelerated computing demands more from the infrastructure underneath it — more compute performance, higher rack density and more efficient, scalable power distribution. The bottleneck isn’t just wattage. It’s how power gets from the grid to the GPU. In traditional power delivery, electricity travels from the grid as an alternating current (AC) […]

ai-infrastructurehardwareexcerpt only · body stays at the source
From the web
17 shown

Privacy choices

Reading never requires analytics. These choices last 90 days on this browser.

Essential sign-in and security storage always stays on. Read the privacy notice.