Performance
We track 87 posts about Performance from 47 engineering blogs. Most active: Vanilla Java, Hayden James, Cloudflare. Latest post: Oct 6, 2026.
Companies writing about Performance
Recent posts
A faster, lighter C# Dev Kit (opens on the source site)
.NET ·
C# Dev Kit has been rebuilt to open solutions in about a second and run on a fraction of the memory, and this post walks through what changed for you. The post A faster, lighter C# Dev Kit appeared first on .NET Blog.
Scaling and Operating a Large dbt Project on Databricks: IFCO's Data Team on Performance, Visibility, and Debugging (opens on the source site)
AbstractIFCO runs one of the world's largest reusable packaging pools with hundreds of millions of crates and pallets...
2026 Birthday week: network performance update (opens on the source site)
Cloudflare now ranks as the fastest provider across 74% of the top 1,000 global networks. By incorporating background telemetry from Cloudflare Challenge Pages, we have expanded our real-user measurement scale while maintaining user privacy.
Four months of VoidZero at Cloudflare: making the open-source JavaScript toolchain faster for all humans and agents (opens on the source site)
Since joining Cloudflare, VoidZero has delivered more than 80 releases that drastically speed up JavaScript compilation, linting, and testing. From a 10x faster React compiler to Vite+ 1.0, here’s how we are building faster tools for developers and AI agents.
Improving site performance by shipping more CSS (opens on the source site)
A look at how a design system shipped major changes without breaking the world. The post Improving site performance by shipping more CSS appeared first on The GitHub Blog.
How I massively improved my AI inference performance without buying new hardware (opens on the source site)
Red Hat ·
Let me paint you a picture. You have 16 NVIDIA H200 GPUs spread across 2 nodes. That is, conservatively, several 100,000 dollars of silicon sitting in a data center, connected by RDMA/InfiniBand, running Kubernetes, and serving a large language model. You're living the agentic dream (not really, but it's a great start). Except your time to 1st token (TTFT) is thousands of milliseconds. The post How I massively improved my AI inference performance without buying new hardware appeared first on Red Hat Developer.
We just shipped support for the ugliest part of HTTP: Vary (opens on the source site)
Vary support is now available in Cache Rules on every plan. You can normalize known negotiation headers, pass exact values through to the origin when those small differences matter, or bypass cache when the variation is too unpredictable.
Open-Sourcing Rebalancer: A Generic, High-Performance Library for Solving Assignment Problems (opens on the source site)
Facebook ·
We’re open-sourcing Rebalancer, the assignment-problem solver that has been used to solve resource allocation problems throughout Meta for over nine years. Rebalancer separates several related concerns: how to specify an assignment problem, how to store it efficiently in memory, how to solve it, and how to debug it. This separation of concerns is crucial to [...] Read More... The post Open-Sourcing Rebalancer: A Generic, High-Performance Library for Solving Assignment Problems appeared first on Engineering at Meta.
Postgres on NVMe: performance and the convergence of transactions and analytics (opens on the source site)
See how local NVMe transforms Postgres transaction performance—and why ClickHouse remains essential for fast analytics as workloads scale.
Saving another 100TB of RAM with math (and Rust) (opens on the source site)
Cloudflare's global network is immense but not limitless. As we look for small ways to trim our resource usage, we sometimes get lucky and we can cut significantly more. Here’s how we reduced one of our Pingora-based service's RAM usage with statistics.
From hours to minutes: Optimizing Red Hat Developer Hub performance testing with immutable LDAP images (opens on the source site)
Red Hat ·
At Red Hat, our CI/CD pipelines are the heartbeat of our development process. However, as we scaled the Red Hat Developer Hub performance testing framework, we hit a wall. Our environment setup time was ballooning, turning what should have been a seamless verification step into a long waiting game. For performance engineering in catalog-dependent applications, deployment and catalog population speed directly dictate your feedback loops. The post From hours to minutes: Optimizing Red Hat Developer Hub performance testing with immutable LDAP images appeared first on Red Hat Developer.
Catching poor LLM performance and accuracy before deployment (opens on the source site)
Red Hat ·
vLLM is the leading open source inference and serving engine for LLMs. Averaging 64 commits a day with bi-weekly releases, the open source project goes through a significant amount of code changes rapidly. Red Hat AI Inference offers an integrated inference platform powered by vLLM, llm-d and vLLM's llm-compressor. The post Catching poor LLM performance and accuracy before deployment appeared first on Red Hat Developer.
NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut (opens on the source site)
Nvidia ·
System performance, efficient infrastructure scaling and continuous software optimization are key levers that determine AI inference economics. Higher system performance means more tokens generated, resulting in higher revenue. Efficient scaling means throughput grows proportionally as hardware gets added, requiring fewer resources to serve users at scale. Continuous optimization means generating more value from infrastructure investments. […]
Performance Improvements in .NET 11 (opens on the source site)
.NET ·
Take a tour through hundreds of performance improvements in .NET 11. The post Performance Improvements in .NET 11 appeared first on .NET Blog.
Understanding W8A8 INT8 LLM quantization: Accuracy and performance results (opens on the source site)
Red Hat ·
In Understanding W8A8 INT8 LLM quantization: Half the size, better performance, same accuracy, we compressed a Llama 3.1 8B Instruct model from 14.9 GB to 8.0 GB using 8-bit integer (INT8) W8A8 quantization with SmoothQuant and Generative Pre-trained Transformer Quantization (GPTQ). The post Understanding W8A8 INT8 LLM quantization: Accuracy and performance results appeared first on Red Hat Developer.
Linux server performance: Is disk I/O slowing your application? (opens on the source site)
If your Linux server is bogged down by disk I/O, your first step may often be to use the top command in the terminal to check load averages.Continue reading...
Eloquent Performance and Database Design: Evidence Before Eager Loading (opens on the source site)
A deep dive into Eloquent performance, from detecting N+1 queries to choosing aggregates, indexes, query plans, pagination, chunking, and transaction boundaries for a growing team dashboard. Read more
ClickHouse Cloud vs. Snowflake: What drives the real-time performance-per-dollar gap (opens on the source site)
ClickHouse Cloud delivered 412× better performance per dollar than Snowflake in CostBench. We trace the gap from fresh data arriving to fast answers coming back.
Automatic Key Exchange: faster, post-quantum secure origin handshakes for 45 billion daily connections (and counting) (opens on the source site)
Automatic Key Exchange probes TLS 1.3-capable customer origins to learn which key agreement algorithms they support. We then lead with the most secure algorithm when connecting to the origin, preferring post-quantum connections wherever the origin supports it.
Measuring real-time performance per dollar under continuous load: CostBench’s first end-to-end results (opens on the source site)
CostBench puts cloud data warehouses under continuous load. Across the complete path from fresh data to fast answers, ClickHouse Cloud delivers 412–1,996× better performance per dollar.
Understanding W8A8 INT8 LLM quantization: Half the size, better performance, same accuracy (opens on the source site)
Red Hat ·
Large language models are expensive to serve. A model like Llama 3.1 8B in Bfloat16 (BF16) precision occupies roughly 15 GB of GPU memory. In BF16, each of the 8 billion parameters takes 2 bytes to store, which adds up to roughly 15 GB for the weights—and that's not all. The GPU needs memory for the key-value (KV) cache to store context for active requests, alongside intermediate tensor outputs (activations, as we call them) generated during inference. The post Understanding W8A8 INT8 LLM quantization: Half the size, better performance, same accuracy appeared first on Red Hat Developer.
The Real Python Podcast – Episode #310: Performance Engineering: Profiling and Making Apps Fast by Default (opens on the source site)
How do you plan for the performance of your Python applications? What does a performance budget entail, and where should you spend your resources? This week on the show, we speak with Den Odell about his new book "Fast by Default: Practical Performance Engineering."
Parquet File Write Support, Bloom Filters, Improved Performance: Hardwood 1.1.0.Beta1 Is Out (opens on the source site)
Table of Contents Write Support Query Evaluation: Bloom Filters and Dictionary-Based Row-Group Pruning Performance Improvements Hardwood CLI Closing Thoughts "When is write support gonna land in Hardwood?" That’s probably the most common question I got over the last few months. As of today, I am very happy to share that the answer has changed from "It’s coming soon" to "A first cut is there, give it a try" — the first Beta of Hardwood 1.1 is out! This is a major milestone for the project, marking the first step in evolving Hardwood from being solely a Parquet parser to a complete library for…
PyTorch vs. TensorFlow: Differences, Performance, and How to Choose (opens on the source site)
Toptal ·
This comprehensive guide explores how PyTorch and TensorFlow shape deep-learning work in 2026, from experimentation and model design to production workflows, ecosystem tooling, and infrastructure considerations.
Linux swap Commands: Create and Manage Swap Files/Partitions (opens on the source site)
A practical guide to creating, managing, and tuning swap files and swap partitions on Linux using mkswap, swapon, and swapoff, with real commands and common troubleshooting fixes.Continue reading...
The Cloudflare Blog – Brought to you by EmDash (opens on the source site)
We migrated the Cloudflare Blog to EmDash to prove our stack at massive scale. Here is how we stress-tested performance, safely routed production traffic, and redesigned the frontend experience.
Migrating a Synology NAS to a UniFi UNAS Pro 8 with Robocopy, SMB Multichannel, and Surprising Performance Traps (opens on the source site)
I’ve had a Synology NAS for a very long time, and recently I started moving its contents to a new Ubiquiti UniFi UNAS Pro 8. This seemed like it ought to be a fairly boring operation. Both devices speak SMB, I have a fast network (recently upgraded to 10 gigabit internally), and Windows has had tools for copying files reliably between machines for decades. Naturally, it turned into a whole evening of learning things I thought I already knew, which is why I started a blog lol. There was a nice bit of history here for me because back in 2007 (good lord!) I wrote a post called “XCopy considered…
Linux Kernel Parameters Tuning for Better Performance (opens on the source site)
Learn how to tune Linux kernel parameters using sysctl and /proc/sys for better network throughput, memory management, file I/O, and server security. Includes a ready-to-use baseline config.Continue reading...
Get in, human: cut Rails boot time with require-profiler and this guide (opens on the source site)
Rails boot time is a DX metric in the AI age: meet require-profiler, learn to actually read sampling profilers, and see the pit stop that cut a 200-component monolith's boot by 40%.
The Real Python Podcast – Episode #307: Improving NumPy Performance on Free-Threaded Python (opens on the source site)
What bottlenecks were preventing NumPy from scaling on free-threaded Python? Christopher Trudeau is back on the show this week with another batch of PyCoder's Weekly articles and projects.
Related topics
This page is generated automatically from the engineering blogs we follow. Every post links to its source, where it was published. See all sources.