Cost
We track 28 posts about Cost from 23 engineering blogs. Most active: GitHub, Kevin Burke, Red Hat. Latest post: Oct 8, 2026.
Companies writing about Cost
Recent posts
Part 5: Operating an LLM system: observability, cost, routing, and the platform underneath (opens on the source site)
Your service can be 100% up and still quietly approving the wrong things, burning its budget, or failing over into untested quality. Level 5 is the infrastructure that lets you see your decisions, bound your spend, route and fail over between models, kill bad behavior in seconds — and the platform that makes all of it possible.
How Instagram Direct engineers built AI-native UI architecture with Jetpack Compose and reduced token cost per agent session by 33% (opens on the source site)
Android ·
Posted by Pavlo Stavytskyi, Software Engineer, Meta and Rebecca Franks, Developer Relations Engineer, Google This blog post is written in collaboration with the Meta team. Instagram Direct is one of the core surfaces on Instagram, handling billions of user messages every single day. Over years of iteration, the team squeezed every micro-optimization possible out of the legacy Android View system. However, maintaining and expanding a heavily optimized legacy surface creates significant technical debt and engineering overhead, especially as teams increasingly adopt declarative UI and AI coding…
How to rank fraud detection models using custom cost metrics (opens on the source site)
Red Hat ·
When building fraud detection systems, standard training steps often hide expensive failures. A model can show 96% accuracy on paper while letting every single fraudulent transaction pass unchecked. In fraud detection, a common first pass uses logistic regression with basic features and a 0.5 threshold. which often yields an accuracy of 96%. The post How to rank fraud detection models using custom cost metrics appeared first on Red Hat Developer.
Build a Cost-Aware LLM Router in Node.js with Claude Opus 5.5 and GPT-6 Sol (opens on the source site)
Build a TypeScript LLM router in Node.js that balances Claude Opus 5.5 and GPT-6 Sol with token budgeting, latency circuit breakers, and tier fallbacks. Continue reading Build a Cost-Aware LLM Router in Node.js with Claude Opus 5.5 and GPT-6 Sol on SitePoint.
What did AI cost you this quarter? (opens on the source site)
Red Hat ·
"What did AI cost us this quarter?" That question lands in every platform team's lap eventually, and most teams can't answer it. The CFO isn't trying to be difficult when they ask, but they do need a number, broken down by department, and they need it before the next board meeting. The post What did AI cost you this quarter? appeared first on Red Hat Developer.
Knowing When Your Composite Index Earns Its Write Cost (opens on the source site)
Composite indexes charge every insert and pay back only some queries. Four steps to price both sides and audit the indexes you already run.
New low-cost burstable Amazon EC2 T8i instances are generally available (opens on the source site)
AWS ·
AWS introduces new low-cost burstable Amazon EC2 T8i instances powered by custom sixth generation Intel Xeon Scalable Processors (Granite Rapids), available only on AWS. T8i instances are among the lowest-cost EC2 instances and deliver up to 30% better price performance over previous generation T3 instances.
The Real Cost of a Single Insert in Postgres (opens on the source site)
PostgreSQL write amplification explained: why indexes multiply WAL volume, how to measure it on your tables, and four ways to reduce it without schema changes.
Why Short AI Coding Prompts Can Cost You More Time (opens on the source site)
Ever tried to save time by keeping a prompt short - only to spend the next ten minutes answering follow-up questions because the AI didn't have what it needed?Yeah, me too.I recently came across a really interesting post from the GitHub Copilot team about making
How we make AI coding more cost efficient without sacrificing task quality (opens on the source site)
GitHub ·
Why shorter outputs can cost more, and how GitHub Copilot reduces wasted work across the complete coding task. The post How we make AI coding more cost efficient without sacrificing task quality appeared first on The GitHub Blog.
Gisting: Compressing LLM Agent context to ↑ throughput and ↓ cost (opens on the source site)
Shopify ·
Gisting compresses context into a set of learned tokens, preserving its quality while making the model faster and cheaper.
How canvases make agentic workflows visible, steerable, and cost-efficient (opens on the source site)
Chat is great for intent, but agent work gets lost in the scroll. Here is how I use canvases with my agentic workflows—and why your workflow also deserves a canvas. The post How canvases make agentic workflows visible, steerable, and cost-efficient appeared first on The GitHub Blog.
The real cost of a failed card payment (opens on the source site)
Quick summary Recurring card payments fail for reasons that have nothing to do with whether a customer can pay: expired cards, re-authentication friction, and insufficient funds at the moment of collection. Failed payments are just the beginning of an admin chain: dunning emails, hours of customer service time, and often involuntary churn. Card payment issues cost businesses around 3.5% of monthly revenue, and 42% of business leaders spend three or more hours a week managing failed card payments. Recurring Pay by Bank checks funds are available before collecting, so fewer payments fail in the…
How Lawmatics Cut CI Compute Cost by 39.3% and Shortened Pipeline Time by 15.8% (opens on the source site)
CI optimization is easiest to reason about when the problem is concrete: one pipeline, one critical path, and one cost model. Lawmatics reached out to us with that kind of problem. Their application pipeline was already parallelized and already using a sensible CI structure. The remaining question was whether the most expensive part of the […] The post How Lawmatics Cut CI Compute Cost by 39.3% and Shortened Pipeline Time by 15.8% appeared first on Semaphore.
The cost of saying yes has changed (opens on the source site)
GitHub ·
The cost of writing code dropped; the cost of owning it didn't. A framework for deciding which changes are actually cheap in the AI era. The post The cost of saying yes has changed appeared first on The GitHub Blog.
Best CI/CD Tools in 2026: Performance and Cost Compared (opens on the source site)
Choosing a CI/CD tool in 2026 isn’t just a technical decision — it’s a cost and productivity decision. Slow pipelines mean slower shipping, and opaque pricing means budget surprises at the end of the month. We benchmarked five of the most widely used CI/CD platforms against the same workload: a real Ruby on Rails application […] The post Best CI/CD Tools in 2026: Performance and Cost Compared appeared first on Semaphore.
Cost Attribution in Discord’s API (opens on the source site)
Discord ·
Discord's API spans 1700+ endpoints across hundreds of Kubernetes deployments. The challenge: tracking per-feature hosting costs without restructuring. Jim Benton helps explain how Discord tackled the situation.
Making User-Sequence Data More Cost-Efficient, Faster, and Easier to Use (opens on the source site)
Authors (listed alphabetically)Ads Feature Engineering Infra team: Ajay Venkatakrishnan, Le ZhangCore ML Infra team: Eric Shang, Pihui WeiML Data team: Connor Votroubek, Yi HeUser Understanding team: Camilo Munoz, Simin LiIf you work on ranking, retrieval, or recommendation systems, you’ve probably asked for some version of the same thing: “Give me the last N meaningful actions this user took, with the right enrichments, in a format that’s easy to train and serve ML models.”On paper, that sounds simple. In practice, “user sequences” often become one of the most expensive and fragile parts of…
How We Cut HAProxy Fleet GCP Cost in Half by Moving from N2D to C4D (opens on the source site)
Criteo ·
Author: Stanislav GlukhovWhen you run a large production footprint in Google Cloud, changing a VM family is never just a hardware refresh. In our case, HAProxy sits on a critical path of the platform, serving as part of the traffic layer that hundreds of downstream systems quietly depend on every day. That means even a seemingly straightforward migration from one instance type to another has to be treated as a reliability exercise first and an infrastructure optimization second.We decided to migrate our HAProxy fleet from N2D to C4D, not because the old setup was failing, but because at our…
The Hidden Cost of Convenience: Rethinking Old ORM Patterns for Scale (opens on the source site)
Ever been here before? Stuck with a job that needs to be continually revisited because its performance gets worse with every passing day, and each attempt at improving said performance yields diminishing returns? This is the situation we found ourselves in with the portfolio balance calculation system—the code responsible for aggregating data from multiple sources... Read more
Optimizing Third-Party Content Delivery: A Deep Dive into Preconnect’s Performance and Call Cost Implications (opens on the source site)
This document details how preconnect improves web performance, especially for Bazaarvoice's third-party content, by accelerating connection setups and reducing LCP. Crucially, internal testing confirmed preconnect operations are not counted as API calls, validating a "no count, low cost" model—a key insight for our developer blog.
Cost Optimisation in ECS: Integrating Spot Instances at Scale (opens on the source site)
At Deliveroo, we’re always refining how we scale - especially when it comes to managing compute costs in the cloud. After optimising our Amazon ECS workloads with Reserved Instances and Savings Plans, we saw an opportunity to push further using EC2 Spot Instances, which offer up to 90% savings compared to On-Demand prices. But Spot comes with challenges: Their availability can fluctuate, and they can be terminated with just a two-minute warning. To unlock these savings without compromising service stability, we had to engineer a robust solution across infrastructure, workload qualification,…
How to Avoid Thread-Safety Cost for Functions' static Variables (opens on the source site)
In this blog post, we’ll look at static variables defined in a function scope. We’ll see how they are implemented and how to use them. What’s more, we’ll discuss several cases where we can avoid extra thread-safety cost. Let’s start. Introduction As you may know, C++ offers a way to define static variables in a function/block scope: void foo() { static int counter = 0; ++counter; } Above, the counter variable will be initialized and created when foo() is invoked for the first time. In other words, a static local variable is initialized lazily. The counter is kept “outside” the function’s…
Cities Can Cost Effectively Start Their Own Utilities Now (opens on the source site)
Most PG&E ratepayers don't understand how much higher the rates they pay are than what it actually costs PG&E to generate and transmit the electricity to their house. When I looked into this recently I was shocked. The average PG&E electricity charge now starts at 40 cents per kilowatt hour and goes up from there. […]
Leveraging Spark 3 and NVIDIA’s GPUs to Reduce Cloud Cost by up to 70% for Big Data Pipelines (opens on the source site)
Paypal ·
By Ilay Chen and Tomer AkiravAt PayPal, hundreds of thousands of Apache Spark jobs run on an hourly basis, processing petabytes of data and requiring a high volume of resources. To handle the growth of machine learning solutions, PayPal requires scalable environments, cost awareness and constant innovation. This blog explains how Apache Spark 3 and GPUs can help enterprises potentially reduce Apache Spark’s jobs cloud costs by up to 70% for big data processing and AI applications.Our journey will begin with a brief introduction of Spark RAPIDS — Apache Spark’s accelerator that leverages GPUs…
Why public sector needs AI-powered observability: Cost savings, ROI, and analyst efficiency (opens on the source site)
Elastic ·
Elastic Observability helps public sector customers solve problems like IT modernization, data silos, and data inaccessibility. Customers gain significant value from Elastic Observability, including increased ROI, analyst efficiency and cost savings.
How much does a San Francisco Chronicle subscription cost? (opens on the source site)
The San Francisco Chronicle charges for subscriptions. How much does a subscription cost? This is an impossible question to answer, even for current subscribers. The Chronicle advertises several different prices for new subscribers. The only public information the Chronicle shares about its permanent subscription rates raises more questions than answers. No one at the Chronicle […]
The Real Cost of Technical Debt (opens on the source site)
Writing software is an iterative process. Rarely is software written and then never revisited. When this iteration occurs, a software engineer is presented with multiple options. Usually, a single objectively correct option is present but may not be chosen due to time constraints or other outside pressures. When shortcuts are taken to alleviate these external pressures, technical debt is often accrued. Permit Today, Pay Tomorrow Technical debt is referred to as such due to its similarities with financial debt. A credit card enables a person to complete a purchase they may not otherwise be…
Related topics
This page is generated automatically from the engineering blogs we follow. Every post links to its source, where it was published. See all sources.