Observability
We track 37 posts about Observability from 23 engineering blogs. Most active: Elastic, Clickhouse, Expedia. Latest post: Oct 9, 2026.
Companies writing about Observability
Recent posts
Taking a look under your agent’s hood (opens on the source site)
Ryan is joined by Yanbing Li, Chief Product Officer at Datadog, to talk about applying observability to non-deterministic AI agents, blurring the boundaries between software development and production workflows, and navigating emerging challenges in AI security and tokenomics.
Part 5: Operating an LLM system: observability, cost, routing, and the platform underneath (opens on the source site)
Your service can be 100% up and still quietly approving the wrong things, burning its budget, or failing over into untested quality. Level 5 is the infrastructure that lets you see your decisions, bound your spend, route and fail over between models, kill bad behavior in seconds — and the platform that makes all of it possible.
Rossoctl for agent discoverability, security, and observability on Kubernetes (opens on the source site)
Red Hat ·
Imagine deploying a pair of agents to Red Hat OpenShift and watching them thrive. But fast-forward 6 months and you find your environment cluttered with 40 of them, their purposes largely unknown. Redundant agents emerge across namespaces, duplicating work. Meanwhile, a security review uncovers agents utilizing static API keys over insecure HTTP. When an agent falters, the lack of tracing makes it impossible to distinguish between model hallucinations, tool errors, or downstream failures. The post Rossoctl for agent discoverability, security, and observability on Kubernetes appeared first on…
8 major updates to Cloudflare Observability (opens on the source site)
Cloudflare is launching eight major updates that bring logs, traces, analytics, alerts, dashboards, querying, and telemetry export into one observability platform, with simpler and more predictable pricing.
From reactive to resilient: Findings from the Total Economic Impact of Elastic Observability 2026 (opens on the source site)
Elastic ·
A Forrester Consulting study finds that enterprises deploying Elastic Observability achieve a 362% ROI, $28.1 million in three-year benefits, and payback in under six months, all while reclaiming thousands of engineering hours from incident response.
Datadog alternatives: which observability platform fits your needs? (opens on the source site)
If you've ever responded to a late-night service outage, then you know how important it is to have the right tooling to get everything back up and running. Platforms like Datadog handle monitoring along with data collection and analysis across all kinds of vectors, spanning infrastructure monitoring, network performance, and security compliance. The insights Datadog provides enable actionable improvements to your technical architecture, resulting in satisfied customers and better business outcomes. While Datadog is one of the dominant players in the observability space, that doesn't mean it's…
Detect and send production issues straight to your agent (opens on the source site)
You can now use built-in error monitoring in Cloudflare Workers to group production failures and send stack traces, logs, traces, and application context directly to a coding agent to investigate further and open a pull request.
Vercel Sandbox now supports memory observability (opens on the source site)
Vercel ·
Vercel Sandbox observability now includes memory usage data. You can access sandbox memory usage data in the dashboard and through the CLI via the vercel metrics command. Sandbox observability memory in the dashboard The Memory Usage card reports average, P75, and P95 memory across your sandboxes, alongside the existing CPU usage and data transfer metrics on the sandbox overview page and the project and team-level pages. On the sandbox detail page, charts are designed to show how close a sandbox is running to its memory ceiling by: Automatically scaling the y-axis to sandbox's memory limit…
Now on Amazon CloudWatch Omni: collaborative AI-powered observability for your applications (opens on the source site)
AWS ·
Amazon CloudWatch Omni is the next evolution of CloudWatch — unified observability that brings your applications and AI agents into one reimagined experience, with auto-discovered topology, natural language queries, and AI-guided investigation powered by AWS DevOps Agent.
Beyond the 200 OK: Architecting Observability for AI (opens on the source site)
Traditional monitoring tools, such as application performance monitoring (APM), were engineered to monitor deterministic software where specific inputs reliably lead to predictable outputs through hard-coded logic. When a traditional API fails, it usually throws a 500 Internal Server Error. But when an AI agent fails, it might return a perfectly healthy 200 OK status code ...
The security attack that hid inside your observability data (opens on the source site)
Elastic ·
Your ops team sees a CPU spike. Your security team sees nothing. The cryptominer runs for six hours. Avoid duplication of cost and time and see how a unified observability and security platform with grounded AI closes the gap.
How Uken Games reduces observability costs by 87% with ClickHouse (opens on the source site)
Uken Games replaced Datadog with an open-source observability stack built on ClickHouse, cutting costs by 87% while storing every trace on a single node.
Data Mesh at Grab (Part III): Operationalizing data reliability with automated DPIs (opens on the source site)
Grab ·
Introduction In the first two parts of this series, we described how Grab approaches data mesh through the Signals Marketplace: a way for teams to publish, discover, and reuse trusted data products across domains. Part II introduced the foundational tools behind certification: Hubble for metadata and ownership, Genchi for data quality observability, and the Data Contract Registry for explicit producer-consumer guarantees. Certification is the starting point for a trusted data marketplace. It gives downstream consumers confidence in an asset’s ownership, documentation, lineage, and quality…
So, is ClickHouse winning the observability wars? (opens on the source site)
After claims that ClickHouse is “winning the observability wars” sparked debate, we reflect on why it has become a leading storage and query engine, where it still falls short, and why winning the database layer isn’t the same as winning observability.
How AI observability works with MLflow (opens on the source site)
Red Hat ·
As I'm sure you already know, AI responses aren't always correct (your favorite LLM probably has a disclaimer saying the same thing). The post How AI observability works with MLflow appeared first on Red Hat Developer.
Shopify powers observability for global-scale commerce with ClickHouse (opens on the source site)
Shopify unified global-scale observability on ClickHouse, achieving up to 30x faster queries while ingesting 100 million events per second at peak.
7 lessons for IT leaders on using observability to monitor AI applications (opens on the source site)
Elastic ·
Discover the seven lessons we learned as we evolve our LLM observability practice to monitor and improve our AI applications.
bitdrift brings ClickHouse-powered mobile observability to ClickStack (opens on the source site)
bitdrift joins ClickHouse’s House Mates program with a mobile observability integration, bringing mobile-native telemetry, tracing, and debugging to ClickStack.
Automatic Logging for Faster, Secure Debugging (opens on the source site)
Pulumi ·
Pulumi v3.254.0 introduces automatic logging: every operation is logged in an encrypted log file that can optionally be shared with the Pulumi team for inspection. No more re-running commands just to get logs to the Pulumi team for debugging; instead you can share existing logs securely. You might have been in a situation where pulumi hit an error for an unexpected reason, or did something that was not quite right. Currently the process for trying to resolve that is to try and reproduce the error, ideally now with logging enabled. Sometimes the error doesn’t reproduce, or the state pulumi was…
Elastic named a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms (opens on the source site)
Elastic ·
Elastic was named a Leader in the Gartner® Magic Quadrant™ for Observability Platforms. In our opinion, Elastic Observability helps teams investigate faster with agentic AI with full context and reduces costs with efficient metrics and logs storage.
Building Service Topology at Scale: Architecture, Challenges, and Lessons Learned (opens on the source site)
Netflix ·
By Parth Jain, Rakesh Sukumar, Yingwu Zhao, Renzo Sanchez-Silva & Nathan FisherA deep dive into the engineering challenges of building a real-time service dependency map at Netflix scale: from streaming architectures and distributed aggregation pipelines to time-travel queries and the methodology that made it work.IntroductionIn our first post, we introduced the problem: engineers at Netflix needed a unified, real-time view of service dependencies to troubleshoot faster, understand blast radius, and navigate our distributed architecture. We described our multi-source approach, combining eBPF…
Linux Server Health Checks: 10 Metrics Every Sysadmin Should Monitor (opens on the source site)
Servers give you warnings before they fail. Most sysadmins performing Linux server monitoring miss them because they're watching the wrong numbers. The metrics that actually matter are one level deeper: iowait instead of CPU percentage, active swap paging instead of memory usage, inode counts instead of just disk space.Continue reading...
Building a Centralized Alerting Framework for Data Quality Monitoring and Incident Management (opens on the source site)
Before We Knew BetterAs our data platform grew, so did the number of pipelines, scheduled tasks, and data quality checks running every day.While Snowflake provided a reliable platform for storing and processing data, operational monitoring was fragmented across multiple systems. Data quality failures were often discovered only after downstream reports showed inconsistencies. Pipeline issues sometimes required engineers to manually inspect logs, query tables, and trace execution paths before identifying the root cause.The challenge wasn’t detecting failures — we already had mechanisms to…
Managing Elasticsearch Reindex at Scale: Performance, Reliability, and Observability (opens on the source site)
Palantir ·
Editor’s Note: This is the fourth post in a series exploring how Palantir customizes infrastructure software for reliable operation at scale.The following is a guest contribution to the Foundations series from the Gotham Core Platform organization, which builds and maintains the bedrock for mission-critical applications within the Gotham ecosystem. This blog post by Kevin Liang, a backend developer based in CA, highlights the design considerations and improvements made to the Elasticsearch reindex machinery — which broadly aims to provide an easy-to-use, performant, reliable, and observable…
The Most Expensive Milliseconds Are Unmeasured (opens on the source site)
Expedia ·
Expedia Group Technology — EngineeringHow a screen-level performance metric reshaped platform decisions, engineering ownership, and release disciplinePhoto by Pietro De Grandi on UnsplashFor the last few years, my responsibility has been straightforward to state but hard to execute — owning the traveler login experience across mobile platforms.Not just whether a feature works, but whether it feels responsive, predictable, and trustworthy in the moments that matter most. During login, those moments are unforgiving: if a login screen hesitates travelers don’t interpret it as ‘a slow render’,…
Expedia’s Service Telemetry Analyzer (opens on the source site)
Expedia ·
Expedia Group Technology — EngineeringA system that facilitates investigation of service degradations and outages using service telemetry data and AIPhoto by Evangelos Mpikakis on Unsplash.The recent advancements in the artificial intelligence space make us re-evaluate how work is done. From programming, to designing systems, or even operating them in production. While there is considerable focus on automating programming, one area which could undergo transformation is how we monitor and operate our systems and services.A few of us came together and designed Expedia’s® Service Telemetry…
ClickHouse Monitoring and Observability Decision Points (opens on the source site)
Given ClickHouse’s ability to execute complex analytical queries across terabytes of data in a single operation, proper monitoring and observability is critical. Its distributed architecture and scalability add layers of complexity, as multi-node clusters require careful coordination monitoring across shards and replicas to ensure data consistency and availability. Adding to the operational pressure is users’ […] The post ClickHouse Monitoring and Observability Decision Points appeared first on Severalnines.
Beyond the demo: Why agentic evaluation matters (opens on the source site)
Criteo ·
Author: Fabian HöringAgentic systems powered by LLMs can be incredibly impressive in demos. With a few well-crafted prompts, they can demonstrate reasoning, calling tools, and solving complex tasks [1]. Demos are effective at showcasing what’s possible. Production environments, however, are where those capabilities are tested at scale and under real-world conditions.The same agent that performs perfectly on curated examples can behave unpredictably when exposed to real users. Inputs vary widely, conversations evolve over multiple turns, and small prompt changes can lead to unexpected…
From Custom to Open: Scalable Network Probing and HTTP/3 Readiness with Prometheus (opens on the source site)
Slack ·
The Problem: Legacy Tooling and Its Limitations Currently, Slack utilizes a hybrid approach to network measurement, incorporating both internal (such as traffic between AWS Availability Zones) and external (monitoring traffic from the public internet into Slack’s infrastructure) solutions. These tools comprise a combination of commercial SaaS offerings and custom-built network testing solutions developed by our…
Maple: an open-source observability platform built with Tinybird's TypeScript SDK (opens on the source site)
Tinybird ·
David Granzin built Maple, an open-source observability platform for metrics, logs, and traces, using Tinybird's TypeScript SDK. Zero infrastructure to manage, AI agents accelerating development, and two projects shipped simultaneously.
Related topics
This page is generated automatically from the engineering blogs we follow. Every post links to its source, where it was published. See all sources.