560 blogs tracked4,950 posts indexed

#observability

15 posts · 12 companies · newest first

Curated view of this subject:Topic: Observability37
4

Data Mesh at Grab (Part III): Operationalizing data reliability with automated DPIs (opens on the source site)

Introduction In the first two parts of this series, we described how Grab approaches data mesh through the Signals Marketplace: a way for teams to publish, discover, and reuse trusted data products across domains. Part II introduced the foundational tools behind certification: Hubble for metadata and ownership, Genchi for data quality observability, and the Data Contract Registry for explicit producer-consumer guarantees. Certification is the starting point for a trusted data marketplace. It gives downstream consumers confidence in an asset’s ownership, documentation, lineage, and quality…

datadatabaseexcerpt only · body stays at the source
From the web
5

Automatic Logging for Faster, Secure Debugging (opens on the source site)

Pulumi v3.254.0 introduces automatic logging: every operation is logged in an encrypted log file that can optionally be shared with the Pulumi team for inspection. No more re-running commands just to get logs to the Pulumi team for debugging; instead you can share existing logs securely. You might have been in a situation where pulumi hit an error for an unexpected reason, or did something that was not quite right. Currently the process for trying to resolve that is to try and reproduce the error, ideally now with logging enabled. Sometimes the error doesn’t reproduce, or the state pulumi was…

loggingobservabilityexcerpt only · body stays at the source
From the web
6

Building Service Topology at Scale: Architecture, Challenges, and Lessons Learned (opens on the source site)

By Parth Jain, Rakesh Sukumar, Yingwu Zhao, Renzo Sanchez-Silva & Nathan FisherA deep dive into the engineering challenges of building a real-time service dependency map at Netflix scale: from streaming architectures and distributed aggregation pipelines to time-travel queries and the methodology that made it work.IntroductionIn our first post, we introduced the problem: engineers at Netflix needed a unified, real-time view of service dependencies to troubleshoot faster, understand blast radius, and navigate our distributed architecture. We described our multi-source approach, combining eBPF…

backend-developmentdistributed-systemsexcerpt only · body stays at the source
From the web
7

Linux Server Health Checks: 10 Metrics Every Sysadmin Should Monitor (opens on the source site)

Servers give you warnings before they fail. Most sysadmins performing Linux server monitoring miss them because they're watching the wrong numbers. The metrics that actually matter are one level deeper: iowait instead of CPU percentage, active swap paging instead of memory usage, inode counts instead of just disk space.Continue reading...

blogguestsexcerpt only · body stays at the source
From the web
8

Building a Centralized Alerting Framework for Data Quality Monitoring and Incident Management (opens on the source site)

Before We Knew BetterAs our data platform grew, so did the number of pipelines, scheduled tasks, and data quality checks running every day.While Snowflake provided a reliable platform for storing and processing data, operational monitoring was fragmented across multiple systems. Data quality failures were often discovered only after downstream reports showed inconsistencies. Pipeline issues sometimes required engineers to manually inspect logs, query tables, and trace execution paths before identifying the root cause.The challenge wasn’t detecting failures — we already had mechanisms to…

snowflakeincident-managementexcerpt only · body stays at the source
From the web
9

The Most Expensive Milliseconds Are Unmeasured (opens on the source site)

Expedia Group Technology — EngineeringHow a screen-level performance metric reshaped platform decisions, engineering ownership, and release disciplinePhoto by Pietro De Grandi on UnsplashFor the last few years, my responsibility has been straightforward to state but hard to execute — owning the traveler login experience across mobile platforms.Not just whether a feature works, but whether it feels responsive, predictable, and trustworthy in the moments that matter most. During login, those moments are unforgiving: if a login screen hesitates travelers don’t interpret it as ‘a slow render’,…

time-to-interactivesoftware-engineeringexcerpt only · body stays at the source
From the web
10

Expedia’s Service Telemetry Analyzer (opens on the source site)

Expedia Group Technology — EngineeringA system that facilitates investigation of service degradations and outages using service telemetry data and AIPhoto by Evangelos Mpikakis on Unsplash.The recent advancements in the artificial intelligence space make us re-evaluate how work is done. From programming, to designing systems, or even operating them in production. While there is considerable focus on automating programming, one area which could undergo transformation is how we monitor and operate our systems and services.A few of us came together and designed Expedia’s® Service Telemetry…

observabilitysoftware-engineeringexcerpt only · body stays at the source
From the web
11

ClickHouse Monitoring and Observability Decision Points (opens on the source site)

Given ClickHouse’s ability to execute complex analytical queries across terabytes of data in a single operation, proper monitoring and observability is critical. Its distributed architecture and scalability add layers of complexity, as multi-node clusters require careful coordination monitoring across shards and replicas to ensure data consistency and availability. Adding to the operational pressure is users’ […] The post ClickHouse Monitoring and Observability Decision Points appeared first on Severalnines.

monitoring---alertingclickhouseexcerpt only · body stays at the source
From the web
12

Beyond the demo: Why agentic evaluation matters (opens on the source site)

Author: Fabian HöringAgentic systems powered by LLMs can be incredibly impressive in demos. With a few well-crafted prompts, they can demonstrate reasoning, calling tools, and solving complex tasks [1]. Demos are effective at showcasing what’s possible. Production environments, however, are where those capabilities are tested at scale and under real-world conditions.The same agent that performs perfectly on curated examples can behave unpredictably when exposed to real users. Inputs vary widely, conversations evolve over multiple turns, and small prompt changes can lead to unexpected…

langfuseagentic-aiexcerpt only · body stays at the source
From the web
13

From Custom to Open: Scalable Network Probing and HTTP/3 Readiness with Prometheus (opens on the source site)

The Problem: Legacy Tooling and Its Limitations Currently, Slack utilizes a hybrid approach to network measurement, incorporating both internal (such as traffic between AWS Availability Zones) and external (monitoring traffic from the public internet into Slack’s infrastructure) solutions. These tools comprise a combination of commercial SaaS offerings and custom-built network testing solutions developed by our…

uncategorizedgolangexcerpt only · body stays at the source
From the web
14

From Data to Insight: Helpshift’s Journey with ML Observability (opens on the source site)

IntroductionIn an age where artificial intelligence (AI) and machine learning (ML) are integral to almost every aspect of our lives, ensuring the effectiveness, fairness, and reliability of ML models is paramount. Observability plays a crucial role in maintaining the performance of these models, allowing us to detect and resolve issues promptly. At Helpshift, we recognized the need for robust ML observability to keep our models running smoothly and efficiently.This blog post explores our journey in building a custom ML observability solution tailored to our specific needs. We’ll delve into…

analyticsartificial-intelligenceexcerpt only · body stays at the source
From the web
15

Heartbeats (opens on the source site)

Heartbeats: How Synthetic Traffic Keeps Us RunningLet me take you on a journey of how we came to use heartbeats in our application design. It’s a happy story of love and no broken hearts along the way.What are heartbeats?What my teams have called heartbeats are a form of synthetic traffic generated by the application itself. The deployed application periodically generates heartbeats at a defined schedule.Heartbeats provide guaranteed regular traffic. In the cases I’ve used them, they have been low volume. In contrast to application traffic, which could vary massively from zero to huge…

observabilitymonitoringexcerpt only · body stays at the source
From the web
15 shown

Privacy choices

Reading never requires analytics. These choices last 90 days on this browser.

Essential sign-in and security storage always stays on. Read the privacy notice.