Ryan is joined by Yanbing Li, Chief Product Officer at Datadog, to talk about applying observability to non-deterministic AI agents, blurring the boundaries between software development and production workflows, and navigating emerging challenges in AI security and tokenomics.
Cloudflare is launching eight major updates that bring logs, traces, analytics, alerts, dashboards, querying, and telemetry export into one observability platform, with simpler and more predictable pricing.
You can now use built-in error monitoring in Cloudflare Workers to group production failures and send stack traces, logs, traces, and application context directly to a coding agent to investigate further and open a pull request.
Introduction In the first two parts of this series, we described how Grab approaches data mesh through the Signals Marketplace: a way for teams to publish, discover, and reuse trusted data products across domains. Part II introduced the foundational tools behind certification: Hubble for metadata and ownership, Genchi for data quality observability, and the Data Contract Registry for explicit producer-consumer guarantees. Certification is the starting point for a trusted data marketplace. It gives downstream consumers confidence in an asset’s ownership, documentation, lineage, and quality…
datadatabaseexcerpt only · body stays at the source
Pulumi v3.254.0 introduces automatic logging: every operation is logged in an encrypted log file that can optionally be shared with the Pulumi team for inspection. No more re-running commands just to get logs to the Pulumi team for debugging; instead you can share existing logs securely. You might have been in a situation where pulumi hit an error for an unexpected reason, or did something that was not quite right. Currently the process for trying to resolve that is to try and reproduce the error, ideally now with logging enabled. Sometimes the error doesn’t reproduce, or the state pulumi was…
By Parth Jain, Rakesh Sukumar, Yingwu Zhao, Renzo Sanchez-Silva & Nathan FisherA deep dive into the engineering challenges of building a real-time service dependency map at Netflix scale: from streaming architectures and distributed aggregation pipelines to time-travel queries and the methodology that made it work.IntroductionIn our first post, we introduced the problem: engineers at Netflix needed a unified, real-time view of service dependencies to troubleshoot faster, understand blast radius, and navigate our distributed architecture. We described our multi-source approach, combining eBPF…
Servers give you warnings before they fail. Most sysadmins performing Linux server monitoring miss them because they're watching the wrong numbers. The metrics that actually matter are one level deeper: iowait instead of CPU percentage, active swap paging instead of memory usage, inode counts instead of just disk space.Continue reading...
Before We Knew BetterAs our data platform grew, so did the number of pipelines, scheduled tasks, and data quality checks running every day.While Snowflake provided a reliable platform for storing and processing data, operational monitoring was fragmented across multiple systems. Data quality failures were often discovered only after downstream reports showed inconsistencies. Pipeline issues sometimes required engineers to manually inspect logs, query tables, and trace execution paths before identifying the root cause.The challenge wasn’t detecting failures — we already had mechanisms to…
Expedia Group Technology — EngineeringHow a screen-level performance metric reshaped platform decisions, engineering ownership, and release disciplinePhoto by Pietro De Grandi on UnsplashFor the last few years, my responsibility has been straightforward to state but hard to execute — owning the traveler login experience across mobile platforms.Not just whether a feature works, but whether it feels responsive, predictable, and trustworthy in the moments that matter most. During login, those moments are unforgiving: if a login screen hesitates travelers don’t interpret it as ‘a slow render’,…
Expedia Group Technology — EngineeringA system that facilitates investigation of service degradations and outages using service telemetry data and AIPhoto by Evangelos Mpikakis on Unsplash.The recent advancements in the artificial intelligence space make us re-evaluate how work is done. From programming, to designing systems, or even operating them in production. While there is considerable focus on automating programming, one area which could undergo transformation is how we monitor and operate our systems and services.A few of us came together and designed Expedia’s® Service Telemetry…
Given ClickHouse’s ability to execute complex analytical queries across terabytes of data in a single operation, proper monitoring and observability is critical. Its distributed architecture and scalability add layers of complexity, as multi-node clusters require careful coordination monitoring across shards and replicas to ensure data consistency and availability. Adding to the operational pressure is users’ […] The post ClickHouse Monitoring and Observability Decision Points appeared first on Severalnines.
Author: Fabian HöringAgentic systems powered by LLMs can be incredibly impressive in demos. With a few well-crafted prompts, they can demonstrate reasoning, calling tools, and solving complex tasks [1]. Demos are effective at showcasing what’s possible. Production environments, however, are where those capabilities are tested at scale and under real-world conditions.The same agent that performs perfectly on curated examples can behave unpredictably when exposed to real users. Inputs vary widely, conversations evolve over multiple turns, and small prompt changes can lead to unexpected…
The Problem: Legacy Tooling and Its Limitations Currently, Slack utilizes a hybrid approach to network measurement, incorporating both internal (such as traffic between AWS Availability Zones) and external (monitoring traffic from the public internet into Slack’s infrastructure) solutions. These tools comprise a combination of commercial SaaS offerings and custom-built network testing solutions developed by our…
IntroductionIn an age where artificial intelligence (AI) and machine learning (ML) are integral to almost every aspect of our lives, ensuring the effectiveness, fairness, and reliability of ML models is paramount. Observability plays a crucial role in maintaining the performance of these models, allowing us to detect and resolve issues promptly. At Helpshift, we recognized the need for robust ML observability to keep our models running smoothly and efficiently.This blog post explores our journey in building a custom ML observability solution tailored to our specific needs. We’ll delve into…
Heartbeats: How Synthetic Traffic Keeps Us RunningLet me take you on a journey of how we came to use heartbeats in our application design. It’s a happy story of love and no broken hearts along the way.What are heartbeats?What my teams have called heartbeats are a form of synthetic traffic generated by the application itself. The deployed application periodically generates heartbeats at a defined schedule.Heartbeats provide guaranteed regular traffic. In the cases I’ve used them, they have been low volume. In contrast to application traffic, which could vary massively from zero to huge…
Optional Google Analytics helps us understand visits. Microsoft Clarity records masked interactions to improve the site. Optional tools stay off unless you choose them. Privacy details.