Kafka
We track 20 posts about Kafka from 13 engineering blogs. Most active: Gunnar Morling, Etsy, Trivago. Latest post: Sep 22, 2026.
Companies writing about Kafka
Recent posts
Data liberation: Apache Kafka's native cluster mirroring (opens on the source site)
Red Hat ·
Apache Kafka excels at moving data within a cluster. Leaders replicate to followers, consumers pull from any replica, and the entire machinery runs with minimal operational overhead. Moving data between clusters has never been that simple. The post Data liberation: Apache Kafka's native cluster mirroring appeared first on Red Hat Developer.
Why your Kafka topic ignores retention.ms (and how to fix it) (opens on the source site)
Red Hat ·
A customer opened a support case with a deceptively simple complaint: a Kafka topic was configured with a 12-hour retention (retention.ms), yet messages produced on July 24 were still readable 4 days later, on July 28. Nothing was broken. The broker logged no errors. The retention policy was working as designed, but the segment layout and continuous write pattern delayed when the old records could actually be removed. The post Why your Kafka topic ignores retention.ms (and how to fix it) appeared first on Red Hat Developer.
Kafka App? There’s a Skill for That (opens on the source site)
Etsy ·
Etsy is home to over 100 million listings from 5.6 million active sellers. Because the items for sale are unique and creative, there is no standard product catalog that tells us what they are. When someone searches for “light linen dress for summer,” our models must infer what the shopper means and what the listings contain. They do that using data from user visits and Etsy listings. We use Kafka for streaming data. These data streams provide fresh features and embeddings that power our machine learning models. If a shopper favorites a hand-painted ceramic mug, that action can shape their…
Kafka App? There’s a Skill for That (opens on the source site)
Etsy ·
Etsy is home to over 100 million listings from 5.6 million active sellers. Because the items for sale are unique and creative, there is no standard product catalog that tells us what they are. When someone searches for “light linen dress for summer,” our models must infer what the shopper means and what the listings contain. They do that using data from user visits and Etsy listings. We use Kafka for streaming data. These data streams provide fresh features and embeddings that power our machine learning models. If a shopper favorites a hand-painted ceramic mug, that action can shape their…
Tableflow: Turn Kafka Topics into Iceberg Tables (opens on the source site)
Learn how Confluent Tableflow turns Kafka topics into Iceberg tables for zero-ETL analytics with automatic schema evolution and open catalog access.
How LivePerson optimized Logstash and Kafka performance on Google Cloud through benchmarking (opens on the source site)
Elastic ·
LivePerson found that benchmarking Google Cloud machine types cut Logstash costs by over half using AMD Milan instances, while Kafka compression codec selection significantly boosted throughput.
How We Cut Kafka Consumer Deployment Costs by 83% (opens on the source site)
Trivago ·
This post walks through a layered performance investigation that cut PSE-kafka’s infrastructure costs by 83% and ended a run of 19 P1 incidents. Problem Background PSE-kafka (price-search-engine-kafka) is an internal microservice. It consumes hotel price Kafka messages and pushes ads to an external Ads service, which drives real revenue for trivago. But the service had a few persistent problems. CPU usage in production sat very low, around 10%. On startup, it took a long time to join its Kafka consumer group, which dragged out every deployment. It suffered from high consumer lag. The replica…
From Always-On to On-Demand: Scaling Kafka Sinks with KEDA (opens on the source site)
Trivago ·
Introduction / Context Kafka sits at the heart of how we move data between systems at trivago. Many teams publish changes to Kafka, and downstream services consume those changes to keep user-facing features up to date—things like accommodation reviews, highlights, and other derived attributes. To make that possible at low latency, we run a fleet of Kafka consumers we call sinks. Each sink takes events from Kafka, applies the necessary business logic (filtering, normalization, policy checks, etc.), and writes the result into a service-local database as a materialized view. This pattern works…
"You Don't Need Kafka, Just Use Postgres" Considered Harmful (opens on the source site)
Looking to make it to the front page of HackerNews? Then writing a post arguing that "Postgres is enough", or why "you don’t need Kafka at your scale" is a pretty failsafe way of achieving exactly that. No matter how often it has been discussed before, this topic is always doing well. And sure, what’s not to love about that? I mean, it has it all: Postgres, everybody’s most favorite RDBMS—check! Keeping things lean and easy—sure, count me in! A somewhat spicy take—bring it on!
Heartbeats (opens on the source site)
Zendesk ·
Heartbeats: How Synthetic Traffic Keeps Us RunningLet me take you on a journey of how we came to use heartbeats in our application design. It’s a happy story of love and no broken hearts along the way.What are heartbeats?What my teams have called heartbeats are a form of synthetic traffic generated by the application itself. The deployed application periodically generates heartbeats at a defined schedule.Heartbeats provide guaranteed regular traffic. In the cases I’ve used them, they have been low volume. In contrast to application traffic, which could vary massively from zero to huge…
Using Clojure channels to increase throughput (opens on the source site)
When building systems that process large volumes of messages synchronously, performance bottlenecks can quickly become a challenge specially with single-threaded designs. In this post, we’ll look at how leveraging worker threads in a Clojure-based Kafka consumer can significantly boost throughput & reduce total processing time. Using simple concurrency primitives, it’s possible to achieve parallelism & scale gracefully, all while keeping the codebase clean & maintainable. We’ll start with a baseline, introduce worker threads using Clojure’s core.async & measure the impact.Setup & ContextKafka…
What If We Could Rebuild Kafka From Scratch? (opens on the source site)
The last few days I spent some time digging into the recently announced KIP-1150 ("Diskless Kafka"), as well AutoMQ’s Kafka fork, tightly integrating Apache Kafka and object storage, such as S3. Following the example set by WarpStream, these projects aim to substantially improve the experience of using Kafka in cloud environments, providing better elasticity, drastically reducing cost, and paving the way towards native lakehouse integration. This got me thinking, if we were to start all over and develop a durable cloud-native event log from scratch—Kafka.next if you will—which traits and…
Building a Real-Time AI Fraud Detection System with Spring Kafka and MongoDB (opens on the source site)
In this tutorial, we'll build a real-time fraud detection system using MongoDB Atlas Vector Search, Apache Kafka, and AI-generated embeddings. We'll demonstrate how MongoDB Atlas Vector Search can be ... The post Building a Real-Time AI Fraud Detection System with Spring Kafka and MongoDB appeared first on DEV.
A Deep Dive Into Ingesting Debezium Events From Kafka With Flink SQL (opens on the source site)
Table of Contents Flink SQL Connectors for Apache Kafka The Apache Kafka SQL Connector in Append-Only Mode The Apache Kafka SQL Connector As a Changelog Source The Upsert Kafka SQL Connector Summary Over the years, I’ve spoken quite a bit about the use cases for processing Debezium data change events with Apache Flink, such as metadata enrichment, building denormalized data views, and creating data contracts for your CDC streams. One detail I haven’t covered in depth so far is how to actually ingest Debezium change events from a Kafka topic into Flink, in particular via Flink SQL. Several…
Building a Native Binary for Apache Kafka on macOS (opens on the source site)
Table of Contents KIP-974: Docker Image for GraalVM based Native Kafka Broker With help of the GraalVM configuration developed for KIP-974 (Docker Image for GraalVM based Native Kafka Broker), you can easily build a self-contained native binary for Apache Kafka. Read on to learn how you can build a native Kafka executable yourself, starting in milli-seconds, making it a perfect fit for development and testing purposes. When I wrote about ahead-of-time class loading and linking in Java 24 recently, I also published the start-up time for Apache Kafka as a native binary for comparison. This was…
Let's Take a Look at... KIP-932: Queues for Kafka! (opens on the source site)
Table of Contents Towards Queue Support in Kafka—Introducing Share Groups Share Groups in Action Retry Behavior and State Management Share Group State Persistence Summary and Outlook In the "Let’s Take a Look at…!" blog series I am going to explore interesting projects, developments and technologies in the data and streaming space. This can be KIPs and FLIPs, open-source projects, services, and more. The idea is to get some hands-on experience, learn about potential use cases and applications, and understand the trade-offs involved. If you think there’s a specific subject I should take a…
Migrating to Confluent Kafka: A Comprehensive Guide (opens on the source site)
VNGRS ·
Confluent Kafka is an enterprise-ready distribution of Apache Kafka, providing a reliable and scalable platform for real-time data streaming. It includes additional features such as schema registry, REST proxy, connectors, and advanced monitoring tools.And migrating your existing Kafka cluster can be driven by several factors, including the need to upgrade hardware, move to a cloud-based infrastructure, enhance performance, improve scalability, or optimize resource utilization. This migration can help ensure your data infrastructure remains robust and capable of handling future…
Data replication across backend services with Kafka and Protobuf (opens on the source site)
The Jobteaser application contains a lot of different relatively independent modules to help universities provide career guidance to students: a job board, a career event management system, a career advice appointment management system…When we decided to migrate our application’s backend from a monolith to a service-oriented architecture, we strived to keep each module as isolated as possible from the others in the event of an incident. If the career appointment system was down, students should still be able to browse and apply to job ads.That isolation is achieved through what we’ve called…
Tiny Letter from Kafka (opens on the source site)
This article discusses the powerful design choice of Apache Kafka, “an open-source distributed event streaming platform,” and gives a sneak…
Kafka Consumer memory usage (opens on the source site)
I’m working with Kafka for more than 2 years and I wasn’t sure if Kafka Consumer eats more RAM memory when it has more partitions. I couldn’t find any useful information on the internet, so I decided to measure everything by myself.
Related topics
This page is generated automatically from the engineering blogs we follow. Every post links to its source, where it was published. See all sources.