Apache
We track 19 posts about Apache from 12 engineering blogs. Most active: Gunnar Morling, AWS, Confluent. Latest post: Sep 30, 2026.
Companies writing about Apache
Recent posts
Amazon S3 Tables now support all Apache Iceberg V3 data types (opens on the source site)
AWS ·
Amazon S3 Tables now offer complete support for the Apache Iceberg V3 spec. Upgrade existing V2 tables or create new V3 tables with built-in compaction, maintenance, and engine compatibility.
Amazon Aurora PostgreSQL now supports direct querying of Apache Iceberg and Parquet data in your data lake (opens on the source site)
AWS ·
Amazon Aurora PostgreSQL now lets you directly query Apache Iceberg and Parquet data stored in your data lake alongside live operational data—no ETL pipelines required. Powered by DuckDB embedded within Aurora, this capability enables single queries that join transactional and historical data using familiar PostgreSQL syntax. It supports AWS Glue Data Catalog, S3, and S3 Tables, with optimizations like predicate pushdown and caching for efficient performance.
ClickHouse now writes Apache Iceberg tables to Microsoft OneLake (opens on the source site)
ClickHouse today announced the ability to write directly to Microsoft OneLake, the unified data lake service within Microsoft Fabric.
Data liberation: Apache Kafka's native cluster mirroring (opens on the source site)
Red Hat ·
Apache Kafka excels at moving data within a cluster. Leaders replicate to followers, consumers pull from any replica, and the entire machinery runs with minimal operational overhead. Moving data between clusters has never been that simple. The post Data liberation: Apache Kafka's native cluster mirroring appeared first on Red Hat Developer.
Rerouting the Stream: How Lyft Moved to the Apache Flink Operator (opens on the source site)
Lyft ·
Written by Maheep Myneni, Arda Kuyumcu, and Prem Santosh Udaya Shankar at Lyft.Why We Migrated: Technical Debt Meets Modern Streaming DemandsOver the past several quarters, Lyft’s Streaming Compute team retired our internally developed Flink Kubernetes operator and moved our entire streaming fleet onto the open-source Apache Flink Kubernetes operator. This post is about why we made the switch, how we pulled it off incrementally without disrupting users, and the follow-on work it took to actually get the benefits we were after.Back in 2020, when we first architected the Lyft Flink Kubernetes…
AWS Glue 6.0 now available with 30% lower price and full Apache Iceberg v3 support (opens on the source site)
AWS ·
AWS Glue 6.0 is built on a fully modernized runtime, Apache Spark 4.1, Python 3.13, and Scala 2.13, delivering 30% lower pricing than previous AWS Glue versions.
Confluent Cloud for Apache Flink: Engine for Mission-Critical, Real-Time Operational Systems and dbt/SQL-Native Home for Data Science and AI (opens on the source site)
Flink now acts as a robust engine for developers through Table API, UDFs, and PTFs while offering a SQL-native, dbt-integrated platform for data science and AI teams.
Announcing Confluent Platform 8.3: Powerful Apache Flink® SQL operations, Easier KRaft Migrations, Expanded Monitoring and more. (opens on the source site)
Announcing Confluent Platform 8.3: Powerful Apache Flink® SQL operations, Easier KRaft Migrations, Expanded Monitoring and more
Scaling Grab's Data Lake: Our journey to Apache Iceberg adoption (opens on the source site)
Grab ·
Introduction: The evolution of Grab’s Data Lake At Grab’s scale, managing petabytes of data across billions of S3 objects demands more than a storage layer. It demands a robust architectural primitive that supports the high-concurrency needs of a modern “Lakehouse.” Our goal is full storage-compute separation, leveraging S3 as an elastic foundation for both near-real-time metrics and large-scale batch transformations. For years, the vast majority of our tables were Hive Parquet, managed through the Hive Metastore with a directory-based layout. This model served us well, but as data volume…
Security Baked Into the JVM: why fork Apache River and OpenJDK? (opens on the source site)
The more distributed a system, the harder it is to secure. Code crosses JVM boundaries. Objects are serialized across trust boundaries. Third-party proxies run inside your process. The usual answer is a network firewall. It helps, but it operates at the wrong level. Java 17 deprecated the SecurityManager, Java 24 put the final nail in its coffin. Most developers didn’t notice.
Hardwood 1.0: A Fast, Lightweight Apache Parquet Reader for the JVM (opens on the source site)
Table of Contents Why Hardwood What’s in Hardwood 1.0 Performance The Hardwood CLI Building Open-Source With AI A Big Thank You What’s Ahead Hardwood is a new Parquet library for the JVM, written from scratch to do one thing well: read (and soon, write) Apache Parquet files fast, with no mandatory dependencies. It is performance-focused and multi-threaded at its core, fanning page decoding out across all your CPU cores by default. Today, Hardwood reaches 1.0. After five preview releases since the start of the year (Alpha1, Beta1, Beta2, CR1, CR2), we now consider Hardwood ready for…
Extract Text from Your PDF and Image Files with Apache Tika (opens on the source site)
With AI becoming increasingly popular in everything, and retrieval-augmented generation (RAG) becoming a requirement in everyone's organization, how you're providing context to the AI tools becomes im... The post Extract Text from Your PDF and Image Files with Apache Tika appeared first on The Polyglot Developer.
Migrating from a Monolithic Orchestrator to Apache Airflow (opens on the source site)
Photo by Corinne Kutz on UnsplashBefore we knew betterOur orchestration system started as a simple internal solution to manage event pipelines and trigger downstream jobs. Over time, as more workflows and dependencies were added, it gradually evolved into a tightly coupled monolithic scheduler that became increasingly difficult to understand and maintain.Understanding how a workflow executed often meant looking through multiple files, configurations and database tables.For newer team members, onboarding into the system took time because much of the workflow context was distributed across…
Hardwood: A New Parser for Apache Parquet (opens on the source site)
Table of Contents Why Hardwood? Hello, Hardwood! Parsing Performance Built With AI, Not By AI What’s Next? Today, it’s my great pleasure to announce the first public release of Hardwood, a new parser for the Apache Parquet file format, optimized for minimal dependencies and great performance. Hardwood is open-source (Apache License 2.0) and supports Java 21 or newer. You can grab it from Maven Central and start parsing your Parquet files with ease and efficiency.
[Math] Intervals: Apache Maven version numbers (opens on the source site)
How intervals denoted formally: [a .. b] closed interval: {x | a ≤ x ≤ b} (a .. b) open interval: {x | a < x < b} [a .. b) half-open interval: {x | a ≤ x < b} (a .. b] half-closed interval: {x | a < x ≤ b} ( Donald E. Knuth - TAOCP, vol.2, 3rd ed. ) Apache Maven versions numbering inherited from mathematical notation: Range Meaning (,1.0] x <= 1.0 ... [1.0] Exactly 1.0 [1.2,1.3] 1.2 <= x <= 1.3 [1.0,2.0) 1.0 <= x < 2.0 [1.5,) x >= 1.5 (,1.0],[1.2,) x <= 1.0 or x >= 1.2. Multiple sets are separated by a comma. (,1.1),(1.1,) This excludes 1.1 if it is known not to work in combination with the…
Building a Native Binary for Apache Kafka on macOS (opens on the source site)
Table of Contents KIP-974: Docker Image for GraalVM based Native Kafka Broker With help of the GraalVM configuration developed for KIP-974 (Docker Image for GraalVM based Native Kafka Broker), you can easily build a self-contained native binary for Apache Kafka. Read on to learn how you can build a native Kafka executable yourself, starting in milli-seconds, making it a perfect fit for development and testing purposes. When I wrote about ahead-of-time class loading and linking in Java 24 recently, I also published the start-up time for Apache Kafka as a native binary for comparison. This was…
Get Running with Apache Flink on Kubernetes, part 2 of 2 (opens on the source site)
Table of Contents Fault Tolerance and High Availability Manually Triggering Savepoints Observability Bonus: Managing Flink Jobs With the Heimdall UI Summary and Discussion This post originally appeared on the Decodable blog. All rights reserved. Welcome back to this two-part blog post series about running Apache Flink on Kubernetes, using the Flink Kubernetes operator. In part one, we discussed installation and setup of the operator, different deployment types, how to deploy Flink jobs using custom Kubernetes resources, and how to create container images for your own Flink jobs. In this part,…
Get Running with Apache Flink on Kubernetes, part 1 of 2 (opens on the source site)
Table of Contents Installation and Setup Deployment Types Deploying Your First Flink Job on Kubernetes Building Custom Job Images This post originally appeared on the Decodable blog. All rights reserved. Kubernetes is a widely used deployment platform for Apache Flink. While Flink has had native support for Kubernetes for quite a while, it is in particular the operator pattern which makes deploying Flink jobs onto Kubernetes clusters a compelling option: you define jobs in a declarative resource, and a control loop running in a component called a Kubernetes operator takes care of provisioning…
Use Apache HTTP Client on Android SDK 23 (opens on the source site)
With the Android M SDK (API 23), Google removed the Apache HTTP Client library. It was deprecated since API 22, and Google recommended to use HttpURLConnection instead since API 9. While the classes are still bundled in Android 6 ROMs, it won't be long until we see them completely go away. Some applications are still relying on this library, and need to be updated to use the SDK 23, without having time/budget/whatever required to switch from HTTP Client. While I strongly recommend you to still take time to move to something else (there are many high-level libraries, like OkHttp or Ion, or you…
Related topics
This page is generated automatically from the engineering blogs we follow. Every post links to its source, where it was published. See all sources.