Data
We track 156 posts about Data from 76 engineering blogs. Most active: Real Python, Doximity, Elastic. Latest post: Oct 9, 2026.
Companies writing about Data
Recent posts
How iSAM Funds built an options research platform for a 10,000x data problem with ClickHouse Cloud (opens on the source site)
iSAM Funds uses ClickHouse Cloud to research full options history, achieving 35–50x compression and 100x faster single-core ingestion while simplifying its data pipelines.
Announcing Workday Data Connect federation in Unity Catalog (opens on the source site)
We're excited to announce the publicly available Beta of the Workday Data Connect federation connector for Unity Catalog...
JSON-LD for Developers: Building Structured Data That Search and AI Systems Can Parse (opens on the source site)
null Continue reading JSON-LD for Developers: Building Structured Data That Search and AI Systems Can Parse on SitePoint.
Continuous Aggregate Refresh Policies for Solar and Wind SCADA Data: Choosing the Right Policy Shape (opens on the source site)
Five refresh policy shapes for solar and wind SCADA: frequent, daily, hierarchical, retention-aware, and manual. Pick the one that fits each job.
★ Keeping Laravel and TypeScript in sync with data objects (opens on the source site)
At Spatie we build SaaS apps with Laravel on the backend and Inertia with React on the frontend. That split means two type systems have to agree about every payload that crosses between them. Laravel data going out, TypeScript picking it up. There There, the helpdesk we're building, is built exactly like that. Tickets, messages, members, custom sidebar sections, all of it crosses the bridge. We'd go insane maintaining those types by hand. So we don't. There There is in private beta right now, and you can apply for early access at there-there.app. Two Spatie packages make this painless.…
Beyond synthetic testing: Capturing and replaying real database workloads at Airbnb (opens on the source site)
Airbnb ·
How we capture real production database traffic at Airbnb and replay it offline to load-test, plan capacity, and de-risk upgrades.By: Zuofei Wang, Erluo LiIntroductionAt Airbnb, MySQL-compatible databases are a critical backbone of our online database infrastructure: a fleet of hundreds of clusters supporting thousands of use cases at millions of queries per second (QPS). Operating databases at scale brings hard problems, including sizing clusters for future growth, keeping behavior consistent across version upgrades and migrations, and reproducing production incidents well enough to debug…
Scaling and Operating a Large dbt Project on Databricks: IFCO's Data Team on Performance, Visibility, and Debugging (opens on the source site)
AbstractIFCO runs one of the world's largest reusable packaging pools with hundreds of millions of crates and pallets...
Smarter personalization: How property data helps us understand user price sensitivity (opens on the source site)
Grab ·
Introduction Existing systems estimate user price sensitivity primarily from spending behavior or demographic proxies. They do not systematically account for residential property values, which can indicate a user’s financial circumstances. This omission creates three limitations: Limited property insights: Existing profiles do not account for property values. One-dimensional profiles: Users with similar spending patterns but different living standards receive the same classification. Regional variation: Manual classification does not adapt well to differences between property markets. This…
AI-ready data: The anthology - Part 2 of 2 (opens on the source site)
In part 1 of the AIRD anthology, our focus was on foundational prerequisites for an AI-ready platform. This included ensuring the data fueling the platform was high quality and governed securely so it could be trusted; using metadata and cataloguing to improve efficiency across a data-first organization; and establishing a semantic business layer so AI systems could identify and use the right dataset for each task. Once a solid foundation is built, more sophisticated components and topics can be introduced. This is the overarching focus of Part two. A to B accessibility As this series…
Decision Models vs. CVE Data (opens on the source site)
Ollama’s new decision endpoint, Cloudflare’s Clef models on a MacBook, and two questions asked of all 27,489 CVEs published in the last 60 days: does the description say why the bug matters, and is it clear? I read the announcement and had a weekend project before I finished it, because the question I care about ... Read more
Strings and Character Data in Python (opens on the source site)
Python strings are a sequence of characters used for handling textual data. You can create strings in Python using quotation marks or the str() function, which converts objects into strings. Strings in Python are immutable, meaning once you define a string, you can’t change it. By the end of this tutorial, you’ll understand that: The str() function converts objects to their string representation. String interpolation in Python allows you to insert values into {} placeholders in strings using f-strings or the .format() method. You access string elements in Python using indexing with square…
Quiz: Python's collections: A Buffet of Specialized Data Types (opens on the source site)
In this quiz, you’ll test your understanding of Python’s collections: A Buffet of Specialized Data Types. Python’s collections module provides specialized container data types that approach specific programming problems more clearly and more efficiently than the general purpose built-ins do. You’ll revisit namedtuple() for readable code, deque for queues and stacks, defaultdict for missing keys, OrderedDict for ordering, Counter for tallies, ChainMap for multiple mappings, and the wrapper classes for customizing built-ins. [ Improve Your Python With 🐍 Python Tricks 💌 – Get a short & sweet…
Basic Data Types in Python: A Quick Exploration (opens on the source site)
Python data types are fundamental to the language, enabling you to represent various kinds of data. You use basic data types like int, float, and complex for numbers, str for text, bytes and bytearray for binary data, and bool for Boolean values. These data types form the core of most Python programs, allowing you to handle numeric, textual, and logical data efficiently. Understanding Python data types involves recognizing their roles and how to work with them. You can create and manipulate these data types using built-in functions and methods, and convert between them when necessary. By the…
Postcodes for Laravel: GB Postcode Lookup and Geography Data (opens on the source site)
Laravel ·
Postcodes for Laravel adds typed GB postcode lookups, validation, geography data, distance searches, and test fakes through the GB Postcodes API. The post Postcodes for Laravel: GB Postcode Lookup and Geography Data appeared first on Laravel News. Join the Laravel Newsletter to get Laravel articles like this directly in your inbox.
SQLite Change Column Type Without Losing Data: The Safe Rebuild Guide (opens on the source site)
SQLite does not support altering a column type directly. Learn the safe table rebuild procedure to migrate column types, preserve data, restore indexes and triggers, and avoid silent data loss from foreign key cascades. Continue reading SQLite Change Column Type Without Losing Data: The Safe Rebuild Guide on SitePoint.
Amazon S3 Tables now support all Apache Iceberg V3 data types (opens on the source site)
AWS ·
Amazon S3 Tables now offer complete support for the Apache Iceberg V3 spec. Upgrade existing V2 tables or create new V3 tables with built-in compaction, maintenance, and engine compatibility.
WebRTC Data Channels vs WebSockets for Real-Time Browser Games (opens on the source site)
The transport you choose shapes how a multiplayer browser game plays. How WebRTC data channels and WebSockets differ, and when to use each. Continue reading WebRTC Data Channels vs WebSockets for Real-Time Browser Games on SitePoint.
Amazon Aurora PostgreSQL now supports direct querying of Apache Iceberg and Parquet data in your data lake (opens on the source site)
AWS ·
Amazon Aurora PostgreSQL now lets you directly query Apache Iceberg and Parquet data stored in your data lake alongside live operational data—no ETL pipelines required. Powered by DuckDB embedded within Aurora, this capability enables single queries that join transactional and historical data using familiar PostgreSQL syntax. It supports AWS Glue Data Catalog, S3, and S3 Tables, with optimizations like predicate pushdown and caching for efficient performance.
Can you use autoregressive diffusion to generate market data? (opens on the source site)
The following is part of a series of posts about 2026 summer intern projects – for more, see “What the interns have wrought, special jumbo 2026 edition”
Building Faster Prefix Search With the Trie Data Structure (opens on the source site)
Toptal ·
Returning suggestions after every keystroke is expensive when a system stores millions of entries. Find out how the trie data structure narrows each search and the most effective ways to build a working implementation in Java.
The 24 most commonly misunderstood marketing data terms (opens on the source site)
Imagine you’re a marketer planning a win-back campaign and you ask your data team for a list of “inactive customers.” Y...
How AI Let a Non-Technical Delivery Lead Work in the Data Layer (opens on the source site)
As a delivery lead, data-layer troubleshooting has historically been developer work: SQL fluency, comfort operating in a production database, and the confidence to investigate without breaking anything. For an accounting-based client, when their data didn’t reconcile or make sense, that question went to a developer. AI changed that for me in one engagement this year. […] The post How AI Let a Non-Technical Delivery Lead Work in the Data Layer appeared first on Atomic Spin.
How Cloudflare addressed a cross-tenant data exposure vulnerability in Containers (opens on the source site)
External security researchers at Accomplish identified a vulnerability in Cloudflare Containers that could expose residual disk data from previous workloads. We explain how the issue worked, how we investigated it, and the steps we took to remediate it.
Data modernization: A practical guide for getting it right (opens on the source site)
Warehouse, lake or lakehouse: How to pick Data warehouse: structured, schema-on-write, built for business intelligence (BI) and reporting with mature SQL tooling. A good fit when your workloads are well-defined, mostly structured and governance/consistency matter more than raw flexibility. Data lake: cheap storage, schema-on-read, holds structured, semi-structured and unstructured data side by side. Great for data science and exploratory work, but without discipline, it turns into a data swamp nobody trusts. Lakehouse: the industry's answer to not wanting to run and reconcile two separate…
Collected, stored, and useless: The lifecycle of most customer data (opens on the source site)
Twilio ·
You probably have more data than you know what to do with. The real problem is that the data is sitting in a system that doesn’t talk to the ones that need it.
How Fountain rebuilt its data plane on ClickHouse Cloud to power Cue, the Frontline Superintelligence (opens on the source site)
Fountain cut analytics latency from three hours to under two minutes and costs by 66% by rebuilding Cue’s data plane on ClickPipes and ClickHouse Cloud.
Data liberation: Apache Kafka's native cluster mirroring (opens on the source site)
Red Hat ·
Apache Kafka excels at moving data within a cluster. Leaders replicate to followers, consumers pull from any replica, and the entire machinery runs with minimal operational overhead. Moving data between clusters has never been that simple. The post Data liberation: Apache Kafka's native cluster mirroring appeared first on Red Hat Developer.
Stopping a crawl when the data stops looking real with Jev (opens on the source site)
A Scrapy pipeline that asks a fast, calibrated AI model whether each scraped field still looks real, and stops the crawl when too many don't. What it caught, what it misses, and whether building it was worth it.
The 2026 CDP Shift: Why AI Agents Are Only as Smart as Your Data Infrastructure (opens on the source site)
Twilio ·
Modern AI agents are only as effective as the data infrastructure behind them. Discover how unified CDPs eliminate AI silos, deliver real-time context, and power seamless, personalized customer experiences at scale. Read the 2026 CDP Report.
More data won't make your AI smarter (here's what will) (opens on the source site)
Twilio ·
Data collection was never the problem. 48% of companies say connecting that data across channels is their biggest obstacle.
Related topics
This page is generated automatically from the engineering blogs we follow. Every post links to its source, where it was published. See all sources.