← All writing

[ Engineering · · 12 min read ]

Real-Time Feature Engineering: Patterns for Low-Latency ML Systems

When your model needs features computed in real time at sub-100ms latency, standard batch approaches collapse. Here are the patterns that scale.

Batch feature engineering is a solved problem. Tools like dbt, Spark and Airflow make it straightforward to compute features on a schedule and serve them from a feature store. But an increasing number of ML applications — fraud detection, dynamic pricing, recommendation engines, security anomaly detection — require features that reflect the state of the world right now, not as of the last batch run. Real-time feature engineering at low latency and high throughput is a fundamentally different engineering challenge that demands different tools, architectures and design patterns.

The core architectural pattern for real-time features is the dual-compute model: batch features are pre-computed and stored in a low-latency serving layer (typically Redis, DynamoDB or a purpose-built feature store), while real-time features are computed on-the-fly from streaming event data using a stream processor like Apache Flink, Kafka Streams or Spark Structured Streaming. At inference time, both feature sets are combined to form the complete feature vector. This pattern lets you get the best of both worlds — the efficiency of batch computation for slowly-changing features and the freshness of stream computation for rapidly-changing ones.

The trickiest aspect of this architecture is maintaining consistency between training and serving. When you train a model on historical data, the real-time features must be reconstructed as they would have appeared at each historical point in time — a process known as point-in-time correct feature computation. Getting this wrong introduces temporal leakage: the model sees information during training that would not have been available at prediction time, leading to inflated offline metrics and degraded production performance.

The second major challenge is managing state in your stream processor. Many useful real-time features are aggregations over time windows — such as the count of transactions in the last five minutes or the rolling average of request latency. These windowed aggregations require the stream processor to maintain state, which introduces complexity around fault tolerance, exactly-once processing and state migration during code deployments. The key is to keep your streaming state as small as possible, use incremental aggregation algorithms that do not require raw event replay, and invest in robust checkpointing and state recovery mechanisms.

Written by Ganesh Khetawat, founder of Aletheia AI

Need this built? See our full-stack development work, or tell us what you’re building.

[ Your turn ]

Have a hard problem?
Let’s build the answer.