Point-in-Time Correctness
This guide explains how to prevent data leakage when datasets for analytics or machine learning are built from an event history. It shows what leakage looks like, why Event Sourcing makes it avoidable, and which discipline it takes to actually avoid it.
Data leakage happens when a model is trained with information that would not have been available at the moment of the prediction. The result is treacherous: the model performs exceptionally well during training and evaluation, and then falls short in production, where the future is – as it should be – unknown.
An event-sourced system holds the complete, ordered history needed to avoid this trap. But the history alone does not prevent leakage. Projections and features have to be built with point-in-time correctness as a core principle.
What Leakage Looks Like#
Suppose a library wants to predict, at the moment a book is borrowed, whether it will be returned late.
The obvious form of leakage is to include information about the outcome itself. If the training data for a loan contains the late-fee-charged event of that very loan, the model is given the answer in advance. It will learn to rely on it and reach near-perfect accuracy in training – and fail in production, where no such event exists yet.
Most leakage is less obvious:
- Aggregates over the wrong period. A reader's on-time return rate, computed over all their loans, includes loans that happened after the one being predicted.
- Values that were corrected later. If a projection holds only the latest, corrected value of an attribute, it shows what is known today, not what was known back then.
- Global statistics. Normalizing a feature with the average of the entire dataset lets information from the future flow into every single example.
In each case, the dataset reflects today's knowledge, while the model will have to work with the knowledge of its moment.
Why Event Sourcing Helps#
An event store keeps events in strict chronological order and never changes them. This makes it possible to:
- Rebuild the state of any subject as it was at any past moment
- Generate projections that contain only what was known at a given time
- Avoid future events in features by design, not by luck
A traditional database cannot do this, because it has overwritten the past. An event-sourced system has kept it.
Try it with EventSourcingDB
In EventSourcingDB, the upperBound option limits a read to the events up to a given event ID – see Fine-Tuning Reads.
Stopping the Clock#
The core technique is to stop the clock for every example in a dataset. For each prediction the model is meant to make, define the cut-off: the moment at which the decision takes place. Then:
- Compute the features only from events before the cut-off.
- Derive the label – the outcome the model should learn – only from events after it.
For the late-return example, the cut-off is the book-borrowed event. The reader's on-time rate counts only returns that happened before it. Whether the book comes back late is determined by the book-returned event that follows.
Remember that events describe what happened, but not necessarily what was known at the time. If facts can be recorded late – a return that is entered into the system two days after it happened – decide which moment counts: when the fact occurred or when the system learned about it. For predictions, it is usually the latter, because a model can only use what the system knew.
The Same Rules in Training and Production#
Point-in-time correctness has to hold in two places: when the training data is built, and when the model is used. If the feature logic differs between them, the model sees one kind of data in training and another in production – a problem known as training-serving skew.
The most reliable way to avoid it is to use the same feature logic for both. In production, a feature is computed from the history up to now. In training, it is computed from the history up to each cut-off. With an event history, both are the same operation, just stopped at different points in time.
Honest Evaluation#
The same principle applies to evaluation. Randomly splitting a dataset into training and test data lets the model learn from the future of the examples it is tested on. A temporal split is more honest: train on everything up to a certain date, test on what came after.
Event Sourcing also makes backtesting straightforward. You can replay the history, let a model make the predictions it would have made at the time, and compare them with what actually happened – under fully controlled conditions, as often as needed.
A Discipline, Not a Switch#
Point-in-time correctness is not a feature that can be switched on. It requires:
- Thinking carefully about what the system knew, and when it knew it
- Making sure that every projection, dataset, and feature respects these boundaries
- Enforcing the same rules in training and in production
Event Sourcing provides the raw material for this discipline: a history that can be read up to any moment. Applying it consistently is a matter of practice. For how to query the history over time, see Patterns for Temporal Queries; for building the datasets themselves, see Designing Analytical Projections and Features from Events.