Concepts / Analyzing / Features from Events
Analyzing

Features from Events

This guide explains how to derive features – the measurable variables a model learns from – from an event history. It covers what makes a good feature, which kinds of features events make possible, and how to keep training data reproducible.

Analytical projections provide structured datasets, but models do not learn from projections directly. They learn from features: values that capture a meaningful aspect of the domain and carry a signal for the prediction at hand. Turning projections into features is where domain understanding meets the craft of feature engineering.

Why Features Matter More Than Algorithms#

The choice and quality of features often have a greater impact on a model than the choice of algorithm. A feature that reflects a genuine business signal can turn a mediocre model into a good one, while a misleading or unstable feature undermines even the most sophisticated method.

Good features are:

  • Relevant – they represent something that actually influences the outcome.
  • Consistent – they mean the same thing across the entire dataset and over time.
  • Explainable – domain experts understand what they mean and where they come from.

What an Event History Adds#

With an event history, features are not limited to what the current state happens to contain. Because the history can be replayed up to any point in time, features can be:

  • Historically accurate – computed from what actually happened, not from what is left of it.
  • Reproducible – the same events always yield the same values.
  • Point-in-time correct – computed only from what was known at the moment of the prediction – see Point-in-Time Correctness.

In a library, this makes features possible such as a reader's punctuality rate derived from all their previous loans, a seasonal demand score per genre derived from a time series, or the average reading time per genre derived from pairs of book-borrowed and book-returned events.

Kinds of Event-Based Features#

Simple counts and averages often miss the context in which things happen. Events allow features that capture not only what happened, but also when and in relation to what:

  • Lag features measure the time since a key event – for example, the days since a reader's last loan.
  • Frequencies count specific event types within a sliding window – for example, late returns in the past six months.
  • Derived states compute a status from past events – for example, the number of active loans or the amount of outstanding fees.
  • Sequences keep the order of events for models that work on sequences, capturing patterns such as repeated extensions followed by a late return.
  • Context joins combine events with external data, such as a calendar of school holidays.

These kinds can be combined, depending on what the model needs. Each of them has to respect the same boundary: only what was known at the moment of the prediction. For context joins, this includes the external data – a holiday calendar published after the fact is no problem, a forecast revised later may be.

Example: Predicting Late Returns#

A model that predicts, at the moment of borrowing, whether a book will be returned late could use:

  • Lag: days since the reader's last loan
  • Frequency: number of late returns in the past six months
  • Derived state: number of currently active loans
  • Seasonality: whether the loan period overlaps with the summer holidays
  • Relation: how the demand for similar titles has developed recently

Together, these give the model a far more context-aware view than a single aggregate could – and each of them can be explained to a librarian in one sentence.

The Role of Domain Experts#

Feature engineering is not just a matter of transforming numbers. Two features can look similar in the statistics, while one is a strong predictor and the other is noise. Domain experts can:

  • Identify the signals that are likely to matter
  • Interpret what a feature means and when it might mislead
  • Make sure that features align with how the business actually works

Because events are named in the language of the domain, this conversation is easier than with columns of an operational database. It also lays the ground for explainable AI: every input of a model can be traced back to the events that produced it. Whether a feature actually influences an outcome, or only happens to correlate with it, is a question for Causal Inference from Event Streams.

Reproducible Training Data#

A dataset derived from events is versioned by nature: it is defined by the events it was computed from and the logic that computed it. This makes it possible to:

  • Run experiments repeatedly under identical conditions
  • Compare models trained at different points in time
  • Attribute differences in performance to the model, not to changes in the data

Not every dataset has to end in a deployed model. The same data supports statistical analyses, correlation studies, and backtesting – activities that often reveal patterns which lead to new features or sharper business rules.

Designing for Flexibility#

A feature that seems irrelevant today may become important tomorrow. With an event history, you can always derive new features from old events, iterate on a model without collecting new data, and experiment safely, because the source never changes.

A few habits help to make the most of it:

  • Start with a hypothesis for why a feature should improve the prediction.
  • Validate it against a baseline model to confirm that it adds value.
  • Keep the feature logic modular, so that features can evolve independently of each other.
  • Document how each feature is derived, for transparency and audits.

Features turn an event history from a record of facts into the input of learning systems. How their predictions become actions – and how those actions flow back into the history – is described in Closing the Loop.