Concepts / Analyzing / Causal Inference from Event Streams
Analyzing

Causal Inference from Event Streams

This guide explains how an event history helps to tell cause and effect apart from mere correlation. It covers why causality matters for decisions, what an event history contributes, and how a typical analysis proceeds.

Correlations are easy to find. Causation is what decisions actually depend on. If late fees and frequent borrowing tend to occur together, does one lead to the other – or are both driven by something else, such as a seasonal promotion? Acting on a correlation alone can lead to interventions that achieve nothing, or even do harm.

Causal inference is the discipline of determining whether one thing actually influences another, and by how much. It marks the step from describing what happened to knowing what to do.

Why Causality Matters#

Many of the questions a business asks are causal questions in disguise:

  • Did sending reminders actually reduce late returns?
  • Did the new loan policy increase how many readers stay members?
  • Did promoting certain genres change what people borrow?

Each of these asks what would have happened without the intervention – something that can never be observed directly. Causal inference estimates it, by comparing what did happen with a carefully chosen reference. Without understanding the direction and mechanism of an effect, decisions rest on guesswork.

What an Event History Contributes#

Causal analysis depends on accurate sequences and context. An event history provides both:

  • Precise timelines. A cause precedes its effect. Because every event has its place in an unchangeable order, it is clear what happened before what.
  • Full context. Not only the event of interest is recorded, but also everything around it – which makes it possible to account for other factors that might explain an outcome.
  • Recorded interventions. When decisions and actions are stored as events of their own – reminder-sent, loan-policy-changed, genre-promoted – it is known exactly who was affected, and when. An intervention that was never recorded cannot be analyzed later.
  • Reconstructable states. Because the state at any past moment can be rebuilt, groups can be compared as they were at the time of an intervention, not as they are today.

With this, established methods – such as difference-in-differences, propensity score matching, or causal graphs – can be applied to clean, reliable data.

Example: Do Reminders Work?#

Suppose a library introduces a reminder two days before the due date. Over the following months, late returns go down. Is this the reminder's doing, or the result of something else – shorter loan periods, a quieter season, a different mix of readers?

With the event history, the analysis can proceed step by step:

  1. Identify the groups. Which loans were followed by a reminder-sent event, and which were not?
  2. Account for other factors. Compare loans that are similar in book type, season, and reader history, so that these differences cannot explain the result.
  3. Compare the outcomes. Measure the late-return rate of both groups over the same period.

If the history also covers the time before the reminders were introduced, the development of both groups can be compared before and after – the idea behind difference-in-differences. This structured approach replaces an impression with evidence.

Experiments and What-If Questions#

The most reliable way to establish a cause is an experiment: the intervention is applied to a randomly chosen group only, and the outcomes of both groups are compared. Event Sourcing supports this well, because the assignment to a group can itself be recorded as an event, and every outcome can be traced back to it.

Where an experiment is not possible, the history allows what-if analyses of a limited but useful kind. Replaying past events with a different decision rule shows exactly whom a rule would have affected – for example, which readers a stricter reminder policy would have reached. What those readers would then have done is not in the history; it has to be estimated with the methods above.

Causality and Machine Learning#

Causal insights also make models better:

  • They help to identify features that actually influence outcomes, rather than features that merely correlate with them – see Features from Events.
  • They support policy evaluation, so that interventions triggered by predictions are backed by evidence – see Closing the Loop.
  • They make simulations of changes more trustworthy before those changes reach production.

Models built on causally validated features tend to be more stable and trustworthy, because they rely less on patterns that break as soon as circumstances change.

Good Practices#

  • Combine domain expertise with statistics. Domain experts know which factors could explain an outcome; methods alone cannot find what nobody thought to account for.
  • Validate with more than one method. If different approaches point in the same direction, the conclusion is stronger.
  • Keep event definitions clear. When an event type changes its meaning over time, causal relationships blur – see Versioning Events.
  • Record decisions as events. Every intervention that is stored explicitly can be evaluated later.
  • Revisit results. New policies and changing behavior can shift relationships, so causal analysis is an iterative process.

An event history holds not just data, but the narrative of a domain. That makes it possible to uncover what really drives the numbers – and to design interventions that reliably change them.