# Causal Inference from Event Streams

This guide explains how an event history helps to tell **cause and effect** apart from mere correlation. It covers why causality matters for decisions, what an event history contributes, and how a typical analysis proceeds.

Correlations are easy to find. **Causation** is what decisions actually depend on. If late fees and frequent borrowing tend to occur together, does one lead to the other – or are both driven by something else, such as a seasonal promotion? Acting on a correlation alone can lead to interventions that achieve nothing, or even do harm.

Causal inference is the discipline of determining **whether one thing actually influences another**, and by how much. It marks the step from describing what happened to knowing what to do.

## Why Causality Matters

Many of the questions a business asks are causal questions in disguise:

- *Did sending reminders actually reduce late returns?*
- *Did the new loan policy increase how many readers stay members?*
- *Did promoting certain genres change what people borrow?*

Each of these asks what would have happened **without** the intervention – something that can never be observed directly. Causal inference estimates it, by comparing what did happen with a carefully chosen reference. Without understanding the **direction and mechanism** of an effect, decisions rest on guesswork.

## What an Event History Contributes

Causal analysis depends on accurate sequences and context. An event history provides both:

- **Precise timelines.** A cause precedes its effect. Because every event has its place in an unchangeable order, it is clear what happened before what.
- **Full context.** Not only the event of interest is recorded, but also everything around it – which makes it possible to account for other factors that might explain an outcome.
- **Recorded interventions.** When decisions and actions are stored as events of their own – `reminder-sent`, `loan-policy-changed`, `genre-promoted` – it is known exactly who was affected, and when. An intervention that was never recorded cannot be analyzed later.
- **Reconstructable states.** Because the state at any past moment can be rebuilt, groups can be compared **as they were** at the time of an intervention, not as they are today.

With this, established methods – such as **difference-in-differences**, **propensity score matching**, or **causal graphs** – can be applied to clean, reliable data.

## Example: Do Reminders Work?

Suppose a library introduces a **reminder two days before the due date**. Over the following months, late returns go down. Is this the reminder's doing, or the result of something else – shorter loan periods, a quieter season, a different mix of readers?

With the event history, the analysis can proceed step by step:

1. **Identify the groups.** Which loans were followed by a `reminder-sent` event, and which were not?
2. **Account for other factors.** Compare loans that are similar in book type, season, and reader history, so that these differences cannot explain the result.
3. **Compare the outcomes.** Measure the late-return rate of both groups over the same period.

If the history also covers the time before the reminders were introduced, the development of both groups can be compared **before and after** – the idea behind difference-in-differences. This structured approach replaces an impression with evidence.

## Experiments and What-If Questions

The most reliable way to establish a cause is an **experiment**: the intervention is applied to a randomly chosen group only, and the outcomes of both groups are compared. Event Sourcing supports this well, because the assignment to a group can itself be recorded as an event, and every outcome can be traced back to it.

Where an experiment is not possible, the history allows **what-if analyses** of a limited but useful kind. Replaying past events with a different decision rule shows exactly **whom a rule would have affected** – for example, which readers a stricter reminder policy would have reached. What those readers would then have done is not in the history; it has to be estimated with the methods above.

## Causality and Machine Learning

Causal insights also make models better:

- They help to identify **features that actually influence outcomes**, rather than features that merely correlate with them – see **[Features from Events](/concepts/features-from-events)**.
- They support **policy evaluation**, so that interventions triggered by predictions are backed by evidence – see **[Closing the Loop](/concepts/closing-the-loop)**.
- They make **simulations** of changes more trustworthy before those changes reach production.

Models built on causally validated features tend to be more **stable and trustworthy**, because they rely less on patterns that break as soon as circumstances change.

## Good Practices

- **Combine domain expertise with statistics.** Domain experts know which factors could explain an outcome; methods alone cannot find what nobody thought to account for.
- **Validate with more than one method.** If different approaches point in the same direction, the conclusion is stronger.
- **Keep event definitions clear.** When an event type changes its meaning over time, causal relationships blur – see **[Versioning Events](/concepts/versioning-events)**.
- **Record decisions as events.** Every intervention that is stored explicitly can be evaluated later.
- **Revisit results.** New policies and changing behavior can shift relationships, so causal analysis is an iterative process.

An event history holds not just data, but the **narrative** of a domain. That makes it possible to uncover what really drives the numbers – and to design interventions that reliably change them.
