Concepts / Analyzing / LLMs and Event Streams
Analyzing

LLMs and Event Streams

This guide explains how large language models (LLMs) can work directly with an event history – answering questions in plain language, exploring patterns, and helping to investigate what happened. It shows where this complements the classic pipeline of projections, features, and models, and which guardrails it needs.

The classic way to learn from events is a pipeline: events flow into projections, projections become features, features feed models, and models produce predictions. This works well for automated, repeatable analyses at scale. But it has a price: before a question can be answered, someone has to build the projection for it.

LLMs open up a different way. Instead of building a pipeline first, you can ask the event history a question – in plain language – and get an answer.

Beyond the Pipeline#

The pipeline is designed for repeatability and scale. Once set up, it runs continuously and produces consistent results. It is the right choice for automated decisions, production models, and recurring reports.

But many valuable questions are ad hoc: a product manager looking into an unusual trend, a domain expert checking whether a new pattern has emerged, a developer trying out an analytical idea. Building a projection for each of these questions would take longer than the question is worth.

For these cases, an LLM with access to the event store offers a faster path. You describe what you want to know, the LLM translates your intent into queries, retrieves the relevant events, and presents the result as an answer you can follow up on. This does not replace the pipeline – it complements it. Pipelines handle production workloads; conversation handles exploration and discovery.

Why Events and LLMs Fit Together#

Events are particularly well suited for working with language models:

  • Business language. Events such as book-borrowed, late-fee-charged, or membership-renewed are already expressed in terms a person – and a language model – can understand. There is no translation layer between the data and the question.
  • Facts, not states. Events describe things that actually happened, in order and unchanged. This gives the model a reliable basis for reasoning, instead of mutable state that may have changed since.
  • Context. Each event carries information about what happened, to which subject, and when, which helps the model understand the circumstances, not just the outcome.
  • Narrative structure. An event stream reads like a sequence of things that happened – a story. This aligns well with how language models process sequences and text.

As a result, an LLM does not need extensive preprocessing or feature engineering to work with event data. It can engage with the events as they are.

What Becomes Possible#

When a language model can access an event history, several new ways of working emerge:

  • Ad-hoc exploration – ask questions that nobody anticipated when the projections were designed, without building a new read model first.
  • Conversational analytics – people without technical background explore the data themselves, instead of waiting for a report.
  • Rapid prototyping – test an analytical hypothesis in minutes before investing in a pipeline.
  • Incident investigation – walk through a sequence of events interactively to understand what happened and why.
  • Scenario data – describe a business scenario and let the model write matching events, for tests, simulations, or demos.

Example: Late Returns in a Library#

A librarian notices that more books are being returned late and wants to understand why. With a pipeline, this would mean defining a projection of late returns by genre, period, and reader group, implementing and deploying it, and then interpreting the results.

With a language model connected to the event history, the librarian can simply ask:

"Which genres had the most late returns last quarter? Were there spikes in particular weeks?"

The model translates the question into queries, retrieves the matching events, and answers – perhaps pointing out that one genre stands out in the last two weeks of the quarter, which coincide with the school holidays. The librarian can follow up naturally: "How does that compare to the same period last year?"

If the analysis proves valuable enough to run regularly, it can then be turned into a proper projection. The conversation served as a fast prototype.

Pipeline or Conversation#

Knowing where each approach is strong helps to choose the right one:

  • Purpose – conversation is ideal for exploration and prototypes; pipelines are built for automated decisions at scale.
  • Setup – a language model can answer right away; a pipeline has to be designed and built first.
  • Consistency – answers of a language model can vary from one run to the next; a pipeline produces the same result every time.
  • Audience – conversation serves domain experts, analysts, and developers; pipelines feed systems and dashboards.

Use conversation to discover what matters, and pipelines to operationalize what you have found.

Try it with EventSourcingDB

The MCP server for EventSourcingDB connects language models to an EventSourcingDB instance, so they can read events, list subjects, and run EventQL queries on your behalf – see Introduction to the MCP Server.

Guardrails#

An interface that understands plain language needs the same care as any other interface to your data – and in some respects more:

  • Scope access. Apply the same access controls as for any other client, and grant only what the task requires. Reading is not harmless when events contain personal data – see GDPR Compliance.
  • Govern writing. If a model may write events at all, keep it away from production write paths, validate what it writes, and make every written event clearly attributable. See Versioning Events for keeping event types consistent.
  • Keep an audit trail. Record what was asked, what was queried, and what was returned, so that answers can be traced and checked.
  • Treat answers as hypotheses. Validate important findings through a proper analysis before acting on them.
  • Invest in clear events. The more consistent and well-named the event types are, the better a model can reason about them.

Combining the rigor of pipelines with the flexibility of conversation gives you a system that supports both continuous automation and exploration on demand – for technical and non-technical people alike.