Concepts / Integrating / Data Mesh and Data Products
Integrating

Data Mesh and Data Products

This guide explains Data Mesh, an approach to sharing data across an organization, and the data products it is built on. It shows why event-sourced systems are a natural fit for it, and how to get started without heavy tooling.

In many organizations, data flows into a central data team that builds and maintains the pipelines for everyone else. As the number of sources and questions grows, this team becomes a bottleneck, and it is responsible for data it does not really understand – the meaning of a field lives with the team that produces it, not with the team that transports it.

Data Mesh turns this around: the teams that own a part of the business also own the data it produces, and they share it as products that others can find, understand, and rely on.

The Core Principles#

Data Mesh is usually described by four principles:

  • Domain-oriented ownership. The team that owns a business domain also owns the data that domain produces, including its quality and documentation.
  • Data as a product. Data is shared with the same care as any other product: with a clear purpose, a defined quality, documentation, and versioning.
  • Self-serve infrastructure. A common platform lets teams publish, discover, and consume data products without waiting for a central team.
  • Federated governance. Global standards – for formats, identifiers, security, and privacy – ensure that data products work together, without centralizing every decision.

What Makes a Data Product#

A data product is more than a dataset that happens to be accessible. It is shared on purpose, and it comes with what a consumer needs to use it without asking:

  • An owner who is responsible for it and can be asked about it
  • A description of what it contains, what each field means, and how it is derived
  • Quality commitments, such as how fresh, complete, and accurate it is
  • A stable interface through which it can be accessed
  • A version, so that it can evolve without breaking its consumers

Consumers should be able to find it, understand it, and trust it – without reverse-engineering someone else's database.

Why Event Sourcing Fits#

Event-sourced systems are unusually well prepared for Data Mesh:

  • Clear semantics. Events are named in the language of the domain, so the data carries its meaning with it.
  • A reliable source. The event history is immutable and complete, which makes it a sound foundation for data products.
  • Rebuildable products. A data product derived from events can be regenerated at any time – with new logic, over the entire history, or exactly as it was at a specific point in time.
  • Natural ownership. The team that owns the events already owns the knowledge of what they mean.

From Events to Data Products#

With Event Sourcing, a data product typically comes into being in a few steps:

  1. Record domain events in the event store.
  2. Build projections tailored to the questions consumers ask – see Designing Analytical Projections.
  3. Publish them as data products, with documentation, quality commitments, and a version.
  4. Offer an interface through which other teams can access them – a query endpoint, files, or a stream of events.

The events themselves can be a data product too. A carefully designed stream of integration events is exactly what Data Mesh calls a source-aligned data product: the facts of a domain, shared on purpose – see Domain Events and Integration Events.

Example: A Library#

The lending team of a library owns the events around loans. From them, it could offer several data products:

  • Borrowing history per reader, for analyses of how readers engage with the library
  • Demand per title over time, for forecasting and acquisitions
  • Overdue patterns, for improving the service and targeting reminders

A single dataset, such as weekly borrowing figures per title enriched with seasonal demand per genre, can serve several consumers at once: reading recommendations, demand forecasts, and long-term research. Because it is defined once and owned by the team that understands it, everyone works with the same numbers and the same definitions.

Governance and Privacy#

Sharing data widely makes some rules non-negotiable. Federated governance means agreeing on them once and applying them everywhere:

  • Common identifiers, so that data products of different teams can be combined
  • Common formats and metadata, so that tooling works across products
  • Access rules and privacy protection, so that personal data is shared only where it is allowed – usually not at all, in pseudonymized or aggregated form instead – see GDPR Compliance

Start Small#

Many discussions of Data Mesh revolve around organizational charts and platforms. None of that is needed to begin. The core idea is simple: provide high-quality, domain-owned data in a form others can easily use.

An event-sourced system already has most of what it takes. Pick one question another team keeps asking, build a projection that answers it, document it, give it an owner and a version – and you have your first data product. The rest can grow as the demand for it does.