Data Mesh and Data Products
This guide explains Data Mesh, an approach to sharing data across an organization, and the data products it is built on. It shows why event-sourced systems are a natural fit for it, and how to get started without heavy tooling.
In many organizations, data flows into a central data team that builds and maintains the pipelines for everyone else. As the number of sources and questions grows, this team becomes a bottleneck, and it is responsible for data it does not really understand – the meaning of a field lives with the team that produces it, not with the team that transports it.
Data Mesh turns this around: the teams that own a part of the business also own the data it produces, and they share it as products that others can find, understand, and rely on.
The Core Principles#
Data Mesh is usually described by four principles:
- Domain-oriented ownership. The team that owns a business domain also owns the data that domain produces, including its quality and documentation.
- Data as a product. Data is shared with the same care as any other product: with a clear purpose, a defined quality, documentation, and versioning.
- Self-serve infrastructure. A common platform lets teams publish, discover, and consume data products without waiting for a central team.
- Federated governance. Global standards – for formats, identifiers, security, and privacy – ensure that data products work together, without centralizing every decision.
What Makes a Data Product#
A data product is more than a dataset that happens to be accessible. It is shared on purpose, and it comes with what a consumer needs to use it without asking:
- An owner who is responsible for it and can be asked about it
- A description of what it contains, what each field means, and how it is derived
- Quality commitments, such as how fresh, complete, and accurate it is
- A stable interface through which it can be accessed
- A version, so that it can evolve without breaking its consumers
Consumers should be able to find it, understand it, and trust it – without reverse-engineering someone else's database.
Why Event Sourcing Fits#
Event-sourced systems are unusually well prepared for Data Mesh:
- Clear semantics. Events are named in the language of the domain, so the data carries its meaning with it.
- A reliable source. The event history is immutable and complete, which makes it a sound foundation for data products.
- Rebuildable products. A data product derived from events can be regenerated at any time – with new logic, over the entire history, or exactly as it was at a specific point in time.
- Natural ownership. The team that owns the events already owns the knowledge of what they mean.
From Events to Data Products#
With Event Sourcing, a data product typically comes into being in a few steps:
- Record domain events in the event store.
- Build projections tailored to the questions consumers ask – see Designing Analytical Projections.
- Publish them as data products, with documentation, quality commitments, and a version.
- Offer an interface through which other teams can access them – a query endpoint, files, or a stream of events.
The events themselves can be a data product too. A carefully designed stream of integration events is exactly what Data Mesh calls a source-aligned data product: the facts of a domain, shared on purpose – see Domain Events and Integration Events.
Example: A Library#
The lending team of a library owns the events around loans. From them, it could offer several data products:
- Borrowing history per reader, for analyses of how readers engage with the library
- Demand per title over time, for forecasting and acquisitions
- Overdue patterns, for improving the service and targeting reminders
A single dataset, such as weekly borrowing figures per title enriched with seasonal demand per genre, can serve several consumers at once: reading recommendations, demand forecasts, and long-term research. Because it is defined once and owned by the team that understands it, everyone works with the same numbers and the same definitions.
Governance and Privacy#
Sharing data widely makes some rules non-negotiable. Federated governance means agreeing on them once and applying them everywhere:
- Common identifiers, so that data products of different teams can be combined
- Common formats and metadata, so that tooling works across products
- Access rules and privacy protection, so that personal data is shared only where it is allowed – usually not at all, in pseudonymized or aggregated form instead – see GDPR Compliance
Start Small#
Many discussions of Data Mesh revolve around organizational charts and platforms. None of that is needed to begin. The core idea is simple: provide high-quality, domain-owned data in a form others can easily use.
An event-sourced system already has most of what it takes. Pick one question another team keeps asking, build a projection that answers it, document it, give it an owner and a version – and you have your first data product. The rest can grow as the demand for it does.