Series "Data Architectures: From Warehouse to Mesh"

Explains why the main architectures emerged, what problem each one solves, what it costs in complexity and organization, and how to choose a proportionate starting point.

Series "Data Architectures: From Warehouse to Mesh"

Who should read it?

Data and analytics leaders, architects, engineers, product owners, and technically curious business readers who need to choose an architecture without being trapped by vendor language.

Important

There is no universally best architecture. Architecture follows use case, workload, data maturity, skills, governance needs, latency, budget, and organizational structure.

Series

  1. Your Data Architecture Is a Business Decision, Not a Shopping List
  2. Data Warehouse, Data Lake, or Both? The Architecture Timeline That Makes It Clear
  3. Data Fabric Is Not Magic: What It Adds to a Modern Data Warehouse
  4. The Data Lakehouse Promise: One Repository, Fewer Pipelines, New Trade-offs
  5. Data Mesh Is an Operating Model, Not a Tool You Can Install
  6. Data Architecture Decision Matrix: Choose the Smallest System That Can Work

What we cover?

Article 1 - Your Data Architecture is a Business Decision

The architecture decision starts before the technology.

Main Topics:

  • Data as a strategic asset
  • Data architecture as a high-level blueprint for data, technology, and flow
  • OLTP versus OLAP
  • Schema-on-write versus schema-on-read
  • The five common elements of an architecture: sources, ingestion, storage, processing, and consumption
  • The architecture design session as a practical way to align business, technical, and organizational concerns
  • Why architecture should be designed from use cases and constraints, not from fashionable products

Key takeaway:

Write down the users, decisions, latency, data types, quality expectations, compliance constraints, and team capabilities before selecting tools.

Article 2 - Data Warehouse, Data Lake, or both?

The evolution from relational systems to modern warehouses.

Main topics:

  • relational databases for operational transactions
  • relational data warehouses for analytical reporting
  • data lakes for raw, flexible, multi-format storage
  • schema-on-write and schema-on-read as a trade-off between control and flexibility
  • ETL, ELT, and data virtualization
  • relational, dimensional, and Data Vault modeling
  • top-down, bottom-up, and hybrid implementation approaches
  • the Modern Data Warehouse as a combination of lake flexibility and warehouse serving, security, and compliance

Key takeaway:

a data lake is not automatically an analytics platform, and a warehouse is not automatically obsolete. Each belongs to a different part of the journey.

Article 3 - Data Fabric is not Magic

Modern Data Warehouse and Data Fabric.

Main topics:

  • the Modern Data Warehouse data journey: ingestion, storage, transformation, modeling, and visualization
  • separation of storage and compute
  • the role of the data lake, relational serving layer, and sandbox
  • how Data Fabric adds data access policies, metadata catalog, lineage, MDM, virtualization, APIs, and real-time processing
  • why governance is part of architecture, not paperwork added at the end
  • the people and skills required to make discoverability and self-service real
  • when the additional platform complexity is justified

Key takeaways:

A fabric creates value only when users can find, understand, access, and trust data. More components do not automatically create more value.

Article 4 - The Data Lakehouse Promise

Data Lakehouse

Main topics:

  • why organizations wanted to reduce duplicated warehouse and lake storage
  • a lakehouse as lake storage plus a transactional table layer
  • Delta Lake, Apache Iceberg, and Apache Hudi as examples of table formats or transactional layers, not storage by themselves
  • ACID transactions, schema enforcement, time travel, compaction, and unified batch and streaming pipelines
  • the difference between physical storage and a relational serving layer
  • dashboard latency and concurrency as reasons a relational serving layer may still be needed
  • operational and skills trade-offs, including ecosystem compatibility and governance

Key takeaways:

A lakehouse simplifies some architecture decisions, but it does not eliminate modeling, serving, governance, or operational work.

Article 5 - Data Mesh Is an Operating Model, Not a Tool You Can Install

Data Mesh and the organizational choice.

Main topics:

  • data mesh as a decentralized architecture and operating model
  • the four principles: domain ownership, data as a product, self-serve data infrastructure, and federated computational governance
  • why decentralizing ownership changes incentives, skills, interfaces, and accountability
  • the difference between federated ownership and uncontrolled data silos
  • myths and concerns: organizational maturity, duplicated capabilities, interoperability, standards, and governance
  • when a mesh is appropriate, and when a well-governed centralized or hybrid model is the better choice
  • a practical readiness test based on domains, product thinking, platform capability, governance, and leadership support

Key takeaways:

Data mesh is not a shortcut around central data problems. It moves responsibility closer to domains and therefore requires stronger product discipline and coordination.

Article 6 - Data Architecture Decision Matrix

Comparison of relational warehouse, data lake, Modern Data Warehouse, Data Fabric, Data Lakehouse, and Data Mesh. Score options against workload, data types, latency, governance, scale, team skills, organization, cost, and maturity.