Urgent.News

What's breaking now, across thousands of outlets.

Tech

Medallion architecture on Databricks: what actually matters

Every Databricks pitch deck has the same three-layer diagram: bronze for raw data, silver for cleaned data, gold for business-ready tables. The diagram is fine. The problems start when teams treat it as an architecture instead of what it really is — a naming convention for a set of decisions you still have to make. Here's where those decisions actually bite, based on the lakehouse builds we've…

Databricks' medallion architecture consists of three layers: raw data, cleaned data, and business-ready tables. Teams should treat it not as an architecture, but a naming convention for decisions. A common mistake is cleaning data before it reaches the raw layer, as this makes it difficult to reprocess history if parsing logic turns out to be wrong.

The raw layer should be append-only and a schema-on-read record of exactly what the source system sent, with ingestion metadata. The cleaned layer is where a data model is committed, with one row per entity, resolved keys, and enforced schemas. Bad records should be quarantined. Two rules are to have expectations for every silver table and to let it survive a BI-tool migration untouched.

The gold layer can be denormalized, redundant, and shaped for one consumer. Unity Catalog should be implemented from day one for governance. Cost visibility should come before cost optimization, with jobs tagged by pipeline and team. A reprocessing story should be decided before launch, whether it's full replays, partition-scoped replays, or versioned logic.

Getting the boring parts right is the real lesson, as discipline, not technology, determines the success of a medallion architecture.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

More from Wednesday 12 August →