Post Title

Sree Hari Subhash • September 27, 2026

Most data platforms don't fail because of missing tools. They fail because nobody trusts the numbers. The medallion architecture is a simple, proven way to fix that on Azure Databricks.

Bronze: land everything, change nothing

Ingest raw files and streams into Delta tables exactly as they arrive, with ingestion timestamps and source metadata. Auto Loader handles incremental file discovery, so new data is picked up without custom bookkeeping.


Silver: clean, conform, and deduplicate

This is where data types are enforced, bad records are quarantined, and entities from different systems are joined into a consistent model. Lakeflow Spark Declarative Pipelines (formerly Delta Live Tables) let you express these steps declaratively, with built-in data quality expectations.


Gold: model for the business

Aggregate silver data into tables shaped around real questions: revenue by region, customer lifetime value, inventory risk. These are the tables Power BI, Databricks SQL, and ML models should read from.


Govern it once with Unity Catalog

Catalogs, schemas, and fine-grained permissions sit on top of all three layers, and lineage shows exactly which gold metric came from which source file. That makes audits, impact analysis, and access reviews far easier.


Our takeaway

Start small. Pick one high-value domain, build bronze-to-gold end to end, prove the numbers match the business, then expand. A working slice beats a perfect blueprint.

Analytics dashboard illustrating data science driving business decisions
By Sree Hari Subhash • September 27, 2026
Why the most valuable data science projects start with a decision, not a dataset, and how to get from prototype to measurable business impact.
Abstract artificial intelligence visualization representing machine learning in production
By Sree Hari Subhash • September 27, 2026
A practical path for moving machine learning models out of notebooks and into reliable, monitored production with MLflow, Databricks, and Azure.
Code on a screen representing reliable data engineering pipelines
By Sree Hari Subhash • September 27, 2026
Five practical habits that keep Azure and Databricks data pipelines reliable: incremental loads, data quality checks, orchestration, CI/CD, and cost tuning.