{"id":3236,"date":"2026-10-08T11:45:25","date_gmt":"2026-10-08T11:45:25","guid":{"rendered":"https:\/\/www.examtopics.info\/blog\/microsoft-dp-600-medallion-architecture-in-microsoft-fabric\/"},"modified":"2026-10-08T11:45:25","modified_gmt":"2026-10-08T11:45:25","slug":"microsoft-dp-600-medallion-architecture-in-microsoft-fabric","status":"publish","type":"post","link":"https:\/\/www.examtopics.info\/blog\/microsoft-dp-600-medallion-architecture-in-microsoft-fabric\/","title":{"rendered":"Microsoft DP-600: Medallion Architecture in Microsoft Fabric"},"content":{"rendered":"<h2>Microsoft DP-600: Medallion Architecture in Microsoft Fabric<\/h2>\n<p>Medallion architecture organizes analytical data into progressively refined layers, usually called bronze, silver, and gold. The labels are simple, but the value comes from the contracts between them. Raw source evidence is preserved, quality and conformance are applied in reusable intermediate data, and curated business structures are published for analytics. Microsoft recommends the pattern for Fabric lakehouse implementations because it fits naturally with OneLake, Delta tables, pipelines, notebooks, warehouses, and semantic models.<\/p>\n<p>The current <a href=\"https:\/\/www.examtopics.info\/dp-600\">DP-600<\/a> scope covers lakehouses, warehouses, data preparation, star schemas, semantic models, governance, and lifecycle management, while <a href=\"https:\/\/www.examtopics.info\/dp-700\">DP-700<\/a> focuses more heavily on ingestion, transformation, security, monitoring, and optimization. Medallion architecture provides a useful way to connect those responsibilities without forcing every workload into the same compute engine.<\/p>\n<p>A strong Fabric medallion design is not simply three folders with different names. Each layer should have a purpose, ownership, quality expectation, access model, and rebuild strategy.<\/p>\n<h3>Use bronze to preserve source-aligned evidence<\/h3>\n<p>The bronze layer is the landing zone for raw or minimally changed source data. Its main purpose is recoverability and traceability. If a downstream transformation is later found to be wrong, engineers should be able to rebuild silver and gold from a durable representation of what the source actually delivered.<\/p>\n<p>Preserving source fidelity does not mean bronze must be chaotic. Add ingestion metadata such as source identifier, load timestamp, batch or event identifier, file name, and schema version. Those fields make it possible to reconcile a downstream row with the arrival that produced it and to distinguish a source correction from a transformation defect.<\/p>\n<p>Where Fabric shortcuts can provide governed access to data already stored in OneLake-compatible or supported external locations, copying may not always be necessary. The architecture decision should consider source ownership, retention, performance, and whether a durable local landing is needed for replay.<\/p>\n<p>Bronze retention should be long enough to support the rebuild and audit requirements of the domain. Keeping everything forever is rarely necessary, but deleting raw evidence too early can make downstream corrections impossible. Define retention with legal, operational, and cost requirements rather than an arbitrary storage period.<\/p>\n<p>Schema drift should be visible in bronze even when the pipeline can technically ingest it. Record new or missing fields and route breaking changes for review. Silent acceptance may preserve bytes while still leaving silver transformations operating on an outdated assumption.<\/p>\n<p>Bronze should make replay possible. Preserve source identifiers, ingestion time, batch or file lineage, and enough original structure to reconstruct what arrived even if downstream rules later change. When source systems issue corrections, decide whether bronze is append-only with correction events or whether controlled updates are permitted; the choice affects auditability and reprocessing. The layer is valuable because it separates \u201cwhat the source delivered\u201d from \u201cwhat the organization currently believes is clean.\u201d<\/p>\n<h3>Use silver for reusable quality and conformance<\/h3>\n<p>Silver is where raw values become reliable analytical entities. Common work includes schema enforcement, type conversion, deduplication, null handling, reference-data mapping, standardization, identity resolution, and business-rule validation. The output should be detailed enough for multiple downstream uses rather than tailored to one report.<\/p>\n<p>Quarantine invalid records instead of silently dropping them when the business needs evidence of what failed. Store rejection reason, source identifier, rule version, and processing time. Quality metrics such as duplicate rate, null rate, invalid-code rate, or late-arrival count can reveal source drift before it contaminates gold data.<\/p>\n<p>Silver is also where cross-system conformance often belongs. Customer identifiers from CRM and billing systems may be resolved into a shared key, currencies standardized, and product codes mapped to a common hierarchy. This creates a reusable analytical foundation instead of making every report repeat the same reconciliation logic.<\/p>\n<p>Data contracts can make the silver boundary more explicit. Define required columns, data types, uniqueness, allowed values, timeliness, and ownership. Producers can then be measured against a stable expectation, and downstream consumers know which guarantees they can rely on.<\/p>\n<p>Do not hide all errors through aggressive cleansing. Some invalid values carry business meaning and should be surfaced to source owners. Silver should improve trust while preserving enough lineage to explain what was corrected, rejected, or inferred.<\/p>\n<p>Conformance rules belong here when multiple consumers need the same interpretation. Standardize identifiers, units, reference data, deduplication, slowly changing dimensions, and domain-specific validity rules once rather than rebuilding them inside every semantic model. Record rejected or quarantined records with a reason code so quality failures are observable and recoverable instead of disappearing during transformation.<\/p>\n<h3>Use gold to publish business-ready data products<\/h3>\n<p>Gold organizes data around consumption. It can contain dimensional models, curated facts and dimensions, aggregates, data marts, or other structures optimized for reporting and analytical workloads. Business meaning should be explicit: measures, grain, keys, and definitions are designed around questions users actually need to answer.<\/p>\n<p>A star schema is often a strong fit for Power BI because clear fact and dimension tables simplify relationships and improve semantic-model performance. Gold should not become a dumping ground for every transformed column. Publish the structures that deserve a stable business contract and keep exploratory or highly detailed data in earlier layers when appropriate.<\/p>\n<p>Define grain before building each gold fact. \u201cOne row per order line\u201d or \u201cone row per account per day\u201d sounds basic, but ambiguous grain is a major source of double counting and confusing semantic models. Dimensions should carry stable business attributes while measures and aggregates align to the intended analytical questions.<\/p>\n<p>Multiple gold products can legitimately serve different domains. Finance may need a controlled monthly ledger-oriented model while operations needs near-real-time process metrics. Reuse conformed silver entities where possible, but do not force every consumer into one oversized universal mart.<\/p>\n<p>The relationship between Fabric and reporting described in <a href=\"https:\/\/www.examtopics.info\/blog\/comparing-microsoft-fabric-and-power-bi-features-benefits-and-use-cases-explained\/\">Microsoft Fabric and Power BI<\/a> becomes concrete at the gold boundary. Curated Delta tables or warehouse objects can feed semantic models through Direct Lake, Import, or DirectQuery depending on the workload.<\/p>\n<h3>Choose lakehouse and warehouse engines by workload<\/h3>\n<p>Medallion architecture does not require every layer to use the same Fabric item type. Microsoft\u2019s guidance supports lakehouses for each layer or combinations in which bronze and silver use lakehouses while gold uses a warehouse. SQL-first structured workloads may also use warehouse-oriented patterns.<\/p>\n<p>Lakehouses are attractive for open Delta storage, Spark processing, large-scale files, semistructured data, and data-science integration. Warehouses are attractive when teams need a strongly relational SQL development experience, T-SQL objects, stored procedures, and warehouse-oriented serving patterns. OneLake provides the common storage foundation that allows Fabric engines to interoperate more closely than separate traditional platforms.<\/p>\n<p>Choose the engine that matches the transformation and consumption contract. Do not add a warehouse simply because gold is \u201csupposed\u201d to be SQL, and do not force SQL analysts into Spark for structured transformations when a warehouse fits better.<\/p>\n<p>A medallion design does not require one engine for every layer. A lakehouse can suit large-scale engineering and Delta-based processing, while a Fabric warehouse can provide a strong SQL serving surface for a curated gold domain. The important boundary is the data product contract, not loyalty to one storage interface. Teams can combine engines when lineage, ownership, and refresh behavior remain explicit.<\/p>\n<h3>Standardize on Delta where reliability and interoperability matter<\/h3>\n<p>Fabric lakehouse tables use Delta Lake as the default table format. Delta combines Parquet data with transaction logs and metadata that support ACID behavior, schema management, history, and reliable batch or streaming updates. Those properties make it well suited to data that moves through medallion layers.<\/p>\n<p>Transactions do not replace business logic. Engineers still need to decide how late records are handled, how source corrections update silver, whether gold facts are restated, and how slowly changing dimensions are represented. Delta gives the storage layer a consistent foundation on which those rules can operate.<\/p>\n<p>Optimize file layout as data grows. Too many small files and poorly shaped tables can increase processing and query overhead. Since semantic models such as Direct Lake consume Delta data directly, physical data engineering choices can influence report performance downstream.<\/p>\n<h3>Move data between layers with explicit transformation contracts<\/h3>\n<p>Fabric offers multiple transformation options: notebooks and Spark, Dataflow Gen2, SQL, dbt jobs, materialized lake views, and other workload-specific tools. The medallion pattern should not prescribe one tool universally. What matters is that each transformation has defined inputs, outputs, quality checks, and failure behavior.<\/p>\n<p>Pipelines can orchestrate movement and transformation across those components. A bronze ingestion may complete, a silver notebook may validate and merge changes, a gold SQL process may build dimensions, and a semantic model may update after all data-quality gates pass. The architecture remains understandable because each stage owns a clear boundary.<\/p>\n<p>The broader data-engineering foundations in <a href=\"https:\/\/www.examtopics.info\/blog\/data-engineering-for-absolute-beginners\/\">data engineering<\/a> still apply: reliable pipelines, idempotent writes, metadata, monitoring, and documented contracts matter more than the names assigned to the layers.<\/p>\n<h3>Design security and workspaces around layer responsibilities<\/h3>\n<p>Bronze often contains the most detailed and sensitive source data, while gold may expose a controlled subset for a broader business audience. Access should therefore become more intentional as data progresses. Do not assume that every user who can read a gold report should also browse raw bronze records.<\/p>\n<p>Separate workspaces can provide stronger governance and deployment boundaries between layers, teams, or domains. Microsoft guidance notes that layer separation can improve control. The right structure depends on ownership, scale, compliance, and operational model, but permissions should follow the data product rather than convenience.<\/p>\n<p>Apply sensitivity labels, row or object controls, and item permissions where appropriate. Security transformations such as masking or removal of sensitive fields can occur before gold publication, but they should complement\u2014not replace\u2014source and workspace access controls.<\/p>\n<h3>Operate the layers with quality, lineage, and freshness metrics<\/h3>\n<p>Monitor the flow between layers as an operational system. Track ingestion freshness, processing duration, row counts, quality-rule failures, late records, rejected records, and the timestamp of the latest successfully published gold data. A green pipeline is insufficient if the output contains half the expected transactions.<\/p>\n<p>Lineage should let operators identify which bronze source and transformation version produced a silver or gold artifact. This matters for impact analysis, debugging, audits, and schema changes. If a source column changes, teams should know which downstream tables and semantic models depend on it before deployment.<\/p>\n<p>Quality expectations should increase across the layers. Bronze may permit source defects because it preserves evidence. Silver should enforce reusable technical and business validity. Gold should publish only data that meets the contract required by its consumers. That progressive quality is the real meaning of the medallion model.<\/p>\n<p>Operational metrics should be layer-specific. Bronze can track source arrival completeness and ingestion lag; silver can track duplicate rate, rejected records, conformance failures, and reconciliation to source totals; gold can track freshness, semantic test results, and consumer-facing service objectives. When a dashboard is stale, those metrics show whether the cause began at ingestion, transformation, or publication instead of forcing operators to inspect every pipeline manually.<\/p>\n<p>Plan for late data and backfills across the whole chain. A corrected bronze partition may require a deterministic silver rebuild and a targeted gold refresh. Document the dependency path and make the reprocessing scope visible before execution so an operator understands which downstream products will change. Reliable medallion architecture is as much about repeatable repair as it is about the normal forward flow.<\/p>\n<h3>Design gold for semantic consumption rather than copying reporting logic<\/h3>\n<p>Gold tables should support reusable analysis, while the semantic model adds business calculations, relationships, hierarchies, and presentation logic. Avoid embedding every report-specific measure into the storage layer or recreating the same transformation independently in each semantic model.<\/p>\n<p>The skills in <a href=\"https:\/\/www.examtopics.info\/blog\/dp-600-exam-focus-designing-and-building-semantic-models-for-analytics-engineers\/\">enterprise semantic-model design<\/a> complement medallion architecture because a clean gold layer is only one part of the serving stack. Storage mode, DAX design, relationships, security, and report usage still determine the final analytics experience.<\/p>\n<p>Direct Lake can make the boundary especially efficient because a semantic model can consume Delta tables in OneLake without a traditional import copy. That does not eliminate the need to design gold grain, star schemas, or physical table layout. It makes those upstream choices even more visible to the reporting layer.<\/p>\n<h3>Keep the medallion pattern adaptable to the domain<\/h3>\n<p>Bronze, silver, and gold are design responsibilities, not a requirement to create exactly three physical items for every source. Some domains may need an additional quarantine zone, a highly restricted sensitive-data layer, several gold products, or streaming paths that converge with batch data. The architecture should remain simple enough to explain while adapting to real requirements.<\/p>\n<p>Domain ownership can also coexist with the pattern. Sales, finance, operations, and customer-service teams may each maintain their own medallion data products within governed Fabric domains. Shared reference data and enterprise definitions should be managed deliberately so decentralization does not recreate incompatible silos.<\/p>\n<p>For teams in the <a href=\"https:\/\/www.examtopics.info\/microsoft-exams\">Microsoft<\/a> ecosystem, medallion architecture is valuable because it creates explicit quality boundaries across Fabric rather than because bronze, silver, and gold are fashionable labels. Preserve recoverable raw data, publish reusable conformed data, curate business-ready products, choose the right Fabric engine per layer, secure each boundary, and operate the flow with measurable quality. That is what turns OneLake into a dependable analytical platform.<\/p>\n<p>Use the layers as responsibilities rather than as a rigid requirement to copy every record three times. Some domains may need multiple silver products, a warehouse-backed gold layer, or shortcuts to governed data that is already curated elsewhere. The pattern remains useful when bronze preserves evidence, silver establishes reusable trust, and gold publishes consumer-ready meaning\u2014even if the physical implementation varies by workload.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Microsoft DP-600: Medallion Architecture in Microsoft Fabric Medallion architecture organizes analytical data into progressively refined layers, usually called bronze, silver, and gold. The labels are simple, but the value comes from the contracts between them. Raw source evidence is preserved, quality and conformance are applied in reusable intermediate data, and curated business structures are published [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12,1],"tags":[],"class_list":["post-3236","post","type-post","status-publish","format-standard","hentry","category-ai-data","category-uncategorized"],"_links":{"self":[{"href":"https:\/\/www.examtopics.info\/blog\/wp-json\/wp\/v2\/posts\/3236","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.examtopics.info\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.examtopics.info\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.examtopics.info\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.examtopics.info\/blog\/wp-json\/wp\/v2\/comments?post=3236"}],"version-history":[{"count":0,"href":"https:\/\/www.examtopics.info\/blog\/wp-json\/wp\/v2\/posts\/3236\/revisions"}],"wp:attachment":[{"href":"https:\/\/www.examtopics.info\/blog\/wp-json\/wp\/v2\/media?parent=3236"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.examtopics.info\/blog\/wp-json\/wp\/v2\/categories?post=3236"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.examtopics.info\/blog\/wp-json\/wp\/v2\/tags?post=3236"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}