INSIGHTS
AI & Data

Microsoft DP-700: Fabric Lakehouse Data Quality

In this article
  1. Define quality in terms of the data contract
  2. Use the medallion pattern to separate trust levels
  3. Combine Delta reliability with explicit quality rules
  4. Validate schema without assuming schema is quality
  5. Measure quality over time, not only at load time
  6. Design quarantine and replay paths
  7. Connect quality controls with governance and security
  8. Publish quality expectations with the data product

A lakehouse becomes valuable only when people can trust the data that moves through it. Microsoft Fabric gives teams a unified storage and analytics platform, but simply landing data in OneLake does not make the data accurate, complete, timely, or fit for downstream use. Data quality has to be engineered into the lakehouse lifecycle.

That work starts with a useful distinction: storage quality and business quality are not the same thing. Delta tables can provide transactional reliability, schema information, and table history, yet a technically valid row can still contain an impossible quantity, an unknown customer, a duplicated event, or a timestamp that arrived two days late. Engineers need controls at several layers, from ingestion through curation.

This is directly relevant to DP-700, where Microsoft expects data engineers to ingest and transform data, manage analytics solutions, and monitor and optimize them. The goal is not to memorize a list of checks. It is to understand where a quality rule belongs, what should happen when it fails, and how the evidence should be surfaced to operators and consumers.

Define quality in terms of the data contract

“Clean data” is too vague to be operational. A useful quality model translates business expectations into checks that can be measured. Common dimensions include completeness, validity, uniqueness, consistency, freshness, and referential integrity, but the exact rules should come from the purpose of the dataset.

For an orders table, a quality contract might require a non-null order identifier, a positive quantity, a recognized currency code, a valid customer key, and an order timestamp within an expected window. For telemetry, uniqueness may be defined by device plus event time rather than a single key. For slowly changing reference data, late arrival may be acceptable as long as effective dates remain consistent.

Make the contract explicit enough that an engineer can implement it and an analyst can understand what a failure means. A generic “quality score” is less useful than knowing that 0.8 percent of rows were dropped because quantity was negative or that freshness exceeded the agreed threshold by 47 minutes.

Use the medallion pattern to separate trust levels

Fabric supports lakehouse designs that organize data through bronze, silver, and gold layers. The pattern is useful because it separates preservation from validation and business presentation. Bronze normally retains source data with minimal alteration. Silver applies cleansing, deduplication, validation, and conformance. Gold provides curated structures for reporting, analytics, or machine learning.

The important idea is not the layer names. It is the trust boundary. Raw data should be recoverable even when a quality rule rejects it from a curated table. If the only copy of an invalid row is deleted during cleansing, investigation becomes difficult. Preserving source evidence allows a team to reproduce the failure, repair logic, and replay affected data.

The existing overview of Microsoft Fabric and Power BI helps frame where this curated data eventually goes. Power BI consumers generally should not have to interpret raw ingestion errors. The lakehouse layers should absorb that complexity and publish data at an appropriate quality level for the semantic model or report.

Combine Delta reliability with explicit quality rules

Fabric lakehouses use Delta Lake as the default table format, which provides ACID transactions, transaction logs, table history, and support for both batch and streaming patterns. Those capabilities help protect the physical and transactional integrity of the data, but they do not replace domain rules.

Fabric materialized lake views add a useful quality mechanism for declarative lakehouse transformations. Constraints can express Boolean conditions that rows must satisfy, and a violation can either fail the refresh or drop the offending row, depending on the intended policy. That makes the quality decision part of the transformation definition rather than an informal check that runs somewhere else.

Choosing between fail and drop is a business decision. A failed refresh is appropriate when publishing any bad row would make the result unsafe, when completeness is mandatory, or when a violation indicates broken upstream logic. Dropping may be reasonable when a small number of independently bad records should be quarantined while the rest of the dataset remains useful. In either case, the decision should be visible in monitoring.

Validate schema without assuming schema is quality

Schema validation answers important questions: are required columns present, are data types compatible, and has the source introduced a structural change? But a correct schema does not prove that values are meaningful. A string can contain an invalid product code, an integer can contain a negative quantity, and a timestamp can be structurally valid but outside the acceptable business window.

Use schema checks as the first gate, then apply domain checks. Some teams also separate hard and soft failures. A missing primary key may block the dataset, while an optional demographic field with a higher-than-normal null rate may trigger an alert without stopping the refresh. This keeps quality controls proportional to the impact of the defect.

Engineers new to the discipline can connect these practices with broader data engineering fundamentals: ingestion is only one stage. Reliable data products need transformation, validation, storage, observability, and repeatable recovery.

Measure quality over time, not only at load time

One successful validation run does not establish long-term trust. Quality drifts as source systems change, user behavior changes, integrations are replaced, or business definitions evolve. Track quality metrics as time series so that gradual degradation becomes visible.

Useful measures include null rate, duplicate rate, invalid-domain rate, referential-integrity failures, rows dropped by a constraint, late-arrival percentage, freshness, and reconciliation differences between source and destination. Fabric materialized lake views provide a data quality report that can surface violation counts and trends, while other quality logic can be recorded to operational tables or monitoring systems.

Thresholds should be tied to business tolerance. A customer-contact dataset might allow a small percentage of missing optional phone numbers but zero duplicate customer identifiers. A financial fact table may tolerate late source delivery but not an unreconciled total. Avoid one universal quality score that hides which rule is actually failing.

Design quarantine and replay paths

Rejected data needs an operational destination. A quarantine area should preserve enough context to explain the failure: the original record or file, ingestion time, source, rule identifier, batch or run ID, and the reason for rejection. Do not put secrets or unnecessary personal data into diagnostic fields, but keep enough evidence to investigate.

Replay is the second half of quarantine. After a source defect or transformation rule is fixed, engineers should be able to reprocess only the affected data when possible. That requires stable identifiers and idempotent logic so that replay does not duplicate records already accepted into the silver or gold layer.

The broader topic of data storage and processing matters here because storage architecture determines whether raw, rejected, validated, and curated states can coexist without confusion. Clear zones and naming make investigation faster and prevent analysts from accidentally querying quarantine data as if it were trusted.

Connect quality controls with governance and security

Quality rules often inspect sensitive fields, and quality evidence can itself contain sensitive information. Access to bronze data, rejected records, profiling outputs, and diagnostic logs should therefore be governed. The people who can view a production report may not need access to raw personal data used to validate that report.

Retention also matters. A team might need enough history to analyze quality trends or satisfy audit requirements, but retaining raw or rejected personal data indefinitely increases risk. Treat quality artifacts as part of the secure data lifecycle, with explicit classification, access, retention, and deletion rules.

In the broader Microsoft data ecosystem, Fabric can bring engineering and analytics closer together. That is useful only when trust travels with the data. Quality rules should be visible, testable, monitored, and connected to recovery—not buried in one notebook cell or remembered by one engineer.

Publish quality expectations with the data product

Downstream consumers need to know more than a table name. A mature lakehouse publishes the assumptions that make a dataset safe to use: freshness target, grain, key definitions, accepted delay, known exclusions, and the quality checks enforced before publication. That information belongs beside the data product, not only in an incident ticket after something goes wrong.

This is especially important when Fabric data feeds semantic models and analytical solutions aligned with roles such as the DP-600 Fabric Analytics Engineer path. Engineers and analysts share responsibility for trustworthy analytics, but they see different parts of the lifecycle. The data engineer should make quality state explicit before the data reaches the modeling layer.

Fabric lakehouse data quality is therefore not a cleanup stage at the end of a pipeline. It is an architectural property. Preserve raw evidence, separate trust levels, enforce domain rules, monitor trends, quarantine failures, support replay, and publish the assumptions that define trustworthy data. Those practices make the lakehouse easier to operate and far safer to use.

Data-quality ownership should follow the meaning of the rule. Platform engineers can enforce technical constraints such as schema, nullability, and freshness, but business rules often require domain ownership. A value may be syntactically valid and still be impossible for the business process it represents. Assigning rule ownership prevents a central data team from becoming the sole judge of semantics it does not control.

Quality metrics are most useful when they preserve history. A one-time dashboard that says a table is 98 percent valid does not reveal whether quality is improving, degrading, or oscillating after upstream releases. Store rule outcomes with dataset version, processing time, source, and severity so teams can correlate a new defect with the pipeline or source-system change that introduced it.

Recovery procedures should distinguish correction from concealment. Replacing invalid values with defaults can make a metric look healthier while destroying evidence of the original defect. Where possible, quarantine bad records, preserve the raw source, and repair them through a controlled path. This keeps the lakehouse trustworthy for audit and makes recurring defects easier to eliminate at their source.

Quality rules should be tested as carefully as transformation logic. Create representative records that should pass, records that should fail each important rule, and boundary cases such as empty strings, unexpected codes, duplicate business keys, or timestamps at the edge of an accepted window. The objective is not only to prove that a rule can reject bad data, but also to prove that it does not reject legitimate variation. When rules change, compare the new result with the previous rule set so teams can see how many records would move between accepted and quarantined states.

Quality service levels can make those results actionable. A critical dimension table might require zero duplicate keys, while a large event stream may tolerate a small percentage of delayed records as long as freshness remains within its agreed window. Document the threshold, severity, owner, and response for each important measure. That turns quality metrics into operational controls instead of passive dashboard numbers. Record exceptions so temporary waivers remain visible and reviewable.

Filed under AI & Data