INSIGHTS
AI & Data

Databricks Data Engineer Professional: Lakehouse Change Data Capture

In this article
  1. Choose CDC because the source changes, not because it sounds modern
  2. Define business keys and sequence information carefully
  3. Use declarative AUTO CDC where it fits
  4. Decide whether consumers need current state or history
  5. Make deletes explicit and test them
  6. Control schema evolution instead of accepting every drift
  7. Design replay and recovery before the first incident
  8. Observe lag, throughput, quality, and change anomalies
  9. Keep CDC models understandable to downstream users

Change data capture is the point where a lakehouse stops behaving like a sequence of full reloads and starts behaving like a continuously maintained representation of source state. The difficult parts are not reading a change feed; they are ordering, duplicate handling, deletes, late events, schema evolution, replay, and preserving the history that downstream consumers actually need.

The current Databricks Certified Data Engineer Professional scope emphasizes production-grade pipelines, Delta Lake, Auto Loader, Lakeflow, streaming, reliability, observability, governance, and optimization. Databricks now recommends the Lakeflow AUTO CDC APIs for declarative CDC and notes that they replace the older APPLY CHANGES naming while preserving the same core purpose.

Choose CDC because the source changes, not because it sounds modern

CDC is appropriate when source systems insert, update, and delete records and downstream systems need those changes without repeatedly scanning and replacing an entire dataset. Transaction databases, SaaS platforms, and operational systems are common sources, but the design should begin with change semantics rather than connector selection.

A source that is append-only may be better served by incremental file or event ingestion. A source that produces ordered row changes can feed a CDC flow directly. A source that only exposes periodic snapshots may require snapshot comparison instead. The ingestion method should match what the source can state reliably about change.

The broader foundations in data engineering still apply: establish source ownership, keys, latency expectations, and error handling before choosing implementation details.

Define business keys and sequence information carefully

A CDC system needs to know which logical entity a change belongs to and how competing changes should be ordered. Business keys identify the entity; sequence columns determine which event wins when records arrive out of order.

Do not assume ingestion time is the same as business event time. Network delays, connector retries, batch exports, and replay can all cause an older business change to arrive after a newer one. When the source supplies a transaction sequence, log position, change version, or trustworthy timestamp, preserve it through ingestion.

If ordering is ambiguous, document that limitation explicitly. A pipeline cannot manufacture source truth that the upstream system never emitted.

Use declarative AUTO CDC where it fits

Lakeflow Spark Declarative Pipelines provides AUTO CDC and AUTO CDC FROM SNAPSHOT so teams can express keys, sequencing, deletes, and slowly changing dimension behavior without hand-writing every merge edge case. This reduces the amount of custom state-management code in the pipeline.

For a change feed, AUTO CDC processes the incoming events into a streaming target. For ordered snapshots, AUTO CDC FROM SNAPSHOT can derive changes by comparing snapshots. The current APIs support SCD Type 1 and Type 2 patterns, and newer capabilities also support bitemporal tracking for workloads that need both valid-time and system-time history.

Declarative automation does not eliminate design responsibility. Teams still need correct keys, sequence semantics, schema expectations, data-quality rules, and recovery procedures.

Decide whether consumers need current state or history

SCD Type 1 overwrites prior values and is appropriate when consumers only need the latest state. SCD Type 2 preserves historical versions so analysts can answer questions such as what a customer segment, product classification, or account status was at a particular time.

History adds storage and query complexity, so it should reflect a real requirement rather than a default preference. If only a handful of attributes need historical tracking, consider whether the model can separate slowly changing dimensions from high-volume facts.

For audit-sensitive domains, bitemporal designs may be required to distinguish when a fact was true in the business from when the platform learned or stored it. That distinction is valuable when corrections arrive late.

Make deletes explicit and test them

Deletes are frequently the weakest part of a CDC design. Some sources emit delete tombstones, some set a status flag, and some omit deleted rows entirely from snapshots. The pipeline must translate the source convention into a deliberate downstream rule.

Physical deletion can simplify current-state tables but may conflict with audit and replay needs. Soft deletion preserves evidence but requires consumers to filter consistently. In regulated or analytical environments, retaining the change history while presenting a current-state view can satisfy both needs.

Test deletion paths separately from inserts and updates. A pipeline that handles 99 percent of changes but silently misses deletes will accumulate increasingly misleading data.

Control schema evolution instead of accepting every drift

CDC streams often run for years, which means source schemas will change. New nullable columns may be harmless, while type changes, key changes, and semantic redefinitions can invalidate downstream assumptions.

Use explicit schema controls, rescued-data patterns, or staged evolution so unexpected fields do not quietly corrupt trusted tables. The pipeline should record what changed, when it changed, and whether downstream models were validated against the new structure.

A schema change is not merely a parser event. It can affect deduplication keys, generated columns, data-quality expectations, and historical compatibility.

Design replay and recovery before the first incident

Checkpointing lets streaming processing resume efficiently, but a checkpoint is not a complete recovery strategy. Teams need to know whether they can replay source changes, rebuild targets from bronze data, or rehydrate from a snapshot when state is lost.

Preserve a durable raw layer when the cost and compliance model allow it. A raw history gives engineers the ability to correct transformation logic without asking the source system to regenerate old changes it no longer retains.

The professional discipline described in advanced data engineering is visible here: recovery is a design feature, not a ticket opened after corruption occurs.

Observe lag, throughput, quality, and change anomalies

A healthy CDC pipeline is not defined only by whether the job is green. Monitor source lag, backlog growth, records processed, duplicate rates, late-event frequency, delete volume, expectation failures, and target freshness.

Unexpected change patterns can reveal upstream incidents. A sudden collapse in update volume may indicate a stalled connector; a spike in deletes may represent a source bug; a widening lag may signal insufficient compute or a downstream sink bottleneck.

Operational metrics should be tied to service expectations. If a business dashboard requires data within fifteen minutes, alert on breach of that freshness objective rather than on arbitrary CPU thresholds.

Keep CDC models understandable to downstream users

CDC plumbing should not leak unnecessary complexity into every consumer query. Curated silver and gold models should present clear current-state and historical semantics, including effective dates, delete status, and lineage.

Document whether timestamps are source event time, ingestion time, or processing time. Clarify which column represents the business key and how rekeys or merges are handled. These details are part of the data contract.

Across Databricks certifications, CDC is a useful synthesis topic because it combines ingestion, Delta Lake, streaming, modeling, data quality, and operations. The implementation succeeds when downstream users can trust both the values and their history.

Before implementing CDC, inventory the source system’s retention guarantees. Some transaction logs keep only a limited history, and some managed connectors can resume only while that history remains available. Recovery objectives must fit within that window or the design needs a fallback full snapshot. A pipeline that promises seven-day recovery while the source keeps changes for twelve hours has an architectural contradiction.

Initial loading deserves a separate plan from steady-state capture. The first snapshot may contain billions of rows while ongoing CDC carries only a small fraction of that volume. Keep the transition between initial hydration and change processing deterministic so records are not missed or duplicated at the boundary. Newer Lakeflow CDC patterns can combine one-time flows with continued changes, but the sequence still needs validation.

Deduplication should be based on source event identity or sequencing where possible, not on every column of a record. Full-row deduplication becomes expensive and may incorrectly collapse legitimate repeated business events. Preserve source transaction metadata so engineers can distinguish a retried delivery from two identical changes that truly occurred at different times.

Late corrections are common in business systems. A customer record may be backdated, a financial transaction may be corrected after close, or a source may replay an older update after reconnection. The target model must define whether sequence order, business effective time, or source commit time wins. That decision belongs in the data contract because different consumers may care about different time semantics.

SCD Type 2 tables need consistent effective and end timestamps, a clear current-row indicator where used, and rules for overlapping versions. Test boundary conditions such as two changes with the same sequence value, deletes after updates, and corrections to previously closed history. Historical correctness is difficult to repair after consumers build reports on inconsistent intervals.

When a CDC feed crosses regions or clouds, network interruption and connector restarts become routine rather than exceptional. Monitor the last acknowledged source position and the oldest unprocessed change. This makes it possible to distinguish a quiet source from a disconnected pipeline and estimate whether retention risk is increasing.

Governance should follow source sensitivity through the change stream. Raw CDC payloads can contain fields that are removed from curated tables but remain present in bronze history. Apply catalog permissions, retention, masking, and deletion obligations to the raw layer as carefully as to the final model.

Backfills should be observable as separate operational events. A historical replay can create a legitimate spike in throughput, writes, and cost that resembles an incident if monitoring has no awareness of maintenance activity. Tag or annotate backfill runs, bound the date range, and verify downstream expectations before the replay starts.

Performance tuning should begin with the highest-volume state transition. Merge-heavy CDC can become expensive when target files are poorly organized or keys provide weak data skipping. Use current Delta optimization and clustering capabilities before inventing complex manual partition schemes that will be difficult to maintain.

Finally, test the pipeline against synthetic adversarial sequences: duplicate events, out-of-order changes, a delete followed by a late update, schema additions, source pauses, and replay after checkpoint loss. Happy-path tests prove syntax; adversarial tests prove the semantics that make CDC trustworthy.

CDC also changes how privacy deletion is implemented. If a regulated request requires erasing personal data, historical bronze records, SCD2 versions, checkpoints, and derived tables may all need consideration. Retention design should reconcile replay capability with legal deletion obligations rather than assuming raw history can be kept forever.

For multi-table transactions, row-level CDC can expose intermediate states that never existed from the application’s business perspective. If consistency across entities matters, preserve transaction identifiers or design downstream logic that waits for a complete transactional boundary where the source provides one.

Schema registries or contract repositories can help teams coordinate source evolution before it reaches CDC. Producers should communicate breaking changes, while consumers validate them in lower environments using representative change sequences. A CDC pipeline should not become the first place anyone discovers that a primary key changed.

When CDC feeds multiple downstream targets, centralize normalization and ordering once rather than letting every consumer independently interpret the raw change stream. A curated silver change model reduces duplicate logic and creates one place to fix source-specific quirks.

Metrics should include the distance between source sequence and target sequence, not only wall-clock lag. A busy source can process thousands of changes per second, so being “two minutes behind” may represent millions of unprocessed records and a significant recovery risk.

Use sampling and reconciliation jobs to compare current source state with target state periodically. Even a well-designed CDC system can suffer silent connector bugs, missed partitions, or incorrect delete interpretation. Reconciliation provides an independent control that validates the accumulated result.

Teams should document when a consumer may query historical versions directly and when it should use a curated current-state view. Exposing raw SCD2 structures without guidance encourages analysts to double-count entities or join against overlapping effective periods.

Source-system maintenance can create discontinuities that look like real business change. Bulk updates, backfills, deduplication projects, and database migrations may emit enormous change volumes or reset sequence behavior. Coordinate these events with downstream teams so capacity, retention, and reconciliation can be adjusted deliberately.

The Data Engineer Associate path is also relevant because foundational ingestion, data loading, monitoring, and troubleshooting skills form the base on which production CDC is built. Professional-level design adds the harder concerns of scale, history, deployment, and recovery.

For high-value tables, store reconciliation results as a governed dataset. Historical mismatch rates, missing-key counts, and correction runs become evidence that the CDC system is preserving source state over time rather than merely appearing healthy in the latest dashboard.

When the source emits multiple event types, keep transformation from transport parsing separate. Normalize inserts, updates, deletes, and metadata into a clear internal change model first, then apply business rules. This makes connector replacements easier because downstream logic depends on one stable representation instead of every source-specific payload shape.

CDC success should be measured against source reconciliation, not only pipeline uptime. Schedule periodic checks that compare sampled keys, counts, or business aggregates between source and target. A healthy job that has been applying the wrong ordering rule for a month is still a data incident.

Use CDC to reduce unnecessary work, but do not trade batch simplicity for opaque streaming complexity. The right design makes source change semantics explicit and keeps recovery possible.

When keys, ordering, deletes, schema evolution, and observability are treated as first-class requirements, a CDC pipeline becomes a dependable data product rather than a fragile sequence of merges.

Filed under AI & Data