INSIGHTS
AI & Data

Microsoft DP-700: Spark Notebooks in Microsoft Fabric

In this article
  1. Use notebooks for transformations that benefit from Spark
  2. Separate exploratory cells from production logic
  3. Make inputs explicit with parameters and configuration
  4. Use Fabric environments to control runtime and dependencies
  5. Design reads and writes around lakehouse reliability
  6. Control performance instead of relying on Spark defaults
  7. Turn notebook output into observable operations
  8. Orchestrate notebooks as reusable production steps

Notebooks are one of the most flexible ways to work with Apache Spark in Microsoft Fabric. They let a data engineer explore a dataset, develop PySpark or Spark SQL transformations, inspect intermediate results, document assumptions, and then turn the same logic into a repeatable production step. That combination makes notebooks useful, but it also creates a trap: code that is convenient during exploration can become difficult to operate if it is promoted to production without stronger structure.

For engineers working toward DP-700, notebooks sit inside a broader Fabric data-engineering workflow. The current exam scope expects candidates to understand how data is ingested and transformed, how an analytics solution is secured and managed, and how it is monitored and optimized. A notebook is therefore not an isolated coding exercise. It is one execution surface inside a lakehouse, pipeline, environment, workspace, and operational model.

Readers who are newer to Spark may also benefit from grounding the workflow in data engineering fundamentals. The key production question is not whether a notebook can produce the right output once. It is whether the notebook can produce the same trustworthy result when it is parameterized, scheduled, retried, promoted, monitored, and maintained by somebody other than its original author.

Use notebooks for transformations that benefit from Spark

Fabric supports several data-processing experiences, so a notebook should be chosen because Spark is appropriate for the workload, not simply because notebooks are familiar. Spark is well suited to distributed transformations over large files and Delta tables, joins across substantial datasets, feature engineering, semi-structured data processing, and workloads that need the broader Spark ecosystem. For smaller relational transformations, a warehouse or SQL endpoint may express the logic more directly. For visual low-code transformation, Dataflow Gen2 can be more maintainable for some teams.

A useful decision test is to ask what the transformation needs from its execution engine. If the workload benefits from distributed processing, programmatic control, reusable functions, custom libraries, or Spark-native APIs, a notebook is a strong candidate. If the work is primarily a handful of SQL statements over curated relational tables, moving it into Spark may add operational complexity without adding value.

That choice matters because every execution surface creates its own support burden. A mature Fabric design minimizes unnecessary engine switching. The existing comparison of Microsoft Fabric and Power BI provides broader platform context, but the notebook decision should remain workload-driven: choose the tool that makes the data contract easiest to understand, test, and operate.

Separate exploratory cells from production logic

Interactive exploration is a notebook strength. An engineer can sample records, display distributions, test a parsing expression, compare joins, and inspect a schema in minutes. Those exploratory cells are valuable during development, but they should not automatically become part of the production path. Debug displays, temporary filters, one-off counts, hard-coded sample paths, and experimental variables can make an operational notebook unpredictable.

Before a notebook is scheduled, reduce it to a deliberate execution sequence. Group configuration near the beginning, define reusable functions clearly, isolate transformations into understandable stages, and make the final write behavior explicit. Cells should be able to run in the intended order without depending on an undocumented manual step performed earlier in the authoring session.

Statefulness deserves special attention. Notebook sessions can preserve variables and temporary objects during interactive work. A production run, however, should assume a clean start unless the design intentionally uses persistent storage. If a notebook succeeds only after a developer has executed cells in a particular informal order, it is not ready for orchestration.

Make inputs explicit with parameters and configuration

Hard-coded dates, paths, table names, and workspace-specific values are among the fastest ways to turn a useful notebook into a fragile one. Production notebooks should receive the values that legitimately vary between runs. Common examples include a processing date, source table, destination table, watermark boundary, batch identifier, environment name, or business-unit code.

Parameters make the execution contract visible to an orchestration layer. A Fabric pipeline can call a notebook with a defined processing window instead of editing code for each run. The notebook can validate those values at startup and fail with a clear message when a required parameter is missing or malformed. This makes reruns and historical backfills much safer than relying on a developer to modify literals inside cells.

Configuration that is stable within an environment but different across environments should also live outside core transformation logic. Connection details, storage locations, feature flags, and operational thresholds may come from managed configuration rather than source code. Secrets should never be embedded directly in notebook cells. The broader discipline of Azure data storage and processing is easier to maintain when compute logic is separated from the environment-specific details that tell it where to read and write.

Use Fabric environments to control runtime and dependencies

A notebook can be correct and still fail when its runtime changes. Python packages, Spark settings, library versions, and other dependencies need to be treated as part of the workload. Fabric environments provide a reusable way to define runtime and library configuration so multiple notebooks can execute against a more consistent dependency set.

This is especially important when notebooks move from development to scheduled production. Ad hoc package installation inside a notebook may be convenient for experimentation, but it can make execution slower and less reproducible. A controlled environment reduces the chance that two notebooks nominally doing the same type of work are actually running against different dependency versions.

Dependency management should also be conservative. Every extra library increases the surface area for compatibility issues, security review, and future upgrades. Prefer platform-provided capabilities and well-justified dependencies. When a custom package is required, record why it is needed and test upgrades deliberately rather than allowing the production environment to drift.

Design reads and writes around lakehouse reliability

Many Fabric notebooks read from or write to OneLake-backed lakehouse tables. The notebook should make those data boundaries clear. Read only the columns and partitions needed for the transformation where practical, apply filters early when they reduce work safely, and avoid repeatedly scanning a large source simply because separate cells were written at different times.

Writes require an even stronger contract. Decide whether a run appends new records, overwrites a partition, merges changed records, or rebuilds a complete table. The same processing window should not produce duplicate business rows if the notebook is retried. Stable keys, deterministic partitioning, merge conditions, and batch metadata can make execution idempotent.

Delta tables are valuable because their transaction model provides a stronger consistency boundary than unmanaged file replacement. Even so, a transaction cannot decide the business semantics for you. If a late record belongs to yesterday’s partition, or a source system changes an existing customer record, the notebook still needs a defined rule for how that change is incorporated.

Control performance instead of relying on Spark defaults

Spark can process large workloads efficiently, but poorly shaped transformations can waste substantial capacity. Start by examining data volume and partitioning rather than immediately increasing compute. Excessive small files, unnecessary shuffles, wide joins without selective filtering, repeated actions on the same uncached intermediate data, and skewed keys can all dominate runtime.

Push filters and column selection as close to the read as possible when semantics permit. Be cautious with operations that move large amounts of data between executors. Repartitioning can help when the current layout is inappropriate for a downstream operation, but repartitioning itself is work, so it should solve a measurable problem rather than become a default step in every notebook.

Caching also needs discipline. Persisting a dataset can accelerate repeated reuse in one execution, but caching everything consumes memory and can reduce performance. Use it where a costly intermediate result is reused enough to justify the storage cost, then release it when it is no longer needed. Performance decisions should be driven by observed execution behavior, not generic Spark folklore.

Fabric also supports high-concurrency notebook scenarios that can reduce startup overhead for compatible workloads by sharing Spark resources. That capability is useful for throughput, but it does not remove the need to understand isolation, resource competition, and the actual concurrency pattern. Consolidating sessions is beneficial only when the workload remains stable under shared capacity.

Turn notebook output into observable operations

A scheduled notebook needs operational signals beyond a final success or failure flag. Log a run identifier, input window, key record counts, rejected-row counts, destination version or partition, and major stage timings. These signals help an operator distinguish between a compute failure and a successful run that produced suspicious data.

Do not use logging as an excuse to dump whole records or secrets. Operational evidence should reveal enough to diagnose the run without exposing sensitive values. This is where notebook engineering intersects with the secure data lifecycle: logs, temporary files, checkpoints, and rejected records are data assets too, and their access and retention need intentional controls.

Monitoring should also include duration and resource behavior. If the same notebook gradually takes longer for a similar input size, that trend can reveal file fragmentation, data skew, a dependency change, or a capacity bottleneck before the job breaches a service-level expectation.

Testing becomes easier when the notebook’s transformation logic can run against a small controlled dataset. Keep pure transformation functions separate from I/O where practical, and create representative cases for nulls, duplicates, malformed values, late records, and boundary conditions. A notebook that can be validated only against a full production-scale dataset will be slower to change and harder to troubleshoot.

Source control should preserve both code history and review. Notebook formats can produce noisy diffs, so teams should agree on a workflow that makes meaningful changes visible and avoids committing large outputs or sensitive results. The aim is the same as with any production code: another engineer should be able to review what changed before the scheduled workload begins using it.

Orchestrate notebooks as reusable production steps

Fabric pipelines can invoke notebooks as activities, which makes notebooks useful building blocks in broader workflows. The pipeline should own dependencies, scheduling, retry policy, and branching. The notebook should own its transformation contract. That separation makes it possible to rerun one failed stage without embedding the entire orchestration graph in notebook code.

Return status information in a form the orchestrator can interpret. A notebook that detects a data-quality failure should not quietly print a warning and exit as though everything succeeded. Conversely, a recoverable condition should not always crash the whole workflow if the agreed contract is to quarantine a bad subset and continue. The failure behavior belongs to the design, not to ad hoc exception handling.

For engineers in the Microsoft ecosystem, the operational lesson is that a notebook is most valuable when it becomes predictable. Parameters define its inputs, environments control dependencies, Delta-based writes support consistent storage, monitoring reveals both execution and data state, and pipelines place the notebook inside a repeatable workflow. That is what turns an interactive Spark workspace into a maintainable component of a Fabric data platform.

Filed under AI & Data