Batch and real-time inference solve different operational problems. Batch scoring processes many records on a schedule or as a dataset operation, while real-time inference responds to individual requests through a low-latency endpoint. The model may be identical, but the surrounding requirements for freshness, throughput, features, reliability, and cost are different enough that serving mode should be an architecture decision rather than an afterthought.
The active Databricks Machine Learning Associate scope explicitly includes choosing among batch, real-time, and streaming serving approaches. The useful question is not which method is more modern. It is which method matches when predictions are needed and how quickly input state changes.
Start with the business decision window
If predictions are consumed once per day for planning, a low-latency API provides little value. Batch inference can score a full population efficiently and write results to a governed table for downstream use. By contrast, fraud checks, recommendation requests, or interactive applications may need a prediction at the moment a user acts.
Translate business timing into a service-level objective. “Real time” can mean tens of milliseconds, a few seconds, or simply fresher than the hourly batch. Precise timing requirements prevent teams from paying for online infrastructure when scheduled processing would satisfy the use case.
Design batch inference for throughput and reproducibility
Batch inference should process a defined input snapshot or interval and produce an output whose model and feature versions are known. This makes large-scale scoring reproducible and easier to reconcile. Spark and Delta tables are natural tools when predictions cover millions of records.
Partition work according to data size and downstream write patterns. The model should load efficiently across workers, and the output should include prediction time, model version, and any identifiers needed to audit the score later.
Design real-time inference for latency and concurrency
Online serving must stay responsive as concurrent request volume changes. Databricks Model Serving handles endpoint infrastructure, but teams still need to measure latency percentiles, error rates, request rates, and resource pressure. A model that performs well in a notebook may be too large or slow for an interactive service.
Optimize the full request path. Feature lookup, preprocessing, model execution, and post-processing all contribute to latency. A small model waiting on a remote database can be slower than a larger model with fast local features.
Keep feature definitions consistent
Training-serving skew occurs when a feature is computed differently during training and inference. Databricks Feature Store is designed to centralize feature definitions so models can reuse the same governed feature logic. For batch scoring, features can be joined from feature tables; online scenarios can use low-latency feature stores or request-time inputs.
The Machine Learning Professional path extends this concern into enterprise-scale MLOps, where automated feature pipelines and production serving need consistent lineage and change control.
Use streaming when predictions must follow events continuously
Streaming inference sits between periodic batch and request-response APIs. It is appropriate when events arrive continuously and should be scored as part of a data pipeline rather than through a user-facing endpoint. Examples include sensor streams, click events, or transaction feeds.
Streaming systems need state, checkpointing, backpressure handling, and late-event semantics. Do not select streaming solely because the source is continuous; if the business only consumes hourly aggregates, a micro-batch or scheduled batch may be simpler and cheaper.
Compare cost using the utilization pattern
Batch processing concentrates compute into predictable windows. Online endpoints maintain capacity to serve unpredictable requests, and low utilization can make cost per prediction high. Conversely, a huge batch can require expensive peak compute and delay decisions that would benefit from online scoring.
Measure cost per successful business outcome. Include feature retrieval, endpoint capacity, data movement, and operational support, not only the raw model execution. The least expensive model call is not necessarily the least expensive serving architecture.
Plan retries and idempotency differently
A failed batch can usually retry a partition or interval if writes are idempotent. Online requests may be retried by clients, which creates the risk of duplicate side effects if prediction is coupled to an action. Keep model inference separate from irreversible business actions or use request identifiers to deduplicate them.
For streaming, checkpoint and replay behavior must be tested explicitly. If an event is processed twice after a restart, the prediction record should not corrupt downstream state.
Monitor quality by serving mode
Batch systems can compare prediction distributions across large cohorts and reconcile outputs after each run. Online systems need sampled request/response logging, endpoint health metrics, and latency monitoring. Streaming systems add processing lag and backlog as important signals.
Quality monitoring should still use common model metrics where labels become available. The general operational discipline in incident response applies when production inference degrades: responders need ownership, evidence, severity, and a tested recovery path.
Choose the simplest mode that meets the requirement
Many systems combine modes. A daily batch may score the full customer base while a real-time endpoint recalculates risk for users who perform a new action. Shared feature definitions, model versions, and governance prevent the two paths from drifting apart.
The Databricks platform supports all three patterns, so the architectural decision should be driven by freshness, latency, throughput, feature availability, and cost. A reliable design does not maximize real-time technology; it uses real time only where the business decision actually benefits from it.
Model size can influence the serving choice. Large models may have startup, memory, or GPU requirements that make continuous online capacity expensive, while batch jobs can provision specialized compute only when scoring runs. Smaller models may be inexpensive enough to keep online continuously. Architecture should account for the operational footprint of the chosen model, not only its accuracy.
Batch inference also supports easier reconciliation. Because a run processes a defined population, teams can compare the number of inputs and outputs, verify segment coverage, and rerun a failed interval deterministically. This is valuable for regulated or financial workflows where every expected record must receive a score.
Real-time systems trade some of that simplicity for freshness. Requests may arrive with incomplete fields, malformed payloads, or identities that lack feature access. Validate inputs before model execution and make client errors distinguishable from internal serving failures so operational metrics remain meaningful.
Feature freshness should be specified per feature rather than per model. A credit score might rely on a slowly changing customer profile plus a transaction count updated every minute. Publishing every feature through the same low-latency path can increase complexity unnecessarily. Use online storage only for features whose staleness would materially affect the decision.
Cold starts and scaling events can affect tail latency. Measure first-request behavior after deployment or scale-down and decide whether the application needs warm capacity. For user-facing systems, consistent p95 or p99 latency may matter more than low average cost.
Batch jobs should record model version in every output or in immutable run metadata linked to the output. Without that connection, later users cannot determine which model produced a stored prediction after the active model has changed. The same principle applies to real-time predictions recorded for audit.
Streaming inference can combine event processing and model scoring in one path, but state growth and backpressure need attention. If input rate exceeds processing capacity, lag can rise continuously while the pipeline remains technically running. Monitor input rate, processing rate, and event-time lag as separate signals.
Serving fallbacks should be designed for the business decision. If the online endpoint is unavailable, can the application use the most recent batch score, a rules-based decision, or a “try later” response? Not every use case should fail open or fail closed in the same way.
Privacy and security differ by mode. Batch datasets can contain millions of sensitive records in one processing job, while online endpoints expose a request surface that must resist abuse. Apply least privilege, encryption, logging controls, and retention appropriate to each path.
Testing should include volume and edge conditions. Batch tests should cover the largest expected partitions and historical backfills. Online tests should cover burst traffic, invalid requests, concurrent clients, and downstream feature-store delays. Streaming tests should include restart and replay behavior.
Hybrid architectures should share prediction semantics. If the same customer can be scored in batch and online, both modes should use compatible feature definitions, model version rules, and output schema. Otherwise teams may see two different “current” scores without understanding why.
The final choice is therefore about decision timing and operational guarantees. Batch maximizes efficient throughput and reproducibility, online serving minimizes response delay, and streaming scores events continuously. Using the simplest mode that satisfies the decision window usually produces the most reliable system.
Data arrival patterns also influence the choice. If source features are updated once overnight, a millisecond endpoint cannot make the underlying information fresher. Online serving only adds value when either request-time context or frequently updated features make the prediction materially different from the most recent batch score.
Batch output tables should have lifecycle rules. Decide how long historical predictions are kept, whether reruns overwrite or version prior results, and how downstream users distinguish corrected scores. Prediction history can become important evidence during audits or model investigations.
Online endpoints should expose enough metadata for callers to troubleshoot safely, such as request identifier and model version, without leaking implementation details or sensitive data. Those identifiers let support teams trace one user-visible decision through serving logs and feature lookups.
Streaming systems need explicit watermark and late-data policy when event time affects features. A prediction generated before a late event arrives may differ from one generated after replay. Decide whether historical predictions are corrected, annotated, or left unchanged according to the business contract.
Capacity planning should account for peak rather than average demand. Interactive traffic often arrives in bursts tied to business hours or events. Load testing should reproduce concurrency and request size so autoscaling behavior is understood before a peak period becomes the first real test.
Operational ownership differs by mode as well. Batch scoring is often owned with data pipelines, while online inference behaves more like an application service. Define who responds to failed batch runs, endpoint latency, feature-store outages, and bad predictions so incidents do not bounce between teams.
Deployment testing should verify that both modes produce compatible outputs for the same representative records when they are intended to use the same model. Small numerical differences can be acceptable, but large discrepancies often reveal preprocessing or feature mismatch that should be fixed before release.
Finally, document why the selected serving mode exists. If business timing changes, future teams should know whether an online endpoint was required for real decision latency or simply chosen because it was convenient during development. Clear rationale makes it easier to simplify or scale the architecture later.
Serving-mode decisions should be revisited as traffic and business timing evolve. A model that begins as nightly batch may later need real-time scoring for one narrow workflow, while other consumers can stay on the cheaper batch path.