Model selection in Microsoft Foundry is not a popularity contest between model names. The best deployment is one that performs the application’s task within its permitted latency, region, cost, safety and data-handling boundaries. A classification workflow with bounded categories may be well served by a smaller model, while a multimodal troubleshooting assistant needs to reason over images and prose. A model that produces polished answers but consistently misses required identifiers is not a good choice for an extraction job, regardless of its leaderboard position.
For the planning portion of AI-103, an engineer should be able to justify which model and deployment mode to use, how to estimate capacity, and how to react when real traffic differs from a small test. Start from a workload contract: input size, output structure, throughput, maximum acceptable delay, allowed regions, cost ceiling, and consequences of mistakes.
Translate product requirements into model capabilities
Write down what counts as a correct result. A customer-service application may need a clear answer with citations, whereas an automated routing service may need only one of six category labels and a reason code. A content-moderation path may prioritize false-negative control over a natural conversational style. These differences affect whether to choose a general-purpose LLM, smaller language model, multimodal model or a specialized Foundry Tool.
Model capability is also constrained by what the target service actually exposes. Not every catalog model supports every modality, structured-output mode or tool calling method. Confirm the deployment’s supported features in the region and API version being used. The AI-103 skills guide explicitly includes model selection across large, small and multimodal models; it does not promise a single preferred model for all tasks.
Create an evaluation set before comparing deployments
Choose representative inputs before testing a candidate model. For an insurance-document assistant, include clean forms, blurry scans, partially missing records, conflicting amounts, adversarial instructions embedded in attachments, and queries where no approved source exists. Record expected behavior for each. Without negative and edge cases, the model that provides the most fluent answer may appear strongest because the evaluation never checks whether it should have refused to guess.
Evaluate each candidate with the same task definition and with conditions that resemble production. Compare exact-field accuracy, appropriate refusal, citation correctness, safety violations, median and tail latency, and cost per completed task. Rerun the set after prompt and retrieval changes; a larger context window or a different tokenizer can alter performance. Avoid comparing two models on differently sized or differently ranked retrieval snippets and calling the result a model-only improvement.
Distinguish context window from useful evidence
A large context window makes it possible to include more information, but it does not guarantee that the model will identify the right fact. A full policy library inserted into one request can contain conflicting versions, irrelevant data and text that attempts to instruct the model. It also increases input tokens and can make the answer slower and more expensive.
Select evidence before generation. Use document ownership, revision timestamps and access-control filters to identify what may be retrieved, then rank only the remaining relevant passages. For long analytical tasks, consider a staged workflow that extracts facts and compares them against original sources before writing a summary. More context is beneficial only when it improves supported decisions, not when it replaces retrieval quality with raw volume.
Interpret rate limits as production capacity constraints
Model deployment capacity is subject to rate and usage limits. Requests per minute and token throughput can affect a workload differently: many short classification calls may hit request pressure, while fewer long document analyses may exhaust tokens. Regional quota, deployment model and subscription capacity can also matter. A load test using only brief prompts hides the second failure mode.
Design the client to distinguish a rate-limit response from a transient service error or a permanent authorization failure. Back off with jitter for retryable errors, cap retry attempts, and respect request deadlines. For a user-facing chat, place a bound on total wait time and offer a meaningful failure message. For a batch job, use a queue with a controlled concurrency limit so bursts do not amplify the problem. Microsoft’s Foundry quota documentation should be checked against the actual model family and deployment, because capacity rules evolve.
Estimate costs by completed workload, not only tokens
An application that fails at the final extraction step may consume several model calls and retrieve the same evidence again on retry. Cost estimation should include failed requests, embedding creation, retrieval and reranking, evaluation calls, tracing retention, tool usage and any agent runtime charges, not simply tokens in a successful answer. A nominally cheaper model that triggers twice as many repairs may have a worse cost per valid task.
Make cost visible at the workflow level. Log request category, model deployment, approximate token usage, retrieval operations, retry count, outcome and trace ID. Use aggregated reports rather than publishing entire prompts into accounting telemetry. Set budgets and alert thresholds appropriate to the environment; a prototype allowed to cost twenty dollars per week is different from a customer-facing application with a contractual load forecast.
Choose deployment isolation and routing deliberately
Separate nonproduction test deployments from the production service that users depend on. A developer trying a new model version should not silently change the behavior of an unrelated application’s endpoint. Keep deployment names and model versions in configuration, and expose the selected version in troubleshooting metadata so operators can correlate a behavior change with a rollout.
Consider failure handling before designing an automatic fallback. Routing a request from one model to another may reduce outages, but the alternate model might have different tool-calling support, safety rules, region restrictions or output consistency. Test compatibility before enabling fallback. When the alternate cannot meet the task contract, returning a clear unavailable message is safer than delivering a plausible but incorrectly structured action.
Test tail latency and interactive behavior
Median response time rarely captures the cases users complain about. A retrieval-heavy agent can respond quickly most of the time yet become very slow when search, a tool call and one model retry line up. Set separate time budgets for retrieval, generation and each tool. Instrument these components so an operator can distinguish model slowness from index degradation or an external API timeout.
Streaming may improve perceived responsiveness, but it does not solve latency for workflows that must validate a complete JSON object before acting. Do not stream unvalidated transaction instructions directly into a payment API. Interactive output and consequential actions have different correctness requirements. If the response must pass schema validation or human approval, design that boundary independently from how the user sees partial text.
Make quota incidents diagnosable
When requests begin failing during a campaign launch, inspect which dimension changed: requests, token volume, model output size, retrieval fan-out or agent tool retries. Compare deployment rate-limit responses with overall service telemetry. If traffic is within configured budgets but a subset fails, check whether those users send much larger attachments or whether one workflow enters a repeated tool loop. Blindly requesting more quota will not fix an infinite retry or oversized document ingestion.
Practice a simulation with three traffic patterns: high request count of small tasks, low count of large tasks, and mixed user requests with a slow downstream tool. Measure queue depth, accepted throughput, tail latency and user-visible failure. Write the tradeoff down: which calls are delayed, which requests are shed, and which operations must not be retried without an idempotency key. A production model selection decision includes this operating behavior, not merely an accuracy score.
Consider a customer-support application with two deployments: a conversational model for drafting answers and a compact model for ticket classification. Under a burst of customer traffic, the conversational endpoint begins returning 429 responses while the classification endpoint remains healthy. The wrong reaction is to retry every request immediately or move all work to a second deployment without checking whether the deployments share the same quota pool. Foundry quota accounting depends on the model family, deployment type and subscription-level scope; Microsoft has been transitioning some model types to shared subscription pools since 2026. The current quota page and deployment settings, not an old per-region assumption, must determine the mitigation.
A useful operational response starts by measuring input and output tokens, request concurrency, queued work and the proportion of requests that genuinely need the expensive model. Use a bounded exponential backoff honoring any Retry-After signal. An interactive user can receive a retryable, truthful service-unavailable response when its latency budget expires; a background batch task can wait longer if the business workflow permits. Distinguish a retry of an inference request from a retry of a completed downstream action. Only the former can be repeated without risking a duplicate ticket or payment.
If a lower-cost model passes the appropriate evaluation suite, route uncomplicated classifications there and reserve higher-capacity inference for ambiguous questions. That is capacity engineering based on task complexity, not a blind instruction to use the cheapest model. Compare response quality under a production-shaped test set including long prompts, non-English input and missing evidence. The fallback path must not silently reduce authorization filtering or bypass content safety because it serves fewer tokens.
Finally, record capacity as an application requirement rather than an administrator’s tribal knowledge: peak requests per minute, typical prompt size, accepted p95 latency, retry budget, per-region support and the procedure for requesting more quota. A model can be excellent in an isolated benchmark but unsuitable when its throughput cannot satisfy that service’s traffic pattern.
Use evidence to decide when to change models
A model change is justified when measurements show that a different deployment solves a real limitation under the required constraints. If retrieval returns irrelevant documents, changing generation models may improve style while leaving the source error intact. If output schema failures are common, tighten the schema and validation path before buying more expensive inference. If hallucinations persist despite strong evidence, evaluate prompt design and grounding behavior explicitly.
Maintain a release report containing the evaluated task set, model and prompt versions, latency distribution, safety checks, estimated cost and any known unsupported cases. Then canary a limited share of traffic and compare observed outcomes before a wider rollout. This approach turns model choice from an aesthetic preference into a repeatable engineering decision.