INSIGHTS
Cloud Computing

Microsoft AZ-305: Highly Available Azure Applications

In this article
  1. Define availability in business terms
  2. Use availability zones to survive datacenter-scale faults
  3. Remove single points of failure from every tier
  4. Health probes must represent application usefulness
  5. Scale design should include failure headroom
  6. Multiregion design addresses regional outages
  7. Design applications to tolerate transient faults
  8. Deployments must preserve the redundancy you designed
  9. Prove recovery through controlled failure testing

High availability in Azure is not a single feature that can be enabled after an application is deployed. It is an end-to-end property created by redundant components, failure-aware traffic routing, resilient state, health monitoring, and recovery procedures. The architecture topics measured by AZ-305 require architects to decide which faults the workload must survive and then choose Azure services and deployment patterns that match those requirements.

Microsoft’s current reliability guidance recommends multiple availability zones for production workloads where the region and service support them, and it suggests both multi-zone and multiregion design for mission-critical workloads. That does not mean every application should run everywhere. Reliability has cost and complexity. The design should buy enough redundancy to satisfy the service objective without creating an operational model the team cannot maintain.

Define availability in business terms

Begin with the user-visible service. Which functions must remain available, how much downtime is acceptable, and what degraded behavior can the business tolerate? A single uptime percentage is often too vague. Login, checkout, reporting, and administrative operations can have different criticality even inside the same application.

Translate those requirements into service level objectives, recovery time objectives, and recovery point objectives. Then identify the dependencies that can prevent the application from meeting them. Compute may be redundant while a single database, DNS dependency, secret store, or hybrid network link remains a point of failure.

The value of availability targets is not the number of nines itself. The target forces architecture and business teams to agree on how much resilience is worth funding and operating.

Error budgets can help connect reliability to engineering decisions. If the service is operating well inside its allowed downtime, teams can take more deployment risk or focus on features. If it is consuming the error budget quickly, reliability work should take priority. This creates a practical feedback loop between the availability target and day-to-day delivery.

Use availability zones to survive datacenter-scale faults

Availability zones are physically separate groups of datacenters within an Azure region with independent power, cooling, and networking. Zone-redundant or multi-zonal deployment reduces the chance that a single datacenter-scale event takes down the workload. For many production applications, this is the first major resilience layer.

Different Azure services use zones differently. Some are zone-redundant services where Azure distributes replicas across zones. Others are zonal resources that the customer places into specific zones. For virtual machines, that can mean distributing instances across zones and placing them behind a load balancer. For managed data services, the service may provide a zone-redundant configuration option.

The concepts behind Azure regions and availability zones matter because a zone is not a second region. Zone resilience protects against one class of failure while keeping the workload inside the same regional boundary.

Zone selection should be tested for service support. A region may provide availability zones while a specific SKU or service feature does not. Architects should verify zone behavior for every critical dependency rather than assuming the regional capability applies uniformly. Where a service is not zone-resilient, compensate through another design or explicitly accept the risk.

Remove single points of failure from every tier

Redundant front ends are useful only if the application’s state and dependencies are equally resilient. Review the architecture tier by tier: ingress, web or API compute, messaging, databases, caches, storage, secrets, DNS, outbound connectivity, monitoring, and identity integration. For each component, ask what happens if an instance, zone, or service dependency becomes unavailable.

State is usually the hardest part. Stateless application instances can be replaced quickly, while transactional data must remain consistent. Use managed database replication and service-native high availability where possible. If sessions are stored locally on one web server, load balancing more instances will not make the application truly resilient.

External dependencies need the same scrutiny. An application can be fully redundant in Azure and still fail because it depends on one on-premises API reachable through one device.

Shared dependencies deserve special scrutiny because they can defeat redundancy across multiple applications at once. A central DNS server, identity bridge, firewall, build system, or secret store may be more critical than any single application component. Platform services should have reliability objectives proportional to the number of workloads that depend on them.

Health probes must represent application usefulness

Load balancers and global routing services need health signals to decide where traffic should go. A probe that checks only whether a TCP port opens can report a component as healthy even when the application cannot serve a valid transaction. Health endpoints should reflect the minimum dependencies needed for useful operation.

Design separate liveness and readiness concepts where appropriate. A process can be alive but not ready to receive new requests. During startup, deployment, or dependency failure, removing an instance from traffic before it begins returning user errors can preserve service quality.

Cloud load balancing distributes traffic, but it does not create healthy capacity by itself. The application must expose a trustworthy signal, and the platform must have enough surviving instances to absorb demand when unhealthy ones are removed.

Health endpoints must also avoid making the problem worse. If every probe runs an expensive database query every few seconds from every instance, the health system can add load during an incident. Prefer lightweight checks that prove essential dependencies without creating a new performance bottleneck.

Scale design should include failure headroom

Capacity plans often assume all instances are available. High-availability planning assumes some capacity is missing. If a three-zone service loses one zone, the remaining zones must handle the traffic. That might require overprovisioning, autoscaling, or a deliberately higher minimum instance count.

Autoscaling should be fast enough to respond without creating instability. If the workload suddenly moves from three zones to two, the surviving resources may need to scale before they saturate. At the same time, scale limits should prevent a runaway signal from creating uncontrolled cost or overwhelming downstream services.

Test the capacity model under failure. Normal load tests do not prove availability if they are run with every dependency healthy and every instance online.

Capacity headroom should be explicit in autoscale thresholds. Scaling only after average CPU reaches a very high level can leave too little time to add instances when one zone fails. Use application latency, queue depth, request rate, or other leading indicators where they better represent demand. Failure scenarios should be part of load testing.

Multiregion design addresses regional outages

Availability zones do not protect against a full-region outage. Workloads with sufficiently strict continuity requirements need a second region. That can be active-active, where both regions serve traffic, or active-passive, where a standby region takes over after failure. Each model changes cost, data replication, traffic routing, and operational complexity.

Global routing can use Azure Front Door or Traffic Manager depending on the protocol and architecture. Regional ingress might use Application Gateway or Load Balancer. Data needs a cross-region strategy that matches the RPO. DNS, certificates, secrets, private networking, and hybrid connectivity must also exist or be reproducible in the secondary location.

The choice of secondary region should be requirement-driven. Paired regions can be useful, but newer Azure regions may not be paired and Azure supports resilient multiregion designs without a pair.

Regional data architecture deserves a separate decision. Some services provide built-in geo-replication; others require application-managed replication or restore. The design should state which region is authoritative, how writes are prevented or reconciled during failover, and how the application returns to the preferred region after recovery.

Design applications to tolerate transient faults

Cloud applications experience short-lived network delays, throttling, instance restarts, and service transitions even when there is no major outage. Resilient applications treat transient failures as expected. Use bounded retries with backoff for operations that are safe to retry, timeouts that fail before user requests hang indefinitely, and circuit breakers where repeated failures would otherwise overload a dependency.

Retries must consider idempotency. Repeating a payment or order creation call can be harmful if the operation is not designed to detect duplicates. Messaging systems may deliver more than once, so consumers should be able to recognize already processed work where the business process requires exactly-once effects.

Availability is therefore partly an application concern. Infrastructure can keep endpoints online, but code determines whether a brief dependency fault becomes a minor delay or a cascading failure.

Queue-based load leveling can improve availability when downstream systems are temporarily slow. Instead of failing user requests immediately, the application can accept work and process it asynchronously where the business process allows. This decouples traffic spikes from backend capacity and gives transient failures time to clear.

Deployments must preserve the redundancy you designed

A highly available architecture can still suffer downtime during releases if deployment procedures update every instance at once. Use rolling deployment, deployment slots, canary patterns, or blue-green techniques appropriate to the compute platform. Preserve enough healthy capacity while a new version is being introduced.

Schema changes need special care because old and new application versions may run at the same time. Backward-compatible database changes reduce the chance that a release creates a coordinated outage. Configuration changes should be versioned and reversible.

Administrators working in the scope of AZ-104 are part of this availability system because they operate resource health, monitoring, scaling, identity, and network configuration. Architecture succeeds only when deployment and operations preserve its assumptions.

Configuration should be externalized and versioned so replacement instances can start without manual intervention. Certificates, feature flags, connection settings, and secrets need a controlled distribution model. If an instance can only be restored by copying local files from a surviving server, the architecture still has hidden state.

Dependency timeouts should be coordinated across tiers. If an upstream service waits 60 seconds for a dependency that itself retries for 50 seconds, failures can accumulate until every worker is blocked. Timeouts, retry counts, circuit-breaker thresholds, and queue limits should be designed as one system so a localized failure remains localized.

Use bulkheads where one class of work could consume all capacity. Separate worker pools, queues, or resource limits can prevent a noisy background process from starving interactive traffic. This is especially important in shared compute platforms where unrelated workloads otherwise compete for the same scale unit.

Prove recovery through controlled failure testing

Reliability improves when teams test the failure modes they claim to survive. Take instances out of rotation, simulate zone loss where practical, disconnect a dependency in a test environment, and rehearse regional failover. The objective is to observe whether health probes, routing, scaling, alerts, and runbooks behave as designed.

Disaster recovery testing should measure application outcomes, not only infrastructure state. Can users authenticate? Are transactions consistent? Do background jobs resume? Are operators receiving the right alerts? How long did the failover take, and how much data was lost?

Highly available Azure applications emerge from aligned layers: resilient code, redundant compute, durable data, intelligent routing, observability, and tested recovery. Design for the specific faults the business must survive, use zones and regions where the objective requires them, and make failure behavior part of normal engineering practice. Availability is credible when the workload can demonstrate it, not when the diagram merely contains duplicate boxes.

Reliability reviews should include human response. Alerts need an owner, escalation path, and runbook. An architecture that can fail over automatically but leaves operators unable to understand what happened can create long recovery and unsafe manual actions after the initial fault. Good availability design combines resilient technology with observable, rehearsed operations.

Maintenance windows and platform updates should be included in availability assumptions. Zone-redundant and managed services reduce disruption, but application deployments, schema changes, certificate rotation, and network maintenance can still create downtime if they are coordinated poorly. Reliability engineering should treat planned change as a failure mode to be controlled, not as an exception to the availability target.

Chaos testing can be introduced gradually. Start with low-risk exercises such as terminating one instance or blocking a noncritical dependency in a test environment. As confidence grows, test zone-aware behavior, database failover, and regional routing. The goal is not dramatic failure injection; it is repeated evidence that the workload behaves predictably when components disappear.

Recovery priorities should reflect business transactions in flight. During failover, decide what happens to requests that were accepted but not completed, messages that were dequeued but not acknowledged, and users with active sessions. Resilience is stronger when the application has explicit semantics for interrupted work rather than relying on users to retry blindly.

Post-incident reviews should compare actual behavior with the architecture assumptions. If a dependency recovered more slowly than expected or a manual step dominated RTO, update the design and runbook. High availability is maintained through learning; it degrades when diagrams remain unchanged while production reality evolves.

Data protection remains part of availability even when the application is continuously online. Accidental deletion, corruption, or a bad deployment can propagate through redundant replicas. Backups, versioning, point-in-time recovery, and immutable recovery options protect against logical failures that high-availability replicas alone cannot solve.

Availability reviews should include dependency diversity, not just instance count. Two application instances in different zones are still vulnerable if both depend on one external API, one certificate authority endpoint, or one on-premises firewall. Map shared dependencies explicitly and decide whether each needs redundancy, graceful degradation, or a documented acceptance of risk.

Filed under Cloud Computing