INSIGHTS
Cloud Computing

Microsoft AZ-305: Storage Design for High-Scale Azure Applications

In this article
  1. Start with access patterns, not product names
  2. Relational services are still the default for strong transactions
  3. Cosmos DB fits globally distributed, partitioned access patterns
  4. Blob and Data Lake storage are better for large objects and analytical files
  5. Separate the system of record from search and cache models
  6. Partitioning is the core scale mechanism
  7. Consistency should match business invariants
  8. Performance and cost need a common capacity model
  9. Operate storage as a dependency with explicit limits

Storage design at scale is not a choice between “SQL” and “NoSQL,” nor is it solved by picking the service with the largest published throughput number. High-scale Azure applications usually have several distinct data patterns: transactional records, documents, files, events, caches, search indexes, telemetry, and analytical history. The architect’s task is to map each pattern to a storage model with acceptable consistency, latency, durability, scale, governance, and cost.

That technology-selection discipline is central to AZ-305. Microsoft’s current Azure Architecture Center recommends identifying access patterns first, mapping them to data models, shortlisting services, and then applying requirements such as performance, security, consistency, and cost. This naturally leads many large systems toward polyglot persistence rather than one database forced to solve every problem.

Start with access patterns, not product names

Write down how the application reads and writes data. Does it need point lookups by key, multi-row ACID transactions, full-text search, time-window aggregation, streaming ingestion, large object delivery, or analytical scans? Different data models are optimized for different shapes of work.

A transactional order system often benefits from a relational database because relationships and consistency are central. A content service might store large media objects in Blob Storage and metadata in a database. A telemetry system might use a time-series or analytical store. A global product catalog with flexible schema might fit a document model.

The overview in block, file, and object storage illustrates the same principle at the storage layer: the way applications access data should determine the storage abstraction.

Classify data by lifecycle as well as access pattern. Some records are authoritative operational state, some are derived analytics, and some are immutable evidence retained for compliance. Those classes can justify different stores and retention rules even when they originate from the same business transaction.

Relational services are still the default for strong transactions

Azure SQL Database, Azure SQL Managed Instance, Azure Database for PostgreSQL, and Azure Database for MySQL provide managed relational options. They are strong candidates when the workload depends on transactional consistency, joins, constraints, mature SQL tooling, and well-understood schemas.

Scale requirements should be examined carefully rather than assuming relational means “small.” Azure SQL Database includes service tiers and architectures for significant scale, while PostgreSQL and MySQL services support read scaling and other performance options. The limit is often not the database category but the application’s partitioning, indexing, transaction shape, and connection behavior.

Use IaaS-hosted databases only when the workload needs operating-system or engine-level control that managed services cannot provide. The additional control comes with patching, backup, high availability, and maintenance responsibilities.

Relational service choice should include compatibility and operational requirements. Azure SQL Managed Instance can suit applications that depend on SQL Server features that are harder to move directly to Azure SQL Database, while PostgreSQL or MySQL may match open-source application stacks. The goal is to minimize unnecessary engine ownership without forcing a rewrite that does not deliver business value.

Cosmos DB fits globally distributed, partitioned access patterns

Azure Cosmos DB is designed for horizontally partitioned, globally distributed data with low-latency access and multiple consistency options. It is a strong fit for document and key-value style workloads that can be modeled around partition keys and that benefit from global replication.

The partition key is an architectural decision. A poor key can create hot partitions, uneven storage, and throughput bottlenecks. Choose a key that distributes load and aligns with common queries. Cross-partition queries are possible, but a design that constantly scans many partitions can undermine the scale benefits.

The existing Cosmos DB provides useful context, but high-scale design still requires workload-specific modeling, consistency decisions, and capacity planning.

Cosmos DB consistency and region settings affect both user experience and cost. Stronger consistency can increase cross-region coordination, while additional regions add replicated capacity. Model the read and write distribution of users so that global replication supports a real latency or resilience requirement rather than being enabled everywhere by default.

Blob and Data Lake storage are better for large objects and analytical files

Azure Blob Storage is a scalable object store for unstructured data such as images, documents, backups, logs, and application artifacts. Azure Data Lake Storage adds hierarchical namespace capabilities that suit analytics and large-scale data processing. These services should not be treated like general-purpose relational databases.

Object stores work well when the application can address whole objects or ranges and when metadata queries are relatively simple. They also support lifecycle tiers, which can reduce cost for data that becomes colder over time. For analytical systems, separating low-cost durable storage from compute can make large historical datasets economical.

The data-engineering perspective in Azure data storage and processing is useful because high-scale applications often feed downstream analytics even when the operational workload itself is transactional.

Object naming and lifecycle design matter at scale. A flat namespace can contain enormous numbers of blobs, but applications still need predictable prefixes, metadata, and retention rules. Data Lake hierarchical namespace can make directory-like operations more efficient for analytics workflows. Choose the namespace behavior that matches the processing model rather than only the storage price.

Separate the system of record from search and cache models

Search indexes and caches are usually derived stores, not the authoritative source of business data. Azure AI Search can provide full-text and semantic search capabilities. Azure Managed Redis can provide low-latency cache access. Both can improve application performance when they are fed from a durable system of record.

Treating a cache as authoritative creates difficult recovery semantics. Treating a search index as the primary database makes transactional updates and consistency harder than necessary. Instead, design a synchronization path so that derived stores can be rebuilt if they are lost.

This principle also supports independent scaling. Search query volume can increase without forcing the transactional database to serve complex text queries, and a cache can absorb hot reads without changing the durable storage model.

Derived stores need a rebuild strategy. If the search index is lost, can it be recreated from the system of record? If a cache is flushed, will the database survive the burst of misses? If analytics data is regenerated from an event stream, how long does replay take? Designing recovery for derived state prevents performance components from becoming hidden authoritative dependencies.

Partitioning is the core scale mechanism

Large systems eventually reach the point where one compute or storage partition cannot handle all traffic. Partitioning distributes data and load across scale units. The partition key should minimize hot spots, keep related operations together where useful, and support the most important queries.

Relational systems may use sharding, table partitioning, read replicas, or scale-out service tiers. Cosmos DB uses logical and physical partitions. Blob Storage naturally distributes objects but applications still need naming and access patterns that avoid bottlenecks. The exact mechanism changes, but the architectural question remains: how is load distributed as the dataset and request rate grow?

High-scale storage is easier when the application understands partition ownership. A design that requires every request to broadcast to every partition will struggle even if the backend can technically scale.

Partition keys should be tested against both present and future skew. A customer ID may distribute load well until one enterprise customer generates half of all traffic. Composite or synthetic partitioning can spread that workload, but it may make queries more complex. Load tests should use realistic tenant and key distributions rather than uniform synthetic data.

Consistency should match business invariants

Strong consistency is valuable when users must immediately see the latest committed state or when conflicting updates would violate business rules. Eventual or bounded-staleness models can improve availability and global performance where brief divergence is acceptable. The correct choice depends on the data, not on a general preference for one consistency model.

Different parts of the same application can use different guarantees. A financial ledger may require strict transactional behavior, while a recommendation feed can tolerate stale data. Architecture should identify the invariants that cannot be violated and spend consistency cost there.

Multi-region replication also affects consistency. Synchronous coordination across distant regions increases latency. Asynchronous replication improves local responsiveness but creates a recovery point gap. Those tradeoffs belong in the application design and recovery objectives.

Schema and consistency choices interact. Denormalizing data can reduce cross-partition reads and joins, but it creates multiple copies that must be updated. Event-driven propagation can keep those copies synchronized asynchronously. The architecture should identify which fields can be eventually consistent and which must change atomically.

Performance and cost need a common capacity model

Storage cost comes from more than capacity. Transactions, provisioned throughput, replicas, indexes, backups, data transfer, and performance tiers can all matter. A design should estimate both data growth and operation rate. One terabyte of rarely accessed archives has a different cost profile from one terabyte receiving millions of transactional reads and writes.

Use hot, cool, cold, or archive tiers where the service supports lifecycle management. Remove unnecessary secondary indexes. Compress data where appropriate. Consider whether every replica, region, or backup copy has a business requirement. At the same time, do not reduce redundancy below the recovery objective merely to lower the bill.

The concept of intelligent data storage is essentially this alignment between performance, placement, and cost. Storage architecture should adapt to the value and access pattern of the data.

Performance testing should measure latency percentiles, not only averages. A store that responds quickly most of the time but has severe tail latency can create poor user experience and timeouts under load. Test at expected data volume, concurrency, partition distribution, and regional topology so the benchmark resembles production behavior.

Write paths should be designed for idempotency when retries are possible. A transient timeout can leave the client uncertain whether a write succeeded. Idempotency keys, conditional updates, transaction identifiers, or upsert semantics can prevent retries from creating duplicate business records.

Event-driven architectures should decide where the event log fits in the data model. A durable event stream can support asynchronous projection into search, cache, and analytics stores, but it introduces ordering, replay, and schema-evolution concerns. The event system should not become an accidental second system of record without clear ownership.

Operate storage as a dependency with explicit limits

Every Azure data service has quotas, scale units, throttling behavior, connection limits, and regional characteristics. Architects should identify the limits that could become bottlenecks and expose them through monitoring. Application code should handle transient throttling safely rather than assuming every request will succeed immediately.

Backup, restore, encryption, private networking, identity, and observability also belong in the storage design. Administrators working within AZ-104 responsibilities need to operate those controls after deployment. A storage service that performs well in a benchmark but cannot meet recovery, security, or operational requirements is not a valid architecture choice.

High-scale Azure storage works when data models follow access patterns, partitioning distributes load, consistency reflects business invariants, and derived stores are separated from systems of record. Choose relational, document, object, search, cache, and analytical services for the jobs they are designed to perform. The result is usually not one universal database but a small set of purposeful stores that can scale independently without losing governance or recoverability.

Data governance should remain consistent across polyglot storage. Classification, encryption, private access, backup, retention, audit, and deletion requirements apply whether the data is in SQL, Cosmos DB, Blob Storage, or a search index. A multi-store architecture is powerful only when the organization can still answer where customer data exists, who can access it, how long it is kept, and how it is recovered.

Backup and point-in-time recovery need application-level validation. Restoring a database is not enough if dependent blobs, search indexes, or messages are from a different point in time. Document which stores must be recovered together and how the application reconciles state after restoration. Distributed storage increases the importance of coordinated recovery testing.

Finally, plan for data migration between models. Requirements change, and a service selected today may become a constraint later. Encapsulate storage access behind clear application boundaries, keep schemas versioned, and maintain export paths for critical data. Good high-scale storage design optimizes for the current workload without making future evolution unnecessarily expensive.

Regional service availability should be checked before standardizing the storage stack. A workload deployed globally may find that a preferred database tier, redundancy option, or analytics service is not identical in every region. Region selection and data architecture therefore need to be reviewed together rather than sequentially.

When several stores participate in one user request, define how partial failure is handled. A successful database write followed by a failed search-index update should not leave the application permanently inconsistent. Outbox, change feed, or retry patterns can make derived updates recoverable and observable.

Operational ownership should be clear for every store. A polyglot design can reduce technical coupling while increasing platform responsibility. Assign owners for schema evolution, backup, capacity, security, and incident response so no datastore becomes an orphaned component that everyone depends on but no team actively maintains.

Service-level objectives for storage should cover more than durability. Define expected latency, throughput, recovery time, recovery point, and acceptable throttling behavior for each important data path. This gives monitoring a meaningful target and helps architects decide when a store needs scaling, repartitioning, a different tier, or a different technology altogether.

Filed under Cloud Computing