Azure monitoring becomes much easier once metrics, logs, and alerts are treated as parts of one operating system rather than as three unrelated portal features. Azure Monitor collects and exposes platform telemetry. Log Analytics is the query experience used to investigate records in Azure Monitor Logs. Alert rules evaluate metrics, logs, and activity events and turn conditions into actionable signals. The administrator’s job is to connect those pieces so that a symptom can be detected, investigated, and acted on without drowning the operations team in noise.
That model sits directly inside the current AZ-104 scope. The exam expects administrators to monitor resources and maintain Azure environments, but the same design choices determine whether a real production team sees a problem before users report it. The useful goal is not “collect everything.” It is to collect the telemetry that answers operational questions, query it efficiently, and alert only when a condition deserves attention.
Metrics, logs, and activity records answer different questions
Metrics are numerical time-series signals such as CPU percentage, request count, latency, or available capacity. They are optimized for fast evaluation and trend visualization. Logs are records with richer context: events, resource details, application messages, security data, and other fields that can be filtered and correlated. The Azure Activity Log is a separate control-plane record of subscription-level operations such as resource creation, configuration changes, and administrative actions.
The distinction matters because an operations question should lead to the right data source. “Is CPU above 90 percent right now?” is a metric question. “Which hosts produced a specific error code after a deployment?” is a log question. “Who deleted this resource group?” is an activity question. When teams send every question to the same telemetry store, they either pay to retain data that does not need rich analysis or lose context by relying only on precomputed metrics.
General network-device logging illustrates the same principle outside Azure: raw events become useful only when collection, timestamps, fields, retention, and investigation workflows are deliberate. Azure Monitor provides the platform, but administrators still have to define what each signal means operationally.
Platform metrics are also available much sooner than many log streams because they are precomputed for monitoring. That makes them attractive for fast operational alerts. Logs trade some immediacy for richer fields and correlation. A monitoring standard should state which signal is authoritative for each critical condition so teams do not create two alerts for the same failure from different stores and then page twice for one incident.
A Log Analytics workspace is a data and access boundary
A Log Analytics workspace stores tables that contain collected log data. Resources can send data into a workspace through supported collection mechanisms, and operators can query that data from Azure Monitor, from the workspace, or from a resource-centric Logs experience. Where Log Analytics is opened affects the default query scope, which is important when a query unexpectedly returns too much or too little data.
Workspace architecture should follow operational and governance requirements. One workspace can simplify cross-resource investigation, shared dashboards, and centralized alerting. Multiple workspaces may be justified by data residency, security boundaries, cost ownership, tenant design, or operational separation. The correct answer is not automatically “one per subscription” or “one for the whole enterprise.” A workspace is both a data store and an access/cost boundary.
Data retention and table configuration also matter. High-volume telemetry can become expensive if every verbose category is ingested indefinitely. Conversely, aggressive retention can make incident reconstruction impossible. Design retention around investigation windows, compliance needs, and the value of each data source. A monitoring architecture should be able to explain why a table is collected and how long it is retained.
Collection design determines what can be queried later
Monitoring failures often begin before anyone writes a query. If diagnostic settings, agents, or data collection rules do not send the necessary records to the intended destination, Log Analytics cannot reconstruct them after the fact. Administrators should therefore treat telemetry onboarding as part of resource deployment rather than an optional post-production step.
For supported platform resources, diagnostic settings can route resource logs and metrics to destinations such as Log Analytics workspaces, storage accounts, or event streaming destinations. Azure Monitor Agent and data collection rules are used for supported guest and custom collection scenarios. The exact mechanism depends on the resource and signal, but the design principle is consistent: define the signal, destination, transformation, and retention deliberately.
Current Azure Monitor guidance also makes cost awareness increasingly important. Microsoft expanded diagnostic-settings export billing in 2026 for remaining Azure resources, so “send every category everywhere” is not a cost-neutral default. Collection should support specific operational, security, or compliance outcomes. If a data stream is never queried, never alerted on, and never used for an audit requirement, its continued ingestion should be challenged.
Network telemetry illustrates the value of intentional collection. Flow information can reveal which endpoints communicated and at what volume, while a resource log may explain why a connection was rejected. Techniques described in NetFlow monitoring show how traffic records become useful when they are collected with a clear investigative purpose. Azure collection should follow the same discipline rather than ingesting every available category by default.
Log Analytics turns records into operational evidence
The Log Analytics tool supports a simple mode for common filtering and aggregation and a KQL mode for full Kusto Query Language analysis. KQL is where richer investigations become possible: filter by time and resource, project useful columns, summarize counts, join related records, parse fields, and build a query that describes the failure rather than merely listing rows.
A good query starts narrow. Select a relevant time range, constrain the resource or subscription context, and filter early on fields that reduce the dataset. Then summarize or correlate only what is needed. This improves readability and performance, and it makes the query easier to turn into a reusable workbook or alert rule. A giant query that tries to answer every incident type is harder to validate and easier to break.
Concepts from NetFlow analysis provide a useful analogy: the value is not the existence of telemetry but the ability to reduce a large event stream into patterns that explain behavior. KQL gives Azure operators that reduction layer for logs. Save queries that repeatedly answer real questions, document their assumptions, and test them against known incidents.
Query results should be designed for the person who will use them during an incident. Include resource identifiers, timestamps, severity, and the fields needed to pivot to the next question. A query that returns a beautifully aggregated count but hides which resources produced the errors may be useful for a dashboard and useless for repair. Keep separate queries for executive trends and operator diagnosis when the audiences need different levels of detail.
Metric alerts and log search alerts solve different detection problems
Metric alerts are appropriate when the signal already exists as a metric and the condition can be evaluated directly. They are efficient for thresholds such as CPU utilization, availability, queue depth, latency, and other numerical time-series data. Azure Monitor also supports dynamic thresholds for suitable scenarios, allowing historical behavior to influence the expected range rather than relying on one fixed number.
Log search alerts use a Log Analytics query and are better when detection requires richer logic. A KQL query can correlate fields, count matching events, evaluate errors across multiple resources, or detect patterns that are not available as a native metric. Simple log search alerts can evaluate individual rows for fast event-oriented scenarios, while traditional log search alerts are useful for aggregations and more advanced logic.
The choice should follow the signal, not the administrator’s favorite tool. If a native metric expresses the condition, using a log query may add latency and cost without adding insight. If the condition requires contextual fields or cross-resource logic, a metric alert may be too shallow. The alerting architecture is stronger when each rule uses the least complex signal that still represents the actual failure.
Action groups separate detection from notification and response
An alert rule should describe what condition is significant. An action group describes what should happen when the condition is met. That separation allows the same notification or automation pattern to be reused across multiple alert rules and lets operations teams change response channels without redesigning every detection.
Action groups can notify people and invoke supported automation or integration targets. The important design question is not how many channels can be configured; it is who needs the signal and what action is appropriate. A production outage may page an on-call engineer and start an automation workflow. A capacity warning may create a ticket without waking anyone. A compliance event may be routed to a security team with a different escalation path.
Alert processing rules can further control notification behavior, including planned maintenance or routing adjustments. This is more sustainable than disabling alert rules whenever a deployment or maintenance window begins. Detection can remain intact while notification behavior changes according to operational context.
Response automation should be idempotent and safe to repeat. Alerts can change state, rules can reevaluate, and integrations can retry. An automation action that restarts a service, scales capacity, or modifies a resource should first verify the current condition and avoid making the incident worse through repeated execution. Use automation for well-understood, reversible remediation and keep ambiguous failures in a human decision path.
Alert quality depends on noise control and state behavior
An alert that fires constantly and rarely requires action trains people to ignore it. Noise usually comes from thresholds that do not reflect normal workload behavior, rules that monitor symptoms with no owner, duplicated detections, or alerts that fire separately for hundreds of resources when one aggregated signal would be more useful.
Azure Monitor distinguishes stateful and stateless behavior for several alert types. Stateful alerts remain fired until the condition is considered resolved, helping prevent repeated notifications for one continuing problem. Activity log alerts are event-driven and represent discrete events rather than an ongoing metric condition. Understanding those semantics helps explain why some alerts resolve automatically while others require a human response state.
Use dimensions and resource scoping carefully. Splitting by resource or another dimension can provide the exact affected instance, but it can also multiply alert volume and cost. Review noisy rules after incidents: Was the threshold useful? Did the alert arrive early enough? Did it contain enough context? Was the responder able to act? Monitoring should improve from operational evidence, not remain frozen after initial deployment.
Ownership is another noise-control mechanism. Every production alert should have a team that can act on it and a documented first response. If no one can name an action for the alert, it is probably telemetry for a dashboard or report rather than a page-worthy condition. This test prevents monitoring programs from measuring success by the number of alert rules instead of the number of useful detections.
Monitoring access and cost are architecture decisions
Telemetry often contains sensitive operational information. Workspace permissions should therefore be designed around who needs to query which data. Azure RBAC can control access to workspaces and resources, and some organizations need stronger separation for security or compliance teams. Broad subscription permissions should not be granted merely because a user needs to run one query.
Cost is similarly structural. Log ingestion, retention, alert frequency, query dimensions, and the number of monitored time series all affect spend. The answer is not to disable monitoring; it is to use metrics for simple numerical conditions, collect logs that support real investigations, set appropriate retention, and tune alert evaluation frequency to the urgency of the signal.
Monitoring also benefits from an administrator who understands the wider Azure estate. A practical discussion of Azure administrator responsibilities shows why visibility crosses compute, networking, storage, identity, and governance. The monitoring platform should reflect those dependencies instead of creating isolated dashboards for each service with no shared operational story.
Architects should also decide where long-term evidence belongs. Some telemetry may need to remain queryable in a workspace for active operations, while other records can be archived or exported according to compliance and cost requirements. The architecture concerns reflected in AZ-305 become relevant because monitoring data is itself a workload with residency, retention, security, and resilience requirements.
A useful monitoring design closes the loop from signal to action
The strongest way to review an Azure Monitor design is to trace an incident from beginning to end. A resource emits a metric, log record, or activity event. The signal reaches the expected store. A query or alert rule evaluates it. The alert contains enough context to identify the affected resource. An action group reaches the correct responder. The responder can use Log Analytics, metrics, and change history to confirm the cause and recovery.
If any step is missing, the monitoring architecture has a gap. Collecting data without detection leaves users as the monitoring system. Alerting without investigative context creates slow incidents. Dashboards without ownership become decorative. Notifications without a runbook create escalation loops in which everyone sees the alert but nobody knows what to do.
Azure Monitor, Log Analytics, and alerts are therefore best designed as one feedback system. Collect only useful signals, preserve enough context for investigation, choose the alert type that matches the data, route it to an owner, and review the rule after real incidents. That approach produces operational visibility instead of a large telemetry bill and a crowded alert list.