INSIGHTS
Technology Fundamentals

ISACA CISA: IT Operations and Service Management Controls

In this article
  1. Define services, owners, and business criticality
  2. Control availability, capacity, and performance as related risks
  3. Audit incident and problem management as different control loops
  4. Govern changes, configurations, releases, and patches together
  5. Control automation, job scheduling, and system interfaces
  6. Evaluate service levels, support processes, and suppliers
  7. Protect operational logs, monitoring, assets, and databases
  8. Address shadow IT and end-user computing pragmatically
  9. Test operational controls through evidence and trends

IT operations controls are the disciplines that keep production technology dependable after systems have been designed and deployed. They cover the daily work of running services, maintaining infrastructure, processing scheduled jobs, responding to incidents, changing configurations, monitoring capacity, protecting operational logs, and meeting service commitments. For an auditor, the central question is not whether an organization uses a particular framework. It is whether operational practices consistently support business requirements and produce evidence that problems, changes, failures, and exceptions are controlled.

The current CISA examination content outline gives Information Systems Operations and Business Resilience a major role and explicitly includes availability and capacity management, incident and problem management, change, configuration and patch management, operational logging, service-level management, database management, automation, interfaces, asset management, and shadow IT. Those topics make service management an audit subject rather than merely an operations methodology. The broader ISACA certifications ecosystem also connects operational control to governance, security management, and technology risk.

Define services, owners, and business criticality

Operational control becomes difficult when teams manage servers, applications, cloud resources, and tickets without a shared definition of the business services those components support. A useful service model identifies the customer or business process, accountable owner, supporting applications and infrastructure, expected hours of operation, dependencies, recovery requirements, data sensitivity, and the teams that provide support.

Service criticality should drive the strength of controls. A payroll service, payment platform, identity provider, or production manufacturing system usually requires tighter availability, monitoring, escalation, and change controls than a low-impact internal utility. Auditors should verify that criticality is based on business impact rather than technical visibility or the preferences of the operations team.

Ownership must extend beyond a name in a configuration database. Service owners should participate in decisions about risk, maintenance windows, accepted downtime, vendor dependencies, lifecycle investment, and recovery. Technical teams may operate the platform, but someone representing the business outcome needs authority to decide what level of service is acceptable.

The concepts behind modern IT service management are useful here because service orientation turns infrastructure work into measurable value delivery. The audit focus remains evidence: whether the organization can connect operational activities to the services and outcomes they are meant to protect.

Audit scope improves when service maps are reconciled to technical inventories rather than accepted as documentation alone. The auditor can select a critical business service, trace it to applications, infrastructure, identities, databases, interfaces, and suppliers, then compare those dependencies with monitoring and support coverage. Gaps often reveal components that are operationally important but missing from ownership records, maintenance schedules, or recovery plans.

Availability is not simply uptime. A service can be technically reachable while response time, queue depth, transaction failure, or downstream dependency problems make it unusable. Operational monitoring should therefore include health indicators that reflect the user experience and business workload, not only device status.

Capacity management should anticipate growth, seasonal peaks, new products, data retention, and architectural constraints before resources become saturated. Auditors can compare forecasts, utilization trends, scaling thresholds, and actual incidents to determine whether capacity decisions are proactive or merely a reaction to outages.

Resilience mechanisms such as clustering, failover, redundant links, autoscaling, and multi-region deployment require testing. The presence of redundant architecture is weak evidence if failover has never been exercised, dependencies are not redundant, or operating teams cannot explain the conditions under which traffic moves to the alternate path.

Performance thresholds should trigger action rather than populate dashboards indefinitely. Repeated warnings, chronic resource pressure, and near-capacity conditions can reveal deferred risk. Management should know when a capacity issue becomes a funding decision, an architecture change, or a formal risk acceptance.

Capacity and availability controls should also account for planned change. Product launches, data migrations, acquisitions, seasonal demand, or infrastructure consolidation can invalidate historical baselines. Management should model the expected effect, define thresholds, and know what scaling or rollback action will occur if assumptions are wrong. Auditors can compare forecast decisions with actual utilization and incidents to determine whether planning information was reliable.

Audit incident and problem management as different control loops

Incident management restores normal service as quickly and safely as possible. Problem management seeks the underlying causes of recurring or significant incidents and prevents recurrence. Organizations often perform the first activity reasonably well because outages create immediate pressure, while the second receives less attention once service is restored.

Auditors should review classification, severity, ownership, escalation, communications, timelines, evidence preservation, and closure criteria. High-impact incidents should show who declared the incident, when technical and business stakeholders were engaged, what decisions were made, and how the organization confirmed that recovery was stable.

Problem records should connect repeated symptoms to root-cause analysis and corrective action. If the same batch failure, certificate expiry, storage exhaustion, or access outage appears every quarter, closing each ticket independently hides a control weakness. Trend analysis can reveal where temporary workarounds have become the permanent operating model.

The operational practices in IT crisis management reinforce the need for clear authority and communication during high-pressure events. For CISA-oriented audit work, the important distinction is whether the incident process restores service and whether the problem process reduces the chance of recurrence.

Govern changes, configurations, releases, and patches together

Change control should provide enough discipline to prevent unauthorized or poorly understood production modifications without turning routine engineering into bureaucracy. Risk-based change models can distinguish standard, normal, and emergency changes, but every category should have defined criteria, accountable approval, implementation evidence, and a way to detect changes that bypass the process.

Configuration management gives change records context. Without a reliable view of what is deployed, who owns it, and which relationships matter, reviewers cannot assess impact. Configuration repositories do not need to be perfect to be useful, but critical services should have accurate enough records to support troubleshooting, vulnerability response, dependency analysis, and recovery.

Patch management is another change discipline with a stronger security driver. The practices discussed in patch management automation are valuable only when the organization also controls asset coverage, testing, maintenance windows, failed deployments, exceptions, and remediation deadlines. A high patch percentage can be misleading if unknown assets or critical exceptions are excluded.

Organizational pressure for faster delivery explains why effective change management must be designed around risk rather than paperwork. Auditors should test production changes against independent deployment or configuration evidence so that unauthorized activity is not invisible simply because no ticket exists.

Control automation, job scheduling, and system interfaces

Many critical business processes depend on scheduled jobs, orchestration, file transfers, integration queues, and automated workflows that receive little attention until they fail. Controls should identify which jobs are critical, who can modify schedules or scripts, how failures are detected, and how reruns are approved when duplicate processing could create financial or operational errors.

Interfaces require reconciliation because one system can successfully transmit data while the receiving process rejects, truncates, duplicates, or misclassifies it. Control design should consider counts, control totals, exception queues, acknowledgments, retry logic, and the handling of records that cannot be processed automatically.

Automation credentials deserve privileged-access treatment. Service accounts, API keys, workload identities, certificates, and secrets can operate continuously and at scale. Their owners, permissions, rotation or federation method, and monitoring should be known. A job scheduler with broad credentials can become a high-impact control point even if no human administrator logs in directly.

Auditors should also understand how manual intervention is controlled. Emergency reruns, file edits, queue manipulation, and direct database updates may be legitimate recovery techniques, but they should create evidence and independent review proportionate to the risk of altering business records.

Batch and interface controls become especially important around financial close, billing, payroll, inventory, and regulatory reporting because a failed or duplicated process can create business records that appear valid. Auditors can inspect restart procedures, idempotency, sequence controls, and reconciliation to determine whether operators can recover safely without introducing duplicate transactions or silent omissions.

Evaluate service levels, support processes, and suppliers

Service-level management translates business expectations into measurable commitments. Useful service levels define the service, measurement method, reporting period, exclusions, target, and responsibility for corrective action. A percentage without a clear calculation or business meaning can create the appearance of control while hiding recurring degradation.

Internal operational teams should also define support expectations such as response and restoration priorities, escalation paths, support hours, and handoffs between groups. Poorly designed handoffs can cause incidents to bounce between teams while clocks continue running. Auditors can use ticket histories to test whether routing rules and escalation actually work.

Third-party services should be incorporated into the same operating model. Cloud platforms, managed services, SaaS, network carriers, and outsourced support may hold contractual service commitments, but the customer still needs monitoring, incident escalation, change awareness, and continuity planning. A vendor SLA does not automatically equal the business requirement.

CRISC provides a useful risk lens when supplier performance, concentration, or repeated service failure creates exposure above tolerance. Service-level exceptions should therefore feed risk management when operational teams cannot resolve them within their own authority.

Service reviews should distinguish chronic underperformance from isolated events. Repeated SLA misses, rising ticket age, or deteriorating vendor response can indicate that the contracted service model no longer matches business demand. Management should decide whether to add capacity, renegotiate terms, redesign the dependency, or formally accept the exposure rather than normalizing repeated exceptions.

Protect operational logs, monitoring, assets, and databases

Operational logs support troubleshooting, security monitoring, audit evidence, capacity analysis, and incident reconstruction. Controls should define which events are collected, how clocks are synchronized, how long records are retained, who can alter them, and whether important sources stop reporting without detection.

Monitoring should be tuned to actionable conditions. Excessive alert volume teaches teams to ignore warnings, while narrow monitoring misses slow deterioration. Auditors can compare incidents to alert history and ask whether indicators existed before the failure, whether thresholds were meaningful, and whether alerts reached people who could act.

Asset management is foundational because operations cannot patch, monitor, retire, or recover technology that is not known. Inventories should cover cloud resources, virtual assets, endpoints, appliances, software, and material end-user computing. Reconciliation against network, cloud, finance, or identity sources can identify blind spots.

Database operations deserve specific attention because changes to privileges, schemas, backup settings, replication, maintenance, and direct data access can affect both availability and integrity. Administrative activity should be attributable and high-risk changes should follow controlled procedures rather than informal production access.

Address shadow IT and end-user computing pragmatically

Shadow IT is not only unauthorized cloud software. It can include departmental databases, scripts, spreadsheets, low-code applications, local automation, unregistered SaaS, and integrations created outside central technology governance. Some of these solutions may be essential to business operations despite weak visibility and support.

An audit should assess business dependency before demanding immediate removal. The better control response may be to identify the owner, document data flows, restrict access, establish backup, improve change control, and move the solution into a supported environment. Forcing useful tools underground can reduce visibility without reducing risk.

End-user computing requires proportionate control. A spreadsheet used for personal analysis does not need the same treatment as a workbook that calculates regulatory capital or drives a revenue process. Criticality, complexity, data sensitivity, and the possibility of undetected error should determine review, versioning, testing, and access requirements.

Management should also investigate why shadow solutions appear. Slow delivery, missing integrations, excessive procurement friction, or weak central reporting may be root causes. Correcting the demand problem can be more effective than repeated policy enforcement.

Critical end-user tools should have an inventory and owner even when central IT does not operate them. Version control, peer review, protected formulas, input validation, backup, and controlled distribution can reduce risk without forcing every spreadsheet or low-code workflow into enterprise software development. The level of control should follow materiality and complexity.

Operations audits are strongest when they combine design review with evidence that controls operated over time. Samples can include incidents, changes, patches, failed jobs, capacity exceptions, service reports, privileged activities, and vendor escalations. Population analytics can then identify unusual patterns such as emergency-change concentration, chronic overdue patches, or repeated incidents tied to the same component.

Metrics should be decision-oriented. The ideas in IT performance management apply when indicators have owners, definitions, thresholds, reliable data, and expected actions. Mean time to restore, change failure rate, backlog age, patch exposure, availability, capacity headroom, and recurring-incident rates can all be useful when interpreted in context.

Auditors should test the integrity of reported metrics. Excluded systems, paused tickets, changing severity definitions, or incomplete monitoring can improve a dashboard without improving service. Recalculating a sample and reconciling source populations helps determine whether management reports reflect operational reality.

Effective IT operations control produces a traceable pattern: services are owned, risks are understood, changes are authorized, failures are detected, incidents are resolved, recurring causes are addressed, suppliers are managed, and evidence supports management decisions. That pattern matters more than whether every operational practice uses the same terminology or tool.

Repeat findings deserve special attention because they can indicate that management is treating symptoms rather than changing the operating system that produces them. If emergency changes remain high, patch exceptions recur, or the same service repeatedly breaches capacity thresholds, the auditor should consider whether governance, architecture, staffing, or funding is the real control issue rather than another isolated operational failure.

Filed under Technology Fundamentals