Infrastructure as code applies software-engineering practices to network intent: configuration and policy are represented in machine-readable files, reviewed through source control, validated automatically, and applied through repeatable tooling. In Cisco environments, that can include direct device automation, Catalyst Center APIs and providers, Ansible collections, Terraform providers, and controller-specific models. The point is not to force every network change into one tool. It is to make intended state reproducible, reviewable, and traceable.
The current 350-401 ENCOR v1.2 blueprint includes automation, and 300-435 ENAUTO focuses on automating and programming enterprise solutions. Cisco DevNet also maintains infrastructure-as-code resources across Catalyst Center, IOS XE, SD-WAN, data center, service provider, and security products. A useful IaC program therefore starts with ownership and source-of-truth design before it starts with syntax.
Define which network state belongs in code
Not every operational value should become an IaC variable. Stable intent such as VLAN definitions, routing policy, interface roles, site templates, device credentials references, and controller objects can be good candidates. Transient state such as learned neighbors, counters, or automatically assigned runtime values belongs in observation rather than desired configuration. Mixing the two makes plans noisy and tempts automation to “correct” healthy dynamic behavior.
Start with a bounded service. For example, manage access-switch interface intent or a set of Catalyst Center site objects before attempting the entire enterprise. Define which fields are authoritative in code and which may still be changed by another system. The infrastructure-as-code model is strongest when the source file expresses a clear desired state, not when it becomes a dump of everything a device can configure.
A schema for the source of truth prevents every team from inventing different names for the same concept. Define types for site, device role, interface role, prefix, peer, virtual network, security zone, and other objects the automation consumes. Validate relationships such as “this subnet belongs to exactly one site” or “this peer address is not inside a local interface range.” Good data modeling prevents more incidents than clever templating.
Use version control as the operational history of intent
Git or another version-control system should record who changed the desired network state, what changed, why it changed, and which review approved it. Small commits make differences understandable. Branches and pull requests allow proposed network changes to be reviewed before they reach a maintenance pipeline. Tags or release references can tie a deployment to the exact code version that generated it.
Do not treat the repository as a file share. Require useful commit messages, peer review for high-impact paths, and protected branches for production. The concepts behind Git workflows become operational controls in network engineering: a diff can show an intended route-policy change before it affects traffic, and history can identify when an unsafe value entered the configuration model.
Reviewers need context beyond a raw diff. A change request should describe expected forwarding impact, failure domain, maintenance window, and validation plan. CI systems can attach rendered configurations or Terraform plans to the pull request so a network engineer can see the practical effect. This keeps code review grounded in network behavior instead of turning approval into a syntax exercise.
Separate data, templates, modules, and secrets
IaC becomes difficult to maintain when site-specific values are copied into large templates or when credentials are embedded directly in configuration files. Separate reusable logic from environment data. A module can describe how a standard branch is built, while inventory or variables provide the site code, prefixes, peers, or interface mappings. The separation reduces copy-and-paste drift and makes changes to the standard easier to review.
Secrets should be references to a managed secret store or protected pipeline variable, not plaintext committed with the rest of the repository. Rotate credentials without rewriting every template. Limit automation identities to the scope they require. Security best practices for Terraform and other tooling matter because the state file itself can contain sensitive values even when the source configuration does not.
Environment boundaries should be explicit. Development, lab, staging, and production may use the same modules but different inventories, credentials, approval rules, and provider endpoints. Prevent a developer’s local variables from targeting production by mistake. Pipeline identity and protected environment settings should make the production destination deliberate and auditable rather than a command-line flag that can be changed casually.
Choose Ansible for procedural orchestration when that model fits
Ansible is effective for workflows that gather facts, render templates, execute ordered tasks, call APIs, or perform conditional steps across groups of devices. Network collections can provide structured modules rather than requiring raw CLI for every action. Playbooks also express sequencing clearly, which is useful for rolling changes where one device must be validated before the next one is touched.
Prefer idempotent modules and explicit checks over unconditional command pushes. If a task must use CLI, verify the current state and account for platform differences. The practical distinction in Ansible and its orchestration layer matters less than maintaining predictable playbooks: inventory must be correct, variables must be controlled, and every task should have a reason to change the device.
Facts gathered at the start of a play can become stale during long workflows. If a later decision depends on routing or interface state after an earlier task, gather the relevant state again instead of relying on the original snapshot. Treat network automation as interaction with a changing distributed system, not as file manipulation. This is especially important during rolling maintenance where neighboring devices are changing between stages.
Choose Terraform when resources and state management fit the problem
Terraform is designed around resources, providers, desired configuration, and state. Cisco provides Terraform providers for multiple platforms, including Catalyst Center. This can be useful when the network object has a lifecycle that maps cleanly to create, read, update, and delete operations. A plan can show proposed changes before apply, giving reviewers a structured view of how the managed resources will change.
State management is part of the architecture. Protect remote state, control concurrent access, and understand which values the provider stores. Import existing resources carefully before declaring them managed. Do not use Terraform to recreate resources that another system owns, because two controllers competing over the same object create drift loops. IaC is about one authoritative intent per object, even when several tools exist in the organization.
Use controllers as the abstraction layer when they own higher-level intent. Catalyst Center can expose APIs and Terraform resources for sites, credentials, templates, device operations, and other controller-managed objects. Managing those through the controller can be safer than configuring every device independently because the controller understands relationships and supported workflows. The code describes controller intent, and Catalyst Center translates that intent into device operations.
Keep the abstraction boundary clear. If Catalyst Center owns an SD-Access fabric, do not also let a separate direct-device pipeline modify fabric-specific CLI without a deliberate exception process. The controller’s database and the Git repository must not disagree about who owns the object. Cisco’s IaC tooling is most effective when each layer—source control, orchestrator, controller, and device—has a defined responsibility.
A Terraform plan is valuable but not infallible. Providers may compute values at apply time, APIs can reject valid-looking requests, and external changes can make the state stale between plan and apply. Treat the plan as a proposed transaction, then run operational postchecks. Network service acceptance still depends on routing, reachability, and policy behavior after the provider reports success.
Validate changes before production with linting, schemas, and labs
Pull-request automation should check syntax, schema validity, required variables, naming standards, forbidden commands, duplicate addresses, and other rules that can be tested without touching production. Where possible, render the candidate configuration or API payload and compare it with policy. More advanced pipelines can run integration tests against virtual devices, sandboxes, or a lab that represents production software releases.
Tests should reflect network risk rather than imitate software testing superficially. A route-policy change might require verifying that an expected prefix remains preferred and a forbidden prefix is still denied. A campus interface change may require authentication and VLAN checks. The broader automation-versus-orchestration distinction is useful here: a pipeline should coordinate validation, deployment, and evidence, not merely call a configuration script.
Policy-as-code checks can block known dangerous patterns before they reach review. Examples include allowing a default route from an untrusted peer, disabling an access protection, using an unapproved VLAN range, or targeting both members of a redundant pair in one stage. Keep policies explainable and versioned so engineers understand why a pipeline rejected a change and can propose a controlled exception when needed.
Detect drift without automatically overwriting legitimate emergency work
Drift occurs when actual state differs from declared intent. It may indicate an unauthorized manual change, a failed deployment, a software default, or a legitimate emergency fix. Continuous comparison is valuable because it exposes differences before they surprise the next deployment. However, blindly forcing every difference back to the repository can erase a temporary change that is protecting production.
Define an emergency-change path that records the manual action and creates a follow-up task to reconcile code. Classify drift by risk: some differences can be corrected automatically, while routing, security, or interface changes may require review. The goal is not zero human CLI activity under all circumstances. The goal is that every lasting production state eventually returns to an approved, reproducible source of truth.
Some controllers and providers maintain their own state, while a Git repository expresses desired intent at a higher level. Reconciliation should compare the correct layer. A difference between generated device CLI and Git may be harmless if the controller intentionally renders release-specific commands. Compare semantic intent where possible, and avoid drift alarms that teach operators to ignore the system because every harmless formatting change appears critical.
Infrastructure as code does not guarantee that every network device remains identical to the repository. Emergency CLI changes, failed jobs, software upgrades, feature defaults, and controller-side modifications can create drift. Decide which source is authoritative and schedule comparisons between intended state and observed state. Not every difference should be overwritten automatically: some values are device-generated, some are operational, and some emergency changes may need to be preserved until reviewed. A useful drift report classifies differences by ownership and risk so engineers can reconcile them deliberately instead of simply forcing the last committed template back onto the network.
Validation should progress through layers. Start with linting, schema checks, and unit tests for variables and templates. Then render intended configuration or API payloads and verify them against platform capabilities in a lab or simulation environment. In production, use a small canary scope before broad deployment and run post-change checks that test routing adjacencies, reachability, policy, and telemetry rather than only confirming that an API returned success. Secrets, tokens, and private keys must stay outside repositories and logs; use a controlled secret store and short-lived credentials where the platform supports them. This turns IaC into a governed delivery system instead of a faster way to push configuration.
Roll out changes by failure domain and preserve rollback evidence
IaC makes it easy to apply one change to hundreds of objects, which increases both consistency and blast radius. Use canary sites, maintenance groups, or staged waves so that an error is discovered before it reaches the entire enterprise. Sequence redundant devices so both members are not modified simultaneously. Pause between stages long enough to evaluate routing, client, and application health.
Rollback should be designed before deployment. A previous Git commit is helpful only if the automation can safely return the network to that state and if external systems have not changed in the meantime. Capture prechange state, deployment logs, postchecks, and the exact code version applied. The network-automation promise becomes credible when recovery is as repeatable as deployment.
Approval gates should become stricter as blast radius increases. A description change on one access port does not need the same review as a BGP policy change on every Internet edge. Encode risk categories into the pipeline so low-risk routine work remains efficient while high-impact network changes require additional reviewers, maintenance timing, and prechange evidence. Good IaC improves speed by automating governance as well as configuration.
Measure IaC success by change quality, not by the percentage of automated commands
A mature program should reduce configuration drift, failed changes, time spent on repetitive tasks, and recovery time after mistakes. It should also improve review quality and make network intent easier for another engineer to understand. Automating a fragile process can simply make failures happen faster. Track which classes of change benefit from IaC and which still require interactive engineering or vendor-specific tooling.
Keep platform teams and network engineers jointly responsible for the system. Software engineers bring testing and pipeline skills; network engineers understand protocol behavior and failure domains. Neither role can safely replace the other. Infrastructure as code succeeds when a proposed change can be reviewed as intent, tested as data, deployed through controlled tooling, and verified against the actual forwarding behavior of the Cisco network.
Feed production telemetry back into the delivery process. If a rollout increases interface errors, client onboarding time, route churn, or application latency, the pipeline should stop before the next wave. This is where IaC and assurance reinforce each other: code defines intended configuration, while network telemetry verifies that the intent produced acceptable behavior. Automation maturity is the closed loop between declaration, deployment, and observed outcome.
Documentation can be generated from the same data model when useful. Site inventories, peer lists, address allocations, and policy matrices derived from code reduce the chance that operational documents drift from production. Generated documentation should still be readable by engineers during an outage without requiring them to understand the entire pipeline. The source of truth earns trust when it helps both automation and humans answer operational questions.