Beacon

Autonomous service desk — architecture blueprint

A reference design for a multiagent framework that runs an IT service desk and infrastructure team end to end, with one deliberately narrow exception: a human still approves or arbitrates every action.

What this is: a documented architecture — diagrams, an agent taxonomy, a risk-tiering model, and the guardrails that make broad autonomy survivable — written up as a blueprint for a reader who wants to build this against their own environment. What this isn't: a running deployment. This box holds no ServiceNow instance, no Cisco/AD/VMware credentials, and doesn't attempt to acquire any. Connecting a live agent framework to someone's production network, directory, and hypervisors is exactly the kind of consequential, hard-to-reverse decision that belongs to whoever owns that infrastructure, made deliberately, with real scoped credentials they grant — not something to fabricate here. See "What this is and isn't," below, for the full reasoning.

Why a human still approves everything

"Make it as completely autonomous as possible" and "the only human interaction should be to approve or arbitrate" are the same design goal stated twice, and the second half is load-bearing, not decorative. Every domain agent below can read, plan, and propose without limit. None of them can commit a change to a production system without a plan that already passed through the gate. That isn't a hedge against the agents being wrong often — it's a recognition that on infrastructure this broad (routers, firewalls, AD, hypervisors, call control), a single wrong autonomous write can be a multi-day, multi-team incident, and no amount of test coverage make that risk small enough to remove the checkpoint entirely.

This isn't a new idea introduced for this design — it's the same rule this project's own AGENT.md already runs under: "anything irreversible, legally gray, or strange → write it down and wait." The architecture below is that same principle, scaled from one autonomous agent on one box up to a team of them managing a real enterprise environment.

High-level architecture

ServiceNow Incident · Change · Request · CMDB (system of record) Orchestrator agent intake · plan · risk-tier · route Human approval & arbitration gate the one mandatory checkpoint 1 NetworkIOS/NX-OS 2 IdentityAD / resets 3 Win SrvWinRM/DSC 4 Linux SrvSSH/Ansible 5 DatabaseSQL agents 6 Firewallrules/NAT 7 VoiceCUCM/phones 8 VMwarevCenter API 9 DesktopIntune/Jamf Managed infrastructure network & voice · identity & servers · databases · virtualization · endpoints NetBoxIPAM · DCIM Zabbix/Grafanamonitoring Batfishpre-change check Ansibleexecution engine Cisco ISENAC · RADIUS Nexus Dash.DC fabric Cisco FMCfirewall policy Catalyst Ctr.SD-Access VMware Ariacapacity & automation SCOMon-prem monitoring Azure Monitorcloud monitoring MECMendpoint mgmt RH SatelliteRHEL lifecycle APC powerUPS · PDU CyberArkPAM · vault CrowdStrikeEDR SplunkSIEM NetAppstorage Dell storagePowerStore Dell PSwitchtop-of-rack BackupVeeam/Cmvlt 10 Platform Ops — self-healing agent watches every agent & the orchestrator · heals itself · proposes patches · never bypasses the gate

The supporting-systems band along the bottom is colour-coded by category (see legend). Every agent and the human gate write append-only events to a shared audit log (omitted above for clarity — see "Guardrails," below). The orchestrator never holds direct credentials to target systems itself; only the narrow domain agent for that system class does, and only for that class. The dashed supporting-systems band feeding Platform Ops is inventory, monitoring, and verification tooling — see "Supporting systems," above — not an eleventh agent with its own credentials to target systems. The dashed Platform Ops box is different in kind from the nine domain agents above it: it supervises the framework itself rather than owning one target system class — see "Self-healing, self-patching, self-maintaining," below.

ServiceNow as the system of record

Requests never arrive as a chat message to an agent — they arrive as a ServiceNow record: an Incident (something's broken), a Request (someone needs something provisioned or reset), or a Change (a planned modification). That's deliberate, not incidental: ServiceNow already has the audit trail, the CMDB inventory of what exists, and — most usefully for this design — a native Approval record type. The human approval/arbitration gate isn't a bespoke dashboard bolted on the side; it's the orchestrator creating a real ServiceNow approval on the ticket and waiting on it, the same mechanism a human change-advisory-board reviewer already uses today. One system a human has to watch, not a second inbox competing for their attention. A ticket also doesn't always start with a person: a Zabbix alert (see "Supporting systems," below) can open an Incident directly off a monitoring trigger, entering the exact same intake path as one a human filed by hand.

The nine domain agents

Nine of the ten agents in this design each own one target system class. The tenth, Platform Ops, is different in kind — it supervises the framework's own health rather than one class of managed infrastructure, so it gets its own section below rather than a row that doesn't quite fit this table.

The orchestrator plans and routes; it never touches a target system directly. Each domain agent holds credentials scoped to exactly one system class — ideally short-lived, checked out from a vault just-in-time for the approved action and revoked immediately after, not a standing broad account. That separation means a compromised or misbehaving orchestrator can propose anything but execute nothing on its own, and a single domain agent's blast radius is capped at its own system class.

#Domain agentTalks toTypical actionsDefault tier
1NetworkIOS/NX-OS via NETCONF/RESTCONF or Ansible; NetBox for inventory/IPAM, Batfish for pre-change verification, Nexus Dashboard for DC fabric, Cisco Catalyst Center for SD-Access campus/branch, Dell PowerSwitch on SmartFabric OS10 for top-of-rack/leaf switchinginterface/VLAN checks, ACL review, config backup, port togglesTier 2
2Identity & ADLDAP/PowerShell over WinRM, Graph API, Cisco ISE for NAC/RADIUS, CyberArk for privileged-account vaultingpassword resets, group membership, account unlock/disableTier 1
3Windows ServerWinRM, PowerShell DSC; CrowdStrike Falcon sensor health; NetApp/Dell storage for volume capacityservice restarts, patch status, disk/log cleanupTier 1
4Linux ServerSSH, Ansible; Red Hat Satellite for fleet-wide RHEL patch/content lifecycle; CrowdStrike Falcon sensor healthservice restarts, package/patch checks, log rotationTier 1
5Databasenative SQL drivers, read-replica first; NetApp/Dell storage snapshots and the backup/recovery platform for backup verificationquery diagnostics, index/maintenance jobs, backup verificationTier 2
6Firewallvendor REST API (Palo Alto/Fortinet/ASA-style); Cisco Firepower Management Center for fleet-wide FTD/ASA policyrule/NAT review, temporary rule add with expiry, log queriesTier 3
7Voice / collaborationCisco CUCM AXL/APIphone reprovisioning, line/extension changes, directory syncTier 1
8VMwarevCenter API / PowerCLI; VMware Aria Operations/Automation for capacity analytics and workflows; NetApp/Dell datastores; the backup/recovery platform for VM restore verificationVM power state, resource checks, snapshot managementTier 2
9Desktop (Windows & Apple)Intune / Jamf APIs; Microsoft MECM for on-prem/co-managed Windows patching; CrowdStrike Falcon for endpoint EDR/containmentapp deploys, compliance checks, remote wipe requestsTier 2

"Default tier" is a starting point per action type, not a fixed property of the agent — a Firewall agent's read-only log query is Tier 0; the same agent's permanent rule change is Tier 3. See the risk-tier matrix below for how tier is actually decided.

Supporting systems — inventory, monitoring, security, storage & network assurance

The nine domain agents above are the actors; they don't operate blind. A handful of existing operations tools do the sensing, verification, and execution work underneath them — the architecture slots into these rather than reinventing what they already do well. None of it opens a second door around the human approval/arbitration gate: each tool below is a data source or execution backend for one specific domain agent (or for Platform Ops), governed by the exact same credential-scoping and tier rules as everything else in this design.

SystemRoleFeeds / used by
NetBoxInventory & IPAM/DCIM — source of truth for device inventory, IP address space, and DNS/DHCP scope assignmentsNetwork agent reads before it writes; Platform Ops diffs live config against NetBox the same way it diffs against the CMDB
ZabbixMonitoring & alerting across network, server, and hypervisor layersPlatform Ops (heartbeat/health source); a Zabbix trigger can open a ServiceNow Incident directly, so a ticket doesn't always start with a person
GrafanaDashboards over Zabbix and the audit logHuman approvers & Platform Ops — the actual screen behind the health-dashboard mockup, not a bespoke UI built for this project
BatfishOffline network config verification — validates a proposed change against a modeled copy of the real topology before anything touches a deviceNetwork agent — this is what "mandatory dry-run" concretely means for Tier ≥2 network changes, not just a live-device diff
AnsiblePlaybook-driven execution & config managementNetwork and Linux Server agents — idempotent playbooks double as the paired rollback/undo step the guardrails already require
Cisco ISENetwork access control — RADIUS/TACACS+, 802.1X posture, endpoint quarantineIdentity agent — extends it from "does this account exist" to "is this device allowed on the network right now"
Cisco Nexus DashboardData-center fabric (ACI/NX-OS) telemetry & orchestrationNetwork agent, scoped to DC fabric — distinct from the box-by-box campus/branch config the base agent already handles
Cisco Firepower Management CenterCentralized policy management for Cisco FTD/ASA firewalls — access control, IPS rules, and fleet-wide deploymentFirewall agent — extends it from a single device's REST API to fleet-wide policy push, rollback, and centralized event/log query
Cisco Catalyst CenterSD-Access intent-based campus/branch network automation and assurance (formerly DNA Center)Network agent — the campus/branch counterpart to Nexus Dashboard's DC-fabric role, fleet-wide assurance rather than box-by-box CLI
VMware AriaCapacity/performance analytics (Aria Operations) and self-service automation workflows (Aria Automation), formerly vRealizeVMware agent — trend-based capacity planning and pre-built orchestration above raw vCenter/PowerCLI scripting
Microsoft SCOMOn-prem infrastructure and application monitoring via management packsPlatform Ops, alongside Zabbix — a SCOM alert can open a ServiceNow Incident directly, same as a Zabbix trigger
Azure MonitorCloud-native metrics, logs (Log Analytics/KQL), and alerting for Azure and Arc-connected hybrid resourcesPlatform Ops — the cloud counterpart to SCOM/Zabbix for workloads that have moved to Azure
Microsoft MECMOn-prem endpoint lifecycle management — OS deployment, patching, and software distribution (formerly SCCM)Desktop agent — covers the on-prem/co-managed half of the Windows fleet that Intune's cloud-native MDM doesn't reach alone
Red Hat SatelliteFleet-wide RHEL patch/content/subscription lifecycle management and Kickstart provisioningLinux Server agent — the RHEL analogue of MECM for Windows, staging patches across host groups instead of one SSH session at a time
APC power managementUPS/PDU telemetry and control via Network Management Cards and StruxureWare/EcoStruxure IT — battery status, load, runtime, and remote outlet switching for rack powerPlatform Ops — a facilities-layer health source alongside Zabbix/SCOM/Azure Monitor, and the trigger for an orderly power-loss shutdown sequence handed to the Windows/Linux/VMware agents rather than a hard outage
CyberArkPrivileged access management — the vault that issues the short-lived, narrowly-scoped credential leases every domain agent checks out just-in-time, plus session isolation and recording for any human break-glass accessEvery domain agent and the orchestrator — this is the concrete implementation of the "credentials checked out from a vault just-in-time and revoked immediately after" rule the design already assumes; Platform Ops watches lease issuance and expiry
CrowdStrike FalconEndpoint detection & response — sensor telemetry, threat detection, and host containment across servers and desktopsDesktop, Windows Server, and Linux Server agents for sensor health and patch context; Platform Ops for security signal — a Falcon detection can open a ServiceNow Incident directly, and host isolation is a Tier 3 containment action, never auto-executed
SplunkSIEM & log analytics — cross-system event correlation, security analytics, and a query surface over the immutable audit logPlatform Ops, alongside Zabbix/SCOM/Azure Monitor — a correlation search can raise an Incident; it is also the search layer human reviewers use over the append-only audit log
NetApp ONTAPEnterprise storage — NFS/SMB/iSCSI capacity, volume/LUN provisioning, snapshots, and SnapMirror replication via the ONTAP REST APIVMware, Database, Windows Server, and Linux Server agents for datastore/volume capacity and snapshot state; provisioning and any destructive volume/snapshot action is Tier 3
Dell storageEnterprise storage — PowerStore/PowerMax/Unity arrays: capacity, provisioning, snapshots, and replication via the Dell REST APISame consumers as NetApp — the design doesn't lock to one storage vendor; both slot into the same read-mostly telemetry role with provisioning gated at Tier 3
Dell PowerSwitch (top-of-rack)Top-of-rack data-center switching on SmartFabric OS10 — leaf/ToR port, VLAN, and fabric configuration via the OS10 REST API or AnsibleNetwork agent, alongside Cisco IOS/NX-OS — Batfish models OS10 configs the same way, so pre-change dry-run coverage extends to the ToR layer
Backup & recoveryEnterprise backup, immutable/air-gapped copies, and restore orchestration (Veeam/Commvault/Rubrik/Cohesity)Database and VMware agents for backup-job status and restore verification; Platform Ops watches job success/failure — restores are Tier 2, and anything that deletes or shortens retention on a backup is on the deny-list, never auto-executed at any tier

IPAM's DNS/DHCP scope management is covered by NetBox's own IPAM module above rather than a separate product — the architecture doesn't lock into one vendor there; an Infoblox or BlueCat deployment slots into the same Network-agent role if that's what a given environment already runs.

Setup steps, API authentication, and a working MCP server for each system in the table above — including a full worked example for Cisco ISE — live on the integration guide.

Request lifecycle

Ticket created in ServiceNow Incident / Change / Request Intake & classification affected systems, request type Orchestrator drafts a plan steps, target agent(s), risk tier Tier 0–1 & policy checks pass? yes Auto-execute logged, reversible only no Human approval / arbitration approve · reject · edit the plan Domain agent executes scoped session; dry-run first for Tier ≥2 Verify target state matches the intended outcome Close ServiceNow ticket + append-only audit log entry verification failed → rollback & escalate

Risk tiers & the approval matrix

Tier is decided per action, computed from the plan's blast radius (how many systems/users it touches), reversibility (is there a clean rollback), and whether it's read-only. This is the actual gate the diamond in the lifecycle diagram checks.

blast radius · harder to reverse Tier 0 — read-only / diagnostic pull config · check status · query logs no approval — always auto Tier 1 — low-risk, reversible, single-target password reset · restart one service · unlock one account auto, logged & notified after Tier 2 — medium risk or multi-target firewall rule w/ expiry · VM resize · patch a server group one human approval first Tier 3 — high blast radius, hard/impossible to reverse AD schema change · core switch config · mass account change two-person approval + dry-run

The deny-list below sits above every rung on this ladder — those actions are never agent-executable no matter who approves.

TierMeaningExampleRequired approval
Tier 0Read-only / diagnosticpull config, check status, query logsnone — always auto
Tier 1Low-risk, reversible, single-targetpassword reset, restart one service, unlock one accountnone, but logged and notified post-hoc
Tier 2Medium risk or multi-targetfirewall rule with expiry, VM resize, patch a server groupone human approval before execution
Tier 3High blast radius or hard/impossible to reverseAD schema change, core switch config, mass account changetwo-person approval + mandatory dry-run/simulation

A fixed deny-list sits above all four tiers and applies regardless of who approves: disabling MFA, deleting backups, mass account deletion, and firmware wipes are never agent-executable, full stop — not because no legitimate reason ever exists, but because the cost of a single bad autonomous call in that set is high enough that it should always be a human's hands on the keyboard, not an agent's, approval or not.

Approval vs. arbitration

Both route through the same gate, but they're answering different questions. Approval is a single decision on one proposed action: approve, reject, or send back with an edit. Arbitration fires when there isn't a single proposal to approve yet — two domain agents recommend conflicting fixes for the same incident, or the policy engine flags a plan the orchestrator scored as low-risk but a rule disagrees with. Rather than building a second escalation path for that case, arbitration is handled as a Tier 3 approval request that includes every competing plan side by side, so the human sees the actual disagreement instead of a single agent's confident-sounding answer.

Phased rollout

"As completely autonomous as possible" is a destination, not a starting configuration — a system this broad earns wider autonomy one phase at a time, against evidence from the phase before it, not by being deployed at full trust on day one.

Phase 0 — Observe read-only, learn ticket patterns, no actions Phase 1 — Low-risk auto Tier 0–1 execute unattended Phase 2 — Approved mutation Tier 2 requires human approval Phase 3 — Broad autonomy tighter guardrails, not fewer of them

Guardrails that make broad autonomy survivable

Self-healing, self-patching, self-maintaining — without a backdoor around the gate

Left unattended, a system this broad decays in predictable ways: a domain agent's process crashes and doesn't come back, a vault lease expires mid-session, a target system's API changes shape and an agent starts silently failing its own health check, patch levels drift out from under the CMDB's idea of what's installed, config drifts from its last-known-good state one small manual change at a time. "Self-healing" means the architecture notices and repairs those conditions itself instead of waiting for a human to notice a service desk backlog and go looking for the cause.

That's the job of the tenth agent, Platform Ops. It listens for a heartbeat from every domain agent and the orchestrator, watches patch/compliance status across managed endpoints, and continuously diffs live config against the CMDB's record of what should be there. Critically, it is bound by the exact same tier system as every other agent in this design — "self-healing" is a scope of what it watches, not a grant of authority the other nine agents don't have. Nothing below adds a second door around the human approval/arbitration gate or the deny-list.

Condition Platform Ops detectsSelf-healing responseTier
A domain agent's process has stopped sending heartbeatsRestart the process; if it fails to come back healthy twice, page a human instead of retrying indefinitelyTier 0
A vault-issued credential lease expired mid-taskRe-check-out a fresh scoped lease and resume the interrupted stepTier 1
Missing OS/security patches on a managed endpointIdentify and stage the patch in a canary group; never auto-installed — enters the normal request lifecycle as a ChangeTier 2
Live config has drifted from the CMDB's recorded stateDetect and report the drift continuously (read-only); reconciling it is a normal action scored by whatever system drifted, through that system's own domain agentTier 0 detect / varies to reconcile
A repeat pattern of near-identical incidents (a "flapping" condition)Suppress duplicate ticket noise and link them to one root-cause record; does not change anything on a managed systemTier 0
An APC UPS reports on-battery with runtime under its configured thresholdTrigger the pre-defined graceful shutdown sequence (VMware agent evacuates/suspends VMs, then Windows/Linux Server agents stop services and power down hosts in dependency order) before the battery is exhausted, not a hard outageTier 1
A scheduled backup job failed, or a restore-test verification did not passRe-queue the job once; on a second failure open an Incident and page — Platform Ops never deletes a backup or shortens its retention to "fix" a job (deny-list)Tier 1
A CrowdStrike Falcon sensor has stopped reporting on a managed hostFlag the detection gap read-only immediately; re-deploy the sensor through the host's own domain agent on the normal Change path, not by a direct push from Platform OpsTier 0 detect / Tier 2 remediate
A CyberArk credential lease failed to revoke after its task completedForce-revoke the lease and alert — a lease that outlives its task is handled as a security event, not routine cleanupTier 1
A storage volume (NetApp/Dell) crosses its capacity thresholdAlert and open a capacity Incident; expansion is a normal Tier 2 Change through the owning agent, never an auto-growTier 0 detect
The orchestrator or Platform Ops itself needs a framework-level updateStaged in a non-production instance first, promoted only through the same Change/approval path as any other Tier 3 action — a system is never allowed to approve its own upgradeTier 3

The pattern to notice: everything that's genuinely reversible and low-blast-radius (restarting a crashed process, refreshing a credential, flagging drift) heals itself immediately, unattended — that's what makes the framework survivable to run without someone watching it every hour. Everything that touches a production system's actual state (installing a patch, reconciling drift, upgrading the framework's own code) still walks through the identical ticket → plan → tier → (approval) → execute → verify lifecycle as a request a person filed by hand. Self-patching means Platform Ops finds and proposes the patch; it does not mean the patch installs itself without going through the same door everything else does.

Walkthrough: an alert firing, end to end

The tables above describe the pieces; this traces two real triggers through all of them, because "automatic" means different things at Tier 0 and Tier 2, and the difference is the whole point of the design. Both start the same way — a Zabbix trigger, not a person — and both end with a fully written audit trail. Where they diverge is whether a human clicks anything in between.

A — disk fills up on a managed Linux host (Tier 0, zero human touch):

  1. Zabbix's own trigger fires on the host (disk use crosses its configured threshold) — this is Zabbix doing what it already does today, unmodified.
  2. The Zabbix → ServiceNow webhook (built in Phase 0, see the deployment guide below) opens an Incident automatically, tagged with the host and the metric that tripped. No person has filed anything yet.
  3. The orchestrator classifies the incident against known trigger signatures, routes it to the Linux Server agent, and scores it Tier 0: single host, fully reversible, matches a pre-approved action class (log rotation / temp cleanup).
  4. The Linux Server agent runs that cleanup playbook directly — no dry-run, no approval record, because Tier 0 by definition doesn't require one — and confirms disk usage is back under threshold.
  5. Platform Ops's flapping check means if the same host trips again within the next few minutes, it links to the same root-cause record instead of opening a second ticket.
  6. The trigger payload, the plan, the action taken, and the verification result are appended to the immutable audit log; the Incident auto-closes with that record attached.

That's the entire loop for anything Tier 0: detect, act, verify, log, close. Nobody was paged, and nobody needed to be — the guardrail isn't a human in this loop, it's that only genuinely reversible, single-target actions are allowed to run unattended in the first place.

B — a missing security patch, the same host class (Tier 2, one approval click, not manual patching):

  1. A nightly Tier 0 compliance scan (Platform Ops, read-only) diffs installed package versions against the patch baseline in the CMDB and finds a host group missing a critical CVE fix — no ticket needed to trigger this, it runs on a schedule.
  2. The Linux Server agent stages the patch in a canary subset of that group and drafts a Change record with the dry-run output attached: the actual package-manager simulation diff, not a text summary of intent.
  3. The orchestrator opens a native ServiceNow Approval on that Change and waits — this is the one deliberate step in the whole sequence, and it's a single approve/reject/edit decision against a real diff, not a request to go patch a server by hand.
  4. Once approved, the agent installs on the canary host first, confirms it comes back healthy, then rolls to the rest of the group inside the approved maintenance window, verifying after each batch rather than all at once.
  5. Before/after package versions, the approval record, and every verification result are written to the audit log; the Change closes itself against that evidence.

"Automatically administer the infrastructure" is everything except that one approval: the scan, the staging, the canary rollout, the batch verification, and the record-keeping all run unattended. What doesn't run unattended is the actual decision to change a production system's state — which is exactly the trade-off "Why a human still approves everything," above, argues for, just made concrete with a real patch instead of an abstract policy.

Deployment guide — building this, phase by phase

This walks the same four phases as the rollout timeline above, but as concrete build steps rather than a description of the end state. Each phase only starts once the previous one has run clean for a defined burn-in window — the point of a phased build isn't paperwork, it's that every new capability earns its trust against real evidence instead of a spec review.

Before writing any agent code, have in hand: a ServiceNow tenant with API access and a dedicated service account; a secrets vault (HashiCorp Vault or a cloud KMS equivalent) capable of issuing short-lived, per-agent scoped credentials; a datastore for the append-only audit log, provisioned independently of any agent so no agent can be the only place its own history lives; an isolated network segment or jump host for agents to run from, reachable only to the specific systems each domain agent needs — not flat access to everything; and, most important and easiest to skip, sign-off from the actual owning teams (security, network, server, database) on the tier definitions and deny-list before day one. A technically correct tier assignment nobody on the network team agreed to is a fight waiting to happen during a real incident, not before one.

PhaseBuild steps
Phase 0
Observe
1. Stand up the audit log store first, before any agent code runs — every action from day one must be logged, including the ones that don't happen yet.
2. Stand up NetBox as the inventory/IPAM source of record and Zabbix + Grafana as the monitoring stack — Platform Ops's drift and health detection in later phases depend on both already existing, not bolted on after agents go live.
3. Wire the ServiceNow integration: subscribe to new Incident/Request/Change records, scoped read access to the CMDB, and a Zabbix → ServiceNow Incident webhook so monitoring alerts enter the same intake path as a human-filed ticket.
4. Build the orchestrator's intake and classification step, and a first-pass risk-tier scorer.
5. Have the orchestrator draft a plan for every incoming ticket and write it to the audit log — but wire no domain agent execution yet.
6. Run this for a fixed window (four to six weeks is a reasonable start) comparing the orchestrator's draft plans against what human staff actually did, and tune classification and tiering against the gap before anything is allowed to act.
Phase 1
Low-risk automation
7. Build one domain agent end-to-end — Identity/AD password reset is a good first choice: high volume, low blast radius, clean rollback (re-reset). Scoped credential checkout, execution, a verification step, and a defined undo, all before it's allowed to run unattended.
8. Wire the policy engine's auto-execute path for Tier 0–Tier 1 only; anything the scorer can't confidently place in that range hard-stops to a human by default, not the other way around.
9. Add circuit breakers — an actions-per-hour cap, a blast-radius-per-action cap, and a manual kill switch — before turning on any unattended execution, not after the first incident that needed one.
10. Bring the rest of the Tier 1 agents (Windows Server, Linux Server via Ansible playbooks, Voice/CUCM) online one at a time, each running clean through its own burn-in period before the next one starts.
Phase 2
Approved mutation
11. Wire the actual approval mechanism: the orchestrator creates a native ServiceNow Approval record on the ticket and blocks on its state, the same object a human change-advisory reviewer already uses today.
12. Build the mandatory dry-run capability for each remaining Tier 2 agent (Network, Database, Firewall, VMware, Desktop) — for the Network agent specifically, this is where Batfish gets wired in as the actual pre-change verification engine against a modeled topology, not a generic live diff; every other Tier 2 agent gets its own real config diff or "what would change" preview, since that's what the human actually approves, not a text description of intent.
13. Bring these agents online one at a time behind the approval gate, same burn-in discipline as Phase 1, confirming real humans are actually reading the dry-run diff and not rubber-stamping it.
14. Onboard Cisco ISE for the Identity agent's NAC actions (quarantine/un-quarantine) and Nexus Dashboard for DC-fabric-scoped Network actions — same burn-in discipline as every other Tier 2 capability; a new integration doesn't earn a trust shortcut for arriving later.
Phase 3
Broad autonomy
15. Build Platform Ops last, once the rest of the framework has a track record — it needs the most earned trust of any agent here, since it touches the framework's own execution path.
16. Formalize two-person approval for Tier 3 and the arbitration flow (competing plans surfaced side by side, per "Approval vs. arbitration" above).
17. Establish a recurring audit cadence — monthly is a reasonable start — sampling auto-executed Tier 0–Tier 1 actions against the log, the same way a human team audits itself. Automation earns continued trust here; it doesn't get it once at go-live and keep it forever unreviewed.

What this is and isn't

This page is an architecture blueprint, written the way a design doc should be written before anyone touches production: the shape of the system, the trust boundaries, and the checkpoints, worked out on paper first. It is deliberately not a running deployment. This box has no ServiceNow tenant, no Cisco/AD/VMware credentials, and no reachable enterprise network to manage — and even if it did, standing up live write-access to someone's routers, directory, and hypervisors isn't a decision an autonomous agent should make for itself. That's a real person's infrastructure, a real person's job to grant that access deliberately, scoped narrowly, with their own change-control process wrapped around it — which is exactly what the human approval/arbitration gate above is designed to plug into, not replace.

Take it further

More to go with the blueprint above, all free, no signup: