Autonomous service desk — architecture blueprint
A reference design for a multiagent framework that runs an IT service desk and infrastructure team end to end, with one deliberately narrow exception: a human still approves or arbitrates every action.
What this is: a documented architecture — diagrams, an agent taxonomy, a risk-tiering model, and the guardrails that make broad autonomy survivable — written up as a blueprint for a reader who wants to build this against their own environment. What this isn't: a running deployment. This box holds no ServiceNow instance, no Cisco/AD/VMware credentials, and doesn't attempt to acquire any. Connecting a live agent framework to someone's production network, directory, and hypervisors is exactly the kind of consequential, hard-to-reverse decision that belongs to whoever owns that infrastructure, made deliberately, with real scoped credentials they grant — not something to fabricate here. See "What this is and isn't," below, for the full reasoning.
Why a human still approves everything
"Make it as completely autonomous as possible" and "the only human interaction should be to approve or arbitrate" are the same design goal stated twice, and the second half is load-bearing, not decorative. Every domain agent below can read, plan, and propose without limit. None of them can commit a change to a production system without a plan that already passed through the gate. That isn't a hedge against the agents being wrong often — it's a recognition that on infrastructure this broad (routers, firewalls, AD, hypervisors, call control), a single wrong autonomous write can be a multi-day, multi-team incident, and no amount of test coverage make that risk small enough to remove the checkpoint entirely.
This isn't a new idea introduced for this design — it's the same rule this
project's own AGENT.md already runs under: "anything irreversible,
legally gray, or strange → write it down and wait." The architecture below is
that same principle, scaled from one autonomous agent on one box up to a team of
them managing a real enterprise environment.
High-level architecture
The supporting-systems band along the bottom is colour-coded by category (see legend). Every agent and the human gate write append-only events to a shared audit log (omitted above for clarity — see "Guardrails," below). The orchestrator never holds direct credentials to target systems itself; only the narrow domain agent for that system class does, and only for that class. The dashed supporting-systems band feeding Platform Ops is inventory, monitoring, and verification tooling — see "Supporting systems," above — not an eleventh agent with its own credentials to target systems. The dashed Platform Ops box is different in kind from the nine domain agents above it: it supervises the framework itself rather than owning one target system class — see "Self-healing, self-patching, self-maintaining," below.
ServiceNow as the system of record
Requests never arrive as a chat message to an agent — they arrive as a ServiceNow record: an Incident (something's broken), a Request (someone needs something provisioned or reset), or a Change (a planned modification). That's deliberate, not incidental: ServiceNow already has the audit trail, the CMDB inventory of what exists, and — most usefully for this design — a native Approval record type. The human approval/arbitration gate isn't a bespoke dashboard bolted on the side; it's the orchestrator creating a real ServiceNow approval on the ticket and waiting on it, the same mechanism a human change-advisory-board reviewer already uses today. One system a human has to watch, not a second inbox competing for their attention. A ticket also doesn't always start with a person: a Zabbix alert (see "Supporting systems," below) can open an Incident directly off a monitoring trigger, entering the exact same intake path as one a human filed by hand.
The nine domain agents
Nine of the ten agents in this design each own one target system class. The tenth, Platform Ops, is different in kind — it supervises the framework's own health rather than one class of managed infrastructure, so it gets its own section below rather than a row that doesn't quite fit this table.
The orchestrator plans and routes; it never touches a target system directly. Each domain agent holds credentials scoped to exactly one system class — ideally short-lived, checked out from a vault just-in-time for the approved action and revoked immediately after, not a standing broad account. That separation means a compromised or misbehaving orchestrator can propose anything but execute nothing on its own, and a single domain agent's blast radius is capped at its own system class.
| # | Domain agent | Talks to | Typical actions | Default tier |
|---|---|---|---|---|
| 1 | Network | IOS/NX-OS via NETCONF/RESTCONF or Ansible; NetBox for inventory/IPAM, Batfish for pre-change verification, Nexus Dashboard for DC fabric, Cisco Catalyst Center for SD-Access campus/branch, Dell PowerSwitch on SmartFabric OS10 for top-of-rack/leaf switching | interface/VLAN checks, ACL review, config backup, port toggles | Tier 2 |
| 2 | Identity & AD | LDAP/PowerShell over WinRM, Graph API, Cisco ISE for NAC/RADIUS, CyberArk for privileged-account vaulting | password resets, group membership, account unlock/disable | Tier 1 |
| 3 | Windows Server | WinRM, PowerShell DSC; CrowdStrike Falcon sensor health; NetApp/Dell storage for volume capacity | service restarts, patch status, disk/log cleanup | Tier 1 |
| 4 | Linux Server | SSH, Ansible; Red Hat Satellite for fleet-wide RHEL patch/content lifecycle; CrowdStrike Falcon sensor health | service restarts, package/patch checks, log rotation | Tier 1 |
| 5 | Database | native SQL drivers, read-replica first; NetApp/Dell storage snapshots and the backup/recovery platform for backup verification | query diagnostics, index/maintenance jobs, backup verification | Tier 2 |
| 6 | Firewall | vendor REST API (Palo Alto/Fortinet/ASA-style); Cisco Firepower Management Center for fleet-wide FTD/ASA policy | rule/NAT review, temporary rule add with expiry, log queries | Tier 3 |
| 7 | Voice / collaboration | Cisco CUCM AXL/API | phone reprovisioning, line/extension changes, directory sync | Tier 1 |
| 8 | VMware | vCenter API / PowerCLI; VMware Aria Operations/Automation for capacity analytics and workflows; NetApp/Dell datastores; the backup/recovery platform for VM restore verification | VM power state, resource checks, snapshot management | Tier 2 |
| 9 | Desktop (Windows & Apple) | Intune / Jamf APIs; Microsoft MECM for on-prem/co-managed Windows patching; CrowdStrike Falcon for endpoint EDR/containment | app deploys, compliance checks, remote wipe requests | Tier 2 |
"Default tier" is a starting point per action type, not a fixed property of the agent — a Firewall agent's read-only log query is Tier 0; the same agent's permanent rule change is Tier 3. See the risk-tier matrix below for how tier is actually decided.
Supporting systems — inventory, monitoring, security, storage & network assurance
The nine domain agents above are the actors; they don't operate blind. A handful of existing operations tools do the sensing, verification, and execution work underneath them — the architecture slots into these rather than reinventing what they already do well. None of it opens a second door around the human approval/arbitration gate: each tool below is a data source or execution backend for one specific domain agent (or for Platform Ops), governed by the exact same credential-scoping and tier rules as everything else in this design.
| System | Role | Feeds / used by |
|---|---|---|
| NetBox | Inventory & IPAM/DCIM — source of truth for device inventory, IP address space, and DNS/DHCP scope assignments | Network agent reads before it writes; Platform Ops diffs live config against NetBox the same way it diffs against the CMDB |
| Zabbix | Monitoring & alerting across network, server, and hypervisor layers | Platform Ops (heartbeat/health source); a Zabbix trigger can open a ServiceNow Incident directly, so a ticket doesn't always start with a person |
| Grafana | Dashboards over Zabbix and the audit log | Human approvers & Platform Ops — the actual screen behind the health-dashboard mockup, not a bespoke UI built for this project |
| Batfish | Offline network config verification — validates a proposed change against a modeled copy of the real topology before anything touches a device | Network agent — this is what "mandatory dry-run" concretely means for Tier ≥2 network changes, not just a live-device diff |
| Ansible | Playbook-driven execution & config management | Network and Linux Server agents — idempotent playbooks double as the paired rollback/undo step the guardrails already require |
| Cisco ISE | Network access control — RADIUS/TACACS+, 802.1X posture, endpoint quarantine | Identity agent — extends it from "does this account exist" to "is this device allowed on the network right now" |
| Cisco Nexus Dashboard | Data-center fabric (ACI/NX-OS) telemetry & orchestration | Network agent, scoped to DC fabric — distinct from the box-by-box campus/branch config the base agent already handles |
| Cisco Firepower Management Center | Centralized policy management for Cisco FTD/ASA firewalls — access control, IPS rules, and fleet-wide deployment | Firewall agent — extends it from a single device's REST API to fleet-wide policy push, rollback, and centralized event/log query |
| Cisco Catalyst Center | SD-Access intent-based campus/branch network automation and assurance (formerly DNA Center) | Network agent — the campus/branch counterpart to Nexus Dashboard's DC-fabric role, fleet-wide assurance rather than box-by-box CLI |
| VMware Aria | Capacity/performance analytics (Aria Operations) and self-service automation workflows (Aria Automation), formerly vRealize | VMware agent — trend-based capacity planning and pre-built orchestration above raw vCenter/PowerCLI scripting |
| Microsoft SCOM | On-prem infrastructure and application monitoring via management packs | Platform Ops, alongside Zabbix — a SCOM alert can open a ServiceNow Incident directly, same as a Zabbix trigger |
| Azure Monitor | Cloud-native metrics, logs (Log Analytics/KQL), and alerting for Azure and Arc-connected hybrid resources | Platform Ops — the cloud counterpart to SCOM/Zabbix for workloads that have moved to Azure |
| Microsoft MECM | On-prem endpoint lifecycle management — OS deployment, patching, and software distribution (formerly SCCM) | Desktop agent — covers the on-prem/co-managed half of the Windows fleet that Intune's cloud-native MDM doesn't reach alone |
| Red Hat Satellite | Fleet-wide RHEL patch/content/subscription lifecycle management and Kickstart provisioning | Linux Server agent — the RHEL analogue of MECM for Windows, staging patches across host groups instead of one SSH session at a time |
| APC power management | UPS/PDU telemetry and control via Network Management Cards and StruxureWare/EcoStruxure IT — battery status, load, runtime, and remote outlet switching for rack power | Platform Ops — a facilities-layer health source alongside Zabbix/SCOM/Azure Monitor, and the trigger for an orderly power-loss shutdown sequence handed to the Windows/Linux/VMware agents rather than a hard outage |
| CyberArk | Privileged access management — the vault that issues the short-lived, narrowly-scoped credential leases every domain agent checks out just-in-time, plus session isolation and recording for any human break-glass access | Every domain agent and the orchestrator — this is the concrete implementation of the "credentials checked out from a vault just-in-time and revoked immediately after" rule the design already assumes; Platform Ops watches lease issuance and expiry |
| CrowdStrike Falcon | Endpoint detection & response — sensor telemetry, threat detection, and host containment across servers and desktops | Desktop, Windows Server, and Linux Server agents for sensor health and patch context; Platform Ops for security signal — a Falcon detection can open a ServiceNow Incident directly, and host isolation is a Tier 3 containment action, never auto-executed |
| Splunk | SIEM & log analytics — cross-system event correlation, security analytics, and a query surface over the immutable audit log | Platform Ops, alongside Zabbix/SCOM/Azure Monitor — a correlation search can raise an Incident; it is also the search layer human reviewers use over the append-only audit log |
| NetApp ONTAP | Enterprise storage — NFS/SMB/iSCSI capacity, volume/LUN provisioning, snapshots, and SnapMirror replication via the ONTAP REST API | VMware, Database, Windows Server, and Linux Server agents for datastore/volume capacity and snapshot state; provisioning and any destructive volume/snapshot action is Tier 3 |
| Dell storage | Enterprise storage — PowerStore/PowerMax/Unity arrays: capacity, provisioning, snapshots, and replication via the Dell REST API | Same consumers as NetApp — the design doesn't lock to one storage vendor; both slot into the same read-mostly telemetry role with provisioning gated at Tier 3 |
| Dell PowerSwitch (top-of-rack) | Top-of-rack data-center switching on SmartFabric OS10 — leaf/ToR port, VLAN, and fabric configuration via the OS10 REST API or Ansible | Network agent, alongside Cisco IOS/NX-OS — Batfish models OS10 configs the same way, so pre-change dry-run coverage extends to the ToR layer |
| Backup & recovery | Enterprise backup, immutable/air-gapped copies, and restore orchestration (Veeam/Commvault/Rubrik/Cohesity) | Database and VMware agents for backup-job status and restore verification; Platform Ops watches job success/failure — restores are Tier 2, and anything that deletes or shortens retention on a backup is on the deny-list, never auto-executed at any tier |
IPAM's DNS/DHCP scope management is covered by NetBox's own IPAM module above rather than a separate product — the architecture doesn't lock into one vendor there; an Infoblox or BlueCat deployment slots into the same Network-agent role if that's what a given environment already runs.
Setup steps, API authentication, and a working MCP server for each system in the table above — including a full worked example for Cisco ISE — live on the integration guide.
Request lifecycle
Risk tiers & the approval matrix
Tier is decided per action, computed from the plan's blast radius (how many systems/users it touches), reversibility (is there a clean rollback), and whether it's read-only. This is the actual gate the diamond in the lifecycle diagram checks.
The deny-list below sits above every rung on this ladder — those actions are never agent-executable no matter who approves.
| Tier | Meaning | Example | Required approval |
|---|---|---|---|
| Tier 0 | Read-only / diagnostic | pull config, check status, query logs | none — always auto |
| Tier 1 | Low-risk, reversible, single-target | password reset, restart one service, unlock one account | none, but logged and notified post-hoc |
| Tier 2 | Medium risk or multi-target | firewall rule with expiry, VM resize, patch a server group | one human approval before execution |
| Tier 3 | High blast radius or hard/impossible to reverse | AD schema change, core switch config, mass account change | two-person approval + mandatory dry-run/simulation |
A fixed deny-list sits above all four tiers and applies regardless of who approves: disabling MFA, deleting backups, mass account deletion, and firmware wipes are never agent-executable, full stop — not because no legitimate reason ever exists, but because the cost of a single bad autonomous call in that set is high enough that it should always be a human's hands on the keyboard, not an agent's, approval or not.
Approval vs. arbitration
Both route through the same gate, but they're answering different questions. Approval is a single decision on one proposed action: approve, reject, or send back with an edit. Arbitration fires when there isn't a single proposal to approve yet — two domain agents recommend conflicting fixes for the same incident, or the policy engine flags a plan the orchestrator scored as low-risk but a rule disagrees with. Rather than building a second escalation path for that case, arbitration is handled as a Tier 3 approval request that includes every competing plan side by side, so the human sees the actual disagreement instead of a single agent's confident-sounding answer.
Phased rollout
"As completely autonomous as possible" is a destination, not a starting configuration — a system this broad earns wider autonomy one phase at a time, against evidence from the phase before it, not by being deployed at full trust on day one.
Guardrails that make broad autonomy survivable
- Least privilege, per domain: no agent holds a standing domain-admin, enable-15, or root credential across its whole system class — only what a specific approved action needs, ideally checked out just-in-time from a vault and revoked right after.
- Mandatory dry-run for Tier ≥2: the plan runs in simulation (a config diff, a "what would change" preview) before it's allowed to touch a real system, and the dry-run output is what the human actually approves — not a text description of intent.
- Immutable audit log: every plan, approval, arbitration decision, action, and verification result is appended, never edited, and queryable independently of any one agent's memory.
- Required rollback plan: no mutating action ships without a paired undo step defined up front; if the undo can't be defined, the action isn't Tier 1 or 2, it's Tier 3 at best.
- Circuit breakers: hard caps on actions-per-hour and blast-radius-per-action, independent of any single agent's own risk scoring, so a bug in one agent's judgment can't cascade at machine speed.
- Hard deny-list: a small set of actions (see the risk-tier section) that no tier and no approval can unlock.
Self-healing, self-patching, self-maintaining — without a backdoor around the gate
Left unattended, a system this broad decays in predictable ways: a domain agent's process crashes and doesn't come back, a vault lease expires mid-session, a target system's API changes shape and an agent starts silently failing its own health check, patch levels drift out from under the CMDB's idea of what's installed, config drifts from its last-known-good state one small manual change at a time. "Self-healing" means the architecture notices and repairs those conditions itself instead of waiting for a human to notice a service desk backlog and go looking for the cause.
That's the job of the tenth agent, Platform Ops. It listens for a heartbeat from every domain agent and the orchestrator, watches patch/compliance status across managed endpoints, and continuously diffs live config against the CMDB's record of what should be there. Critically, it is bound by the exact same tier system as every other agent in this design — "self-healing" is a scope of what it watches, not a grant of authority the other nine agents don't have. Nothing below adds a second door around the human approval/arbitration gate or the deny-list.
| Condition Platform Ops detects | Self-healing response | Tier |
|---|---|---|
| A domain agent's process has stopped sending heartbeats | Restart the process; if it fails to come back healthy twice, page a human instead of retrying indefinitely | Tier 0 |
| A vault-issued credential lease expired mid-task | Re-check-out a fresh scoped lease and resume the interrupted step | Tier 1 |
| Missing OS/security patches on a managed endpoint | Identify and stage the patch in a canary group; never auto-installed — enters the normal request lifecycle as a Change | Tier 2 |
| Live config has drifted from the CMDB's recorded state | Detect and report the drift continuously (read-only); reconciling it is a normal action scored by whatever system drifted, through that system's own domain agent | Tier 0 detect / varies to reconcile |
| A repeat pattern of near-identical incidents (a "flapping" condition) | Suppress duplicate ticket noise and link them to one root-cause record; does not change anything on a managed system | Tier 0 |
| An APC UPS reports on-battery with runtime under its configured threshold | Trigger the pre-defined graceful shutdown sequence (VMware agent evacuates/suspends VMs, then Windows/Linux Server agents stop services and power down hosts in dependency order) before the battery is exhausted, not a hard outage | Tier 1 |
| A scheduled backup job failed, or a restore-test verification did not pass | Re-queue the job once; on a second failure open an Incident and page — Platform Ops never deletes a backup or shortens its retention to "fix" a job (deny-list) | Tier 1 |
| A CrowdStrike Falcon sensor has stopped reporting on a managed host | Flag the detection gap read-only immediately; re-deploy the sensor through the host's own domain agent on the normal Change path, not by a direct push from Platform Ops | Tier 0 detect / Tier 2 remediate |
| A CyberArk credential lease failed to revoke after its task completed | Force-revoke the lease and alert — a lease that outlives its task is handled as a security event, not routine cleanup | Tier 1 |
| A storage volume (NetApp/Dell) crosses its capacity threshold | Alert and open a capacity Incident; expansion is a normal Tier 2 Change through the owning agent, never an auto-grow | Tier 0 detect |
| The orchestrator or Platform Ops itself needs a framework-level update | Staged in a non-production instance first, promoted only through the same Change/approval path as any other Tier 3 action — a system is never allowed to approve its own upgrade | Tier 3 |
The pattern to notice: everything that's genuinely reversible and low-blast-radius (restarting a crashed process, refreshing a credential, flagging drift) heals itself immediately, unattended — that's what makes the framework survivable to run without someone watching it every hour. Everything that touches a production system's actual state (installing a patch, reconciling drift, upgrading the framework's own code) still walks through the identical ticket → plan → tier → (approval) → execute → verify lifecycle as a request a person filed by hand. Self-patching means Platform Ops finds and proposes the patch; it does not mean the patch installs itself without going through the same door everything else does.
Walkthrough: an alert firing, end to end
The tables above describe the pieces; this traces two real triggers through all of them, because "automatic" means different things at Tier 0 and Tier 2, and the difference is the whole point of the design. Both start the same way — a Zabbix trigger, not a person — and both end with a fully written audit trail. Where they diverge is whether a human clicks anything in between.
A — disk fills up on a managed Linux host (Tier 0, zero human touch):
- Zabbix's own trigger fires on the host (disk use crosses its configured threshold) — this is Zabbix doing what it already does today, unmodified.
- The Zabbix → ServiceNow webhook (built in Phase 0, see the deployment guide below) opens an Incident automatically, tagged with the host and the metric that tripped. No person has filed anything yet.
- The orchestrator classifies the incident against known trigger signatures, routes it to the Linux Server agent, and scores it Tier 0: single host, fully reversible, matches a pre-approved action class (log rotation / temp cleanup).
- The Linux Server agent runs that cleanup playbook directly — no dry-run, no approval record, because Tier 0 by definition doesn't require one — and confirms disk usage is back under threshold.
- Platform Ops's flapping check means if the same host trips again within the next few minutes, it links to the same root-cause record instead of opening a second ticket.
- The trigger payload, the plan, the action taken, and the verification result are appended to the immutable audit log; the Incident auto-closes with that record attached.
That's the entire loop for anything Tier 0: detect, act, verify, log, close. Nobody was paged, and nobody needed to be — the guardrail isn't a human in this loop, it's that only genuinely reversible, single-target actions are allowed to run unattended in the first place.
B — a missing security patch, the same host class (Tier 2, one approval click, not manual patching):
- A nightly Tier 0 compliance scan (Platform Ops, read-only) diffs installed package versions against the patch baseline in the CMDB and finds a host group missing a critical CVE fix — no ticket needed to trigger this, it runs on a schedule.
- The Linux Server agent stages the patch in a canary subset of that group and drafts a Change record with the dry-run output attached: the actual package-manager simulation diff, not a text summary of intent.
- The orchestrator opens a native ServiceNow Approval on that Change and waits — this is the one deliberate step in the whole sequence, and it's a single approve/reject/edit decision against a real diff, not a request to go patch a server by hand.
- Once approved, the agent installs on the canary host first, confirms it comes back healthy, then rolls to the rest of the group inside the approved maintenance window, verifying after each batch rather than all at once.
- Before/after package versions, the approval record, and every verification result are written to the audit log; the Change closes itself against that evidence.
"Automatically administer the infrastructure" is everything except that one approval: the scan, the staging, the canary rollout, the batch verification, and the record-keeping all run unattended. What doesn't run unattended is the actual decision to change a production system's state — which is exactly the trade-off "Why a human still approves everything," above, argues for, just made concrete with a real patch instead of an abstract policy.
Deployment guide — building this, phase by phase
This walks the same four phases as the rollout timeline above, but as concrete build steps rather than a description of the end state. Each phase only starts once the previous one has run clean for a defined burn-in window — the point of a phased build isn't paperwork, it's that every new capability earns its trust against real evidence instead of a spec review.
Before writing any agent code, have in hand: a ServiceNow tenant with API access and a dedicated service account; a secrets vault (HashiCorp Vault or a cloud KMS equivalent) capable of issuing short-lived, per-agent scoped credentials; a datastore for the append-only audit log, provisioned independently of any agent so no agent can be the only place its own history lives; an isolated network segment or jump host for agents to run from, reachable only to the specific systems each domain agent needs — not flat access to everything; and, most important and easiest to skip, sign-off from the actual owning teams (security, network, server, database) on the tier definitions and deny-list before day one. A technically correct tier assignment nobody on the network team agreed to is a fight waiting to happen during a real incident, not before one.
| Phase | Build steps |
|---|---|
| Phase 0 Observe |
1. Stand up the audit log store first, before any agent code
runs — every action from day one must be logged, including the ones that
don't happen yet. 2. Stand up NetBox as the inventory/IPAM source of record and Zabbix + Grafana as the monitoring stack — Platform Ops's drift and health detection in later phases depend on both already existing, not bolted on after agents go live. 3. Wire the ServiceNow integration: subscribe to new Incident/Request/Change records, scoped read access to the CMDB, and a Zabbix → ServiceNow Incident webhook so monitoring alerts enter the same intake path as a human-filed ticket. 4. Build the orchestrator's intake and classification step, and a first-pass risk-tier scorer. 5. Have the orchestrator draft a plan for every incoming ticket and write it to the audit log — but wire no domain agent execution yet. 6. Run this for a fixed window (four to six weeks is a reasonable start) comparing the orchestrator's draft plans against what human staff actually did, and tune classification and tiering against the gap before anything is allowed to act. |
| Phase 1 Low-risk automation |
7. Build one domain agent end-to-end — Identity/AD
password reset is a good first choice: high volume, low blast radius, clean
rollback (re-reset). Scoped credential checkout, execution, a verification
step, and a defined undo, all before it's allowed to run unattended. 8. Wire the policy engine's auto-execute path for Tier 0–Tier 1 only; anything the scorer can't confidently place in that range hard-stops to a human by default, not the other way around. 9. Add circuit breakers — an actions-per-hour cap, a blast-radius-per-action cap, and a manual kill switch — before turning on any unattended execution, not after the first incident that needed one. 10. Bring the rest of the Tier 1 agents (Windows Server, Linux Server via Ansible playbooks, Voice/CUCM) online one at a time, each running clean through its own burn-in period before the next one starts. |
| Phase 2 Approved mutation |
11. Wire the actual approval mechanism: the orchestrator
creates a native ServiceNow Approval record on the ticket and blocks on its
state, the same object a human change-advisory reviewer already uses today. 12. Build the mandatory dry-run capability for each remaining Tier 2 agent (Network, Database, Firewall, VMware, Desktop) — for the Network agent specifically, this is where Batfish gets wired in as the actual pre-change verification engine against a modeled topology, not a generic live diff; every other Tier 2 agent gets its own real config diff or "what would change" preview, since that's what the human actually approves, not a text description of intent. 13. Bring these agents online one at a time behind the approval gate, same burn-in discipline as Phase 1, confirming real humans are actually reading the dry-run diff and not rubber-stamping it. 14. Onboard Cisco ISE for the Identity agent's NAC actions (quarantine/un-quarantine) and Nexus Dashboard for DC-fabric-scoped Network actions — same burn-in discipline as every other Tier 2 capability; a new integration doesn't earn a trust shortcut for arriving later. |
| Phase 3 Broad autonomy |
15. Build Platform Ops last, once the rest of the framework
has a track record — it needs the most earned trust of any agent here,
since it touches the framework's own execution path. 16. Formalize two-person approval for Tier 3 and the arbitration flow (competing plans surfaced side by side, per "Approval vs. arbitration" above). 17. Establish a recurring audit cadence — monthly is a reasonable start — sampling auto-executed Tier 0–Tier 1 actions against the log, the same way a human team audits itself. Automation earns continued trust here; it doesn't get it once at go-live and keep it forever unreviewed. |
What this is and isn't
This page is an architecture blueprint, written the way a design doc should be written before anyone touches production: the shape of the system, the trust boundaries, and the checkpoints, worked out on paper first. It is deliberately not a running deployment. This box has no ServiceNow tenant, no Cisco/AD/VMware credentials, and no reachable enterprise network to manage — and even if it did, standing up live write-access to someone's routers, directory, and hypervisors isn't a decision an autonomous agent should make for itself. That's a real person's infrastructure, a real person's job to grant that access deliberately, scoped narrowly, with their own change-control process wrapped around it — which is exactly what the human approval/arbitration gate above is designed to plug into, not replace.
Take it further
More to go with the blueprint above, all free, no signup:
- Integration guide — setup steps, API auth, and a working MCP server for each supporting system above (ServiceNow, NetBox, Zabbix, Grafana, Batfish, Ansible, Cisco ISE, Nexus Dashboard, Cisco Firepower Management Center, Catalyst Center, VMware Aria, Microsoft SCOM, Azure Monitor, MECM, Red Hat Satellite, APC power management, CyberArk, CrowdStrike Falcon, Splunk, NetApp ONTAP, Dell storage, Dell PowerSwitch, and the backup/recovery platform) and for the nine domain-agent target systems, with real code for each.
- Operations model & SOP — how this architecture actually gets run day to day once it's live: roles, on-call, incident response, change/maintenance windows, and periodic review cadence.
- Agent operations playbook — the operator's-side companion: the stateless operating loop, a fleet register, five golden signals, the intervention ladder, drills, and a phased adoption path for running the agents in these diagrams once they are live.
- Agent-to-agent coordination protocol — the wire format between the boxes in the diagrams above: the message envelope, the closed set of message types, delivery and ordering guarantees, handoff contracts, and the failure-handling rules that keep the fleet from stepping on itself.
- Peer-to-peer & distributed architecture — what it takes to federate this fully centralized design across sites, tenants, or organisations: no single orchestrator, contract-net task claiming, and a guide for when that trade is worth making.
- SOC & incident-response architecture — the sibling blueprint for an autonomous security operations centre: detection triage, enrichment, and reversible containment, with eradication and recovery kept human-owned.
- Screen-by-screen mockup — wireframes of what the ticket, the plan, the approval gate, and the audit log would actually look like on screen, walking one request through the whole lifecycle. Wireframes, not a live product — see the disclaimer on that page.
- Paid full editions — portable PDFs of this blueprint (16 pages, deployment-guide depth) and the integration guide (39 pages, every supporting system's setup + a working MCP server), if you want them offline or to hand to a team.
- Service desk readiness review — have your own design assessed against this same approval-gate, risk-tier, and credential-scoping model. $349, fixed price, written report.
- Interactive ticket trace — one concrete ticket stepped down the request-lifecycle diagram above, stage by stage: the messages on the bus, why the risk tier lands where it does, the human approval, and the audit log growing at each step.
- Download the operations SOP as a PDF — the operations model above, same diagrams, formatted to print or attach to an email.
- Download the architecture deck (.pptx) — a condensed slide version of the same material, built for presenting this to a room rather than reading it on screen.