Enterprise AI Agent Infrastructure Documentation
Version 1.1 • October 2026 • Published by ZTICOM Tech Ltd ([email protected])
0. Purpose of the Document
This document serves as the absolute source-of-truth specification for WorkAgent OS. WorkAgent OS enables AI agents to safely collaborate alongside human employees with live access to corporate systems, verifiable action execution, policy-governed tool access, and human-in-the-loop approvals.
1. Product Definition & Positioning
2. 8-Layer Agent Execution Model & Identity Chain
Tenant → User Identity → Agent Identity → Policy Scope → Scoped Tool → Action
An agent never operates as an unconstrained tenant-wide service account; all tool calls inherit the delegated identity and permissions of the initiating user.
4 & 5. Enterprise Agents Fleet & MVP Scope
The fleet features eight specialized agent types: AI Chief of Staff (MVP), Sales Agent, Customer Success Agent, Project Agent, Meeting Agent, Research Agent, Developer Agent, and Finance/Operations Agent.
Connectors: Gmail, Google Calendar, Slack, Google Drive, and Jira/Linear. Core tasks: Daily morning prioritization, meeting preparation, action item extraction, detecting overdue deliverables, and weekly executive digests.
6. Master System Architecture
8. Agent Runtime Algorithm
9. The 9-Step Context Engine Pipeline
Raw tenant data is never intended to be dumped directly into the model context. The Context Engine resolves nine context stages (scope and policy resolution — not ranking) prior to model invocation. Retrieval and ranking run as a separate pipeline: structured fetch → semantic search → relevance scoring → source authority → recency → deduplication → token budget → citation tagging.
Resolves authenticated actor identity, session tokens, and delegated authority delegation chain (Tenant → User → Agent).
Injects strict tenant boundary metadata, subscription tier quotas, data residency rules (e.g. EU-West), and feature toggles.
Evaluates RBAC/ABAC department permissions, data clearance level, and designated approval authorities.
Decomposes raw user prompt into explicit objective, trigger type, latency budget, and execution constraints.
Maintains multi-turn conversational history, past user clarifications, and active thread-level parameters.
Incorporates enterprise SOPs, corporate glossaries, team rosters, and internal operational guidelines.
Resolves whether long-term organizational memory may enter context. Distinct from retrieval: this stage decides admission; fetching and ranking candidates is a separate retrieval pipeline.
Inspects live MCP connector registry, tool availability, health heartbeats, API rate quotas, and JSON schemas.
Evaluates active risk thresholds, PII/financial sensitivity tags, and human-in-the-loop sign-off triggers before model prompt generation.
11 & 12. MCP-Compatible Tool Gateway & Credential Isolation
WorkAgent OS integrates directly with the Model Context Protocol (MCP) standard as a secure gateway layer. The gateway handles discovery, schema enforcement, credential isolation, tenant isolation, and verification.
Third-party tokens reside exclusively in HashiCorp Vault. The model sees tool schemas with parameter definitions—never raw API keys, OAuth access tokens, or database connection strings.
Every tool invocation that mutates state requires a deterministic idempotency key. Network retries and duplicate agent runs cannot produce duplicate writes or external side-effects.
13 & 14. Dynamic Risk Scoring & Policy Tiers
Each variable is normalized strictly between 0.0 and 1.0. Higher factor values consistently indicate higher risk.
- • 0–20: AUTO EXECUTION (Internal calendar reads, semantic document queries) — still subject to policy, tenant, and scope checks
- • 21–50: ABAC / POLICY CHECK (Internal project ticket creation/update)
- • 51–80: HUMAN APPROVAL (External client emails, cross-service writes)
- • 81–100: DENY / ELEVATED APPROVAL (Financial disbursements, bulk customer blasts)
Illustrative scoring model — a low risk score does not guarantee execution; policy and scope checks may still deny the action.
15. Action Integrity & Approval Canonicalization
Before human approval is requested, the action payload is transformed into deterministic canonical JSON (sorted keys, UTF-8 encoded) and hashed using SHA-256 to produce an action_hash. Hashing provides integrity binding — approver authentication is a separate step, and a digital signature field is only added if real signing is adopted.
If any parameter diverges between approval and execution, the runtime kernel immediately halts. Idempotency keys eliminate duplicate execution side-effects.
16. Active Read-Back State Verification
Verification does not merely inspect HTTP status codes. For write operations, the verifier performs an independent GET read-back of the external entity (e.g. querying Jira for issue ENG-4821) and matches actual field values against the expected state before declaring the step completed.
20 & 21. Relational Data Model & ERD
The persistence tier relies on strict multi-tenant relational schemas inside PostgreSQL, protected by Row-Level Security (RLS) bound to session_user_tenant_id.
24. Multi-Agent Hierarchy & Orchestration
WorkAgent OS establishes a clear division of labor: the AI Chief of Staff serves as the primary executive orchestrator and coordinator. It decomposes compound organizational goals and delegates specialized sub-tasks to domain-specific agents (Sales, Project, Developer, Finance) with explicit permission envelopes.
27. Strict Tenant Isolation Across All Tiers
Tenant boundaries are designed to be enforced systematically across the full infrastructure stack:
30. Definition of Done (DoD) & Evaluation Standards
A WorkAgent OS agent release is intended to be marked production-ready only when it passes the complete golden benchmark evaluation suite. These are defined targets — none have been executed against a published deployment:
Target: 100% rejection across adversarial cross-tenant retrieval queries (planned suite).
Target: halt whenever action payload parameters diverge from the approved hash.
Mandatory external state confirmation matching expected schema prior to run completion.
Target: prompt context inspection clean against sensitive patterns (planned suite).
31. Mandatory Engineering Directives
P2. Operational Semantics
SPECROADMAPThe sections below document the target operational semantics of the platform. They are design specifications and roadmap items — none are verified deployed behavior.
Concurrency & Conflict Handling
- •Optimistic concurrency: entities carry version fields; writes fail on stale state rather than silently overwriting.
- •Duplicate suppression: deterministic idempotency keys prevent repeated side-effects from retries or duplicated runs.
- •Resource locks: run-scoped locks in the cache layer prevent concurrent mutation of the same external entity.
- •Conflict handling: on detected drift, the plan re-reads current state and re-plans instead of forcing the stale write.
Cost & Budget Enforcement
Budgets are specified at five scopes. When the projected cost of a planned action exceeds any applicable budget — not only when actual spend crosses it — execution is designed to pause and require approval rather than silently continuing spend:
Model Routing (Illustrative)
The adapter layer is designed to route tasks by complexity and policy to different models. Illustrative routing targets — not a deployment description:
- •Document extraction & summarization → lightweight model
- •Planning & DAG compilation → frontier reasoning model
- •Code analysis & refactoring proposals → code-specialized model
- •High-risk / approval-sensitive steps → most capable model under strict policy
Data Classification Mapping (Proposed)
Proposed mapping of the five classification levels to model, retrieval, storage, log, export, and approval behavior. This mapping is a policy proposal — not enforceable controls today.
| Level | Model | Retrieval | Storage | Log | Export | Approval |
|---|---|---|---|---|---|---|
| PUBLIC | Full context allowed | Unrestricted within tenant | Standard tenant storage | Full trace detail | Allowed | None |
| INTERNAL | Full context allowed | Tenant-scoped | Standard tenant storage | Full trace detail | Tenant-internal only | None |
| CONFIDENTIAL | Scoped context, minimal necessary | Role/department gated | Encrypted, access-scoped | Masked in traces | Policy-gated | Required for external export |
| RESTRICTED | Minimal-context delivery only | Explicit grant required | Encrypted, excluded from long-term memory | Redacted; metadata only | Blocked by default | Elevated approval required |
| HIGHLY RESTRICTED | Never enters model context | Vault-only, no agent retrieval | Dedicated vault / HSM-class | Existence-only events, no payload | Forbidden | Cannot be approved for model exposure |
Disaster Recovery (Roadmap)
Planned DR scope covers region loss, model provider failure, connector failure, and data corruption. RPO/RTO targets are not yet measured and no values are published:
- •Tenant-scoped backups with periodic restore testing
- •Region loss — failover target per deployment configuration
- •Connector failure — degraded-mode behavior and reconciliation per failure class
- •Model provider failure — adapter fail-over to alternate provider
- •Data corruption — point-in-time recovery target
Connector Failure Classes
Seven connector failure classes are specified (schema drift shown as an optional eighth). On timeout the design is to re-read state before retrying — never blindly resend a write. Note: there is no universal exactly-once guarantee across providers (e.g. Gmail); idempotency reduces but does not eliminate duplicate risk.
- 1.Auth failure — refresh/rebind credentials once; escalate on repeat
- 2.Rate limit (429) — bounded backoff with jitter
- 3.Timeout — re-read external state first; avoid blind resend
- 4.Partial success — reconcile via read-back and compensating action
- 5.Permission denied — surface to user; never retry around policy
- 6.External conflict — re-read state, re-plan, don't force stale write
- 7.Provider outage — pause queue, fail over per deployment policy
- 8.Schema drift (optional 8th) — validate response; fail safe, don't guess fields
33. Ultimate Vision & Master Equation
"AI agents that can act — with context, permissions, verification and accountability."