WorkAgent OS • Master Specification Document (v1.1)

Enterprise AI Agent Infrastructure Documentation

Version 1.1 • October 2026 • Published by ZTICOM Tech Ltd ([email protected])

Specification Versionv1.1 Master Release
Publication DateOctober 2026
Engineering OwnerZTICOM Tech Ltd
Contact / Feedback[email protected]

0. Purpose of the Document

This document serves as the absolute source-of-truth specification for WorkAgent OS. WorkAgent OS enables AI agents to safely collaborate alongside human employees with live access to corporate systems, verifiable action execution, policy-governed tool access, and human-in-the-loop approvals.

1. Product Definition & Positioning

WorkAgent OS is not a chatbot. It is a multi-tenant execution platform where human knowledge workers and AI agents collaborate through live company context, policy-controlled tools, verifiable execution, and human-in-the-loop approvals.
Core Abstraction: The model is not the product. Models are interchangeable reasoning modules inside the Agent Runtime.

2. 8-Layer Agent Execution Model & Identity Chain

01 Identity & Context → 02 Agent Runtime → 03 Planning DAG → 04 Tool Selection → 05 Risk & Policy → 06 Execution → 07 Verification → 08 Audit
Identity Delegation Chain:

Tenant → User Identity → Agent Identity → Policy Scope → Scoped Tool → Action

An agent never operates as an unconstrained tenant-wide service account; all tool calls inherit the delegated identity and permissions of the initiating user.

4 & 5. Enterprise Agents Fleet & MVP Scope

The fleet features eight specialized agent types: AI Chief of Staff (MVP), Sales Agent, Customer Success Agent, Project Agent, Meeting Agent, Research Agent, Developer Agent, and Finance/Operations Agent.

MVP Scope: AI Chief of Staff

Connectors: Gmail, Google Calendar, Slack, Google Drive, and Jira/Linear. Core tasks: Daily morning prioritization, meeting preparation, action item extraction, detecting overdue deliverables, and weekly executive digests.

6. Master System Architecture

┌─────────────────────────────────────────────────────────────┐ │ WORKAGENT OS │ ├─────────────────────────────────────────────────────────────┤ │ Presentation: Employee Workspace • Admin / IT Console │ ├─────────────────────────────────────────────────────────────┤ │ Gateway: Identity & Auth • Tenant Resolver • Rate Limit │ ├─────────────────────────────────────────────────────────────┤ │ AGENT RUNTIME │ │ Context Builder • Planning DAG • Interchangeable Models │ │ Policy Engine • Human Approval • Verification / Recovery │ ├─────────────────────────────────────────────────────────────┤ │ MCP-COMPATIBLE TOOL GATEWAY │ │ Gmail • Calendar • Slack • Drive • Jira • GitHub • CRM │ ├─────────────────────────────────────────────────────────────┤ │ PERSISTENCE LAYER │ │ PostgreSQL (RLS) • pgvector • Redis Cache • S3 Object Store │ ├─────────────────────────────────────────────────────────────┤ │ OBSERVABILITY & EVALUATION │ │ OpenTelemetry Traces • Token Accounting • Golden Eval Suites │ └─────────────────────────────────────────────────────────────┘

8. Agent Runtime Algorithm

function run_agent(request, user, tenant): # 1. Identity & Tenant Authentication authorize_user(user, tenant) # 2. Context Retrieval & Filtering context = build_context(user_profile, history, memory, sources, request) policy = load_policy(tenant, agent_type) validate_context_access(context, policy) # 3. Planning DAG Compilation task = classify_task(request) plan = model.plan(task, context, available_tools) # 4. Tool Execution Loop for step in plan: tool = resolve_tool(step) if not policy.allows(tool, user, context): return blocked("permission_denied") risk = risk_engine.assess(step, tool, context) if risk.requires_human_approval: approval = request_approval(step, risk) if not approval.granted: audit("approval_denied") continue result = execute_tool(tool, step.parameters) # 5. Independent Read-Back State Verification if not verifier.validate(result, step.expected_outcome): recovery = planner.replan(step, result) if recovery.available: continue return safe_failure() write_trace(step, result) update_short_term_memory(step, result) persist_relevant_long_term_memory() return final_response()

9. The 9-Step Context Engine Pipeline

Raw tenant data is never intended to be dumped directly into the model context. The Context Engine resolves nine context stages (scope and policy resolution — not ranking) prior to model invocation. Retrieval and ranking run as a separate pipeline: structured fetch → semantic search → relevance scoring → source authority → recency → deduplication → token budget → citation tagging.

Step 01Source & Scope
Identity Context

Resolves authenticated actor identity, session tokens, and delegated authority delegation chain (Tenant → User → Agent).

{"actor_id": "usr_991", "delegated_user": "[email protected]", "role": "VP_PRODUCT"}
Step 02Source & Scope
Tenant Context

Injects strict tenant boundary metadata, subscription tier quotas, data residency rules (e.g. EU-West), and feature toggles.

{"tenant_id": "tenant_89f2a", "residency": "EU_WEST_1", "tier": "ENTERPRISE_DEDICATED"}
Step 03Source & Scope
User / Role Context

Evaluates RBAC/ABAC department permissions, data clearance level, and designated approval authorities.

{"clearance": "CONFIDENTIAL", "can_trigger_external_writes": true, "department": "Product"}
Step 04Source & Scope
Task Context

Decomposes raw user prompt into explicit objective, trigger type, latency budget, and execution constraints.

{"objective": "MEETING_PREPARATION", "priority": "HIGH", "target_deadline": "14:00:00Z"}
Step 05State & Memory
Conversation Context

Maintains multi-turn conversational history, past user clarifications, and active thread-level parameters.

{"turn_count": 3, "resolved_entities": ["Acme Corp", "Q4 Proposal"]}
Step 06State & Memory
Organizational Context

Incorporates enterprise SOPs, corporate glossaries, team rosters, and internal operational guidelines.

{"sop_reference": "SOP-EXEC-PREP-04", "escalation_contact": "[email protected]"}
Step 07State & Memory
Memory Context

Resolves whether long-term organizational memory may enter context. Distinct from retrieval: this stage decides admission; fetching and ranking candidates is a separate retrieval pipeline.

{"episodic_hits": 2, "top_relevance": 0.94, "memory_id": "mem_acme_q3_sla"}
Step 08Policy & Safety
Tool / System State

Inspects live MCP connector registry, tool availability, health heartbeats, API rate quotas, and JSON schemas.

{"available_tools": ["mcp_calendar", "mcp_salesforce", "mcp_jira"], "registry_status": "HEALTHY"}
Step 09Policy & Safety
Policy / Risk Context

Evaluates active risk thresholds, PII/financial sensitivity tags, and human-in-the-loop sign-off triggers before model prompt generation.

{"risk_tier": "MEDIUM", "requires_canonical_approval": true, "active_policy": "ABAC-V2"}
Data Flow: Raw Inputs → Context Resolution → Policy Filtering → Relevance Ranking → Context Assembly → Model

11 & 12. MCP-Compatible Tool Gateway & Credential Isolation

WorkAgent OS integrates directly with the Model Context Protocol (MCP) standard as a secure gateway layer. The gateway handles discovery, schema enforcement, credential isolation, tenant isolation, and verification.

Zero Raw Credential Exposure

Third-party tokens reside exclusively in HashiCorp Vault. The model sees tool schemas with parameter definitions—never raw API keys, OAuth access tokens, or database connection strings.

Idempotency Key Enforcement

Every tool invocation that mutates state requires a deterministic idempotency key. Network retries and duplicate agent runs cannot produce duplicate writes or external side-effects.

13 & 14. Dynamic Risk Scoring & Policy Tiers

RiskScore = Impact × Irreversibility × Externality × Sensitivity × Uncertainty × 100

Each variable is normalized strictly between 0.0 and 1.0. Higher factor values consistently indicate higher risk.

  • • 0–20: AUTO EXECUTION (Internal calendar reads, semantic document queries) — still subject to policy, tenant, and scope checks
  • • 21–50: ABAC / POLICY CHECK (Internal project ticket creation/update)
  • • 51–80: HUMAN APPROVAL (External client emails, cross-service writes)
  • • 81–100: DENY / ELEVATED APPROVAL (Financial disbursements, bulk customer blasts)

Illustrative scoring model — a low risk score does not guarantee execution; policy and scope checks may still deny the action.

15. Action Integrity & Approval Canonicalization

Before human approval is requested, the action payload is transformed into deterministic canonical JSON (sorted keys, UTF-8 encoded) and hashed using SHA-256 to produce an action_hash. Hashing provides integrity binding — approver authentication is a separate step, and a digital signature field is only added if real signing is adopted.

Canonical Action Payload → UTF-8 → SHA-256 → action_hash → Human Sign-off → Recompute Hash on Execution → Match → Execute

If any parameter diverges between approval and execution, the runtime kernel immediately halts. Idempotency keys eliminate duplicate execution side-effects.

16. Active Read-Back State Verification

Verification does not merely inspect HTTP status codes. For write operations, the verifier performs an independent GET read-back of the external entity (e.g. querying Jira for issue ENG-4821) and matches actual field values against the expected state before declaring the step completed.

20 & 21. Relational Data Model & ERD

The persistence tier relies on strict multi-tenant relational schemas inside PostgreSQL, protected by Row-Level Security (RLS) bound to session_user_tenant_id.

┌──────────────┐ ┌──────────────┐ ┌─────────────────┐ │ Tenants │◄──────┤ Users │◄──────┤ Agent Runs │ │ (id, config) │ │ (id, tenant) │ │ (id, plan, DAG) │ └──────┬───────┘ └──────────────┘ └────────┬────────┘ │ │ ▼ ▼ ┌──────────────┐ ┌──────────────┐ ┌─────────────────┐ │ MCP Creds │ │ Approvals │◄──────┤ Audit Events │ │ (Vault Ref) │ │ (hash, sign) │ │ (hash chained) │ └──────────────┘ └──────────────┘ └─────────────────┘

24. Multi-Agent Hierarchy & Orchestration

WorkAgent OS establishes a clear division of labor: the AI Chief of Staff serves as the primary executive orchestrator and coordinator. It decomposes compound organizational goals and delegates specialized sub-tasks to domain-specific agents (Sales, Project, Developer, Finance) with explicit permission envelopes.

27. Strict Tenant Isolation Across All Tiers

Tenant boundaries are designed to be enforced systematically across the full infrastructure stack:

API Gateway → Auth → Tenant Resolver → Queries → PostgreSQL RLS
→ Vector Metadata (tenant_id) → Object Store (tenant/{id}/) → Redis Cache → Connectors

30. Definition of Done (DoD) & Evaluation Standards

A WorkAgent OS agent release is intended to be marked production-ready only when it passes the complete golden benchmark evaluation suite. These are defined targets — none have been executed against a published deployment:

1. Strict Tenant Isolation

Target: 100% rejection across adversarial cross-tenant retrieval queries (planned suite).

2. Hash Integrity Binding

Target: halt whenever action payload parameters diverge from the approved hash.

3. Read-Back Verification

Mandatory external state confirmation matching expected schema prior to run completion.

4. Zero Raw Credential Leak

Target: prompt context inspection clean against sensitive patterns (planned suite).

31. Mandatory Engineering Directives

1. Do not state technical claims without verification tests
2. Bind human approvals to canonical action payload hashes
3. Do not treat SHA-256 as authentication by itself
4. Enforce tenant isolation across all storage and execution layers
5. Scope every tool call to Tenant + User + Agent boundaries
6. Require human sign-off for high-impact actions
7. Enforce read-back state verification on write actions
8. Emit tamper-evident audit records for every critical event
9. Maintain strict model provider abstraction
10. Scope all MCP connectors to tenant-specific credentials
11. Treat external content as untrusted input
12. Prompt injection must never alter tool permission scopes
13. Security controls must be core execution pipeline components
14. Execute regression evaluation suites across all agent releases

P2. Operational Semantics

SPECROADMAP

The sections below document the target operational semantics of the platform. They are design specifications and roadmap items — none are verified deployed behavior.

Concurrency & Conflict Handling

  • •Optimistic concurrency: entities carry version fields; writes fail on stale state rather than silently overwriting.
  • •Duplicate suppression: deterministic idempotency keys prevent repeated side-effects from retries or duplicated runs.
  • •Resource locks: run-scoped locks in the cache layer prevent concurrent mutation of the same external entity.
  • •Conflict handling: on detected drift, the plan re-reads current state and re-plans instead of forcing the stale write.

Cost & Budget Enforcement

Budgets are specified at five scopes. When the projected cost of a planned action exceeds any applicable budget — not only when actual spend crosses it — execution is designed to pause and require approval rather than silently continuing spend:

Tenant budget (monthly aggregate)
Agent budget (per-agent allowance)
Task / run budget (per execution cap)
Model budget (per-provider token ceilings)
Tool budget (per-connector spend limits)

Model Routing (Illustrative)

The adapter layer is designed to route tasks by complexity and policy to different models. Illustrative routing targets — not a deployment description:

  • •Document extraction & summarization → lightweight model
  • •Planning & DAG compilation → frontier reasoning model
  • •Code analysis & refactoring proposals → code-specialized model
  • •High-risk / approval-sensitive steps → most capable model under strict policy

Data Classification Mapping (Proposed)

Proposed mapping of the five classification levels to model, retrieval, storage, log, export, and approval behavior. This mapping is a policy proposal — not enforceable controls today.

LevelModelRetrievalStorageLogExportApproval
PUBLICFull context allowedUnrestricted within tenantStandard tenant storageFull trace detailAllowedNone
INTERNALFull context allowedTenant-scopedStandard tenant storageFull trace detailTenant-internal onlyNone
CONFIDENTIALScoped context, minimal necessaryRole/department gatedEncrypted, access-scopedMasked in tracesPolicy-gatedRequired for external export
RESTRICTEDMinimal-context delivery onlyExplicit grant requiredEncrypted, excluded from long-term memoryRedacted; metadata onlyBlocked by defaultElevated approval required
HIGHLY RESTRICTEDNever enters model contextVault-only, no agent retrievalDedicated vault / HSM-classExistence-only events, no payloadForbiddenCannot be approved for model exposure

Disaster Recovery (Roadmap)

Planned DR scope covers region loss, model provider failure, connector failure, and data corruption. RPO/RTO targets are not yet measured and no values are published:

  • •Tenant-scoped backups with periodic restore testing
  • •Region loss — failover target per deployment configuration
  • •Connector failure — degraded-mode behavior and reconciliation per failure class
  • •Model provider failure — adapter fail-over to alternate provider
  • •Data corruption — point-in-time recovery target

Connector Failure Classes

Seven connector failure classes are specified (schema drift shown as an optional eighth). On timeout the design is to re-read state before retrying — never blindly resend a write. Note: there is no universal exactly-once guarantee across providers (e.g. Gmail); idempotency reduces but does not eliminate duplicate risk.

  1. 1.Auth failure — refresh/rebind credentials once; escalate on repeat
  2. 2.Rate limit (429) — bounded backoff with jitter
  3. 3.Timeout — re-read external state first; avoid blind resend
  4. 4.Partial success — reconcile via read-back and compensating action
  5. 5.Permission denied — surface to user; never retry around policy
  6. 6.External conflict — re-read state, re-plan, don't force stale write
  7. 7.Provider outage — pause queue, fail over per deployment policy
  8. 8.Schema drift (optional 8th) — validate response; fail safe, don't guess fields

33. Ultimate Vision & Master Equation

Identity + Company Context + Agents + MCP / Tools + Memory + Permissions + Approvals + Verification + Audit = WorkAgent OS

"AI agents that can act — with context, permissions, verification and accountability."