WorkAgent OS Architecture
A decoupled, modular execution engine designed from the ground up for strict multi-tenant boundaries, verifiable tool orchestration, and model-agnostic reasoning.
The 6-Layer Platform Architecture
WorkAgent OS decouples interchangeable reasoning models from business rules, persistence, and tool integrations. The 6-Layer Platform Architecture provides the multi-tenant infrastructure host, while the 8-Layer Execution Model governs each agent action lifecycle.
Presentation & Workspace Layer
Modern Next.js 16 web application, role-based Employee Workspace, Admin Console, and Human-in-the-Loop review portals.
API Gateway & Security Boundary
Enterprise gateway (SPEC) designed for authentication, request-signature validation where applicable, tenant isolation, and rate limiting. Request-signature checks (e.g. JWT or webhook verification) are distinct from action payload hashes.
Agent Runtime & Execution Engine
The core reasoning engine. Features modular LLM adapters (Claude, GPT, DeepSeek), 9-step Context Engine, Dynamic Risk Engine, and State Verification.
MCP-Compatible Tool Gateway
Bi-directional connector gateway providing tenant-scoped, permission-controlled tool access to enterprise SaaS and internal databases.
Enterprise Multi-Tenant Storage Layer
Strictly segregated multi-tenant persistence layer combining relational data, vector embeddings, and tamper-evident audit logging.
Observability, Cost & Evaluation Layer
Full OpenTelemetry-compliant execution tracing, real-time cost accounting, and automated regression evaluation suites.
Agent Runtime & Execution Engine
The core reasoning engine. Features modular LLM adapters (Claude, GPT, DeepSeek), 9-step Context Engine, Dynamic Risk Engine, and State Verification.
Core Modules & Protocols:
9-Stage Context Resolution
Context resolution stages: Identity, Tenant, User, Task, Conversation, Org, Memory, Tool, Policy. Retrieval and ranking (structured fetch, semantic search, relevance, source authority, recency, dedup, budget, citations) run as a separate pipeline.
Planning & Tool Router
Decomposes requests into step-by-step DAGs, selects verified tools, validates JSON schemas.
Policy & Risk Engine
Normalized formula: Impact × Irreversibility × Externality × Sensitivity × Uncertainty × 100. 0-100 tiered action gates.
State Verification & Recovery
Re-reads external state post-execution, verifies outcome schema against expected entity, and triggers compensation, retry, reconciliation, or escalation depending on connector capability.
The Agent Runtime Algorithm
Every run is designed to follow a fixed, verifiable execution algorithm. The model never communicates directly with raw databases or unverified APIs; every action is mediated by the Policy Engine, Risk Scorer, and Verifier.
function run_agent(request, user, tenant):
# Step 1: Authentication & Tenant Resolution
authorize_user(user, tenant)
# Step 2: Context Retrieval & Filtering
context = build_context(user_profile, history, memory, sources, request)
policy = load_policy(tenant, agent_type)
validate_context_access(context, policy)
# Step 3: Planning & Decomposition
task = classify_task(request)
plan = model.plan(task, context, available_tools)
# Step 4: Step-by-Step Tool Execution Loop
for step in plan:
tool = resolve_tool(step)
if not policy.allows(tool, user, context):
return blocked("permission_denied")
risk = risk_engine.assess(step, tool, context)
if risk.requires_human_approval:
approval = request_approval(step, risk)
if not approval.granted:
audit("approval_denied")
continue
result = execute_tool(tool, step.parameters)
# Step 5: Verification & State Recovery
if not verifier.validate(result, step.expected_outcome):
recovery = planner.replan(step, result)
if recovery.available:
continue
return safe_failure()
write_trace(step, result)
update_short_term_memory(step, result)
persist_relevant_long_term_memory()
return final_response()The 9-Step Context Engine Pipeline
AI agents should never be flooded with indiscriminate data dumps. The Context Builder applies rigorous semantic, temporal, and authority filters to curate the exact minimum required tokens.
Intent & Entity Decomposition
Parses prompt for core intent, entity references, and relevant time window.
Tenant Permission Scoping
Defines strict tenant and department boundaries before any database querying.
Structured Deterministic Fetch
Queries deterministic sources: Calendar attendees, CRM opportunities, user profile.
Semantic Vector Search
Performs pgvector cosine similarity search across Drive docs and past emails.
Multi-Factor Scoring
Calculates composite score: Recency (30%) + Semantic Relevance (40%) + Source Authority (30%).
Deduplication & Noise Pruning
Discards redundant email chains and low-confidence document fragments.
Context Budget Allocation
Caps assembled context within the pre-allocated token budget (e.g. 8,000 tokens).
Citation Metadata Tagging
Attaches persistent source IDs and timestamps to each context fragment.
Prompt Context Delivery
Transmits verified, citation-backed context payload to the model reasoning adapter.
Verification & Self-Healing Recovery
Traditional agents assume tool execution succeeded if an HTTP 200 was returned. WorkAgent OS enforces active state verification: reading back the external resource to confirm changes took effect.
POST /jira/rest/api/3/issue (idempotency: 77a1bc)GET /jira/rest/api/3/issue/ENG-4821Status == "OPEN" & Assignee == "[email protected]" ✓Connector Verification & Recovery Capabilities
The specification classifies each external connector by its independent read-back verification rule, proposed retry policy, compensating capability, and reconciliation audit. These are design examples — provider behavior must be verified per deployment.
| Connector | Read Verification Strategy | Retry Strategy | Rollback Capability | Compensation Mechanism | Reconciliation Audit |
|---|---|---|---|---|---|
| Google Calendar | GET event by ID verify status === 'confirmed' and start/end matches | Proposed: exponential backoff (3 attempts, max 10s) | Compensating | Compensating action: delete created event ID or restore previous event payload snapshot | Etag comparison against audit log snapshot |
| Gmail | Verify draft ID exists or sent message ID in sent folder | Proposed: linear retry on 429 / 503 (max 2 attempts) | None (Manual) | Irreversible external send; pre-execution hash approval required; send follow-up cancellation if configured | Message-ID and thread-ID logged in append-only audit chain |
| Jira / Linear | GET issue by issue_key verify fields, status, and assignee | Proposed: exponential backoff on 429 rate limit (3 attempts) | Partial / Compensating | Transition issue to 'Cancelled' or revert custom fields to cached pre-execution state | Changelog API diff comparison against execution intent |
| Slack | conversations.history by message ts verify text and attachments | Proposed: exponential backoff with jitter on HTTP 429 | Compensating | Compensating action: chat.delete by channel and ts or chat.update with redaction notice | Message timestamp and channel ID mapped in execution store |
| Salesforce / CRM | SOQL query by record ID verify updated field values and SystemModstamp | Proposed: exponential backoff on transient network / lock errors (3 attempts) | Partial / Compensating | Revert updated fields to snapshot state captured in pre-execution context | Field history tracking cross-checked against audit store |
Risk Scoring Specification
SPECIllustrative scoring model — policies, tenant configuration, and scope checks can still deny low-scoring actions. The canonical formula multiplies five normalized factors:
| Factor | Range | Meaning |
|---|---|---|
| Impact | 0.0 – 1.0 | Blast radius of the action if executed wrongly |
| Irreversibility | 0.0 – 1.0 | Cost/difficulty of undoing the effect |
| Externality | 0.0 – 1.0 | Effect on parties/systems outside the tenant |
| Sensitivity | 0.0 – 1.0 | Data-classification level touched |
| Uncertainty | 0.0 – 1.0 | Model/policy confidence margin |
| Score | Tier | Behavior |
|---|---|---|
| 0 – 20 | Auto Execution | Eligible for autonomous execution in the illustrative model — still subject to policy, tenant, and scope checks. |
| 21 – 50 | ABAC / Policy Check | Policy evaluation before execution. |
| 51 – 80 | Human Approval Gate | Requires hash-bound human approval. |
| 81 – 100 | Deny / Elevated Approval | Denied or routed to elevated approval. |
Enterprise Trust Zones & Data Flow Topology
A simplified perspective for enterprise security architects showing how user requests flow safely from identity providers into policy-bounded agent runtimes and external enterprise SaaS tools.
Human & Identity Scope
Enterprise SSO (SAML/OIDC), session resolution, and delegated execution authority binding (Tenant → User → Agent).
Policy & Risk Kernel
Deterministic ABAC/RBAC rules, Quantitative Risk Formula (0–100), and canonical SHA-256 action hash approval gates.
Context Engine & Memory
PostgreSQL RLS, pgvector semantic search, working memory, and token budget allocation with source citations.
Model-Agnostic Adapter
Stateless commercial foundation model inference (target); retention terms depend on the provider contract.
MCP Tool Gateway & Vault
Credential isolation inside Vault, idempotency control, external state read-back, and append-only audit ledger commit.