Every arrow is a control point, not a pipe.
Onam Operations is a governed agent platform built above the products that already know your cloud. It reasons over their resolved data instead of re-scanning, sends every call through one tool gateway and one model gateway, and puts a person between any recommendation and any change. This page walks through the design, block by block.
- Person
- Agent
- Multi-agent plan
- Workflow
- Tool
- Cloud
- Validation
- Evidence
Each step can refuse. Nothing reaches your cloud without passing back through a person.
Eight blocks on the request path, four beside it
A request descends one block at a time and never skips one. Beside every block runs a control plane — governance, policy, approval and audit — that is in the call path rather than advisory: each can refuse, and a refusal stops the request. Below the capability block the path forks three ways: reading the estate, executing an approved change, and calling a model. Only the execution fork can change anything.
| Block | Contains | Must never |
|---|---|---|
| AExperience | Workspace UI, embedded panels, API clients | Hold a credential, or call a product backend directly |
| BEntry | The existing platform API gateway and the Operations API | Accept a tenant ID from a request body |
| COrchestration | Orchestrator, task service, workflow engine | Read estate data itself, or decide whether a control applies |
| DAgents | Agent runtime, agent registry, context engine, memory | Call a tool directly, or modify its own definition or scope |
| ECapability | Skill runtime, tool gateway, model gateway | Let unsanitised text re-enter the model’s context |
| FIntelligence | Cloud Estate Intelligence: estate query, graph, identity resolver | Write anything, or guess an identity mapping |
| GExecution | The execution sandbox | Start without an approval record, or reach a host outside its allowlist |
| HModels | Approved model providers | Receive data outside the permitted provider or region |
The “must never” column is the useful one. A diagram that only says what each block does cannot be enforced; one that says what each block is forbidden to do becomes a review checklist and a set of automated build checks.
Agents reason over resolved truth — they do not rediscover your cloud
The products already scan your cloud. Cloud Estate Intelligence is a read-only layer over what they found: assets, relationships, findings and, in time, cost and recovery. Agents query it instead of calling cloud APIs, which is the difference between an answer in seconds and an answer after a fresh scan. An agent that goes and scans your cloud itself is treated as a bug.
It reads published, read-only views of each product’s data — never a product’s own tables — under a role that can only select.
An identity resolver maps each product’s resource IDs to one canonical ID, so answers join on the same resource rather than a similar name.
Results carry the time they were observed and the scan they came from, so an answer can say how old it is.
If a source is unavailable, the answer says which domain is missing and what is therefore unknown, instead of quietly returning less.
| Data source | What agents see | Status |
|---|---|---|
| Onam Security — inventory, assets, findings | What exists and what is wrong with it | Early access |
| Onam Security — relationships, exposure, attack paths, compliance results | How it connects, what is reachable, which controls hold | In development |
| Cost data (Onam FinOps) | What it costs and what is wasted | On the roadmap |
| Recovery data (Onam DRM) | What would come back, and how fast | On the roadmap |
One turn: budgeted, scope-checked, recorded
Each agent turn runs the same eight steps. Two orderings matter more than the rest: the permission check happens before the tool call, and sanitising happens before the result goes back to the model. Reversing either would open a privilege-escalation or an injection path.
No component holds a model SDK. All inference goes through one gateway that enforces the approved-model list, your region, per-task budgets, and that only one customer’s data is in any call. Today it routes to Amazon Bedrock in-region.
Agents that propose or act do not fall back to a weaker model when the preferred one is unavailable; they fail safe and say so.
Every level has a ceiling on tool calls and wall-clock time, and every task a cost ceiling. Repeated identical calls are detected and the task is stopped.
Eleven words that are never used interchangeably
Agent platforms go wrong when a dozen concepts collapse into three words and nobody can reason about permissions. These are fixed.
- Agent
- A specialist with a declared purpose, level, skills, tools and permissions — a versioned definition, not a prompt.
- Skill
- A declared, versioned capability with typed input and output, a risk level and an evidence contract. Agents invoke skills.
- Tool
- A concrete integration that touches something outside the platform. Only skills invoke tools; only the tool gateway runs them.
- Permission
- Authority for one agent to perform one action on one class of resource through one tool, in one organisation.
- Policy
- A rule, held as data, that decides whether something may happen and how many approvals it needs.
- Task
- One unit of work with a lifecycle, an owner and a result — visible in the task centre.
- Workflow
- A versioned multi-step procedure that can outlive a single request, such as preparing an account for audit.
- Action
- A single attempted change to the outside world. Only the automation actor performs one, only from an approval.
- Approval
- A recorded human decision authorising one specific change at one specific version.
- Evidence
- The data behind a claim — query, rows, scan and time — retained and addressable.
- Memory
- What an agent keeps between turns, deliberately narrow. It never holds an estate fact or a secret.
An agent cannot exceed the person it acts for
A permission is a tuple — agent, organisation, accounts, resource class, tool, action — and the effective permission is the intersection of what the agent is granted, what you are allowed to do, your organisation’s policy and the scope of the conversation. An agent acting for a read-only viewer can do no more than that viewer.
| Level | Writes | Approval |
|---|---|---|
| L1 · Investigator | Its own investigation notes | None — it cannot change anything |
| L2 · Proposer | Change proposals only | Every proposal needs a human decision before anything is applied |
| L3 · Actor | The customer cloud, scoped to one approved change | Mandatory, per action. Starts only from an approval record |
| L4 · Orchestrator | Nothing | Routes work; never answers or acts itself |
| Action | Meaning | Who may hold it |
|---|---|---|
| read | Retrieve existing data | Any agent |
| analyse | Compute over retrieved data | Any agent |
| recommend | Produce a proposal | Any agent — applying it needs approval |
| simulate | Dry-run with no side effect | Any agent |
| execute · modify | Apply or alter a change | L3 actor only, per risk class |
| delete | Remove a resource | L3 actor only, always high or critical |
| approve | Authorise an action | Never an agent — people only |
An agent’s level is fixed in its definition. Promotion is a governance event that needs re-certification — scope review, security review including the prompt-injection test corpus, an evaluation suite, and documentation of its limits — never a runtime decision.
Seven gates between a proposal and your cloud
A change must pass entitlement, agent scope, the actor level, an execute permission, a policy-assigned risk class, a person approving the exact change, and a drift check on the target. Any one of them holding prevents an unauthorised change. Below the application, the database will not store an executed action without a matching approval.
The execution sandbox
On the roadmapBuilt and tested in the platform; enabled for no customer yet. When it is, it will require a write role you create in your own account and an autonomy ceiling you raise past “propose”.
- Its own isolated namespace — the only place holding a write credential
- Network egress default-deny, allowlisted to the endpoints the action declares
- Short-lived role scoped to the target account and operation
- One action per container, an ephemeral filesystem, a hard time limit
- Every command and API call captured into the evidence package
- Validation after every change; failure rolls back and raises an incident
A new surface on the existing platform, not a new silo
Onam Operations runs as its own service group on the Onam platform, behind the same gateway, session and organisation model as Onam Security. It introduces no new identity or tenancy mechanism: identity reaches it only from the gateway, every table carries the tenant, and its data stays in the region agreed with you.
Never touches a product database directly — all estate reads go through Cloud Estate Intelligence
Nothing reaches a customer cloud except through the tool gateway and the sandbox
No component holds a model SDK; inference goes through the model gateway
Identity and tenant come only from the platform gateway
Agent records live alongside the platform’s own, under the same naming and tenancy rules