Onam Security
Onam Operations
How it is designedEarly access

Every arrow is a control point, not a pipe.

Onam Operations is a governed agent platform built above the products that already know your cloud. It reasons over their resolved data instead of re-scanning, sends every call through one tool gateway and one model gateway, and puts a person between any recommendation and any change. This page walks through the design, block by block.

  1. Person
  2. Agent
  3. Multi-agent plan
  4. Workflow
  5. Tool
  6. Cloud
  7. Validation
  8. Evidence

Each step can refuse. Nothing reaches your cloud without passing back through a person.

Block architecture

Eight blocks on the request path, four beside it

A request descends one block at a time and never skips one. Beside every block runs a control plane — governance, policy, approval and audit — that is in the call path rather than advisory: each can refuse, and a refusal stops the request. Below the capability block the path forks three ways: reading the estate, executing an approved change, and calling a model. Only the execution fork can change anything.

Block architecture: workspace, entry, orchestration, agents and capability blocks descend to Cloud Estate Intelligence, the execution sandbox and approved models, with a control-plane column of governance, policy, approval and audit beside them.
Products feed Cloud Estate Intelligence; agents reach it only through the tool gateway; inference goes only through the model gateway; changes go only through the approval gate and the sandbox.
BlockContainsMust never
AExperienceWorkspace UI, embedded panels, API clientsHold a credential, or call a product backend directly
BEntryThe existing platform API gateway and the Operations APIAccept a tenant ID from a request body
COrchestrationOrchestrator, task service, workflow engineRead estate data itself, or decide whether a control applies
DAgentsAgent runtime, agent registry, context engine, memoryCall a tool directly, or modify its own definition or scope
ECapabilitySkill runtime, tool gateway, model gatewayLet unsanitised text re-enter the model’s context
FIntelligenceCloud Estate Intelligence: estate query, graph, identity resolverWrite anything, or guess an identity mapping
GExecutionThe execution sandboxStart without an approval record, or reach a host outside its allowlist
HModelsApproved model providersReceive data outside the permitted provider or region
Governance
Kill switch, organisation settings, agent / skill / tool registries
Must never: be bypassed at runtime.
Policy engine
Evaluates the permission tuple and assigns the risk class
Must never: be influenced by prompt content.
Approval service
Records a person’s decision bound to the change’s hash
Must never: be held or exercised by an agent.
Audit and evidence
Append-only, hash-chained record of every call, success or failure
Must never: be updated or deleted.

The “must never” column is the useful one. A diagram that only says what each block does cannot be enforced; one that says what each block is forbidden to do becomes a review checklist and a set of automated build checks.

Cloud Estate Intelligence

Agents reason over resolved truth — they do not rediscover your cloud

The products already scan your cloud. Cloud Estate Intelligence is a read-only layer over what they found: assets, relationships, findings and, in time, cost and recovery. Agents query it instead of calling cloud APIs, which is the difference between an answer in seconds and an answer after a fresh scan. An agent that goes and scans your cloud itself is treated as a bug.

Read through contracts

It reads published, read-only views of each product’s data — never a product’s own tables — under a role that can only select.

One identity per resource

An identity resolver maps each product’s resource IDs to one canonical ID, so answers join on the same resource rather than a similar name.

Freshness on every row

Results carry the time they were observed and the scan they came from, so an answer can say how old it is.

Honest about gaps

If a source is unavailable, the answer says which domain is missing and what is therefore unknown, instead of quietly returning less.

Data sourceWhat agents seeStatus
Onam Security — inventory, assets, findingsWhat exists and what is wrong with itEarly access
Onam Security — relationships, exposure, attack paths, compliance resultsHow it connects, what is reachable, which controls holdIn development
Cost data (Onam FinOps)What it costs and what is wastedOn the roadmap
Recovery data (Onam DRM)What would come back, and how fastOn the roadmap
Agent runtime and model gateway

One turn: budgeted, scope-checked, recorded

Each agent turn runs the same eight steps. Two orderings matter more than the rest: the permission check happens before the tool call, and sanitising happens before the result goes back to the model. Reversing either would open a privilege-escalation or an injection path.

Eight steps of an agent turn: governance check, context assembly, skill choice through the model gateway, policy check, tool gateway, estate query, sanitising, and answer with evidence.
If the model is persuaded by injected text to try something else, the policy engine and tool gateway still refuse it.
Model gateway

No component holds a model SDK. All inference goes through one gateway that enforces the approved-model list, your region, per-task budgets, and that only one customer’s data is in any call. Today it routes to Amazon Bedrock in-region.

No silent downgrade

Agents that propose or act do not fall back to a weaker model when the preferred one is unavailable; they fail safe and say so.

Budgets and loops

Every level has a ceiling on tool calls and wall-clock time, and every task a cost ceiling. Repeated identical calls are detected and the task is stopped.

Concepts

Eleven words that are never used interchangeably

Agent platforms go wrong when a dozen concepts collapse into three words and nobody can reason about permissions. These are fixed.

Agent
A specialist with a declared purpose, level, skills, tools and permissions — a versioned definition, not a prompt.
Skill
A declared, versioned capability with typed input and output, a risk level and an evidence contract. Agents invoke skills.
Tool
A concrete integration that touches something outside the platform. Only skills invoke tools; only the tool gateway runs them.
Permission
Authority for one agent to perform one action on one class of resource through one tool, in one organisation.
Policy
A rule, held as data, that decides whether something may happen and how many approvals it needs.
Task
One unit of work with a lifecycle, an owner and a result — visible in the task centre.
Workflow
A versioned multi-step procedure that can outlive a single request, such as preparing an account for audit.
Action
A single attempted change to the outside world. Only the automation actor performs one, only from an approval.
Approval
A recorded human decision authorising one specific change at one specific version.
Evidence
The data behind a claim — query, rows, scan and time — retained and addressable.
Memory
What an agent keeps between turns, deliberately narrow. It never holds an estate fact or a secret.
Permissions and levels

An agent cannot exceed the person it acts for

A permission is a tuple — agent, organisation, accounts, resource class, tool, action — and the effective permission is the intersection of what the agent is granted, what you are allowed to do, your organisation’s policy and the scope of the conversation. An agent acting for a read-only viewer can do no more than that viewer.

LevelWritesApproval
L1 · InvestigatorIts own investigation notesNone — it cannot change anything
L2 · ProposerChange proposals onlyEvery proposal needs a human decision before anything is applied
L3 · ActorThe customer cloud, scoped to one approved changeMandatory, per action. Starts only from an approval record
L4 · OrchestratorNothingRoutes work; never answers or acts itself
ActionMeaningWho may hold it
readRetrieve existing dataAny agent
analyseCompute over retrieved dataAny agent
recommendProduce a proposalAny agent — applying it needs approval
simulateDry-run with no side effectAny agent
execute · modifyApply or alter a changeL3 actor only, per risk class
deleteRemove a resourceL3 actor only, always high or critical
approveAuthorise an actionNever an agent — people only

An agent’s level is fixed in its definition. Promotion is a governance event that needs re-certification — scope review, security review including the prompt-injection test corpus, an evaluation suite, and documentation of its limits — never a runtime decision.

The approval gate and execution

Seven gates between a proposal and your cloud

A change must pass entitlement, agent scope, the actor level, an execute permission, a policy-assigned risk class, a person approving the exact change, and a drift check on the target. Any one of them holding prevents an unauthorised change. Below the application, the database will not store an executed action without a matching approval.

Read path with three gates; write path with seven gates, one a person, then sandbox execution, validation, evidence and rollback on failure.
The asymmetry is the product: reading is fast because it is safe; writing is deliberately expensive.
Approval lifecycle: proposal, policy risk class, pending with simulate, modify, reject or approve, approval bound to hash and expiring, drift check before execution.
An approval is bound to one change, expires, and is superseded if the target changes before execution.

The execution sandbox

On the roadmap

Built and tested in the platform; enabled for no customer yet. When it is, it will require a write role you create in your own account and an autonomy ceiling you raise past “propose”.

  • Its own isolated namespace — the only place holding a write credential
  • Network egress default-deny, allowlisted to the endpoints the action declares
  • Short-lived role scoped to the target account and operation
  • One action per container, an ephemeral filesystem, a hard time limit
  • Every command and API call captured into the evidence package
  • Validation after every change; failure rolls back and raises an incident
Where it runs

A new surface on the existing platform, not a new silo

Onam Operations runs as its own service group on the Onam platform, behind the same gateway, session and organisation model as Onam Security. It introduces no new identity or tenancy mechanism: identity reaches it only from the gateway, every table carries the tenant, and its data stays in the region agreed with you.

Commitment 1

Never touches a product database directly — all estate reads go through Cloud Estate Intelligence

Commitment 2

Nothing reaches a customer cloud except through the tool gateway and the sandbox

Commitment 3

No component holds a model SDK; inference goes through the model gateway

Commitment 4

Identity and tenant come only from the platform gateway

Commitment 5

Agent records live alongside the platform’s own, under the same naming and tenancy rules