Onam

What is agentic AIOps?

In short

Agentic AIOps applies AI agents to IT and cloud operations. Instead of only correlating alerts, agents plan and carry out multi-step work — investigating an incident, gathering evidence, proposing a fix — using tools with defined permissions. In a human-in-the-loop design, a person approves any change before an agent makes it.

From AIOps to agentic AIOps

AIOps — a term coined by the analyst firm Gartner — describes applying machine learning and analytics to operations data: correlating alerts, detecting anomalies, reducing noise. It tells an operator what is happening faster.

Agentic AIOps goes a step further. An AI agent, usually built on a large language model, is given a goal ("why did checkout latency rise?"), a set of tools (query the inventory, read the logs, check recent changes) and permission boundaries. It plans the steps, calls the tools, reasons about the results and reports back — or proposes an action.

The difference is between a dashboard that highlights an anomaly and an assistant that investigates it.

What can AI agents do in cloud operations?

TaskWhat the agent does
InvestigationGathers configuration, metrics, logs and recent changes across systems and summarises what it finds
Inventory questionsAnswers "what do we run, where, and who owns it?" from live data
Security triageExplains a finding, checks whether it is reachable or exploitable, and groups related issues
Change proposalsDrafts a fix — a configuration change, a pull request, a runbook — for a person to review
ReportingProduces summaries and evidence for audits or reviews

Why human-in-the-loop approval matters

An agent that can only read is low risk. An agent that can change production is a new kind of privileged identity — one that can be wrong with confidence, or be manipulated through the data it reads.

Human-in-the-loop design keeps a person in the decision: the agent investigates and proposes, and a named human approves before anything changes. Good implementations make the decision easy to make well:

  • the proposal states what will change, why, and what evidence supports it;
  • it shows how to undo the change;
  • approval is a recorded act by an accountable person, not a default.

Levels of autonomy

A useful way to think about agent authority is as a ladder an organisation climbs deliberately:

  1. Answer — read-only questions and explanations.
  2. Investigate — multi-step, read-only analysis with evidence.
  3. Propose — draft changes for a person to approve.
  4. Act with approval — execute a change after explicit approval.
  5. Act within policy — execute pre-approved, low-risk changes on its own, inside strict limits.

Most organisations should start at the bottom and move up one level at a time, per environment and per kind of change.

How to keep AI agents safe in operations

  • Least privilege. Give each agent only the tools and permissions its job needs, scoped to the accounts it works on.
  • Evidence on every claim. Require agents to cite the data behind their conclusions, so a person can check them.
  • Audit trail. Record every question, tool call, proposal and approval.
  • A kill switch. Be able to stop all agent activity immediately.
  • An autonomy ceiling. Set, per organisation or environment, the highest level an agent may operate at.
  • Treat inputs as untrusted. Logs, tickets and resource tags can contain text written to manipulate a model. Security guidance such as the OWASP Top 10 for LLM applications lists prompt injection and excessive agency among the main risks.

Next steps

Frequently asked questions

What is the difference between AIOps and agentic AIOps?

AIOps applies machine learning to operations data to correlate alerts and detect anomalies. Agentic AIOps adds AI agents that plan and carry out multi-step tasks with tools — investigating an issue, gathering evidence, proposing a fix — rather than only surfacing signals for a person to investigate.

What does human-in-the-loop mean for AI agents?

It means a person approves an agent's proposed action before it is carried out. The agent can investigate and recommend, but a named, accountable human decides whether a change is made, and that decision is recorded.

Is it safe to let AI agents change cloud infrastructure?

Only with controls: least-privilege permissions, human approval for changes, an audit trail, a kill switch and a ceiling on autonomy. Many organisations keep agents read-only or proposal-only until they have evidence about how the agents behave in their environment.

What is prompt injection in AIOps?

Prompt injection is when text in the data an agent reads — a log line, a ticket, a resource tag — contains instructions that try to change the agent's behaviour. Agents in operations read a lot of untrusted text, so their permissions and approval steps must assume some of it is hostile.

See it on your own cloud

Onam AIOps is in early access: agents investigate with evidence and propose; a person approves every change.