Guide

How to assess and oversee an AI agent

Assessing an AI agent is the risk work you would do for any consequential system, extended to cover the fact that the agent acts. How to run the assessment and put the oversight in place.

How do you assess and oversee an AI agent?

Assessing an AI agent is the risk work you would do for any consequential system, extended to cover the fact that the agent acts. This guide sets out how to run that assessment and how to put the oversight in place once you have, anchored to Australia's Voluntary AI Safety Standard and the NIST AI Risk Management Framework.

It assumes you already have the agent recorded. If you do not, start with what to record about an agent and come back.

Work through it for each agent that warrants the attention. Not every agent needs a full assessment, but any agent that can take a consequential action on its own does.

Before you assess: know what the agent can do

An assessment is only as good as your picture of what the agent actually touches. Before you weigh the risk, write down four things about the agent.

Write down its purpose, which is the goal the agent is set and the scenarios it runs in.

Write down its permissions and reach, meaning the systems it can read from and write to, the credentials it holds and the other agents, tools or services it can call. This is the surface area, and it is where most agent risk lives.

Write down its autonomy. This is which actions it may take on its own and which actions stop for a person to approve. An agent that can draft a message and one that can send it are different systems, and the line between them is the thing you are governing.

Write down its owner, the named person accountable for the agent under Guardrail 1 of the Voluntary AI Safety Standard.

Step 1: Map how the agent could cause harm

Take the agent’s purpose and reach and ask, concretely, what could go wrong when it acts. The NIST AI Risk Management Framework calls this Map, building a clear picture of the risks in context before you try to measure or manage them.

Cover the harms an advisory system could cause, a wrong or biased output, a privacy breach, a poor decision, and then add the harms that only appear because the agent acts. OWASP’s agentic threat work is a useful prompt here. It points to the agent misusing the tools it has been given, acting on a poisoned or manipulated context, holding more access than its task needs and chains of actions that carry a small error a long way before anyone sees it. For each, ask what the agent could do in your systems and who would be affected.

Step 2: Weigh likelihood and consequence

For each harm you have mapped, weigh how likely it is and how serious it would be if it happened. The NIST framework calls this Measure. Two factors specific to agents raise the weight.

Reach raises consequence. An agent that can change customer records, move money or send external communications can do more damage per action than one confined to a sandbox. The more the agent can touch, the higher the consequence side of every harm.

Autonomy raises likelihood of an unreviewed harm. The more an agent does without a person approving each step, the more likely a harm reaches the world before anyone catches it. An action that always stops for approval carries a human check. An action the agent takes on its own does not.

This step tells you which agents are high-risk and warrant a full impact assessment, and which are contained enough to record and monitor without one.

Step 3: Set the oversight that lets a person intervene

Guardrail 5 of the Voluntary AI Safety Standard asks for meaningful human oversight and the ability to intervene in an AI system. For an agent, oversight is not a sign-off at the start. It is a person being able to see what the agent is doing and being able to stop or override it while it runs. Put four things in place.

Set an approval line. Decide which actions the agent may take on its own and which always stop for a person. Set the line by consequence, so the actions that are hard to undo or that reach outside the organisation are the ones that wait for approval.

Bound the agent’s permissions. Give the agent only the access its task needs, and no standing access it does not use. This limits the surface area you assessed in the first place and contains the harm if the agent behaves in a way you did not expect.

Build a way to see and stop it. A person needs a live view of what the agent is doing and a way to halt it. Oversight that cannot intervene until after the fact is monitoring, not control.

Set a monitoring and review cadence. Decide who watches the agent in operation, and set a date to reassess it. Guardrail 4 of the standard asks you to monitor a system once it is deployed, and agents change as their prompts, tools and permissions change, so a review that was accurate six months ago may not describe the agent running today.

Step 4: Record the assessment and keep it current

An assessment that lives in someone’s head is not evidence. Guardrail 9 of the Voluntary AI Safety Standard asks you to keep records that let a third party assess your work. Record the assessment against the agent, note the risk decision and who made it, then set the date it is next reviewed. When the agent’s permissions or autonomy change, the assessment is due again.

Where Aicura fits

Aicura’s AI Register holds the picture the assessment depends on, each agent’s purpose, permissions, autonomy and owner, with its disclosure and model card beside it, versioned as the agent changes and prompting you when a review is due. Incidents give you the place to record when an agent does something it should not have, so the picture stays true to how the agent actually behaves.

Aicura surfaces the picture and holds the evidence. It does not run the assessment for you, decide whether an agent is safe or sign off its use. You weigh the risk and the accountable owner makes the call. Aicura gives them the record to make it on and to show later.

Sources

  • Voluntary AI Safety Standard, Department of Industry, Science and Resources, industry.gov.au/publications/voluntary-ai-safety-standard
  • AI Risk Management Framework (AI RMF 1.0) and Generative AI Profile, NIST, nist.gov/itl/ai-risk-management-framework
  • Agentic AI: Threats and Mitigations (v1.0), OWASP GenAI Security Project, genai.owasp.org

A note on this page

This guide is general information on how to assess and oversee an AI agent against current Australian and allied guidance. It is not legal advice and it does not tell you whether a particular agent meets a particular obligation. For how the guidance applies to your organisation, read the primary sources above and take your own professional advice.

When you need to show your work

Aicura is your AI Register, and Pro adds the assurance layer, meaning impact assessments, incident records, attestation and the Evidence Vault, which seals each record so it can be verified without taking anyone's word for it, ours included. Pro is sales-led, so the best next step is a conversation and a walkthrough.