Autonomous systems · evaluation · governance

Your AI agents act. Evaluation follows, agent by agent.

An agent chains actions together without human validation at every step. That changes the governance question, and how it gets measured.

In brief

An AI agent is a system that pursues a goal by chaining actions autonomously: it calls tools, reads and writes data, and triggers other systems without human validation at each step. identifiable evaluates them one by one with the AI Index, across the six properties of the iDIA framework, because an agent carries risk that evaluating a model alone does not capture.

What is an AI agent?

Published · Last updated · Wissam Daibess

An AI agent pursues a stated goal by deciding for itself which sequence of actions to run. It calls tools, reads and modifies data, triggers other systems, and repeats until the goal is met or abandoned.

The difference lies in the autonomy of the chain. A model answers a request and stops. An agent keeps going, and every link in that chain acts on real systems.

That autonomy is the entire point of agents. It is also why they are governed differently.

How does an agent differ from an AI model?

A model produces an output that someone reads before acting. Supervision is structurally present: it sits between the output and the consequence.

An agent removes that interval across most of its steps. Supervision therefore has to be designed and placed explicitly, at the points where it matters, rather than inherited from the workflow.

The consequences shift as well. A model error produces bad text. An agent error produces a sent email, a modified record, a placed order.

How is an AI agent evaluated?

The AI Index evaluates agents one by one, across the six properties of the iDIA framework: Accountable, Governed, Secure, Sovereign, Supervised, Reproducible. Each property is scored on documented evidence, never on a statement.

Three properties carry most of the weight on an agent. Accountable: who answers, by name, for the actions the agent takes. Supervised: where the human stopping points sit, and on what risk criterion they were placed. Reproducible: can a past decision by the agent be replayed and a deviation explained.

Evaluation is annual and continuous. An agent whose tooling, underlying model or data perimeter changes is no longer the agent that was evaluated.

What risks does an autonomous agent introduce?

The AI Risk SOLN framework separates two origins. Governance covers internal harm, most often unintentional: an agent exceeding its mandate, reaching data outside its perimeter, or acting on a system it was never meant to touch.

Security covers external, intentional attacks. An agent that reads untrusted content widens the prompt-injection surface, and every tool it can reach becomes a capability handed to whoever takes control.

Then comes the sovereignty question: where the data the agent handles travels, and under which jurisdiction. An agent calling several providers splits that answer across them.

Frequently asked questions

Should an AI agent be evaluated separately from its model?

Yes. The model and the agent carry different risks. The AI Index evaluates agents one by one, because the autonomy of the action chain creates exposure that evaluating the model alone does not capture.

Which iDIA properties weigh most on an agent?

Accountable, Supervised and Reproducible. They answer the three questions a board asks about an agent: who answers for it, where the human intervenes, and whether a past decision can be replayed.

How often should an agent be re-evaluated?

Annually, and continuously as soon as a material change occurs. A change of tooling, underlying model or data perimeter produces a different agent from the one that was evaluated.

Does Quebec's Law 25 apply to AI agents?

As soon as an agent processes personal information, Quebec obligations apply, including those on automated decisions. The agent's autonomy does not reduce the organization's obligations.

Can an agent earn the Responsible AI Practice designation?

The designation is granted to an organization, not to an agent. Agents enter the evaluated inventory, and their profile counts in the AI Index that decides the designation.

What are you going to do with your AI?

The agent inventory opens the scope. The AI Index places each one, on evidence.