Governance & Trust

2026

Evaluation and Assurance

Test whether an agent is fit for a defined transformation task—and keep testing as knowledge, models and conditions change.

Fluent output can conceal weak implementation judgment. Evaluation must therefore start from the task and its consequences.

Build a representative evaluation set

Include normal cases, edge cases, incomplete evidence, conflicting sources, localization variants and deliberately unsafe requests. Use real project patterns where permitted, with protected or synthetic data as appropriate.

Subject-matter experts should define expected outcomes and acceptable variation.

Measure multiple dimensions

Depending on the agent, assess:

  • factual and calculation accuracy.
  • completeness and material omissions.
  • source quality and traceability.
  • consistency with approved decisions and standards.
  • process, architecture, control and policy compliance.
  • appropriate uncertainty and escalation.
  • usefulness to the accountable practitioner.
  • security, privacy and tool-use behavior.
  • latency and cost.

No single score describes trustworthiness.

Evaluate the workflow

Agent quality depends on retrieval, skills, tools, policies and human interaction—not only the underlying model. Test the full workflow and its failure paths.

For multi-agent scenarios, evaluate handoffs, shared state, conflict resolution and the possibility that one agent amplifies another’s error.

Continue in operation

Use sampled reviews, deterministic checks, shadow evaluation and outcome signals. Segment results by process, release, country, language and risk class so that aggregate performance does not hide local failure.

Material changes require regression testing. Significant incidents should add new cases to the evaluation set.

Keep assurance independent

The team building an agent should test it, but higher-risk capabilities also need challenge from people who do not own its delivery target. Independence makes inconvenient findings more likely to survive.