Governance & Trust
2026Evaluation and Assurance
Test whether an agent is fit for a defined transformation task—and keep testing as knowledge, models and conditions change.
Fluent output can conceal weak implementation judgment. Evaluation must therefore start from the task and its consequences.
Build a representative evaluation set
Include normal cases, edge cases, incomplete evidence, conflicting sources, localization variants and deliberately unsafe requests. Use real project patterns where permitted, with protected or synthetic data as appropriate.
Subject-matter experts should define expected outcomes and acceptable variation.
Measure multiple dimensions
Depending on the agent, assess:
- factual and calculation accuracy.
- completeness and material omissions.
- source quality and traceability.
- consistency with approved decisions and standards.
- process, architecture, control and policy compliance.
- appropriate uncertainty and escalation.
- usefulness to the accountable practitioner.
- security, privacy and tool-use behavior.
- latency and cost.
No single score describes trustworthiness.
Evaluate the workflow
Agent quality depends on retrieval, skills, tools, policies and human interaction—not only the underlying model. Test the full workflow and its failure paths.
For multi-agent scenarios, evaluate handoffs, shared state, conflict resolution and the possibility that one agent amplifies another’s error.
Continue in operation
Use sampled reviews, deterministic checks, shadow evaluation and outcome signals. Segment results by process, release, country, language and risk class so that aggregate performance does not hide local failure.
Material changes require regression testing. Significant incidents should add new cases to the evaluation set.
Keep assurance independent
The team building an agent should test it, but higher-risk capabilities also need challenge from people who do not own its delivery target. Independence makes inconvenient findings more likely to survive.