Governance & Trust
2026Risk, Failure and Trust
Design for uncertainty, detect failure early and make recovery part of the agentic operating model.
Trust does not mean believing that an agent will never fail. It means knowing which failures are plausible, how they will be detected and what happens next.
Recognize different failure modes
An agent may:
- invent or misinterpret evidence.
- apply the wrong release, country or client context.
- omit a material dependency.
- overstep its authority.
- expose restricted information.
- execute the correct action at the wrong time.
- create a loop or amplify another agent’s error.
- degrade after a model, knowledge or tool change.
Controls should target the specific failure, not “AI risk” in the abstract.
Design graceful degradation
When evidence is missing, confidence is low or a control service is unavailable, the workflow should stop, reduce autonomy or route to a person. A fallback should be safer than the normal path, even if it is slower.
Prepare incident response
Define severity, ownership and communication before launch. The team needs the ability to:
- contain the affected capability.
- preserve relevant evidence.
- determine scope and impact.
- correct downstream records or actions.
- inform affected owners.
- improve controls and evaluation.
Calibrate user trust
Both over-trust and under-trust are operational risks. Interfaces and training should communicate capability boundaries, evidence and uncertainty. A highly polished answer should not receive more authority than its evidence warrants.
Learn without normalizing failure
Corrections and near misses are essential learning inputs. They should improve the agent and its controls, but recurring material failures require reduced autonomy or withdrawal—not another disclaimer.