Governance & Trust

2026

Risk, Failure and Trust

Design for uncertainty, detect failure early and make recovery part of the agentic operating model.

Trust does not mean believing that an agent will never fail. It means knowing which failures are plausible, how they will be detected and what happens next.

Recognize different failure modes

An agent may:

  • invent or misinterpret evidence.
  • apply the wrong release, country or client context.
  • omit a material dependency.
  • overstep its authority.
  • expose restricted information.
  • execute the correct action at the wrong time.
  • create a loop or amplify another agent’s error.
  • degrade after a model, knowledge or tool change.

Controls should target the specific failure, not “AI risk” in the abstract.

Design graceful degradation

When evidence is missing, confidence is low or a control service is unavailable, the workflow should stop, reduce autonomy or route to a person. A fallback should be safer than the normal path, even if it is slower.

Prepare incident response

Define severity, ownership and communication before launch. The team needs the ability to:

  1. contain the affected capability.
  2. preserve relevant evidence.
  3. determine scope and impact.
  4. correct downstream records or actions.
  5. inform affected owners.
  6. improve controls and evaluation.

Calibrate user trust

Both over-trust and under-trust are operational risks. Interfaces and training should communicate capability boundaries, evidence and uncertainty. A highly polished answer should not receive more authority than its evidence warrants.

Learn without normalizing failure

Corrections and near misses are essential learning inputs. They should improve the agent and its controls, but recurring material failures require reduced autonomy or withdrawal—not another disclaimer.