
For your company
Governance & Compliance
Module · Prove what the AI did
The AI Evidence & Documentation Check
When a regulator, a court, or a customer asks you to justify an AI-driven decision, memory and good intentions are not evidence. You either kept the record at the time or you did not, and no amount of reconstruction after the fact fully closes the gap. This module checks the five things that make an AI decision defensible: decision logs, versioned prompts and configs, dated approvals, reproducibility, and a retention policy that keeps the evidence long enough to matter.
What the five levels look like
Every dimension in this assessment is scored 1 to 5. This is what the levels mean, dimension by dimension. The graded report diagnoses where your own answers land and what to do about it.
Decisions are logged
- 1Nothing logged
- 2Outcome only
- 3Partial logs
- 4Inputs and output logged
- 5Full decision record
At the low end: An AI decision with no log is one you cannot explain, defend, or learn from. Start capturing the inputs, the recommendation, and the action taken for your higher-stakes decisions now. What good looks like: A full decision record is what turns 'trust us' into evidence. Keep it tamper-evident and tied to the decision, so a record cannot be quietly edited after the fact.
Prompts and configs versioned
- 1Not tracked
- 2Latest only
- 3Some version notes
- 4Versioned informally
- 5Fully version-controlled
At the low end: Untracked prompts and settings mean every past output is unexplainable the moment the configuration changes. Put prompts and key parameters under version control so each output ties to a known state. What good looks like: Version-controlled prompts and configs let you tie any output back to the exact setup that made it. Link the version to the decision log so the two records reinforce each other.
Approvals are dated
- 1No approvals recorded
- 2Informal sign-off
- 3Emails somewhere
- 4Recorded approvals
- 5Dated approvals with basis
At the low end: Undated, unrecorded approvals leave you unable to show that any AI deployment was ever consciously authorised. Start a simple approvals record: what was approved, by whom, on what date. What good looks like: Dated approvals with the evidence they rested on show a governance process that actually operated. Keep the register current so a new deployment cannot go live without its entry.
Past outputs reproduce
- 1Impossible
- 2Roughly, by memory
- 3Some cases
- 4Most recent outputs
- 5Reproducible on demand
At the low end: If you cannot reproduce a past output, you cannot prove it was reasonable when it was made. Start keeping the model version, prompt and inputs together so key decisions can be re-run. What good looks like: On-demand reproducibility is the strongest evidence you can offer that a decision was sound. Test it periodically on an old case; a pipeline that quietly changed may have broken your ability to reproduce without telling you.
Evidence is kept long enough
- 1No retention plan
- 2Deleted ad hoc
- 3Kept, no policy
- 4Policy exists
- 5Policy matched to risk
At the low end: With no retention plan, your evidence may be gone exactly when a challenge arrives. Set how long each type of AI record must survive, driven by how long the underlying decision can be contested. What good looks like: Retention matched to the risk window means the evidence is there when challenged and gone when it is only liability. Review the periods as regulations and contracts change; the right window shifts under you.