World Model Readiness
Engraved governance instrument

For your company

Governance & Compliance

Module · Prove what the AI did

The AI Evidence & Documentation Check

When a regulator, a court, or a customer asks you to justify an AI-driven decision, memory and good intentions are not evidence. You either kept the record at the time or you did not, and no amount of reconstruction after the fact fully closes the gap. This module checks the five things that make an AI decision defensible: decision logs, versioned prompts and configs, dated approvals, reproducibility, and a retention policy that keeps the evidence long enough to matter.

Question 1 of 5 · Decisions are logged

For an AI-influenced decision, can you show what the system recommended and what was done?

The defensible record is not 'the AI suggested it' but the actual inputs, the recommendation, and the human action that followed. Without a decision log captured at the time, you are reconstructing from memory under pressure. Auditors and courts read contemporaneous records, not later explanations.

Question 2 of 5 · Prompts and configs versioned

Can you tell which prompt and settings produced an output six months ago?

The same model gives different answers as prompts, parameters and system messages change. If those are not versioned, you cannot say what configuration produced a given result, and you cannot explain a past decision. A prompt edited in place erases its own history.

Question 3 of 5 · Approvals are dated

When someone signed off on deploying an AI system, is that decision recorded and dated?

Governance that leaves no dated trace is invisible to an auditor. Who approved a deployment, on what date, on what evidence: that record is what shows a decision was made deliberately rather than drifted into. A verbal yes in a meeting is not documentation.

Question 4 of 5 · Past outputs reproduce

If challenged, could you reproduce an AI output your company relied on a year ago?

Reproducibility is the hardest evidence and the most convincing: showing that the same inputs and configuration still yield the same result. It depends on having kept the model version, the prompt, the data and the settings together. Without it, you can describe a decision but not demonstrate it.

Question 5 of 5 · Evidence is kept long enough

Do your AI records survive as long as the decisions they justify can be challenged?

Evidence that is deleted before the limitation period ends is no evidence at all. Retention has to match how long a decision can come back: contract terms, regulatory windows, statutory limits. Too short and you are exposed; forever and you create a different liability. Someone has to have decided.

For the statistics · one click each

Three questions for the public picture

These do not affect your score. They feed the anonymised, aggregated statistics; groups under 8 respondents are never shown.

Do you keep decision logs for AI-influenced decisions?

None
For some decisions
For most decisions
For all high-stakes decisions
Not sure

Could you reproduce an important AI output from a year ago?

No
Roughly
For some
Yes, reliably
We have never tried

Have you ever been asked to evidence an AI decision?

Never
Internally only
By a customer
By a regulator or court
Prefer not to say

Your context

Used to calibrate the report. Company size and sector remain in the anonymized dataset; your email does not.

What the five levels look like

Every dimension in this assessment is scored 1 to 5. This is what the levels mean, dimension by dimension. The graded report diagnoses where your own answers land and what to do about it.

Decisions are logged

  1. 1Nothing logged
  2. 2Outcome only
  3. 3Partial logs
  4. 4Inputs and output logged
  5. 5Full decision record

At the low end: An AI decision with no log is one you cannot explain, defend, or learn from. Start capturing the inputs, the recommendation, and the action taken for your higher-stakes decisions now. What good looks like: A full decision record is what turns 'trust us' into evidence. Keep it tamper-evident and tied to the decision, so a record cannot be quietly edited after the fact.

Prompts and configs versioned

  1. 1Not tracked
  2. 2Latest only
  3. 3Some version notes
  4. 4Versioned informally
  5. 5Fully version-controlled

At the low end: Untracked prompts and settings mean every past output is unexplainable the moment the configuration changes. Put prompts and key parameters under version control so each output ties to a known state. What good looks like: Version-controlled prompts and configs let you tie any output back to the exact setup that made it. Link the version to the decision log so the two records reinforce each other.

Approvals are dated

  1. 1No approvals recorded
  2. 2Informal sign-off
  3. 3Emails somewhere
  4. 4Recorded approvals
  5. 5Dated approvals with basis

At the low end: Undated, unrecorded approvals leave you unable to show that any AI deployment was ever consciously authorised. Start a simple approvals record: what was approved, by whom, on what date. What good looks like: Dated approvals with the evidence they rested on show a governance process that actually operated. Keep the register current so a new deployment cannot go live without its entry.

Past outputs reproduce

  1. 1Impossible
  2. 2Roughly, by memory
  3. 3Some cases
  4. 4Most recent outputs
  5. 5Reproducible on demand

At the low end: If you cannot reproduce a past output, you cannot prove it was reasonable when it was made. Start keeping the model version, prompt and inputs together so key decisions can be re-run. What good looks like: On-demand reproducibility is the strongest evidence you can offer that a decision was sound. Test it periodically on an old case; a pipeline that quietly changed may have broken your ability to reproduce without telling you.

Evidence is kept long enough

  1. 1No retention plan
  2. 2Deleted ad hoc
  3. 3Kept, no policy
  4. 4Policy exists
  5. 5Policy matched to risk

At the low end: With no retention plan, your evidence may be gone exactly when a challenge arrives. Set how long each type of AI record must survive, driven by how long the underlying decision can be contested. What good looks like: Retention matched to the risk window means the evidence is there when challenged and gone when it is only liability. Review the periods as regulations and contracts change; the right window shifts under you.