World Model Readiness
Engraved governance instrument

For your company

Governance & Compliance

Module · Models decay the day they ship

The Model Risk Management Check

A model that passed its demo can quietly rot in production as the world it learned from moves on. The failure is rarely dramatic: accuracy slips a few points a quarter until the recommendations are wrong more often than right, and nobody notices because nobody is watching. This module checks the five disciplines that keep a live model honest: validation before launch, drift monitoring, performance thresholds with alerts, retraining, and knowing which models you even run.

Question 1 of 5 · Validated before launch

Before a model goes live, does anyone independent check that it actually works?

A model that scores well on the data it was built with can still fail on the cases that matter. Independent pre-deployment validation, on data the builder did not touch, is the difference between a tested system and a hopeful one. The team that trained it is the worst-placed to judge it.

Question 2 of 5 · Drift is watched

Would you notice if a live model slowly got worse over the next six months?

Models degrade as the world drifts away from their training data, and the decline is usually gradual enough to miss. Without monitoring on inputs and outputs, the first signal you get is a customer complaint or a bad quarter. Silence is not the same as everything being fine.

Question 3 of 5 · Thresholds trigger alerts

Is there a performance line a model can cross that actually triggers action?

Monitoring only helps if a number crossing a line does something. A defined threshold, with an owner and an alert, turns 'the model is slipping' into 'someone is now looking at it'. A dashboard nobody watches is monitoring in name only.

Question 4 of 5 · Retraining is deliberate

When a model needs refreshing, does that happen by decision or by accident?

Retraining is where model risk is renewed or reduced. Done on a schedule and revalidated, it keeps the model current; done ad hoc under pressure, it can ship an untested model straight into production. A retrain is a new model and deserves the same gate as the first one.

Question 5 of 5 · You know your models

Could you produce a list of every model making decisions in your company today?

You cannot manage risk in models you cannot name. Between bought tools, embedded features and things a team built quietly, the real inventory is usually longer than anyone expects. The models nobody listed are exactly the ones nobody is monitoring.

For the statistics · one click each

Three questions for the public picture

These do not affect your score. They feed the anonymised, aggregated statistics; groups under 8 respondents are never shown.

How many models are making or shaping decisions in your company right now?

None yet
A handful
Dozens
More than we can easily count
We do not know

Do you monitor live models for performance decline?

No monitoring
Manual checks
Some models monitored
Automated across the board
No models in production

Has a model degrading in production already caused a problem for you?

Not that we know of
We suspect so
Yes, minor
Yes, serious
No models in production

Your context

Used to calibrate the report. Company size and sector remain in the anonymized dataset; your email does not.

What the five levels look like

Every dimension in this assessment is scored 1 to 5. This is what the levels mean, dimension by dimension. The graded report diagnoses where your own answers land and what to do about it.

Validated before launch

  1. 1No validation
  2. 2Builder self-checks
  3. 3Informal review
  4. 4Independent validation
  5. 5Validated against holdout

At the low end: Shipping a model nobody validated means production is your test environment and customers are your test cases. Add a validation gate on held-out data before the next model goes live. What good looks like: Independent validation on a clean holdout is the standard most teams skip. Keep the validation set genuinely separate; the moment it leaks into training, the check stops meaning anything.

Drift is watched

  1. 1No monitoring
  2. 2Checked yearly
  3. 3Occasional spot checks
  4. 4Automated monitoring
  5. 5Monitored with baselines

At the low end: An unmonitored model in production is decaying at a rate you have chosen not to measure. Start tracking its inputs and outputs against launch-day baselines so decline becomes visible. What good looks like: Automated drift monitoring against baselines is what lets you trust a model between retrains. Review the baselines periodically; a drift alarm calibrated to last year's world will cry wolf or stay silent.

Thresholds trigger alerts

  1. 1No thresholds
  2. 2Numbers on a dashboard
  3. 3Thresholds, no alerts
  4. 4Alerts, unclear owner
  5. 5Alerts reach an owner

At the low end: Without a defined threshold, there is no moment at which anyone is obliged to act, so nobody does. Set a minimum performance line per model and decide what happens when it breaks. What good looks like: Thresholds that alert a clear owner turn model decline into a routine ticket instead of a crisis. Rehearse the response occasionally; an alert with no agreed next step just adds noise.

Retraining is deliberate

  1. 1Never retrained
  2. 2Only after failure
  3. 3Ad hoc retrains
  4. 4Scheduled retraining
  5. 5Scheduled and revalidated

At the low end: A model that is never refreshed drifts until it is actively misleading. Decide a retraining trigger, whether a calendar date or a drift threshold, before the decline forces your hand. What good looks like: Scheduled retraining with revalidation keeps the model current without smuggling in untested changes. Treat every retrain as a new deployment; skipping validation because 'it is just a refresh' is how regressions ship.

You know your models

  1. 1No inventory
  2. 2In people's heads
  3. 3Partial, outdated
  4. 4Documented inventory
  5. 5Living, owned inventory

At the low end: Without an inventory, every model on this checklist is one you cannot confirm you are managing. Start the list this week: what each model decides, who owns it, where it runs. What good looks like: A living, owned inventory is the foundation every other control here sits on. Tie it to deployment so a new model cannot go live without an entry; unlisted models are unmonitored models.