World Model Readiness
Engraved team-functions instrument

For your team

Teams by Function

Module · Whether AI features ship honest

The Product Team AI Check

Shipping an AI feature is easy; shipping one that keeps working is the hard part. A model that demoed beautifully can degrade in the wild, delight the wrong users, or quietly optimise for a number that does not matter. This module checks the five disciplines that separate a durable AI feature from a lucky demo: evals before launch, real feedback loops, success metrics beyond engagement, degradation monitoring, and a roadmap that tells the truth about what the model cannot do.

Question 1 of 5 · Evals before launch

Does your team run an eval suite before an AI feature ships?

A model that looks great in a handful of hand-picked demos can fail on the long tail nobody tried. An eval suite is a repeatable test set with a known answer key, so you measure quality before users do. Vibes are not a launch gate.

Question 2 of 5 · Feedback loops close

Does real user feedback actually make it back into the model or prompts?

A thumbs-down button that nobody reads is theatre. A closed loop means the signal from users, corrections, complaints, abandoned sessions, reaches the people who can change the prompt, the retrieval, or the model, and does. Otherwise the same failure ships forever.

Question 3 of 5 · Metrics beyond engagement

Do you measure whether the AI feature actually helps, not just whether people use it?

Engagement is easy to grow and easy to fake: a confusing feature drives clicks too. The metrics that matter are whether users got the outcome they came for, task completion, resolution, time saved, trust. Optimising engagement alone can make a feature more addictive and less useful at once.

Question 4 of 5 · Degradation is watched

Would your team notice if the AI feature quietly got worse in production?

Models drift: inputs shift, a provider updates a version, retrieval goes stale, and quality slides without a single error in the logs. Without monitoring on output quality, the first person to notice degradation is an angry user, months in.

Question 5 of 5 · Roadmap is honest

Is your roadmap honest about what the AI cannot reliably do yet?

Pressure to promise AI magic is intense, and a roadmap that oversells sets the team up to ship something that cannot deliver. Honesty means the limits are named in planning: what the model gets wrong, where it needs a human, what is genuinely not ready. Silence about limits becomes a launch commitment.

For the statistics · one click each

Three questions for the public picture

These do not affect your score. They feed the anonymised, aggregated statistics; groups under 8 respondents are never shown.

How many AI features has your team shipped to users?

None yet
One
A handful
Many, across the product

How does your team decide an AI feature is good enough to ship?

Gut feel and demos
Manual spot checks
A repeatable eval set
An eval that gates launch
Not shipped one yet

What does your team mainly measure to judge an AI feature's success?

Usage and engagement
Usage plus some outcomes
User outcomes primarily
Nothing formal yet
Not shipped one yet

Your context

Used to calibrate the report. Company size and sector remain in the anonymized dataset; your email does not.

What the five levels look like

Every dimension in this assessment is scored 1 to 5. This is what the levels mean, dimension by dimension. The graded report diagnoses where your own answers land and what to do about it.

Evals before launch

  1. 1No evals
  2. 2Manual spot checks
  3. 3Ad hoc test set
  4. 4Eval suite runs
  5. 5Eval gate, versioned

At the low end: Shipping an AI feature with no eval is launching blind and hoping. Build a small test set with expected answers and run it before the next release. What good looks like: A versioned eval suite that gates launch is what lets your team ship AI features with a straight face. Grow the set as you find new failure modes; an eval that never changes stops catching new bugs.

Feedback loops close

  1. 1No feedback path
  2. 2Collected, ignored
  3. 3Read occasionally
  4. 4Reviewed regularly
  5. 5Drives model changes

At the low end: A feature that cannot hear its users repeats its mistakes indefinitely. Add a feedback path and, more importantly, someone whose job is to act on it. What good looks like: A loop that actually drives changes is how an AI feature gets better in the wild instead of decaying. Keep the cycle short; feedback that takes a quarter to land teaches the model slowly.

Metrics beyond engagement

  1. 1Engagement only
  2. 2Vanity metrics
  3. 3Some outcome data
  4. 4Outcome metrics tracked
  5. 5Outcomes drive decisions

At the low end: Measuring only engagement tells you the feature is used, not that it works. Define an outcome metric, task completion or resolution, and track it alongside usage. What good looks like: Outcome metrics driving decisions keep your team building features that help rather than features that merely hold attention. Watch for metrics gaming; any number you optimise hard eventually gets gamed.

Degradation is watched

  1. 1No monitoring
  2. 2Users report it
  3. 3Manual checks
  4. 4Quality monitored
  5. 5Monitored with alerting

At the low end: Silent degradation is the default failure mode of a shipped model, and you have no way to see it. Put a quality signal in production before the next provider update moves the ground under you. What good looks like: Monitored quality with alerting means degradation is a page, not a surprise. Keep the baseline current; a monitor calibrated to last year's model misses this year's slide.

Roadmap is honest

  1. 1Overpromises freely
  2. 2Limits unspoken
  3. 3Limits known internally
  4. 4Limits documented
  5. 5Limits shape the roadmap

At the low end: A roadmap that promises what the model cannot do commits your team to shipping a disappointment. Name the known limits in planning before they become deadlines. What good looks like: A roadmap that lets the model's real limits shape what you commit to is how a product team keeps its credibility. Revisit the limits as the models improve; yesterday's hard no can become today's cautious yes.