Framework

PRISM decides which AI capabilities are allowed to ship.

PRISM is the Product Risk and Impact Selection Model. It is not a compliance review that happens after the roadmap is set. It is a selection function that runs before commitment, and its job is to tell you which bets to fund, which to narrow, and which to kill while killing them is still cheap.

Responsible AI is usually sold as a tax on delivery. That framing is why it loses every argument it enters. A model that only ever says no gets routed around. PRISM earns its place by improving the selection, not by slowing the queue: teams that score capabilities on five dimensions before they build ship fewer things, ship them sooner, and withdraw almost nothing after launch.

The five dimensions

Score every candidate on all five. A low score anywhere is a veto.

Each dimension is scored one to five by the people who would carry the consequence, not by the team proposing the work. The total is not an average. Reversibility and impact radius are vetoes: a capability that scores well on purpose and badly on reversibility does not proceed in its current shape, however loudly it is wanted.

P

Purpose

Which decision does this capability change?

  • Name the decision, the person who makes it today, and what they do differently once the model is in the loop.
  • A capability that changes no decision is a demo. It scores zero here regardless of how good the output looks.
R

Reversibility

What does it cost to undo a wrong action?

  • Score the write, not the read. Drafting text is reversible. Sending it, pricing with it, or filing it is not.
  • Reversibility is the dimension that kills the most attractive bets, and it is the one most roadmaps never score.
I

Impact radius

How far does a single wrong output travel?

  • One user, one account, one segment, or every customer at once. Batch and background jobs widen the radius quietly.
  • A capability with a wide radius needs a narrower launch, not a longer review.
S

Supervision surface

Where does a human sit, and can they actually see?

  • Approval buttons are not supervision. Supervision means the reviewer is shown the evidence that would change their answer.
  • If the reviewer approves ninety-eight percent of items without reading them, the surface is decorative. Design it out or design it properly.
M

Measurement

What is the eval set that defines correct?

  • Ten labelled examples of a correct output and ten of an incorrect one, written by the product manager before build starts.
  • No eval set, no acceptance criteria. No acceptance criteria, no ship date.

The sequence

Discover, validate, ship.

The three phases run in order and each one has an exit condition that can be failed. A framework where nothing ever fails a gate is a decoration.

01

Discover

Two weeks, no build

Collect candidate capabilities from support tickets, operator workarounds, and the decisions leadership already reviews by hand. Score each one on all five dimensions in a single room with product, engineering, and whoever carries the risk: legal, clinical, credit, or safety. The output is a ranked list with an explicit kill column. Most teams discover that their most-demanded capability is a low-purpose, high-radius write, and that the one nobody asked for is the reversible read that removes a week of manual review.

02

Validate

Four to six weeks

Build the eval set before the feature. Run the candidate against real production examples, not benchmarks, and hold the result against the tolerance the team wrote down in discovery. Validation ends in one of three states: the capability clears its tolerance, it clears with a narrower scope, or it is withdrawn. The third state has to be used at least once for the first two to mean anything.

03

Ship

Staged, with gates you own

Ship behind the same controls you would apply to any critical dependency: staged rollout, canary comparison against the previous behavior, structured logging of outputs rather than inputs alone, and a named owner for the model-version gate. Shipping is not the end of the assessment. Each model upgrade re-enters validation, because the vendor changing a model is a change to your product whether or not anyone filed a ticket.

Measurement

Evals are the acceptance criteria.

The most common failure in AI delivery is not a bad model. It is a feature that was never specified. A traditional spec says what the screen does. An AI feature has no equivalent surface to describe, so teams write intent and hope the model meets it. Then nobody can say whether launch was a success, because nobody wrote down what correct looked like.

The eval set closes that gap, and it belongs to product, not to the machine learning team. If the product manager cannot produce ten examples of a correct output and ten examples of a wrong one, drawn from real workload rather than imagination, the feature is not specified and the estimate is fiction. Writing those twenty examples usually takes an afternoon and routinely changes the scope, because the act of labelling forces the team to decide what the capability is actually for.

Once the set exists it does the work acceptance criteria always did. It gates the build, it gates the release, and it gates every model upgrade that follows. It also gives the board a defensible answer to the only question that matters about an AI feature: how do you know it works. Not a benchmark score. Your workload, your examples, your tolerance, checked on a schedule.

Generic benchmarks are the trap here. They measure a model against tasks someone else chose. Your business logic, your data formats, and your edge cases are not in that distribution, and the difference between the two is where production incidents live.

Impact radius

Blast radius is a product decision, not an infrastructure one.

The failure mode of production AI is not an outage. It is silent degradation: no error thrown, no alert fired, outputs quietly different after a model upgrade nobody in product signed off on. The infrastructure is healthy throughout. The decisions made on top of it are not.

PRISM scores impact radius early so that the controls are sized to the consequence rather than applied uniformly. A model drafting an internal summary needs logging. A model that writes to a customer record, moves money, or feeds a leadership dashboard needs traced calls, an owned upgrade gate, a workload eval, structured output instrumentation, and a written tolerance agreed before launch instead of argued after an incident.

The two arguments this page rests on are set out in full elsewhere: why every production AI system needs a blast radius plan covers the operational discipline, and the next AI category is not agents, it is the control plane covers the governed execution layer that makes those controls enforceable rather than aspirational.

A worked example: the bet that scores well and still gets killed.

A B2B platform proposes an agent that resolves billing disputes end to end. Purpose scores five: it removes a queue that costs a full-time analyst and delays revenue. Measurement scores four: past resolutions give a strong eval set. Supervision scores three, because a reviewer could approve each resolution, though at volume they would not read them.

Then reversibility scores one. A wrong credit is a ledger entry, a customer communication, and in some jurisdictions a filing. Impact radius scores two, because the agent runs as a nightly batch across every open dispute. The capability does not ship as proposed. It ships as a recommendation the analyst accepts or rejects, with the write left in human hands until the eval set clears its tolerance for three consecutive months. Same value thesis, different shape, and nothing to unwind.

The useful output of PRISM is rarely a no. It is a narrower yes, arrived at before anyone built the wide one.

Selection, not permission

PRISM runs before commitment, so risk shapes the bet instead of arriving as a veto after the demo.

Owned by product

The scores, the eval set, and the tolerances belong to the product leader. Risk and legal advise the scoring; they do not run the queue.

Re-run on every upgrade

A vendor changing a model is a change to your product. The assessment re-opens whether or not anyone filed a ticket.

In practice

Where this has been run.

Context

European automotive retail SaaS, four product tribes, dealer and OEM customers.

Situation

AI, machine learning, and conversational AI work was running across product lines without a shared thesis or a single owner.

Intervention

Consolidated the AI work behind a smaller set of product bets and named an owner for each.

Outcome

A single AI thesis with a named owner behind each bet, and a shorter list of committed initiatives.

The first step

Ninety minutes on the decision you are actually stuck on.

Most advisory relationships start with a discovery call that discovers nothing. This one starts with work. Send the context beforehand — the roadmap, the org chart, the three AI initiatives nobody owns. We spend ninety minutes on the single decision blocking the others. You leave with a written point of view: what you should stop, what you should own, and what the next ninety days look like.

Format

90 minutes, one decision

Deliverable

Written point of view within 48 hours

Price

$1,500, credited against any engagement

Not a fit if you want a vendor evaluation, a staffing plan, or a deck to circulate. If ninety minutes shows this is not a fit, Kevin will say so and point you somewhere better.