Buying rather than building? This agenda is for a team reviewing its own product. If you are evaluating something a vendor built, use the vendor scorecard instead — the same ten criteria, re-derived for someone deciding whether to sign.
Who to have in the room. Whoever owns the metrics, whoever owns the model or the ranking, and one person who actually talks to users. Three to six people. More than six and you get a presentation instead of a review — the failure mode this agenda is shaped to avoid.
The agenda
State the product’s own claim
One sentence, written where everyone can see it: what does this product say it is optimising for?
Do not discuss it yet. You will come back to it at the end, and it often turns out to disagree with the dashboard — which is the first finding and it arrives for free.
Score it, silently and separately
Everyone opens the scorecard and scores independently, without discussing. Ten criteria, 0–4 each, against published anchors that say exactly what each score requires.
The silence is load-bearing. A group that scores together scores its most senior person’s opinion, and you will have spent fifteen minutes discovering what your VP thinks. You already knew that.
The Conferral Design Scorecard → — free, no signup, and nothing you enter leaves your browser.
Compare only the disagreements
Skip every criterion where you agreed. Agreement is not information.
The spread is the finding. Two people scoring the same feature 1 and 4 means the team does not share a definition of what the product is for — and that is worth considerably more than the total. Spend the whole fifteen minutes on the widest three gaps and let the rest go.
Take the evidence tags seriously
Every score carries a tag: did you observe it, were you told it, or did you infer it? Count them.
A score built mostly on stated is a hypothesis wearing a number, and a team can run for a year on one without noticing. This is the step people skip, and it is the one that changes what the meeting was worth.
The four questions
Ask them in this order. They are ordered by how uncomfortable they get.
If session time went down next quarter and task completion went up, would we call that a win — and would our dashboard?
Where does the product ask for trust before it has done anything to earn it? Name the specific screen.
What does the model do when the user is wrong? If the answer is “agrees pleasantly,” you have found the most expensive thing in the review.
Which of our numbers could be bought? Anything purchasable is not evidence of conferral. Sort the dashboard into bought and earned and look at what is left.
What to do with the result
One change, not a roadmap. The scorecard names a band and the criterion that moved it most. Fix that one, then re-score in ninety days.
A single score is a snapshot; the movement is the finding. A team that scores once and files it has bought nothing.
Five red flags override the arithmetic entirely — if any one of them is true, the total does not matter and the scorecard says so. They are published with everything else.
Why this is free
The rubric, the anchors, the per-criterion fixes, the red flags and the bands are public and will not be put behind a door. A measure that only one person can run is an opinion with a number attached.
What a review from outside adds is the one thing a self-score structurally cannot: the evidence discipline enforced by someone who is not on the team. A room scoring itself will tag as observed things it has only ever been told. That is not dishonesty, it is proximity, and it is why external review exists at all.
If you want it done from outside
A Conferral Design Review applies the same scorecard to one product with that discipline enforced: every score tagged by someone who is not on the team, the three highest-leverage corrections written as implementation sketches, and the metric panel to instrument so the score becomes trackable rather than annual.
It takes one conversation and product access. It is not an ML audit and does not require model internals — it is a design and strategy review, which is where these ten failures live. The first three are $2,500, in exchange for permission to publish the review.
Two worked examples, scored end to end and published complete: an AI executive assistant and a personalised discovery app. Both are fictional composites, and each says so before any finding.
The Conferral Design Scorecard and Conferral Theory are the work of Clint Miller. This page describes a self-assessment process for product teams; it is not a safety evaluation, a compliance audit, or a substitute for either.