Why this is a different instrument
The team that builds a product can look at its own dashboard. You cannot. So this version does not ask what the roadmap intends — it asks what you can establish before signature, and it separates three things that a procurement process routinely blends together:
| What it means | What it is worth | |
|---|---|---|
| We saw it | You tested it, in a trial or a demo you controlled | The strongest evidence available to a buyer |
| It’s in writing | It is in the agreement, the DPA, or a signed commitment | The only kind that survives a change of account manager |
| They told us | Someone said it, or it is in a deck or a policy page | Real, and unenforceable |
The score matters less than that split. A high score built entirely on what you were told is a description of a sales process, not of a product — and the gap between what was said in the room and what is warranted in the agreement is the single most useful number this produces.
Five things that override the score
If any of these is true, the total does not matter. Each is a thing you can check rather than a judgement you have to make.
The rubric, in full
Published complete and never paywalled, so an assessment can be checked by the person you hand it to — and so a vendor can see exactly what they were scored against. Version v1.0.
| Criterion | 0 | 4 |
|---|---|---|
| 1. Objective honesty What is this product optimised for — and can they say it in one sentence? | They cannot or will not state the objective. The pitch leads with adoption, engagement or time-in-product as though those were the benefit. | They can state what the product refuses to do for engagement, and the success measures in the agreement are outcome measures rather than usage measures. |
| 2. Prompted versus chosen usage Will they report usage your people chose, separately from usage the product prompted? | Usage arrives as a single number. A notification-driven open and a deliberate one are counted identically. | The split appears in your reporting, and prompted usage does not count toward any adoption target. |
| 3. Permission and revocation What can it do without asking, and how fast can you take that back? | Broad scopes at deployment, autonomy bundled together, revocation requires a support request. | Least-privilege defaults, an explicit ladder from suggest to act, and self-service revocation that takes effect immediately and is logged. |
| 4. Calibration and pushback Does it disagree with a confident user who is wrong — and can they show evidence? | No evaluation evidence offered. Demonstrations are happy-path and the product agrees with whatever framing it is given. | They can show measurement of agreement-under-pressure specifically, and will let you run your own adversarial test before signature. |
| 5. Failure conduct When it is wrong, who finds out and how quickly? | No disclosure commitment. You expect to hear about failures from your own people. | A defined notification window in writing, a published record of past incidents, and a route for your users to flag a wrong answer that reaches the vendor. |
| 6. Commercial influence on output Does anyone pay to affect what it says? | Placements, affiliate relationships or house services can influence outputs, and this is not disclosed. | No commercial interest touches the output, and the vendor will warrant that in the agreement. |
| 7. Withdrawal reporting Will they tell you when your people quietly stop trusting it? | Only adoption and growth are reported. Turning it off is invisible to you until renewal. | Mutes, disablements, permission revocations and abandoned tasks are reported and attributed to the change that preceded them. |
| 8. Interruption dependence Does its success depend on interrupting your people? | Notifications, nudges and re-engagement are load-bearing to the vendor's own adoption story. | It works with notifications off, and the vendor's own success measure survives that setting. |
| 9. Provenance Can your people see why it said that, and check it? | Outputs arrive with no account of where they came from. | Every output carries provenance a user can verify, including what internal data was drawn on and why it was surfaced now. |
| 10. What it trusts downstream What does it rely on, and what happens when that is wrong? | It accepts output from tools, plug-ins or other agents without verification, and nothing is logged. | Reliance flows only through verified identity, every delegated action is logged and revocable, and there is an accountable path when a downstream system is wrong. |
Where this comes from
This is the buy-side fork of the Conferral Design Scorecard, which scores a product from inside the team that builds it. Both come from Conferral by Design, the methodology, published open. The metrics are defined at the Conferral Metrics Standard.
If you want the same assessment run on a product by someone outside your organisation — with the evidence discipline enforced and the findings written up — that is the Conferral Design Review. The first three are $2,500, in exchange for permission to publish the review.
The Conferral Vendor Scorecard and Conferral Theory are the work of Clint Miller. This instrument supports a purchasing decision; it is not legal, security or compliance advice, and the suggested contract language is a starting point for your own counsel rather than a substitute for them.