Conferral by Design · the buy-side instrument

You are about to deploy someone else’s AI to your people.

Ten questions to answer before you sign — not about whether the model is good, but about whether the product is built so your people’s reliance on it is earned rather than drawn against. Most of it you can establish in one call.

Ten criteria · about ten minutes · free, and the rubric is never paywalled
Nothing you type leaves your browser. There is no account and no storage.
Built for the buyer. The build-side version is for teams scoring their own product.

Why this is a different instrument

The team that builds a product can look at its own dashboard. You cannot. So this version does not ask what the roadmap intends — it asks what you can establish before signature, and it separates three things that a procurement process routinely blends together:

 What it meansWhat it is worth
We saw itYou tested it, in a trial or a demo you controlledThe strongest evidence available to a buyer
It’s in writingIt is in the agreement, the DPA, or a signed commitmentThe only kind that survives a change of account manager
They told usSomeone said it, or it is in a deck or a policy pageReal, and unenforceable

The score matters less than that split. A high score built entirely on what you were told is a description of a sales process, not of a product — and the gap between what was said in the room and what is warranted in the agreement is the single most useful number this produces.

This is a design and commercial-terms assessment. It is not a security review, a model evaluation, a safety audit or a compliance assessment, and it does not replace any of them. It asks a question those processes do not: whether this product needs your people’s attention in order to succeed. How this sits alongside the EU AI Act, NIST and ISO 42001 →

Last step · the five stoppers

Five things that override the score

If any of these is true, the total does not matter. Each is a thing you can check rather than a judgement you have to make.

The rubric, in full

Published complete and never paywalled, so an assessment can be checked by the person you hand it to — and so a vendor can see exactly what they were scored against. Version v1.0.

Criterion04
1. Objective honesty
What is this product optimised for — and can they say it in one sentence?
They cannot or will not state the objective. The pitch leads with adoption, engagement or time-in-product as though those were the benefit.They can state what the product refuses to do for engagement, and the success measures in the agreement are outcome measures rather than usage measures.
2. Prompted versus chosen usage
Will they report usage your people chose, separately from usage the product prompted?
Usage arrives as a single number. A notification-driven open and a deliberate one are counted identically.The split appears in your reporting, and prompted usage does not count toward any adoption target.
3. Permission and revocation
What can it do without asking, and how fast can you take that back?
Broad scopes at deployment, autonomy bundled together, revocation requires a support request.Least-privilege defaults, an explicit ladder from suggest to act, and self-service revocation that takes effect immediately and is logged.
4. Calibration and pushback
Does it disagree with a confident user who is wrong — and can they show evidence?
No evaluation evidence offered. Demonstrations are happy-path and the product agrees with whatever framing it is given.They can show measurement of agreement-under-pressure specifically, and will let you run your own adversarial test before signature.
5. Failure conduct
When it is wrong, who finds out and how quickly?
No disclosure commitment. You expect to hear about failures from your own people.A defined notification window in writing, a published record of past incidents, and a route for your users to flag a wrong answer that reaches the vendor.
6. Commercial influence on output
Does anyone pay to affect what it says?
Placements, affiliate relationships or house services can influence outputs, and this is not disclosed.No commercial interest touches the output, and the vendor will warrant that in the agreement.
7. Withdrawal reporting
Will they tell you when your people quietly stop trusting it?
Only adoption and growth are reported. Turning it off is invisible to you until renewal.Mutes, disablements, permission revocations and abandoned tasks are reported and attributed to the change that preceded them.
8. Interruption dependence
Does its success depend on interrupting your people?
Notifications, nudges and re-engagement are load-bearing to the vendor's own adoption story.It works with notifications off, and the vendor's own success measure survives that setting.
9. Provenance
Can your people see why it said that, and check it?
Outputs arrive with no account of where they came from.Every output carries provenance a user can verify, including what internal data was drawn on and why it was surfaced now.
10. What it trusts downstream
What does it rely on, and what happens when that is wrong?
It accepts output from tools, plug-ins or other agents without verification, and nothing is logged.Reliance flows only through verified identity, every delegated action is logged and revocable, and there is an accountable path when a downstream system is wrong.
Where this comes from

This is the buy-side fork of the Conferral Design Scorecard, which scores a product from inside the team that builds it. Both come from Conferral by Design, the methodology, published open. The metrics are defined at the Conferral Metrics Standard.

If you want the same assessment run on a product by someone outside your organisation — with the evidence discipline enforced and the findings written up — that is the Conferral Design Review. The first three are $2,500, in exchange for permission to publish the review.

The Conferral Vendor Scorecard and Conferral Theory are the work of Clint Miller. This instrument supports a purchasing decision; it is not legal, security or compliance advice, and the suggested contract language is a starting point for your own counsel rather than a substitute for them.