Conferral by Design · the public scoring instrument
The Conferral Design Scorecard
Score any AI product — a recommender, an agent, or an agent platform — on ten criteria, 0–4 each. The anchors tell you exactly what each score requires, so two people scoring the same product are answering the same question rather than trading impressions.
What this is. A structured comparison against ten published criteria — a decision aid, not a measure of trustworthiness, a certification, an audit or a prediction. Validity is not established for the total or for any interpretation of it, and no reliability study has been run yet. Read the criterion profile before the number. What changed in v2.0 and why.
Every score also carries an evidence tag: did you observe it, were you told it, or did you infer it? That second layer is the point. A score built on what the team says is a hypothesis, not a finding — and a self-score can be mostly hypothesis without anyone noticing.
Ten criteria · about ten minutes · free, and the definitions are never paywalled
The same scores always produce the same result — the arithmetic is deterministic and published, and the ten judgements feeding it are yours
Whether different people score it the same way is being measured, in public
Your ten scores never leave your browser — the scoring runs here and there is nothing to upload. The optional reminder at the end sends your address and your score, and nothing else.
Evaluating a product someone else built? The buy-side version asks the same ten questions from outside the company, where you cannot see the roadmap and have to decide whether to sign.
Last step · the override
Five red flags
A product can score decently and still have one of these. They are reported separately from the total and do not change it — each describes a structural condition a good average elsewhere does not compensate for, and each needs its own answer.