← Conferral by Design

Conferral by Design · pre-registration

Can two people apply this rubric and land in the same place?

The strongest objection to any scoring rubric is that it only produces the score its author wanted. The way to answer that is not argument; it is to hand the rubric to people who did not write it, have them use it without conferring, and publish how much they agreed — including if the answer is embarrassing.

Amended 31 August 2026, before any data was collected. This study was pre-registered against Conferral Design Scorecard v1.0. Later the same day, a counter-case defeated the interpretation behind criterion 2, and the rubric was revised to v2.0 — re-anchoring criteria 2 and 8, removing the band labels, and marking four further criteria as under revision. No ratings had been collected and no raters had been recruited, so nothing is being discarded and nothing is being re-run. The pre-registration below is preserved exactly as published, as the record of what was fixed in advance. A study against the v2.0 anchors will be pre-registered separately before any rater is approached; a result obtained under v1.0 anchors could not have supported a claim about v2.0 in any case. The corrections record.

And what this study cannot establish. Agreement between raters would show that the rubric can be applied consistently under the stated conditions. It would not show that the rubric measures what its criteria are named for, that the total predicts anything, or that the 0/2/4 anchors and the weighting are justified. Reliability is a necessary condition, not a sufficient one, and it is not validity. Nothing on this site should be read as claiming otherwise.

Pre-registered 31 August 2026 · results pending

This page was published before any data was collected. The method, the agreement threshold and the condition that would count as failure are all fixed below, and were fixed before a single rating existed. If the study fails its own threshold, this page will say so and the anchors will be revised in public. A pre-registration that can only report success is a press release.

1. The question

Does the Conferral Design Scorecard produce consistent scores when different people apply it to the same evidence?

Note what that does and does not ask. It does not ask whether the rubric measures something worth measuring — that is a validity question and a harder one. It asks the prior question: is this an instrument, or is it one person’s judgement with numbers attached? If independent raters working from identical material land in different places, nothing downstream matters, and the honest thing is to find that out early and publicly.

2. Design

Unit of analysisOne criterion applied to one product. Ten criteria × the number of products in the set.
RatersAt least four, independent, not involved in writing the rubric. They do not confer, do not see each other’s scores, and do not discuss the products until every sheet is submitted.
MaterialsEvery rater scores from an identical evidence pack. This is the design decision that makes the study mean anything: if raters gathered their own evidence, the study would measure what each of them happened to find, not whether the rubric produces agreement. Holding evidence constant isolates the rubric.
ProductsChosen to span the range — at least one expected to score high, one low, and two in the middle. A set clustered at one end cannot distinguish a reliable rubric from a lucky one.
Scale0 / 2 / 4, the published anchors, unmodified. Raters may not use intermediate values.
BlindingRaters are not told the study’s hypothesis, are not told what the author expects any product to score, and submit before any discussion.

3. The statistic, and why this one

Krippendorff’s alpha with the ordinal difference function.

Three reasons it is the right choice here rather than a percentage. It handles more than two raters without averaging pairwise figures. It handles missing data, which matters because a rater may reasonably decline a criterion the evidence pack cannot support. And, decisively, the ordinal form knows that 0 and 2 are closer together than 0 and 4 — a nominal statistic would treat a near-miss as identical to a complete disagreement, which for a rubric with ordered anchors is simply wrong.

Raw percentage agreement is reported alongside it, because alpha is hard to picture and a reader deserves a number they can.

4. The threshold, fixed in advance

AlphaReadingWhat happens next
≥ 0.80ReliableRaters applied the anchors consistently under these conditions. Publish, and cite this study wherever a score is used — as evidence of consistency, never of validity.
0.667 – 0.80TentativeUsable for drawing tentative conclusions only, and every published score must say so. Identify the criteria dragging it down and revise their anchors.
< 0.667Not establishedThe anchors do not yet produce consistent judgements. Publish that finding, name the criteria responsible, rewrite their anchors, and re-run. Do not use scores in commercial work until it clears.

These are Krippendorff’s own customary cut-offs. They are stated here so they cannot be chosen after the fact to flatter a result.

The commitment. Whatever comes back gets published on this page, including a failure, including the per-criterion breakdown that shows which parts of my own rubric other people could not apply consistently. A study that would only have been published if it succeeded is not evidence of anything.

5. What a failure would mean — and what it would not

A low alpha would not mean the underlying distinction is wrong. It would mean the anchors are underspecified: that two careful readers can look at the same evidence and reasonably land on different rungs. That is a writing problem, it is fixable, and finding it is the point of doing this.

It would also predict which criteria fail. My own expectation, recorded here in advance so it can be checked: criteria 1 and 2 — objective honesty and granted-share awareness — will show the weakest agreement, because both ask about something a rater usually cannot observe directly and must infer. If the result matches that prediction it is mildly reassuring about the method; if the disagreement lands somewhere I did not expect, that is more interesting and more useful.

6. The instruments

Both tools run entirely in this page. Nothing you enter is transmitted or stored, and closing the tab discards everything — which is also why each rater must copy their export string out and send it to the coordinator themselves.

Rating sheet

Score the product in front of you against the ten published anchors. Do not discuss it with anyone until you have submitted. If the evidence pack cannot support a criterion, leave it blank — a declined criterion is better data than a guessed one.

Nothing on the sheet yet.

Send the export string to the coordinator. It contains your rater code, the product codes and your scores — nothing else.

Agreement analysis

Paste every rater’s export string, one per line. At least two raters are needed; the pre-registered design calls for four.

7. Results

Pending. No data has been collected. When the study runs, the alpha, the percentage agreement, the per-criterion breakdown, the number of raters and the product set will be published here, with the date — and this section will replace itself whether the result is good or bad.

8. Limitations, stated before the result

9. Run it on something else

The tools above are not specific to this rubric in any deep way — ten ordinal items on a 0/2/4 scale is a common enough shape. If you want to check the reliability of a rubric of your own, use them. The alpha implementation is in the source of this page and is tested against ten fixtures; open ?selftest to run them in your own browser.

Conferral Theory and the Conferral Design Scorecard are the work of Clint Miller. This page is a study pre-registration; no results are reported at the date of publication, and nothing here should be cited as a finding until the results section replaces its pending notice.