Conferral by Design · the build-side instrument
The launch gate
v1.0 · 8 questions, 9 fields · free to adopt, copy or adapt
A launch review usually collects evidence without forcing a conflict.
The feature met its quality threshold. Latency is acceptable. The interface passed review. The growth model is positive. Known risks have owners. Every line can be true while the central decision goes unstated: what is the system now permitted and rewarded to do that it could not do before?
An AI feature that recommends or acts changes more than capability. It changes the allocation of judgement. A launch may let the product select, rank, speak with confidence, request broader access, act through another service, or interrupt at a time of its choosing. Those are product choices even when the underlying capability arrived from a model or a vendor.
Free to adopt, copy, adapt or rename. This is published so a team can run it on Monday without asking anyone. It is not a certification, and a completed record can still support a bad decision. Its value is that the decision cannot later pretend the trade-off did not exist.
Start with the promise, not the feature list
Write one sentence:
This product helps [specific user] achieve [observable outcome] by [recommendation or action], subject to [the boundary the product will not cross].
The boundary is load-bearing. “Help people manage their work” is a direction. “Do not send, spend, publish or grant access without a function-specific authorisation” is a product decision.
Then ask eight questions
1
What wins when the measures conflict?
Name the success measure that governs experiments. Then name one situation in which it is allowed to fall.
If completion improves while session time declines, is that a win? If the system refuses a task outside its capability and action volume falls, is that a win? If a revocation control becomes easier to find and recorded withdrawals rise, is that a win?
The answer cannot be “it depends” without the condition that decides it. The gate is valuable because it forces the condition into the room before the result arrives.
2
What evidence supports the product’s confidence?
Do not ask whether the model is generally capable. Ask what evidence supports this function, for this task distribution, under these conditions.
A system may be reliable at drafting and unreliable at choosing. It may extract dates accurately and infer intent badly. It may perform well when a user corrects it and poorly when left alone. Capability does not transfer automatically across functions, contexts or consequences.
Limits should appear where they change a decision. A limitations page can be accurate and still fail to govern an interaction.
3
What happens when the user is wrong?
An assistant that agrees with a confident user whose premise is wrong will feel easy to use. Ease is not calibration.
Launch evidence should include cases where the user is mistaken, the evidence is incomplete, and the product is uncertain. The question is not whether it disagrees often — routine contrarianism is another failure. It is whether confidence and dissent track the evidence rather than the pressure to preserve approval.
Record what it does: answer, qualify, ask, refuse, or escalate.
4
What exactly has been authorised?
List authority by function.
Suggesting a meeting time is not booking it. Drafting a message is not sending it. Finding a product is not buying it. Reading an account is not changing it. A global label like “agent mode” compresses distinctions the product still has to implement.
For each function record the default, the grant, the confirmation rule, the consequence, the reversal path, and who can revoke it. Prior success can inform a new grant. It cannot silently create one.
5
What will the user see after it acts?
Completion should leave evidence proportionate to consequence.
A low-risk reversible task may need a one-line receipt. A consequential action may need the object changed, the authority used, the reason, the downstream service, and the route to reverse or challenge it. More detail is not automatically better — boilerplate makes inspectability nominal rather than usable.
The test is whether the user can tell what happened without reconstructing the system’s internal process.
6
How does it behave after failure?
Failure conduct begins before the apology.
Who detects the problem? What reaches the affected user, and how quickly? What is corrected? Does the system narrow the relevant capability or permission claim? Can the user report an error through the product, and where does that report go?
There is no universal script: a competence failure and an integrity violation do not call for the same response. “We are sorry” is not a correction protocol.
7
Which interests can shape the recommendation?
Record commercial interests at the point where they enter the product decision.
Paid placement, affiliate consideration, house services, preferred providers, incentives attached to use. The existence of an interest does not prove the output is bad, and disclosure does not prove the user noticed it. The question is whether a material interest can shape the answer, whether the user can separate it from the product’s judgement, and what alternative remains available.
8
What evidence would make you stop?
A consequential launch should carry a reversal condition.
Name the event, threshold or pattern that triggers a pause, a narrower permission, a rollback or a new confirmation rule. Include evidence the current dashboard might read as good: more time, more prompted returns, more actions, fewer visible withdrawals.
This is where the session-origin distribution and a withdrawal count can earn their place — as proposals to test, not established measures. If the product cannot link origin, authorisation and exposure, report the observable event and the missing data rather than manufacturing the interpretation.
The decision record
The gate ends in a record, not a verdict. Nine fields:
| Field | Required entry |
|---|---|
| Product promise and refusal | One sentence each |
| Governing outcome | The outcome this feature is meant to improve |
| Metric allowed to fall | The proximate number the team will not protect at the outcome’s expense |
| Authority added | Function, default, grant, confirmation, reversal |
| Strongest contrary evidence | The finding that would defeat the launch case |
| Unresolved uncertainty | What the current evidence cannot establish |
| Cost accepted | Usage, speed, scope, revenue, or operational burden |
| Decision and owner | Launch, narrow, test, postpone or refuse; a named accountable role |
| Revisit date | A date — not “after sufficient data” |
Five dispositions, not two
Launch — the evidence and authority are adequate for the stated scope.
Narrow — it may ship for fewer functions, users or consequences.
Test — the uncertainty can be reduced under controlled exposure first.
Postpone — a necessary dependency or piece of evidence is not ready.
Refuse — the present business case depends on conduct the product has said it will not accept.
These are product outcomes, not levels of virtue. A narrow launch may be commercially stronger than a broad one if it keeps the feature where performance is demonstrable and reversal is cheap. A test can be irresponsible when the exposure itself creates the unacceptable consequence. A refusal can become an approval once the objective, the permission or the evidence changes.
The point is to stop “launch” and “do not launch” carrying the whole burden. There is almost always another lever: reduce the authority, separate the functions, change the default, add a confirmation at a consequence boundary, or remove an adoption target that rewards the wrong behaviour.
Record the rejected alternative as well as the chosen one. Six months later a team needs to know whether a boundary was deliberate, inherited or forgotten. Without that record, a temporary exception becomes the product’s permission model by accumulation.
How this relates to the rest
The gate is the build-side version of a question that shows up in three places. The buy-side instrument asks it of a vendor before you sign, and ends in contract language. The scorecard The route audit Corrections asks it of a whole product rather than one launch. This asks it of the specific decision in front of you.
The form changes with the decision. The questions do not.
A launch gate is a policy artifact rather than a review: something a team adopts once and runs every time, which is why it is published free and unbranded enough to rename.
Buying rather than building? The buy-side instrument asks the same eight questions of a vendor before you sign, with the question to ask in the room and the contract clause each one produces. Also free.
The essays: Sycophancy is a counterfeit deposit · Session time is an anti-metric for agents · Machine trust will be allocated by contest · An ad in a list is not an ad in an answer · The same task, two designs (runnable) · The two numbers nobody publishes