Conferral by Design · essay
The two numbers nobody publishes
v1.1 · amended 31 August 2026 · sample of six, one date, public sources only · free to quote with attribution
Corrected 31 August 2026, the day it was published. This page originally proposed a measure called Initiated Share, defined as a binary split between sessions a person began and sessions the product prompted. A counter-case defeated that definition within hours. The correction is the first section below, and the replacement is specified at the session-origin distribution. The rest of the argument stands, and one further claim in it has been narrowed.
In August 2026 I checked six major AI assistant surfaces against the ten criteria of the Conferral Design Scorecard, using public disclosure only — documentation, policy pages, official announcements and on-the-record reporting. No internal access, no claims about anyone’s private practice.
Eight of the ten criteria produced something to score. Two produced nothing at all, on any surface in the sample: nothing about how a session began, and nothing about people reducing or ending the product’s access to them.
I am scoping that carefully, because it is a claim about a sample rather than an industry: six surfaces, checked on one date, public sources only. Somebody may publish one of these tomorrow, and if they have already and I missed it, I want to know.
The correction
This is a correction, not a clarification.
Version 1.0 of this page defined Initiated Share as the fraction of sessions a person began without the product having prompted them. A session opened within a stated window of any outbound approach — a notification, an email, a badge, a re-engagement message — counted as prompted. Everything else counted as initiated. The interpretation was that initiated use was closer to wanted use, and prompted use closer to captured use.
That interpretation fails.
A person can ask an AI product to do something later. The product does the work and sends a notification when the result is ready. The person opens it. Under the v1.0 definition that session is prompted, and Initiated Share falls precisely because the delegated service worked as requested.
This is not a hypothetical edge case. OpenAI’s official documentation describes one-time and recurring tasks in ChatGPT, monitoring for changes, and task notifications. Google’s official documentation describes user-scheduled recurring actions in Gemini, prepared in the background, with a notification when the result is ready.1 2
The prompt is the delivery mechanism of a prior grant. It is not evidence of capture. The binary erased the question that actually matters — who authorised the product to act, for what purpose, and whether that authorisation is still current.
So the v1.0 measure is retired. Unprompted initiation survives as one useful class of behaviour. It does not survive as a direction of product quality. What replaces it is a session-origin distribution: eligible sessions reported across eight disclosed origin classes, with no weighted aggregate and no assumption that any class is intrinsically good or bad. The correction costs the simplicity of one number and buys back the distinction the number was erasing.
Why these two, and not the other eight
They are the two that would tell you whether reliance on the product is being earned or spent.
Every other criterion can be satisfied by a product that is nonetheless drawing down the trust it started with. You can label your advertising, document your permissions, publish a model card and disclose your incidents while still running a product whose growth depends on interrupting people. Those criteria describe conduct.
These two describe the balance. One says where your usage came from; the other says how fast people are taking their reliance back. Unlike the other eight, neither can be satisfied by writing a better policy page.
The absence is structural, not lazy
It would be easy and wrong to read this as an oversight. Consider what these two have in common that the published numbers do not.
Every metric these products publish is one that goes up. Weekly actives, sessions, messages, adoption, time in product, developers on the platform, tokens served. The direction is the point: a number that rises is a number that can be announced. That is a claim about what is published, not about what is measured — teams do track churn, revocation, complaints and intervention internally. The gap is that none of it is joined into a public account.
Session origin and withdrawal are different in kind. They are not merely capable of falling — they fall precisely when the growth tactics are working. Ship a more insistent re-engagement campaign and your sessions rise while the marketing-prompted share of them rises too. Widen a default and your engagement rises while your revocation rate does. A company publishing both would be publishing, every quarter, an account of what its own growth had cost.
That is not an oversight. That is a coherent reason not to.
What it would actually cost
Here the v1.0 page made its second error, and the correction is worth stating separately, because it cuts one of the argument’s legs and leaves the other standing.
The original claim was that both numbers already exist inside these companies, so publishing them is a matter of willingness rather than engineering. That is true of withdrawal and false of origin.
Withdrawal events are close to already-instrumented. A product cannot process a permission revocation, a disabled notification or a cancelled subscription without recording it; assembling those into a rate is a query and a definition, not a project.
A session-origin distribution is a project. It needs a controlled prompt taxonomy, a link from each delivery back to the authorisation that created it, cross-surface identity, notification exposure states, a versioned attribution window, and an honest way to report the cases it cannot classify. Some products will have most of that already. Many will not.
So the willingness argument holds for one of the two measures. For the other, the honest answer is that it costs real engineering, and the first company to do it will be building something rather than merely disclosing it.
Why publishing first is worth something — and what it takes
The instinct is that publishing a falling number is a pure cost. It is not, but the reason is narrower than the v1.0 page claimed.
A commitment is credible in proportion to how expensive it would be to fake. Any company can say it puts users first; the sentence costs nothing, which is exactly why nobody believes it. But a self-defined number, published once, is also cheap — the definition can move, an unflattering quarter can quietly not appear, and nothing follows from misstating it. First publication on its own does not separate a trustworthy product from an untrustworthy one.
What creates the separation is the apparatus around the number: a definition fixed in advance and versioned when it changes, unknowns reported rather than absorbed, bad periods left visible, corrections preserved with dates, and eventually assurance by somebody who is not the publisher. Those are expensive for a company that would need to hide something and cheap for one that would not. That is the costly signal, and a competitor cannot neutralise it by saying anything — only by building the same apparatus and publishing their own numbers.
It is also cheaper than it looks in one specific way: the first number does not have to be good. Nobody has a baseline, so there is nothing to be worse than. Whoever publishes first defines the scale and how it is computed. That window closes the moment somebody else goes first.
The falsifiable version
This piece makes a claim that can be shown wrong, and the way to show it wrong is to point me at a product that publishes either measure. I will correct it here and say who.
The replacement measure can fail too, in at least four ways: independent classifiers may not assign the origin classes consistently from the same evidence; real product schemas may be unable to establish origin without unacceptable missingness; the classes may fail to separate cases that need different decisions; or a simpler established measure may answer the question with less burden. Any of those should narrow, revise or retire it. The specification states them in full.
The prediction recorded in v1.0 stands, adjusted to the new terms: the first of these two to be published anywhere will be a session-origin figure, not a withdrawal rate — because origin can be framed as a quality metric a confident product might be proud of, and withdrawal cannot be framed as anything except a cost.
How this was arrived at
Six assistant surfaces, chosen to span business models rather than sampled from any defined population, scored against published anchors using public documentation and named, dated reporting. Disclosure only: what a company says, never what it privately does. “No disclosure found” is a finding about the public record, not an accusation about conduct.
The scored, named version of that review is not published, and may not be. This — the part about the category rather than about any company in it — is the half worth having.
Whether the rubric behind it produces consistent scores in other people’s hands is a fair question and is pre-registered to be measured, with the threshold fixed in advance and a commitment to publish a failure. Consistency is not the same as validity: neither the scorecard’s total nor the measures proposed here have been shown to measure what they are named for.
Every substantive change to a published claim on this site is dated and logged at the corrections record. This page is the first entry.
Sources
1. OpenAI, “Scheduled tasks in ChatGPT”, official Help Center, retrieved 31 August 2026.
2. Google, “Schedule actions in Gemini Apps”, official Gemini Apps Help, retrieved 31 August 2026.
These are criterion 2 and criterion 7 on the scorecard — the two that ask what a product knows about where its own usage came from, and about reliance being taken back. Both anchors were revised on 31 August 2026 to match this correction.
Score your own product against this. The Conferral Design Scorecard turns the whole argument into ten criteria with published anchors — free, and nothing you enter leaves your browser. It is a structured comparison, not a measure of trustworthiness. The complete methodology behind it is at Conferral by Design.
The other essays: Sycophancy is a counterfeit deposit · Session time is an anti-metric for agents · Machine trust will be allocated by contest · An ad in a list is not an ad in an answer · The same task, two designs (runnable) · Where these ideas come from