← Conferral by Design

Conferral by Design · essay

Session time is an anti-metric for agents

v1.0 · about 800 words · free to quote with attribution

Somewhere right now, a team building an AI agent is celebrating a chart. Weekly sessions are up. Time-in-product is up. The retention curve has that shape investors screenshot. And every number on that chart is measuring the product's failure.

Here is the test. Ask what the perfect version of the product looks like — the version so good that every user's ideal outcome is realized. For a social feed, entertainment app, or game, the perfect version plausibly involves more time: the value is the attention itself. For an agent — a system a user delegates tasks to — the perfect version is nearly invisible. It receives the task, completes it correctly, reports in a sentence, and recedes. The user checks in briefly, trusts the report, and goes back to their life. Perfect agent, minimal session. If your product's success state is the user leaving, then time-in-product is not a KPI. It is an anti-metric: a number whose growth should worry you.

Why do agent sessions actually run long? Audit the minutes and they decompose into three parts, none of them value. Supervision — the user watching the agent work because they don't yet trust it unwatched. Correction — the user fixing, redirecting, re-prompting. Capture — the user pulled back by a notification, a streak, a re-engagement nudge, because someone was goaled on the chart above. Supervision is trust not yet earned. Correction is reliability not yet delivered. Capture is the engagement economy's old reflex, pasted onto a product category it actively damages. A rising session-time line for an agent means one of these three is growing — and the dashboard is applauding.

The metric confusion isn't cosmetic; it compounds, because teams build what dashboards reward. Goal an agent team on engagement and they will — each decision locally reasonable — add the daily-summary notification, the streak, the chatty confirmation loop that turns a completed task into a conversation. Every one of those manufactures presence. And presence is precisely what delegation is purchased to eliminate. A user who wanted to spend time managing their inbox already had an inbox. The product ends up optimized against its own reason to exist: maximizing the supervision its success was supposed to remove.

What should the chart show instead? The metrics that fit systems-that-act are mostly inversions of the engagement panel:

  • Granted share of sessions — the fraction the user initiated on purpose, versus arrivals prompted by notification or streak. High and rising is health. (A session count that grows while granted share falls is capture, not growth.)
  • Tasks completed without intervention — the deposit line.
  • Permission-ladder progression against revocation rate — are users granting more autonomy over time, and how often do they take it back? Revocation is the drawdown line, and it should be attributed to the event that caused it.
  • Reliance retention after errors — does honest failure-handling preserve the account?
  • And the headline: time-to-left-alone — oversight per task, trending down as trust accrues. The one chart where down is the celebration.

Notice what these numbers have in common: they measure the accumulation of trust rather than the consumption of attention. That is the whole correction in one contrast. Engagement metrics read the income statement — attention extracted this period. Conferral metrics read the balance sheet — reliance earned and retained. The engagement economy ran a decade of record income statements into a solvency crisis because nobody was reading the other document. Agents, holding calendars and inboxes and payment methods, do not get a decade of float.

There is a genuinely hard version of the objection: "we can't take less usage to our board." True — which is why the reframe matters more than the metric. An agent's expansion story was never minutes; it is scope: more task types delegated, higher autonomy rungs granted, reliance retained through errors, users who came back after being left alone because leaving them alone is what worked. Those curves grow for years. Session time for an agent peaks early — at maximum distrust — and its decline is the revenue signal, because the user who no longer supervises is the user about to delegate more.

The success state of an agent is being trusted enough to be left alone. Build the dashboard that can see that, or the dashboard will quietly rebuild the product into something that can't be.

This is Principle 5, it drives criteria 2 and 5 on the scorecard, and it is the metric panel in both sample reviews.

Score your own product against this. The Conferral Design Scorecard turns the whole argument into ten scored criteria with published anchors — free, deterministic, and nothing you enter leaves your browser. The complete methodology behind it is at Conferral by Design.

The other essays: Sycophancy is a counterfeit deposit · Machine trust will be allocated by contest · An ad in a list is not an ad in an answer · The same task, two designs (runnable) · The two numbers nobody publishes · Where these ideas come from