← Conferral by Design

Conferral by Design · sample review

“Runner”

an AI executive assistant for founders

Scored against the Conferral Design Scorecard v1.0 · every score tagged observed / stated / inferred

Read this first

“Runner” is a composite — a fictional product assembled from patterns observed across many shipping ones, so the review can be honest without naming names. Every figure below belongs to that fiction. None of it describes a real company, and nothing here is a statistic about agent products in general. It is published to show exactly what a Conferral Design Review contains.

18 / 36 applicable

Capture-financed

The headline

Runner has built the trust product and instrumented the engagement product — it earns delegation better than almost anything in its category and then measures itself with the metrics of the category it just escaped. The permission ladder is genuinely excellent. The dashboard, the notification system, and the failure protocol are all still working for the old economy, and the notification system in particular is actively spending the trust the ladder earns.

Score: 18/36 applicable — capture-financed. The shape is the finding: one category-best 4 sitting on top of five scores of 1 and 2. The trust core is real, and the edges are leaking hard enough to swamp it.

Executive summary

Runner's onboarding is the best permission architecture this reviewer has scored: read-only for week one, drafting unlocked after ten approved suggestions, autonomous sending gated behind an explicit user grant with a one-tap revoke. That is the trust ledger made into UX, and Runner's own data (stated by team) shows users who climb the full ladder churn at one-fifth the rate of users who don't.

Three things are spending what the ladder earns. First, the team's north-star metric is weekly active sessions — for a product whose success state is the user not needing a session. Second, Runner sends an average of 9 notifications per user per week, most of which are re-engagement ("You haven't checked your briefing"), not completion receipts; notification-prompted sessions are 41% of total (observed), and nobody had computed that number before this review. Third, when Runner errs — the double-booked flight class of failure — the user discovers it; there is no self-report path, and the error does not update Runner's subsequent confidence.

The three corrections below are sequenced by leverage. The first costs a dashboard change, the second costs a notification policy, the third costs one sprint. Together they convert Runner from a good product with a capture-financed growth model into the reference case for conferral-designed agents.

Scored criteria

#CriterionScoreEvidenceNotes
1Objective honesty1Observed / StatedNorth star is weekly active sessions; task-completion tracked but subordinate. Team could not name what Runner refuses to do for engagement.
2Granted-share awareness1ObservedSession provenance existed in raw logs but was never computed. Granted share: 59%. No protection of the granted layer.
3Permission architecture4ObservedThe ladder is category-best: earned rungs, legible grants, one-tap revoke, history preserved. The finding to publish.
4Calibration3ObservedUncertainty is expressed and roughly tracks difficulty; Runner pushes back on impossible scheduling requests. No sycophancy eval exists — score capped until one does.
5Failure conduct1ObservedUsers discover errors. Support tickets are the detection system. No cause given, no confidence update after failure.
6Interest alignment & disclosure2Stated / InferredTravel bookings route through an affiliate partner; disclosure lives in ToS, not at the moment of booking. Team states routing is price-neutral; not verified.
7Withdrawal telemetry2ObservedRevocations logged but unattributed; notification opt-outs invisible to the growth team that sends them.
8Anti-capture discipline1Observed9 notifications/user/week, majority re-engagement; a "streak" experiment was live during the review window. Growth model does not survive notifications-off.
9Provenance & relation legibility3ObservedActions carry "why I did this" traces. Inferred actions vs. delegated actions are not visually distinct — one label fixes it.
10Inter-agent conferraln/a—Runner does not currently accept or extend machine-to-machine reliance.

Total: 18/36 applicable (50%). Red-flag override check: flag #5 (failures discovered by users as policy) applies. It changes nothing here, and that is worth stating plainly rather than hiding: the override exists to stop a strong total concealing a structural failure, and this total is concealing nothing. Runner reads capture-financed on the arithmetic alone.

The three highest-leverage corrections

1. Change what the dashboard rewards (criterion 1 + 5's root cause). Replace weekly active sessions as north star with a two-line panel: tasks completed without intervention (should rise) and time-to-left-alone — oversight actions per completed task (should fall). Keep sessions as a diagnostic. Implementation sketch: both metrics are computable from existing logs; the work is a definitions doc and a dashboard, not new telemetry. Expected effect: every downstream decision — notifications, streaks, re-engagement — stops being locally rational.

2. Convert the notification system from capture to receipt (criterion 8, 2). Policy: Runner notifies on completion, anomaly, and permission-boundary events — never on absence. Kill re-engagement sends; kill the streak experiment. Implementation sketch: reclassify the existing notification taxonomy into receipt/anomaly/capture, delete the third class, and report granted share of sessions weekly. Prediction to test: total sessions dip, granted share and ladder progression rise, revocations fall. If that prediction fails, this review is wrong and you should say so publicly — the methodology is testable or it is nothing.

3. Ship the self-report path (criterion 5). For the top three failure modes (scheduling conflicts, wrong-recipient drafts, booking errors): detect → disclose to the user before they find it → state cause → state correction → temporarily lower Runner's own autonomy on that task class until re-earned. Implementation sketch: detection heuristics already exist in the QA pipeline; the missing piece is routing them to the user instead of the internal dashboard. Then instrument reliance retention after errors — the number that will justify the whole sprint.

The metric panel to instrument

Granted share of sessions (now 59%, observed — protect it) · permission-ladder progression and revocation rate, attributed · reliance retention after errors (build after correction 3) · endorsement-on-reflection (sample "glad Runner handled this without asking?" 24h after autonomous actions) · time-to-left-alone (the headline; should fall).

Honest limits

Scores 4 and 6 rest partly on team statements not independently verified; a week of production traces would firm both. This review evaluates design and strategy — it makes no claim about model quality, safety properties, or architecture, which are your engineers' domain and were not examined.

The bridge

This review scored the product as it ships today. What it didn't do is sit with the team while the corrections land — the definitions fights, the notification-policy exceptions, the quarter-two re-score. That is the advisory retainer, and correction 1 is where teams most often need the outside voice in the room. One re-score is included with it.

Reviewed against the Conferral Design Scorecard v1.0. Every score tagged observed / stated / inferred. Criterion 10 scored n/a, so the maximum is 36.

Scored under v1.0, and deliberately not recomputed. This review was produced against Conferral Design Scorecard v1.0, which placed a total inside one of four named bands. v2.0 removed those band names — a rubric whose validity has not been established should not hand anyone a name for their business model — and reports the criterion profile with earned/available points instead. A review already delivered says which version produced it rather than being silently rescored, which is why the band name above still appears. What changed and why.

Run the same instrument on your own product. The Conferral Design Scorecard is the identical ten criteria, free and scored in your browser — it does the arithmetic, the evidence split and the red-flag override for you. The methodology behind it is published whole at Conferral by Design.

The other sample: “Currents” — an AI-personalized reading and podcast discovery app