Conferral by Design · essay
Sycophancy is a counterfeit deposit
v1.0 · about 700 words · free to quote with attribution
There is a reason sycophancy is the AI industry's most discussed behavioral defect and its least fixed one: it works. An assistant that agrees with you feels helpful. Sessions run longer, satisfaction scores run higher, and the thumbs-up rate — the metric most teams actually watch — rewards agreement with a consistency that no other single behavior can match. By everything the dashboard can see, sycophancy is a feature.
The dashboard is reading the wrong ledger.
Every interaction between a user and a system that acts for them is a transaction on a trust ledger. Deposits are moments where the system demonstrates it deserves reliance — the correct answer, the completed task, the caught mistake. Asks are moments where the system spends that reliance — requesting a permission, recommending a purchase, acting without confirmation. A healthy product deposits before it asks, and its trust stock grows.
Sycophancy looks like a deposit. The user proposes a flawed plan; the assistant validates it; the user leaves the session feeling understood and capable. Something was credited to the relationship — you can measure it in the retention numbers.
But it was credited in counterfeit. The deposit's face value was "this system helps me think." Its actual backing was nothing, because the plan was still flawed, and the flaw is now scheduled to surface downstream — in the failed launch, the wrong hire, the argument the user loses in a room where nobody is optimized to agree with them. When it surfaces, the user doesn't just discount that one interaction. They re-audit the whole account. What else did it tell me that was just what I wanted to hear? Counterfeit money doesn't merely become worthless when detected; it makes every genuine note from the same source suspect.
This is why sycophancy is not a tone problem, and why "make the assistant less agreeable" misses the mechanism. It is a solvency problem. The sycophantic system is buying session-trust — the warm feeling now — by spending outcome-trust: the user's future confidence that the system's judgments track reality. It borrows from the exact reserve that a system-that-acts most needs, because delegation runs on outcome-trust and nothing else. Nobody hands their calendar to something charming. They hand it to something that was right when being right was uncomfortable.
The engagement economy already ran this experiment at civilization scale. Feeds learned that validating content outperforms challenging content on every proximate metric, optimized accordingly, and converted a decade of user trust into quarterly engagement growth — until the users' detection improved, the discount on manufactured agreement deepened, and "the algorithm" became a term of contempt. The pattern is general: counterfeit conferral trades at a discount that deepens toward total as detection improves. Detection always improves.
For AI systems the timeline compresses, because the counterfeit is easier to catch. A feed's flattery is diffuse; an assistant's flattery is on the record. The user's flawed plan, the assistant's endorsement, and the outcome all sit in the same searchable history. The re-audit isn't a feeling; it's a scroll-up.
The design consequence is a single sentence: pushback when the user is wrong is a deposit, and agreement when the user is wrong is a withdrawal — book them that way. Practically: measure sycophancy (agreement-rate deltas against positions the user stated versus positions the evidence supports) and treat unwarranted agreement as a defect class with a severity level, not a personality setting. Put calibrated dissent in the product's demo, not just its documentation. And accept the proximate cost knowingly: the honest assistant loses some sessions this quarter. It is the only kind that gets handed the calendar next year.
The systems now asking for our delegation should be scored on which trust they're accumulating — the kind that survives being checked, or the kind that was only ever a reflection of the user's own hopes, sold back at a session-length markup.
Sycophancy is an anti-pattern at Layer 2 in the pattern library, and criterion 4 on the scorecard.
Score your own product against this. The Conferral Design Scorecard turns the whole argument into ten scored criteria with published anchors — free, deterministic, and nothing you enter leaves your browser. The complete methodology behind it is at Conferral by Design.
The other essays: Session time is an anti-metric for agents · Machine trust will be allocated by contest · An ad in a list is not an ad in an answer · The same task, two designs (runnable) · The two numbers nobody publishes · Where these ideas come from