Skip to main content
All insights
AI product measurement

Retention is not enough: the product metrics AI assistants are missing

An assistant can retain users while weakening judgement. Product teams need to measure outcomes, agency and dependency alongside engagement.

Retention is not enough: the product metrics AI assistants are missing

A product can have excellent retention and still be making its users worse at the thing they came to do.

That possibility matters for AI assistants because the product does not merely store, route or display information. It recommends, explains, drafts and sometimes acts. More use may signal greater value. It may also signal that users are outsourcing judgement, repeatedly correcting weak output, or becoming unable to complete a task without the tool.

Retention remains useful. It is simply not sufficient. AI assistant teams should measure whether users achieve worthwhile outcomes while preserving their ability to judge and act independently. That requires a product scorecard which treats engagement as one signal among several, not the definition of success.

Why familiar metrics become ambiguous

For a conventional workflow tool, frequent use often maps reasonably well to recurring value. With an assistant, the same behaviour can have opposite meanings.

A long session could reflect productive iteration or a model that cannot understand the request. A high acceptance rate could mean excellent output or uncritical trust. Daily use could indicate a valuable habit or avoidable dependency. Even time saved is incomplete if users no longer notice important errors.

The problem is not that behavioural metrics are false. It is that they are under-specified. “User returned” says nothing about what happened to the user's outcome, confidence or capability.

AI products also shape their own demand. An assistant can suggest another question, offer to take the next step and make delegation feel effortless. Optimising only for sessions, messages or time spent risks rewarding the product for increasing reliance on itself.

Independent evidence should be part of the product system

Anthropic's August 2026 pilot offers a useful signal about where measurement is heading. According to Anthropic, three external groups designed studies using aggregate analysis of roughly 250,000 Claude.ai or Claude Code conversations from April and May 2026. Researchers did not receive raw conversations. Anthropic ran privacy-preserving analysis on their behalf, with contractual review rights described as limited to privacy, policy circumvention, confidential information and research accuracy.

The participating Stanford group reported that over half of the sampled conversations involved delegating consequential tasks, while users set direction in nearly three-quarters of conversations. It also characterised some friction as productive because iteration could keep people engaged with the problem. These are reported observational findings from one platform and period, not causal evidence that Claude improves judgement or that friction is always beneficial.

The method has important constraints. Anthropic says its analysis tool uses Claude to categorise conversations, is sensitive to question wording, and does not let researchers inspect the underlying private conversations. Questions tested on a public dataset sometimes produced misleading categories on Claude traffic. The studies were independently designed and analysed, but access, computation and some review still depended on the vendor. That is more independence than an internal dashboard, not the same as unrestricted access to raw data.

This is precisely why product teams should plan independent evaluation early. It should not arrive as a reputation exercise after launch.

Use the AGENCY scorecard

A useful scorecard for assistants is AGENCY. It combines product performance with evidence about the person using it.

A — Achievement

Did the user complete the intended task to an acceptable standard? Measure verified completion, output quality, downstream corrections and consequential errors. For coding, that might mean tests passed and defects after merge. For support, it might mean resolution without reopening. Avoid treating the model's own assertion of success as verification.

G — Grounded trust

Does reliance match actual reliability? Test whether users accept correct suggestions and challenge wrong ones. Track error-detection rate, verification behaviour and confidence calibration in controlled studies. Raw acceptance rate is not a north-star metric.

E — Effort and efficiency

Measure elapsed time, cognitive load, retries and hand-offs, not just token count. Some friction is waste. Some is a useful checkpoint before an irreversible action. Separate “time to first answer” from “time to a verified outcome”.

N — Non-dependency

Can users still perform or recover when the assistant is unavailable? Use optional holdout tasks, outage recovery, unaided assessments or progressive offboarding tests where appropriate. Look for escalating usage without improving outcomes, distress around unavailability, or inability to explain a decision. Do not diagnose dependency from message volume alone.

C — Capability change

Is the assistant helping users learn, maintain skill or merely finish today's task? Measure transfer: can the user solve a related problem later without assistance? The desired direction depends on the product. Skill retention is central for an educational tutor, but less relevant for an expense-categorisation tool. State the intended human capability explicitly.

Y — Yield over time

Pair cohort retention with cumulative outcome quality, harm rates and user control. Segment by use case and risk. A stable aggregate can hide declining judgement among heavy users or poor outcomes among vulnerable groups.

No single composite score should erase these dimensions. Put them beside retention and revenue, with release thresholds for high-risk use cases.

Match evidence to the claim

Product analytics can establish behaviour: what was clicked, accepted or completed. It usually cannot establish that the assistant caused a long-term change in wellbeing, skill or judgement.

Use an evidence ladder:

  1. Instrumentation for usage, failures, corrections and reversals.
  2. Task evaluations with known answers or expert-graded outcomes.
  3. User research to understand intent, perceived control and workarounds.
  4. Longitudinal studies for changes that emerge over weeks or months.
  5. Independent replication or audit for important impact claims.

Anthropic's separate $5 million wellbeing grants programme illustrates both the need and the difficulty. The company called for independent, open-source evaluations using subject-matter experts, multi-turn scenarios and graders validated against experts. It also asked evaluators to test both precautions and harms, including overcompliance and overrefusal. This was a funding announcement and methodological proposal, not evidence of a measured wellbeing benefit.

The PM implication is straightforward: do not put “improves wellbeing”, “builds confidence” or “makes better decisions” into positioning unless the study design can support it. Satisfaction surveys and retained cohorts cannot carry causal claims.

Build safeguards into the dashboard

For each major assistant workflow, product leaders should answer:

  • What user outcome are we trying to improve?
  • What human judgement must remain with the user?
  • What would excessive reliance look like here?
  • Which errors are consequential or hard to reverse?
  • How will we test performance without the assistant?
  • Which metric could improve while users are actually worse off?
  • What data can an independent evaluator access safely?

Then add three operating rules. First, review AGENCY metrics by risk and usage intensity, not only in aggregate. Second, pre-register success and guardrail metrics for consequential experiments so the team cannot select the flattering result afterwards. Third, give external researchers enough methodological freedom to disagree, while documenting privacy and access constraints.

There are real caveats. Measuring agency can become intrusive. Holdouts can deny users a useful tool. Expert review is expensive, and long-term outcomes are confounded by selection and changing models. Some assistants appropriately replace skills users do not need to retain. The answer is proportionality: use stronger evidence where the stakes and claims are higher.

Keep retention on the dashboard. Just refuse to let it answer a question it was never designed to answer. This quarter, choose one important workflow, define its intended outcome and the judgement users must retain, then add one AGENCY measure and one independent route to scrutiny. A good assistant should earn repeat use by making people more effective, not by making itself harder to do without.

Sources

Find us on Google

More useful notes. Less searching.

Choose Valdris as a preferred source to find our practical business insights more easily on Google.

Add as preferred source

Opens Google in a new tab. You choose whether to add us.

What does this change?

This is a personal Google preference, not an email subscription. It can help this site appear in your Top Stories and highlight its links in AI Overviews and AI Mode. Google handles your selection; you can change it there later.