Shipping
Adoption Is Not an AI Value Metric

While helping operate a portfolio of dozens of internal products, I learned how quickly a roadmap update can become a catalogue of activity. We could count launches, processed work, active users, and blockers. Those numbers were useful, but they did not answer the decision the update needed to support: which products were creating enough value to justify more investment?
The measurement work became a product problem of its own. We built an internal framework that considered throughput, cycle time, output, capacity, and the assumptions behind each calculation. Just as importantly, we kept validated measures separate from estimates that still needed review. The exercise made one gap impossible to ignore: usage could describe exposure to a product without proving that someone finished the job they came to do.
Adoption is evidence that people used the feature, but it is insufficient evidence of user or business value. An AI product needs a metric tree that connects output quality to task completion, human effort, downstream outcome, risk, latency, and cost. If those links are missing, a rising usage chart can hide a product that creates more work than it removes.
Adoption and value can move in different directions
Recent research makes the gap visible at the firm level. A February 2026 NBER working paper on firm-level AI use surveyed almost 6,000 executives across four countries. Around 70 percent of firms reported active AI use, while more than 80 percent reported no impact on employment or productivity over the prior three years. The study's unit is a firm, not an individual feature, and the survey cannot explain why a particular product did or did not create value. I use it only as context for questioning the shortcut between reported use and reported impact.
A more focused field study shows why measurement design matters. In Generative AI at Work, researchers studied 5,172 customer-support agents and defined productivity as resolutions per hour. That measure combined speed with successful resolution. They also tracked handle time, resolution rate, and customer satisfaction. A count of AI suggestions accepted would have described interaction with the tool. It would not have captured whether customer problems were actually solved.
The lesson is not that adoption metrics are useless. They answer important questions about discovery, access, and repeat use. The mistake is promoting them to a value metric before proving the causal path underneath them.
The useful unit is a successful outcome
For a conventional workflow, teams can often count completed transactions. AI complicates the picture because the system can produce an answer that looks complete while shifting verification and correction to the user.
That is why I would anchor the product around cost per successful outcome:
Total model, infrastructure, review, and correction cost divided by the number of qualifying user outcomes.
The denominator does most of the intellectual work. A team must define what qualifies as success in observable product terms. “Generated a response” is usually too weak. “The user accepted a response” may still be weak if acceptance is a one-click action with no downstream signal.
For a career-writing tool, a stronger outcome might be: the user creates an achievement statement, reviews it, makes no material factual correction, and exports it for a resume or performance review. That is a proposed definition, not a claim about any live product. The right definition will vary by task, and it should change when the user or business outcome changes.
The AI outcome metric tree
The following artifact is a proposed operating model. It is designed to expose where apparent value can leak out of an AI workflow.
The AI outcome metric tree
- Successful user outcomeDefine the job the user completed, specify the observation window, and exclude outputs needing a material factual correction.
- Workflow completionTask completion rate, time to complete including review, abandonment and retry, escalation to a human.
- Output fitnessGrounded correctness, relevance to stated intent, required edits before use, and abstention when evidence is missing.
- Downstream valueThe user action that follows, the business outcome connected to it, and a comparison against the prior workflow.
- Risk and trustCritical privacy or safety failures, unsupported claims, uneven performance across slices, user corrections and reversals.
- Delivery economicsCost per attempt, review and correction cost, median and tail latency, and cost per successful outcome once earlier gates pass.
This tree is intentionally wider than a single north-star number. The NIST AI Risk Management Framework Core recommends selecting metrics for the most significant mapped risks and documenting human oversight. Product measurement should carry those constraints beside value and cost, instead of treating them as a separate compliance dashboard.
A synthetic dashboard can reveal a bad win
Consider a tabletop example for an achievement-writing assistant. The inputs and thresholds below are synthetic. They demonstrate the method and say nothing about Bragora's production performance.
Suppose weekly generations increase by 40 percent. At the same time:
The adoption chart says growth. The outcome tree says the product is shifting work into review and correction. Lower model cost does not rescue the result because the relevant unit is the successful outcome, not the attempt.
The product response would be equally specific: inspect correction themes, improve or constrain the weakest task slices, and test whether completion time and material-correction rate recover. Marketing the higher generation count would be premature.
A metric tree is still a hypothesis
There is a reasonable counterargument: downstream outcomes can take weeks or months, depend on factors outside the product, and arrive too late for daily decisions. A job-search tool cannot claim credit for an interview without accounting for role fit, labor-market conditions, and the rest of the application.
That limitation is real. Use a chain of evidence with different speeds:
- Leading signals: output fitness, correction burden, and task completion.
- Behavioral outcomes: export, reuse, or another task-specific action.
- Lagging outcomes: the real-world result, measured with an explicit attribution limitation.
I would change my conclusion if a team could show that adoption alone consistently predicts a defined downstream outcome across new user cohorts and important task slices. Until then, adoption belongs near the top of the funnel, not at the center of the value story.
Bring evidence to the metrics review
Before approving more investment in an AI feature, require two artifacts: a written definition of the completed user outcome, and an evidence chain showing where review, correction, risk, latency, and cost enter the result.
Then put the weakest link on the agenda. Which connection between use and downstream value remains an assumption? If the team cannot name it, the dashboard is measuring activity without showing where value could disappear. That may be enough for discovery. It is not enough for an investment decision.


