Introduction
Product managers today feel constant pressure to add AI to every new feature. Yet the conversation rarely moves beyond whether AI can perform a task to what happens when it gets things wrong. This article introduces a practical decision tree helping product teams evaluate AI-driven features before they ship.
What Happened
The discourse around AI in products has split into two extremes. One side publishes governance frameworks and ethics principles that matter but rarely answer whether a specific feature should launch. The other shares demos that make every capability look shippable, obscuring real trade-offs. Neither gives product managers a repeatable method to assess risk cost or user impact. Drawing from years in fintech product and engineering the author reframes the problem: the interesting question is whether a model can perform a task but what occurs when it errs and how that failure compares to the human process it replaces. Automation does not necessarily reduce the error rate it reshapes its shape. A tired analyst makes isolated mistakes that can be questioned. A miscalibrated model can produce thousands of identical errors before lunch with no one to ask why. This insight forms the core of the decision tree that follows.
Why This Matters
To help teams decide where a feature lands on the risk spectrum the framework poses three sequential questions followed by a veto question that can halt a feature regardless of its position. Question 1: Can the user see and correct the mistake? When a user can inspect the output and fix it before anything takes effect the cost of an error collapses. AI drafting an email is the canonical example the model writes the user reads the user edits or deletes and the message sends. The user catches the mistake because they have the most context and incentive. What looks like roadmap friction is actually a free error-correction layer. Contrast this with AI silently reordering a merchants payout schedule or auto-archiving support tickets. If the user cannot see the decision they cannot correct it and errors accumulate quietly until they surface as churn or audit failures. Question 2: Is the mistake reversible? Reversibility varies widely. A wrong content recommendation is easily shrugged off after two seconds of mild annoyance. An auto-reply that missed the mark can be corrected with a follow-up message. But in payments money movement is irreversible. Ledgers update funds leave accounts and third parties have no obligation to reverse the transaction. Features that delete data cancel accounts or push funds across borders often fall into this irreversible category. Reversibility is also a design choice adding a holding period or staged execution step can convert a permanent action into an undoable one frequently delivering more value than trying to improve the model itself. Question 3: Will someone demand an explanation? Some decisions carry an implicit or explicit requirement for justification. A denied loan application a benefit rejection or a fraud flag that freezes funds all trigger an expectation of explanation. Simply stating the model scored it below threshold satisfies no one and may violate regulatory obligations. In lending and regulated industries the explanation requirement functions as a hard functional requirement comparable to a latency budget and its specifics vary by jurisdiction in ways the model cannot anticipate. If a feature carries explanation debt the bar is not the AI is usually right. The bar becomes we can show our work every time to someone hostile. Running a feature through these three questions sorts it into one of three buckets. AI decides: The model output lands in front of a user who cannot correct it no one demands an explanation and the mistake is forgettable. Think recommendation engines surfacing the wrong article search ranking adjustments or summarization. The output is visible the error is easily dismissed and no one escalates. Features in this bucket can ship quickly with measurement as the primary post-launch activity. AI suggests human decides: The model performs the heavy lifting flagging a suspicious transaction drafting an email suggesting a refund amount while a human reviews and owns the final call. The division of labor is clear AI handles volume the human owns the commit and the defense of that commit if questioned. This design captures most of the value of automation while preserving a safety net. It is also typically the cheaper path because the human review layer means the feature does not need the largest most expensive model to clear its bar. AI never touches it: Features that involve irreversible actions missing explanation paths or no user recourse fall here. Auto-rejected loan applications accounts closed overnight with no appeal refunds issued at scale with no hold. A better model will not fix the fundamental risk these features pose. Moving a feature out of this bucket requires changing reversibility or explanation design not upgrading the model. The framework also introduces a fourth question the veto. It asks what does the feature cost when the model is right? Traditional software has near-zero marginal user cost. AI does not. Every query incurs a bill. A feature can pass all three earlier questions safe reversible explainable and still die because its unit economics never close. Token prices may have fallen but cost per completed task has risen especially in agentic workflows that chase answers with multiple failed attempts. The cost gap traces back to whether the operator can distinguish a correct output from a confident one both on the way in and on the way out. AI suggests human decides is positioned not as a compromise but as the better product choice. It captures most of the value of automation while protecting the one thing full autonomy cannot easily give back: the perception that a decision was taken away from the model and handed to a human which users experience as a downgrade. Many features should graduate from suggests to decides as evidence accumulates. Plenty never should and that is acceptable.
Key Takeaways
- The capability bar is useless because models can almost always perform the task. The real bar is what the feature costs when it is wrong.
- Three questions user correction reversibility and explanation demand sort features into three buckets AI decides AI suggests with human oversight or AI never touches it.
- The veto question cost when the model is right often kills features that pass the first three tests. Token economics and per-task cost matter as much as accuracy.
- Human-in-the-loop designs are usually the stronger product choice and the cheaper one because they capture automations value while preserving a safety net and user trust.
- Features that cannot justify their cost per completed job should not ship regardless of how capable the underlying model appears.
Conclusion
The next time a roadmap review asks where is the AI the answer should not be a list of capabilities. It should be a structured assessment of error cost reversibility explanation requirements and unit economics. By sorting features through the decision tree product teams can defend their sorting veto what cannot pay for itself and ship AI features that survive their second year not because they were bold but because they were priced.




Discussion
Join the conversation
Thoughtful reactions, questions, and follow-up ideas help shape the next story.