Shipping
Faster Prototypes Need Better Evidence

Generative AI has made it possible to turn a rough product idea into screens, copy, code, and a clickable flow before the next meeting. That is real leverage. It is also an easy way to make weak assumptions look finished.
A polished prototype creates momentum. Stakeholders can see it, react to it, and imagine it in production. But visual completeness is not evidence that the problem matters, that the workflow fits the user's context, or that the system can deliver the promised outcome reliably.
My product rule is this: AI should reduce the cost of testing an assumption, not reduce the standard of evidence needed to make a decision. The faster the artifact appears, the more deliberate the team must be about what it proves.
Artifact speed and learning speed are different
The value of a prototype is not the amount of interface it contains. Its value is the uncertainty it removes.
An AI-generated onboarding flow might help a team compare two information architectures in an afternoon. It cannot show whether new users understand the product promise unless people from the intended audience use it. A generated dashboard might expose missing states or unclear hierarchy. It cannot prove that the underlying data exists, arrives on time, or supports the decision shown on screen.
First-party research on GitHub Copilot reported faster completion in a bounded coding task. That is useful evidence for that task and setting. It is not evidence that AI automatically improves product discovery, roadmap quality, or time to market. Producing code faster and learning the right thing faster are different outcomes.
The distinction protects teams from a common failure mode: building a high-fidelity answer before agreeing on the question.
Start with the assumption, not the tool
Before generating anything, write the assumption the prototype is meant to test. A useful statement names the user, context, behavior, and consequence.
Weak assumption
A claim nothing can disprove
“Users will like an AI career coach.” Nothing here names a user, a context, a behavior, or a consequence, so no prototype can fail it.
Stronger assumption
A claim a prototype can actually fail
“Professionals preparing for a performance review can verify and reuse an AI-drafted achievement faster than writing one from scratch, without accepting unsupported claims.” It requires source evidence, editing, rejection, and recovery.
The stronger assumption changes the prototype. It requires source evidence, editing, rejection, and recovery. A beautiful chat screen without those states would test presentation while skipping the risky part of the product.
This is why prompts cannot substitute for discovery. AI can generate plausible flows for missing requirements. Plausibility makes the gap harder to notice.
The Prototype Evidence Card keeps the artifact honest
I would use a five-part Prototype Evidence Card before a prototype enters a roadmap discussion:
The Prototype Evidence Card
- Assumption: which belief about the user, workflow, feasibility, or value is being tested?
- Fidelity: which parts behave realistically, and which parts are simulated or absent?
- Evidence plan: who will use it, what task will they attempt, and what will be observed?
- Decision rule: which result supports, changes, narrows, or stops the idea?
- Discard rule: when does this artifact expire so it cannot quietly become production code?
The fidelity field matters because AI makes simulation cheap. A prototype may use fabricated data, a human behind the scenes, or a fixed response disguised as model behavior. Those choices can be legitimate when the team states them. They become dangerous when a stakeholder mistakes simulated capability for tested capability.
The discard rule matters for a different reason. Prototype code is optimized for learning. Production code needs maintainability, security, accessibility, observability, and recovery. Keeping the artifact can be appropriate, but that should be a new engineering decision rather than the default reward for a convincing demo.
Evaluation belongs inside the prototype plan
The NIST AI TEVV program frames evaluation around explicit tasks, metrics, testbeds, methods, and known limitations. The NIST Generative AI Profile adds lifecycle concerns such as confabulation, privacy, testing, and human oversight. Neither source prescribes a product-prototyping process. Together, they reinforce a useful constraint: an AI demo needs a defined method for deciding whether it worked.
For variable outputs, test more than the happy path. Current evaluation guidance recommends typical, edge, and adversarial cases with human calibration. The vendor tooling may change, but the product principle holds.
If the prototype drafts achievement statements, include a clear note, an ambiguous note, a contradiction, missing evidence, and a request to invent impact. Observe factual preservation, required edits, abstention, time to completion, and whether the user understands what the system did.
The result is not one usability score. It is a map of where the product is ready, where the scope should narrow, and which uncertainty requires another test.
AI can improve roadmap conversations without writing the roadmap
Generative tools are useful for exploring consequences. A PM can ask for alternative sequencing, dependency questions, risk scenarios, or counterarguments to a plan. That can widen the option set before the team commits.
The output should enter the roadmap as a proposal with provenance, not authority. Which source informed it? Which capacity or dependency assumption did it make? Which stakeholder has not been represented? What would invalidate the sequence?
A roadmap is a set of commitments under uncertainty. AI can help produce options and inspect consistency. It cannot own the tradeoff between customer harm, strategic value, technical risk, and organizational capacity.
High fidelity can be the wrong next step
There is a counterargument that rough prototypes create weak feedback because people cannot imagine the finished experience. Sometimes higher fidelity is necessary, especially when trust, timing, or interaction detail is the assumption.
The answer is not to keep every prototype rough. Match fidelity to uncertainty. Use a sketch to test information order, a clickable flow to test navigation, a simulated AI response to test comprehension, and a working model path to test variable behavior. Do not pay for realism that does not change the evidence.
Before generating the next prototype, complete the Evidence Card in one page. If the team cannot name the assumption and decision rule, the artifact is not accelerating discovery. It is accelerating production of something that looks decided.
When a prototype earns promotion to a launch candidate, the discipline cannot stop at discovery. The AI Feature Ship Gate carries the same evidence standard into the release decision.


