The Most Valuable Prototype Is the One That Can Be Wrong
Why AI has made prototypes cheap, demos dangerous, and evidence the only thing worth funding.

The demo ran on a Tuesday afternoon in a room booked for forty-five minutes. The agent took a messy operational case, pulled the right playbook, explained what it thought had gone wrong, and proposed the next three steps. Someone at the back asked whether it could do that for the harder case types. It could.
The sponsor asked for the recording before the meeting ended. Two people stayed behind to talk about which teams should see it next. On the way out, one of the engineers mentioned that the coffee machine on that floor had finally been fixed.
Eleven weeks later I asked what had happened to it. The honest answer was: nothing. It had been shown nine more times. It had changed no decision, had no owner, and had never been tested against a number anyone had agreed to in advance.
It was a very good demo. It was not an experiment.
The persuasion artifact
Most enterprise AI programs are not short of prototypes. They are short of decisions. Demonstrations accumulate, sponsors are impressed, and the organization is no better at answering the questions that govern capital: will this work for these users, on this data, under these controls, at this cost?
The reason is structural. The prototype was built to persuade, not to learn. I call that a persuasion artifact, and its defining property is that it cannot be wrong in any useful way. Nobody declared what result would have killed it, so no result can.
I have built these myself. For years I measured a prototyping team by how many rooms it could win, and I was good at it. The rooms were real. The learning was not.
Why this is getting worse, not better
For most of the history of enterprise R&D, the people who investigate were separated from the people who build, because building was expensive. You studied the market, wrote the requirements, modeled the solution, and only then commissioned a build. The uncertainty persisted for the whole sequence, and that was acceptable because there was no alternative.
That constraint is gone. A bounded, working agentic system can be stood up in days. The prototype is no longer the output of research. It is the instrument of research, or it could be.
Which means the cost of a persuasion artifact has gone up, not down. When a demo took three months, a bad one wasted three months. When a demo takes four days, an organization can produce a dozen of them a quarter, each one impressive, each one unfalsifiable, each one accumulating in the same drawer as the last. Cheap prototypes made demo theater affordable at scale, and AI-accelerated bureaucracy is still bureaucracy.
In my last piece I argued that AI makes individual performance visible, because the path from idea to outcome has compressed. The same compression is now happening to ideas themselves. The question has shifted from who was in the room to what the organization now knows that it did not know before.
Output speed and learning speed are not the same thing
The conflation sits inside the phrase everyone uses: rapid prototyping. It bundles two different things. One is the speed at which you can produce an artifact. The other is the speed at which you can produce evidence you would act on.
AI has made the first nearly free. It has done almost nothing for the second, because the second depends on decisions made before the first line of code: what must be true for this idea to work, what would prove it false, who owns the answer either way.
A clickable mock-up demonstrates an intended experience. A bounded, instrumented prototype tests behavior under realistic conditions. The difference is the difference between an opinion and an observation, and only one of them changes a decision.
The unit of progress is no longer the artifact completed. It is the uncertainty resolved.
The rule: build the smallest thing that can be wrong
The operating model I now run has six stages, and the whole of it fits in one sentence. Hypothesize, materialize, experiment, generate intelligence, decide, industrialize. But the mechanism that makes it work is a single decision rule applied before anything is built.
Name the riskiest assumption. Then build the smallest artifact capable of proving that assumption wrong, and write down, before the demo, what result would make you stop.

Run the rule and the shape of the prototype changes. A multi-agent architecture is almost never the smallest thing that can be wrong about anything, so it rarely survives the first pass. A triage agent that advises specialists is a much smaller instrument than one that resolves cases, and it tests the assumption that actually matters, which is whether the explanations are good enough for a human to act on. Deterministic logic replaces a model wherever repeatability is what you need to prove.
The rule also changes what counts as finishing. Every experiment ends in one of four outcomes: iterate, pivot, transfer to product engineering, or stop. Stopping early is one of the most valuable results an R&D capability can produce. An organization that cannot celebrate a well-evidenced stop will, over time, stop producing honest evidence.
And then the done-test, which is the part people skip. Can the team point to a decision that the evidence changed? A budget line moved, a scope narrowed, a use case retired, a handoff accepted by someone who now owns it. If yes, it was an experiment. If no, it was a demo, however good it looked.
The questions I ask before a prototype is funded
These are the questions I put to any team proposing to build something, including my own. Each one exposes a different failure.
“What has to be true for this to work?” — exposes whether there is a hypothesis at all, or only an idea.
“What result would make us stop?” — exposes whether the prototype can be wrong. If the room goes quiet, it cannot.
“What is the baseline we are trying to beat?” — exposes whether anyone measured the current process before proposing to replace it. Usually nobody did.
“What is the smallest thing that tests that?” — exposes agent maximalism: architecture chosen for sophistication rather than for the assumption under test.
“Who owns the answer either way?” — exposes the shadow product organization: a prototype with no receiving owner becomes production by accident.

Notice what is missing. I no longer ask how impressive the demo will be. I have stopped asking because the answer has stopped correlating with anything I care about.
What I got wrong
I used to think faster prototyping was the point. When the tools arrived I measured my teams on cycle time to a working artifact, and the numbers were extraordinary. Things that took a quarter took a week. I took that as proof the model was working.
It took an uncomfortable review of what those artifacts had actually changed to see the problem. Speed of output had gone up by an order of magnitude. Speed of trustworthy learning had barely moved, because we were still deciding what to build the way we always had, and still declaring success the way we always had, with applause.
The shift was to stop treating the prototype as the deliverable and start treating it as the instrument. Rapid prototyping is a method. Accelerated learning is the outcome. I had been optimizing the method and reporting it as the outcome, and I suspect most AI programs still are.
My stake in this
I run an AI and forward-deployed engineering function. A model that funds experiments on evidence rather than on demos favors the kind of engineer I hire and the kind of work my teams do, and it makes a certain kind of innovation theater, which I have also been paid to produce, harder to sustain. Discount accordingly. Then run the done-test on the last three prototypes your organization funded and see what you find.
“Demos get funded. Evidence gets questioned.”
This is the objection I hear most from people who have run innovation inside large enterprises, and it is at least half right. Budgets move on stories. A compelling demonstration in front of the right executive can release investment faster than an evidence dossier. A prototype reporting a 25 percent improvement invites questions about the baseline, sample, and method. One that simply works in the room invites a conversation about rollout.
So the honest version of the objection is political: declaring what would make you stop can hand opponents a weapon, and nobody wants to write down the number that killed their own initiative. Two responses. First, evidence is harder to inflate than a demo, and beyond a single quarter that matters. A prototype can be run. A deck can only be presented. A team that brings a predeclared threshold and clears it has something a better demo cannot manufacture, and sponsors burned by a drawer full of persuasion artifacts learn the difference.
Second, this is a leadership decision, not a technology one. If leaders keep funding on applause, they will keep getting applause and penalize the teams that stop weak ideas early. Measurement has to move upstream, to resolved uncertainty rather than demos built. It moves only when someone senior decides it will. That person is reading this.
The question I would ask first
Before any AI initiative is funded, before the architecture and the vendor and the roadmap, I would ask one thing: what would have to happen for us to stop?
If nobody can answer, there is no experiment yet, only an intention. The scarce thing in the agentic era was never the ability to build. It is the discipline to build something that can be wrong, and the judgment to act on what it tells you.
Access to AI made prototypes cheap.
Evidence decides which of them deserve to exist.
Decisions are what finally make the evidence worth producing.
So, in your organization: of the last ten prototypes that impressed a room, how many can you point to a decision they changed?
That handoff—from evidence-producing experiment to production-owned product—is a discipline of its own, and one I’ll address separately.
About the author
Rodnei Connolly is a product, AI, and enterprise transformation leader focused on turning emerging technology into measurable business outcomes.
Authorship note: This piece reflects Rodnei Connolly's personal perspective based on experience leading AI-enabled transformation and enterprise technology initiatives. AI tools were used as part of the research, ideation, and editorial process; the thesis, perspectives, judgment, and conclusions are his own.
