A corporate venture team enters its investment review with a sharper market analysis, a broader competitor map, a polished business case and a credible experiment plan.
The work looks as though several specialists spent weeks producing it. Much of it was completed in two days with generative AI.
The presentation is impressive.
The evidence is not.
The team still has not shown that the customer problem is urgent, that customers will change their behavior or that anyone will pay. AI improved the quality of the artifacts before the underlying uncertainty fell. Because the work now looks stronger, another round of funding becomes easier to justify.
AI did not create output theatre.
It made it cheaper, faster and harder to detect.
Weak OKRs were already confusing completed work with achieved progress. “Launch the assistant,” “produce the market report,” “conduct 20 interviews” and “deliver the prototype” may all describe necessary work. They do not show that anything meaningful changed.
Generative AI can now produce these visible signs of activity at a speed that most management systems were not designed to handle.
The problem is not that AI makes teams more productive.
The problem is that many organizations were already measuring the production of work as though it were the purpose of the work.
AI has exposed the difference.
The productivity gain is real
It would be a mistake to dismiss generative AI as a machine for producing inexpensive text and attractive presentations.
In a field experiment involving 758 Boston Consulting Group consultants, participants using GPT-4 completed 12.2 percent more tasks and worked 25.1 percent faster on tasks within the model’s capability frontier. The quality of their work also improved substantially.[1]
The effect was not uniform.
For a complex task deliberately selected to sit outside that frontier, consultants using AI were 19 percent less likely to reach the correct answer than those working without it. The AI did not merely fail to help. Its plausible but incorrect analysis pulled people toward the wrong conclusion.[1]
A second field experiment examined 791 professionals at Procter & Gamble working on real product-development challenges. Individuals using AI produced results comparable in quality to two-person teams working without it. AI also helped commercial specialists produce more technically balanced proposals and technical specialists produce more commercially balanced ones.[2]
This matters.
AI can give an individual access to some of the perspectives, analytical support and production capacity that previously required a larger or more diverse team. It can help someone examine an issue from another discipline, challenge an argument and produce work that would otherwise require much more time.
For innovation teams with limited access to specialist expertise, that can be a significant advantage.
But the advantage concerns the ability to produce and improve an output.
It does not establish that the output is correct, that the underlying assumptions are valid or that the work has created the outcome the organization needs.
The distinction becomes more important as the output improves.
A polished artifact can conceal unchanged uncertainty
At their core, large language models generate responses by predicting sequences of tokens from patterns learned during training. The GPT-4 Technical Report describes GPT-4 as a Transformer-based model pre-trained to predict the next token in a document and explicitly notes that the system is not fully reliable and can produce fabricated or inaccurate content.[3]
GPT-4o extended the approach across text, images, audio and video, but its system card continued to treat model behavior as something requiring evaluation, safeguards and human judgement.[4]
The mechanism does not make the output trivial. AI-generated analysis can be genuinely useful, original and better than what a person would have produced alone.
It does make one point unavoidable:
A plausible output is not self-validating evidence.
The model can draft an experiment plan. It cannot establish that the experiment tests the decisive assumption.
It can summarize customer interviews. It cannot determine whether the sample represents the market.
It can construct a business case. It cannot make the assumed adoption rate true.
It can generate a persuasive recommendation. It cannot accept the consequences when that recommendation is wrong.
The organization still has to judge the work.
Unfortunately, better presentation can make that judgement less critical precisely when it should become more demanding.
AI can increase confidence without increasing validity
The problem becomes sharper when AI does not merely recommend an answer but also explains why the answer should be trusted.
A field experiment involving 228 evaluators examined 3,002 decisions about 48 real early-stage innovations. Participants either evaluated the opportunities without AI, received AI recommendations without explanations or received recommendations accompanied by narrative rationales.[5]
AI assistance improved screening performance overall.
The explanations created a separate effect. Narrative rationales increased evaluators’ alignment with the AI recommendations by 19 percentage points without producing a corresponding improvement over recommendations presented without explanations. The effect was especially strong when the AI recommended rejecting an opportunity.[5]
The explanation made the recommendation more persuasive.
It did not make it more accurate.
That finding is particularly relevant for corporate innovation because market analyses, interview summaries, business cases and experiment reports are rarely neutral documents. They are part of an argument for what the organization should fund, stop or build next.
AI can now make those arguments more coherent, more comprehensive and more difficult to challenge.
The investment committee sees a well-structured market narrative, a clear problem statement, a detailed experiment plan and a professional financial model. The quality of the package raises confidence in the initiative.
But the critical assumptions may be exactly as uncertain as they were before the package was produced.
What improved was the argument.
What did not necessarily improve was the decision.
AI exposed what weak OKRs were measuring all along
The problem is not primarily the technology.
It is the management system receiving its output.
Objectives and Key Results are intended to connect work with observable progress. The Objective describes what the organization wants to accomplish. The Key Results indicate how it will know whether that state has been achieved or meaningfully advanced.
In practice, many Key Results describe work instead:
launch the prototype;
complete the market study;
interview 20 customers;
deploy the AI assistant;
create the investment proposal;
test three pricing models.
These items are concrete, measurable and controllable. They are also attractive because teams can complete them.
That does not make them evidence that the Objective was achieved.
The OKR literature distinguishes between inputs, outputs and outcomes. Inputs describe actions under the team’s control. Outputs describe what those actions produce. Outcomes describe the change created by the work. Outcome-based Key Results are often more useful because they preserve flexibility over how the desired change is achieved.[6]
Outputs can sometimes be legitimate Key Results. Delivering a defined capability may itself be strategically important. In highly regulated, infrastructural or operational work, completing something may be a necessary and material result.
The danger is not the presence of an output.
It is the absence of any connection between that output and the outcome it is supposed to create.
AI makes this weakness harder to ignore because it dramatically increases the supply of what is easy to produce and count.
A team can now generate reports, prototypes, summaries, recommendations, market maps and implementation plans much faster than before. If the OKR system rewards those objects, AI will make the organization look increasingly successful without necessarily making it more effective.
This is a familiar management failure. Research on goal setting has warned that narrow targets can focus attention on the measured objective while causing people to neglect important consequences outside it.[7] Steven Kerr described the broader organizational pattern as rewarding one behavior while hoping for another.[8]
Organizations say they want customer value, reduced uncertainty and better investment decisions.
They reward completed interviews, delivered prototypes and polished business cases.
AI did not create that contradiction.
It industrialized it.
An initiative can complete every Key Result and still fail its Objective
Consider an innovation team with the following Objective:
Build confidence that an AI assistant can improve how account managers prepare for important customer meetings.
Its Key Results are:
interview 20 account managers;
produce a market and competitor analysis;
build a working prototype;
pilot it with three sales teams;
prepare a business case for scaling.
The team may complete every Key Result.
The Objective may remain unresolved.
Twenty interviews do not establish that the problem is urgent or that current preparation causes meaningful commercial harm.
A competitor analysis does not show that customers will adopt another tool.
A prototype does not prove that behavior will change.
A pilot does not establish repeatable value merely because people agreed to participate.
A business case does not become evidence because its calculations are detailed.
Every Key Result describes progress through a plan.
None necessarily demonstrates progress toward the decision.
AI can now help the team complete all five more quickly and professionally. The OKR score may improve while the uncertainty remains unchanged.
This is why an initiative output cannot automatically be treated as an innovation outcome.
The relevant question is not only:
Did we do what we planned?
It is:
Did the work change what we know, what customers do or what commitment is now justified?
Outcome-based OKRs need a different interpretation in innovation
Outcome-based Key Results should remain the default.
Early innovation creates a legitimate complication: final commercial outcomes may not yet be observable.
A pre-revenue venture cannot always measure recurring revenue. A discovery team may not yet be able to demonstrate retention. An early technical experiment may need to establish feasibility before customer behavior can be tested.
This does not justify replacing outcomes with activity.
It means that the organization needs stage-appropriate evidence.
A useful innovation Key Result may measure whether a decisive assumption has crossed a predefined evidence threshold.
Compare these statements:
Conduct 20 customer interviews.
and:
Demonstrate that at least 12 of 20 customers in the target role experienced the problem during the past three months, used a costly workaround and will commit time, data or budget to a next test.
The first measures activity.
The second defines evidence that could strengthen or weaken the problem thesis.
Compare:
Launch a prototype.
with:
Demonstrate that at least 70 percent of target users complete the critical workflow without assistance and choose the new process over their current alternative in repeated use.
The first measures an output.
The second measures observable behavior under defined conditions.
Compare:
Build the AI assistant.
with:
Reduce meeting-preparation time by 40 percent without reducing the accuracy, relevance or commercial usefulness of the resulting briefing.
The first says what the team will produce.
The second says what must become better.
The distinction is not semantic hygiene.
It determines whether the organization funds work because a team produced something or because the work created evidence that justifies another commitment.
AI can produce the output. It cannot own the outcome.
An AI system can draft, calculate, compare, classify, recommend and explain.
It can help a team formulate an Objective, rewrite a Key Result, generate possible thresholds and challenge whether the selected measure is vulnerable to gaming.
It cannot own the outcome.
Ownership requires more than producing the work. It includes deciding which trade-offs are acceptable, determining what evidence is sufficient, accepting responsibility for false positives and false negatives, and answering for the consequences of the commitment.
AI cannot carry that responsibility.
A model does not lose credibility when a venture receives another €2 million and fails. It does not explain the decision to the board, the employees who joined the project or the customers affected by the result.
The accountable executive does.
This is why inserting AI into an OKR system does not reduce the need for ownership. It makes ownership more important because the system can now produce a much larger quantity of persuasive work without increasing the number of people prepared to stand behind its validity.
AI did not break OKRs. It exposed whether they were measuring completed work or meaningful change.
When output becomes cheap, output-based confidence should become more expensive.
The organization needs a higher standard for what counts as progress, not a larger volume of visible production.




