INNOVATION&

INNOVATION&

B(AI)AS

When GenAI improves the quality of your work faster than your ability to judge it

Yetvart Artinyan's avatar
Yetvart Artinyan
Oct 08, 2026
∙ Paid

I am currently using GenAI to rethink the SEO and GEO strategy for INNOVATION&.

I would not describe myself as an expert in either field. I know enough to understand the basic mechanisms, follow an argument, question recommendations, compare alternatives, and recognise advice that is too generic to be useful. However, I do not possess the depth of experience that would allow me to judge every technical claim independently or immediately recognize every important omission.

This is precisely why GenAI is so valuable to me.

What would previously have required weeks of research, conversations with several specialists, or the hiring of external expertise can now begin with a dialogue. I can ask the model to assess my current strategy, explain unfamiliar concepts, compare different approaches, identify contradictions, propose priorities, and translate all of this into concrete changes. I can then challenge the result, add context, reject recommendations, request evidence, introduce another perspective, and revise the strategy again.

I rarely accept the first answer. The process is not a single prompt followed by passive acceptance. It is an extended conversation in which I repeatedly question what I receive. Sometimes we make clear progress. Sometimes a later response exposes a weakness in an earlier one. Sometimes we move backwards because a recommendation that initially seemed convincing no longer survives closer examination. Occasionally, we move in circles and return to an idea that had already been rejected.

The result of this process often looks impressive. At least, it looks impressive to me.

That qualification matters because it leads to a question that extends far beyond SEO or GEO.

Am I using GenAI to reason, reflect, judge, and decide more effectively, or am I gradually becoming biased towards its output because I lack enough expertise to recognize where the reasoning is incomplete?

The longer the conversation continues, the more difficult this distinction becomes. Each iteration makes the output more specific to my situation. The model adopts my terminology, incorporates my objections, and increasingly reflects the way I frame the problem. The result therefore feels less like a generic machine response and more like a jointly developed strategy.

Perhaps that means the strategy is improving.

Perhaps I am simply becoming more invested in it.

INNOVATION& | Better Strategic Decisions Under Uncertainty is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

GenAI can raise the floor, raise the ceiling, or conceal the gap

Most discussions about GenAI at work begin with productivity. Can it help people write faster, analyze more information, produce more ideas, or perform tasks that previously required specialist support?

There is growing evidence that it can improve immediate task performance in several domains, although the magnitude and direction of the effect depend heavily on the task, the user, and the way performance is measured. The more important question for innovation work, however, is not simply whether the output improves. It is what happens to the expertise and judgement behind that output.

GenAI can influence expertise in at least three different ways.

It can raise the floor by helping people with limited experience produce work that is more complete, structured, or credible than what they could produce unaided. A non-specialist can create a reasonable first market analysis, a more coherent interview guide, a clearer opportunity statement, or a better-structured experiment.

It can also raise the ceiling by allowing experienced practitioners to explore more alternatives, examine more evidence, pressure-test assumptions, and make connections they may not have reached alone. In this mode, the model does not substitute for expertise. It expands the area over which expertise can operate.

There is also a third possibility. GenAI can raise the apparent quality of the work without raising the user’s ability to judge it by the same amount. The output becomes more sophisticated, while the person evaluating it remains dependent on the same system that produced it.

This third possibility is what I call B(AI)AS.

B(AI)AS is not an established scientific term. It is my name for a possible judgement gap: the distance between the apparent quality of an AI-assisted output and the user’s independent ability to evaluate whether that output is correct, complete, relevant, and appropriate for the decision at hand.

The risk is not limited to people who accept the first response without thinking. It can also emerge through long and apparently critical interactions. A user may challenge the model repeatedly and still become increasingly dependent on its framing, evidence selection, and interpretation.

The output improves. Confidence increases. Yet the underlying judgement capability may remain unchanged.

Why fluent reasoning feels like reliable reasoning

GenAI is persuasive in a way that previous decision-support systems were not. It does not merely produce a recommendation. It can explain the recommendation, defend it, criticize it, compare it with alternatives, acknowledge uncertainty, and reformulate the conclusion after receiving new information.

This creates the experience of reasoning.

The conversation may contain disagreement, revision, qualification, and reflection. These are normally signals we associate with thoughtful human deliberation. When the model responds to objections and changes its conclusion, it can appear as though the original proposition has survived a serious examination.

However, a conversation can become more coherent without becoming more accurate. The same system may generate the proposition, the counterarguments, the evaluation criteria, and the final conclusion. It can simulate opposition without introducing genuinely independent evidence.

The wording becomes sharper. Contradictions disappear. The argument becomes easier to follow. None of this proves that the underlying claims are correct.

Research on automation bias provides part of the scientific basis for this concern. Studies conducted long before the current generation of language models showed that people sometimes follow automated recommendations despite contradictory information, particularly when they perceive the system as competent or when independent verification requires additional effort [1]. More recent research has reframed the challenge as one of appropriate reliance: people need to know when AI advice should be accepted, rejected, or verified rather than simply learning to trust it more or less [2].

Generative AI complicates appropriate reliance because the recommendation is delivered through dialogue. The system appears responsive to the user’s situation and can provide an explanation for almost every answer. This responsiveness can increase the perceived legitimacy of the output even when the explanation does not provide new evidence.

Research also suggests that explanations do not automatically solve this problem. Depending on how they are designed and used, explanations can help people identify errors, but they can also increase reliance by making an incorrect recommendation appear more understandable [2]. An explanation should therefore not be confused with verification.

The capacity to explain an answer is one of GenAI’s most useful characteristics. It is also one of the reasons why an unsupported answer can become unusually convincing.

Iteration is better than accepting the first answer, but it is not independent validation

A common response to concerns about GenAI is that people need to learn better prompting. They should reject the first answer, request alternatives, ask the model to identify its assumptions, and instruct it to argue against its own conclusion.

This is sound practical advice. It is still incomplete.

Iteration can produce genuine improvement when the user contributes new information, corrects the model, introduces external evidence, changes the framing, and applies independent evaluation criteria. In that situation, the conversation is not merely generating more text. It is helping the user examine a larger space of possibilities.

Iteration can also produce what I would call iterative dependence. The user remains active, but the model continues to supply most of the hypotheses, objections, alternatives, and evaluation criteria. The interaction becomes more detailed and sophisticated, yet the epistemic loop remains largely closed.

This form of dependence is difficult to recognize because it does not feel lazy. The user may spend hours working with the model. Several alternatives may be compared, multiple objections considered, and numerous revisions made. The final output may contain assumptions, scenarios, risks, and a convincing decision matrix.

It looks as though the conclusion has been earned.

However, the amount of work invested in refining an argument is not evidence that the argument is better. Repeatedly asking the same underlying system to reconsider its answer does not create the same degree of independence as introducing customer observations, operational data, external research, domain expertise, or evidence that could contradict the conclusion.

Research on the metacognitive demands of GenAI helps explain why this matters. Metacognition refers broadly to the ability to monitor and regulate one’s own thinking. Tankelevitch and colleagues argue that GenAI places substantial metacognitive demands on users because they must define the task, monitor the quality of the interaction, evaluate outputs, and decide when further investigation is necessary [3].

The tool may reduce the cognitive effort required to generate material, while increasing the importance of monitoring and evaluation. This becomes particularly difficult when the user lacks the domain knowledge needed to evaluate the output. To identify an omission, one must usually know enough to understand that something is missing.

A study of 319 knowledge workers found that greater confidence in GenAI was associated with lower self-reported critical-thinking effort, while greater confidence in one’s own ability was associated with more critical evaluation [4]. The authors did not conclude that GenAI simply eliminates critical thinking. Their findings suggest that it changes where critical thinking occurs. Less effort may be required to produce a first output, while more effort is needed to verify, integrate, and govern what the system produces.

This is the central paradox behind B(AI)AS. The less expertise we possess, the more GenAI can raise our floor. Yet the less expertise we possess, the less capable we may be of recognizing where the output is wrong.

The problem becomes more consequential in innovation work

This judgement gap matters in many forms of knowledge work, but it becomes particularly expensive in innovation because the outputs are used to make decisions under uncertainty.

Innovation teams increasingly use GenAI to identify opportunity areas, synthesize interviews, formulate customer jobs, create personas, analyze markets, generate business models, design experiments, interpret evidence, and prepare investment cases. None of these activities is inherently inappropriate. Many can benefit substantially from AI assistance.

They do not, however, carry the same decision risk.

Using GenAI to improve the wording of a workshop invitation is different from using it to conclude that a customer problem is important. Generating an initial list of competitors is different from determining that a market is attractive. Drafting an experiment is different from deciding that a critical assumption has been validated.

The closer an AI-assisted output moves towards an irreversible or expensive commitment, the more important independent judgement becomes.

Consider a team evaluating a new opportunity. GenAI can identify relevant trends, propose customer struggles, formulate market segments, generate value propositions, describe adoption barriers, compare competitors, and recommend experiments. Within an afternoon, the team can produce a document that appears more comprehensive than what it might previously have created in several weeks.

The question is what the team has actually learned.

The model may have organized existing information, revealed connections, or introduced plausible hypotheses. It may have improved the team’s preparation for customer research and helped the team articulate what it needs to learn.

It has not observed a customer experiencing the problem.

It has not seen how people currently improvise around an inadequate solution. It has not watched a purchasing decision stall, identified who controls the budget, or established which constraint prevents switching. It has not demonstrated that the proposed problem is important enough and underserved enough to change behavior.

GenAI can make a hypothesis more articulate. It cannot turn that hypothesis into empirical evidence.

Yet articulation matters psychologically. A hypothesis expressed in professional language, supported by a coherent causal story, and embedded in a polished framework is easier to mistake for an informed conclusion.

This distinction is particularly important when AI is used to synthesize customer interviews. A model can identify themes across transcripts much faster than a human team. It can also flatten contradictions, remove contextual details, overstate weak patterns, or organize responses around categories introduced by the prompt. The resulting synthesis may appear clearer than the underlying evidence actually is.

The risk is not that GenAI always produces a poor synthesis. The risk is that the team loses visibility into the interpretive decisions through which the synthesis was produced.

Raising the floor does not necessarily build expertise

Experimental research provides credible evidence that GenAI can improve some forms of individual performance. Doshi and Hauser, for example, found that access to AI-generated ideas improved how short stories were evaluated. The stories were judged to be more creative, better written, and more enjoyable, with particularly strong benefits for participants who initially scored lower on creativity [5].

This is a clear example of the expertise floor being raised.

The same study also found that AI-assisted stories became more similar to one another [5]. Individual outputs improved, while the collective diversity of outputs declined. Other research into AI-supported ideation has reported related homogenization effects, suggesting that users can converge towards similar semantic territory when drawing on similar model-generated suggestions [6].

For innovation work, this creates an important distinction between the quality of an individual idea and the breadth of the search process.

A team may generate ideas that are more detailed, convincing, and professionally presented than before. At the same time, the overall set of ideas may become less varied. Each individual proposal appears stronger, while the organization explores a narrower opportunity space.

This pattern would be easy to overlook if each proposal were evaluated separately. The team sees that the ideas have improved. It does not see the alternatives that were never generated because multiple participants started from similar AI-produced anchors.

The technology may therefore raise the floor while narrowing the field.

That does not mean AI-generated ideas should be avoided. It means that teams need deliberate mechanisms for preserving variation, introducing independent starting points, and looking beyond the first plausible categories offered by the model.

The first answer does not have to be accepted to become influential. It can act as an anchor around which every later answer is organized.

Expertise changes the nature of the collaboration

A novice and an expert can use the same model, enter similar prompts, and receive outputs of comparable apparent quality. The value they derive from the interaction may nevertheless be fundamentally different.

A novice often uses GenAI to supply missing knowledge. An expert is more able to use it to interrogate existing knowledge.

The novice may ask what the innovation strategy should be. The expert is more likely to examine which assumptions the proposed strategy depends on, what evidence would contradict them, which alternatives have been excluded, and how the recommendation changes under different constraints.

This difference is often described as prompting skill, but prompting is only part of it. Experts possess reference points outside the conversation. They recognize patterns, but they also understand when a familiar pattern has been transferred into the wrong context. They can identify variables that have been omitted, causal claims that are unsupported, or recommendations that remain generic despite appearing highly specific.

Most importantly, they can reject a convincing answer for reasons that the model was never asked to consider.

The phrase “human in the loop” therefore provides little reassurance by itself. A human can be present while contributing almost no independent judgement. The relevant question is whether the person reviewing the output has the knowledge, incentives, time, and authority required to recognize when it should not be trusted.

Research on human–AI collaboration supports this more cautious view. A systematic review and meta-analysis by Vaccaro, Almaatouq, and Malone found that human–AI combinations did not, on average, outperform the better of the human or AI working alone. The combined performance was significantly worse than the best individual performer, with more favorable results in content-creation tasks and greater performance losses in decision tasks [7].

This finding does not show that human–AI collaboration is generally ineffective. It shows that complementarity should not be assumed.

For a combined system to outperform its individual parts, the human must understand when their own judgement is more reliable, when the model is more reliable, and how disagreement should be resolved. Without this calibration, the combination can inherit the weaknesses of both.

Innovation work makes such calibration especially difficult because feedback is often delayed and ambiguous. If AI misclassifies an image, the error may be detected quickly. If it helps a team misdiagnose why customers are not switching, the consequences may unfold over months.

The team may build experiments around the wrong assumption, interpret ambiguous evidence as support, and increase its commitment before reconsidering the original diagnosis. By the time the error becomes visible, capital, credibility, and organisational energy may already be locked in.

The cost of B(AI)AS is therefore not merely one incorrect answer. It is a sequence of reasonable-looking decisions built on an answer that was never independently established.

From first-answer dependence to expert expansion

The useful distinction is not between using GenAI and avoiding it. It is between different modes of reliance.

At one end is first-answer dependence, in which the user accepts the first plausible output with little challenge or verification.

A more sophisticated form is iterative dependence. Here, the user questions and refines the answer, but the model continues to provide most of the reasoning, alternatives, and criteria through which its own work is assessed.

The next mode is informed augmentation. The user brings external evidence, defines evaluation criteria independently, understands important limitations, and has enough knowledge to reject the model’s framing.

At the other end is expert expansion, in which GenAI allows a capable practitioner to examine more possibilities, connect more information, and test more interpretations than would otherwise be practical.

These modes cannot be distinguished by looking at the final output.

Two market analyses may appear equally sophisticated. One may be the result of extensive domain knowledge expanded through GenAI. The other may be the result of a non-expert spending several hours refining the model’s assumptions.

This is why the apparent quality of an output is becoming a weaker indicator of the capability behind it.

GenAI can improve the visible work faster than it improves the user’s ability to judge that work. It may make limited expertise look like deep expertise, both to the reader and to the person who produced the output.

The central question is therefore no longer simply whether GenAI improved the answer.

It is whether our ability to evaluate the answer improved at the same rate.

Thanks for reading INNOVATION&! This post is public so feel free to share it.

Share


Paid learning and application section


From understanding B(AI)AS to evaluating a real decision

The purpose of the B(AI)AS Check is not to determine whether every AI-assisted output is true or false. That would require domain-specific verification. Its purpose is to determine whether an output is sufficiently trustworthy to influence a consequential innovation decision.

Choose one current AI-assisted output. It might be an opportunity assessment, a customer-research synthesis, a market analysis, a business model recommendation, an experiment plan, or an investment case.

Then examine how the conclusion was produced.

User's avatar

Continue reading this post for free, courtesy of Yetvart Artinyan.

Or purchase a paid subscription.
© 2026 Yetvart Artinyan · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture