INNOVATION&

INNOVATION&

Permanent Beta: When Software Updates Start Changing Judgement

Continuous updates made software easier to improve. AI makes it harder to know what changed, whether it still works, and who carries the cost when it does not.

Yetvart Artinyan's avatar
Yetvart Artinyan
Aug 13, 2026
∙ Paid

I do not remember the last app update note I read carefully.

Not because I am indifferent to what changes inside the products I use, but because most update notes have trained me not to expect meaningful information.

“Bug fixes and performance improvements.”

“Stability improvements.”

“We regularly update the app to make it better.”

These statements communicate that something changed without telling me what changed, which problem was corrected, which new risk was introduced or whether the product is now safer, faster or simply different in a way I will discover later.

I accept the update anyway.

That acceptance is the part worth examining.

We no longer use finished software. We live inside systems that continue changing while asking us to continue trusting them.

For much of consumer software, this arrangement has been enormously productive. Defects can be repaired quickly, security vulnerabilities can be addressed without waiting for the next major release, and products can respond to real behavior rather than relying entirely on assumptions made before launch.

But continuous improvement also changed the relationship between provider and user.

Software used to arrive with a clearer boundary between development and use. Permanent beta gradually dissolved that boundary. AI now dissolves another one: the boundary between software that performs a function and software that makes or influences a judgement.

That makes the old bargain much harder to evaluate.

INNOVATION& | Better Strategic Decisions Under Uncertainty is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

Software used to have stronger edges

Earlier software releases were not necessarily better.

Bugs could remain unresolved for months. Installation was cumbersome. Companies delayed improvements until enough changes justified a new version. Enterprise releases often required testing, documentation, migration planning and nervous weekends for everyone involved.

The system nevertheless had a visible identity.

You installed a particular version. Changes occurred at recognizable moments. Someone had decided that the release was sufficiently complete and stable to leave the development environment.

The rise of internet services changed that rhythm. In his description of Web 2.0, Tim O’Reilly called one of its defining patterns the perpetual beta: software developed in the open and updated monthly, weekly or even daily.[1]

The product no longer needed to be complete before it reached the user. Use itself became part of product development.

This brought genuine benefits. Instead of attempting to anticipate every need, companies could release earlier, observe behavior and improve the product through feedback.

The user became a participant in the learning system.

Less visibly, the user also became part of quality assurance.

Features could be tested in production. Edge cases could emerge through actual use. Problems that had not appeared in controlled environments could be discovered across thousands or millions of customers.

In low-consequence settings, that can be a reasonable trade.

A music recommendation can be poor. A photo filter can distort the wrong part of an image. A productivity app can rearrange its interface and irritate users for a week.

The defect is visible, the consequence is limited and the next update can often reverse it.

The logic becomes more difficult when software is embedded in infrastructure on which other organizations depend.

When an update becomes the failure

On July 19, 2024, CrowdStrike distributed a routine content-configuration update to Windows systems using its Falcon security platform.

According to CrowdStrike’s own post-incident review, a defect in its Content Validator allowed problematic content to pass validation. When the affected file was processed, it triggered Windows system crashes. CrowdStrike reversed the update within roughly 80 minutes, but the consequences had already spread through organizations around the world.[2]

Microsoft estimated that approximately 8.5 million Windows devices were affected. That represented less than one percent of Windows machines globally, yet many of those devices were operated by enterprises providing critical services.[3]

Airlines, hospitals, financial institutions and public services experienced disruption.

The mechanism designed to protect infrastructure had become the infrastructure failure.

CrowdStrike is not an argument against continuous updating. Security software must respond rapidly to changing threats. A system unable to update quickly would create risks of its own.

The incident reveals a more important principle:

The faster and more widely a change can propagate, the stronger the release controls, containment mechanisms and recovery options must become.

Speed does not remove the obligation to prove reliability.

It increases it.

A product used by a few willing beta testers can tolerate a different level of uncertainty from software embedded across critical operational systems. The release mechanism may look technically similar, but the consequences are not.

Permanent beta therefore needs a boundary.

The difficulty is deciding where to place it.

The permanent-beta bargain

The bargain behind continuous software development can be stated simply:

The provider receives real-world learning. The user receives faster improvement.

The bargain remains defensible when several conditions are present.

The failure is visible. The consequence is limited. The update can be rolled back. A workable alternative remains available. The affected user can recognize the problem and seek correction.

When those conditions weaken, “we will improve it in the next release” becomes less reassuring.

A blocked payment may be corrected later, but the missed deadline remains. A patient may eventually receive the right appointment, but the delay has already occurred. A job applicant may appeal a ranking, but the hiring process may have moved on.

Not every error is erased by correcting the system that produced it.

This is the point at which permanent beta stops being only a product-development philosophy and becomes a governance question.

Who is permitted to learn in production?

Who carries the cost of that learning?

Which consequences can be reversed, and which become part of someone’s life before the next version arrives?

AI changes the shape of failure

Traditional software failures often declare themselves.

The application crashes. The button does nothing. The calculation produces an error. The payment does not complete.

The user may not know why the system failed, but at least the failure is visible.

AI can fail without appearing broken.

The chatbot responds. The ranking appears. The summary is concise. The recommendation is formatted correctly. The fraud alert looks authoritative.

Nothing crashes.

The failure may be an omitted fact, a weak inference, an outdated policy, a biased proxy or a confident answer derived from insufficient context.

The system continues operating, and the output looks like success.

This is what makes AI a different form of permanent beta. Its errors are often semantic rather than mechanical. The system does not merely change how a function is performed. It changes the information, interpretation or recommendation on which a person may act.

A customer-service system may produce an incorrect interpretation of a policy. A recruitment tool may convert limited historical data into a candidate ranking. An internal copilot may omit the one commercial risk that should have changed the recommendation.

The failure does not feel technical.

It feels cognitive.

Fluency changes trust before reliability has been earned

A clear answer is easier to trust than a confused one.

This is normally useful. Clear communication helps people understand reasoning and act on information.

With generative AI, clarity can become misleading. A well-structured answer may be based on incomplete evidence. A confident explanation may disguise uncertainty. A precisely formatted score may rest on a poor proxy.

Presentation quality and epistemic quality are not the same.

Research on automation has documented the risk of automation bias: people may over-rely on automated recommendations, particularly when their attention is divided or the system has performed reliably in previous interactions.[4]

AI adds natural language to that dynamic.

A conventional system might produce a number. Generative AI can produce the number, explain it, defend it and rewrite the explanation in the language of the organization using it.

The output feels less like a machine signal and more like judgement.

That makes human oversight more difficult, not less necessary.

The reviewer must ask whether the answer is correct, whether the evidence is sufficient, whether the model had access to the relevant context and whether the claim should be made at all.

In many consequential settings, the affected person cannot perform that audit.

A candidate cannot inspect the assumptions behind a proprietary hiring score. A customer cannot verify the training data behind a fraud decision. A patient cannot independently evaluate every medical inference generated by a model.

The system produces the judgement.

The person carries the consequence.

Capability is accelerating faster than inspectability

The Stanford AI Index reported in 2026 that documented AI incidents increased from 233 in 2024 to 362 in 2025.[5]

At the same time, transparency among major foundation-model providers declined. The Foundation Model Transparency Index rose from an average score of 37 in 2023 to 58 in 2024, then fell to 40 in 2025. Significant gaps remained around training data, compute, model use and post-deployment impact.[5][6]

This does not prove that every less transparent model is unsafe or that transparency alone ensures reliability.

It does mean that organizations are increasingly integrating systems whose behavior, provenance and downstream consequences may be difficult to evaluate independently.

A model can improve rapidly while the organization using it understands less about what changed between versions.

The interface may retain the same name while the underlying model, retrieval process, safeguards, system instructions and response behavior change. A workflow that was tested in March may no longer operate identically in July.

With conventional software, an update may alter the function.

With AI, an update may alter the judgement.

That requires a different release question.

It is no longer enough to ask:

Does the system still run?

Leaders must also ask:

Does it still produce sufficiently valid and reliable outputs in the context where we use it?

Post-deployment monitoring is necessary—and still immature

Pre-release evaluation remains important, but it cannot reproduce every condition a deployed AI system will encounter.

Inputs change. User behavior changes. surrounding processes change. Data distributions move. Model responses may vary even when prompts look similar.

NIST’s 2026 report on monitoring deployed AI systems argues that post-deployment monitoring is necessary to confirm real-world reliability, detect unforeseen outputs and identify unexpected consequences in changing contexts. It also concludes that common terminology, validated methods and monitoring practices remain immature and fragmented.[7]

This creates a tension.

Organizations cannot prove every aspect of AI behavior before deployment. Some learning must occur in use.

Yet the fact that monitoring is necessary does not mean every system should be allowed to learn through live consequences.

The appropriate question is:

Which uncertainty may be managed in production, and which uncertainty must be reduced before the system receives authority?

That distinction is missing from many AI strategies.

A prototype, internal assistant, customer chatbot and hiring recommendation system may all be described as AI use cases. They should not inherit the same release philosophy.

Human oversight cannot mean human decoration

The standard reassurance is that a human remains in the loop.

That statement is almost meaningless until the loop is examined.

Can the person understand the system’s relevant capabilities and limitations?

Can they recognize when the output may be wrong?

Do they have access to the evidence necessary to challenge it?

Do they have enough time to review it?

Are they permitted to override it?

Will reversing the AI recommendation create additional work, social friction or managerial suspicion?

A human who clicks “approve” because the system has already shaped the workflow is not meaningful oversight. They are the final interface element.

The EU AI Act’s Article 14 requires more for high-risk systems. Human overseers must be enabled to understand relevant capabilities and limitations, monitor for unexpected performance, remain alert to automation bias, interpret outputs and decide not to use, override or reverse them.[8]

This is not simply a requirement to place a person after the model.

It is a requirement to preserve human authority and make that authority operationally usable.

An organization has not created meaningful oversight merely because an employee is technically able to reject the output. The employee needs knowledge, time, evidence and institutional permission to do so.

Otherwise, judgement has already been delegated, even when the governance diagram says it has not.

Update culture has moved into judgement

Update culture trained us to accept instability in software.

The application may behave differently tomorrow. A feature may disappear. A recommendation may improve. A defect may be corrected after users discover it.

AI risks extending that acceptance into decisions.

The model may interpret the case differently tomorrow. A ranking may change because the provider updated the system. A summary may omit different information. A safeguard may become stronger in one context and weaker in another.

This is not automatically unacceptable. Human judgement also varies, improves and fails.

The difference is scale, opacity and authority.

An AI system can reproduce one weak assumption across thousands of decisions before the pattern becomes visible. The person affected may not know that AI influenced the result. The organization may not know that the system’s behaviour changed.

NIST’s AI Risk Management Framework treats validity and reliability, safety, accountability, transparency and explainability as characteristics that must be managed throughout the AI lifecycle, including after deployment.[9]

The underlying principle is straightforward:

A system should not be trusted because it is continually improving. It should be trusted only to the extent that its current performance, limitations and controls justify the authority it has been given.

Not everything belongs in the same beta

This is not an argument against updates, experimentation or AI.

It is an argument against applying one release philosophy to every level of consequence.

Some systems can reasonably operate in open permanent beta. Their mistakes are visible, limited and easily reversed.

Some require controlled beta. They may learn through use, but only within boundaries that include monitoring, escalation, rollback, documentation and a workable human alternative.

Some should not make live consequential decisions while still being treated as experimental. They require stronger evidence before authority is granted because correction after the event cannot fully restore what was lost.

The free diagnosis is complete:

Permanent beta is acceptable only when the organization has matched the speed of learning with the consequence of being wrong.

The central leadership decision is not whether a system may continue improving after deployment.

Almost every useful digital system will.

The decision is how much authority it may receive before its reliability, limits and recovery mechanisms have been demonstrated in the context where the consequences occur.

The Beta Boundary Review

User's avatar

Continue reading this post for free, courtesy of Yetvart Artinyan.

Or purchase a paid subscription.
© 2026 Yetvart Artinyan · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture