Capability&Consequence

Essay

The Verification Shadow

AI can collapse generation time while leaving the obligation to prove the output almost untouched.

Document
CAC-002
Issue
1.0
Published
May 25, 2026
Reading time
9 minutes
Review due
November 23, 2026

TL;DR

Generative AI reduces the time required to produce code, analysis, decisions and documents. It does not automatically reduce the obligation to check them. In consequential workflows, faster generation creates a verification shadow: review, testing, traceability, provenance, exception handling and independent challenge that sit behind the visible saving. The right productivity metric is not output generated per hour. It is accepted, evidenced and supportable output after the full checking burden—and that burden varies radically by workflow.

The fastest part of many AI workflows is becoming the least important number.

A model can now draft code, a credit memo, a clinical summary, a contract analysis or a board presentation in minutes. The comparison is usually made against the hours a person once spent producing the first version. The difference becomes the headline productivity gain.

But the output is not valuable when it is generated. It becomes valuable when the organisation is willing to accept, use and stand behind it.

Between those two moments sits a second body of work: reviewing, testing, tracing, reconciling, challenging, documenting, approving and handling exceptions. That work is the verification shadow.

The shadow is easy to miss because it often belongs to a different person, budget, system or stage than the generation saving. It also grows in the places where being wrong is expensive.

Generation cost and acceptance cost are different curves

Suppose AI reduces the production time for an artifact from ten hours to one. That does not imply a 90 per cent productivity gain.

The relevant equation is closer to:

Net productivity gain = production time removed − new verification, integration and exception work.

If the old process required ten hours of production and two hours of review, while the new process requires one hour of generation and seven hours of checking, the improvement is real—but it is not the one advertised.

The distinction becomes more important as generation approaches zero. Once the first draft takes minutes, almost all remaining cost sits in whether the organisation can trust and operationalise it.

This is not an argument that AI produces poor output. It is an argument that the obligation to verify is not determined only by average output quality. It is also determined by consequence, regulation, accountability and the evidence required after the event.

Functional safety already contains the pattern

Every assurance regime rests on a trace: requirements to design, design to implementation, implementation to test and test to evidence. The trace usually assumes that a named person or controlled tool stood at each link and can account for the result.

Generative implementation changes that assumption.

Functional-safety practice does not answer this with a universal ban. Tool-confidence methods ask two broad questions: how serious an undetected tool error could be, and how likely the surrounding process is to detect it. Independent measures outside the tool—static analysis, unit testing, integration testing and review—can reduce reliance on the generator itself.

Read that in commercial terms: the organisation may not need to qualify every generative tool as though its output were inherently trusted. It can verify the output more aggressively instead.

The cost has not vanished. It has moved from generation or tool qualification into downstream verification.

The least expensive qualification path in mature tool regimes often depends on a record of successful use under stable conditions and a stable version. Fast-changing generative tools struggle to accumulate that history before the version changes. The route that relies on long prior use is therefore structurally difficult for the category.

The practical result is uneven economics. Generative assistance can be highly attractive in low-integrity parts of a system and heavily gated in high-integrity ones. A single productivity assumption across an engineering organisation building both is almost certainly wrong.

The same pattern appears outside code

The vocabulary changes by industry. The economic mechanism does not.

In banking, a model-generated decision memo may be fast to produce, but consequential decisions still require data lineage, policy compliance, independent challenge, version identification and evidence an examiner can inspect.

In healthcare, summarisation may save clinician time, but a recommendation that affects diagnosis or treatment creates a different review obligation from a draft discharge note. Contraindications, thresholds and clinical responsibility do not disappear because the average model improved.

In legal work, generation may collapse the time required to assemble a first argument. Citation checking, privilege review, jurisdictional fit and professional accountability remain.

In sales, a model can draft a proposal almost instantly. The real bottleneck may be solution feasibility, pricing approval, contractual risk and whether the organisation can deliver what the generated proposal promised.

In finance, an AI-created analysis can be numerically polished and still require reconciliation to source systems, control evidence and an accountable owner.

The key variable is not whether the workflow is “creative” or “analytical.” It is the cost of accepting an error and the evidence demanded before and after the decision.

Better models do not eliminate every shadow

Improving model accuracy should reduce some verification work. Fewer defects mean fewer corrections. Better retrieval and structured outputs can make checking faster. Provenance tooling can automate parts of the record.

But three obligations may remain even when average quality rises.

Independence. Some decisions require challenge by a person or system that did not produce the original output.

Reproducibility. The organisation may need to show which model, data, policy and version produced the result, particularly if the decision is contested later.

Accountability. Someone must still decide that the output is acceptable for use. A better model can change the evidence available to that person; it does not necessarily eliminate the role.

This creates a floor under the verification cost of certain workflows. Model quality can improve rapidly while the minimum defensible process changes slowly.

The productivity percentage is a portfolio, not a constant

The phrase “AI makes developers 30 per cent more productive” is not a usable operating assumption. Neither is its equivalent in finance, legal, healthcare or sales.

A real productivity model separates at least four populations:

  1. Low-consequence generation. Errors are cheap and easily reversed. Verification can be light.
  2. High-volume structured work. Automated tests or reconciliation can check output cheaply.
  3. Consequential judgment. Errors are expensive and independent review remains material.
  4. High-integrity or regulated implementation. Traceability and evidence are part of the product, not overhead added later.

The same model may be economically transformative in the first two and marginal in the fourth. The organisation-level result depends on the mix.

This is why pilots mislead. They often measure the speed of producing a representative artifact under observation. Production economics depend on repeated acceptance, exceptions, integration and evidence across the full distribution of cases.

How to measure the shadow

For each workflow, collect six numbers:

  • Time to produce the artifact before AI
  • Time to review and accept it before AI
  • Time to generate it with AI
  • Time to verify, correct and approve it with AI
  • Exception rate and exception-handling time
  • Rework discovered after acceptance

Then distinguish time saved from capacity converted. A person recovering three hours a week has not automatically created three hours of economic value. The operating model must redirect, aggregate or remove that capacity.

Finally, identify which controls can move outside the model. Deterministic checks, reconciliation, policy rules, tests and permission limits can make verification cheaper and more stable across model changes.

The goal is not to minimise checking at all costs. It is to design the workflow so the cheapest reliable verifier handles each obligation.

What this predicts

AI will automate fastest where output can be checked mechanically, mistakes are reversible and the evidence required for acceptance can be generated automatically.

It will move more slowly where judgment must be independent, consequences are asymmetric, provenance matters years later or the organisation cannot define acceptance clearly.

That does not mean the second category will remain untouched. It means its economics will be driven less by raw model capability and more by the architecture of verification.

The generation revolution is visible. The verification redesign is the work that determines who captures the value.

The decision this should change

Stop applying one AI-productivity percentage across a function or codebase. Separate workflows by consequence, define the evidence required for acceptance, measure generation and verification time independently, and automate only where the total accepted-output economics improve.

What this adds

Prevailing consensus

As models become more accurate and generation becomes cheaper, the human review burden should decline proportionately and productivity gains should spread broadly across knowledge work.

What this challenges

The duty to verify is often created by the consequence of error, not merely the model's average quality. Higher model performance can reduce defects while leaving traceability, independence, provenance and accountability obligations substantially intact.

New contribution

The essay defines the verification shadow as a separate cost curve, shows why it differs by workflow and explains why generative assistance is cheap in low-consequence work but gated, uneven or uneconomic in high-integrity work.

What would weaken the argument

Reliable external evidence that generative systems can supply accepted provenance, reproducibility and assurance artifacts at materially lower total cost—or rules that permit organizations to rely on model performance without independent verification—would weaken the thesis.

Sources and references

  1. ISO 26262-8 — Supporting processes for functional safety
  2. NIST — AI Risk Management Framework
  3. Federal Reserve — SR 26-2, Revised Guidance on Model Risk Management