ART-018
The Most Valuable AI Workflow Should Need Less AI Over Time
Elsewhere I argued for placing probabilistic reasoning only where uncertainty still requires it, and deterministic execution where the organization already knows the answer—and for refusing the AI-versus-rules camp fight so both mechanisms can sit in the same workflow. That was placement at a point in time. This is about what happens next. If the same bounded family of cases keeps getting solved correctly under the same conditions, and the organization still pays full probabilistic rediscovery every time, I do not call that maturity.
I call it expensive amnesia with a dashboard. I am not anti-AI. I am against treating model calls as the scoreboard.
The operating claim is blunt:
The most valuable AI workflow should need less AI over time—where the work is understood.
Not less intelligence overall. Less repeated reasoning on problems that experience has already earned the right to treat as known. Less probabilistic rediscovery on familiar work—not less intelligence when novelty is real.
Success is not maximizing model calls
In the AI roadmaps I see, volume still too often gets treated as virtue. More agents. More autonomy.
More invocations. More “AI in the loop” on every step, including the ones that already have a contract. That is a procurement story, not an operating model.
If a schema check already enforces a known postcondition, wrapping it in a model does not make the organization smarter. It makes the gate optional in practice and theatrical in presentation. If a CI path already makes a security scan non-optional, asking an agent whether this change probably deserves one is not flexibility. It is a retreat from a constraint you already own.
If a BPM path already encodes a stable approval sequence with clear ownership, rediscovering that sequence in chat every week is not adaptation. It is unpaid rework dressed as intelligence. The useful score is whether the workflow finishes trustworthy work—and whether successful judgment leaves durable capability behind.
History is not the same thing as learning. A system that solves the same bounded problem correctly many times, then starts from nearly the same blank slate again, has accumulated runs. It has not necessarily accumulated organizational capability.
Maturity is temporal, not tribal
Use probabilistic intelligence where uncertainty requires reasoning, and deterministic mechanisms where the answer is already sufficiently known. Do not force an entire workflow onto AI or rules as a loyalty identity. Type consequential steps by uncertainty, attach the matching mechanism, and keep promotion and demotion owned.
Time is the missing axis. Early in a new process, exploration is honest.
Agents investigate. Humans correct them. Exceptions appear.
Patterns emerge—if anyone is watching for them. Later, if a bounded pattern keeps producing validated outcomes under stable conditions, continuing to pay for full rediscovery is not “keeping the system intelligent.”
It is refusing to cash the learning. Promotion is the governed path from repeated successful reasoning to deterministic capability. A rule.
A validation check. A decision table.
A parameterized query. A reusable transformation. A workflow fragment that sits beside the CI gates, runbooks, rules engines, and BPM paths you already trust.
This is not rediscovering continuous improvement, MLOps promote-to-prod, or RPA maturity under a new label. Those practices already cash process learning and model promotion into known paths. The distinctive claim is narrower: convert repeated successful AI judgment—specifically—into governed deterministic gates the organization already owns, and keep demotion reversible when evidence, assumptions, or conditions change.
None of that requires a greenfield platform story. It requires refusing to treat every Tuesday as day zero.
Promote into what you already own
Serious organizations already have deterministic machinery. CI/CD gates. Schema and contract tests. Rules engines. BPM case paths. Runbooks. Change boards. Permission systems.
Audit logs. Maturity should start by telling the truth about those assets. If a gate already enforces a known postcondition, keep it—and stop asking a model to re-derive the same certainty for aesthetics.
If a rules estate covers a bounded family well, do not demote it to “legacy” merely because an AI program needs a narrative of progress. If a stable routing sequence already has owners, do not launder it into an unowned chat plan and call the laundering innovation. Coexistence is not nostalgia. It is how you avoid paying twice for the same certainty—once in software you trust, and again in probabilistic rediscovery that pretends the software never existed.
Promotion into that estate should feel boring when it is done well. Boring is often the point.
Recognition is the hard part—and it is owned
The hard part is not agreeing that familiar work should cost less repeated reasoning. The hard part is recognizing, honestly, when a once-agentic decision is ready to become deterministic—and when it is not. I do not think there is a universal formula for that judgment.
Volume alone is not enough. Quiet quarters hide fragile assumptions. A large pile of similar-looking successes can still rest on conditions nobody wrote down.
Before promoting repeated reasoning into a governed deterministic mechanism, I want the organization to be able to answer questions like these:
- Evidence: What outcomes, reviews, or checks actually validated the pattern—not just that it ran often?
- Similarity bounds: Under what input conditions does the pattern apply? What variations are still “the same problem”?
- Exclusions: Which cases must remain on a reasoning path even when they look adjacent?
- Invalidation signals: What policy change, error rate, upstream dependency shift, or human override should reopen the question?
- Review authority: Who approves the promotion, and who can force demotion when those signals appear?
Those dimensions do not solve recognition by themselves. They make recognition an explicit organizational decision instead of an accidental side effect of another model call. In practice, the people closest to the workflow should propose promotion or demotion. Architecture or engineering standards should review the boundary.
Operators and domain experts closest to outcomes should be able to raise invalidation signals without waiting for a roadmap cycle. On Monday, that ownership can look ordinary: an evidence packet with outcomes, similarity bounds, and exclusions; a named approver in architecture review, change board, or standards forum; and a demotion path that suspends or narrows the gate and returns affected cases to investigation or on-call judgment when invalidation signals fire. If nobody owns those decisions, the maturity claim is only a slogan.
Less AI on familiar work is not anti-AI
The phrase “need less AI over time” can sound like abolition. It is not. If a problem is novel, ambiguous, high-impact, contested, or poorly understood, I want the system to use every appropriate reasoning capability available.
Retrieve evidence. Ask another specialist. Simulate alternatives. Challenge the first answer. Escalate to a human.
Take more time. AI is valuable precisely because those problems do not fit neatly into rigid software. That is also why we should not waste AI on the parts that do.
The goal is maximizing the value of scarce intelligence—human and artificial—not minimizing model usage as a vanity metric. Low token counts that produce wrong outcomes are not maturity. They are cheap failure. High token counts on work the organization already knows how to check are not sophistication.
They are directed waste.
Deterministic does not mean permanent
There is an obvious danger. Today’s correct rule can become tomorrow’s legacy defect. A payer changes documentation requirements.
A jurisdiction adopts a new code. A vendor changes an API.
A workload changes shape. A regulatory requirement appears. A business objective changes. If learned behavior becomes deterministic automation and then loses its why, the organization has not gained maturity. It has gained confident wrongness with a ticket queue attached.
I have seen the premature version. A team promotes a “stable” routing rule after a quiet quarter. Then the world underneath it moves.
The rule still fires confidently. Exceptions pile up in a queue nobody designed for demotion. Recovery means treating the rule as suspect again: suspend or narrow it, return affected cases to investigation or human judgment, and rebuild only after new evidence earns trust.
So every promoted mechanism needs an answer to another question:
What would have to change before we should stop trusting this?
Without that answer—and without an owned path back—promotion is just process debt with better marketing.
Score the trajectory, not the theater
A practical Monday test for a consequential AI workflow—remembering that “less AI on familiar work” is not “less AI when novelty is real”:
1. Which familiar cases still pay full probabilistic rediscovery even though outcomes have been validated under stable conditions? 2. What deterministic gates do you already own that could absorb those cases without a greenfield rewrite? 3. What evidence, bounds, exclusions, and invalidation signals would justify promotion—not mere volume? 4. Who can approve promotion, and who can force demotion without waiting for a roadmap cycle? 5. If chat history disappeared tomorrow, would the why of each promoted path still be recoverable from owned artifacts? 6. Are you scoring the program by model calls and agent seats—or by trustworthy completion and retained capability?
Those questions do not require a product category name. They do require an operating habit: explore where uncertainty is real, cash learning into governed certainty where certainty is earned, coexist with the deterministic estate you already have, and keep the frontier movable when reality moves. For transformation offices, the scorecard implication is the same shape. Stop treating invocation volume as proof of AI progress.
Ask teams to show which familiar work got cheaper in reasoning cost because capability was promoted—and which novelty still correctly burns intelligence. Ask them to show demotion ownership, not only promotion slides. At a stage gate, the artifact can be a short register—not a platform purchase—listing which familiar families were promoted into existing gates, which novelty still correctly burns intelligence, and who can force demotion.
Intelligence should leave something behind
I am not claiming a formula for when promotion is ready. I am claiming a direction. Early, reason honestly. Later, stop rediscovering what experience has already paid for.
Always keep a path back when the frontier moves. The organization that treats model calls as the product will keep buying more of them. The organization that treats trustworthy completion—and retained capability—as the product will often need less probabilistic AI on familiar work as it matures. That is not a call to eliminate intelligence.
It is a call to stop wasting it. The next question is sharper still: when a once-stable path no longer knows, how should the organization admit that—and force honest demotion—before confident wrongness compounds?
