The Trouble with Tokenmaxxing: When More AI Produces Less Enterprise Value

- September 1, 2026

Ganesh Padmanabhan

CEO and Co-founder of Autonomize AI

An Executive Point of View on Controlling AI Costs, Reducing Variability, and Delivering Consistent Outcomes at Scale

Abstract

Enterprises are spending heavily to make AI smarter and may be making it less dependable. When a system falls short, teams add more context, more reasoning, larger models, and more agents, a practice known as tokenmaxxing. The result is not only a larger bill. Every additional reasoning step gives the system another opportunity to drift, turning the same facts into different conclusions for different people. In healthcare and other regulated industries, that is not ordinary software variation. It is an operational, compliance, and trust failure.

Tokenmaxxing is often misdiagnosed as a cost problem when it is fundamentally a consistency problem. The critical question is not simply how much each answer costs, but whether the same evidence, policy, and circumstances reliably produce the same decision. This article makes the case for reasoning in fewer places, grounding verified knowledge for reuse, and reserving powerful models for work that truly requires them. The next winners in enterprise AI will not be the companies that consume the most intelligence, but those that make it repeatable, traceable, and trusted.


Over the past two years, an additive mindset has taken hold in enterprise AI. When a system underperforms, teams expand the context, add another reasoning step, use a larger model, or introduce a second agent to check the first. The market refers to this as tokenmaxxing.

The term first described companies maximizing token consumption and treating usage as evidence of AI adoption or productivity. Like measuring developers by lines of code, it was easy to track and reward but did not necessarily reflect the value of the work produced. Now tokenmaxxing has become an architectural habit. Teams add more intelligence before determining where intelligence is actually needed.

And, market incentives encourage this behavior. Vendors are often paid based on consumption, longer context windows are marketed as greater capability, and frontier models become the default. Adding another model call is fast, but determining what the system should stop doing takes discipline.

The result is predictable. Previously processed context is repeatedly sent to models. Verified facts are re-derived for every case. Routine tasks run on the most expensive models available. AI spending grows faster than measurable value. These approaches may shine in demonstrations or isolated pilots, but in production they quickly become expensive, unpredictable, and difficult to govern.

And, that distinction between a promising pilot and a dependable production system matters, especially in healthcare. At Blueprint 2026, McKinsey partner Jessica Lamb, who leads the healthcare and public sector vertical for QuantumBlack, described an industry at a genuine inflection point. For the first time, a majority of healthcare organizations surveyed by McKinsey reported implementing, rather than merely piloting, a generative AI solution. Yet most still could not quantify the return.

As Lamb put it, “People are saying we definitely think it’s positive, but when you ask them to quantify it, most cannot.”

That gap between confidence and evidence should concern enterprise leaders. More deployments, more agents, and more tokens do not necessarily mean more value. The real test is whether AI improves the economics, consistency, and outcomes of the work.

Building enterprise AI for healthcare has made the issue increasingly clear: the money spent on excess inference is significant, but the organizational cost of inconsistent decisions is much larger.

The real enterprise risk is inconsistent results

Every leader is aware of the token costs accumulating on an AI bill. What many may be missing is the larger risk already entering their operations: the same case can produce a different answer the next time it runs.

Every reasoning step creates another opportunity for the system to produce a different result. Run the same case twice and you may get different wording, different interpretations of the evidence, different reasoning paths, or different conclusions. A ten-step chain has ten places to drift. The agent added to check the work introduces another probabilistic judgment rather than eliminating the first.

This is not a theoretical edge case. AI engineer and author Chip Huyen describes the problem plainly: “LLMs are stochastic. There’s no guarantee that an LLM will give you the same output for the same input every time.”

OpenAI’s Shyamal Anadkat makes the same point from the model-provider perspective: “The Chat Completions and Completions APIs are non-deterministic by default.” Even with matched parameters and controls intended to improve reproducibility, OpenAI notes that determinism is not guaranteed.

The enterprise consequence is far more serious than inconsistent prose. If an AI system summarizes the same clinical evidence differently, applies a coverage policy differently, assigns a different risk score, or recommends a different next action, the variation moves downstream. It affects which cases are escalated, which claims receive scrutiny, how quickly care advances, and whether two people in materially similar circumstances receive materially similar treatment.

At sufficient scale, small variations become systemic inconsistency. One answer may fall within policy while another does not. One case may be approved while a comparable case is delayed. One reviewer may accept an AI-generated recommendation that another reviewer never sees. The average accuracy of the system can remain impressive while individual decisions become impossible to defend.

That is why “it was right most of the time” is not an adequate enterprise standard. Organizations must be able to demonstrate that comparable inputs receive comparable treatment, that deviations are detected, and that every material decision can be traced to the evidence and rules that governed it.

This is the deeper cost of tokenmaxxing. Teams add reasoning to improve average quality, but the additional reasoning can widen the range of possible outcomes. The system becomes more capable while also becoming less dependable, a distinction that is easy to miss in a demonstration and dangerous to ignore in production.

In my conversations with healthcare leaders, I see the same pattern repeatedly: organizations are not short on AI pilots. What they lack is the operating model to unite business and technology, redesign work around AI, and turn promising experiments into measurable impact in production and at scale.

Tokenmaxxing can mask that operating-model gap. When a workflow is poorly defined, institutional knowledge is fragmented, or accountability is unclear, adding more model reasoning may make the demonstration look better without fixing the underlying process. The organization is using AI to compensate for ambiguity it has not resolved.

In regulated work, that ambiguity becomes risk. When a clinician, auditor, regulator, or member asks how a decision was reached, the organization needs a clear path that will hold up months later. A collection of changing model outputs is not an auditable decision record.

The legal implications are beginning to receive more attention as well. Writing in the Yale Law Journal, Trent S. Kannegieter warns: “Even when given identical inputs, LLMs produce varied, unbounded outputs that cannot be predicted in advance or from prior testing.” His argument is aimed at AI liability, but the operational lesson is immediate. Testing a system once and observing the right answer does not prove it will produce that answer consistently in production.

Variation that seems harmless in testing can become inconsistent member treatment, rework, financial leakage, compliance exposure, or an audit finding at scale. The issue is no longer whether the model can produce a correct answer. It is whether the enterprise can depend on that answer recurring when the facts have not changed.

Point solutions optimize steps. Enterprises must redesign domains.

Healthcare’s AI adoption challenge is magnified by a long-standing productivity problem. Lamb highlighted that healthcare has produced essentially no per-person productivity gain over the past 25 years while other major industries have benefited from technology. The opportunity is enormous, but the industry has limited muscle memory for translating technology investment into operational improvement.

This helps explain why tokenmaxxing is so tempting. It lets organizations layer intelligence onto existing processes without making harder decisions about how the work itself should change.

Horizontal assistants can help employees become comfortable with AI. Point solutions can improve an isolated task. Neither necessarily changes the economics of an end-to-end domain. Lamb described the next era as domain reimagination: redesigning claims, credentialing, network contracting, care management, or another function around what today’s technology makes possible.

That is also the right level at which to confront tokenmaxxing. The question is not, “How can we make this individual model call smarter?” It is, “Why does this workflow require the same policy to be interpreted 25 times? Why are verified facts being reconstructed for every case? Which decisions require judgment, and which should follow an established path?”

Real productivity gains come from removing unnecessary work, not simply surrounding it with more AI.

Reason once, then reuse what is known

Enterprises can reduce cost and variability by limiting where open-ended model reasoning occurs. Once a fact, rule, or procedure has been validated, it can be encoded with its source, meaning, and applicable constraints. Agents can then retrieve and apply that knowledge rather than reconstructing it for each new case. This replaces repeated model calls with a consistent, inspectable execution path.

Consider a common pattern inside large healthcare enterprises. A model processes the full standard operating procedure for every case. It identifies the relevant section, interprets the logic, decides which step applies, and generates an answer. The process may require 20 or 25 sequential model calls and is repeated for every case, even when the procedure itself has not changed.

Using recursive language model techniques, the procedure can instead be compiled once into a structured representation the system can navigate deterministically. Model calls during case processing can decline from 20 or 25 to zero, while processing time drops from minutes to seconds. Compiling each document version may require approximately one million tokens, but that cost is incurred once per version rather than repeatedly for every case.

The efficiency gains are substantial, but consistency and traceability are more important. The compiled procedure produces repeatable results through an execution path that can be inspected, audited, and defended.

Related research from Autonomize reinforces the value of separating flexible reasoning from deterministic control. In RLMOpt: Adaptive Prompt Optimization via Recursive Language Models, a recursive language model directs exploration while a deterministic harness owns objective scoring, selection, and regression constraints. The application is different, but the architectural principle is the same: use model reasoning where adaptability creates value, and use deterministic mechanisms where the enterprise requires consistency and control.

This is what it means to turn intelligence into an enterprise asset. The organization pays to reason once, validates the result, preserves the provenance, and makes that knowledge reusable across cases, workflows, and agents.

Match the intelligence to the task

The goal is not to minimize token use everywhere. It is to match the intelligence to the task and reserve expensive reasoning for the problems that genuinely require it.

Novel questions, open-ended research, ambiguous cases, and high-value work without a predefined path can benefit from deeper reasoning and more powerful models. In those situations, reasoning is what the organization is paying for, and it can materially improve the outcome.

The mistake is applying the same approach to routine, repeated, and well-understood work, which represents much of what an enterprise runs. Classifying documents, extracting fields, applying established policies, retrieving verified facts, and routing cases can often use deterministic logic, grounded retrieval, or smaller models. These workflows do not need AI to repeatedly reconsider what the enterprise already knows. They need the system to retrieve the right answer and apply it consistently.

This distinction also provides a more useful way to evaluate AI investments. Instead of tracking token consumption, model size, or agent count, leaders should measure operational outcomes: cycle time, accuracy, acceptance rates, rework, financial impact, workforce capacity returned, and the consistency of decisions across comparable cases.

Consistency must be measured directly, not inferred from a single accuracy score. Enterprises should repeatedly run identical and near-identical cases, measure decision disagreement rates, test policy conformance at every material branch, and monitor whether model or prompt changes alter previously accepted outcomes. A system can achieve high average accuracy and still be unsafe if its errors move unpredictably from case to case.

Technology is only part of the investment

Across industries, organizations achieving measurable value from AI often invest four to five times as much in operationalizing the technology as they do in the technology itself. The real challenge is not whether AI can perform. It is whether the organization can redesign workflows, prepare its workforce, drive adoption, and embed AI into everyday operations.

The ratio will vary by organization and use case, but the principle is sound. An AI architecture cannot be separated from the operating model around it. Employees need to understand when to trust the system, when to challenge it, and when to escalate. Business owners must help define the rules and outcomes. Technology teams must establish evaluation, monitoring, and governance. Leaders must redesign roles and incentives so that time saved becomes productive capacity rather than an unmeasured benefit.

This is another reason tokenmaxxing is such a costly habit. Every unnecessary reasoning step adds technical complexity that must then be tested, monitored, explained, and supported. Simpler execution paths do not just reduce inference costs; they reduce the organizational burden of making AI safe and useful at scale.

Conclusion: Ask what the system already knows

If your AI bill is rising, cost is only the most visible symptom. The more important question is how often your systems are reasoning through something your organization has already established, verified, and paid to determine.

In my experience, the answer is far more often than most leaders realize. That repeated reasoning drives cost, but it also introduces inconsistency into decisions that should be repeatable, traceable, and defensible.

Enterprises should not measure AI maturity by how many tokens they consume, how many pilots they launch, or how many agents they deploy. They should measure whether the same facts and rules produce the same material decision, how reliably AI applies what the organization knows, how selectively it reasons when the answer is not known, and how much trusted business value it produces in return.

The first phase of enterprise AI was about proving organizations could scale consumption. The next will be about proving they can scale intelligence without scaling cost, inconsistency, and risk along with it. In a regulated enterprise, intelligence that cannot be reproduced cannot be trusted.

Learn More

RLMOpt Explained: A Smarter, Adaptive Way to Optimize AI Prompts

Building Healthcare AI That Works in the Real World

Request a Meeting

About the Author

Ganesh Padmanabhan is the CEO and founder of Autonomize AI, where he is building AI-native solutions for healthcare enterprises. With over 20 years of experience across AI, cloud, and enterprise transformation, he focuses on embedding governed intelligence into workflows to improve outcomes and reduce operational friction. A recognized voice in responsible AI and healthcare innovation, Ganesh is a frequent speaker, advisor, and advocate for using technology to tackle humanity’s biggest challenges.