Join our Newsletter — 33% off our NHI Course
Home› FAQ› Architecture & Implementation› Why do production agents become expensive even when…
Architecture & Implementation

Why do production agents become expensive even when they use the same model and tools?

← Back to all FAQ
By NHI Mgmt Group Editorial Team Updated September 30, 2026 Domain: Architecture & Implementation

The model usually sets the ceiling, but the harness determines how much work it takes to reach it. If the loop loads too much context, retries clumsily, or keeps unnecessary tool output inline, token usage climbs fast. Well-designed harness controls can cut cost without changing the underlying model, because they reduce wasted tokens and unnecessary execution steps.

Why the same model can still produce very different costs

The model is only one part of the bill. In production, the harness around it often dominates spend through prompt assembly, tool orchestration, retries, post-processing, and the amount of intermediate state kept alive from turn to turn. When those mechanics are inefficient, the agent pays for repeated reasoning work even if the underlying model and tools never change.

The practical implication is that “same model” does not mean “same workload.” Two agents can call the same endpoint but generate very different token footprints if one is disciplined about context growth, tool-output trimming, and stopping conditions while the other keeps appending everything it sees. That difference compounds quickly under real traffic.

Cost also rises when the harness turns uncertainty into extra steps. If the loop retries on weak signals, asks the model to restate work it already did, or over-serializes tool traces back into the next prompt, it creates self-inflicted token inflation. The expensive part is often not the answer, but the repeated path taken to get there.

Where token waste usually enters the production loop

Three patterns show up most often: oversized context windows, unnecessary tool chatter, and unstable control flow. Oversized context happens when every prior result, log, or scratch note is fed forward even though only a small subset is needed. Tool chatter happens when verbose tool responses are preserved inline instead of summarized or normalized. Unstable control flow happens when the harness cannot decide whether a result is good enough, so it keeps looping.

This is why cost can rise without any visible change in product feature use. A single request may trigger more prompt tokens, more completion tokens, and more execution passes than expected. If the agent is multi-step, the multiplier effect is severe: each extra step not only adds its own token use, it often increases the size of the next step as well.

Production agents are especially sensitive to observability and audit of agent actions because token spikes are frequently a symptom of hidden retries, loops, or unnecessary tool fan-out rather than model quality issues. Teams that can trace each step usually find the waste path faster than teams that only watch aggregate spend.

How to control cost without changing the model

The best savings come from controlling the harness before you try to “optimize the model.” Start by limiting what enters the prompt, then reduce what gets echoed back, and only then tune the step count. In practice that means summarizing tool output, pruning irrelevant history, and making the agent stop when it has enough confidence to act or answer.

A useful rule is to treat every extra token as a design decision. If a step exists only because the harness cannot safely bound the agent’s behavior, that step is likely a cost leak. If a tool response is useful but not needed verbatim, convert it into a compact representation rather than carrying the full payload forward.

For agent systems that can act through browsers or desktops, browser and computer-use agent controls matter because uncontrolled interaction can inflate both execution time and token usage. The same is true for task-scoped AI agent authorisation, where tighter permissions reduce pointless exploration and keep the agent from wandering into expensive, low-value actions.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Non-Human Identity Top 10 and OWASP Agentic AI Top 10 address the attack and risk surface, while NIST SP 800-53 Rev 5 sets the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
OWASP Non-Human Identity Top 10NHI-02 — Secret LeakageLong context and tool output often retain sensitive material, increasing spend and exposure.
Recommendation — Trim retained tool output and context so secrets do not propagate into later turns.
OWASP Agentic AI Top 10ASI02 — Tool MisuseWasteful retries and overuse of tools drive extra steps and token cost in agent loops.
ASI03 — Identity & Privilege AbuseOver-broad agent authority can trigger expensive, unnecessary actions and retries.
Recommendation — Constrain tool invocation paths to reduce unnecessary calls and repeated execution. Scope agent permissions tightly so the harness cannot fan out into avoidable work.
NIST SP 800-53 Rev 5AU-6 — Audit Record Review, Analysis, and ReportingMonitoring execution traces helps find token waste, retries, and loop pathologies.
CM-2 — Baseline ConfigurationA stable harness baseline is needed to control prompt growth and execution drift.
Recommendation — Review agent execution records to identify recurring cost inflation patterns. Baseline the agent harness and prevent uncontrolled prompt or workflow expansion.

Practitioner Guidance

What to measure: Track tokens per successful task, retries per task, tool calls per task, and the share of prompt text coming from history versus fresh user intent. Those four signals usually show where the harness is leaking cost before the total bill does.

Decision rule: If cost is rising faster than task volume, inspect the loop for repeated context expansion and retry churn before blaming the model. If the same task path regularly needs more turns than expected, redesign the stop condition or trim the tool contract.

Common mistake: Teams often optimize only the model choice or price tier, then leave verbose tool output and long-lived context untouched. That usually preserves the underlying inefficiency and makes the spend problem recur at scale.

Practitioner takeaway: In production, cost control is mostly harness control, the cheapest token is the one you never add to the next turn.

Deepen Your Knowledge

Sign up to our weekly newsletter — get 33% off our NHI Foundation Level Course

    NHIMG Editorial Note
    Reviewed and updated by the NHIMG editorial team on September 30, 2026.
    NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org