Software optimisation creates more value because it changes how much work each prompt requires, while hardware upgrades mostly improve capacity at the margin. The report cites 23x gains from model architecture improvements versus 1.4x from hardware utilisation alone. That difference matters when budgets are rising quickly and inference volume keeps expanding.
Why software-level optimisation outperforms hardware upgrades for AI workloads
Software changes the economics of every request. Better model architecture, token handling, batching, caching, quantisation, routing and scheduling reduce the work needed per prompt, so the same hardware does more useful inference. Hardware upgrades still matter, but they usually improve throughput or latency only at the margin unless the software stack can already use that capacity efficiently.
The practical lesson is that AI spend is dominated by recurring inference cost, not one-time cluster refreshes. When the workload grows faster than the infrastructure budget, improving software efficiency compounds across every request, while a larger accelerator fleet only stretches the same inefficiency over more machines.
That is why optimisation at the model and serving layer often produces a larger return: it lowers unit cost, reduces queue pressure, and extends the useful life of existing infrastructure without forcing a platform rewrite.
Where the value comes from in practice
Software optimisation creates value by reducing wasted computation. A smaller or more efficient model can answer with fewer FLOPs, a smarter serving layer can batch requests and avoid idle time, and a better retrieval or routing strategy can send only the hard cases to expensive models. Each of those changes improves cost per request, not just raw capacity.
Hardware upgrades mainly help when the system is already close to saturation. If the bottleneck is memory bandwidth, context length, or inefficient execution, a faster GPU may still leave much of the cost structure intact. The gain is real, but it is bounded by the underlying software path that determines how much work each prompt triggers.
For AI teams, this distinction matters because optimisation targets the full request path, from prompt preparation to model selection to inference serving. The most effective improvements usually come from removing unnecessary work before asking for more compute.
Why the advantage compounds as inference scales
Inference is a volume business. Once usage grows across products, internal tools, and agentic workflows, even small savings per request create large aggregate effects. That makes software gains multiplicative: a 10% reduction in tokens, a 15% improvement in batching, and a better fallback policy all apply across every call, every day.
Hardware scaling does not compound in the same way. Additional chips or instances increase the ceiling, but they do not change the workload shape, the number of tokens generated, or the amount of orchestration overhead. As a result, hardware often shifts the cost curve upward, while software optimisation bends the curve downward.
That is especially important in AI deployments where latency, reliability, and cost all matter at once. A more efficient serving design can improve all three, whereas raw hardware expansion often improves only one dimension and may increase operating complexity elsewhere.
What leaders should optimise first
Start with the parts of the stack that determine how much compute each prompt consumes. Model architecture, prompt design, routing, caching, batching, quantisation, and serving policy usually offer the best first-order savings. Only after those levers are mature does hardware expansion become the more attractive lever for a specific bottleneck.
Hardware is still the right answer when the system is genuinely compute-bound after optimisation, or when a workload requires more parallel capacity than the current fleet can provide. But buying more compute before fixing software inefficiency usually locks in a higher cost base and delays the point at which the platform becomes economical.
In other words, optimise the request path before scaling the machine pool. That sequence gives you lower unit cost, better elasticity, and more evidence about the true capacity limit before you commit to a larger infrastructure footprint.
Risk and Threat Considerations
When AI workloads are scaled by hardware alone, organisations can mask inefficiency instead of fixing it. The result is higher operating cost, faster budget burn, and a larger blast radius if demand spikes or a model change increases token usage unexpectedly.
Failure mechanism: Inefficient prompts, oversized models, poor batching, or weak routing keep consuming expensive compute on every request, so capacity expands faster than value delivered.
Impact: Cost per inference rises, performance tuning stalls, and teams may overprovision infrastructure simply to preserve service levels, which makes the environment harder to scale sustainably.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST SP 800-53 Rev 5, CIS Controls v8 and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST SP 800-53 Rev 5 | SC-38 — Operations Security | AI serving efficiency depends on controlled compute pathways and constrained execution. |
| Recommendation — Control execution paths to reduce unnecessary processing and resource waste. | ||
| CIS Controls v8 | CIS-12 — Network Infrastructure Management | AI workloads benefit from efficient resource and capacity management across the stack. |
| Recommendation — Tune infrastructure capacity and scheduling to support measured workload demand. | ||
| NIST CSF 2.0 | GV.RM-01 — Risk Management Strategy | The question is about choosing the higher-value cost-control strategy for AI operations. |
| Recommendation — Set investment priorities based on measured unit-cost reduction, not headline capacity. | ||
Practitioner Guidance
What to prioritise: Measure cost per successful inference, not just GPU utilisation. If token counts, cache hit rates, and batching efficiency are not visible, you are likely optimising the wrong layer.
Decision rule: If the workload still shows large variance in prompt size, model selection, or routing efficiency, treat software optimisation as the first capital-saving move; reserve hardware expansion for a proven capacity bottleneck.
What good looks like: The team can explain which software change reduced compute per request, how much it saved, and whether the saving held across production traffic rather than only in benchmark runs.
Practitioner takeaway: Hardware buys headroom, but software decides the unit economics, so the highest-value AI performance work is the work that reduces waste before adding more machines.
Related resources from NHI Mgmt Group
- Why do static secrets create more risk for AI agents than for traditional workloads?
- Why do AI workloads create a bigger identity risk than ordinary service accounts?
- When does agentic AI create more risk than value?
- Why do static credentials create more risk for AI agents than for traditional workloads?