Cost per correct answer is a practical evaluation metric that combines API charges, token usage, extraction costs, and model inference costs, then divides the total by the number of correct responses. It helps teams compare retrieval setups on business value, not just raw accuracy.
How Cost Per Correct Answer Works
Cost per correct answer is a decision metric, not just an accuracy metric. It combines the full run cost of a retrieval or generation setup, including API charges, token consumption, extraction work, and inference spend, then divides by the number of correct answers to show what useful output really costs.
This matters because two systems with the same accuracy can have very different economics. A cheaper prompt, a smaller model, better chunking, or a cleaner retrieval path can improve cost per correct answer even when headline accuracy barely changes.
What It Reveals About System Trade-offs
The metric exposes the trade-off between quality and efficiency. A setup that is slightly less accurate but dramatically cheaper may be the better business choice, while a high-accuracy setup that burns tokens or makes repeated model calls can be poor value.
It also helps teams separate hidden cost drivers. Retrieval depth, reranking, repeated tool calls, and large context windows can all increase cost even when they appear to improve answer quality only marginally.
How Teams Should Interpret the Metric
Cost per correct answer is most useful when compared across the same task set, the same evaluation standard, and the same definition of correctness. Without that consistency, the number can mislead because a system can look efficient simply by answering fewer hard questions or by changing what counts as correct.
The metric becomes stronger when paired with supporting measures such as latency, coverage, and failure modes. That combination helps teams understand whether a lower cost reflects genuine efficiency or simply a narrower, less capable system.
Where the Metric Is Most Useful
This metric is especially valuable in retrieval-augmented workflows, support automation, and other settings where each answer may trigger multiple paid calls or preprocessing steps. It gives product, engineering, and finance teams a shared way to judge whether a design choice improves business value, not just model scores.
It is also useful for comparing vendors or architectures, because it normalizes cost against correctness instead of treating price and quality as separate conversations. That makes it easier to identify when a modest quality gain is worth the additional spend.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST CSF 2.0 provides the primary governance reference for this term.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST CSF 2.0 | GV.RM-01 — Risk Management Strategy | Connects cost efficiency with risk and value trade-offs in security decisions |
| GV.OV-01 — Oversight of the Cybersecurity Program | Supports oversight of how metrics inform program decisions and investment priorities | |
| ID.RA-01 — Asset Vulnerabilities Are Identified and Recorded | Applies when evaluation cost must be balanced against quality and coverage in analysis pipelines | |
| Recommendation — Define cost-value evaluation criteria for controls and automation choices. Use metrics like cost per correct answer to inform oversight and spending decisions. Record evaluation assumptions and cost drivers alongside measured outcomes. | ||
Related resources from NHI Mgmt Group
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on September 30, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org