Evaluation becomes incomplete when the judge cannot see the original request, the system prompt, or supporting tool evidence. A response may look adequate in isolation while missing the actual user ask, and a correct refusal may be scored as failure if the criterion only checks resolution. For outcomes like refunds or account changes, the evidence must include the relevant transaction or tool result.
Why partial evaluation breaks the judgment model
An evaluator needs the full request context to judge whether a response actually answered the task. If the judge only sees the agent output, it cannot tell whether the agent followed the user’s intent, handled hidden constraints, or refused appropriately. That creates a false sense of quality: fluent text can score well even when it is incomplete, misdirected, or irrelevant to the real ask.
Complete evaluation depends on the same context the agent had at decision time. The original request, system prompt, and any tool evidence are what make the output interpretable. Without them, the evaluator is scoring a surface artifact rather than the action that produced it.
When the task involves a transaction, account change, or other stateful action, the tool result is part of the answer itself. A response that says an action succeeded or failed should be judged against the underlying evidence, not against wording alone. That is the difference between evaluating language and evaluating execution.
Why good refusals and good outcomes can both be mis-scored
A refusal can be the correct behavior even when it looks unhelpful in isolation. If the evaluator cannot see the original ask, a refusal may appear evasive when it was actually the safest or most policy-compliant outcome. The inverse also happens: an answer can sound complete while omitting the actual user objective, because the missing context was never visible to the judge.
This is especially important when the request includes hidden dependencies or tool-mediated evidence. For example, a user may ask for a refund, access change, or account update, but the real quality signal is whether the agent verified the relevant record before acting. If that evidence is absent, the evaluator cannot distinguish a correct resolution from a plausible hallucination.
The practical failure is not only incorrect scoring, but also poor feedback loops. If the evaluation harness rewards polished text without context, agents are trained toward confident narration instead of correct task completion.
What evaluators should compare, not just inspect
Good evaluation compares the response against the full task chain, not against the response in isolation. The judge should be able to see at least four things: the original request, the agent’s instruction context, the tools or evidence used, and the final output. That lets the evaluator check whether the response is faithful, whether any refusal was justified, and whether a state-changing action was supported by evidence.
For outcome-based workflows, the relevant question is usually not “Does the answer sound right?” but “Does the evidence support the outcome?” If the answer cannot be grounded in a tool result, transaction log, or other authoritative signal, the score should not treat it as proven success.
That same rule applies when the agent is judged on helpfulness, safety, or compliance. A concise but context-correct response is better than a detailed answer that looks convincing while missing the actual instruction.
Risk and Threat Considerations
Partial visibility creates a measurement risk: bad answers can look good, and good refusals can look bad. Over time, that produces distorted benchmarks, weak regression testing, and poor operator trust in the evaluation process.
Failure mechanism: The evaluator only sees the final prose, so it cannot verify intent alignment, refusal correctness, or whether a tool-backed action actually occurred. This lets missing context, fabricated confidence, and unsupported claims pass as acceptable output.
Impact: Teams may ship agents that optimize for plausible language instead of correct task completion, while real failures in refunds, account changes, access updates, or other stateful actions remain hidden until production.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST SP 800-53 Rev 5 sets the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST SP 800-53 Rev 5 | AU-10 — Non-repudiation | Context and evidence are needed to verify who did what and whether an action really occurred. |
| AU-6 — Audit Record Review, Analysis, and Reporting | Judging agent outcomes requires reviewing records that show request, tool use, and result. | |
| CA-7 — Continuous Monitoring | Evaluation quality depends on continuous observation of actions and outcomes, not isolated text. | |
| Recommendation — Retain tool and transaction evidence so evaluators can verify outcomes and non-repudiation. Review audit records alongside the response to confirm the action matches the request. Continuously monitor task traces and evidence, not just final responses. | ||
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Correct evaluation must see whether an agent acted with the right authority and evidence. |
| ASI09 — Human-Agent Trust Exploitation | A polished response can mislead evaluators when they cannot see the full request context. | |
| Recommendation — Check per-action authority and evidence before scoring an agent outcome. Require full context so fluent responses do not mask incorrect or unsupported actions. | ||
Practitioner Guidance
What to verify: Build the evaluation record so every scored item includes the user request, system context, tool inputs or outputs, and the final response. If the workflow can change state, retain the transaction evidence alongside the text so the judge can validate the result.
Decision rule: If a response is being scored on correctness, do not accept text-only grading for any task where context or evidence changes the meaning of the answer. If the task is purely stylistic, text-only review may be enough; if it is operational, it is not.
Practitioner takeaway: The safest evaluator is not the one that reads the most polished answer, but the one that can reconstruct the decision from request to evidence to outcome.
Related resources from NHI Mgmt Group
- What breaks when an agent evaluator cannot see the full trace?
- What breaks when organisations map AI risk without a full agent and tool inventory?
- What breaks when AI agent access to ServiceNow is not inspected before the model sees the response?
- What breaks when long-running agent tasks are forced into a request-response pattern?
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on September 30, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org