Join our Newsletter — 33% off our NHI Course
Home› FAQ› AI Security› What breaks when an evaluator only sees the…
AI Security

What breaks when an evaluator only sees the agent response and not the full request or tool evidence?

← Back to all FAQ
By NHI Mgmt Group Editorial Team Updated September 30, 2026 Domain: AI Security

Evaluation becomes incomplete when the judge cannot see the original request, the system prompt, or supporting tool evidence. A response may look adequate in isolation while missing the actual user ask, and a correct refusal may be scored as failure if the criterion only checks resolution. For outcomes like refunds or account changes, the evidence must include the relevant transaction or tool result.

Why partial evaluation breaks the judgment model

An evaluator needs the full request context to judge whether a response actually answered the task. If the judge only sees the agent output, it cannot tell whether the agent followed the user’s intent, handled hidden constraints, or refused appropriately. That creates a false sense of quality: fluent text can score well even when it is incomplete, misdirected, or irrelevant to the real ask.

Complete evaluation depends on the same context the agent had at decision time. The original request, system prompt, and any tool evidence are what make the output interpretable. Without them, the evaluator is scoring a surface artifact rather than the action that produced it.

When the task involves a transaction, account change, or other stateful action, the tool result is part of the answer itself. A response that says an action succeeded or failed should be judged against the underlying evidence, not against wording alone. That is the difference between evaluating language and evaluating execution.

Why good refusals and good outcomes can both be mis-scored

A refusal can be the correct behavior even when it looks unhelpful in isolation. If the evaluator cannot see the original ask, a refusal may appear evasive when it was actually the safest or most policy-compliant outcome. The inverse also happens: an answer can sound complete while omitting the actual user objective, because the missing context was never visible to the judge.

This is especially important when the request includes hidden dependencies or tool-mediated evidence. For example, a user may ask for a refund, access change, or account update, but the real quality signal is whether the agent verified the relevant record before acting. If that evidence is absent, the evaluator cannot distinguish a correct resolution from a plausible hallucination.

The practical failure is not only incorrect scoring, but also poor feedback loops. If the evaluation harness rewards polished text without context, agents are trained toward confident narration instead of correct task completion.

What evaluators should compare, not just inspect

Good evaluation compares the response against the full task chain, not against the response in isolation. The judge should be able to see at least four things: the original request, the agent’s instruction context, the tools or evidence used, and the final output. That lets the evaluator check whether the response is faithful, whether any refusal was justified, and whether a state-changing action was supported by evidence.

For outcome-based workflows, the relevant question is usually not “Does the answer sound right?” but “Does the evidence support the outcome?” If the answer cannot be grounded in a tool result, transaction log, or other authoritative signal, the score should not treat it as proven success.

That same rule applies when the agent is judged on helpfulness, safety, or compliance. A concise but context-correct response is better than a detailed answer that looks convincing while missing the actual instruction.

Risk and Threat Considerations

Partial visibility creates a measurement risk: bad answers can look good, and good refusals can look bad. Over time, that produces distorted benchmarks, weak regression testing, and poor operator trust in the evaluation process.

Failure mechanism: The evaluator only sees the final prose, so it cannot verify intent alignment, refusal correctness, or whether a tool-backed action actually occurred. This lets missing context, fabricated confidence, and unsupported claims pass as acceptable output.

Impact: Teams may ship agents that optimize for plausible language instead of correct task completion, while real failures in refunds, account changes, access updates, or other stateful actions remain hidden until production.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST SP 800-53 Rev 5 sets the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
NIST SP 800-53 Rev 5AU-10 — Non-repudiationContext and evidence are needed to verify who did what and whether an action really occurred.
AU-6 — Audit Record Review, Analysis, and ReportingJudging agent outcomes requires reviewing records that show request, tool use, and result.
CA-7 — Continuous MonitoringEvaluation quality depends on continuous observation of actions and outcomes, not isolated text.
Recommendation — Retain tool and transaction evidence so evaluators can verify outcomes and non-repudiation. Review audit records alongside the response to confirm the action matches the request. Continuously monitor task traces and evidence, not just final responses.
OWASP Agentic AI Top 10ASI03 — Identity & Privilege AbuseCorrect evaluation must see whether an agent acted with the right authority and evidence.
ASI09 — Human-Agent Trust ExploitationA polished response can mislead evaluators when they cannot see the full request context.
Recommendation — Check per-action authority and evidence before scoring an agent outcome. Require full context so fluent responses do not mask incorrect or unsupported actions.

Practitioner Guidance

What to verify: Build the evaluation record so every scored item includes the user request, system context, tool inputs or outputs, and the final response. If the workflow can change state, retain the transaction evidence alongside the text so the judge can validate the result.

Decision rule: If a response is being scored on correctness, do not accept text-only grading for any task where context or evidence changes the meaning of the answer. If the task is purely stylistic, text-only review may be enough; if it is operational, it is not.

Practitioner takeaway: The safest evaluator is not the one that reads the most polished answer, but the one that can reconstruct the decision from request to evidence to outcome.

Deepen Your Knowledge

Sign up to our weekly newsletter — get 33% off our NHI Foundation Level Course

    NHIMG Editorial Note
    Reviewed and updated by the NHIMG editorial team on September 30, 2026.
    NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org