TL;DR: Building two eval systems for AI developer tools, one for a Claude Agent SDK CLI and one for context-loaded skills, showed why pass rates, transcript review, and domain-specific scoring mattered more than intuition or generic “helpfulness” metrics, according to WorkOS. The lesson is that non-deterministic AI outputs need outcome-based governance, not test-style expectations.
At a glance
What this is: This article explains how WorkOS evaluated two AI developer tools and found that pass rates, transcript review, and domain-specific scoring exposed real performance differences that intuition missed.
Why it matters: IAM and AI platform teams need outcome-based evals because autonomous or semi-autonomous tooling can look correct while still producing noisy, harmful, or low-value behaviour in production workflows.
Context
LLM evals for agent tools become necessary when output is non-deterministic and success cannot be judged by whether code merely runs. In this case, the governance problem was not basic functionality but whether an AI-powered developer workflow actually improved the developer outcome across repeated runs.
That is why generic helpfulness checks fail for identity and AI-adjacent workflows: they do not tell you whether the tool preserved correctness, reduced hallucination, or followed framework-specific constraints. For teams building agentic or context-augmented systems, the control question is whether the measured outcome matches the operational intent.
The article’s central lesson is that evaluation design must reflect the task being governed. A CLI agent that edits code and a context-loaded skill that shapes model output need different scoring, but both require measurable evidence rather than intuition.
Key questions
Q: How can organisations know whether AI-assisted finding tools are actually helping?
A: Measure whether they reduce time from validated finding to verified risk reduction. If they only increase alert volume, they are adding overhead. The right signals are fewer duplicate tickets, faster owner assignment, and proof that exposure dropped after remediation, not just that a scan was completed.
Q: What should security teams do first when building LLM evals?
A: Define the smallest pass/fail outcome that proves the workflow worked, then test it across several realistic cases. Once the baseline is stable, add quality dimensions such as completeness, concision, and hallucination avoidance. Starting with a complex rubric before you know the outcome usually creates noise instead of confidence.
Q: Why do transcript reviews matter in AI evals?
A: Scores tell you whether performance changed, but transcripts show why. When a model regresses, the run log reveals whether the issue was missing steps, distracting context, tool misuse, or a bad rubric. Without transcripts, teams can only guess at root cause and often fix the wrong thing.
Q: What is the difference between pass rates and helpfulness scores?
A: Pass rates measure whether a task met a defined operational threshold across many runs. Helpfulness scores are softer judgments that can obscure failure modes, especially when the model produces plausible but incomplete output. For AI evals, pass rates are better for governance because they are tied to observable outcomes, not vibes.
Technical breakdown
Why deterministic tests fail for LLM-driven tools
Traditional tests assume a stable input-output pair. LLM-driven tools do not behave that way because the same prompt can produce different file changes, different wording, and different sequencing on every run. That makes snapshot-style assertions brittle. The more useful unit of analysis is a scenario with a defined outcome, then repeated measurement across many runs. In practice, evals become statistical controls, not binary correctness checks. They tell you whether the distribution of results stays within an acceptable band, whether retries improve quality, and whether the tool is drifting from the desired behaviour.
Practical implication: Use repeated scenario scoring instead of exact-output assertions when judging agentic or context-sensitive workflows.
How outcome-based grading works for AI tool evals
Outcome-based grading separates functional pass/fail from qualitative judgment. In the article, one layer checked concrete project state changes such as required files, imports, and successful builds. A second layer scored code style, minimalism, error handling, and idiomatic fit using rubric-based scoring. That structure matters because a technically valid result can still be operationally poor. The evaluator is not asking how the model arrived at the answer. It is asking whether the final artefact meets the standard that a human reviewer would accept. This is especially important for AI tools embedded in engineering workflows, where correctness and maintainability are not the same thing.
Practical implication: Score both functional correctness and output quality so a working result does not mask a harmful or noisy one.
Why transcripts matter more than scores alone
Scores show that something changed, but transcripts show why it changed. The article’s breakthrough came when side-by-side run logs revealed that a context skill was introducing noise and distracting the model from the core task. That kind of diagnosis is impossible if you only keep aggregates. Transcript review also helps catch cases where the model appears to improve on one dimension while silently regressing on another, such as ordering, completeness, or hallucination avoidance. For eval programmes, transcript retention is not a nice-to-have audit trail. It is the evidence base that lets teams distinguish useful guidance from harmful prompt interference.
Practical implication: Keep full transcripts for every eval run so regressions can be traced to the actual model behaviour, not just the score.
NHI Mgmt Group analysis
Outcome-based evals are now a governance control, not a product hygiene exercise. When LLM output is non-deterministic, intuition cannot establish whether a workflow is safe, useful, or merely plausible. The article shows that pass rates, retry thresholds, and rubric-based scoring create the only defensible evidence of quality. For identity teams, that shifts evals from a software quality concern to a control that determines whether automated assistance is trustworthy enough to use.
Generic usefulness metrics fail because they do not map to the actual security or delivery outcome. A model can be conversationally strong and still produce the wrong imports, wrong flow, or wrong access logic. That means the governance question is not whether the tool sounds helpful, but whether it produces the correct result within the intended boundary. Teams that rely on vague satisfaction metrics will miss failure modes that matter to deployment, access, and developer trust.
Trust is a measured property, not a subjective judgment. The strongest signal in this article is the discipline of comparing with-skill and without-skill runs, then looking for a measurable delta. That pattern is highly transferable to identity governance, where teams often assume a control is helpful because it feels useful. The practitioner standard should be simple: if a control or context layer cannot demonstrate improvement over baseline, it is adding risk as much as value.
Named concept: eval lift debt. This is the gap between a system that appears to add value and a system that proves it through repeated scoring. The article demonstrates that context can be additive, neutral, or actively harmful, and only eval design reveals which case applies. Practitioners should treat every AI workflow, prompt package, or skill bundle as guilty until the data shows lift.
AI-assisted development needs domain-specific evidence, not generic model confidence. The article’s most useful lesson is that the evaluation criteria must mirror the real task, not the model’s self-assessment. In identity and access programmes, the same principle applies to automation around provisioning, policy generation, and developer assistance: if the score does not reflect the operational outcome, it is measuring the wrong control.
From our research library:
- AI-related credential leaks surged 81.5% year-over-year in 2025, with the surrounding AI infrastructure leaking 5x faster than core LLM providers, according to the State of Secrets Sprawl 2026.
- Read next: Agentic AI Security Guide
What this signals
Eval lift debt: context layers, prompt packages, and skills should be treated as unproven until they show measurable lift over baseline. In practice, that means every AI workflow needs a comparison arm, because apparently useful context can still pull the model off task.
A mature programme should expect mixed results from added context. Some skills will improve accuracy, some will do nothing, and some will actively hurt output quality. The governance lesson is to retire anything that does not clear a measured threshold, even when the content looks correct to humans.
For practitioners
- Define pass rates before adding quality rubrics Start with a binary question such as whether the workflow completed the required outcome, then add thresholds for first-attempt success, correction success, and retry success only after the baseline is stable.
- Separate functional checks from quality scoring Use hard checks for required artefacts, correct imports, and build success, then score style, minimalism, and idiomatic fit in a separate rubric so one dimension does not hide another.
- Save full transcripts for every run Retain the raw prompts, tool calls, and outputs for both successful and failed runs so regressions can be explained rather than guessed.
- Compare baseline and assisted output Run the same task with and without the extra context or skill, then measure the delta to decide whether the added material actually improves the result.
Key takeaways
- LLM evals need to prove task outcomes, not just produce plausible-looking outputs.
- Domain-specific scoring exposed regressions that generic helpfulness metrics missed.
- Teams should treat context, skills, and prompt packages as governed inputs that must demonstrate lift over baseline.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST AI RMF and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI06 — Memory & Context Poisoning | The article shows added context can distort agent output and reduce task quality. |
| ASI02 — Tool Misuse | The CLI agent’s tool calls and flow correctness are central to the eval design. | |
| Recommendation — Evaluate added context for unintended influence on agent outputs and remove prompts that degrade task performance. Score whether the agent uses tools in the intended order and refuses hallucinated or off-scope actions. | ||
| NIST AI RMF | MEASURE — AI Measurement and Monitoring | The article is fundamentally about measuring AI quality with repeatable thresholds and rubrics. |
| Recommendation — Define measurable acceptance criteria and monitor whether AI outputs stay within approved performance bands. | ||
| NIST CSF 2.0 | GV.OV-01 — Organizational cybersecurity risk management strategy is informed by risk oversight activities | Eval thresholds and baseline comparison are governance activities for AI-enabled workflows. |
| Recommendation — Use oversight metrics to decide whether AI-assisted workflows are safe enough to ship or need rollback. | ||
Key terms
- Evaluation Threshold: An evaluation threshold is the minimum performance level a model must reach before a team treats it as ready for a specific use case. It gives product and engineering teams a concrete decision point, turning model assessment into a repeatable governance control instead of a subjective judgment.
- Outcome-Based Grading: Outcome-based grading judges the final state of a task rather than the steps used to reach it. For AI agents, that means assessing the resulting code, files, or configuration against task-specific criteria, because the internal reasoning path is often less useful than the observable result.
- Transcript Retention: The practice of saving prompts, tool calls, intermediate outputs, and final responses from an AI run. It creates the evidence needed to explain regressions, compare baselines, and distinguish a genuinely better result from a noisy or misleading one.
- Baseline Comparison: A controlled comparison between a model run with added context, skill, or tooling and a run without it. The comparison shows whether the added material actually improves the result or merely increases complexity, noise, or false confidence.
Deepen your knowledge
NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
Published by the NHIMG editorial team on June 7, 2026.
Updated on October 7, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org