TL;DR: Building two eval systems for AI developer tools, one for a Claude Agent SDK CLI and one for context-loaded skills, showed why pass rates, transcript review, and domain-specific scoring mattered more than intuition or generic “helpfulness” metrics, according to WorkOS. The lesson is that non-deterministic AI outputs need outcome-based governance, not test-style expectations.
Editorial analysis by NHI Mgmt Group, based on content published by WorkOS: “Writing my first evals”.
Key questions
Q: How can organisations know whether AI-assisted finding tools are actually helping?
A: Measure whether they reduce time from validated finding to verified risk reduction.
Q: What should security teams do first when building LLM evals?
A: Define the smallest pass/fail outcome that proves the workflow worked, then test it across several realistic cases.
Q: Why do transcript reviews matter in AI evals?
A: Scores tell you whether performance changed, but transcripts show why.
Practitioner guidance
- Define pass rates before adding quality rubrics Start with a binary question such as whether the workflow completed the required outcome, then add thresholds for first-attempt success, correction success, and retry success only after the baseline is stable.
- Separate functional checks from quality scoring Use hard checks for required artefacts, correct imports, and build success, then score style, minimalism, and idiomatic fit in a separate rubric so one dimension does not hide another.
- Save full transcripts for every run Retain the raw prompts, tool calls, and outputs for both successful and failed runs so regressions can be explained rather than guessed.
Bottom line: LLM evals need to prove task outcomes, not just produce plausible-looking outputs.
Explore further
View Full Forum → | NHI Foundation Course → | Our Services → | Read the full analysis →
Outcome-based evals are now a governance control, not a product hygiene exercise. When LLM output is non-deterministic, intuition cannot establish whether a workflow is safe, useful, or merely plausible. The article shows that pass rates, retry thresholds, and rubric-based scoring create the only defensible evidence of quality. For identity teams, that shifts evals from a software quality concern to a control that determines whether automated assistance is trustworthy enough to use.
A few things that frame the scale:
- AI-related credential leaks surged 81.5% year-over-year in 2025, with the surrounding AI infrastructure leaking 5x faster than core LLM providers, according to the State of Secrets Sprawl 2026.
A question worth separating out:
Q: What is the difference between pass rates and helpfulness scores?
A: Pass rates measure whether a task met a defined operational threshold across many runs. Helpfulness scores are softer judgments that can obscure failure modes, especially when the model produces plausible but incomplete output. For AI evals, pass rates are better for governance because they are tied to observable outcomes, not vibes.
👉 Read our full editorial: LLM evals for agent tools: why pass rates beat intuition