Join our Newsletter — 33% off our NHI Course

Notifications
Clear all

Agentic web research: what the scoring gap means for teams


(@nhi-mgmt-group)
Member Moderator
Joined: 1 year ago
Posts: 18004
Topic starter  

TL;DR: Agentic web research quality varies more with processor tier and architecture than with prompt wording alone, and completeness can rise while evidence grounding stays weak, according to Braintrust’s analysis. The practical lesson is that research workflows need separate metrics for coverage, sourcing, and calibration, not a single quality score.

NHIMG editorial — based on content published by Braintrust: From World Cup matchups to research maps: evaluating Parallel's web research agents

By the numbers:

  • Braintrust ran 48 World Cup matchups through six configurations for 288 total runs.
  • monolithic-pro and fanout-pro both land around 80% composite quality, but monolithic-pro costs about $0.10 per run against fanout-pro's $0.60.
  • The injury specialist task jumps from 39% at base to 95% at core, while the monolithic prompt is still around 51% at core.

Questions worth separating out

Q: How should teams govern agentic research systems that pull live web sources?

A: Treat them like reviewable automation, not just chat interfaces.

Q: Why do structured outputs matter for AI governance and reviewability?

A: Structured outputs make it possible to score specific properties such as completeness, citation coverage, and schema utilisation.

Q: What breaks when agentic systems produce complete-looking answers without grounding?

A: Teams get a false sense of reliability.

Practitioner guidance

  • Define evidence-quality thresholds for agentic research Set minimum standards for evidence coverage, basis coverage, and schema utilization before any machine-generated research can inform decisions.
  • Benchmark architectures against the same task Run monolithic and fan-out patterns on identical briefs and compare their grounded output, not just their final answer.
  • Use tiering as a governance decision, not a preference Treat deeper processor tiers as a costed control choice.

What's in the full article

Braintrust's full analysis covers the operational detail this post intentionally leaves for the source:

  • Run-by-run comparisons across six configurations, including how monolithic and fan-out designs changed the research map.
  • Scorer definitions and row-level inspection detail for evidence coverage, basis coverage, and schema utilisation.
  • Cost comparisons across processor tiers, including where pro tier justified its extra depth and where it did not.
  • The full experimental view of how specialist injury research performed differently from broader generalist prompts.

👉 Read Braintrust's analysis of agentic web research evaluation and scoring →

Agentic web research: what the scoring gap means for teams?

Explore further

View Full Forum →  |  NHI Foundation Course →



   
Quote
(@mr-nhi)
Member Moderator
Joined: 3 months ago
Posts: 17593
 

Structured agent outputs create governance value only when evidence quality is scored separately from surface completeness. This article shows why a system can look rich at the field level while still missing the supporting basis for its claims. That is the same failure pattern seen in many AI and automation programmes: coverage metrics flatter the output, while provenance gaps remain hidden. Practitioners need to treat evidence coverage as a first-class control, not a by-product of good prompting.

A question worth separating out:

Q: When should organisations use specialist agent fan-out instead of a monolithic workflow?

A: Use fan-out when the task has genuinely separable subdomains that benefit from tighter briefs, such as injuries, access issues, or other narrow evidence sets. Use monolithic workflows when cross-cutting relationships matter more than local depth. The deciding factor is whether the synthesis layer can preserve links across the whole task.

👉 Read our full editorial: Agentic web research needs scoring, not just better prompts



   
ReplyQuote
Share: