Agent Seek Eval

We asked 50 questions.

We checked whether Agent Seek's short list was preferred to the same pages in You.com's order. We also checked whether the known answer was in the top three results.

Snip preference

74% preferred 37/50

Agent Seek's short list was preferred to the same pages in You.com's order.

Of 50 questions

  1. Wins3774% of 50
  2. Ties00% of 50
  3. Losses1326% of 50

Preferred on 37 of 50 questions.

Top-3 hit rate

86% 43/50

The known answer was in the top three results.

How we measured

Both published runs measure Agent Seek against same-pool You.com order: mode=snip (product default — UI and API/MCP) and mode=deep (optional richer mode; eval treatment).

Of n, a win rate is that side's wins divided by judged queries, and ties count as non-wins. Excl. ties drops ties and uses decided pairs only. When a run has no ties, those two rates match. Each win rate shows a Wilson 95% CI and a two-sided exact binomial p-value for H0: p=0.5.

Snip preference is 74% preferred (37/50; Wilson 60.4–84.1%; p=0.0009). Blinded pairwise preference vs same-pool You.com order, not top-2 sufficiency. Judge gpt-5.6-sol · reasoning.effort=medium · n=50. Artifact: summary.json.

Top-3 hit rate is 86% (43/50; Wilson 73.8–93.0%; p=<0.0001). Misses 7. Errors 0. n=50. Judge gpt-5.6-sol · reasoning.effort=medium · n=50. k=3. A query hits when any successfully judged Agent Seek snip top-3 page states the locked gold answer. Queries whose page calls all fail are errors, not misses. Artifact: summary.json.

With more text from each page, Agent Seek was 64% preferred (32/50).

Deep preference is 64% preferred (32/50; Wilson 50.1–75.9%; p=0.065). Excluding ties, that is 67% (32/48; Wilson 52.5–78.3%; p=0.029). Judge gpt-5.6-sol · reasoning.effort=medium · n=50. mode=deep (optional richer mode). Artifact: deep summary.json.

These figures come from saved runs. Opening this page does not call OpenAI, You.com, or TypeSafe.

Diagnostics

Snip (product default) flip rate 14.0% (7/50).

Shortlists in this pass were the published URLs (titles and snippets were not stored).

All 50 questions

The frozen query set is balanced across an L1–L3 difficulty mix for coverage. Tier-level win rates are not published because those cells are underpowered at the current n (~15–20 per tier).

Per-query results (provider URLs, snippets, and judge rationales) are not stored in this repo. The figures above come from the published summary.json files. Regenerate a run locally with scripts/eval_llm_judge.py to inspect rows.