---
title: Agent Seek Eval
canonical: /eval.md
---

# Agent Seek Eval

We asked 50 questions.

We checked whether Agent Seek's short list was preferred to the same pages in You.com's order. We also checked whether the known answer was in the top three results.

## Snip preference

**74% preferred** (37/50).

Agent Seek's short list was preferred to the same pages in You.com's order.

Of 50 questions:

- **Wins → 37** (74% of 50)
- **Ties → 0** (0% of 50)
- **Losses → 13** (26% of 50)

Preferred on 37 of 50 questions.

## Top-3 hit rate

**86%** (43/50).

The known answer was in the top three results.

## How we measured

Both published runs measure Agent Seek against same-pool You.com order: `mode=snip` (product default — UI and API/MCP) and `mode=deep` (optional richer mode; eval treatment).

Of n, a win rate is that side's wins divided by judged queries, and ties count as non-wins. Excl. ties drops ties and uses decided pairs only. When a run has no ties, those two rates match. Each win rate shows a Wilson 95% CI and a two-sided exact binomial p-value for H0: p=0.5.

Snip preference is 74% preferred (37/50; Wilson 60.4–84.1%; p=0.0009). Blinded pairwise preference vs same-pool You.com order, not top-2 sufficiency.

Judge gpt-5.6-sol · reasoning.effort=medium · n=50.

Artifact: [`evals/public_v1/published/v1.0.0-snip/summary.json`](/evals/public_v1/published/v1.0.0-snip/summary.json).

Top-3 hit rate is 86% (43/50; Wilson 73.8–93.0%; p=<0.0001). Misses 7. Errors 0. n=50.

Judge gpt-5.6-sol · reasoning.effort=medium · n=50. k=3. A query hits when any successfully judged Agent Seek snip top-3 page states the locked gold answer. Queries whose page calls all fail are errors, not misses.

Artifact: [`evals/public_v1/published/v1.0.0-snip-top3/summary.json`](/evals/public_v1/published/v1.0.0-snip-top3/summary.json).

## Deep preference

With more text from each page, Agent Seek was **64% preferred** (32/50).

Deep preference is 64% preferred (32/50; Wilson 50.1–75.9%; p=0.065). Excluding ties, that is 67% (32/48; Wilson 52.5–78.3%; p=0.029).

Judge gpt-5.6-sol · reasoning.effort=medium · n=50. `mode=deep` (optional richer mode).

Artifact: [`evals/public_v1/published/v1.0.0-first-run/summary.json`](/evals/public_v1/published/v1.0.0-first-run/summary.json).

These figures come from saved runs. Opening this page does not call OpenAI, You.com, or TypeSafe.

## Diagnostics

Snip (product default) flip rate 14.0% (7/50).

Shortlists in this pass were the published URLs (titles and snippets were not stored).

## Per-query

The frozen query set is balanced across an L1–L3 difficulty mix for coverage. Tier-level win rates are not published because those cells are underpowered at the current n (~15–20 per tier).

Per-query results (provider URLs, snippets, and judge rationales) are not stored in this repo. The figures above come from the published summary.json files. Regenerate a run locally with scripts/eval_llm_judge.py to inspect rows.