Agent Evals
Weave publishes a public, read-only eval surface for its core agent behavior. The dashboard lives at /weave-agent-evals/ and shows sanitized results from published eval runs.
What the public eval surface includes
The public surface is the dashboard, its suite detail pages, and the published report artifacts behind them. It is meant to make Weave's current eval coverage visible without exposing private repos, raw harness state, or internal-only data.
Supported suite families
Weave currently publishes seven text-only suite families:
- Loom Routing (
loom-routing) - Tapestry Execution (
tapestry-execution) - Shuttle Execution (
shuttle-execution) - Pattern Planning (
pattern-planning) - Weft Review (
weft-review) - Warp Security (
warp-security) - Spindle Tools (
spindle-tools)
Current limitation
These public evals are text-only. They grade prompt-visible behavior and plain-text outputs from synthetic cases.
They do not claim runtime-backed verification of tool use, browser activity, network activity, repository side effects, or other live execution behavior.
What you can inspect
At a high level, the dashboard lets you inspect:
- overall suite status and pass rates
- recent runs and score trends over time
- suite-by-suite history
- model comparison views for published runs
- per-case outcomes and public report downloads
If you want the live surface first, go straight to /weave-agent-evals/.
