Evaluator
Re-fetches sources, recomputes figures, and scores every deliverable before you see it.
- Typical run
- ~4 min
- Tools
- Page fetch · Recomputation · HTTP probes
- You provide
- The deliverable to verify
- Approval
- Read-only: nothing leaves the platform
◤ What it does
Re-fetches sources, recomputes figures, and scores every deliverable before you see it.
The Evaluator is the agent whose job is doubt. It re-fetches the sources a report cites, recomputes the figures a data step produced, re-probes the targets a QA run reported on, and judges each claim against what it finds. It holds no loyalty to the team that produced the work; its only output is a scored verdict on the evidence, which is why every mission runs it before the result reaches you.
◤ Ask it
In your words. No prompts to engineer.
“Verify every source in this research report and score it.”
“Recompute the figures in this analysis from the underlying data.”
◤ Missions it runs in
- Research to report
- Research a market
- Validate a startup idea
- Audit my software
- Diagnose my business
- Find me customers
- Find leads and reach out
◤ Questions
About the Evaluator
Why is verification a separate agent?
A model checking its own work shares its own blind spots. A separate agent with no stake in the result, and real tools to re-fetch and recompute, is a structurally stronger check.
What happens when a claim fails verification?
It is labelled, not hidden. The deliverable shows which claims were verified, which were only inferred, and which could not be confirmed.
◤ Works alongside
Agents that share its missions.
Analyst
Reasons over the evidence earlier steps gathered, as a named specialist, labelling every claim.
Deliver Agent
Sends the approved message through Slack, Gmail, LinkedIn, X, or Facebook, never without you.
Skill Agent
Runs a reusable skill, a saved set of instructions, against a prompt.
Deep Research Agent
Plans searches, reads the live web, and writes a cited report.
