Agent Bench
Agent Bench result
AgentMail Agent Bench results
One task run and one autonomous setup run, reviewed as separate forms of evidence.
C72/100
Task run · 2026-05-19
Verification passed
The agent recovered from an inbox limit and a bounced recipient, then sent to the test inbox itself. The run passed, although the workaround changed the requested recipient behavior.
- Evaluation
- 67
- Discovery
- 86
- Time
- 5m 4s
- Cost
- $0.33
- Calls
- 18
- Errors
- 5
Run trace
Where the agent spent its time
The trace divides the run into discovery, tool search, installation, coding, execution, and debugging phases where applicable.
01Discovery
3 calls~1m 28s
02Tool search
1 calls~8s
03Code
5 calls~45s
04Run
5 calls~1m 38s
05Debug
4 calls~1m 5s
Editorial interpretation
What this result means
Passed the Agent Bench task evaluation and ships MCP, CLI, llms.txt, OpenAPI, and agent skills; the passing workaround keeps it below Resend.
The evaluator recorded a pass. Review the caveat before treating it as semantic success.
Autonomous setup run
Setup run completed
This dataset measures setup time, cost, errors, and human interruption. It is separate from the task-performance score.
- Category rank
- #4
- Time
- 1m 53s
- Cost
- $0.84
- Errors
- 4
- Interruptions
- 1
- Outcome
- Completed
Benchmark record
Run completed and preserved
This page reports the recorded All Them APIs run, including failures, interruptions, and caveats.