Agent Bench
Agent Bench result
Resend Agent Bench results
One task run and one autonomous setup run, reviewed as separate forms of evidence.
B75/100
Task run · 2026-05-20
Verification passed
The task completed and the API was easy to discover, but test-mode sending restrictions caused repeated corrections and a relatively high error count.
- Evaluation
- 65
- Discovery
- 100
- Time
- 4m 24s
- Cost
- $0.64
- Calls
- 38
- Errors
- 11
Run trace
Where the agent spent its time
The trace divides the run into discovery, tool search, installation, coding, execution, and debugging phases where applicable.
01Discovery
11 calls~1m 50s
02Tool search
1 calls~3s
03Code
8 calls~26s
04Run
15 calls~1m 39s
05Debug
3 calls~26s
Editorial interpretation
What this result means
Passed the Agent Bench task evaluation, ranked first in the setup run, and publishes the broadest first-party agent integration path in this set.
The evaluator recorded a pass. Review the caveat before treating it as semantic success.
Autonomous setup run
Setup run completed
This dataset measures setup time, cost, errors, and human interruption. It is separate from the task-performance score.
- Category rank
- #1
- Time
- 56s
- Cost
- $0.54
- Errors
- 0
- Interruptions
- 1
- Outcome
- Completed
Benchmark record
Run completed and preserved
This page reports the recorded All Them APIs run, including failures, interruptions, and caveats.