Agent Bench

Resend Agent Bench results

One task run and one autonomous setup run, reviewed as separate forms of evidence.

Reviewed Aug 17, 2026 2 recorded runs Tested by All Them APIs
B75/100

Verification passed

The task completed and the API was easy to discover, but test-mode sending restrictions caused repeated corrections and a relatively high error count.

Evaluation
65
Discovery
100
Time
4m 24s
Cost
$0.64
Calls
38
Errors
11

Where the agent spent its time

The trace divides the run into discovery, tool search, installation, coding, execution, and debugging phases where applicable.

01Discovery
11 calls~1m 50s
02Tool search
1 calls~3s
03Code
8 calls~26s
04Run
15 calls~1m 39s
05Debug
3 calls~26s

What this result means

Passed the Agent Bench task evaluation, ranked first in the setup run, and publishes the broadest first-party agent integration path in this set.

The evaluator recorded a pass. Review the caveat before treating it as semantic success.

Setup run completed

This dataset measures setup time, cost, errors, and human interruption. It is separate from the task-performance score.

Category rank
#1
Time
56s
Cost
$0.54
Errors
0
Interruptions
1
Outcome
Completed

Run completed and preserved

This page reports the recorded All Them APIs run, including failures, interruptions, and caveats.

Resend API profile