Agent Bench

Resend Agent Bench results

The task completed and the API was easy to discover, but test-mode sending restrictions caused repeated corrections and a relatively high error count.

Reviewed Aug 17, 2026 2 public datasets Not an All Them APIs rerun
B75/100

Verification passed

The task completed and the API was easy to discover, but test-mode sending restrictions caused repeated corrections and a relatively high error count.

Evaluation
65
Discovery
100
Time
4m 24s
Cost
$0.64
Calls
38
Errors
11

Where the agent spent its time

The source report divides the run into discovery, tool search, installation, coding, execution, and debugging phases where applicable.

01Discovery
11 calls~1m 50s
02Tool search
1 calls~3s
03Code
8 calls~26s
04Run
15 calls~1m 39s
05Debug
3 calls~26s

Integration assets the evaluator found

A strong checklist improves discovery, but it does not guarantee task success. Nylas is the clearest example in this dataset.

Context7llms.txtMCPTyped SDKOpenAPIAgent skillsCLI
Checklist
100
Tokens
782,135
Outcome
Verification passed

What this result means

Passed the public task evaluation, ranked first in the public setup run, and publishes the broadest first-party agent integration path in this set.

The source evaluator recorded a pass. Review the caveat before treating it as semantic success.

Setup run completed

This dataset measures setup time, cost, errors, and human interruption. It is separate from the task-performance score.

Category rank
#1
Time
56s
Cost
$0.54
Errors
0
Interruptions
1
Outcome
Completed

Independent rerun pending

This page summarizes recorded public evidence. We will mark the report when All Them APIs completes its own controlled rerun.

Resend API profile