Resend Agent Bench results
The task completed and the API was easy to discover, but test-mode sending restrictions caused repeated corrections and a relatively high error count.
Verification passed
The task completed and the API was easy to discover, but test-mode sending restrictions caused repeated corrections and a relatively high error count.
- Evaluation
- 65
- Discovery
- 100
- Time
- 4m 24s
- Cost
- $0.64
- Calls
- 38
- Errors
- 11
Where the agent spent its time
The source report divides the run into discovery, tool search, installation, coding, execution, and debugging phases where applicable.
Integration assets the evaluator found
A strong checklist improves discovery, but it does not guarantee task success. Nylas is the clearest example in this dataset.
- Checklist
- 100
- Tokens
- 782,135
- Outcome
- Verification passed
What this result means
Passed the public task evaluation, ranked first in the public setup run, and publishes the broadest first-party agent integration path in this set.
Setup run completed
This dataset measures setup time, cost, errors, and human interruption. It is separate from the task-performance score.
- Category rank
- #1
- Time
- 56s
- Cost
- $0.54
- Errors
- 0
- Interruptions
- 1
- Outcome
- Completed
Independent rerun pending
This page summarizes recorded public evidence. We will mark the report when All Them APIs completes its own controlled rerun.