Agent Bench

AgentMail Agent Bench results

The agent recovered from an inbox limit and a bounced recipient, then sent to the test inbox itself. The run passed, although the workaround changed the requested recipient behavior.

Reviewed Aug 17, 2026 2 public datasets Not an All Them APIs rerun
C72/100

Verification passed

The agent recovered from an inbox limit and a bounced recipient, then sent to the test inbox itself. The run passed, although the workaround changed the requested recipient behavior.

Evaluation
67
Discovery
86
Time
5m 4s
Cost
$0.33
Calls
18
Errors
5

Where the agent spent its time

The source report divides the run into discovery, tool search, installation, coding, execution, and debugging phases where applicable.

01Discovery
3 calls~1m 28s
02Tool search
1 calls~8s
03Code
5 calls~45s
04Run
5 calls~1m 38s
05Debug
4 calls~1m 5s

Integration assets the evaluator found

A strong checklist improves discovery, but it does not guarantee task success. Nylas is the clearest example in this dataset.

llms.txtMCPTyped SDKOpenAPIAgent skillsCLI
Checklist
86
Tokens
308,324
Outcome
Verification passed

What this result means

Passed the public task evaluation and ships MCP, CLI, llms.txt, OpenAPI, and agent skills; the passing workaround keeps it below Resend.

The source evaluator recorded a pass. Review the caveat before treating it as semantic success.

Setup run completed

This dataset measures setup time, cost, errors, and human interruption. It is separate from the task-performance score.

Category rank
#4
Time
1m 53s
Cost
$0.84
Errors
4
Interruptions
1
Outcome
Completed

Independent rerun pending

This page summarizes recorded public evidence. We will mark the report when All Them APIs completes its own controlled rerun.

AgentMail API profile