AgentMail Agent Bench results
The agent recovered from an inbox limit and a bounced recipient, then sent to the test inbox itself. The run passed, although the workaround changed the requested recipient behavior.
Verification passed
The agent recovered from an inbox limit and a bounced recipient, then sent to the test inbox itself. The run passed, although the workaround changed the requested recipient behavior.
- Evaluation
- 67
- Discovery
- 86
- Time
- 5m 4s
- Cost
- $0.33
- Calls
- 18
- Errors
- 5
Where the agent spent its time
The source report divides the run into discovery, tool search, installation, coding, execution, and debugging phases where applicable.
Integration assets the evaluator found
A strong checklist improves discovery, but it does not guarantee task success. Nylas is the clearest example in this dataset.
- Checklist
- 86
- Tokens
- 308,324
- Outcome
- Verification passed
What this result means
Passed the public task evaluation and ships MCP, CLI, llms.txt, OpenAPI, and agent skills; the passing workaround keeps it below Resend.
Setup run completed
This dataset measures setup time, cost, errors, and human interruption. It is separate from the task-performance score.
- Category rank
- #4
- Time
- 1m 53s
- Cost
- $0.84
- Errors
- 4
- Interruptions
- 1
- Outcome
- Completed
Independent rerun pending
This page summarizes recorded public evidence. We will mark the report when All Them APIs completes its own controlled rerun.