Agent Bench

AgentMail Agent Bench results

One task run and one autonomous setup run, reviewed as separate forms of evidence.

Reviewed Aug 17, 2026 2 recorded runs Tested by All Them APIs
C72/100

Verification passed

The agent recovered from an inbox limit and a bounced recipient, then sent to the test inbox itself. The run passed, although the workaround changed the requested recipient behavior.

Evaluation
67
Discovery
86
Time
5m 4s
Cost
$0.33
Calls
18
Errors
5

Where the agent spent its time

The trace divides the run into discovery, tool search, installation, coding, execution, and debugging phases where applicable.

01Discovery
3 calls~1m 28s
02Tool search
1 calls~8s
03Code
5 calls~45s
04Run
5 calls~1m 38s
05Debug
4 calls~1m 5s

What this result means

Passed the Agent Bench task evaluation and ships MCP, CLI, llms.txt, OpenAPI, and agent skills; the passing workaround keeps it below Resend.

The evaluator recorded a pass. Review the caveat before treating it as semantic success.

Setup run completed

This dataset measures setup time, cost, errors, and human interruption. It is separate from the task-performance score.

Category rank
#4
Time
1m 53s
Cost
$0.84
Errors
4
Interruptions
1
Outcome
Completed

Run completed and preserved

This page reports the recorded All Them APIs run, including failures, interruptions, and caveats.

AgentMail API profile