- Eval
- 65
- Discovery
- 100
- Calls
- 38
- Errors
- 11
- Time
- 4m 24s
Email API Agent Bench
How easily can a coding agent discover, connect, and complete a real task with each email API? We compare setup performance with real task outcomes.
Resend ranks first in both public email datasets.
Resend records the fastest autonomous setup and perfect discovery marks. Its lower task score and 11 errors still expose friction that a simple feature checklist would miss.
Autonomous setup results
Autonomous setup runs ranked within the email category by time, cost, errors, and human interruptions.
| Rank | API | Time | Cost | Errors | Interruptions | Outcome |
|---|---|---|---|---|---|---|
| 01 | 56s | $0.54 | 0 | 1 | Completed | |
| 02 | 1m 11s | $0.41 | 0 | 1 | Completed | |
| 03 | 1m 21s | $0.56 | 0 | 1 | Completed | |
| 04 | 1m 53s | $0.84 | 4 | 1 | Completed | |
| 05 | — | — | — | — | Automation blocked |
A lower time, cost, error count, and interruption count is better. SendGrid could not be evaluated because it blocked browser automation.
Task performance results
End-to-end API tasks scored on task evaluation and discovery, with execution cost, calls, errors, and time retained as evidence.
- Eval
- 67
- Discovery
- 86
- Calls
- 18
- Errors
- 5
- Time
- 5m 4s
- Eval
- 21
- Discovery
- 86
- Calls
- 63
- Errors
- 15
- Time
- 12m 4s
Phase timings and task outcomes
Expand an API to inspect calls by phase, integration assets, token usage, and whether the requested behavior was completed.
ResendVerification passedB · 75
The task completed and the API was easy to discover, but test-mode sending restrictions caused repeated corrections and a relatively high error count.
AgentMailVerification passedC · 72
The agent recovered from an inbox limit and a bounced recipient, then sent to the test inbox itself. The run passed, although the workaround changed the requested recipient behavior.
NylasVerification failedD · 40
The agent could not create a real mailbox without an authenticated provider grant. It generated 13 files and a demo-mode output, but did not complete the requested live workflow.
They use different tasks, harnesses, scoring systems, and run conditions. This report preserves each source as a separate lens until All Them APIs runs a controlled benchmark.