API index
Agent BenchPublic evidence edition

Email API Agent Bench

How easily can a coding agent discover, connect, and complete a real task with each email API? We compare setup performance with real task outcomes.

Snapshot fetched Aug 17, 2026 2 independent sources Not tested by All Them APIs
01

Resend ranks first in both public email datasets.

Resend records the fastest autonomous setup and perfect discovery marks. Its lower task score and 11 errors still expose friction that a simple feature checklist would miss.

Autonomous setup results

Autonomous setup runs ranked within the email category by time, cost, errors, and human interruptions.

Claude Code · Opus 4.6
Resend
56s
Mailgun
1m 11s
Postmark
1m 21s
AgentMail
1m 53s
RankAPITimeCostErrorsInterruptionsOutcome
01Resend56s$0.5401 Completed
02Mailgun1m 11s$0.4101 Completed
03Postmark1m 21s$0.5601 Completed
04AgentMail1m 53s$0.8441 Completed
05SendGrid Automation blocked

A lower time, cost, error count, and interruption count is better. SendGrid could not be evaluated because it blocked browser automation.

Task performance results

End-to-end API tasks scored on task evaluation and discovery, with execution cost, calls, errors, and time retained as evidence.

Claude Code · API evaluation
01
Resend
75
B
Eval
65
Discovery
100
Calls
38
Errors
11
Time
4m 24s
02
AgentMail
72
C
Eval
67
Discovery
86
Calls
18
Errors
5
Time
5m 4s
03
Nylas
40
D
Eval
21
Discovery
86
Calls
63
Errors
15
Time
12m 4s

Phase timings and task outcomes

Expand an API to inspect calls by phase, integration assets, token usage, and whether the requested behavior was completed.

ResendVerification passedB · 75

The task completed and the API was easy to discover, but test-mode sending restrictions caused repeated corrections and a relatively high error count.

Discovery11 calls~1m 50s
Tool search1 calls~3s
Code8 calls~26s
Run15 calls~1m 39s
Debug3 calls~26s
Checklist 100782,135 tokensTested 2026-05-20
Context7llms.txtMCPTyped SDKOpenAPIAgent skillsCLI
View Agent Bench results
AgentMailVerification passedC · 72

The agent recovered from an inbox limit and a bounced recipient, then sent to the test inbox itself. The run passed, although the workaround changed the requested recipient behavior.

Discovery3 calls~1m 28s
Tool search1 calls~8s
Code5 calls~45s
Run5 calls~1m 38s
Debug4 calls~1m 5s
Checklist 86308,324 tokensTested 2026-05-19
llms.txtMCPTyped SDKOpenAPIAgent skillsCLI
View Agent Bench results
NylasVerification failedD · 40

The agent could not create a real mailbox without an authenticated provider grant. It generated 13 files and a demo-mode output, but did not complete the requested live workflow.

Discovery11 calls~3m 12s
Tool search1 calls~5s
Install1 calls~12s
Code14 calls~1m 33s
Run27 calls~5m 12s
Debug9 calls~1m 49s
Checklist 861,536,888 tokensTested 2026-05-19
Context7llms.txtMCPTyped SDKAgent skillsCLI
View Agent Bench results
Do not compare the two tables directly.

They use different tasks, harnesses, scoring systems, and run conditions. This report preserves each source as a separate lens until All Them APIs runs a controlled benchmark.