Agent Bench
Agent Bench result
SendGrid Agent Bench results
One autonomous setup run, preserved with its recorded outcome and caveats.
Autonomous setup run
Setup run blocked
This dataset measures setup time, cost, errors, and human interruption. It is separate from the task-performance score.
- Category rank
- #5
- Time
- —
- Cost
- —
- Errors
- —
- Interruptions
- —
- Outcome
- Blocked
Benchmark record
Run completed and preserved
This page reports the recorded All Them APIs run, including failures, interruptions, and caveats.