Agent Bench
Agent Bench result
Postmark Agent Bench results
One autonomous setup run, preserved with its recorded outcome and caveats.
Autonomous setup run
Setup run completed
This dataset measures setup time, cost, errors, and human interruption. It is separate from the task-performance score.
- Category rank
- #3
- Time
- 1m 21s
- Cost
- $0.56
- Errors
- 0
- Interruptions
- 1
- Outcome
- Completed
Benchmark record
Run completed and preserved
This page reports the recorded All Them APIs run, including failures, interruptions, and caveats.