Terminal-Bench v4.0
Score
Measures
Can the agent finish real work in a terminal/CLI harness (commands ↔ container state ↔ verifier pass/fail)? Headline: resolution / score %.
- 1st. GPT-6 Astra (Max) 59.8%
- 2nd. GPT-6 Astra (High) 59.7%
- 3rd. Claude Fable 5 (with fallback) 54.0%