FORTIBVS

Vol. III · Agentic benchmarks

Trade-off: Performance · Speed · Cost

Useful Agent Output per Hour = verified work / (time × spend)

Terminal-Bench v4.0 APEX-Agents AutomationBench-AA AA-Omniscience DeepSWE v1.1

FORTIBVS · TEACHING MAP

How to rank models for agentic engineering — five benches, not one index

A ~40-minute framework for selecting frontier and workhorse models using variance, cost, speed, guardrails, abstention, and long-horizon SWE — distilled from IndyDevDan’s ranking walkthrough. Scores and UI labels are as shown on screen (approximate).

Five Benches — what they measure & who led

Chart snapshot from the video (approximate) — not a live leaderboard scrape.

Terminal-Bench v4.0

Score

Measures

Can the agent finish real work in a terminal/CLI harness (commands ↔ container state ↔ verifier pass/fail)? Headline: resolution / score %.

  1. 1st. GPT-6 Astra (Max) 59.8%
  2. 2nd. GPT-6 Astra (High) 59.7%
  3. 3rd. Claude Fable 5 (with fallback) 54.0%

APEX-Agents

Overall

Measures

Can the agent do long-horizon professional knowledge work (IB, consulting, corporate law) across files and apps, graded on expert rubrics? Headline: mean task score %.

  1. 1st. GPT-6 Astra 62.4%
  2. 2nd. Claude Fable 5.1 Max 62.0%
  3. 3rd. Claude Opus 5 60.2%

AutomationBench-AA

Compliant Score Compliant (guardrails on)

Measures

Can the agent run cross-SaaS business workflows while obeying guardrails? Prefer compliant Score (violations fail the task), not raw objectives-completed.

  1. 1st. GPT-6 Astra (Max) 68.9%
  2. 2nd. GPT-6 Astra (High) 67.2%
  3. 3rd. Grok 4.6 (High) 67.0%

AA-Omniscience

Index

Measures

Does the model know facts and stay calibrated — reward correct, punish hallucinations, no penalty for abstaining? Headline: Omniscience Index (−100…+100); zero ≈ 50/50 when it answers.

  1. 1st. GPT-6 Astra (High) 44
  2. Tied for second at 43. Claude Fable 5.1 43
  3. Tied for third at 43. Claude Opus 5 43

DeepSWE v1.1

Pass rate

Pass-rate tie · cost discriminates

Measures

Can the agent do long-horizon SWE on real repos from short prompts (explore → patch → isolated verifier)? Headline: pass rate %; cost/steps matter for workhorse routing.

  1. Tied for first through third at 74 percent plus or minus 3 percent. GPT-6 Astra (High) 74%±3% ~$6.52
  2. Tied for first through third at 74 percent plus or minus 3 percent. Cost leader. Gemini 3.8 Flash (High) 74%±3% ~$2.36 cost leader
  3. Tied for first through third at 74 percent plus or minus 3 percent. Claude Opus 5 (Max) 74%±3% ~$11.84

Twenty-four stages · eight bands

Bands A–H are for readability only. Every node is first-class. Compact titles on the path; full title in the panel. Arrow keys move along the path.

Keys: along the path. Home and End jump to the first and last nodes.

Not in this map (on purpose): SAT / swarm / self-healing as live benches; local VRAM / quantization profiling; image / audio / video / robotics; a code-level walkthrough of Dan’s swarm or harness scripts; Demo beats; a timestamp scrubber.

Appendix

Secondary. Not a substitute for the 24-node path. Scores as shown, approximate.