skip to content

Stats

The creator of Gantry collects statistics on his own Gantry jobs and presents them here, to give an idea about how Gantry performs in real-life situations.

2026-06-18 - 2026-09-20

28 codebases 473 runs archived

877 milestones 4.5k planned tasks

20k agent sessions

930 plan 4.7k build 154 fix 437 of 473 runs

136.4M tokens out 66k edits

359k messages 385k commands 437 of 473 runs 435 of 473 runs 437 of 473 runs 435 of 473 runs

Execute-agent peak context

4,464 build sessions - exact per-session peaks 415 of 450 runs

The main goal of Gantry is to break work down into chunks small enough for agents to execute without running up large context windows. The key metrics are about context windows, and we focus on the peak number: the largest size the context window reached at any time during the session.

Across 4,464 autonomous execute sessions, the peak context reached stays under the 150k design target in 71% of them - and 88% stay under 200k. Only 0.6% of sessions (25 of 4,464) peaked over 300k.

Under 150k target
71%
3,175 of 4,464 sessions 415 of 450 runs
Under 200k
88%
only 534 sessions above 415 of 450 runs
Median peak
117k
across all execute sessions 415 of 450 runs
Over 300k
0.6%
25 sessions 415 of 450 runs
claude execute 987 sessions 112 of 450 runs
0
25
50
75
100
125
150
175
200
225
250
275
300+
150k 200k
68%< 150k
121kmedian
16%> 200k
2.5%> 300k
542kmax
codex execute 3,477 sessions 309 of 450 runs
0
25
50
75
100
125
150
175
200
225
250
275
300+
150k 200k
72%< 150k
117kmedian
11%> 200k
0%> 300k
246kmax

Each histogram is that harness's own execute sessions, bin width 25k, final bar 300k+; bar height is scaled to the busiest bin in that harness.

Supporting roles

the aside - non-execute agents

Execution is the most important type of agent run and the most context-constrained. The bottleneck we are targeting. Here are the other roles and their context peaks.

Per-role median by harness
Role claude n codex n median peak
environment-build 101 of 450 runs 109,110 5 76,650 360
plan 360 of 450 runs 74,576 304 40,155 850
gate-build 101 of 450 runs 76,502 6 60,747 386
execute 415 of 450 runs 120,627 987 116,521 3,477
troubleshoot 124 of 450 runs 64,152 35 62,180 231
review 411 of 450 runs 64,618 1,028 75,910 4,212
replan 168 of 450 runs 55,289 203 40,577 214
resolve 34 of 450 runs 143,164 4 84,412 59
support 91 of 450 runs 36,895 2 27,950 134
Median over all roles: 79,631 416 of 450 runs. claude peaks higher than codex in every role except review.

Agent runtime

active working time per run and agent runtime by role

How long runs and agents take. The whole-run figure is active time: stopped time is excluded, and parked time is excluded where the journal records it. The harness columns are each harness's median; median, mean and p90 pool both.

Runtime per agent role
Role claude codex median mean p90 n
environment-build 101 of 450 runs 5m 48s 2m 20s 2m 21s 2m 27s 3m 21s 364
plan 101 of 450 runs 3m 51s 1m 49s 1m 50s 1m 58s 2m 50s 401
gate-build 101 of 450 runs 5m 8s 2m 50s 2m 51s 3m 23s 6m 4s 392
execute 106 of 450 runs 16m 24s 6m 23s 6m 25s 8m 8s 14m 17s 1,505
troubleshoot 45 of 450 runs 5m 41s 1m 50s 1m 52s 3m 7s 5m 41s 117
review 101 of 450 runs 3m 35s 2m 6s 2m 6s 2m 40s 5m 7s 2,002
resolve 23 of 450 runs 10m 38s 4m 18s 5m 7s 6m 59s 11m 34s 29
support 91 of 450 runs 32s 32s 32s 36s 49s 153
Task runtime median
11m 25s
4,519 executed task rows - p90 29m 54s
Task runtime p25-p75
7m 31s - 18m 30s
execute + test + review per task
Active working median
2h 9m
whole-run active time
Active working mean
4h 24m
average across measured runs
Active working p90
8h 55m
90th percentile active time
Active working sample
473
runs with active duration
Stopped human waits are excluded for every active-duration observation. Parked retry waits are excluded for 7 newer runs; 466 older runs have stop-corrected active time only because their journals do not record parked waits.

Decomposition

450 structured builds

Gantry sequences work by modelling it as milestones and tasks, dispatching agents to implement task by task while planning on multiple levels. Here is how that structure ends up slicing the work.

Tasks / milestone
4.0
median - p25-p75 3.3-5.0
Milestones / job
3
median of the 315 jobs that split into milestones
Tasks / job
8
median across all 450 builds - max 57
How a plan breaks down
Metricminmedianmeanp90max
Milestones / multi-ms job n=315 1 3 2.8 4 8
Tasks / job n=450 1 8 9.8 20 57
Tasks / milestone n=315 0.5 4.0 4.4 6.0 12.0
Milestones per job

450 structured builds, bucketed by milestone count. A flat build is one whose tasks are a single list, with no milestone layer above them.

flat 135 135 of 450 runs
1 ms 44 44 of 450 runs
2 ms 103 103 of 450 runs
3 ms 90 90 of 450 runs
4 ms 50 50 of 450 runs
5 ms 19 19 of 450 runs
6+ ms 9 9 of 450 runs
Tasks per job

450 structured builds, bucketed by total task count.

1–2 tasks 43 43 of 450 runs
3–5 tasks 126 126 of 450 runs
6–10 tasks 131 131 of 450 runs
11–15 tasks 68 68 of 450 runs
16–25 tasks 61 61 of 450 runs
26–40 tasks 17 17 of 450 runs
41+ tasks 4 4 of 450 runs

Gantry builds Gantry

this repo's own history, through 2026-09-19

Gantry's own repository - the orchestrator and this website - is largely built by Gantry runs. Gantry commits under its own git identity, so its share of the work is measurable: every figure here is read from the repo's commit history and the records its runs left behind.

Runs on this repo
255
216 ran to completion
Agent sessions
6,908
530h 5m of agent runtime
Code lines by Gantry
75.8%
1,533,875 of 2,023,569 changed lines
Commits by Gantry
66%
5,284 of 8,005 commits
Who changed what

Changed lines and files across the whole history, by author identity. Gantry's run records and vendored third-party trees answer nothing about who did the work, so the narrower cuts drop them.

Cut Gantry files Gantry lines other files other lines Gantry share
everything 34,004 6,331,502 33,732 4,546,989 58.2%
project files 3,272 2,110,209 3,653 945,903 69%
code only 1,964 1,533,875 1,674 489,694 75.8%
"Other" is every commit not made under Gantry's identity - including inline agent sessions that commit as the human, so it means not-Gantry, not hand-written. File counts overlap: a file both sides touched counts once for each. Of the 1,180 code files in the tree today, 719 (61%) were last written by Gantry.