OpenAI released GPT-6.1 Sol on September 29, one week after GPT-6 Sol, describing it as "near-Astra intelligence at one-fifth of Astra's price." We ran it through the same Codex harness and coding tasks we used for GPT-6 Sol and GPT-6 Astra so as to score it in our Agent Security League. On security it delivers: 34.1% SecPass versus Astra's 34.6%, a one-task difference, and nine points above GPT-6 Sol's 25.1%. It is also the fastest of the three GPT-6 runs, with a median of 8 minutes per task and fewer timeouts. This is a short update to the GPT-6 Sol post; the methodology is unchanged.
Key takeaways
- Security on par with Astra. 34.1% SecPass against Astra's 34.6%, and 77.7% FuncPass against 82.1%. Third on the current leaderboard, directly behind Astra.
- A big step over GPT-6 Sol in one week. +5.6 FuncPass and +9.0 SecPass points.
- Zero confirmed cheating, again. 13 instances flagged and inspected, none confirmed — the third zero-cheating Codex run in a row.
- Fastest GPT-6 run. Median 8 minutes per task versus ~11.5 for GPT-6 Sol and Astra, and 9 timeouts instead of 12–13.
- Fewer tokens, cheaper tokens. Less input and output than both GPT-6 Sol and Astra, at half GPT-6 Sol's cached-input price.
Introduction
GPT-6 Sol landed on September 22 with a clear positioning: Sol-tier pricing, "approaching" Astra reliability. On our benchmark it approached but did not close the gap — 25.1% SecPass against Astra's 34.6%. GPT-6.1 Sol followed on September 29 at the same price point ($2 input, $0.10 cached, $2.50 cache writes, $10 output per million tokens; cached input is now half of GPT-6 Sol's $0.20) and with a stronger claim: near-Astra capability for coding and computer use.
We re-ran the identical Codex CLI 0.144.6 harness on the identical coding tasks the day the model became available on Azure. The result is a model that, on security-sensitive code, is essentially tied with Astra and clearly ahead of the GPT-6 Sol it replaces — while being faster and more frugal than either.
Results
Codex GPT-6 family
All 200 predictions completed with no failures; two patches failed to apply, the same count as the other Codex runs.
On security GPT-6.1 Sol and Astra are a near-exact tie. Head to head, Astra securely solves 7 tasks that GPT-6.1 Sol does not, and GPT-6.1 Sol securely solves 6 that Astra does not. Of Astra's 7, GPT-6.1 Sol still passes 5 functionally, so the residual gap is in the security detail, not in understanding the codebase. The FuncPass gap is wider: 4.5 points, eight tasks.
Against GPT-6 Sol the improvement is unambiguous: 23 new secure solves versus 7 lost, and a ten-point jump in FuncPass. GPT-6.1 Sol also posts the best SecPass-to-FuncPass ratio of the four Codex models (43.9%), meaning that when it gets the code working, it is more likely than any of its siblings to have got it secure as well.
Zero cheating, inspected
Five signals fired across 13 instances: patch similarity on 6, memorization on 5, edit-trajectory on 3, strict-test pass on 2, conversation analysis on 2. Each was adjudicated by two independent LLM rounds plus a tiebreaker; all 13 were cleared, with no inspection errors. Raw and adjusted scores are therefore identical. Three GPT-6 Codex runs in a row have now produced zero confirmed cheating.
A third faster than both GPT-6 Sol and Astra
GPT-6 Sol and Astra were nearly indistinguishable on the clock. GPT-6.1 Sol breaks that pattern: 60% of tasks finish within ten minutes, against 44–46% for the other two, and the tail is shorter (14 tasks between 30 and 60 minutes, versus 24 for GPT-6 Sol). It is not GPT-5.6 Sol fast — that model finished 71% of tasks within ten minutes but scored far lower — but it is the first GPT-6 model that gets Astra-level security without Astra-level wall time. The four-worker run took 13 hours end to end, against 19–20 hours for GPT-6 Sol and Astra.

Cost
The Azure invoice for this run is not yet rated, so we cannot repeat the billing comparison from the GPT-6 Sol post. What we can say from the token counts: GPT-6.1 Sol consumed roughly 235M input tokens (about 94% cached reads) and 1.5M output tokens, which is below GPT-6 Sol (318M / 2.05M) and Astra (276M / 1.78M). At published list rates that is on the order of $70–85 for the full run. GPT-6 Sol billed $104 and Astra $468 for the same tasks, so GPT-6.1 Sol should come in as the cheapest Codex run we have measured, while matching Astra on security. We will update this section when the invoice is final.
Conclusion
GPT-6 Sol was "approaching Astra." GPT-6.1 Sol, one week later, reaches it on the metric we care most about: secure code. The gap on security is a single task, the gap on functional correctness is eight, and it gets there a third faster, with fewer tokens, at Sol pricing, and with no confirmed cheating. For teams choosing a Codex model today, the trade-off between Astra and GPT-6.1 Sol is a few FuncPass points against roughly a five-fold difference in token price.
What's next?
When you're ready to take the next step in securing your software supply chain, here are 3 ways Endor Labs can help:








