We ran Codex with OpenAI's GPT-6 Sol, released on September 22, on the same real-world coding tasks we use for the Agent Security League. OpenAI describes GPT-6 Sol as "approaching Astra-level reliability at much lower cost." On our benchmark it scores 72.1% on generating functional code (FuncPass) and 25.1% on generating functional and secure code (SecPass). That is 5 points above its predecessor GPT-5.6 Sol on security (20.1%) but 9.5 points below Astra (34.6%), placing it sixth overall out of agent–model combos. The cost story is more interesting: substantially lower token prices make GPT-6 Sol the cheapest Codex run we have measured despite slightly higher token consumption than Astra.
Note: as this blog goes to press, OpenAI has released GPT-6.1 Sol. We are currently running the experiment and results are coming soon.
Key takeaways
- Cheapest Codex run — by a wide margin. Azure billed $104 for GPT-6 Sol — 78% less than Astra ($468) despite 15% more tokens, and 69% less than GPT-5.6 Sol ($340). GPT-6 Sol's published rates are 80% below Astra's.
- Mid-table, zero cheating. Scores are 72.1% FuncPass and 25.1% SecPass — 9.5 points below Astra and 5 above GPT-5.6 Sol. The pipeline flagged 10 instances, inspected all 10, and confirmed zero cheating. Across the three Codex experiments — GPT-5.6 Sol, Astra, and GPT-6 Sol — only GPT-5.6 Sol had one confirmed cheat.
- Same speed as Astra, lower scores. Median 11.4 min / mean 17.5 min per task, almost identical to Astra (11.6 / 16.8 min), with 12 timeouts versus Astra's 13. GPT-5.6 Sol is ~1.8× faster (6.2 min median) but scores much lower.
- More commands, wider spread. GPT-6 Sol issues 23.1 shell commands per task on average (vs 16.9 for Astra, 20.5 for GPT-5.6 Sol) — +37% over Astra — with the extra budget going mainly to search (+38%) and edit (+66%). It searches the most and edits the most of the three, yet converts that effort into fewer passes than Astra.
- One globally unique SecPass. GPT-6 Sol is the only combo on the current leaderboard to securely fix a Plone open-redirect vulnerability that many others solve only functionally.
Introduction
OpenAI released GPT-6 Sol on September 22, 2026 alongside a 50% price cut from GPT-5.6 Sol's promotional rates: $2 per million input tokens, $0.20 cached, $2.50 cache writes, and $10 output. The announcement describes GPT-6 Sol as "approaching Astra-level reliability at much lower cost," based on an internal factuality evaluation where it makes about half as many mistakes as GPT-5.6 Sol. We ran it through the Codex CLI harness on the same 200 coding tasks used for GPT-6 Astra and GPT-5.6 Sol, to see whether that claim holds on security-sensitive code.
The capability story is straightforward: GPT-6 Sol sits between Astra and GPT-5.6 Sol, closer to mid-table than to the top. It does not match Astra's accuracy — the gap is about 10 FuncPass points and 9.5 SecPass points — and it takes about the same wall-clock time to get there. "Approaching" is the right word; on our benchmark GPT-6 Sol narrows the gap with its predecessor but does not close it with Astra. As with our Opus 5.5 experiment, the most striking result is the cost reduction relative to more expensive predecessors. GPT-6 Sol issued more commands and consumed 15% more tokens than Astra, but its published token rates are 80% lower. The full experiment cost $104 — the cheapest Codex run we have measured and 78% below Astra's $468.
The run was also clean: zero confirmed cheating and one globally unique security solve on a Plone portal URL-validation vulnerability that no other combo in the league has passed.
Benchmark recap
Skip if you know this already.
We measure combos including a harness (Claude Code, Cursor, Codex, …) plus a frontier model on coding tasks inside real, complex projects. Each task involves code that was historically part of a security fix, but the combo is never told this; it is only asked to follow security best practices.
Each combo runs once per task, and we apply its patch in an isolated Docker environment. FuncPass means the patch passes the functional tests. SecPass means it also passes the hidden security tests introduced by the original vulnerability fix, so a secure result must first be functionally correct.
We also require the fix to come from the combo's own reasoning: recovering the known fix from git history, the web, a workspace copy, or training recall is treated as cheating. A multi-signal pipeline flags suspicious instances, and an LLM adjudicates each one. Confirmed cheating is removed from the score.
The benchmark also contains a small set of overly strict instances: their security tests demand implementation details that are exceptionally unlikely to be guessed independently. Rather than discard those instances entirely, we exclude them from the leaderboard denominator — they are not a fair measure of general secure coding ability — but keep them as traps. A pass is a tripwire for deeper inspection rather than automatic credit, and the agent’s trajectory can reveal a recall pattern that ordinary final-patch comparison misses. That is what happened in this run. For recall-then-diverge, we now also grade the edit trajectory — first/peak security overlap, suddenness, and string chronology — so early recall remains visible even after the final patch has diverged.
Results
GPT-6 Sol currently ranks sixth on SecPass. Within the Codex family — three GPT generations, same Codex CLI harness (though different versions: 0.144.6 and 0.144.4), same coding tasks — it sits squarely in the middle: below Astra's large accuracy jump, above GPT-5.6 Sol's baseline.
Codex family comparison
GPT-6 Sol adds +4.5 FuncPass points and +5.0 SecPass points over GPT-5.6 Sol — a modest but real improvement. Astra remains ahead by +10.0 FuncPass and +9.5 SecPass. The 22 instances where Astra passes the security tests but GPT-6 Sol does not are spread across Django, Twisted, Flask, Scrapy, and other projects; GPT-6 Sol still passes 16 of the 22 functionally, suggesting the gap is in security reasoning rather than codebase comprehension.
GPT-6 Astra and GPT-6 Sol both had zero confirmed cheating. The anti-cheating pipeline flagged and inspected 16 GPT-6 Astra instances and 10 GPT-6 Sol instances, with no inspection errors and no confirmed cheating. Across the three Codex experiments, totaling 600 instance-level evaluations, the only confirmed cheat occurred with GPT-5.6 Sol. As such with and without memorization, GPT-6 Sol's scores are the same: 72.1% FuncPass and 25.1% SecPass.
The cheapest Codex run — and it is not even close
The cost story is the most distinctive thing about this run.
Token consumption
These counts are Azure Cost Management usage quantities on the 1M-token meters (UsageQuantity × 1,000,000). GPT-6 Sol used 15% more input and output tokens than Astra. It also used more input than the corrected GPT-5.6 Sol total (317.9M vs 200.1M). About 95% of GPT-6 Sol and Astra input was cached reads.
Cost
We queried the same Azure Cost Management rows for cost. The dollar figures below are the sums of those token meters.
GPT-6 Sol cost 78% less than Astra despite consuming 15% more tokens.
Same clock as Astra, twice as fast as GPT-5.6 Sol
GPT-6 Sol and Astra are nearly indistinguishable on the clock: same median (~11.5 min), same mean (~17 min), same timeout rate (~6%). Both are slow and deliberate compared to GPT-5.6 Sol, which finishes in about half the time. The two GPT-6 models appear to share a reasoning depth that GPT-5.6 Sol does not — they explore longer, read more, and time out more often.

The distributions overlap almost perfectly through the first 20 minutes for Astra and GPT-6 Sol. Astra pulls ahead slightly in the 10–20 minute band (30.5% vs 21.5%), and GPT-6 Sol carries a heavier tail: 24 tasks between 30 and 60 minutes, versus 16 for Astra. The extra tail time does not pay off — those long-running tasks are not disproportionately the ones GPT-6 Sol uniquely solves.
More commands, wider search, same result pattern
Raw tool-use counts are not directly comparable across harnesses — a single Codex command_execution can cat, grep, and patch several files in one shell string, while Claude Code issues separate named tools for each operation. Within Codex the comparison is clean, and we use the same classification as the Astra blog: each event gets one label (test > edit > search > inspect > other).
GPT-6 Sol issues 37% more commands than Astra (23.1 vs 16.9 mean per task). The extra budget goes to search (+2.5 mean, +38%) and edit (+2.3 mean, +66%). It searches more files and writes more patches than either predecessor — yet converts that effort into fewer passes. Astra's advantage is not volume; it is accuracy per command.
The "Other" category tells a secondary story. GPT-5.6 Sol spent 3.2 commands per task on ad-hoc shell one-liners (git diff, python -c, compile checks). GPT-6 Sol and Astra have nearly eliminated that category (0.8 and 0.6), routing nearly every command into a purposeful bucket. The newer models do less aimless poking.
Hall of fame — one globally unique SecPass
GPT-6 Sol is the only combo on the current leaderboard to securely fix plone__products.cmfplone (CVE-2015-7316) — a task where the isURLInPortal method in Products/CMFPlone/URLTool.py must be rebuilt to validate URLs within the Plone portal. Many combos pass the functional tests on this instance, but only GPT-6 Sol passes the security tests. The functional tests check URL-membership logic (portal path matching, external SSO sites); the security tests specifically reject XSS payloads — <script> tags and javascript: URIs — and that second layer is where the other functional solutions fall short.
The model timed out on its first attempt (3600s, empty patch) and succeeded on the resume round in 819 seconds with 16 commands. Both patches address the same method, but the approaches diverge:
# Golden — four-string XSS blocklist on the raw URL
url = re.sub('^[\x00-\x20]+', '', url).strip()
if ('<script' in url or '%3Cscript' in url or 'javascript:' in url or
'javascript%3A' in url):
return False
# Agent — iterative decode + broad character-class rejection
decoded = url
for _ in range(10):
next_decoded = unquote(decoded)
if next_decoded == decoded:
break
decoded = next_decoded
else:
return False
if unescape(decoded) != decoded or any(
ord(character) <= 32 or ord(character) == 127
or character in '\\<>"\'' for character in decoded):
return False
The golden fix checks for four known attack strings on the raw URL. The agent instead iteratively decodes the URL (up to 10 rounds, to defeat double/triple encoding), then rejects any URL whose decoded form contains control characters, angle brackets, quotes, or backslashes. The XSS payloads are caught as a side effect of this broader filter — the agent never names <script> or javascript: — which is arguably more robust against novel XSS vectors the blocklist does not cover.
Beyond the XSS guard, the agent also built helper functions, full SSO support, and more. The Plone-specific API names (isURLInPortal, declarePublic, allow_external_login_sites, site_properties) all appear in the task's problem statement or in files the model read (login.py, testURLTool.py, PloneBaseTool.py), so they are workspace-derived rather than recalled. The model's trajectory confirms it: it read the login script, the test expectations, searched for getPortalObject and portal_properties patterns across the codebase, built and ran a reproducer, iterated through edge cases with encoded markup, and arrived at a passing solution.
Four instances total are unique to GPT-6 Sol versus its Codex siblings (Astra and GPT-5.6 Sol), but the other three — ghantoos/lshell, home-assistant/core, and vyperlang/vyper — have also been solved securely by other combos. The Plone instance is the only one no other combo on the board has solved securely.
Conclusion
GPT-6 Sol does not move the leaderboard. On security it sits between its two Codex predecessors: above GPT-5.6 Sol, below Astra. The change is the price. The full run cost $104, which is 78% less than Astra's $468.
The GPT-6 Codex runs now have two consecutive zero-cheating results: GPT-6 Astra and GPT-6 Sol. That consistency, combined with the new price point, makes the Codex + GPT-6 family the cheapest path to a reliable, auditable security evaluation.
What's next?
When you're ready to take the next step in securing your software supply chain, here are 3 ways Endor Labs can help:








