14 August 2026 · parallel feedback · v1.1
Hours after Anthropic published its two-thirds theorem, a repository appeared on GitHub claiming the first improvement of its Theorem D — 67.3008% of zeta's zeros simple and on the critical line, against 67.2501% — and describing itself as a research draft generated by GPT-5.6 Sol, an OpenAI model. We audited it the way we audited the original: every constant recomputed from scratch, the load-bearing matrix inequality checked line by line and stress-tested, both computer-certified minimization targets independently measured against our own optimizers, the verifier read closely and then rerun end to end on this machine — reproducing both committed certificates exactly, to the hash — the one missing proof step reconstructed and supplied, and the result placed against the ceiling of the method, which we hold from the other audit.
The repository ainta/zeta-simple-zeros is a single git commit, timestamped
10 August 2026, 19:55 Pacific — the same day the Anthropic paper went public.
MIT licence, authored and released by
@kaizero_ainta (GitHub user
ainta), and a status line that says exactly what it is: "Research
draft generated by GPT-5.6 Sol. Independent verification and peer review are welcome." An
AI-generated draft extending an AI-generated theorem, landing within hours, on the machine that had
spent the week reproducing and kernel-checking the original. We report that as their statement and
audit the mathematics on its merits.
The claim is an extension of Theorem D. Where Anthropic proves that the proportion of zeros of zeta in a dyadic window that are simple and on the critical line is asymptotically at least H₀ = 3/2 − (1/√2)cot(1/√2) = 0.672500703679…, the draft proves
liminf N₀ˢ(T,2T)/N(T,2T) ≥ (1,345,000·H₀ − 2,680)/1,340,003 = 0.673008527927…
— an improvement of +5.078×10⁻⁴, in two graded steps: a 3-point argument reaching 67.2519767% and a 7-point argument reaching 67.3008528%. The deliverable is small and deliberately so: a short LaTeX paper, a web outline of the proof, a verifier of seven short Python files built on Arb interval arithmetic, and two committed certificate files. The verifier's design principle is stated in their docs and it is the right one: it reconstructs every transcendental enclosure from the formulas on each run, and trusts no cached tables.
github.com/ainta/zeta-simple-zeros · paper 4 pp (7-point argument only) · docs/proof.md (3-point and 7-point) · src/ 7 files · certificates/ 2 files · MIT · one commit, 2026-08-10.
One structural note a referee would raise immediately: the committed paper proves only
the 7-point theorem. The 3-point argument — the sum-free property, the constant
ε4, and the graph inequality it turns on — exists only in the web outline
docs/proof.md, where one of its steps is asserted with the phrase "a dual form of
Ψ" and no proof. We come back to that step below, with the proof it needs.
Measured
Anthropic's rank–trace step reads three numbers off a Hermitian matrix — rank, inertia, two traces — and its equality case would let the vectors attached to simple zeros be mutually orthogonal. But those vectors are not free: for the optimised test family, the inner product of two simple-zero vectors is pinned by the gap between their ordinates — it is the Montgomery–Taylor overlap kernel k evaluated at the gap, and k cannot vanish at u, v and u+v simultaneously. Consecutive zeros therefore always leave a quantitative trace in the Gram matrix, and the draft's stability refinement keeps that trace — a defect term Δ(M) = tr Ψ(M), Ψ(t) = (t−1)² on [0,2] and 2t−3 beyond — which the original two-trace argument throws away. The certified finite inequalities turn the defect into a proportion.
Their Lemma (inequality 2.1): if V has r columns of norm at most one, P = VV*, M = V*V, and Q is Hermitian with at most b positive eigenvalues, then
‖P+Q‖²F ≥ 4 tr(P+Q) − 3r − 4b + tr Ψ(M).
We checked every step by hand: the split Q = Q₋ − Q₋₋ with the dropped cross term tr(PQ₊) genuinely nonnegative; the eigenvalue inequality q² ≥ 4q − 4 on the positive part; the von Neumann pairing on the negative part; the pointwise minimisation minn≥0[(p−n)² + 4n] = 2p − 1 + Ψ(p); and the detail a hasty reader would miss — M's spectrum is P's nonzero spectrum padded with zeros, and Ψ(0) = 1, so collapsed columns pay into the defect too, which the proof handles correctly. Dropping Ψ recovers Anthropic's two-trace bound exactly; the refinement is strictly stronger and slots into their Proposition 4.4 at the matrix level, changing nothing upstream. Verified
Then we attacked it numerically: 3,000 randomised trials over dimensions up to 40, mixed column norms, adversarial choices of Q — minimum slack +1.72, never negative; and the equality probe (orthonormal columns, Q = 0) lands on zero slack to machine precision, which is the two-trace equality case the refinement is built to punish. Measured
The proof's positive-part step, q² ≥ 4q − 4, carries its own slack: (q − 2)² per positive eigenvalue of Q₋. So the lemma is really an inequality with a free extra term Σ⸺(q⸺ − 2)², and its second equality case is every positive eigenvalue sitting at exactly 2. A double zero on the critical line contributes exactly 2. An off-line pair at depth β = σ − ½ contributes a signature-(1,1) block whose positive eigenvalue q(β) — normalised so that an on-line double at the same ordinate gives exactly 2 — has no reason to stay there. This programme priced it, on the only data that can: its own bank of genuine off-line zeros.
The measurement. All 309 off-line pair blocks across three conductors and three heights (argument-principle-balanced censuses, the same bank as the first feedback's off-line control): every block has signature (1,1) as the structure demands, and q > 2 in 309 of 309 cases — down to the shallowest pair at depth 0.006. The law is clean: with L the band width, q(β) − 2 = c·(Lβ)² at small depth, with c constant across all seven fitted data sets to 3%. Measured
The derivation, and it is exact. For an interior pair the whole law collapses to a closed form: q(β) − 2 = 2∫φ²sinh²(βu)du / ∫φ²du, φ the test window — band-limitation makes every lattice sum an exact Poisson integral, the cross term vanishes identically, and the pair block's two eigenvalues obey q₋ + q₋₋ = 2 exactly. The constant c is the window's own second moment; for the Montgomery–Taylor window of this chain it is √2·cot(1/√2) − 3/2 = 0.154998592641…, which is also 2(1 − H₀) − ½ — the draft's own pinned pair energy in disguise. The closed form reproduces the measured per-set table, the deep pairs to a part in a million (the "superquadratic" growth at depth is just the sinh), and, since sinh²(x) ≥ x², the bound q(β) − 2 ≥ c·(Lβ)² is an exact pointwise inequality, not an asymptote. Verified
The honest pricing. The free term cannot raise the unconditional constant: the counting argument's binding configuration — every bad zero a double, exactly on the line — sits precisely at the second equality case, and no unconditional input forbids zeros at depth o(1/log T), where the term dies. What survives, free and strict, is a dichotomy: any configuration within ε of the certified count must keep all but (ε/c²δ⁴)·N of its off-line pairs within depth δ/log T of the critical line (a pair at depth 1/log T costs 0.024; at 2/log T, 0.38) and all but ε·N of its multiple zeros exactly double (a triple costs a full unit). A strict penalty on any configuration that violates the Riemann hypothesis by more than the argument's own blind zone — and a theorem about where the remaining slack in their lemma lives: after this term, the only unpriced step left is the dropped overlap tr(PQ₊) ≥ 0, which no unlabelled certificate can reach. Verified
The draft's kernel is K(x) = ∫−½½ cos(√2 t)cos(2πxt) dt with k = K/K(0). We recomputed the closed form against direct quadrature at 40 digits — worst difference 1.15×10⁻⁴¹ over seven points including the removable singularity — and K(0) = √2 sin(1/√2) exactly. The entire-sinc expression their code evaluates is the same function; our float64 implementation of it agrees with 40-digit arithmetic to 2.2×10⁻¹⁶ across [0, 11.5]. Verified
Its positive zeros are old acquaintances. Solving x·tan(πx) = c, c = tan(1/√2)/(√2π) = 0.192332…, gives z₁ = 1.05727829, z₂ = 2.03006753, z₃ = 3.02024299 — the same zeros our Anthropic-side audit measured for the optimal test function (1.05727830, 2.03006755, 3.02024302 on that page), the ones whose refusal to add up made the pair-correlation ceiling's dual bound unattainable. The draft's 3-point argument runs on exactly that refusal, one level down: if x, y and x+y were all zeros, the tangent addition formula forces x² + xy + y² + c² = 0, which is impossible over the reals. We verified the algebra including the two pole cases the short proof glosses (cos πx = 0 cannot occur at a zero, so the tangents exist), and measured the miss: on the domain u+v ≤ 4 the closest any zero-sum comes is 0.0671, at z₁+z₂ against z₃. Verified
The 7-point bookkeeping is tight and we reproduced all of it: 21 pairwise separations from six gaps with coefficients 2/(7−s); the window-summing lemma (a pair spanning s gaps appears in at most 7−s windows, each gap in at most six, giving Em + span/500 ≥ (19/5000)(m−6)); the choice m = 269 as the largest block with A₀ = (19/5000)·263 = 4997/5000 strictly below one, which is exactly what the defect lemma tr Ψ(G) ≥ min{1, 2Σ|Gij|²} needs so the min never clips; the convex pinching over shifted block partitions; and the final algebra. Every rational checks: 19·263 = 4997, 5000·269 = 1,345,000, 1,345,000 − 4,997 = 1,340,003. The headline constant recomputes at 40 digits as
(1,345,000·H₀ − 2,680)/1,340,003 = 0.6730085279277797613234…
— agreeing with the paper's printed value to all nineteen digits it prints, and the 3-point bound (H₀ − ε/4)/(1 − ε/2) = 0.672519767113… with theirs likewise. The asymptotic inputs the chain imports — Lemma 2.2's completeness identity, the tail estimates of Proposition 4.2, the index bound of Proposition 4.4, the ℓ−2 grid-tail decay — are precisely the components this programme verified numerically in the original audit, inside a development that builds and passes the Lean kernel on this machine. Verified
The web outline's inequality (3.4) — for any graph E of maximum degree two on the columns of V, Δ(V*V) ≥ (3/2) Σ{i,j}∈E |⟨Vi,Vj⟩|² — carries no proof in any file of the repository; the phrase "a dual form of Ψ" is all there is, and the committed paper avoids the step entirely by omitting the 3-point argument. The step is true, and here is the proof it needs. The convex conjugate of Ψ on [−2,2] is Ψ*(y) = y + y²/4. Take Y to be G's off-diagonal restricted to E: by Gershgorin, a degree-two graph with entries bounded by one keeps Y's spectrum inside [−2,2], so Fenchel–Young for spectral functions (von Neumann's inequality plus scalar Young) gives tr Ψ(G) ≥ tr(GY) − tr Ψ*(Y) = 2ΣE|Gij|² − (1/2)ΣE|Gij|² = (3/2)ΣE|Gij|². Two lines, and "dual form of Ψ" was evidently the intended route. Verified
Measured on 300,000 random 3×3 Grams plus a direct minimisation, the sharp constant for the triangles actually used is exactly 5/3 — the minimisation lands on 1.666666667 and the random trials never dip below 1.671 — so the stated 3/2 holds with room to spare, and the argument only ever applies it to disjoint triangles, where pinching reduces everything to the 3×3 case. Measured
The two computer-assisted inputs are finite minimisation claims: the 3-point constant ε4 ≥ 221/10⁶ on the triangle u,v ≥ 0, u+v ≤ 4, and the 7-point inequality F₆ ≥ 19/5000 on all nonnegative six-gap vectors. Their verifier proves these by interval branch-and-bound. We did not rerun their verifier (below); instead we built the gate the other way round, out of quantities their program never produces: independent minimisers hunting for the true minima — if either certified floor sat above the truth, the draft would be wrong, and our optimiser would find the violation.
A 1/500-mesh scan of the triangle, 1,065 basin candidates polished by Nelder–Mead, best refined at 40 digits: the global minimum of k(u)²+k(v)²+k(u+v)² is 2.2214911×10⁻⁴ at (u,v) = (1.05309, 2.01206) — the two gaps parked on the first two kernel zeros, their sum pressed as close to the third as the sum-free obstruction allows. Their certified 221/10⁶ sits 0.52% below the truth: valid, and nearly sharp. Measured
Six hundred thousand sampled gap vectors; then a systematic sweep giving every one of the 4,096 assignments of the six gaps to the first four kernel zeros its own Nelder–Mead descent, plus sixty random-basin descents; best point refined at 40 digits. The global minimum found is 3.82623121×10⁻³, at the gap pattern (1.045, 1.977, 1.042, 1.986, 1.989, 1.046) and its mirror image — three gaps on the first kernel zero, three on the second. Their certified 19/5000 sits 0.69% below the measured truth: valid, with almost nothing left on the table. Power of the search, stated: any counterexample to their certificate must live where the one-body bound U(g) = g/3000 + w(g)/3 fails to reach the target — a region we recomputed independently below and swept assignment by assignment; a violation hiding from this search would need a basin invisible to 4,156 descents and 600,008 samples. Measured
The 7-point verifier first removes every gap cell on which the one-body term already clears the
target, and its certificate records the survivors as three intervals of 1/4000-cells:
[3809,4778];[7221,9363];[10572,44827]. We recomputed that map from our own kernel code
— five-point sampled minima per cell, float64, no Arb — and got
[3809,4778];[7221,9362];[10572,44827]: identical to the cell, except one
boundary cell, which is exactly the direction a rigorous lower bound must differ from a
sampled minimum. The pressure cutoff is exact arithmetic (11.4/3000 = 19/5000, so beyond total gap
11.4 the linear term alone proves the claim), and the certificate files are internally consistent
to the counter: nodes = splits + pruned in both, leaves minus splits equals the 729 initial boxes,
and the three prune counters sum to the recorded total. Their kernel table is not something we
trusted — it is something we reproduced. Verified
The code does what the docs say. The interval conventions are right where they are easy to get
wrong: a sum of s closed cells maps to the inclusive cell range the code uses; boxes crossing the
domain boundary fall back to zero as a lower bound for a nonnegative function rather than reading
out of range; all floating-point combination of Arb endpoints is widened outward with
nextafter; the convex-tangent prune proves its Hessian positive-definite twice, once
in floats as a heuristic and again in Arb before it is believed; and an unresolvable terminal cell
raises — the program cannot return verified=true by silently giving up. The trust
base is stated plainly in their docs — Python and IEEE-754, python-flint/Arb/FLINT, their own
short source, OS and hardware — and after reading all seven files we found nothing smuggled in
against it: no cached enclosures, no sampled optimisation feeding the proof path.
Measured
With the interval library installed, both certificates were rerun end to end on this machine:
verified=true twice, and every recorded field of the committed certificate
files reproduces exactly — both sha256 kernel-table hashes, the 7,157 and 707,901
visited nodes, maximum depths 32 and 37, every prune counter, and the three surviving gap
components. Only the timings differ. The committed certificates are not just internally
consistent; they are what their code produces, byte for byte, on independent hardware. The
computer-assisted step of the draft is now fully closed here: floors measured independently from
outside, and the certifying run reproduced from inside.
Verified
This draft commits its certificates to the repository and rebuilds every transcendental
enclosure from formulas each time the verifier runs. That is precisely the property the Anthropic
development lacks at its one weak link — EnclOK resting on the unpublished
cert_N256_blk_b128m.json, exterior interval arithmetic, "available from the authors"
— which our other page carries as the single most useful thing Anthropic could fix.
The standing ask now has a worked example, produced by the first outside extension of
their own theorem. Conceded — theirs to own, and
done right.
float() on an exact endpoint errs by less than one unit in the last place before
nextafter widens it; that is how correctly-rounded conversion behaves, and the
end-to-end rerun exercised those conversions on this machine's library build without incident, but
pinning python-flint's conversion semantics at the source level was not done. It is the one
remaining unpinned link, and it is a library-reading exercise, not a computation.The refinement reads band-width-one data — the kernel k is the optimised window's overlap — and it holds configuration by configuration. That places it squarely in the certificate class governed by Remark 1.1 of the Anthropic paper, whose ceiling our other audit traced to a theorem in their Lean development: no certificate of this class can certify beyond p₀ = 0.681828687…, the simple-point fraction of an explicit 256-periodic extremal law. The draft's result is consistent with the ceiling, as it must be, and gives its first calibration from outside:
| Quantity | Value | Status |
|---|---|---|
| Theorem D (Anthropic) | 0.672500703679 | proved, kernel-checked here |
| 3-point refinement (draft) | 0.672519767114 | sound at statement level; certificate valid by our measurement; (3.4) proof supplied here |
| 7-point refinement (draft) | 0.673008527928 | sound at statement level; certificate valid by our measurement |
| ceiling of the certificate class | 0.681828687464 | theorem in the Lean development (our prior audit); one link rests on the unpublished certificate file |
The draft advances the constant by 5.078×10⁻⁴ against a remaining headroom of 9.328×10⁻³ — 5.4% of the distance from Theorem D to the ceiling, spent by refining the matrix inequality alone, with no new arithmetic input. Measured
The gap patterns at which F₆ bottoms out — runs of unit-spaced zeros broken by gaps at the second kernel zero — are precisely the unit fence carrying double-plus-hole defects that our period-2-through-128 linear programmes found as the optimal law at every period, and that Anthropic's 256-periodic law exhibits at the ceiling. The draft's local inequality is pricing, over a seven-point window, exactly the configuration class whose infinite-window optimum is the ceiling. That is why the route works — and why its reach is bounded by the ceiling and by nothing else. Measured
The draft's own design is at its measured limit. Its target captures 99.31% of the measured true minimum of its functional, and its pressure coefficient sits at the measured peak of the design curve — of six coefficients we tried from 1/1500 to 1/15000, theirs is the best, and nothing within the seven-zero design moves the endpoint by more than +1.7×10⁻⁵. As an audit verdict: the construction is essentially fully spent at its own choices, and whoever tuned it, tuned it well. Measured
The remaining headroom. Between the draft's 0.673009 and the ceiling's 0.681829 lie 8.8×10⁻³ that no band-one certificate has claimed. Whether any of it is reachable, and at what certification cost, is a question this programme is currently investigating; results will be reported separately, with their own certificates. Exploratory
Past the ceiling. Closed to this entire family of arguments. Beyond 0.6818… the band-width must exceed one, and the cheapest known purchase is the one our other page isolated: a one-sided, shift-averaged bound on a single signed prime-pair sum — materially weaker than the Hardy–Littlewood conjectures the paper names, and still nobody has a route to it. The draft does not change that landscape; nothing at band-width one can. Conceded
The mathematics holds. The stability inequality is correct and strictly stronger than the two-trace bound it refines; the kernel algebra and the sum-free obstruction are right; the block bookkeeping is exact; both certified floors sit just under the true minima we measured independently; the pruning geometry of their verifier reproduces from our own code cell for cell; and the final constants recompute to every printed digit. The one derivation gap — the unproved (3.4) in the web outline — is real but repairable, and this page supplies the repair. Their verifier has been rerun here end to end and reproduces its committed certificates exactly. Subject only to the imported inputs this programme has already verified upstream, we found nothing that threatens the result, and we looked for it adversarially.
And the honest scale. This is a +0.0005 refinement of a 0.6725 theorem, spending one part in eighteen of the headroom its own certificate class leaves; it introduces no new arithmetic and inherits the entire analytic machinery of the original. A polished AI-generated draft is a hypothesis until audited — this one was audited, and it survived. Provenance is worth exactly one sentence of record: the first outside improvement of the first AI-proved major zeta theorem was itself produced by an AI, within hours, and it stands up to an adversarial audit. The record should also say the draft does one thing better than the work it extends: its computer-assisted step ships with published certificates and a verifier that rebuilds every enclosure from scratch.
cert_N256_blk_b128m.json — now has a worked example
from inside their theorem's first extension.No zero is located, excluded or constrained by anything on this page, and nothing here bears on whether the Riemann hypothesis is true. No RH claim is made or implied — not by us, and not by the draft under review, whose result is a theorem about a proportion. The draft's improvement moves that proportion by five parts in ten thousand and moves nothing else.