Artificial IntelligenceQuantified Self-Evaluation51 AI developments · status to 2 October 2026Working Draft2 October 2026

Measuring the Path
to SuperintelligenceEvaluation of 2 October 2026

The necessity of three pillars, and Logarchéon’s current build placed among 2026 AI developments on one measured line

Two questions, answered in order. The first is whether the three pillars of the Logarchéon certification ladder — nonlocal coordination (CEAS), causal reasoning (Ψ) and invariant geometry (GRAIL) — are necessary for superintelligence, as named mechanisms or as functions that other mathematics can supply. The second is where Logarchéon’s current build (CEAS–Ψ–GRAIL) stands among the AI developments of 2026 when each development is scored from measured numbers, reduced to research scale, on a single [0, 100] line.

AI Evaluation · Self-Assessment 51 Developments · 7 Groups 8 Measured Quantities Logarchéon Inc. · William Chuang
00

Abstract

The central question

Are coordination, causal reasoning and invariant geometry necessary on the way to superintelligence, and, if they are required at least as functions, which 2026 AI developments have demonstrated them, measured rather than described?

The thesis advanced here

(i) Causal reasoning at Pearl’s second and third rungs is necessary as a function, by a proven theorem. (ii) Nonlocal coordination is necessary as a function only where communication rounds are counted, and ordinary global attention already supplies it. (iii) Exact invariance is a proven efficiency, not a proven necessity, and becomes necessary only where exactness on unseen group elements is part of the task. (iv) The joint necessity of the three is not proved. (v) When the ladder term is computed from measured numbers instead of judged in prose, Logarchéon’s current build ranks first of 51 developments on both readings of the line, by 0.9 points on the ungated reading and by about 9 points on the strict reading. Each clause is defeasible by a measurement named in the method.

How the thesis is tested

Each pillar’s necessity is checked against published theorems. Eight quantities, each in [0, 1], are computed by stated formulas from measured numbers: coordination C, causal reasoning Ψ, invariance G, the closed self-improvement loop R, control improving across cycles U, held-out verification H, breadth B and scaling S. They are aggregated into a ladder index in which a missing pillar zeroes the core. That index then replaces the one prose-judged ladder criterion in a four-criterion line covering 51 developments. Sensitivity readings bound every result.

Second thesis

The lead measured here is a lead in demonstrated ladder evidence, not in capability. The build’s measurements come from its own synthetic tasks, and most pillar cells of other systems are unmeasured. Running one battery on every runnable system is the only route to a comparison that is measured throughout.

01

Present Position

What this section is for: the state of the field and of the ladder at the date of the study, before any scoring.

As of 2 October 2026 the most capable public systems are OpenAI’s GPT-6 Astra, Anthropic’s Claude Opus 5.5 and Google’s Gemini 4 Argon. On the ladder’s 0–12 Rank scale they sit at Rank 5 (Emerging AGI). On the Certification Levels they sit at Level 1 read literally and at Level 2 read functionally. Global attention gives every transformer the coordination function, while no frontier system has published a group-invariance measurement or a current causal-rung score, and every frontier self-improvement loop still has human gating.

The systems nearest the ladder’s middle rungs are research-scale loops: Meta’s Hyperagents, Google DeepMind’s AlphaEvolve, the Darwin Gödel Machine and the Huxley-Gödel Machine. Each runs twenty or more cycles without manual edits, and none has separable causal, invariance and coordination components. Embodied systems (DeepMind’s self-improving embodied foundation models, Physical Intelligence, Gemini Robotics 2, NVIDIA GR00T) improve control, mostly between human-run releases.

Logarchéon’s current build meets Certification Level 5 under its own certification rule: all 23 test batteries pass at research scale, simulated control improves across self-improvement cycles on validation and on unused hidden seeds, and a hidden benchmark committed by hash in advance is consistent with validation. Two of the ten hard exclusion conditions are not cleared (transfer to an unseen task family; a falling self-improvement rate) and two are partly cleared (one task family per loop; no adversarial benchmark). Level 6 is at the groundwork stage.

The two scales of the ladder

Certification LevelPass criterion (lecture notes)
1 · ANIPasses one task or one domain.
2 · Architecture demonstrationAt least one of CEAS, Ψ, GRAIL beats its baseline independently.
3 · Triadic integrationJ111 below all seven ablation variants.
4 · Recursive triadic improvementAll three modules improve over K ≥ 20 closed-loop cycles, with no manual edits inside the loop.
5 · Physical-ASI seedLevel 4, plus all three Pearl rungs and inverse design, invariant generalisation, nonlocal coordination, control improving across self-improvement cycles, and ablation superadditivity.
6 · ASI claimLevel 5, plus best-baseline-beating performance across at least ten domains, an improvement process that itself improves, scaling exponents that predict larger hardware, and consistent private and adversarial benchmarks.
RankCapability classRankCapability class
0Classical automation7Expert / Exceptional AGI
1Narrow AI8Embodied AGI / Physical AI
2Foundation / frontier AI9Artificial superintelligence
3Agentic AI10Physical ASI
4AI-amplified R&D11Recursive ASI
5Emerging AGI12Complete-spectrum ASI
6Competent AGI

Certification Levels measure evidence about one architecture; Ranks measure capability class. There is no Certification Level 7: a “7” exists only as Rank 7 (Expert AGI) or as RSI Level 7 (closed-loop successor design) in the notes’ separate table of self-improvement levels.

Take-awayThe frontier leads in capability class; the ladder’s middle and upper rungs are populated only by research-scale systems, and Logarchéon’s current build is the one development that claims Level 5.
02

Summary of Findings

  1. Causal reasoning is necessary as a function. The Causal Hierarchy Theorem shows that answering rung-2 and rung-3 questions requires information at that rung or above; an explicit causal engine is one way to supply it, not the only way.
  2. Nonlocal coordination is necessary as a function only in a round-count model, and global attention or a global mean already supplies it. On the build’s own scaling tasks a plain global mean matched CEAS exactly.
  3. Exact invariance is a proven efficiency (a |G|-fold effective sample gain, and excess risk equal to the non-invariant part of the predictor), necessary only where exactness on unseen group elements is required.
  4. The joint necessity of the three pillars is not proved; the independence of the three obstructions is asserted, not shown.
  5. On the measured ladder index, Logarchéon’s current build scores 75.9 (bounded 75.9–88.4; 64.2 under conservative readings) at computed Level 5, and is the only system above zero on the strict, necessity-gated index.
  6. Every other system scores zero on the strict index because its pillars are unmeasured, not because a measurement failed: its bounded intervals reach 67.5–100.
  7. On the 51-entity line, with the ladder term computed from measurements, the build ranks first on both readings: 61.4 against Hyperagents’ 60.5 (ungated), and 61.4 against 52.3 for AlphaEvolve and Hyperagents (strict). With the prose-judged ladder term, it had ranked 9th–11th at 60.0.
  8. The ungated lead is fragile. With the conservative ladder index the build scores 58.7 and falls to second behind Hyperagents.
  9. The build fails Level 6 on four counts: breadth B = 0, no prediction to larger hardware tested, an improver that grew once (+0.0041; sign-test p = 0.40), and no adversarial set.
03

Method

What this section is for: every rule of measurement and aggregation, fixed before any system is scored, with the exclusions stated as prominently as the rules.

Rule 1 · Pillars are required as functions

Following the ladder strictly means requiring the three functions, not the three named modules: any mechanism that provides a pillar’s function by a mathematically equivalent route counts. Each function is tested by a number measurable on any system (Part I).

Rule 2 · Eight quantities, each in [0, 1], from measured numbers only

QuantityFormulaRule for edge cases
C · coordination1 − qq is an architectural fact or a measured round count. Global attention gives q = 0; local message passing in 3-D gives q = 1/3.
Ψ · causalmean of four terms: â2, â3, abstention rate, inverse-design success; â = clip((a − c)/(1 − c), 0, 1)Chance c = 0.5 for binary items (CLadder); c = 0 for numeric or open-ended items (the build’s demos, CausalReasoningBenchmark, inverse design). A pooled score with no rung split and an unverified definition (CausalDS, secondary source) is excluded from the strict reading and shown as a sensitivity.
G · invariancemin(1, accunseen / accseen) on held-out group elements, or 1 when exact invariance holds by constructionARC-AGI is never used for G. Learned or augmented invariance with no measured error counts as unmeasured.
R · loopmin(1, K/20)·(1 − h)K counts cycles that modify the system’s own code, weights or scaffold; h = 0 when the published procedure has no human step inside a cycle.
R2fraction of cycles that improvedUsed only in the Level 4 threshold (≥ 0.6).
U · control across cyclesg·min(1, cycles/2); g = (E0 − E1)/E0, or (P1 − P0)/(1 − P0) for success ratesA single improvement phase counts half; an unknown cycle count is treated as one phase.
H · held-out verification1, 0.5 or 01: scored by a third party on a set the developer could not see, with no third-party finding of overstatement. 0.5: a developer-held hidden set (hash-committed included), a developer-chosen held-out domain, a third-party run on a public set, or a third-party finding of overstatement. 0: none reported.
B · breadthmin(1, n/10)n counts domains with independent measurement at or above the best human.
S · scaling1 if a fitted exponent’s confidence interval excludes 0 against a baseline, else 0Log-linear fits without an exponent and interval score 0.

Rule 3 · Level thresholds

LevelThreshold rule (pillar equivalence allowed)
L1 · ANIAny measured score on one task.
L2 · Architecture demonstrationC ≥ 0.9, or Ψ ≥ 0.5, or G ≥ 0.9.
L3 · Triadic integrationAll three pillars meet their L2 thresholds, and a 23 ablation shows J111 below all seven variants.
L4 · Recursive improvementL3, R = 1 (K ≥ 20, h = 0), R2 ≥ 0.6, and every module improved inside the loop (battery R3).
L5 · Physical-ASI seedL4, every Ψ term ≥ 0.8, G ≥ 0.9 with εtwin ≤ 10−6, C = 1, U > 0, H ≥ 0.5, and ablation superadditivity.
L6 · ASI claimL5, B = 1, S = 1 with a confirmed prediction on larger hardware, a significant measured rise of the improver (Rt+K > Rt), and H = 1 including an adversarial set.

Rule 4 · The ladder index, with necessity built in

I = 100 · P · [1 + R + (U + H)/2 + (B + S)/2] / 4,   P = (C · Ψ · G)1/3

The geometric mean encodes necessity: a zero pillar gives a zero core and a zero index. The bracket follows the ladder’s stages: a complete core alone is worth 25 (Levels 2–3), the loop adds up to 25 (Level 4), control and held-out verification add up to 25 (Level 5), breadth and scaling add up to 25 (Level 6). The higher stages are additive inside the gate, so partial progress counts. Because the formula is monotone, the lower end of each bounded interval equals the strict value.

Alternative A (no gate): arithmetic mean of C, Ψ, G, R, U, H, B, S × 100
Alternative B (proof-weighted gate): I = 100 · √(C·Ψ) · [1 + G + R + (U + H)/2 + (B + S)/2] / 5

Rule 5 · The 51-entity line

Score = (20·B1 + 15·B2 + 15·B5 + 15·Q/10) / 6.5
CriterionWeightWhat it measures
B1 · novelty of the mechanism20A new mechanism, as opposed to an integration of known parts.
B2 · breadth per unit of resource15Capability shown per unit of compute, log-discounted.
B5 · falsifiability and protocol discipline15Pre-registration, stated gates, negative results reported.
Q · measured ladder index15Q-ungated = Alternative A; Q-strict = the gated index I. Scaled from 0–100 to 0–10.

Every development is reduced to research scale: compute, capital and deployment are removed, and each is judged on its own published evidence. The prose-judged ladder criterion of the earlier line (B6) is replaced by Q, so the ladder term is computed from measurements; B1, B2 and B5 are carried over unchanged from the own-evidence line.

Exclusions

No credit for peer review, outside evaluation, replication or adoption on the 51-entity line: every development is judged on its own evidence, so a self-certified build and a peer-reviewed system are compared on the same terms.
Unmeasured counts as zero on the strict reading, and as the interval [0, 1] on the bounded reading.
Instruments and components with no ladder index are left out: METR, the CLadder-type causal benchmarks, the UK AI Security Institute and the Muon optimiser.
Enablers with no mechanism to score are left out: NVIDIA compute, TSMC, Huawei Ascend, national compute programmes, and announced laboratories with no public system.
Systems absent from the measured table are transformer-based with only C = 1 measured, and receive the floor value Q-ungated = 12.5.
B1, B2 and B5 remain judgements recorded in prose; only the ladder term is computed from measurements.
Hidden-benchmark seed values are not reproduced.
Take-awayEvery number on this page follows from the rules above; changing a rule changes the result, and the sensitivity readings in Parts II and IV show by how much.
04

I · Are the Pillars Necessary?

What this section is for: to settle, from published theorems, which pillar is required, in what sense, and by what measurable criterion.

The lecture notes (v21.5) prove three narrow theorems: a round separation (local message passing needs at least D rounds, where D is the diameter, while one all-reduce suffices; physical time is Ω(N1/d) for either model), a two-model non-identifiability example with a worst-case error floor of 1/4, and a finite-group Reynolds risk identity. The broader triadic necessity theorem builds the three capabilities into its definition of scalable superintelligence and asserts that “these three obstructions are independent” without a separating construction; “Architecture from Theory” is labelled a conjecture.

PillarVerdictTheorem behind itCriterion measurable on any system
CEASNecessary as a function, proven only in a unit-cost-broadcast (round-count) model; not necessary in physical timeRound separation (notes); LOCAL t-hop lower bound (Ghaffari); GNN depth×width bounds (Loukas)Rounds exponent q in T(N) ∝ Nq on globally sensitive aggregation: local q = 1/d, functional CEAS q ≈ 0. Any all-reduce qualifies (global attention, a global mean, a bus, a tree).
Ψ · rungs 2–3Necessary as a function, provenCausal Hierarchy Theorem (Bareinboim, Correa, Ibeling & Icard): rung collapse occurs only on a measure-zero set of models; meagre-set version (Ibeling & Icard); robust agents imply an approximate causal model (Richens & Everitt)Rung-2 and rung-3 accuracy above chance; correct abstention on non-identifiable queries; the 1/4 error floor on observationally equivalent pairs.
Ψ · inverse designUnproven as a separate necessity; reduces to rung 2 plus planningNone foundSuccess rate in reaching a target under constraints, against a forward-only baseline.
GRAILEfficient, not strictly necessary; a symmetry prior is necessary only if exact invariance on unseen orbit elements is requiredExcess risk equals ‖f⊥‖2 (Elesedy & Zaidi); |G|-fold sample gain (Tahmasebi & Jegelka); canonicalisation is equivalent (Kaba et al.)Retention on held-out group elements; εinv = E‖F(gx) − F(x)‖2; any map that factors through X/G qualifies.
Triad jointlyUnprovenIndependence asserted, not shownMain effects of a 23 ablation.

Three concessions that bound the verdicts

  1. The notes state that standard transformer self-attention already realises an aggregate primitive, so the round theorem does not by itself distinguish CEAS from transformers.
  2. On the build’s own R4 tasks a global mean matched CEAS exactly, and the entropy-corridor controller failed its CEAS hypothesis test (mean −0.002), while the β-controller passed (+0.092).
  3. GRAIL’s necessity argument is a no-free-lunch argument that does not exclude canonicalisation or learned invariance, which grows with scale in vision models (Gruver et al.).

The Causal Hierarchy Theorem does not require an explicit causal-model engine: causal statements in training text, or experiments run through tools, can carry rung-2 information. Following the ladder strictly therefore means requiring the three functions, and treating GRAIL’s function as required only because the ladder writes exact invariance into its definition.

Take-awayOf the three pillars, only the causal function is proven necessary without a cost-model caveat; a proof-weighted index gates on causal reasoning and coordination, and measured causal accuracy is the pillar frontier systems have yet to report.
05

II · The Measured Ladder Index

What this section is for: the eight quantities computed for each system, the resulting index and level, and how far each result moves under alternative readings.

Cells give the strict value. “·” means unmeasured, so the bounded value runs over [0, 1]; a narrower bounded range is given in brackets. † Level 4 applies if U is read against the loop’s own earlier miss, which rises with wear, instead of against the frozen controller.
SystemCΨGRUHBSPStrict IBounded IAlt ALevel
Logarchéon current build11110.5740.50 (·)1175.975.9–88.475.95 (4†)
Hyperagents1··10.3320.5··000–85.435.42
GPT-6 Astra1··0 (≤0.5)·1··000–87.525.02
GPT-6.1 Sol, GPT-5.6 Sol, Opus 5, Opus 5.5, Opus 4.8, Gemini 3.8 Flash, Grok 4.6, Kimi K3, Qwen3.8 (27B), GLM-5.3, Seed, Sakana AI Scientist, Co-Scientist, AlphaProof, Aristotle1····1··000–10025.02
Darwin Gödel Machine1··1····000–10025.02
MACE / GNoME0.667 (≤1)·1·····000–10020.82 (via G)
DeepSeek V4-Pro1····0.5··000–93.818.82
Huxley-Gödel Machine1··0·0.5··000–93.818.82
Self-improving embodied FMs1··0 (≤0.05)0.300···000–67.516.32
π*0.61··00.25 (≤0.5)···000–93.815.62
AlphaEvolve10.238·0····000–99.615.52
GPT-5.5 / 5.310.085······000–94.213.62
Gemini 4 Argon, Muse Spark, V-JEPA 2-AC, Absolute Zero, Gemini Robotics 2, GR00T, AgiBot GO-2, AlphaFold 31·······000–10012.52
Calibration: GPT-4 on CLadder10.117······000–85.114.02
Calibration: Euclidean transformer inside the build1—0.300——————————

How the non-trivial cells are computed

  • Logarchéon current build. Ψ = mean(24/24, 5/5, 26/26, 100/100) = 1 with c = 0. G = 1.000/1.000 on SL2(ℝ), with εtwin ≈ 10−13. R = min(1, 40/20)·1 = 1, and R2 = 0.80 (31–33 of 40 cycles). U = mean over six seeds of (1 − loop/frozen) = 0.574, multiplied by min(1, 24/2) = 1. S = 1 from the token-axis difference interval [0.692, 0.936]. I = 100·1·(1 + 1 + 1.074/2 + 1/2)/4 = 75.9.
  • Hyperagents. U = (0.372 − 0.060)/0.94 = 0.332 over 100 iterations of robotics reward design. H = 0.5 because the held-out domain was chosen by the developer.
  • Self-improving embodied FMs. g takes three values: (0.875 − 0.63)/0.37 = 0.662 on the real robot, (0.75 − 0.45)/0.55 = 0.545 in simulation, and 0.595 on BananaTable. Their mean is 0.601; one phase halves it to U = 0.300.
  • π*0.6. The primary abstract reports that the method roughly halves the task failure rate, so g = 0.5; the cycle count is unknown, so one phase halves it to U = 0.25.
  • AlphaEvolve. The target was matched or beaten on about 75% + 20% of 50+ problems, so inverse-design success = 0.95 and Ψ = 0.95/4 = 0.238. Crediting a design task run against an exact evaluator as inverse design is a judgement.
  • GPT-5.3. Complete identification specification was 34.1% (c = 0), so Ψ = 0.341/4.
  • GPT-4 on CLadder. Ψ = [(0.628 − 0.5)/0.5 + (0.606 − 0.5)/0.5]/4 = 0.117.
  • GPT-6 Astra. The upper bound on R comes from the report that more than half of 4–8 hour research tasks needed at least one human intervention, which gives h ≥ 0.5.
  • H = 1 assignments come from ARC Prize semi-private scores, live IMO contests, blind peer review and an independent experiment. DeepSeek receives 0.5 because NIST CAISI found it performed worse than its own report suggested.

Sensitivities

ReadingEffect
U read against the loop’s own miss (U = 0)Build falls to 68.8 (Level 4).
Wilson 95% lower bounds on the small Ψ samples (5/5 gives 0.566; Ψ = 0.815)Build 70.9.
Both of the aboveBuild 64.2.
Alternative B (proof-weighted gate)Build 80.7, or 75.0 with U = 0; AlphaEvolve 9.7, GPT-4 6.8, GPT-5.3 5.8 — the only aggregation that lifts other systems above zero.
Unreported h counted as 1Hyperagents’ Alternative A falls from 35.4 to 22.9; the Darwin Gödel Machine’s from 25.0 to 12.5. Strict index unchanged.
CausalDS pass rate read as pooled accuracyOpus 4.8 and GPT-5.5 Ψ = 0.412, Gemini 3.1 Pro 0.383; strict index still 0, because G is unmeasured.

The index line

STRICT INDEX (necessity-gated)
0         20        40        60        80       100
|---------|---------|---------|---------|---------|
X  ← 33 systems at 0 (Astra, Opus 5.5, Gemini 4, Hyperagents, DGM, AlphaEvolve, robots, provers …)
                                      L (Logarchéon build) 75.9
bounded:  build [75.9 ===== 88.4]   EFM [0 ====== 67.5]   Hyperagents [0 ======= 85.4]
          Astra [0 ======= 87.5]   most others [0 ========== 100]

ALTERNATIVE A (arithmetic mean, no necessity gate)
0         20        40        60        80       100
|---------|---------|---------|---------|---------|
      ^12.5 Argon, Muse, V-JEPA, robots, AF3
        ^15.5-16.3 AlphaEvolve, π*0.6, EFM
         ^18.8-20.8 DeepSeek, HGM, MACE
            ^25.0 ARC-verified LLMs, DGM, provers, Co-Scientist, Sakana
                ^35.4 Hyperagents
                                      ^75.9 Logarchéon build
Take-awayThe strict line has one non-zero point because the necessity gate meets unmeasured pillars everywhere else; it records what has been shown, not a ranking of ability.
06

III · The 51-Entity Line

What this section is for: every development on one [0, 100] line, with the ladder term computed from measurements, under both readings.

Score = (20·B1 + 15·B2 + 15·B5 + 15·Q/10)/6.5. Ungated: Q = Alternative A. Strict: Q = the necessity-gated index, so every system with an unmeasured pillar has Q = 0. Before the ladder term was measured, the same line with the prose-judged ladder criterion placed Logarchéon’s current build 9th–11th at 60.0.
Reading

Indicator Table

RANK = capability rank (0–12)  · LEVEL = certification level (functional)  · B1 novelty  · B2 breadth per resource  · B5 protocol discipline  · Q = measured ladder index

#DevelopmentGroupRankLevelB1B2B5Q-ung.UngatedQ-str.Strict

The leaders on each reading

#Ungated readingScore#Strict readingScore
1Logarchéon current build61.41Logarchéon current build61.4
2Hyperagents60.52=AlphaEvolve52.3
3AlphaEvolve55.92=Hyperagents52.3
4Darwin Gödel Machine55.84Darwin Gödel Machine50.0
5Equivariant scientific ML54.05=Equivariant scientific ML49.2
6Self-improving robot models53.05=V-JEPA 2-AC49.2
7AlphaProof52.75=Self-improving robot models49.2
8V-JEPA 2-AC52.18=AlphaProof, Absolute Zero, Huxley-Gödel Machine46.9
Take-awayWith the ladder term measured, Logarchéon’s current build leads both readings; frontier general models fall to the middle of the line (GPT-6 Astra 39.6 ungated, 33.8 strict), behind the research-scale self-improvement and world-model systems.
07

IV · Firmness of the Lead

What this section is for: how much of the first place survives a change of reading.

ReadingLogarchéon buildNearest rivalMarginWhat the margin rests on
Ungated, Q = 75.961.4 (1st)Hyperagents 60.5+0.9Measured pillar evidence; Hyperagents leads on novelty (B1 8 against 4.5).
Ungated, conservative Q = 64.258.7 (2nd)Hyperagents 60.5−1.8Control judged against the loop’s own error, and cautious bounds on the small causal samples.
Strict, Q = 75.961.4 (1st)AlphaEvolve, Hyperagents 52.3about +9The rule that an unmeasured pillar counts as zero: others lose the ladder term because their pillars have not been measured, not because a measurement failed.
Prose-judged ladder term (earlier line)60.0 (9th–11th)Hyperagents 73.1−13.1The ladder criterion judged in prose instead of measured.
Take-awayThe strict first place is robust to the build’s own sensitivities but depends on the zero-for-unmeasured rule; the ungated first place depends on one reading of the control quantity.
08

V · What the Numbers Cannot Say

What this section is for: the exclusions of the evidence, stated as prominently as its findings.

  1. Equal numbers do not mean equal capability. The build’s values of 1.0 are exact results on its own synthetic tasks: 24 causal questions, the SL2(ℝ) and PSL2(ℤ) groups, a two-dimensional point mass, 16 causal blocks and at most 40 cycles on one workstation, measured against correlation-only baselines and SmolLM2-135M. Frontier numbers come from public benchmarks of a different order of difficulty; ARC-AGI-3 measures novel-task generalisation, not group invariance.
  2. Most frontier pillar cells are unmeasured. The strict index records demonstrated ladder evidence; the bounded intervals of 0–100 show how little the public record constrains those systems.
  3. The ladder was written around one architecture. Its Level 5 checks (εtwin ≤ 10−6, a hash-committed hidden set, an R2 rate) are quantities the build was designed to report, and the coordination term gives every transformer the same score as CEAS.
  4. Three of the four line criteria remain prose judgements (B1, B2, B5).
  5. Verification is internal. The build is self-certified; its reviews were carried out by separate AI agents reading the same files, and no outside party has rerun it.
  6. Two hard exclusion conditions of the examination are not cleared (transfer to an unseen task family; a falling self-improvement rate), and two are partly cleared (one task family per loop; no adversarial benchmark). The Level 5 verdict applies the certification report’s rule and is stated together with these conditions.
  7. Level 6 is not met. Breadth B = 0; no prediction to larger hardware has been tested; the improver grew once (+0.0041, 9 of 16 tasks won, one-sided sign-test p = 0.40); no adversarial set exists.
Take-awayThe line measures demonstrated ladder evidence at research scale; it does not measure capability, and it is not an outside verification.
09

VI · Plan of the Study

What this section is for: the one route to a comparison measured throughout — the same battery run on every system that can be run.

  1. Choose models. Open-weight models with ARC Prize scores: Qwen3.8-27B, DeepSeek V4, GLM-5.3 and Kimi K3; frontier APIs where access exists.
  2. C: fit q from the number of reasoning steps needed against N on globally sensitive aggregation and parity tasks.
  3. Ψ: run the 24-question suite, the action-order demonstration, the non-identifiable abstention items, physical inverse design and the per-rung CLadder splits, at the stated chance levels.
  4. G: measure seen/unseen retention and εinv on SL2(ℝ) and PSL2(ℤ) inputs, presented in context and after fine-tuning.
  5. R, R2 and U: place each model as the proposer inside the same 40-cycle gated loop, log h, and run the same control plant.
  6. H: score on the unused hash-committed hidden seeds, plus an adversarial set built by a third party.
  7. S: use the scaling axes of battery R4 with bootstrap intervals.
  8. Release: publish the audit bundle (scripts, logs, Lean files and used seeds) so that every number above can be checked from the files alone.

Runs on 27–30B models fit on one or two large GPUs; the loop runs (40 cycles × 3 seeds per model) are the main cost. Every cell would then be measured on identical items, and the bounded intervals would collapse to points.

Take-awayThe next measurement that can change this page is the same battery on open-weight models.
10

Sources