Abstract
Are coordination, causal reasoning and invariant geometry necessary on the way to superintelligence, and, if they are required at least as functions, which 2026 AI developments have demonstrated them, measured rather than described?
(i) Causal reasoning at Pearl’s second and third rungs is necessary as a function, by a proven theorem. (ii) Nonlocal coordination is necessary as a function only where communication rounds are counted, and ordinary global attention already supplies it. (iii) Exact invariance is a proven efficiency, not a proven necessity, and becomes necessary only where exactness on unseen group elements is part of the task. (iv) The joint necessity of the three is not proved. (v) When the ladder term is computed from measured numbers instead of judged in prose, Logarchéon’s current build ranks first of 51 developments on both readings of the line, by 0.9 points on the ungated reading and by about 9 points on the strict reading. Each clause is defeasible by a measurement named in the method.
Each pillar’s necessity is checked against published theorems. Eight quantities, each in [0, 1], are computed by stated formulas from measured numbers: coordination C, causal reasoning Ψ, invariance G, the closed self-improvement loop R, control improving across cycles U, held-out verification H, breadth B and scaling S. They are aggregated into a ladder index in which a missing pillar zeroes the core. That index then replaces the one prose-judged ladder criterion in a four-criterion line covering 51 developments. Sensitivity readings bound every result.
The lead measured here is a lead in demonstrated ladder evidence, not in capability. The build’s measurements come from its own synthetic tasks, and most pillar cells of other systems are unmeasured. Running one battery on every runnable system is the only route to a comparison that is measured throughout.
Present Position
What this section is for: the state of the field and of the ladder at the date of the study, before any scoring.
As of 2 October 2026 the most capable public systems are OpenAI’s GPT-6 Astra, Anthropic’s Claude Opus 5.5 and Google’s Gemini 4 Argon. On the ladder’s 0–12 Rank scale they sit at Rank 5 (Emerging AGI). On the Certification Levels they sit at Level 1 read literally and at Level 2 read functionally. Global attention gives every transformer the coordination function, while no frontier system has published a group-invariance measurement or a current causal-rung score, and every frontier self-improvement loop still has human gating.
The systems nearest the ladder’s middle rungs are research-scale loops: Meta’s Hyperagents, Google DeepMind’s AlphaEvolve, the Darwin Gödel Machine and the Huxley-Gödel Machine. Each runs twenty or more cycles without manual edits, and none has separable causal, invariance and coordination components. Embodied systems (DeepMind’s self-improving embodied foundation models, Physical Intelligence, Gemini Robotics 2, NVIDIA GR00T) improve control, mostly between human-run releases.
Logarchéon’s current build meets Certification Level 5 under its own certification rule: all 23 test batteries pass at research scale, simulated control improves across self-improvement cycles on validation and on unused hidden seeds, and a hidden benchmark committed by hash in advance is consistent with validation. Two of the ten hard exclusion conditions are not cleared (transfer to an unseen task family; a falling self-improvement rate) and two are partly cleared (one task family per loop; no adversarial benchmark). Level 6 is at the groundwork stage.
The two scales of the ladder
| Certification Level | Pass criterion (lecture notes) |
|---|---|
| 1 · ANI | Passes one task or one domain. |
| 2 · Architecture demonstration | At least one of CEAS, Ψ, GRAIL beats its baseline independently. |
| 3 · Triadic integration | J111 below all seven ablation variants. |
| 4 · Recursive triadic improvement | All three modules improve over K ≥ 20 closed-loop cycles, with no manual edits inside the loop. |
| 5 · Physical-ASI seed | Level 4, plus all three Pearl rungs and inverse design, invariant generalisation, nonlocal coordination, control improving across self-improvement cycles, and ablation superadditivity. |
| 6 · ASI claim | Level 5, plus best-baseline-beating performance across at least ten domains, an improvement process that itself improves, scaling exponents that predict larger hardware, and consistent private and adversarial benchmarks. |
| Rank | Capability class | Rank | Capability class |
|---|---|---|---|
| 0 | Classical automation | 7 | Expert / Exceptional AGI |
| 1 | Narrow AI | 8 | Embodied AGI / Physical AI |
| 2 | Foundation / frontier AI | 9 | Artificial superintelligence |
| 3 | Agentic AI | 10 | Physical ASI |
| 4 | AI-amplified R&D | 11 | Recursive ASI |
| 5 | Emerging AGI | 12 | Complete-spectrum ASI |
| 6 | Competent AGI |
Certification Levels measure evidence about one architecture; Ranks measure capability class. There is no Certification Level 7: a “7” exists only as Rank 7 (Expert AGI) or as RSI Level 7 (closed-loop successor design) in the notes’ separate table of self-improvement levels.
Summary of Findings
- Causal reasoning is necessary as a function. The Causal Hierarchy Theorem shows that answering rung-2 and rung-3 questions requires information at that rung or above; an explicit causal engine is one way to supply it, not the only way.
- Nonlocal coordination is necessary as a function only in a round-count model, and global attention or a global mean already supplies it. On the build’s own scaling tasks a plain global mean matched CEAS exactly.
- Exact invariance is a proven efficiency (a |G|-fold effective sample gain, and excess risk equal to the non-invariant part of the predictor), necessary only where exactness on unseen group elements is required.
- The joint necessity of the three pillars is not proved; the independence of the three obstructions is asserted, not shown.
- On the measured ladder index, Logarchéon’s current build scores 75.9 (bounded 75.9–88.4; 64.2 under conservative readings) at computed Level 5, and is the only system above zero on the strict, necessity-gated index.
- Every other system scores zero on the strict index because its pillars are unmeasured, not because a measurement failed: its bounded intervals reach 67.5–100.
- On the 51-entity line, with the ladder term computed from measurements, the build ranks first on both readings: 61.4 against Hyperagents’ 60.5 (ungated), and 61.4 against 52.3 for AlphaEvolve and Hyperagents (strict). With the prose-judged ladder term, it had ranked 9th–11th at 60.0.
- The ungated lead is fragile. With the conservative ladder index the build scores 58.7 and falls to second behind Hyperagents.
- The build fails Level 6 on four counts: breadth B = 0, no prediction to larger hardware tested, an improver that grew once (+0.0041; sign-test p = 0.40), and no adversarial set.
Method
What this section is for: every rule of measurement and aggregation, fixed before any system is scored, with the exclusions stated as prominently as the rules.
Rule 1 · Pillars are required as functions
Following the ladder strictly means requiring the three functions, not the three named modules: any mechanism that provides a pillar’s function by a mathematically equivalent route counts. Each function is tested by a number measurable on any system (Part I).
Rule 2 · Eight quantities, each in [0, 1], from measured numbers only
| Quantity | Formula | Rule for edge cases |
|---|---|---|
| C · coordination | 1 − q | q is an architectural fact or a measured round count. Global attention gives q = 0; local message passing in 3-D gives q = 1/3. |
| Ψ · causal | mean of four terms: â2, â3, abstention rate, inverse-design success; â = clip((a − c)/(1 − c), 0, 1) | Chance c = 0.5 for binary items (CLadder); c = 0 for numeric or open-ended items (the build’s demos, CausalReasoningBenchmark, inverse design). A pooled score with no rung split and an unverified definition (CausalDS, secondary source) is excluded from the strict reading and shown as a sensitivity. |
| G · invariance | min(1, accunseen / accseen) on held-out group elements, or 1 when exact invariance holds by construction | ARC-AGI is never used for G. Learned or augmented invariance with no measured error counts as unmeasured. |
| R · loop | min(1, K/20)·(1 − h) | K counts cycles that modify the system’s own code, weights or scaffold; h = 0 when the published procedure has no human step inside a cycle. |
| R2 | fraction of cycles that improved | Used only in the Level 4 threshold (≥ 0.6). |
| U · control across cycles | g·min(1, cycles/2); g = (E0 − E1)/E0, or (P1 − P0)/(1 − P0) for success rates | A single improvement phase counts half; an unknown cycle count is treated as one phase. |
| H · held-out verification | 1, 0.5 or 0 | 1: scored by a third party on a set the developer could not see, with no third-party finding of overstatement. 0.5: a developer-held hidden set (hash-committed included), a developer-chosen held-out domain, a third-party run on a public set, or a third-party finding of overstatement. 0: none reported. |
| B · breadth | min(1, n/10) | n counts domains with independent measurement at or above the best human. |
| S · scaling | 1 if a fitted exponent’s confidence interval excludes 0 against a baseline, else 0 | Log-linear fits without an exponent and interval score 0. |
Rule 3 · Level thresholds
| Level | Threshold rule (pillar equivalence allowed) |
|---|---|
| L1 · ANI | Any measured score on one task. |
| L2 · Architecture demonstration | C ≥ 0.9, or Ψ ≥ 0.5, or G ≥ 0.9. |
| L3 · Triadic integration | All three pillars meet their L2 thresholds, and a 23 ablation shows J111 below all seven variants. |
| L4 · Recursive improvement | L3, R = 1 (K ≥ 20, h = 0), R2 ≥ 0.6, and every module improved inside the loop (battery R3). |
| L5 · Physical-ASI seed | L4, every Ψ term ≥ 0.8, G ≥ 0.9 with εtwin ≤ 10−6, C = 1, U > 0, H ≥ 0.5, and ablation superadditivity. |
| L6 · ASI claim | L5, B = 1, S = 1 with a confirmed prediction on larger hardware, a significant measured rise of the improver (Rt+K > Rt), and H = 1 including an adversarial set. |
Rule 4 · The ladder index, with necessity built in
The geometric mean encodes necessity: a zero pillar gives a zero core and a zero index. The bracket follows the ladder’s stages: a complete core alone is worth 25 (Levels 2–3), the loop adds up to 25 (Level 4), control and held-out verification add up to 25 (Level 5), breadth and scaling add up to 25 (Level 6). The higher stages are additive inside the gate, so partial progress counts. Because the formula is monotone, the lower end of each bounded interval equals the strict value.
Rule 5 · The 51-entity line
| Criterion | Weight | What it measures |
|---|---|---|
| B1 · novelty of the mechanism | 20 | A new mechanism, as opposed to an integration of known parts. |
| B2 · breadth per unit of resource | 15 | Capability shown per unit of compute, log-discounted. |
| B5 · falsifiability and protocol discipline | 15 | Pre-registration, stated gates, negative results reported. |
| Q · measured ladder index | 15 | Q-ungated = Alternative A; Q-strict = the gated index I. Scaled from 0–100 to 0–10. |
Every development is reduced to research scale: compute, capital and deployment are removed, and each is judged on its own published evidence. The prose-judged ladder criterion of the earlier line (B6) is replaced by Q, so the ladder term is computed from measurements; B1, B2 and B5 are carried over unchanged from the own-evidence line.
Exclusions
Unmeasured counts as zero on the strict reading, and as the interval [0, 1] on the bounded reading.
Instruments and components with no ladder index are left out: METR, the CLadder-type causal benchmarks, the UK AI Security Institute and the Muon optimiser.
Enablers with no mechanism to score are left out: NVIDIA compute, TSMC, Huawei Ascend, national compute programmes, and announced laboratories with no public system.
Systems absent from the measured table are transformer-based with only C = 1 measured, and receive the floor value Q-ungated = 12.5.
B1, B2 and B5 remain judgements recorded in prose; only the ladder term is computed from measurements.
Hidden-benchmark seed values are not reproduced.
I · Are the Pillars Necessary?
What this section is for: to settle, from published theorems, which pillar is required, in what sense, and by what measurable criterion.
The lecture notes (v21.5) prove three narrow theorems: a round separation (local message passing needs at least D rounds, where D is the diameter, while one all-reduce suffices; physical time is Ω(N1/d) for either model), a two-model non-identifiability example with a worst-case error floor of 1/4, and a finite-group Reynolds risk identity. The broader triadic necessity theorem builds the three capabilities into its definition of scalable superintelligence and asserts that “these three obstructions are independent” without a separating construction; “Architecture from Theory” is labelled a conjecture.
| Pillar | Verdict | Theorem behind it | Criterion measurable on any system |
|---|---|---|---|
| CEAS | Necessary as a function, proven only in a unit-cost-broadcast (round-count) model; not necessary in physical time | Round separation (notes); LOCAL t-hop lower bound (Ghaffari); GNN depth×width bounds (Loukas) | Rounds exponent q in T(N) ∝ Nq on globally sensitive aggregation: local q = 1/d, functional CEAS q ≈ 0. Any all-reduce qualifies (global attention, a global mean, a bus, a tree). |
| Ψ · rungs 2–3 | Necessary as a function, proven | Causal Hierarchy Theorem (Bareinboim, Correa, Ibeling & Icard): rung collapse occurs only on a measure-zero set of models; meagre-set version (Ibeling & Icard); robust agents imply an approximate causal model (Richens & Everitt) | Rung-2 and rung-3 accuracy above chance; correct abstention on non-identifiable queries; the 1/4 error floor on observationally equivalent pairs. |
| Ψ · inverse design | Unproven as a separate necessity; reduces to rung 2 plus planning | None found | Success rate in reaching a target under constraints, against a forward-only baseline. |
| GRAIL | Efficient, not strictly necessary; a symmetry prior is necessary only if exact invariance on unseen orbit elements is required | Excess risk equals ‖f⊥‖2 (Elesedy & Zaidi); |G|-fold sample gain (Tahmasebi & Jegelka); canonicalisation is equivalent (Kaba et al.) | Retention on held-out group elements; εinv = E‖F(gx) − F(x)‖2; any map that factors through X/G qualifies. |
| Triad jointly | Unproven | Independence asserted, not shown | Main effects of a 23 ablation. |
Three concessions that bound the verdicts
- The notes state that standard transformer self-attention already realises an aggregate primitive, so the round theorem does not by itself distinguish CEAS from transformers.
- On the build’s own R4 tasks a global mean matched CEAS exactly, and the entropy-corridor controller failed its CEAS hypothesis test (mean −0.002), while the β-controller passed (+0.092).
- GRAIL’s necessity argument is a no-free-lunch argument that does not exclude canonicalisation or learned invariance, which grows with scale in vision models (Gruver et al.).
The Causal Hierarchy Theorem does not require an explicit causal-model engine: causal statements in training text, or experiments run through tools, can carry rung-2 information. Following the ladder strictly therefore means requiring the three functions, and treating GRAIL’s function as required only because the ladder writes exact invariance into its definition.
II · The Measured Ladder Index
What this section is for: the eight quantities computed for each system, the resulting index and level, and how far each result moves under alternative readings.
| System | C | Ψ | G | R | U | H | B | S | P | Strict I | Bounded I | Alt A | Level |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Logarchéon current build | 1 | 1 | 1 | 1 | 0.574 | 0.5 | 0 (·) | 1 | 1 | 75.9 | 75.9–88.4 | 75.9 | 5 (4†) |
| Hyperagents | 1 | · | · | 1 | 0.332 | 0.5 | · | · | 0 | 0 | 0–85.4 | 35.4 | 2 |
| GPT-6 Astra | 1 | · | · | 0 (≤0.5) | · | 1 | · | · | 0 | 0 | 0–87.5 | 25.0 | 2 |
| GPT-6.1 Sol, GPT-5.6 Sol, Opus 5, Opus 5.5, Opus 4.8, Gemini 3.8 Flash, Grok 4.6, Kimi K3, Qwen3.8 (27B), GLM-5.3, Seed, Sakana AI Scientist, Co-Scientist, AlphaProof, Aristotle | 1 | · | · | · | · | 1 | · | · | 0 | 0 | 0–100 | 25.0 | 2 |
| Darwin Gödel Machine | 1 | · | · | 1 | · | · | · | · | 0 | 0 | 0–100 | 25.0 | 2 |
| MACE / GNoME | 0.667 (≤1) | · | 1 | · | · | · | · | · | 0 | 0 | 0–100 | 20.8 | 2 (via G) |
| DeepSeek V4-Pro | 1 | · | · | · | · | 0.5 | · | · | 0 | 0 | 0–93.8 | 18.8 | 2 |
| Huxley-Gödel Machine | 1 | · | · | 0 | · | 0.5 | · | · | 0 | 0 | 0–93.8 | 18.8 | 2 |
| Self-improving embodied FMs | 1 | · | · | 0 (≤0.05) | 0.300 | · | · | · | 0 | 0 | 0–67.5 | 16.3 | 2 |
| π*0.6 | 1 | · | · | 0 | 0.25 (≤0.5) | · | · | · | 0 | 0 | 0–93.8 | 15.6 | 2 |
| AlphaEvolve | 1 | 0.238 | · | 0 | · | · | · | · | 0 | 0 | 0–99.6 | 15.5 | 2 |
| GPT-5.5 / 5.3 | 1 | 0.085 | · | · | · | · | · | · | 0 | 0 | 0–94.2 | 13.6 | 2 |
| Gemini 4 Argon, Muse Spark, V-JEPA 2-AC, Absolute Zero, Gemini Robotics 2, GR00T, AgiBot GO-2, AlphaFold 3 | 1 | · | · | · | · | · | · | · | 0 | 0 | 0–100 | 12.5 | 2 |
| Calibration: GPT-4 on CLadder | 1 | 0.117 | · | · | · | · | · | · | 0 | 0 | 0–85.1 | 14.0 | 2 |
| Calibration: Euclidean transformer inside the build | 1 | — | 0.300 | — | — | — | — | — | — | — | — | — | — |
How the non-trivial cells are computed
- Logarchéon current build. Ψ = mean(24/24, 5/5, 26/26, 100/100) = 1 with c = 0. G = 1.000/1.000 on SL2(ℝ), with εtwin ≈ 10−13. R = min(1, 40/20)·1 = 1, and R2 = 0.80 (31–33 of 40 cycles). U = mean over six seeds of (1 − loop/frozen) = 0.574, multiplied by min(1, 24/2) = 1. S = 1 from the token-axis difference interval [0.692, 0.936]. I = 100·1·(1 + 1 + 1.074/2 + 1/2)/4 = 75.9.
- Hyperagents. U = (0.372 − 0.060)/0.94 = 0.332 over 100 iterations of robotics reward design. H = 0.5 because the held-out domain was chosen by the developer.
- Self-improving embodied FMs. g takes three values: (0.875 − 0.63)/0.37 = 0.662 on the real robot, (0.75 − 0.45)/0.55 = 0.545 in simulation, and 0.595 on BananaTable. Their mean is 0.601; one phase halves it to U = 0.300.
- π*0.6. The primary abstract reports that the method roughly halves the task failure rate, so g = 0.5; the cycle count is unknown, so one phase halves it to U = 0.25.
- AlphaEvolve. The target was matched or beaten on about 75% + 20% of 50+ problems, so inverse-design success = 0.95 and Ψ = 0.95/4 = 0.238. Crediting a design task run against an exact evaluator as inverse design is a judgement.
- GPT-5.3. Complete identification specification was 34.1% (c = 0), so Ψ = 0.341/4.
- GPT-4 on CLadder. Ψ = [(0.628 − 0.5)/0.5 + (0.606 − 0.5)/0.5]/4 = 0.117.
- GPT-6 Astra. The upper bound on R comes from the report that more than half of 4–8 hour research tasks needed at least one human intervention, which gives h ≥ 0.5.
- H = 1 assignments come from ARC Prize semi-private scores, live IMO contests, blind peer review and an independent experiment. DeepSeek receives 0.5 because NIST CAISI found it performed worse than its own report suggested.
Sensitivities
| Reading | Effect |
|---|---|
| U read against the loop’s own miss (U = 0) | Build falls to 68.8 (Level 4). |
| Wilson 95% lower bounds on the small Ψ samples (5/5 gives 0.566; Ψ = 0.815) | Build 70.9. |
| Both of the above | Build 64.2. |
| Alternative B (proof-weighted gate) | Build 80.7, or 75.0 with U = 0; AlphaEvolve 9.7, GPT-4 6.8, GPT-5.3 5.8 — the only aggregation that lifts other systems above zero. |
| Unreported h counted as 1 | Hyperagents’ Alternative A falls from 35.4 to 22.9; the Darwin Gödel Machine’s from 25.0 to 12.5. Strict index unchanged. |
| CausalDS pass rate read as pooled accuracy | Opus 4.8 and GPT-5.5 Ψ = 0.412, Gemini 3.1 Pro 0.383; strict index still 0, because G is unmeasured. |
The index line
STRICT INDEX (necessity-gated)
0 20 40 60 80 100
|---------|---------|---------|---------|---------|
X ← 33 systems at 0 (Astra, Opus 5.5, Gemini 4, Hyperagents, DGM, AlphaEvolve, robots, provers …)
L (Logarchéon build) 75.9
bounded: build [75.9 ===== 88.4] EFM [0 ====== 67.5] Hyperagents [0 ======= 85.4]
Astra [0 ======= 87.5] most others [0 ========== 100]
ALTERNATIVE A (arithmetic mean, no necessity gate)
0 20 40 60 80 100
|---------|---------|---------|---------|---------|
^12.5 Argon, Muse, V-JEPA, robots, AF3
^15.5-16.3 AlphaEvolve, π*0.6, EFM
^18.8-20.8 DeepSeek, HGM, MACE
^25.0 ARC-verified LLMs, DGM, provers, Co-Scientist, Sakana
^35.4 Hyperagents
^75.9 Logarchéon buildIII · The 51-Entity Line
What this section is for: every development on one [0, 100] line, with the ladder term computed from measurements, under both readings.
Indicator Table
RANK = capability rank (0–12) · LEVEL = certification level (functional) · B1 novelty · B2 breadth per resource · B5 protocol discipline · Q = measured ladder index
| # | Development | Group | Rank | Level | B1 | B2 | B5 | Q-ung. | Ungated | Q-str. | Strict |
|---|
The leaders on each reading
| # | Ungated reading | Score | # | Strict reading | Score |
|---|---|---|---|---|---|
| 1 | Logarchéon current build | 61.4 | 1 | Logarchéon current build | 61.4 |
| 2 | Hyperagents | 60.5 | 2= | AlphaEvolve | 52.3 |
| 3 | AlphaEvolve | 55.9 | 2= | Hyperagents | 52.3 |
| 4 | Darwin Gödel Machine | 55.8 | 4 | Darwin Gödel Machine | 50.0 |
| 5 | Equivariant scientific ML | 54.0 | 5= | Equivariant scientific ML | 49.2 |
| 6 | Self-improving robot models | 53.0 | 5= | V-JEPA 2-AC | 49.2 |
| 7 | AlphaProof | 52.7 | 5= | Self-improving robot models | 49.2 |
| 8 | V-JEPA 2-AC | 52.1 | 8= | AlphaProof, Absolute Zero, Huxley-Gödel Machine | 46.9 |
IV · Firmness of the Lead
What this section is for: how much of the first place survives a change of reading.
| Reading | Logarchéon build | Nearest rival | Margin | What the margin rests on |
|---|---|---|---|---|
| Ungated, Q = 75.9 | 61.4 (1st) | Hyperagents 60.5 | +0.9 | Measured pillar evidence; Hyperagents leads on novelty (B1 8 against 4.5). |
| Ungated, conservative Q = 64.2 | 58.7 (2nd) | Hyperagents 60.5 | −1.8 | Control judged against the loop’s own error, and cautious bounds on the small causal samples. |
| Strict, Q = 75.9 | 61.4 (1st) | AlphaEvolve, Hyperagents 52.3 | about +9 | The rule that an unmeasured pillar counts as zero: others lose the ladder term because their pillars have not been measured, not because a measurement failed. |
| Prose-judged ladder term (earlier line) | 60.0 (9th–11th) | Hyperagents 73.1 | −13.1 | The ladder criterion judged in prose instead of measured. |
V · What the Numbers Cannot Say
What this section is for: the exclusions of the evidence, stated as prominently as its findings.
- Equal numbers do not mean equal capability. The build’s values of 1.0 are exact results on its own synthetic tasks: 24 causal questions, the SL2(ℝ) and PSL2(ℤ) groups, a two-dimensional point mass, 16 causal blocks and at most 40 cycles on one workstation, measured against correlation-only baselines and SmolLM2-135M. Frontier numbers come from public benchmarks of a different order of difficulty; ARC-AGI-3 measures novel-task generalisation, not group invariance.
- Most frontier pillar cells are unmeasured. The strict index records demonstrated ladder evidence; the bounded intervals of 0–100 show how little the public record constrains those systems.
- The ladder was written around one architecture. Its Level 5 checks (εtwin ≤ 10−6, a hash-committed hidden set, an R2 rate) are quantities the build was designed to report, and the coordination term gives every transformer the same score as CEAS.
- Three of the four line criteria remain prose judgements (B1, B2, B5).
- Verification is internal. The build is self-certified; its reviews were carried out by separate AI agents reading the same files, and no outside party has rerun it.
- Two hard exclusion conditions of the examination are not cleared (transfer to an unseen task family; a falling self-improvement rate), and two are partly cleared (one task family per loop; no adversarial benchmark). The Level 5 verdict applies the certification report’s rule and is stated together with these conditions.
- Level 6 is not met. Breadth B = 0; no prediction to larger hardware has been tested; the improver grew once (+0.0041, 9 of 16 tasks won, one-sided sign-test p = 0.40); no adversarial set exists.
VI · Plan of the Study
What this section is for: the one route to a comparison measured throughout — the same battery run on every system that can be run.
- Choose models. Open-weight models with ARC Prize scores: Qwen3.8-27B, DeepSeek V4, GLM-5.3 and Kimi K3; frontier APIs where access exists.
- C: fit q from the number of reasoning steps needed against N on globally sensitive aggregation and parity tasks.
- Ψ: run the 24-question suite, the action-order demonstration, the non-identifiable abstention items, physical inverse design and the per-rung CLadder splits, at the stated chance levels.
- G: measure seen/unseen retention and εinv on SL2(ℝ) and PSL2(ℤ) inputs, presented in context and after fine-tuning.
- R, R2 and U: place each model as the proposer inside the same 40-cycle gated loop, log h, and run the same control plant.
- H: score on the unused hash-committed hidden seeds, plus an adversarial set built by a third party.
- S: use the scaling axes of battery R4 with bootstrap intervals.
- Release: publish the audit bundle (scripts, logs, Lean files and used seeds) so that every number above can be checked from the files alone.
Runs on 27–30B models fit on one or two large GPUs; the loop runs (40 cycles × 3 seeds per model) are the main cost. Every cell would then be measured on identical items, and the bounded intervals would collapse to points.
Sources
- Bareinboim, Correa, Ibeling & Icard, “On Pearl’s Hierarchy and the Foundations of Causal Inference” (causalai.net/r60.pdf); Ibeling & Icard (arXiv 2107.08558); Richens & Everitt (arXiv 2402.10877); Zečević et al. (arXiv 2308.13067).
- Ghaffari, LOCAL model notes (MIT); Loukas (arXiv 1907.03199).
- Elesedy & Zaidi (arXiv 2102.10333); Tahmasebi & Jegelka (arXiv 2303.14269); Kaba et al. (arXiv 2211.06489); Gruver et al. (arXiv 2210.02984).
- CLadder (arXiv 2312.04350); CausalReasoningBenchmark (arXiv 2602.20571); ARC Prize results (arcprize.org).
- Hyperagents (arXiv 2603.19461); self-improving embodied foundation models (project page); π*0.6 (arXiv 2511.14759); AlphaEvolve (Google DeepMind).
- OpenAI research-automation report (Help Net Security, 7 September 2026); NIST CAISI on DeepSeek (The Decoder).
- Logarchéon, Lecture Notes: From Physical AI to ASI, v21.5 (October 2026), including Research Build I–XI; certification battery report and pass-by-pass results record of the build.