How the ARQ is scored
This page is the scoring contract. It exists so the index is auditable: what an entry's numbers mean, how they combine, and what the whole thing cannot tell you. The rubric is stated here on purpose — before you read any particular score — so numbers are not reverse-engineered to fit a wished-for result.
The scale
One number, 0–100. It is a balance between risk signals and mitigation signals in a rolling 14-day window — not a count of news stories.
Red is dangerous, green is good, 50 is a perfectly balanced ledger. Values are whole numbers only — the underlying scores are ordinal judgements, not precise measurements.
The rubric
Every entry is given a risk weight and a mitigation weight, each 0–5. A story that is genuinely both good and bad is given a value on both sides. The definitions are:
Risk weight — how much this raises the ARQ
| Score | Meaning |
|---|---|
| 1 | Minor. Framing or concern only; small trust or stability impact. |
| 2 | Moderate. A setback, minor misuse, or a notable capability bump. |
| 3 | Significant. A real escalation — a breach, a misuse case, a capability jump. |
| 4 | Major. Weaponisation or misuse, a monitorability regression, a top insider warning. |
| 5 | Critical. Documented lethal or autonomous misuse, or a systemic governance failure. |
Mitigation weight — how much this lowers the ARQ
| Score | Meaning |
|---|---|
| 1 | Reassurance or framing only. |
| 2 | A real but partial step — a policy, a safeguard, a disclosure with teeth. |
| 3 | A substantive improvement — governance in force, a verified clinical or scientific result. |
| 4 | A major beneficial development — a breakthrough with evidence. |
| 5 | Transformative — a verified cure or a binding global instrument (rare). |
Two modifiers change how a score is counted: verified (unverified or disputed claims count at half weight) and primary event (a re-report of the same event is not counted as a fresh signal).
The index
ARQ = 100 · (R + α/2) / (R + G + α)
- R — decayed sum of risk weights over the window.
- G — decayed sum of mitigation weights over the window.
- α — a prior of 12 pseudo-events, which pulls thin samples toward 50 and stops a quiet week pegging the scale.
Each event decays with a 7-day half-life and falls out entirely after 14 days. The denominator includes G, so any mitigation at all prevents a saturation at 100 — the index cannot sit at its ceiling during a busy week of good news.
Attribution and the frozen record
"Why the index moved" uses leave-one-out: each figure is what the ARQ would change by if that single entry were removed. It is a local, not a marginal, measure — but it is the honest answer to "what is moving this number?"
The historical line comes from a record that is written down once and never rewritten — currently 258 days, 2025-12-31 → 2026-09-14. A daily job appends each day's value; past days are immutable, so the chart is evidence rather than a recomputation that drifts with every rebuild.
The scorecard (honesty by design)
Every weekly digest states, in advance, what would move the index the following week — then grades its own previous calls as hit / miss / partial. The track record is shown publicly. A publication that will not grade its own predictions should not ask you to trust its index.
What this cannot tell you
- It measures documented discourse and incidents, not raw capability. A rise can come from better disclosure rather than worse reality; the index is a lens, not a radar.
- It is one scorer's judgement. The rubric makes it auditable, not objective. Do not read two decimal places into it — hence whole numbers.
- It cannot see the transition. Superintelligence is partly defined by capability outrunning our ability to check it; a monitorability collapse is precisely what would break an index built from human observations. The number may get noisier on approach and unreliable at the moment it would matter most.
- It is not an assessment of any company, nor financial or investment advice.