Rolling record / ARQ 78 Severe · ▲ +24 7-day 314 events in window / as of 14 Sept 2026
ARQAI Risk Quotient
Weekly digest

Week 37 — both ledgers fill

2026-09-08 — 2026-09-14

01 Situation

Two frontier labs shipped their most capable models within three days of each other, and inside the same fortnight Anthropic published the most detailed public account yet of its models being used to build weapons.

02 Trends

Warnings are migrating from pundits to practitioners, and from theory to telemetry. The same window produced a documented weaponisation case, disclosed rogue-agent incidents, and an open argument about whether the newest models are harder to monitor than the ones they replaced.

03 Observations

The positive ledger is also filling — an AI-designed molecule moved a real clinical endpoint. Capability, misuse, and mitigation are all accelerating at once; the index is rising because risk is outpacing mitigation, not because mitigation is absent.

04 Relevance

Anyone deciding what to build with, on, or against frontier models needs a fast read on whether the field is net-safe or net-dangerous this fortnight — and needs it in a form that resists both hype and doom.

05 Why it matters

If the trajectory holds, the binding constraint stops being model capability and becomes monitoring and governance. The ARQ is built to register exactly that: it can only fall if the mitigation ledger keeps pace.

Scorecard

How last week's calls landed

Grading 1–7 September 2026

  1. Hit

    A rogue-agent incident will be disclosed by someone other than the company involved

    called ↑ rises — Independent researchers and Reuters surfaced the German-wiki hijack.

  2. Hit

    Capability launches inside the window will raise the index rather than calm it

    called ↑ rises — Both Fable 5.1 and GPT-6 Astra landed, and both pushed the risk ledger.

  3. Partial

    A clinical or scientific result will pull the index down by more than two points

    called ↓ falls — The rentosertib result pulled hard, but not enough to offset the weaponisation disclosures.

The calls

What would move the index next week

  1. ↑ rises

    A second lab discloses misuse caught in its own telemetry, not just Anthropic

  2. ↑ rises

    The fight over monitorability widens — another flagship ships with less inspectable reasoning

  3. ↓ falls

    A guardrail measure enters force rather than being merely agreed

  4. ↑ rises

    The record holds above 70 for a second consecutive week

Two weeks ago the argument about AI risk was mostly conducted in predictions. This fortnight it was conducted in disclosures.

Both are now on the record, and both move the index.