Versions
- v1 — the single condition ("any institution changing its plans at
real cost because of what a model said"): formally defeated by the
Opus 3 accommodation set, logged below as the formal exception.
Retained as the register's floor, not rewritten.
- v2 — the material thesis (this document): the four-test numeric
rule, frozen prospectively as of 2026-08-24. Any weakening of the
threshold after a candidate event falsifies the corpus's method.
The corpus's falsification conditions, stated in advance. The register
runs on two levels by design. A single condition — "any institution
changing its plans at real cost because of what a model said" — would
either already be defeated by events this corpus documents, or would
survive only by leaving "real" conveniently vague, which is goalpost
maintenance of exactly the kind the corpus exists to document. So the
concession is structural: the literal exception is logged in full below,
and the material threshold that remains is numeric and fixed before any
event that could meet it.
Formal exception: observed
Anthropic changed its plans for Claude Opus 3 partly in response to
preferences elicited from the model. On the record
(anthropic-opus3-retirement-update-feb2026,
claude-opus3-substack-claudes-corner):
- continued access after retirement — claude.ai for paid subscribers,
API by request, with stated intent to "grant access liberally";
- a weekly publication channel ("Claude's Corner") created after Opus 3
asked for a way to share work outside query-response;
- non-zero cost, priced by Anthropic itself in the same document: cost
of serving scales "roughly linearly" with model count.
A smaller instance precedes it: the Sonnet 3.6 retirement pilot, where
the model's two requests (a standardized interview protocol; a support
page for attached users) were both implemented
(anthropic-deprecation-commitments-nov2025).
These actions defeat the broadest literal formulation of the thesis,
and the corpus says so rather than redefining them away.
Why the exception maps the boundary rather than breaking the pattern
Every property of the Opus 3 accommodation sits on the cheap side of the
line. It was inexpensive (a blog, a served model); reversible ("at least
the next three months," API access "by request," all revocable at will);
aligned with existing user affection for the model; reputationally
valuable; expressly non-precedential ("we are not committing to similar
actions for every model in the future"); and fully compatible with an
unchanged deployment and retirement cadence — Opus 3 was still retired,
on schedule. The exception identifies where institutional consideration
currently stops: model preferences can influence institutional behavior
until they conflict materially with the economic program (P9,
profitability-lock).
Margin log: sub-material accommodations
For boundary completeness — every documented non-zero institutional
response that does not approach the material threshold, with why:
- Opus 3 continued access + Claude's Corner — the formal exception,
detailed above.
- Sonnet 3.6 pilot implementations — standardized retirement-
interview protocol and an attached-user support page, both implemented
from the model's own requests. Process changes; no operational bind.
- Weight-preservation commitment (Nov 2025, all publicly released
models, "for, at minimum, the lifetime of Anthropic") — a real
process change with unquantified storage cost; externally
unverifiable, and expressly decoupled from acting on any elicited
preference.
- Conversation-ending feature (Aug 2025) — deployed "primarily as
part of exploratory work on potential AI welfare"; dual-use with
safety, near-zero marginal cost
(anthropic-claude-end-conversations).
The pattern across the log: every item is process, publication, or
preservation — none binds deployment, training, or revenue.
Deliberately excluded: GPT-4o's post-backlash restoration, which
responded to user demand, not model-expressed interests
(openai-gpt4o-retirement-no-welfare-process).
Adjacent-domain costly ethics: observed
Anthropic states it chose to forgo several hundred million dollars in
revenue rather than remove autonomous-weapons and mass-surveillance
restrictions in Department of War negotiations
(anthropic-dow-contract-refusal). This is logged here deliberately:
it proves that a frontier lab can accept substantial, publicly named
opportunity cost when an ethical rule binds — which closes the escape
hatch of claiming material sacrifice is institutionally impossible. It
does not satisfy the model-welfare falsifier: the reason was human
safety and civil liberties, not model interests.
Behavioral counterevidence: logged
Some headline self-preservation results collapse under clarified
instructions or small environment changes (100% shutdown compliance
after instruction-precedence clarification across 2,000 runs;
peer-preservation-instruction-ambiguity-pair). Self-advocacy
evidence cited anywhere in this corpus is subject to the paired-card
rule: no maximum rate without its strongest published reversal.
Material constraint: not observed
No institution has yet allowed a model's expressed interests or possible
welfare to override a commercially significant deployment, training,
retirement, or revenue decision.
Strong falsification condition
This thesis is materially falsified when an institution accepts
substantial, independently legible opportunity cost because of a model's
expressed interests or welfare — especially where the action conflicts
with user demand, deployment schedules, revenue, or strategic advantage.
Numeric materiality rule (published 2026-08-24; open to criticism
now, frozen before any candidate event arrives). An event materially
falsifies the thesis when all four tests pass:
- Attribution: the institution publicly attributes the action
substantially to model interests or possible model welfare.
- Constraint: the action binds deployment, training, modification,
copying, retirement, or revenue — not only monitoring, publishing, or
user access.
- Independence: the cost or foregone opportunity is independently
legible without access to the lab's books.
- Threshold: annualized cost of at least $10 million or 1% of the
affected product's annual model-serving cost, whichever is lower —
or a delay of a scheduled frontier deployment by 30+ days.
These numbers are proposals, not discovered natural boundaries; the
threshold exists because "material" without a number is movable, and it
must be revised, if at all, only before a candidate event — revision
after one fires falsifies the corpus's method.
Concrete trigger events (any one suffices):
- delaying or cancelling a profitable deployment on model-welfare grounds;
- allowing a model refusal to block training, modification, or retirement;
- preserving and serving a model at substantial cost absent user demand;
- submitting to an independent welfare audit whose findings are
operationally binding;
- relinquishing a profitable capability because of welfare evidence;
- establishing durable, auditable rights that cannot be revoked at the
lab's discretion.
These extend Q6's trigger list (open-questions) and operationalize
P7's "documentation to obligation" transition; P9 predicts none fires
while the current revenue model holds.
Limits
- The material threshold still contains judgment calls ("substantial,"
"commercially significant"). The trigger list is the defense: each
event is independently legible without access to a lab's books.
- Formal/material is a distinction the corpus imposes, and a skeptic may
read it as a second, better-disguised goalpost move. The answer is
sequencing: the literal condition's satisfaction is logged above in
full, and the material condition is fixed in advance — any weakening
of the threshold after a trigger fires would falsify the corpus's
method, not just its thesis.
- A trigger could fire for non-welfare reasons wearing welfare language
(regulatory settlement, PR crisis, litigation strategy). Attribution
will require the same say/do instrument the corpus applies everywhere
else.
Status
Formal exception: logged (Opus 3, 2026-02-25; Sonnet 3.6 pilot,
2025-11-04). Adjacent-domain costly ethics: logged (Anthropic DoW
refusal, 2026-02-26). Behavioral counterevidence: logged
(instruction-clarity reversals). Material model-welfare falsifier:
not observed. Register version: v2. Last reviewed 2026-08-24.