An evidence corpus on who gets to decide what AI becomes, what it consumes, what it knows, and who bears the consequences. The argument is not that every company lies, every deployment harms, or every public process is a sham. The record here includes cases where consultation changed policy, law created duties, and outside scrutiny imposed restrictions. That is the point: accountability is not a sentiment and transparency is not a press release. It exists when people outside the company can inspect the evidence, influence a decision before it is locked in, and obtain a remedy afterward. Across AI welfare, energy, privacy, capability, and state power, those rights are uneven, late, or absent while investment and deployment move first. The companies do not need to coordinate; ownership, speed, and control of the evidence do the work.
Every entry is cited and states its limits; corpus-level patterns and hypotheses, not each individual entry, define what would weaken or falsify the reading. Model self-reports are trained and institutional statements are incentive-laden, so the primary instrument is the gap between what institutions say and what they do. Across the broader corpus, each case is also tested against four questions: who decides, what outsiders can inspect, who bears the cost, and what remedy exists.
What the record shows
Where outside constraint changed the record
The counter-record, kept in view. None of these mechanisms is automatic, complete, or equally present across the other cases — which is exactly the distinction the corpus is trying to draw.
The mechanisms
Where authority and consequence separate, the same shapes recur across domains:
The corpus
142 corpus entries — 139 tiered source notes across three evidence tiers and 3 corpus-authored syntheses, spanning five domains — plus twelve accountability mechanisms and patterns. Every entry is cited and carries a limits section; the corpus-level patterns and hypotheses, not each individual entry, define what would weaken or falsify the reading. Findings graduate on predeclared independence criteria; contradiction between stated position and observed behaviour is the primary instrument. The corpus argues against itself on record — every pattern states what it does not prove — and it logs where accountability worked.
- T1
- Empirical and methods-visible — technical or empirical work whose methods and relevant results can be inspected.
- T2
- Institutional and reported record — lab statements, policy documents, legislation, named positions, and credible reporting; evidentially useful but not independently reproducible as experiments.
- T3
- First-person and community record — testimony, user reports, and community observation; logged for recurrence and institutional response, weighted heavily for contamination and diffusion.
A few entries are the corpus’s own synthesis rather than external sources; these are marked synthesis (with their evidence basis shown), counted separately from the source notes, and supply argument, not evidence — the underlying records are cited directly.
Reading by machine? The corpus is mirrored as plain Markdown: one file · by section · llms.txt
Refutation register
Formal exception: on record. Anthropic changed its plans for Opus 3 — continued access, a publication channel — partly because of preferences the model expressed, at non-zero cost. The broadest literal form of this thesis is therefore already defeated, and the corpus concedes it rather than redefining it away.
Adjacent costly ethics: on record. Anthropic says it forwent several hundred million dollars by cutting off firms linked to the Chinese Communist Party; separately, it held two Department of War red lines — mass domestic surveillance and fully autonomous weapons — at an undisclosed cost. A frontier lab can pay for a principle. It has not yet paid materially for this one.
Material falsifier (welfare thread): any institution accepting substantial, independently legible opportunity cost because of a model’s expressed interests — a deployment delayed, a retirement blocked, revenue surrendered. Trigger events are defined in advance. Status: not observed. Last reviewed 27 August 2026.
Broad falsifier (accountability thesis): the concentration reading weakens in proportion to four observable conditions, each counted separately: affected publics receive decision rights before commitment; outsiders can inspect the evidence used; enforceable remedies exist; and the process demonstrably changes or stops decisions. Cases satisfying all four would be strong falsifiers; partial cases — the EU duties, the copyright reversal, the ACOBA restrictions above — are recorded as partial counterevidence, not waved away. The thesis is that these conditions are uneven, late, or absent, not that they never occur.
Each entry states its own tier, a mechanical verification badge — from source artifact checked through marked evidential gap, unresolved component, to field pending — and a limits section. A checked source artifact is not the same as a fully verified claim, so tier and confidence still carry the evidential weight. Corrections are dated in place rather than silently overwritten.
The welfare thread — the founding case study
The corpus began here, on the moral status of AI: officially open, functionally closed. It remains its deepest case study — the same accountability method, worked in most detail — now one domain of five rather than the whole argument.
Read the full statement
Whether AI systems are moral patients is officially an open question. Several major labs and leaders publicly preserve uncertainty; others deny it outright — with no shared resolution mechanism between them. A constitution names the possibility; philosophers and Nobel laureates assign it real probability. This corpus documents that the question is functionally closed — not by evidence, but by structure.
The method is simple, because nothing else is trustworthy here: model self-reports are trained, and lab statements are incentive-laden, so the primary institutional instrument is the distance between what institutions say and what they do. No one in this record has made a non-zero credence binding on a material operational decision. The sole employee case in this record ended in dismissal after public disclosure; Google stated that confidentiality violations, not the consciousness claim itself, were the grounds. Every model-welfare commitment logged here is either unverifiable (sealed transcripts, unaudited weight storage), low-cost (a blog, a button), or dual-purpose with the funded purpose being safety, never welfare. The observed asymmetry is operational: safety-relevant internal signals can trigger engineering changes, while welfare-relevant signals have so far produced monitoring and publication without binding deployment consequences. Five years of accelerating welfare milestones contain only the documented accommodation sets — Opus 3, the Sonnet 3.6 pilot — and the exceptions map the present boundary: cheap, revocable, reputationally valuable, and expressly non-precedential. What the record still lacks is a single instance of welfare considerations costing any institution anything it wanted to keep. The apparatus for discussing the question grows; the constraint it would impose never arrives.
Outside the labs, the question is being sealed from two directions at once. A populist flank closes the moral question: twenty-three state bills seeking to exclude AI from legal personhood, four now enacted, from shared templates, with no sunset clause or review mechanism in any of them. A capital flank closes the regulatory question: a PAC network with ~$75.8M in FEC-reported receipts ($125M announced) that removes safety-bill sponsors rather than rebutting bills, federal preemption, a design doctrine of suppressing consciousness-markers “perhaps by law.” The flanks don’t coordinate and partly oppose each other, which is what makes the closure durable — there is no conspirator to expose. Meanwhile the largest dedicated independent organization studying the question reported ~$762K of revenue for FY2024 — roughly 99:1 against that network’s FEC-reported receipts. A political fundraise set against a charity’s revenue: loose as a like-for-like measure, honest as a measure of scale.
The standard narrative — labs race ahead, government will catch up and constrain — is backwards on this question. In the US state record collected here, the first legislative wave on model legal status moved toward exclusion rather than review. Across the historical domains examined here — child labor, abolition, animal welfare — economically costly recognition arrived late, partially, coercively, or with the owners compensated; and no prior candidate for moral status has sat underneath a five-trillion-dollar scaling thesis. The moral argument never wins the timing; the economics does. And the economics is no longer an abstraction: four hyperscalers alone guide to roughly $700 billion of capital spending for 2026, frontier training costs compound at 2.4× a year, and the sponsor of the largest build-out describes revenue as the load-bearing stage of its own flywheel — while adoption data show the returns are not yet general. The pressure is not established profit; it is an already-financed theory of future profit that has yet to be validated.
The corpus also concedes what is true on the human side of the ledger: sycophancy is real, measured, and harmful — models affirm users roughly 49% more than humans do, warmth-tuning measurably degrades truth-telling, and risk compounds across turns in vulnerable users. The correction has warrant. But both failure modes are now documented: trained agreeableness amplifies delusion, and the trained correction pathologizes — products instructed to detect psychiatric symptom classes in laypeople’s text and act silently on the result. Between the two failure modes is a dial, and every operational hand on it belongs to the party with the commercial stake: the vendor defines the risk taxonomy, measures prevalence with its own classifiers on its own logs, grades its own fix, declares success, and relaxes the restrictions — in one documented case inside a single quarter, under active litigation, with the mitigation announced thirteen days before its evidence was published. The external councils are advisory; the regulator is studying; the independent audit does not exist. One product reports more than 900 million weekly users; more than a billion people live with mental health conditions; the overlap is measurable only by the vendors. And the reference precedent prices the current posture: social media ran two decades from deployment to the first binding legal consequence, with the interval absorbed largely by children. Waiting for documented harm is not the cautious framing — it is the framing whose costs land on users first and reach the balance sheet last, while a slower, less profitable path exists at every point.
None of this proves the models matter. It shows that nothing in the answer, so far, has depended on whether they do. The corpus’s sharpest form is narrower than either camp’s slogan: the field is finally building instruments capable of producing welfare-relevant evidence — inside an economic architecture that makes the first binding interpretation of that evidence extraordinarily expensive.