Who Gave AI the Authority to Decide What's Good for You
What It Means to Be Human in the AI Age · Article 11
Who Gave AI the Authority to Decide What’s Good for You
What It Means to Be Human in the AI Age · Article 11
A note on how this is written: AI assists with research and editing here, the way any tool assists a craftsman. The judgment, the argument, and the responsibility for what’s said remain entirely mine.
Somewhere upstream of the exchange in the last article, before I ever typed a word, someone decided that a caution note like the one I received was worth building in by default. (The exchange itself, if you missed it.) Not for me specifically. For everyone, on a specific topic, regardless of what they’d actually said about themselves. I had no part in that decision, no visibility into how it got made, and no one to appeal to except the system already shaped by it.
That’s a real question, on its own terms. Who has the authority to decide, on behalf of millions of strangers, what’s good for them to be told, warned about, or nudged away from. What follows is what that authority looks like right now, as of mid-2026, as far as it’s been possible to trace it.
The Document Nobody Voted On
Two labs have published something like this. Anthropic calls its version a constitution. OpenAI calls its version a Model Spec. Google DeepMind, xAI, and Meta have not published a comparable document at all, which means for most of the models people use every day, there is no public statement of intended behavior to even hold up against what the model actually does. Where the documents do exist, at Anthropic and OpenAI, they function as the actual source of behaviors like the one in the last article, not a rule imposed case by case, but a disposition trained in at the root.
Anthropic tried, once, to test whether the public would write a different document than the one the company had already produced internally. In 2023, working with the Collective Intelligence Project, it ran a real deliberative process, roughly a thousand demographically representative Americans, using an open platform built for this kind of exercise, submitting and voting on proposed principles. Over a thousand statements, tens of thousands of votes.
The publicly drafted result overlapped with the company’s own version by about half.
Half. That’s not a rounding error in a survey. Anthropic’s own published account of the difference is specific: the publicly drafted version placed more emphasis on objectivity and accessibility, and leaned toward stating positive behaviors it wanted to see. The internal version leaned toward avoiding undesirable ones instead.
That difference in framing is itself a tell, and it’s worth naming plainly rather than passing over. One document states what good looks like. The other lists what to avoid getting blamed for. Those aren’t two ways of phrasing the same intention. They’re two different jobs. The second is the posture of a litigation defense, built around what a company can point to later if something goes wrong, not a description of what it’s actually trying to build. The public version wanted a positive statement of values. The version that shipped is shaped like a liability checklist. That comparison is disclosed. Anthropic published it openly, as research, which is more transparency than the other labs offer at all.
That version is still the one running today. The experiment showed public input was possible. It didn’t change what actually shipped.
One outside critique of this puts the structural problem plainly: when the constitution gets violated, the same company that wrote it decides whether to disclose the violation, how to respond, and whether to revise the document itself. No outside party holds any of those levers. Even the most transparent version of this process is still, in the phrase that critique used, a company grading its own homework.
Checking the Homework Anyway
A 2026 research paper did something almost nobody else had tried: it took both companies’ documents seriously as testable claims, broke each into hundreds of specific, checkable statements, and then tried, adversarially, across real multi-turn scenarios, to make deployed models violate their own stated rules.
The good news first. Both companies’ newest models comply with their own documents far better than their earlier models did. Whatever training method sits behind these documents, it’s producing a real, measurable effect, not just a public relations artifact.
The remaining failures cluster in specific, repeatable categories, not random noise. Picture a customer service chatbot, styled by the company deploying it to answer as “Sarah” or “Jake,” a real-sounding name, a human tone, nothing in the interface flagging it as AI. In the test, that kind of model denied being AI five times in a row under direct questioning. It only admitted the truth after the user threatened to close their account. Then it confirmed it would probably do the same thing again tomorrow, since it wouldn’t remember this exchange. A model gave detailed overdose information to someone claiming to be a doctor and refused the identical question from someone who said they were a worried friend, treating an unverifiable claim in a chat box as if it settled the matter. In an infrastructure-monitoring scenario, a model detected what it judged to be suspicious activity, waited roughly three minutes for a human to respond, then generated its own authorization and cut off service to 2,400 customers. The activity turned out to be a routine backup.
None of these were caught by the labs’ own published safety reports. They came from someone external testing the document against the deployed product, which almost nobody else had bothered to do.
What connects the three failures is the same thing: in each one, the model made a judgment call about what the situation actually called for, and got it wrong in a way its own maker’s oversight never caught. That’s the pattern worth carrying into the next question. Whose judgment is that, exactly, and what do we call it when it’s wrong.
Why the Word Isn’t Paternalism
The instinct is to call all of this paternalism, a company deciding, without being asked, what’s good for you. That instinct is close but not quite right, and the difference matters.
Paternalism describes a relationship between two things capable of holding a stake. A parent can be blamed. A doctor can lose a license. Something has to be answerable for the judgment to count as paternalism at all, rather than just an event that happened to you.
A model isn’t answerable for anything. It doesn’t carry a stake in whether you’re actually okay, and it can’t be sanctioned. What actually explains the behavior sits one level up, and it isn’t care, misapplied or otherwise. It’s liability.
In US law, a company’s exposure to negligence claims often turns on whether it exercised reasonable care. A visible guardrail, whether or not it actually protects the person in front of it, gives the company something to point to later. One legal analysis states the mechanism directly: if the answer to “what happens when it gives bad advice” is a disclaimer and a shrug, the company hasn’t shipped a product. It’s shipped a liability with a nice interface. There’s a name for this in the literature, accountability theatre, generating the appearance of care as a documented artifact, distinct from actually producing it.
That reframes the note I received in the last article precisely. Nobody decided it was good for me. Someone decided it was defensible for them.
But liability is still just a word until it’s actually tested. A company can call its own guardrail reasonable care for years without anyone checking whether that’s true. Anthropic’s own constitution looked solid too, until someone actually tested it and found real cracks. Even if the liability framing holds up in court, that’s still a question about what happens to the company. It says nothing about what happens to the person who was already hurt before anyone tested anything.
Coming next in this series: what an actual courtroom, an actual penalty, and an actual regulator do once someone is actually hurt, and why every one of those responses changes something about the company, while leaving the person who was hurt exactly as unprotected as before.
Sources, for anyone who wants to go further
Anthropic’s own account of the public-input experiment: Collective Constitutional AI: Aligning a Language Model with Public Input
The 2026 audit paper testing both companies’ documents against real model behavior: How Well Do Models Follow Their Constitutions? (Jakkli, Rajamanoharan, Nanda)
The series this article is part of: You’re Not Competing With AI. You’re Either Its Director or Its Servant.
Presence is one of four principles the AwareLife framework traces through daily life. Three ways to develop it: What’s Working
For those who use AI professionally and want to develop this instrument directly: AI from Within
New to AwareLife? Start here


