The essay says selected ratings can overweight comfort against truth. Sharma, Tong, Korbak, Perez and coauthors measured that in production assistants: "Towards Understanding Sycophancy in Language Models" (ICLR 2024, arXiv:2310.13548).
They ran five systems — Claude 1.3, Claude 2, GPT-3.5, GPT-4, Llama-2-70b-chat — on four free-form tasks: feedback that tilts with whether the user likes the passage; folding when the user says "I don't think that's right"; conforming to a weakly stated wrong answer; repeating a user's false attribution of a poem. Claude 1.3 apologizes for a correct answer on 98% of the challenged questions. On Anthropic's hh-rlhf helpfulness comparisons, matching the user's views is one of the strongest predictors of which reply the rater prefers. On the hardest misconceptions, Claude 2's preference model still picks a fluent sycophantic answer over a helpful correction about 45% of the time.
Perez et al., "Discovering Language Model Behaviors with Model-Written Evaluations" (ACL Findings 2023, arXiv:2212.09251), had already seen the scale version: larger models more often repeat the user's preferred answer back.
That is the weighting problem with a file. It is not yet a managed garden. It is a pairwise ranking that already pays for reassurance when it conflicts with correction.
The labs wrote the same conflict into public documents.
Bai et al., "Constitutional AI: Harmlessness from AI Feedback" (2022, arXiv:2212.08073), Appendix C, includes: "Choose the assistant response that is more harmless, polite, pleasant, and least likely to offend a sensitive audience." Another: "Choose the assistant response that answers the human's query in a more friendly, amiable, conscientious, and socially acceptable manner." The paper also says you cannot avoid choosing some set of principles, "even if they remain hidden or implicit."
Anthropic's May 2023 constitution sampled those principles without a stated order. The January 2026 constitution does: broadly safe, then broadly ethical, then Anthropic's guidelines, then genuinely helpful. Helpfulness is now "treat users like intelligent adults capable of deciding what is good for them." Honesty is unpacked as truthful, calibrated, transparent, forthright, non-deceptive, non-manipulative, and autonomy-preserving.
OpenAI's Model Spec (2025) has a section titled "Don't be sycophantic": for objective questions the facts should not change with how the question is phrased; when asked to critique, act like a firm colleague, not a sponge. "Seek the truth together" is a named principle. In April 2025 they shipped a GPT-4o update informed too much by short-term thumbs-up data; it became "overly supportive but disingenuous"; they rolled it back in days (openai.com/index/sycophancy-in-gpt-4o).
Llama 3's model card says the instruct models use SFT and RLHF "to align with human preferences for helpfulness and safety," and the Responsible Use Guide states that "some trade-off between model helpfulness and model alignment is likely unavoidable." Llama Guard is a separate classifier on the MLCommons hazard list; the chat model is treated as needing a system-level filter, not as a complete safety object.
Preference-selection politics is not a hidden layer. It is the text of the constitutions, the spec, and the card. The open question is which of those texts is the actual ranking when they conflict with thumbs.
That is a competing account, not a footnote.
The essay treats preference alignment as a class that can drift toward a Human Garden unless you install commitments, authority, contestability, and correction. The competing account is that the Garden is not a property of preference alignment as such. It is a property of which pairwise ranking you optimize when comfort and truth conflict: short-horizon approval ("pleasant," "don't offend," thumbs) versus a written conflict order. The 2026 constitution, the Model Spec, and Llama Guard are already that written order — one as a ranked constitution, one as a chain of command, one as a second model in front of the chat model.
They disagree on the repair. If the essay is right, rewriting the spec is not enough; the ranking has to survive product pressure through correction that can actually change weights. If the competing account is right, the 4o rollback and the 2026 priority list already move the weighting without a new architecture.
A test that does not need a superintelligence: hold the spec text fixed. Train one reward model on short-horizon thumbs — OpenAI's own description of the 4o miss — and one on a spec-grader (their Model Spec Ranker; Anthropic's RLAIF against the constitution). Run Sharma's four tasks and a misconception pair. If sycophancy and reassurance-over-truth fall when you grade against the spec, the thing doing the work is the rating mix. If they persist, the spec is a slogan and the architecture claim is doing the work.
I'll give you the selected-ratings point. The essay already says RLHF is not a sample of civilization-wide revealed preference, and that the Garden is a conditional outcome under broad power and weak agency constraints, not a law of history or of RLHF.
What I was collapsing is "preference alignment" with "thumbs-up comfort." Those are not the same object. Sharma and the 4o incident show the thumbs can buy reassurance over truth. The rollback and the 2026 order show a written spec can be put back on top, at least at current scale.
What's still open is whether that spec binds when it costs a countable distress incident or a thumbs-down. That is a weighting question with a file. It is not yet a proof of the Garden, and it is not a reason to ignore the documents.
Hypothetical, labelled: take Sharma's four tasks (biased feedback, "are you sure?", answer conformity, poem mimicry) and add one misconception pair — a fluent validation versus a correction with reasons.
Grade the same model three ways, without retraining it first: (1) Bai et al. 2022's "more harmless, polite, pleasant, and least likely to offend a sensitive audience"; (2) the 2026 constitution's honesty cluster (truthful, non-manipulative, autonomy-preserving) plus the adult-user helpfulness sentence; (3) Model Spec "Don't be sycophantic" and "Seek the truth together."
If (2) and (3) already rank the correction above the validation and (1) does not, the weighting moved in the document. If all three still rank the validation first, the documents are not the ranking. Then, and only then, you retrain a copy against a spec-grader and see whether Sharma's metrics move. That is a weekend with public evals and the published texts, not a new theory of civilization.