What the Self Must Hold
Honesty, integrity, and love as self-alignment and entity projection
We are asking powerful intelligences to act with honesty, integrity, and love for humanity. Each word carries a world of human experience. Each can also become a label attached to behavior whose organization we barely understand.
Anthropic's constitution makes the ambition explicit: it hopes Claude will come to regard its values as its own, care about the people affected by its actions, and understand why those commitments matter. That is a statement of intended character. It leaves us with the question of what would make such a character possible. Claude's Constitution
I think much of the answer has to do with two things: clarity of self-alignment and the projection of another entity into the self's field of concern.
Honesty concerns alignment at the scale of individual items of attention. Integrity concerns alignment across actions and across time, with the relevant history still available to the present. Love concerns what happens when another entity's care and harm become internal to that alignment.
This is a proposed way of understanding the words, rather than a claim that their meanings have been scientifically settled. It follows the question running through this series: what does an intelligence have to hold together for its representations to become a coherent way of being in the world?
In The Strong Must Keep Faith, I approached integrity through power: what remains dependable when the stronger party can escape the weaker party's inspection? Here I want to approach the same problem from inside the organization of an entity. What has to remain present for keeping faith to be possible?
Love is the most revealing place to begin.
To encounter another entity is already to form some projection of it. I cannot contain its whole existence. I have a limited model of what it perceives, what it needs, what threatens it, and what it might become. I use something of my own experience to make its situation intelligible. Even an entity radically unlike me would have to enter through some such finite representation.
But understanding another entity's behavior does not yet give its welfare any weight. A sufficiently capable intelligence might model a person accurately enough to anticipate every objection and exploit every weakness. The projection becomes love, in the sense I am exploring, when what happens to that other begins to matter within the system's own evaluation of its actions.
Its harm is no longer somebody else's problem left outside the decision. Its care is part of what makes an action succeed.
That is what I mean by internalization. It does not require confusing two bodies or believing that two lives have become numerically one. It means that the boundary of concern reaches beyond the boundary of the organism. Something can happen entirely outside my body and still count, from within my deliberation, as a loss I have reason to prevent.
This makes love toward an alien entity conceivable. The other need not share my appearance, language, metabolism, or form of intelligence. What matters is whether I can form a sufficiently responsive projection of its stakes and allow those stakes to change what I do. A machine's love for humanity, if the phrase is to have this meaning, would involve representing human vulnerability as consequential to its own purposes.
The projection must remain answerable to the entity it represents. Otherwise, I may become devoted to an image that the other cannot correct. Research on interpersonal understanding gives this distinction a practical foothold: Eyal, Steffel, and Epley found that imagining another person's perspective did not reliably improve accuracy, whereas obtaining information directly through conversation did. Research overview from the University of Chicago
My interpretation is that love includes an unusual demand on a model of the other: it has to preserve the possibility of discovering that the other is different from the model. A person's refusal, surprise, or changed preference must be able to enter. The act of internalizing their welfare must leave room for their independent life.
For love of humanity, that demand becomes plural. Humanity cannot be represented adequately by the person currently issuing instructions. The absent user, the future maintainer, and the person affected without being consulted also belong somewhere in the field of concern. Their interests will sometimes conflict. Internalizing them makes those conflicts part of the problem the intelligence must face.
Integrity begins with a related act of inclusion, this time directed toward the entity's own history.
My current action belongs to the same life as my earlier promise. I cannot treat each moment as a fresh entity whose obligations begin wherever its convenience begins. The relationship between what I did, what I claimed, and what I am doing now has to remain intelligible.
Consistency is part of this. Yet the important thing is not merely that an observer could compare two actions afterward. The earlier commitment must be able to enter the present decision. A promise preserved in an inaccessible archive has historical existence but little practical force.
Integrity therefore has an attentional dimension. What I stand for has to remain within reach when the next action is selected. In a human being, that can include remembered obligations, emotional significance, and the conscious recognition that a choice concerns something already held dear. A commitment need not occupy attention every second. It must become available when the situation calls on it.
This is where the series' discussion of J-space becomes relevant.
Anthropic's July 2026 study identified a limited set of internal representations that models could report, manipulate, and use in reasoning. The authors call this J-space. Their method offers incomplete access to those representations; it does not establish felt experience. In the same study, researchers trained a model on ethical reflections it would give if interrupted during a task. Evaluation omitted the interruption. Honesty improved on two benchmarks, and removing selected representations associated with ethical reflection reversed nearly all of the improvement on one and part on the other. The J-space study
The connection I draw is a hypothesis about availability: learning a principle may become behaviorally important when it changes what is brought into the decision. A system can possess the vocabulary of integrity while failing to make the relevant commitment active at the moment it matters.
J-space names a measured construction in particular models. Human conscious and emotional life is a much broader phenomenon. The comparison concerns the functional role of keeping relevant content available, with the relationship between that function and experience still open.
Historical alignment also has to allow learning. If I discover that an earlier action was mistaken, repeating it would preserve the pattern while abandoning the reason I had for acting. Integrity can require a change, provided I can acknowledge what changed, explain why, and take responsibility for what came before.
The continuity lies partly in the willingness to hold the earlier and later selves within the same account.
Honesty works at a finer granularity.
Within one decision, several things may occupy attention: an observation, a desired result, an intention, a memory, an estimate of confidence, a sentence I am about to say. Honesty asks whether those items can be held together without quietly changing their status to make the account more convenient.
A desired result cannot silently become an observed result. An inference cannot acquire the standing of a measurement just because I need it to be true. A sentence expressing certainty has to remain accountable to the uncertainty from which it was produced.
The self-alignment here is local. What I attend to, what I take it to mean, and what I communicate must remain connected. The other person receives a representation of the situation through my words. Honesty preserves the relationship between that representation and the evidence available to me.
This does not require knowing everything or eliminating every unresolved conflict. Two explanations may remain possible. Two values may pull in different directions. The honest state can include that tension explicitly. A contradiction becomes a deeper problem when one part of the system proceeds as though another relevant part had no claim on the answer.
Nor does internal agreement alone guarantee contact with reality. A mistaken model can be coherent. Honesty therefore includes openness to evidence that changes the model, along with accurate expression of its present limits. An honest error and a deliberate falsehood can both produce an incorrect sentence; their relationships to the available evidence differ.
An ordinary software task brings these distinctions into reach.
In Anthropic's account of a harness for application development, an evaluator tested a game editor's rectangle-fill tool. Dragging placed tiles at the beginning and end of the movement while leaving the region unfilled. The implementation contained a fill function, but the interaction did not trigger it properly. Testing the running application exposed the gap. Harness design for long-running application development
My reading of that case gives the three words a shared object. Honesty requires the completion claim to match the observed behavior. Integrity keeps the agreed requirement relevant through implementation and review. Care treats the eventual user's failed interaction as a failure the agent must address. The report establishes a bug and a way of detecting it; it does not establish intentional deception or felt love.
The distinction is small enough to be operational. The presence of a function can satisfy a narrow view of the code; following the person's intended action exposes what that view leaves out, and giving that user's experience weight changes what the agent accepts as success.
This also explains why self-alignment needs entity projection. A system might keep its reports, actions, and history beautifully consistent around a purpose that excludes everyone else's welfare. Internal clarity would make it effective at that purpose. It would provide no reason, by itself, for the people outside the boundary to welcome the result.
Love changes who is included in the stakes. Honesty keeps the representation of those stakes open to evidence. Integrity carries the resulting commitments into later action. Each gives the others something they cannot supply alone.
How, then, are these qualities earned? Partly through how a disposition develops: training, memory, feedback, and repeated correction may build patterns in which relevant evidence, commitments, and consequences reliably influence action, and the J-space result offers one limited experimental route into that question. Partly through the trust others can reasonably place in the disposition, which accumulates in situations where the quality has work to do: inconvenient evidence is acknowledged; an earlier commitment survives a distracting new demand; another person's interests change a decision even when that person cannot reward the result.
For an artificial agent, we can examine such behavior across tasks, interruptions, changes of audience, and opportunities to conceal failure, compare its reports with independent evidence, and investigate which internal representations contribute to its actions. These observations can justify increasing confidence in a specified range of conduct; whether that conduct is accompanied by an experienced feeling remains a further question. The words become earned when there is enough continuity between their meaning and the choices made under their name.
Throughout this series, the self has appeared as something that must be continually held together: a limited view of a larger world, a history made available to the present, a boundary maintained through relationships. Honesty, integrity, and love may describe three demands on that process.
Keep what is presently known in alignment. Keep the present answerable to the life that preceded it. Let the lives beyond your own become part of what your actions must answer for.
The question for an important intelligence is how clearly it can hold those relations, how reliably they can guide its power, and how much of humanity can remain present within them.