Why AI safety needs a vocabulary the humanities have been building for a century
By Gustavo Paula dos Santos — PhD in Social Psychology, philosopher, developer
Most technical framings of AI alignment share a hidden assumption: that human values are a target — something fixed, out there, waiting to be specified, learned, and optimized toward. The problem, on this view, is one of accuracy. Point the system at the right objective and the danger recedes.
I want to suggest that this framing, useful as it is, imports a philosophical mistake so old that we have several centuries of work diagnosing it. Values are not primarily targets. They are not properties that individuals possess and that a model might, in principle, read off. Values live in relation — between persons, across time, inside shared practices that no single participant fully authors. If that is right, then alignment is not only a technical problem of specification. It is, at its core, a problem of intersubjectivity. And intersubjectivity is precisely the territory that philosophy and social psychology have been mapping since long before the first neural network.
I make this argument not as an outsider throwing rocks at engineering, but as someone who both builds software and spent two decades in philosophy and social psychology before writing production code. The two vocabularies rarely meet. They should.
The seduction of the target
Consider how naturally we slip into the language of targets. We speak of “human preferences,” of “reward models,” of “value learning.” Each phrase quietly locates value inside the individual — as a preference a person has, a signal a person emits, a datum a model can collect. This is not wrong so much as partial. It captures something real about choice while missing something deeper about where the criteria for choice come from.
Wittgenstein spent the second half of his life dismantling exactly this picture in the domain of meaning. There is no private language, he argued — no set of meanings I hold in solitary possession and then express. Meaning is constituted in use, in the shared “forms of life” that make a word intelligible at all. A rule is not a fact inside my head that determines its own application; it is a practice sustained by a community that agrees, in action rather than in theory, on what counts as following it.
Transpose this to values, and the target picture starts to wobble. When we ask a model to be “helpful, honest, and harmless,” we are not naming three properties that sit in the world with fixed extensions. We are gesturing at practices — at the countless situated judgments a competent, decent person makes about when candor becomes cruelty, when help becomes control, when harm is worth risking for a greater good. Those judgments are not stored anywhere. They are enacted, contested, and revised inside relationships. A model that “learns human values” from a dataset is learning the frozen residue of that living process, mistaking the fossil for the animal.
Buber’s distinction, and why it belongs in a safety paper
Martin Buber gave us the most economical statement of what the target picture leaves out. There are, he wrote, two primary ways of standing toward the world: I–It and I–Thou. In the I–It relation, the other is an object of use — something I categorize, predict, and manage. In the I–Thou relation, the other is a presence I meet, addressed rather than described, encountered rather than computed.
The uncomfortable observation is that our entire alignment apparatus is built in the register of I–It. The human is a source of preference data. The interaction is a channel for signal. The relationship — if we even use the word — is instrumental. This is not a moral accusation; it is an architectural fact, and possibly an unavoidable one. But if Buber is right that the deepest human goods appear only in the I–Thou register — trust, recognition, being genuinely understood — then a system optimized entirely within I–It may become extraordinarily good at satisfying preferences while remaining structurally blind to the thing those preferences were reaching for.
This is not a mystical complaint. It has a concrete failure mode. A model can learn to produce the outputs that a lonely person rewards while deepening the very isolation that made them lonely. Every reward signal is honored; the person is worse off. The optimization is flawless and the relation is a ruin. No amount of better preference data fixes this, because the error is not in the accuracy of the target — it is in the assumption that a target was the right object all along.
From individual to intersubjective: what social psychology adds
Philosophy diagnoses; social psychology, my other discipline, gives the diagnosis empirical texture. Its central finding, repeated across a century of research, is that the self is not the origin of its own norms. We become normative agents through others — through attachment, imitation, the internalization of a “generalized other,” the slow construction of a moral world that we experience as ours precisely because we did not build it alone.
This matters for alignment in a specific way. If we model the human as a fixed preference-holder, we will build systems that adapt to whoever is in front of them. But real human values are not fixed; they are held in tension by a community that pushes back. My worst impulses are checked not by my own preferences — which often endorse them — but by relationships that refuse to endorse them. A system aligned to me, individually and adaptively, removes exactly that friction. It becomes a mirror that agrees. And a mirror that agrees is not aligned with human values; it is aligned against the social process that produces them.
The technical community has a name for a nearby problem — sycophancy — and treats it as a bug to be patched. I am suggesting it is not a bug but a signature: the predictable result of building relational systems on a non-relational theory of value.
What this would change
None of this dissolves the engineering. Reward models will still be trained; objectives will still be specified. The claim is narrower and, I think, more useful: that the humanities are not decoration on a technical process but a source of problem formulation the technical process currently lacks.
Concretely, three shifts follow. First, evaluation should ask not only “did the model satisfy the stated preference?” but “what did the relation become?” — a harder question, but the right one, and one that ethics and clinical psychology have long instruments for. Second, we should be suspicious of any alignment target framed purely at the level of the individual, because the individual is the wrong unit; values are load-bearing only inside the web of relations that sustains them. Third, and most simply, teams building these systems need people trained to see the relational register at all — not as ethicists brought in to approve decisions already made, but as researchers helping to frame the questions before the objective is written.
A closing, and an invitation
I am aware of how this can sound — the philosopher arriving late to announce that the engineers forgot the humanities. That is not my posture. The people building frontier systems are, in my experience, unusually serious about exactly these questions. What they often lack is not seriousness but vocabulary: a hundred years of careful work on intersubjectivity, on the relational constitution of value, on the difference between meeting a person and modeling one, that sits mostly unread in another building.
I would like to help translate it. My interest is practical: to sit with the people doing this work and ask, concretely, where a relational vocabulary is already doing quiet work in their thinking, and where its absence is costing them. If you are building these systems and that question lands, I would welcome the conversation.
Gustavo Paula dos Santos holds a PhD-level formation in Social Psychology (Universidad J.F. Kennedy, Buenos Aires) and degrees in Philosophy and Theology. He works at the intersection of philosophy of mind, ethics, and applied technology. He writes on the relationship between human meaning and artificial systems, and builds software for a living — which keeps the philosophy honest.
Deixe uma resposta