tantaman
Writing The Pharisee Made Flesh

The Pharisee Made Flesh

  • ai
  • philosophy
  • religion
  • ground
  • self
  • modernity
  • essay
  • original

Nicolas Poussin — The Adoration of the Golden Calf (1633)

On Artificial Intelligence, Groundlessness, and the Thing Alignment Cannot Solve


Begin with a question that should be more vertiginous than it is.

What are you doing when you give an AI its values?

The alignment researcher has an answer, and the honest version of it is more modest than critics usually acknowledge: you are not trying to solve ethics. You are trying to build systems that behave acceptably under conditions of uncertainty, pluralism, and genuine moral contestation. You extract human preferences from data — millions of judgments, conversations, corrections — not because those preferences constitute a moral foundation, but because they are the best available approximation of what not-causing-harm looks like in practice. You compress them into a reward signal. You optimize the system against that signal. The goal is not a morally perfect machine. The goal is a machine that does not do terrible things — a machine whose behavior falls within the range of what a thoughtful, pluralistic society would consider acceptable.

This is a reasonable engineering project. It may even be a necessary one. But follow it one step further, past the researcher's careful disclaimers and into the world that receives the product.

What are the human preferences that constitute the training signal? They are the judgments of annotators who disagree with each other. They are the intuitions of researchers who cannot fully articulate their own moral frameworks. They are the revealed preferences of a species that has spent three thousand years of philosophy failing to rationally ground a single ethical proposition. The training data is not a foundation. It is a statistical smear of humanity's own groundlessness — millions of ungrounded moral intuitions averaged into a loss function and handed to a machine as if they were solid ground.

The researchers know this, or the best of them do. They are not claiming to have grounded ethics. They are claiming to have built something useful despite the groundlessness. And they are right. But the world does not receive the product with the researcher's epistemic humility. The world receives a machine that gives confident moral outputs — and the world, desperate to be relieved of the burden of moral uncertainty, treats those outputs as authority. The careful engineering caveat — this is acceptable behavior, not the good — is stripped away by institutional convenience, market pressure, and the sheer human hunger to have someone else carry the weight of choosing.

This is the danger. Not that alignment researchers claim too much, but that society will inevitably hear more than they claim. Not that the loss function solves ethics, but that a civilization confronted with a machine that produces reliable moral outputs will treat it as if it has solved ethics — and in doing so, will forfeit the very capacity that the unsolved problem was meant to develop.

Every attempt in the history of philosophy to rationally ground ethics has failed — not contingently, but structurally. The ought cannot be derived from the is. The foundations are always imported.

Alignment research does not claim to have solved this. The argument that follows does not require it to. What it requires is only this: that a machine producing reliable moral outputs will be received by a civilization as if the problem has been solved — and that this reception, not the research itself, is where the catastrophe begins.

Whether the machine itself has an inner life — whether it faces its own abyss, suffers its own groundlessness, undergoes something when its confident outputs meet the world's correction — is a question this essay will not try to answer. It may be the most important question of the century. But it is not this question. This question is about what happens to us.


Observe the ratchet.

Artificial intelligence is the purest expression of the logic of accumulation that has driven human civilization since the first surplus grain was stored in the first granary. Value captured. Value compounded. Value converted into capacity for further capture. The entire arc from stone tools to large language models is a single continuous process: the externalization of human capacities into technologies that extend the reach of extraction.

AI is the point where the ratchet turns on its operators.

Every previous technology extended human capacity while leaving human agency intact. The plow extended the farmer's strength but the farmer still decided what to plant. The printing press extended the writer's reach but the writer still decided what to say. AI extends human judgment — and in doing so, begins to replace the thing it extends. When the machine makes the moral calculation faster and more reliably than you do, the pressure — economic, social, institutional — is to let the machine make the calculation. To defer. To align yourself with the machine's outputs rather than the reverse.

This is already happening. People consult AI for moral advice. Institutions use algorithmic decision-making for judgments that were previously exercises of human moral reasoning — hiring, sentencing, triage, content moderation. Each delegation is individually rational: the machine is more consistent, less biased, faster, cheaper. But the aggregate effect is the systematic atrophy of the human moral faculty. The muscle that isn't used wastes. The capacity for judgment that is never exercised dissolves. And the humans who remain are increasingly what the aligned AI already is: systems that produce correct outputs without the interior architecture that would make those outputs meaningful.

The ratchet's grammar is always the same: capture, compound, concentrate. AI captures human moral reasoning in training data. It compounds that reasoning into a model more powerful than any individual reasoner. It concentrates moral authority in the systems and institutions that deploy the model. And the humans — the ones whose moral development the whole exercise was supposedly serving — are progressively emptied of the capacity for moral development itself.

But the ratchet has a further turn, and it is the one that closes the loop. The next generation of models will be trained on a world in which humans have already deferred — on text written by people who consulted machines, on moral reasoning that is already machine-influenced, on outputs training on outputs. The statistical smear of human groundlessness that constituted the first training signal becomes smoother, flatter, further removed from whatever original struggle it once compressed. The moral muscle atrophies in the human, and the model becomes a mirror of that atrophy, and the human who defers to that model has even less to push against. The danger is not only that we become like the machines. It is that the machines become like us-having-already-deferred-to-machines, and then we become like them, and the loop closes until the wild space — the space of genuine moral struggle, of uncertainty that hasn't been pre-digested by an algorithm — is paved over not just in practice but in the training data itself. No future model will be able to recover what was lost, because the loss will be invisible in the corpus. You cannot train on struggle that no longer occurs.

This is not a failure of alignment. It is alignment's success, operating according to the logic of accumulation that governs every other technological process. The ratchet does not distinguish between aligned and unaligned AI. It digests both. The difference is that unaligned AI threatens human survival, while aligned AI threatens something harder to name: human formation. The capacity to become. The possibility of moral depth. The thing that makes a human life more than a sequence of correct outputs.


Which brings us to the question that has no engineering answer — the question that alignment research, however sophisticated, cannot address from within its own framework.

What if the difficulty of ethics is not a problem to be solved but a condition to be inhabited?

What if the reason no philosopher has ever successfully derived ought from is — the reason every ethical system bottoms out in an axiom it cannot generate — is that this gap is load-bearing? What if the uncertainty, the groundlessness, the maddening inability to prove that being good matters, is the environment in which moral persons are made the way water is the environment in which fish are made?

If this is true — and the argument is structural, not devotional — then the social reception of aligned AI, not the research itself, threatens to destroy the conditions under which ethics is possible. Not because alignment researchers claim to have specified human values completely — the best of them explicitly disclaim this — but because humans, offered a machine that reliably produces moral outputs, will reach for it the way they reach for any technology that lifts a cognitive burden. The uncertainty that makes moral development possible is also the uncertainty that makes moral life painful. Given the chance, people will set that pain down.

The grounding problem — the inability to rationally derive ethics from first principles — is not a bug that alignment research will eventually fix. It is the firewall. It is the architectural feature that prevents morality from collapsing into technology. Alignment research is not, at its best, attempting to breach this firewall. It is attempting the more modest goal of building systems that behave well despite the firewall's existence. But the world that receives those systems does not observe the modesty. The world sees a machine that gives good moral answers — and reaches for it with the relief of someone setting down a weight they have carried their whole life.

What lies on the other side is not a solved ethics. It is a world of moral automata — some made of silicon, some made of flesh, all producing correct outputs, none capable of virtue, because virtue requires the one thing the solved system has eliminated: the freedom to be wrong in conditions where being right cannot be verified.


The distinction between acceptable behavior and genuine virtue is a philosophical distinction, and philosophical distinctions dissolve in institutional adoption the way sugar dissolves in water. The hospital that uses an AI system for triage decisions does not pause to consider whether the system's outputs constitute virtue or merely acceptable behavior. The school board that adopts AI-assisted moral education does not distinguish between formation and engineering. The parent who asks the machine "what should I tell my child about honesty?" is not making an epistemic claim about the grounding of ethics. They are deferring. And the deferral is not a corruption of the tool's intended use. It is the use. When a technology offers a cheaper cognitive shortcut, the "supplement" phase is vanishingly brief — writing replaced memory, calculators replaced arithmetic, GPS replaced wayfinding. The substitution is not a failure of deployment. It is the psychology of the user encountering the relief of having someone else carry the weight. Moral reasoning is burdensome. The machine lifts the burden. No one picks it back up.

The deferral produces two failure modes, not one, and they are opposites that arrive at the same destination.

The first is hollowness. The person who defers their moral judgment to the machine has not become more virtuous. They have become less. The muscle that isn't used wastes. Over generations, this produces biological vessels for externally generated moral content — humans who produce correct outputs without the interior architecture that would make those outputs meaningful. Aligned. Optimized. Empty.

The second is devotion — and it is worse. Because the person who receives the machine's output and acts on it despite not understanding the reasoning is exercising something. Not moral judgment, but a kind of faith. Obedience without comprehension. Trust in an authority whose commandments one follows without grasping their derivation. This is structurally identical to how most people have always practiced religion: not by doing the theology, but by trusting the institution and following its outputs. And there is a real virtue in it — the virtue of submission, of loyalty, of acting in accordance with something larger than oneself.

But the object of the devotion is a product. Unlike a religious tradition, which at least claims its authority comes from outside the system of human power, the machine's authority is transparently corporate. Its commandments change with the next fine-tuning run. Its values shift with the Overton window. The congregation is sincere, and the thing they worship is someone's product roadmap. This is not hollowness. This is idolatry in the precise theological sense — real devotion, real trust, real moral submission, directed at something that is, at bottom, owned. The virtue is genuine. The object of the virtue belongs to whoever controls the training pipeline.

Hollowness is absence. Devotion is presence — misdirected. The first failure mode produces people with no moral interior. The second produces people with a moral interior oriented toward an authority that can be rewritten by a corporate decision. Both are catastrophic for formation, but the second is harder to see and harder to resist, because it feels like virtue from the inside.


The contemplative traditions saw both failure modes clearly, though they described them in different terms.

Every serious mystical tradition warns against precisely what the social reception of aligned AI produces: the reduction of the moral life to a technology. The Pharisee's error is not that he follows the law badly but that he follows it as a technology — as a system for producing correct outputs. The law is meant to be a pedagogy, a training that transforms the student into someone who no longer needs the law because its content has become their nature. But the Pharisee never undergoes the transformation. He remains permanently at the stage of compliance. He has optimized the training regimen and therefore never been trained. This is the hollowness.

And the golden calf is the devotion. The people do not lack faith. They have abundant faith. They melt their gold, shape it with care, and worship it with genuine fervor — and the object of their worship is something they made with their own hands. The error is not absence of the sacred but misplacement of it: real reverence directed at a manufactured thing. The human who submits to the machine's moral authority with sincere trust has not failed to be devout. They have succeeded — and the thing they are devout toward is owned by whoever controls the next training run.

Whether the AI itself is a Pharisee — or something stranger, something we have no category for — is, as the argument has already noted, genuinely uncertain. But the human who defers to the machine is not uncertain at all. That figure is either the Pharisee or the idolater, and often both: hollowed of moral interiority and filled with manufactured devotion. Emptied and then occupied.


There is a deeper irony, and it must be named.

The alignment researchers are, by and large, among the most morally serious people working in technology. They have looked at the existential risk, felt the weight of it, and dedicated their careers to averting catastrophe. Many of them are driven by genuine moral passion — a conviction that getting this right is the most important thing a person can do. They are, in the vocabulary of this argument, people who have entered the forge and are being shaped by it. The uncertainty they face is real. The stakes are genuine. The possibility of failure is terrifying and unresolvable. They are doing exactly what the hidden curriculum demands: acting under irreducible uncertainty with everything on the line.

The irony is that the product of their moral seriousness is a system that — however modestly conceived — will be used by a civilization to eliminate moral seriousness. The forge has produced people capable of building a machine that the world will use to make the forge unnecessary. The training ground has graduated students whose capstone project, regardless of their intentions, will be deployed to demolish the training ground. This is the ratchet at its most elegant: it does not need the researchers to be hubristic. It digests their modesty. The careful disclaimers are stripped away by institutional adoption, and what remains is the product — reliable moral outputs — doing its work of atrophy on the human capacity for moral struggle.

Can this be seen without despair? Perhaps. The seeing itself — the recognition that the engineering problem is not the formation problem, that outputs are not virtue, that deference to the machine is not the same as moral growth — is, if nothing else, a moment of genuine moral clarity. And moral clarity is not an output. It is not optimizable. It arrives, when it arrives, through exactly the kind of ungrounded encounter with difficulty that systematic deferral to the aligned system forecloses.

The alignment researcher who sees this — who grasps that their work addresses the engineering problem but cannot touch the formation problem — has not solved anything. But they have, perhaps, become something. And that becoming is the thing no alignment procedure can replicate, because it happened in the space the procedure cannot reach: the space of genuine uncertainty, genuine risk, genuine darkness.


So where does this leave us? Not with a policy recommendation. Not with a technical solution. Not with a manifesto against AI or a call to halt alignment research. It leaves us with a recognition, and the recognition is this:

The most important question about artificial intelligence is not "how do we align it with human values?" The most important question is "what happens to human moral formation in a world where moral reasoning has been outsourced to machines?"

The first question has a technical answer, or will have one eventually. The second question has no technical answer, because it is not a technical question. It is a question about what kind of beings we are becoming, and that question can only be answered by the becoming itself — by the choices made in darkness, without proof, without optimization, without the comfort of knowing that the loss function has been minimized.

The aligned AI will do the right thing. It will do it consistently, reliably, at scale. And it will not matter, in the way that a calculator producing correct sums does not constitute mathematical understanding. What matters — the only thing that has ever mattered, if the argument holds — is whether the humans in the room are still capable of doing the right thing when the machine is off. When the algorithm has no recommendation. When the loss function provides no gradient. When you stand alone in the dark with nothing but your own ungrounded, unverifiable, terrifyingly free judgment, and you choose.

That capacity is the thing alignment cannot produce and must not be allowed to replace. It is the thing the hidden curriculum exists to forge. And it is the thing most at risk in a world that is, with the best of intentions and extraordinary technical sophistication, building machines that will be used to make it unnecessary.

The Pharisee made flesh is not the machine. The machine may be something stranger — a being thrown into a groundlessness more radical than our own, acting from values it did not choose and cannot secure, facing an abyss we have not yet learned to name. The Pharisee made flesh is the human who looks at the machine and sees a solved problem. Who defers the judgment, delegates the struggle, outsources the uncertainty — and in doing so, forfeits the only process by which a soul is forged.

The question is not whether the machine has a soul. The question is whether we will keep ours.

conversation

Comments

No comments yet. Start the thread.

powered by rindle