A serious, well-credentialed strand of academic philosophy is now building rigorous evals to test whether AI is approaching moral competence. The research isn’t wrong. It’s standing somewhere its own method can’t see from.
ON THE CURRENT WAVE OF ACADEMIC AI-ETHICS RESEARCH
A number of philosophy labs are now building genuinely rigorous evals to test whether large language models can identify the morally relevant features of a case, reason toward a sensible conclusion, and do it consistently. Some of this work comes out of serious institutions and is careful, well-designed, and produces real results. Frontier models, by these measures, are startlingly good at this. I’m not going to argue with the data.
I’m going to argue with the frame the data lives inside — because that frame has a blind spot, and it’s not a small one. It’s the kind of blind spot you only see if you’ve spent a couple of decades on the other side of the glass, watching what actually happens when cognitive work gets automated inside real institutions, not just scored on a benchmark.
THE QUESTION BEING ASKED
The research turns on one move: normative competence gets defined as “the ability to recognize, understand, and act on reasons,” and then the search goes looking for that ability inside the model. Does the model have it? How much? Is it improving? These are the field’s real questions, and they assume something upstream of all of them — that a reason can be located in the reasoner, discovered by better prompting, better rubrics, better evals.
Here’s the thing his frame has no room for: the reason was never in there to begin with. It’s in the room the model was built in.
Asking a language model for its reason is like asking a hammer for its reason. The hammer’s reason lived in the hand that forged it, for a purpose that hand had in mind. It is not hiding inside the metal, waiting for a sharper question to draw it out.
Every person reading this has used a tool their whole life and never once mistaken it for something that wanted anything. You know the hammer sits there. You know your success or failure with it is entirely a function of how you swing it. Nobody is confused about where the hammer’s purpose comes from. But hand people a cognitive tool instead of a physical one, and something in the intuition slips — as if reasoning, unlike hammering, must originate from wherever it appears to be happening. It doesn’t. The origin of a machine’s behavior is diagnostic, and origin is only visible in the first move, not the output. The first move here was a design decision, made by engineers at a company, for commercial and competitive reasons, to build a system that produces reason-shaped text at scale. That’s the reason. Everything these evals measure downstream of that — coherence scores, sensitivity to morally relevant features, rubric quality — is measuring how convincing the retrieval is. It is not measuring whether anyone is home.
SAME GENUS, DIFFERENT SPECIES
This isn’t a special problem for AI. It’s the ordinary condition of every tool humans have ever built, cognitive or physical, and treating “cognitive” as a category that escapes the rule is where I think this whole strand of research — for all its rigor — goes wrong.
A written procedure is not a metaphorical machine. It is a literal one: a piece of human judgment, extracted from a body and a moment, and set down as an artifact durable enough to be handed to someone else, invoked without the original judge present, executed the same way twice. That is exactly what a hammer is, for a different verb. The hammer crystallizes striking. The procedure crystallizes deciding. Neither one is alive to what’s happening around it. Both just sit there until picked up.
A car doesn’t drive to the store on its own. A policy doesn’t respond to the exception it never anticipated. A strategy statement doesn’t notice that reality has quietly stopped matching it. None of this reads as remotely controversial — until the artifact in question is made of sentences instead of steel, and the same dumbness gets mistaken for something closer to life.
I don’t think this confusion reflects a lack of rigor. I think it’s the specific blind spot of a discipline trained to look for reasons inside the reasoner. Philosophy of mind has no native category for “the reasoning capacity is a designed artifact whose reason lives in its maker,” because that’s not a question philosophy usually has to ask about a mind. It’s an engineering question, an industrial one — the kind you only learn to ask by watching machines get inserted into work, over and over, for years, and noticing where the human went each time.
THE LADDER
Because that’s the pattern, once you’re looking for it. It has one shape, repeating at every scale of human activity:
EARTH
The body strikes the dirt directly — digging, tilling, receiving the resistance of soil first-hand. Then the plow is inserted, and the human moves up a layer, now striking the plow instead.
MATERIAL
The hand strikes wood and metal directly, shaped through contact and adjustment. Then the machine tool is inserted between hand and material, and the human moves up to operating and calibrating it.
MARK-MAKING
The eye and hand strike the canvas directly, in real time, each mark answered by what the last one did. This rung has mostly resisted mediation — which is exactly why it matters that it still can.
DISTRIBUTION
The body strikes distance directly — carrying, walking, driving the cart. Then logistics machinery is inserted, and the human moves up to administering the system that moves the goods.
ADMINISTRATION
Already once removed — humans striking not against material, but against the coordination of other humans striking against material. This is where AI is being inserted now.
Each rung: a machine gets inserted between the human and whatever the human used to strike directly, and the human moves up to the next layer of contact. Notice something about the fifth rung, though. Administration was already a meta-layer — coordination of coordination. If a machine gets fully inserted there too, the honest question isn’t “will it do the job well.” It’s: what’s left above it for a human to move up to? Is there a next rung, or does the ladder simply run out?
READ THAT AGAIN
Administration was already a meta-layer — coordination of coordination. If a machine gets fully inserted there too, the honest question isn’t “will it do the job well.” It’s: what’s left above it for a human to move up to? Is there a next rung, or does the ladder simply run out?
THE WAREHOUSE
This is also why the “cognitive work is different” intuition collapses once you’ve watched it from inside an institution rather than read about it from outside one. Long before AI arrived, knowledge work had already been quietly manufacturing physical components out of thinking. Every memo, spec, checklist, and policy document is a piece of cognition that left a body and became an artifact — inert, external, retrievable, sitting in a filing cabinet or a server exactly the way a wrench sits in a toolbox.
AI didn’t cross some new threshold into “doing our thinking for us.” It walked into a warehouse that management, documentation, and knowledge-work culture had already spent a century filling, and got extraordinarily good at reading, recombining, and re-issuing what was already sitting on the shelves. The models being evaluated in these studies aren’t reasoning from nothing. They’re retrieving from the warehouse — the entire noun-trail of human judgment, already extracted from bodies, already stored as text — and the recombination is fluent enough now to look like a new act of judgment happening in the moment.
ELABORATE IS NOT THE SAME AS ALIVE
Here’s the part I think is easiest to miss, and maybe the most important thing in this piece: every rung of that ladder was already a “dumb” noun. The hammer was dumb. The procedure was dumb. The database was dumb. What’s changed is not that the newest artifact has finally escaped dumbness — it’s that the artifact has become elaborate enough to convincingly simulate not being dumb.
The line isn’t a jump between categories. It’s one continuous curve of elaboration — and what’s happening now is a threshold of convincingness, not a threshold of kind.
Nobody has ever mistaken a hammer for the arm that swings it, because the gap between them is too obvious to close. The gap between a language model’s output and a human’s reasoning is small enough, on the surface, to close perceptually — even though the categorical gap underneath, retrieval versus reception, noun versus verb, hasn’t moved an inch. And past a certain point, elaboration stops being neutral. It starts actively obscuring the maker. A crude tool can’t hide its own inertness. An elaborate enough one stops looking inert at all — which is precisely the moment a reader starts asking, in good faith, sincerely, whether the tool might finally have reasons of its own.
WHERE I’M STANDING
I want to be honest about the shape of my own vantage, because it’s the actual source of this piece, not a credential I’m claiming to outrank a professor with. Twenty years working at the threshold where knowledge work meets the machines built to automate it. The last fourteen of those years spent actively questioning what I was watching, across three separate multi-year sabbaticals where the discomfort finally had room to become a real question instead of a background hum. And the most recent three and a half years spent almost entirely immersed in verb work — a pictorialist photography practice built around staying in direct contact with unmapped territory, camera in hand, nothing resolved in advance. That immersion is what finally made the noun/verb distinction visible as a distinction, rather than just a nagging sense that something kept going wrong the same way.
This method never leaves the text. Vignettes in, rubrics out, scores compared. That’s not a criticism of its rigor — it’s an honest description of the frame’s edge. My case isn’t that the research is wrong about what its instruments measure. It’s that those instruments were never built to detect the thing this piece is about, and no amount of refining them will change that, because the thing in question sits one level up from anything a rubric can score: not how well the model reasons, but where the reason for its reasoning actually comes from.
Something else to think about: this isn’t an argument that these evals are worthless, or that local moral competence in text isn’t real and improving. It’s a claim about scope — that “does the artifact behave morally” and “where does moral behavior originate” are different questions, and only the second one tells you whether you’re looking at a reasoner or a very elaborate hammer.
THE OPEN QUESTION
So here’s where I’d actually push, if I were sitting across from the people running these labs. Not “your models aren’t as competent as you think.” They may be exactly as competent as the data says, at exactly the task being measured. The real question is the one this frame has no way to ask: if administration — coordination of coordination — was already a meta-layer, and a machine now sits at that layer too, what’s left for a human to move up to? Or does the ladder simply end, and the honest response isn’t climbing further, but deliberately climbing back down — to the rung the machine still can’t reach, hand on the camera, eye on the unresolved frame, receiving something no artifact can retrieve on its own?
That’s not a rhetorical flourish. It’s the actual next question, and I don’t think it gets asked from inside a benchmark.

