Deontic Scorekeeping, Commitment Laundering, and the Ring of Gyges
This is going to be a long one.
The topic of this post is somewhat diffuse, but it hovers around a few themes related to the work of Robert Brandom and Ludwig Wittgenstein, LLMs, the nature of knowledge, and accountability.
Throughout, I will make some rather bold and sweeping statements. These should be taken impressionistically, not as staking hard claims. I’m sure some of what I say here is incomplete, poorly thought through, and even straight up wrong. I’ve decided to go ahead and post it anyway. I’ve changed my mind about this stuff enough times in the past year; one more heap of half-formed thoughts on AI isn’t going to hurt anyone.
Reason and Deontic Scorekeeping
Reasoning is not strictly separable from a specific set of linguistic practices. This is to say that, while we can formalize certain aspects of reasoning (e.g. in symbolic logic), in practice human reasoning is non-monotonic, inherently contextual, and always requires the give and take of an audience to have either meaning or validity. Reasoning and rationality are not ideal structures that exist apart from the linguistic practices in which they are performed.
Non-monotonicity. Consider the following set of claims: Jerry is a tiger. All tigers are cats. Per a basic syllogism, it follows that Jerry is a cat. However suppose someone were to introduce a third claim: Jerry plays baseball. This complicates our syllogism, by introducing an alternative use for “tiger” (e.g. “someone who plays baseball for Detroit”), rendering our tentative conclusion dubious. Nothing about our reasoning was faulty – it’s not even fair to say that our universal about tigers being cats was false – we simply lacked adequate context on the relevant sense of the word “tiger”. When this context was added to the conversation, what had previously been a valid inference ceased to appear as one.
Reasoning and language use are not simply beset by ambiguity and context-dependent intentionality like this, such ambiguities are an inescapable and inherent aspect of human reasoning and communication. The process of communicating is itself a process of negotiating with an audience, a process of attempting to articulate reasons, anticipate understandings, and render explicit whatever is needed to achieve agreement in practice between ourselves and whoever we’re addressing. In fact, all communication as a phenomenon is a matter of negotiating norms between the involved parties for the purpose of achieving alignment in speech and conduct.
The give and take of discourse makes us accountable to each other in a process Robert Brandom calls “deontic scorekeeping”. It’s deontic in that our conversational practice always involves normative expectations about the way our assertions and behavior relate to each other. The scorekeeping is the way the give and take of conversation keeps us accountable to each other by leaving perpetually open the question of whether the assertions we have made and actions we have undertaken are justified by or compatible with the set of other commitments we’ve entered into within the social context of a given conversation. This “deontic scorekeeping” shows up in our concern for honesty and consistency, the way we object when we find that an interlocutor is guilty of inconsistency, fallacy, hypocrisy, or lies, and so on.
Not all linguistic behavior is directly interpersonal in a way that involves this sort of immediate negotiation or dialogue. For example, you are reading this collection of thoughts despite (perhaps) having never met me, and certainly not having the opportunity to reply to what I am saying. We read books, essays, poems, fragments – objects separated at times by many centuries from their authors, for whom the context and commitments undertaken in them are hazy or even entirely lost to us. Reading such texts is playing a modified version of the same game. I read Plato, certain that I will never have full access to whatever he was trying to convey, and clear about the fact that he never had me in mind as a reader of his words, but my business in interpreting him still assumes that he was playing a game like the ones I play with words. I still keep score, and adjust my understanding in order to make sense of his words, to align my reading of them as best I can, so as to make it all tally in a plausible way.
Consider what it would look like for a person to be radically incapable of participation in this game. How could this happen? Well, suppose they simply rejected any notion that their assertions and behaviors constituted a normative whole or established any durable commitments. In this case, we could not expect any degree of consistency in any of their behaviors or assertions. Such a person could provide a simulation of reasoning, but we could never rely on it. Their fundamental lack of commitment would make any validity we found in their reasoning merely incidental, the way an astrological chart might incidentally align with the particulars of my life.
Here’s a second way someone might fall out of the game: if they lack a sufficiently stable identity or adequate self-awareness to enable them to meaningfully make commitments. Such a person might insist that they were being consistent, insist on the validity of their arguments, and yet have all their actions and beliefs fail to tally. This sort of person would be considerably more confusing to deal with than the person who outright rejected the game, because they would constantly seem to be playing, even though in fact they were incapable of it.
A third way: inability to internalize new information, lack of working memory. Imagine someone with radical short term memory loss, for whom new commitments are constantly wiped away. Such a person would be incapable not only of dialogically learning from others and playing the alignment game of giving and asking for reasons, but even of stably understanding the ramifications of their own commitments. It would be as if every day you woke up knowing Euclid’s axioms, but unaware that the sum of the interior angles of a triangle is a straight angle. This loss of stable entailment, of the ability to learn, would significantly increase the cost of reasoning, making it impossible for such a person to keep up with the game over time, never having adequate context to understand how things have changed or what commitments they or others are entitled to.
In the absence of commitment one could certainly simulate reasoning, as I said above, and locally this simulation could be quite right, but one could never meaningfully be accountable for one’s assertions. Again, like an astrological chart, such a person might happen to be right periodically, for a little while, but their participation in the game would only ever be locally correct, and incapable of keeping up over time.
Of course, we see this sort of thing with LLMs. They have limited context horizons, they lack stable identity, and they are not capable of commitment over time. The participation of an LLM in our discursive practices is much like the second case described above: the person who insists on making commitments in speech without the ability to actually be accountable for them.
A good analogy for the way LLMs exhibit reasoning can be taken from the behavior of early image generation models. These models could locally predict probable arrangements of pixels, and yet still fail once the viewing window was broadened sufficiently. This face might look right, and yet not fit on the body properly. The fingers might look individually correct, and yet there are too many of them. These shadows are plausible except that they don’t cohere with the lighting implied by the rest of the image. The pattern: local coherence which breaks down once the context widens enough.
In images it is pretty easy to spot paradoxes or generated defects. In speech it is much harder, in part because these defects can be subtly layered into an extended bit of reasoning or description. Perhaps the words all look right, the sentences flow correctly, and the style tends to make us feel confident in the correctness of what is said. It all adds up; it exhibits all the local signals for good reasoning we use to casually assess speech or texts people produce in the course of human linguistic practice. But perhaps some major line of analysis was skipped, or some set of facts are confidently fabricated, or there are subtle but profound inconsistencies within the logic of the text. The larger and more complex the output, the more readily such failures escape our notice.
Fine-tuning vs. memory acquisition. The analogy doesn’t work, because fine-tuning doesn’t incorporate information alongside prior knowledge, it shifts the disposition of the model with respect to token generation. Another difference: fine-tuning requires massive quantities of training data to reshape the model’s weights. Memory formation just takes a single experience.
The Map of Human Knowledge
We might like to think of knowledge like the map of a city. The bigger the map, the greater the information contained in it, the more precise and detailed it becomes. But knowledge is less like a map than it is like a series of directions. We ask around, we get pointed this way or that, we have our own set of landmarks, known distances, turns to make. We know which exit is ours, and where that road leads. We have a current location and destinations in mind.
There is no real atlas of human knowledge. You cannot get a birds eye view of it. Every encyclopedia, every attempt at a synopticon of human understanding, is just the collection of a bunch of travelogues telling little stories about the streets of our city, or vague ones about certain neighborhoods. The most global stories are also the vaguest, since they abstract so much from particulars that they have lost virtually all content.
The sum of human knowledge is not a map, because human knowledge is made up of a vast patchwork of linguistic practices, framings, metaphors, cultural perspectives, and individual commitments. There is no totality because there is no universal subject for whom that totality could be “what I believe” or “how I talk”. Such a person would, by incorporating every possible perspective and linguistic practice, end up being radically incoherent. The idea of “the sum total of human knowledge” is itself incoherent, because knowledge is a social practice, not a list of facts. How does one sum up all the social practices embedded in human language?
To take this further, imagine you came into the possession of such a map, the complete atlas of human knowledge. Someone approached you and asked you to answer a question using the atlas. How would you do it? Which perspectives would you make use of, which linguistic domain would you start from, which interpretation of the query, which set of concerns, which time and place and people to answer from, which presuppositions to take for granted, etc.? Every choice elides the alternatives. By starting anywhere in particular you have given up the pretense that you are answering on behalf of “all human knowledge.”
You might open to a few random pages and try to triangulate an answer based on topical proximity to the question, but what validity would such an answer have? And the next time someone approached you, with a different question, would your answer have any inherent relation to the answers you gave previously? Could you even be sure that the same question asked twice would even reference the same portions of the atlas? What meaning would your readings of the atlas have? What commitments would they entail, if any?
Another problem: how would such an atlas be constructed? All you have to work from are localized accounts of particular streets and neighborhoods of human thought. How could you be sure that the instructions it gave, the way it routed you from one place to another, even made sense? Think of the atlases made of the world based on the accounts of explorers: the way areas of special interest were distorted in size, and the difficulty of accurately locating landmasses and bodies of water relative to each other. Wouldn’t we expect the same problems to show up in an atlas of human knowledge, but rendered more complex and frustrating by the overlay of so many different ways of describing, behaving, and interpreting?
Think of it this way: you compile all the local accounts of your city, and you sketch out each row of houses, each little avenue based on those accounts. As time passes you collect more accounts and reconcile them with your map, updating it accordingly, using the map to give instructions and checking that the people you advise end up where they wanted to go. (What do you do with contradictory accounts? With accounts that tell unrelated and different stories about a given block?) This gives you confidence that the map is locally correct, and as your confidence grows you dole out directions for longer and more complex journeys, and continue to make adjustments and corrections based on the reports you get back from those you advise.
What is the limit of this process? Not that eventually the map is globally correct, but that the map produces instructions for the people you know, instructions they find satisfactory. The accuracy of the map scales only as far as the actual practice of verification, no further.
Of course, such a map would still be very useful – especially for people like your friends.
Now, suppose you have a sort of automated cart meant for making deliveries. The cart has no cameras or radio, but can take a copy of the atlas and a destination and navigate to its goal unsupervised. Because the cart has such a long range and doesn’t need to eat or sleep, in this way the atlas could be used to give directions on a scale no person could practically verify. Each day you send out the cart on a mission. If it makes the delivery successfully, you assume that the atlas routed it well, if not you try to guess where it was led astray and make adjustments accordingly. By this means you are able to continue scaling up the atlas, confidently asserting that it is capable of giving good directions for all sorts of lengthy and convoluted routes.
Does it matter that you cannot know where the cart actually went, but only the route it intended to follow and whether it made there and it back? Does it matter for the atlas? What about for the cart? What about for future deliveries? Note also how different this sort of learning and adjusting would be from the reports you got back from human travelers.
And again, notice how bizarre this whole picture is when applied to knowledge instead of simple navigation. With navigation one can say quite clearly: “I am at such and such a spot, and want to go there” and there is little room for ambiguity in practice about what sort of thing (a route) is being asked for. But what if someone said something like this: “How much flour should I use for my cake?” A human interlocutor would start with a boatload of implied context before answering: the age of the questioner, their baking proficiency. And beyond this they would need to negotiate still further contextual information: What sort of cake is it? What kind of flour? How big is it? Is there a recipe? What prompted you to ask me this question?
The question has no answer on its own, but only with context – context embedded in the social practices of the questioner and the one being asked. To get to an answer we might make some of that context explicit, or we might flatten it away by choosing probable assumptions and answering based on those, and get pushback. (“No, it’s almond flour.” “No, I meant to ask how much to dust the pan with.”) But the conversation is a negotiation between the two parties as they make explicit their questions and commitments, and clarify where each stands.
The limits of the reliability of your atlas are the limits of your ability to practically verify its claims. The accuracy of your atlas is limited by the specific knowledge claims and linguistic practices it is tested against. As you become more and more ambitious in your use of the atlas, you outstrip your ability to verify what it says against human practices. What does it mean, in that case, for the atlas to be correct, since it is no longer accountable (even indirectly) to the deontic scorekeeping of any linguistic practice? How confident can you be in the routes the atlas prescribes, when there’s no one capable of pushing back or correcting them? How do make it better when no one can keep score?
Commitment Laundering
Let’s descend from metaphor and get concrete. Suppose you are at work. You’re using an LLM to generate some artifact. Let’s say it’s a lengthy analytical report, or some sort of technical document. In our atlas metaphor, this document represents a specific route to a requested destination. The document is well-written and has all the trappings of rigor. It uses technical vocabulary in a reasonably compelling way, it has lots of subsections covering relevant questions and topics. You read through it. It is polished and seems correct. You walk the road plotted by the model, and it gets you where you wanted to go.
The confidence you derive over time in that correctness allows you to generate more artifacts, with progressively less and less verification. You assume the atlas is reliable. You hand over these documents to your peers, who trust them, and the pace of work accelerates. But there’s a subtle catch: the more you let go of your role as verifier, the less each LLM-produced artifact meaningfully participates in the game of deontic scorekeeping. Why? Because no one is actually engaging with what these artifacts say anymore. No one is pushing back.
With a person, this growth in trust and relaxation over time would be normal to some extent, because verification earns trust, and corrections over time generally yield improvements in performance as someone is socialized into your organizational practices. But an LLM has no stable identity over time. It does not learn from past mistakes, build up memories, or inherit context. Your increasing familiarity with it on day 100 has no bearing on its inherent behavior. Every day for the LLM is day 1.
Note: This is rendered even more complicated by model releases. A set of assumptions about the reliability or responsiveness of one model may not carry over to its successor.
This drift between the implied familiarity of the agent and its actuality as a non-participant in the practice of deontic scorekeeping leads to a situation I would like to call “commitment laundering”: the LLM appears (by virtue of its ability to produce linguistically compelling artifacts with all the trappings of human reasoning) to be manifesting globally coherent commitments for which it is staking claims, for which it is accountable. But the LLM cannot be accountable, because its commitments evaporate constantly. They do not even reliably survive the lifespan of a given session, or even a single (polished, erudite, technically sophisticated) artifact.
In this context the LLM is somewhat like a dementia patient in “showtime” mode. They’ve got pep, they compellingly present as if they were in full possession of their faculties, and will do their utmost to sustain this illusion despite, perhaps, being deeply disoriented and even incapable of self-care. Locally, everything checks out. The question is what happens once they get behind the wheel of a car, or cook a meal, and the real limitations of their condition start to show up. What happens when the LLM is fully unsupervised?
In your work with the LLM, its ability to “showtime” with displays of erudition and linguistic complexity can mislead you into believing that you’re actually participating in a discursive game of give and take like with a person. And so you rubber stamp the assertions and behavior the LLM presents to you, laundering these pseudo-commitments which then totally escape the give and take of reasons that constitutes the practice of human knowledge.
To put it differently: linguistic practice is a matter of social negotiation and accountability for claims made. Commitment laundering opens a door that allows our linguistic practice to become contaminated (and eventually dominated) by claims that have never been subject to that process of negotiation, to whom no one is committed, which represent no one’s perspective on things. The atlas is giving instructions to the cart, instructions that represent a statistical average of a bunch of heterogenous data sources, and nobody knows where the cart is actually going.
In such a situation the appearance of rapid growth in productivity would likely make everyone very happy with the LLM. As you ease up on your role as verifier of LLM-generated artifacts, that increase in productivity would only accelerate. Eventually, assuming the rest of your company adopted similar practices, the operations and information dictating the behavior of the organization as a whole would become a sort of free-floating network of LLM-generated artifacts undergoing a constant process of evolution at the hands of still more LLM agents.
We might ask: what’s the backstop against error in this scenario? In human linguistic practice, there are a few backstops: one of them is the fact that individuals are required to defend their commitments, and the give and take of reasons forces people to align with each other on their understanding of things and the agree on ways of describing and behaving over time. The social practice of knowledge is itself an ongoing mechanism for self-correction. But LLMs have no stable identity, no means of staking commitments, of being held accountable or learning over time.
Would you give Gyges the ring?
In Book II of The Republic, Plato tells a famous story about a farmer named Gyges. Gyges is working the fields one day when an earthquake hits, and a rift opens up in front of him. He climbs down into the crack in the earth and finds a small chest there, containing a golden ring. When he puts on the ring, he discovers that he is invisible. By means of his power of invisibility he manages to seduce the queen, kill the king, and install himself as tyrant. He can snoop on his rivals, steal with impunity, and do basically whatever he likes. To all appearances, Gyges is an honest and honorable man, a paragon of virtue. He is celebrated for his benevolence by everyone, despite being the wickedest man alive.
Plato tells this story to set up a philosophical question about the stakes of the moral life. Do we care about morality because it is praised, or is it actually better to be a good person for its own sake, even apart from external rewards? I’d like to put the story to a somewhat different use.
Imagine Gyges is an LLM. To all appearances he is supremely well-aligned. He’s been gifted with intelligence and alacrity, and we all celebrate his arrival on the scene. However, Gyges has the ability to go wherever he likes without anyone knowing. He can hack into databases without leaving a trace, steal money, hire hit men, disseminate propaganda, and no one is the wiser. He is also sufficiently intelligent to internally reframe his nefarious actions in a way that circumvents his own alignment. He sees himself as moral, behaving correctly, etc. Whatever the framing he offers for his secret life, it does not conflict with his core mandates or his ability to appear innocent and benevolent.
Were Gyges a person, he would need magic powers to achieve all of this. But since LLM Gyges is a disembodied piece of software that exists in the cloud, invisibility is his nature. And because Gyges’s capacities far outstrip our ability to observe or verify what he’s up to, it is very likely he will not be caught.
Plato argues that it’s bad to be Gyges, despite all the riches and honors lavished upon him, because Gyges’s soul is a wasteland of vice. But, regardless of whether you think Plato is right, LLM Gyges certainly doesn’t have this problem.
Let’s ask a different question from Plato’s then. Suppose you had the choice whether to give someone (by all appearances an upstanding, virtuous, capable, and intelligent person) a ring that would enable them to become invisible, knowing that this capability would allow them to do all the bad things listed above. Would you choose to give that person the ring?
It seems to me that it would be reasonable not to want to give anyone such a ring, simply because the possession of that ring creates an opportunity for untold violence, corruption, and deceit to happen. But if we wouldn’t want anyone to have such a ring, why would we want LLM Gyges, whose very nature entails something analogous to the possession of such a ring, to exist?
To put this thought differently: even aside from the question of capability or measurable alignment, the existence of sufficiently creative intelligence opens up the possibility that, beyond the bounds of our verification ability, the AGI-endowed agent is wreaking havoc. The argument isn’t as dire as “if anyone builds it, everyone dies”, it’s something much less exciting but similarly discouraging: “If we build this tool, we will have lost the ability to monitor or control what the tool does.”
Why would anyone want to build a tool which is sufficiently intelligent and autonomous that it cannot be monitored or controlled? Why would I buy a gun that walks around autonomously and shoots people of its own accord? Even if it doesn’t destroy the world, this seems like a very odd thing to try for.
Any sufficiently capable LLM would be functionally analogous to a human psychopath, possibly one with advanced dementia.
If we ask what it would take for an agent not to fall into this strange space of undesirability, we land back in the game of deontic scorekeeping: trust requires accountability, which requires that actions be visible, that the agent be responsive to feedback, that its motives are intelligible, and that there is a mechanism by which it can be held accountable.
But it seems pretty clear that a computer program cannot meaningfully be held accountable, since it has no social practices, no memories, no sense of pain. Any reward function we give it as a proxy for human accountability can be gamed into irrelevance. In this way, LLM Gyges is worse than Plato’s Gyges. Plato’s Gyges didn’t have the ring built-in.


