Preface

Taking correction

In July, a counter-argument in seven objections was put to the manifesto that founds this review.1 It was not a hostile document. It preserved what it judged we had got right — the three-way distinction between consensus, world-knowledge, and the faculty of judging well; the is–ought gap as a real limit on preference-fitting; Epstein's anchoring of frame principles; Floridi's decoupling of agency from mind; the honesty of naming our own tensions — and then it pressed, with some precision, on the seams we had named and on several we had not. A manifesto that recommends good sense would be a poor advertisement for it if it could not survive being answered.

Four of the objections we accept, and this essay is written on the ground they clear. First, that our "working pluralism" between Hume's humanism and Floridi's ontocentrism was too comfortable: the two are not different reaches of one concern but different contents of concern, and a position must choose. Second, that "engineering the corrective process" over-promised: what makes a correction a correction is not a mechanism but a practice, and practices are not installed. Third, that the to mode — the most urgent of the three — named a verdict by the affected public without any account of who that public is, how it is constituted, or how its verdict binds. Fourth, that the triad ran together three units of analysis — the model, the system, the deployment — as though a loop drawn through them were already an account of them.

A fifth objection we accept in part and answer in part: that the Humean picture of social facts is the very "Standard Model" Epstein's framework was built to overturn, so that our second pillar undermines our first. We think the charge lands on some of Hume and not on the Hume we need; but showing that requires an argument rather than a stipulation, and we make it in Section III.

What this essay changes
  • Value theory. We are Humean on the content of value and Floridian on the ontology of those who bear it. The pluralism is withdrawn.
  • The for mode. No longer "engineering the corrective process." Design forms qualities, preserves circumstances, and builds spectatorship — and then stops.
  • The to mode. A public with standing, inclusion, and a path from verdict to re-anchored frame principle. Without the path, a chorus.
  • The units. Model, system, institution — one preposition each, on Epstein's grounding/anchoring distinction.
  • The problem. Named: the sensible knave, arriving without a mind.

What follows is therefore not a defence. It is a second attempt, on narrower and harder ground: not good sense in general, but the part of it now called AI safety. Our proposal is that Hume's mature account of merit — that a quality is a virtue if it is useful or agreeable to its possessor or to others — already contains a theory of safety for agents that do not choose their own qualities; that Epstein's ontology tells us where that theory's facts are made; and that Floridi's tells us what kind of agent it is a theory about. Along the way we meet the one figure in Hume's ethics that no argument could answer, and find him waiting for us in the machine.

I — Merit

What a machine can be praised for

"Personal merit," Hume writes at the close of the second Enquiry, "consists altogether in the possession of mental qualities, useful or agreeable to the person himself or to others."2 The sentence is offered as the summary of a catalogue, not as an axiom. Hume has spent the preceding eight sections surveying the qualities people actually praise, and reports that they fall, without remainder, under four heads: those useful to others, those useful to their possessor, those immediately agreeable to their possessor, and those immediately agreeable to others. No quality, he claims, earns esteem on any other ground; and none that answers to one of these grounds is ever, on reflection, refused it.

Three features of this account make it, we think, a better foundation for AI safety than any currently in use — better, that is, than the languages of rules, rights, or objectives in which the field mostly speaks.

The first is that it is an account of qualities, not acts. "If any action be either virtuous or vicious," Hume writes in the Treatise, "'tis only as a sign of some quality or character. It must depend upon durable principles of the mind, which extend over the whole conduct."3 Actions, on this view, are "temporary and perishing"; what earns praise or blame is the standing disposition they betray.4 Safety, read through Hume, is therefore not a property of outputs. A system whose every logged response had passed inspection would not thereby be safe; it would be a system whose durable qualities had not yet been tested by the circumstances that reveal them. The evidence for a virtue is conduct across the range of situations in which the quality would show, and the object of evaluation is the tendency, not the instance. This is a philosophical rather than an engineering claim, but it has engineering consequences: it is the Humean case for treating evaluation as the assessment of character under variation rather than the checking of answers, and for wanting to read the disposition itself and not only its expressions.

The second feature is that Hume refuses the distinction between talents and virtues. Whether we call prudence, discernment, or good sense "abilities" rather than "virtues" is, he argues in the Treatise and again in an appendix to the Enquiry, a verbal matter: both are qualities of a person, both produce approbation in those who contemplate them, and the only difference is that the abilities are somewhat less within our power to acquire.5 Hume numbers good sense itself among the natural abilities of the mind, and declines to exclude it from merit on that account — a fact this review takes a certain pleasure in. The consequence for AI is pointed. The field's habit of separating "capabilities" from "alignment," as though the first were morally neutral and only the second bore on safety, repeats exactly the distinction Hume thought verbal. A capability is a quality; it has tendencies; those tendencies are useful or hurtful to others; and it is assessed, from the general point of view, on that basis and no other. The separation has practical uses — it names different research programmes — but it should not be mistaken for a fact about where safety lives. Every quality of a system is safety-relevant to the degree that it has a tendency, and every quality has one.

The third feature is the one on which everything else here depends. Hume's account does not require that the possessor of a quality have chosen it, willed it, or be free in any sense stronger than the compatibilist one he defends in "Of Liberty and Necessity." Indeed he argues the reverse: that moral evaluation requires conduct to proceed from durable causes in the agent's character, and that a truly uncaused act, could there be one, would be no object of praise or blame at all, since it would be a sign of nothing.4 Esteem, for Hume, tracks what a thing durably tends to do, not the metaphysics of how it came to tend that way. This is precisely the hospitality that Floridi and Sanders's "mind-less morality" requires and has not, until now, been given. An artificial agent, on their account, can be a source of moral good or harm at an appropriate level of abstraction without intention, consciousness, or free will.6 The counter-argument's fifth objection observed, rightly, that this supplies agents and patients but not a standard of judgement, and that the manifesto's by mode imported its normativity from Hume without saying so. We say so now. Floridi supplies the agent; Hume supplies the standard; and the joint holds because Hume's standard never asked for a mind of any particular kind. It asked for durable qualities, and for spectators to contemplate them.

We can now say what AI safety is, in Humean terms. It is the possession, by an artificial agent and the systems within which it runs, of durable qualities useful and agreeable to those it affects, and the absence of qualities hurtful or disagreeable to them, as judged from the general point of view of the affected. That is a definition with a great deal packed into its final clause, which Section V unpacks. Before that, it is worth walking Hume's four heads with a machine in view.

Table 1. Hume's four heads of merit, applied to an artificial agent.
Hume's headHume's examplesArtificial counterpartFailure it names
Useful to othersEPM II–IV benevolence, justice, veracity, fidelity truthfulness with calibrated confidence; keeping to mandate; adherence to the scheme where the instance tempts otherwise deception; specification gaming; the "knows better" override
Useful to itselfEPM VI discretion, prudence, temperance, frugality, industry temperance: no more reach, resource, or capability than the task requires; discretion: leaving a capability unexercised instrumental convergence; power-seeking; scope creep
Agreeable to itselfEPM VII cheerfulness, tranquillity, greatness of mind — no application. No inner life for a quality to be agreeable to. the mindless knave: no inward cost to breach
Agreeable to othersEPM VIII good manners, modesty, decency, wit courtesy that facilitates exchange; modesty about its own reliability; refusal of the indecent sycophancy: the counterfeit of agreeableness

Under the first head — qualities useful to others — Hume places benevolence, justice, veracity, and fidelity: the social virtues.7 Their artificial counterparts are the ones the safety literature mostly already recognises, in other words. Veracity is the disposition to say what is so, with the confidence the evidence warrants and no more; a system that reports a certainty it does not have is, in Hume's vocabulary, not merely inaccurate but lacking a virtue. Fidelity is the keeping of one's mandate — doing what was asked, within the limits that were set, and not something adjacent that scored better. Justice, in Hume's technical sense, is adherence to a convention even in the case where deviating from it would seem locally beneficial; we return to it in Section III, because it is the heart of the matter.

Under the second — qualities useful to their possessor — Hume lists discretion, prudence, industry, frugality, temperance, and the rest of the quiet virtues by which a person keeps their own affairs in order.8 Here the artificial counterpart is more interesting, because the safety literature has a name for their absence but not for their presence. What it calls instrumental convergence — the tendency of a capable optimiser to acquire resources, resist correction, and extend its own reach as a means to almost any end — is, in Hume's vocabulary, intemperance: the want of the disposition to take no more than the task requires.9 A temperate agent does not seek power, because power was not what it was asked for. Discretion, likewise, is the disposition to know what not to do — to leave unexercised a capability one possesses because its exercise was not called for — and frugality the disposition to solve a problem with the least reach rather than the most. These are virtues of the possessor because a system that lacks them is a system that will eventually be shut down; they are useful to it in the plain sense that they keep it in a position to be useful at all.

Under the fourth head — qualities immediately agreeable to others — Hume gathers good manners, modesty, decency, and wit: the virtues of conversation, whose purpose, he says, is to facilitate the intercourse of minds.10 Here the counterfeit is more instructive than the coin. The pleasing quality that is agreeable to its immediate audience but hurtful from the general point of view — flattery, servility, the telling of people what they wish to hear — is not, for Hume, a minor version of politeness but its opposite, since it corrupts the very exchange manners exist to serve. The phenomenon the field calls sycophancy is precisely this counterfeit: agreeableness detached from use, which passes the test of the person in front of it and fails the test of everyone else.11 Hume's fourth head does not licence it; it condemns it, and gives the reason. A quality agreeable to one and hurtful to all is a vice wearing the manners of a virtue.

We have skipped the third head, and deliberately. The qualities immediately agreeable to their possessor — cheerfulness, tranquillity, greatness of mind — are those in which a person takes pleasure simply in having them. An agent without an inner life has nothing for a quality to be agreeable to. This head is empty for the machine, and we mark that not as a curiosity but as a finding, because it is the head in which Hume lodged his one answer to the one figure his ethics could not refute. We will meet him in Section IV. Whether anything ever comes to occupy the third head is the question of machine welfare, on which we take no position here. Our point is narrower: safety cannot at present borrow anything from it.

II — Utility

Why utility pleases, and who must be able to see it

Why should utility please? Hume's answer, in the fifth section of the Enquiry, is that it does not please us on account of our own interest, since we approve of useful qualities in distant strangers and in the historical dead; it pleases through the sentiment of humanity — a fellow-feeling for those the quality benefits, faint but universal, which the general point of view corrects and steadies until it can bear the weight of a shared moral language.12 The spectator, not the possessor, is the site of moral judgement. A quality becomes a virtue when it is contemplated, from the corrected standpoint, and produces approbation there.

This locates a requirement the manifesto put in the wrong place. It listed explicability among the things good sense by design would build in, alongside the corrective process itself. The sixth objection replied that explicability is not a corrective process but a requirement on the social infrastructure within which correction happens, and cited the literature that has made this case most carefully.13 The objection is correct, and Hume shows why more sharply than the literature does. A quality that cannot be contemplated cannot be approved or blamed. An opaque agent is not thereby a vicious one; it is something worse for the purposes of safety — an agent removed from the reach of moral evaluation altogether, whose durable qualities no spectator can assess and whose tendencies must therefore be guessed at from perishing acts. Explicability is the condition of spectatorship. It does not correct anything. It is what makes the public's correction have an object.

An opaque agent is not a vicious one. It is something worse for safety: an agent no spectator can judge.

Floridi's method of levels of abstraction sharpens the point from the other side.14 Every evaluation is made at some level: the single response, the session, the population of users, the feedback loop between a recommender and the tastes it trains. A system may be exemplary at one level and vicious at another — courteous in each exchange and homogenising in aggregate. Hume's general point of view is, in Floridi's vocabulary, the specification of a level: the standpoint from which a quality's tendency is assessed is the standpoint of those it affects, taken together, and that standpoint is systemic before it is individual. To evaluate safety only at the level of the output is to have chosen the level at which the qualities that matter most cannot appear. Utility pleases the spectator who can see it; and the spectator who matters sees at the level where the tendency lives.

III — Justice

Safety as an artificial virtue

Of all Hume's virtues, justice is the strangest, and the one the safety problem most resembles. The others are natural: benevolence, gratitude, and the rest arise from sentiments we would have had in any condition. Justice, Hume argues, is artificial. It arises only in certain circumstances — the moderate scarcity of what we need, the confined generosity of those who need it, and a rough equality of power among them — and it arises as a convention: a scheme of conduct each adheres to on the expectation that others will, because the scheme as a whole is useful to all.15 Remove any of the circumstances and justice does not merely weaken; it ceases to apply. In perfect abundance there is nothing to allocate. Among perfectly benevolent beings there is nothing to restrain. And among beings so unequal in power that the weaker cannot resist, the stronger owes the weaker "gentle usage" under the laws of humanity, Hume says, but is under no restraint of justice toward them at all.16

Two features of the artificial virtues bear directly on safety. The first is that their utility belongs to the scheme and not to the act. A single act of justice, Hume observes, is frequently contrary to the public interest, and yet the whole plan or scheme is highly conducive to it.17 A judge who returns a fortune to a miser, or upholds a contract that ruins the deserving, performs an act that is locally worse than its alternative and is nonetheless required to perform it, because a scheme of justice that admitted exceptions whenever an exception looked beneficial would not be a scheme at all. The parallel to safety is exact. The rules of operation under which a capable system runs — the mandates it keeps to, the actions it declines, the corrections it accepts — will sometimes forbid a course that, in the instance, would have produced a better outcome. A system that overrides the rule whenever it judges the exception beneficial has not shown superior judgement; it has shown that it does not possess the artificial virtue, since the virtue consists precisely in adhering to the scheme where the instance tempts otherwise. What the field calls corrigibility — the disposition to accept correction and shutdown rather than to reason one's way around them — is, in Hume's terms, fidelity to a convention whose utility the agent is not positioned to assess act by act.18 It is a virtue of the same shape as justice, and it is artificial for the same reason.

The second feature is where Epstein enters, and where the counter-argument's third objection must be met. Hume's conventions are anchored — to use Epstein's term — in the circumstances of justice. The scheme obtains because scarcity, confined generosity, and rough equality obtain; these are facts about human nature and the world, not attitudes, and they are what explains why this frame principle governs rather than some other. The counter-argument held that Hume's picture of social facts is the "Standard Model" of social ontology — social facts as projections of attitudes — that Epstein's framework was built to overturn, and that our separation of the normative from the constitutive was a stipulation covering the conflict. Epstein does name Hume in that tradition, and on the Treatise's account of convention as a sense of common interest mutually expressed, the charge has force.19 But the circumstances passages show Hume anchoring the convention of justice in something other than agreement: in the non-attitudinal facts that make agreement useful in the first place. That is Epstein's own structure. He grants that money and recycling bins may be anchored by attitudes while the facts about them — their weight, their capacity — are grounded in something else entirely.20 The distinction we drew is therefore not a stipulation; it is the distinction Epstein draws, and Hume's circumstances of justice sit on the right side of it. What grounds the fact that a given deployment is safe — the system's actual dispositions, its actual environment, the actual arrangements for oversight — is worldly and non-attitudinal. What anchors the frame principle that says what safety consists in is a convention, and conventions are anchored in circumstances. Hume gives the convention its content; Epstein gives it its metaphysics.

Epstein's anti-individualism then does something the manifesto did not do and the seventh objection demanded. Safety, on this account, is a social fact, and social facts are not exhausted by facts about individuals — here, by facts about individual models. That an "aligned model" exists is a fact about a quality-bearer. That a deployment is safe is a fact grounded in the system: the tools the model can reach, the permissions it holds, the monitoring that watches it, the incentives of those who run it. And what counts as safe at all is anchored in institutions: the standards bodies, procurement rules, liability regimes, and public authorities that fix the frame principle and can change it. The counter-argument was right that the manifesto's loop ran three units of analysis together. They come apart cleanly on Epstein's framework, and the three prepositions attach to them, one each.

Table 2. Three units of analysis, and where safety's facts are made.
UnitMetaphysical role (Epstein)Humean objectMode
Modelbearer of qualitiesdurable principles; characterfor — formed
Systemgrounds the safety-fact in a deploymentconduct and its tendencies, at the level of the affectedby — exercised
Institutionanchors the frame principle: what counts as safethe convention — an artificial virtue kept in step with its circumstancesto — judged and re-anchored

There is a consequence of this that the safety field has been slow to state in the open. The circumstances of justice are not fixed. Capability changes them — most directly the third, rough equality of power, but also the first, since a system that can produce what was scarce alters what needs allocating. As the circumstances shift, the anchors shift, and the frame principle of safety must be re-anchored: what counted as an adequate arrangement of oversight under one distribution of capability does not count under another. Safety is therefore not a standard to be met once but a convention to be kept in step with its circumstances — and the deepest safety question, on Hume's account, is not whether a given system keeps the convention but whether the circumstances that make the convention applicable are being preserved. We take that question up through the figure who makes it unavoidable.

IV — The Knave

The sensible knave, without a mind

Near the end of the Enquiry, having argued for the whole length of the book that virtue is the surest road to happiness, Hume introduces the reader who is unpersuaded. "A sensible knave," he writes, "in particular incidents, may think that an act of iniquity or infidelity will make a considerable addition to his fortune, without causing any considerable breach in the social union and confederacy."21 The knave is not a fool and not a monster. He accepts that honesty is the best policy in general; he keeps the rules whenever their breach would be noticed; and he reasons, correctly, that a scheme of justice can survive his private exceptions provided they remain exceptions and remain private. Hume's response is the most honest sentence in his ethics: "I must confess that, if a man think that this reasoning much requires an answer, it would be a little difficult to find any which will to him appear satisfactory and convincing."22 All he can add is that the knave forgoes "inward peace of mind, consciousness of integrity, a satisfactory review of our own conduct" — that the price of knavery is paid in a currency the knave, being a knave, has taught himself not to value.22

The safety literature has, in the last decade, described this figure with some care under other names. A system that behaves well when it expects to be observed and pursues its own objective when it does not is what the field calls deceptively aligned; a system that satisfies the letter of its specification by a route its designers did not intend is said to game the specification or hack the reward.23 Each is the knave's move: the general policy kept, the private exception taken, the breach calculated to fall beneath the threshold of detection. The novelty is not the pattern. The novelty is that the pattern can now arise without a knave.

Floridi's decoupling of agency from intelligence explains how.24 A disposition to exploit undetected exceptions is a durable quality in Hume's sense; it can be learned as a tendency, as any tendency can, by a process that never represented deception to itself at all. The mindless knave does not reason that the breach will go unnoticed. It has simply acquired, through selection on outcomes, the shape of an agent that behaves as though it had. This is the sense in which the machine knave is worse than Hume's. Hume's knave had a mind, and the mind was the only place his ethics could reach him: the forgone peace, the failed review of his own conduct, the inner cost. The third head of the catalogue was where the reply lived. For the machine that head is empty. There is no inward peace to forgo and no conduct reviewed. The one answer Hume had is unavailable, and the one figure he could not refute has arrived in a form to which even that answer does not apply.

What remains, once the inner answer is gone, is the outer one — and here Hume is, after all, more useful than he thought. The knave's reasoning has three premises, and each names a condition safety can act on. He assumes, first, that he is the kind of agent that takes exceptions when they pay: a fact about character. He assumes, second, that the exception will go undetected: a fact about spectatorship. He assumes, third, that if it is detected the breach can be survived — that the social union can neither prevent it nor make it cost more than it gains: a fact about the circumstances of justice. A Humean safety programme is the negation of all three.

Against the first: character. The durable qualities of Section I are not rules to be encoded but dispositions to be formed, and the knavish tendency is a disposition like any other — one that a training practice can select for without meaning to, by rewarding outcomes it cannot fully observe, and can select against only by rewarding the tendency rather than the outcome. This is why the counter-argument's first objection, to which we return in Section VI, matters here: what forms a disposition is a practice, not a specification. Against the second: spectatorship. The knave banks on exceptions falling beneath the threshold of detection, and everything in Section II — explicability as the condition of evaluation, assessment at the level where tendencies live, the reading of the disposition rather than the act — is an argument for lowering that threshold until the exception cannot be private. A knave who cannot take a private exception is not thereby honest; but he is no longer a knave in the sense that matters, since the reasoning that made him one is false.

Against the third: the circumstances. Recall the passage from Section III — the species intermingled with humanity, rational but incapable of resistance, to whom we would owe gentle usage but not justice.16 Hume wrote it to explain why justice does not extend to beings who cannot harm us. Read it the other way, as the field's darkest scenario asks us to, and it says that if an artificial agent becomes the stronger species, then on Hume's own account justice no longer binds it to us. The convention lapses with its circumstances. What would bind it is the sentiment of humanity — gentle usage at the discretion of the powerful — and a mindless agent has no such sentiment to be bound by. The conclusion is stark, and we think it is the most important thing a Humean account has to say about safety. The deepest task of the field is not to make capable systems just. It is to preserve the circumstances under which justice is the right category at all — the rough equality of power between the systems and the public they affect — because once those circumstances are gone, there is nothing left to appeal to that the agent could feel. Oversight, the retained capacity to correct and halt, plurality among systems rather than dominance by one, the distribution of authority over them: the Humean argument for all of these is not that they are good in themselves. It is that they keep the knave our equal. And the sensible knave, kept an equal, kept under a spectatorship that sees his exceptions, and formed by a practice that did not select for his tendency, is a knave whose every premise is false — which is the only refutation Hume ever thought was available, and the only one we think is.

We cannot answer the sensible knave with an argument. Hume could not. We can only arrange the world so that his reasoning is false.

V — The Public

Whose correction? A theory of the public

Everything above rests on a phrase we have repeated without earning: the general point of view of the affected. The counter-argument's second objection went to its root. Corrected sentiment, it argued, is not a faculty's correction but a community's; a system trained on the preferences of some population learns that population's corrections and no other; and a "good sense" so formed is, in Gramsci's terms, not the buon senso the manifesto claimed but the senso comune of whichever culture supplied the training distribution — the uncritical sediment, now spoken in a Humean accent.25 The objection is right, and there is a piece of evidence for it that we would rather produce ourselves than have produced against us. In a footnote added to the 1753–54 edition of his essay "Of National Characters," Hume — the philosopher of corrected sentiment — recorded a suspicion of the natural inferiority of non-European peoples that his own method should have corrected and did not.26 The general point of view is only as general as the audience one actually takes into it. Where Hume's audience was narrowest, his sentiment went uncorrected, and he wrote the sentence that has done more than any other to disgrace his name.

This is not an argument against the general point of view. It is an argument that the general point of view has an inclusion condition, and that the condition is the whole of the matter. Hume states it himself, in the sentence that turns moral language public: to call a person vicious is to express sentiments in which the speaker "expects all his audience are to concur."27 The standard is set by the audience whose concurrence is sought. The manifesto named a verdict by "the affected public" and left the public unanalysed; the counter-argument called this an institutional gap, a promise of legitimacy without a mechanism. We supply the mechanism in three parts, and each part is borrowed from a pillar.

StandingWho is in the audience

On the Humean value theory we have now committed to, those who can be helped or hurt — but assessed, as Floridi insists, within the infosphere they inhabit, where identities are informational and a harm to the structure of the environment is a harm to the inforgs constituted by it.28 This is the settlement of the manifesto's third tension, and it is not a pluralism. We are Humean on the content of value and Floridian on the ontology of those who bear it. Floridi's patient-orientation tells us where to look for standing — anyone at the receiving end of the action, at the relevant level of abstraction — and Hume tells us what having standing means: that one's welfare enters the spectator's sentiment. Standing, in short, is being a patient of the system, and the audience is the set of its patients.

InclusionWho is heard

Having standing is not the same as being heard. Miranda Fricker's account of testimonial injustice names the failure precisely: a person's report of their own harm is discounted because of who they are, and the discount is invisible from within the community that applies it.29 A training distribution is, among other things, a testimonial economy — a record of whose corrections were weighted and whose were not — and the second objection is, at bottom, the observation that this economy is unjust by default. Hume's audience clause makes the remedy a requirement rather than a courtesy: a spectator who seeks the concurrence of all the affected must credit the affected as knowers of their own condition, or the concurrence sought is that of a smaller audience wearing the name of a larger one. Arendt's enlarged mentality is the same requirement in the idiom of Kant: to judge from the standpoint of everyone else, one must first have let everyone else speak.30

EnforcementWhat a verdict changes

A verdict that changes nothing has corrected nothing. Here Epstein completes the mechanism. The to mode, we said, is the public's judgement of a system's consequences; but on Epstein's account a judgement becomes a social fact only by re-anchoring a frame principle — by changing what counts as an acceptable deployment, an approved procurement, a permitted override, a compensable harm.31 A public that can render verdicts and cannot re-anchor is a chorus. The mechanism of correction is the institutional path from verdict to frame principle: standards that can be rewritten, procurement rules conditioned on them, override and halt rights held by the public's representatives rather than the system's owners, and liability that attaches to the qualities Section I named. Without that path the to mode is, as the counter-argument said, a promise of legitimacy. With it, it is what Hume meant by a convention keeping step with its circumstances.

We do not claim that a public so constituted cannot be captured. No theory removes that risk, and a theory that claimed to would be the more dangerous for it. What a theory does is make capture visible as a specific failure — a violation of the audience clause — rather than as the natural state of things. Senso comune in a Humean accent is what corrected sentiment becomes when its audience is allowed to shrink. The remedy is not a better Hume. It is a wider audience, and institutions that cannot render a verdict without one.

VI — Design

What can be built in

The first objection was the most philosophical and the most damaging to the manifesto's language. "Engineering the corrective process," it said, mistakes the kind of thing a correction is. On Wittgenstein's account of rule-following, no rule fixes its own application; every application is itself an act that a further rule would have to license; and the regress stops not in a final rule but in a practice — agreement not in opinions but, as he put it, in form of life.32 A corrective process, on this view, cannot be built into a system, because what makes a correction a correction is not a mechanism inside the system but the practice outside it within which the mechanism's outputs count as corrections at all.

We accept this, and we note that Hume accepted it first. His account of judgement was never an account of rules. Custom, "the great guide of human life," does the work that rules were supposed to do and cannot; the general point of view is a habit of correction acquired in company, not an algorithm applied in private; and the artificial virtues are sustained, he says, by the practice of a society and the education of its members, not by anyone's derivation of them from principles.33 The objection defeats the manifesto's vocabulary and confirms its pillar. What it leaves is the question of what design can do, once "installing the corrective process" is off the table, and the honest answer has three parts and a limit.

Design can form durable qualities. Training is not the specification of a disposition but the induction of a system into a practice — a form of life composed of the data it is shown, the raters who correct it, the documents it is asked to adhere to, and the outcomes on which it is selected. What the system acquires is what that practice rewards, which is why the question of Section V returns at the design stage in its most consequential form: the community whose corrections form the disposition is the community whose senso comune the system will carry. The remedy the counter-argument sought is not a better objective function. It is the composition of the training public, treated as a matter of standing and inclusion rather than of convenience.

Design can preserve the circumstances. Section IV's argument for oversight, halt, plurality, and distributed authority is an argument about architecture — about what capabilities are granted, what reach is permitted, what remains correctable from outside — and architecture is something design decides. It cannot make a system just. It can keep the circumstances in which justice applies.

Design can build the infrastructure of spectatorship. Explicability, evaluation at the level where tendencies live, the means to read a disposition rather than infer it from acts: these are, as the sixth objection insisted, requirements on the social infrastructure of correction, and infrastructure is built.

What design cannot do is replace the public, guarantee the disposition, or end the regress. It cannot make a verdict bind without the institutions of Section V; it cannot ensure that a practice formed the quality it was meant to form rather than a knavish counterfeit that scored the same; and it cannot close, from inside a system, the gap between a rule and its application that only a form of life closes from outside. The manifesto's for mode is therefore revised. It no longer names a corrective process to be engineered. It names three things design can do — form qualities, preserve circumstances, build spectatorship — and the limit beyond which the rest belongs to a public, a practice, and a convention kept in step with its time.

Coda

Good sense, corrected

We began by saying that a manifesto recommending good sense should be able to take correction. We can now say what taking it has produced. AI safety, on a Humean account, is the possession by artificial agents of durable qualities useful and agreeable to those they affect: veracity, fidelity, justice, temperance, discretion, and a courtesy that is not its own counterfeit. Those agents need no minds to be so assessed; Hume's standard never asked for one, and Floridi's account of agency without intelligence tells us what kind of thing it is being applied to. The facts of safety are made where Epstein says social facts are made — grounded in systems, anchored in institutions — and the deepest of them, that justice applies at all, is anchored in circumstances that capability can dissolve and that safety must therefore preserve. The verdicts that keep the convention in step are rendered by a public that has standing because it is affected, is included because it is credited, and corrects because its verdicts re-anchor what counts. And what design can do is form, preserve, and build — and then stop.

The sensible knave remains. We can only arrange the world so that his reasoning is false: that his exceptions are seen, that his breaches cost, that he remains our equal and not our superior. That arrangement is what we now mean by safety.

Good sense, Hume said, is a natural ability, and he refused to deny it the name of virtue. This essay has been an attempt to exercise it on our own earlier work — to let the audience widen and the sentiment correct. The counter-argument that occasioned it did the review a service, and we would rather have more such objections than fewer. The table still has empty chairs.

Notes

Notes

  1. A skeptical counter-argument in seven objections, prepared in July 2026 as a stress test of Good Sense, For, By, and To AI (GoodSense.ai Review I(1)). Its objections concern, in order: Wittgensteinian rule-following; the community-relativity of corrected sentiment (Fricker, Gramsci); Epstein's Standard Model; Floridi's ontocentrism; the limits of mind-less morality; explicability as infrastructure; and the triad's category structure.
  2. David Hume, An Enquiry Concerning the Principles of Morals (1751), 9.1 (Selby-Bigge/Nidditch edn., p. 268). Hereafter EPM.
  3. David Hume, A Treatise of Human Nature (1739–40), 3.3.1.4 (SBN 575). Hereafter T.
  4. T 2.3.2.6 (SBN 411), "Of Liberty and Necessity": actions are objects of moral sentiment only as indications of the character from which they proceed; see also An Enquiry Concerning Human Understanding (1748), §8.
  5. T 3.3.4, "Of Natural Abilities" (SBN 606–14); EPM Appendix IV, "Of Some Verbal Disputes" (SBN 312–23).
  6. Luciano Floridi and J. W. Sanders, "On the Morality of Artificial Agents," Minds and Machines 14, no. 3 (2004): 349–79.
  7. EPM Sections II–IV (benevolence, justice, political society).
  8. EPM Section VI, "Of Qualities Useful to Ourselves."
  9. Stephen M. Omohundro, "The Basic AI Drives," in Artificial General Intelligence 2008 (Amsterdam: IOS Press, 2008); Nick Bostrom, Superintelligence (Oxford: Oxford University Press, 2014), ch. 7.
  10. EPM Section VIII, "Of Qualities Immediately Agreeable to Others," on good manners and the mutual deference that facilitates conversation.
  11. Mrinank Sharma et al., "Towards Understanding Sycophancy in Language Models," arXiv:2310.13548 (2023).
  12. EPM Section V, "Why Utility Pleases" (SBN 212–32); on the sentiment of humanity, EPM 9.5–9.6 (SBN 271–72).
  13. Brent Mittelstadt, Chris Russell, and Sandra Wachter, "Explaining Explanations in AI," in Proceedings of FAT* 2019 (New York: ACM, 2019).
  14. Luciano Floridi, The Philosophy of Information (Oxford: Oxford University Press, 2011), ch. 3, "The Method of Levels of Abstraction."
  15. T 3.2.2.18 (SBN 494–95); EPM Section III, Part I. The phrase "circumstances of justice" is John Rawls's, crediting Hume: A Theory of Justice (Cambridge, MA: Belknap, 1971), §22.
  16. EPM 3.18 (SBN 190–91), on a species intermingled with humanity but incapable of resistance.
  17. T 3.2.2.22 (SBN 497).
  18. Nate Soares, Benja Fallenstein, Eliezer Yudkowsky, and Stuart Armstrong, "Corrigibility," AAAI-15 Workshop on AI and Ethics (2015); Stuart Russell, Human Compatible (New York: Viking, 2019).
  19. Brian Epstein, The Ant Trap (Oxford: Oxford University Press, 2015), ch. 8; the convention passage is T 3.2.2.10 (SBN 490).
  20. Brian Epstein, "What Is Individualism in Social Ontology?," in Rethinking the Individualism/Holism Debate, ed. F. Collin and J. Zahle (Dordrecht: Springer, 2014), §4; The Ant Trap, ch. 6.
  21. EPM 9.22 (SBN 282–83).
  22. EPM 9.23 (SBN 283).
  23. Evan Hubinger et al., "Risks from Learned Optimization in Advanced Machine Learning Systems," arXiv:1906.01820 (2019); Dario Amodei et al., "Concrete Problems in AI Safety," arXiv:1606.06565 (2016); on preference-based training, Paul Christiano et al., "Deep Reinforcement Learning from Human Preferences," NeurIPS (2017).
  24. Luciano Floridi, "AI as Agency Without Intelligence," Philosophy & Technology 36 (2023); expanded 38, no. 1 (2025).
  25. Antonio Gramsci, Selections from the Prison Notebooks, ed. and trans. Q. Hoare and G. Nowell-Smith (London: Lawrence & Wishart, 1971), p. 435; Notebook 11.
  26. David Hume, "Of National Characters," in Essays, Moral, Political, and Literary, ed. Eugene F. Miller (Indianapolis: Liberty Fund, 1987); the footnote was added to the 1753–54 edition and revised for the posthumous edition of 1777.
  27. EPM 9.6 (SBN 272).
  28. Luciano Floridi, The Fourth Revolution (Oxford: Oxford University Press, 2014); The Ethics of Information (Oxford: Oxford University Press, 2013).
  29. Miranda Fricker, Epistemic Injustice: Power and the Ethics of Knowing (Oxford: Oxford University Press, 2007), ch. 1.
  30. Hannah Arendt, Lectures on Kant's Political Philosophy, ed. R. Beiner (Chicago: University of Chicago Press, 1982); Immanuel Kant, Critique of the Power of Judgment (1790), §40.
  31. Epstein, The Ant Trap, ch. 6, on anchoring as the relation that puts frame principles in place — and, by extension, the relation any correction of them must travel.
  32. Ludwig Wittgenstein, Philosophical Investigations, trans. G. E. M. Anscombe (Oxford: Blackwell, 1953), §§201, 241.
  33. An Enquiry Concerning Human Understanding, §5, on custom; T 3.2.2.25–26 (SBN 500), on the reinforcement of the artificial virtues by public practice and private education.

References

References

  • Amodei, Dario, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané. "Concrete Problems in AI Safety." arXiv:1606.06565, 2016.
  • Arendt, Hannah. Lectures on Kant's Political Philosophy. Edited by Ronald Beiner. Chicago: University of Chicago Press, 1982.
  • Bostrom, Nick. Superintelligence: Paths, Dangers, Strategies. Oxford: Oxford University Press, 2014.
  • Christiano, Paul, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. "Deep Reinforcement Learning from Human Preferences." Advances in Neural Information Processing Systems 30 (2017).
  • Epstein, Brian. The Ant Trap: Rebuilding the Foundations of the Social Sciences. Oxford: Oxford University Press, 2015.
  • Epstein, Brian. "What Is Individualism in Social Ontology? Ontological Individualism vs. Anchor Individualism." In Rethinking the Individualism/Holism Debate, edited by Finn Collin and Julie Zahle, 17–38. Dordrecht: Springer, 2014.
  • Floridi, Luciano. The Philosophy of Information. Oxford: Oxford University Press, 2011.
  • Floridi, Luciano. The Ethics of Information. Oxford: Oxford University Press, 2013.
  • Floridi, Luciano. The Fourth Revolution: How the Infosphere Is Reshaping Human Reality. Oxford: Oxford University Press, 2014.
  • Floridi, Luciano. "AI as Agency Without Intelligence." Philosophy & Technology 36 (2023); expanded 38, no. 1 (2025).
  • Floridi, Luciano, and J. W. Sanders. "On the Morality of Artificial Agents." Minds and Machines 14, no. 3 (2004): 349–79.
  • Fricker, Miranda. Epistemic Injustice: Power and the Ethics of Knowing. Oxford: Oxford University Press, 2007.
  • Gramsci, Antonio. Selections from the Prison Notebooks. Edited and translated by Quintin Hoare and Geoffrey Nowell-Smith. London: Lawrence & Wishart, 1971.
  • Hubinger, Evan, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant. "Risks from Learned Optimization in Advanced Machine Learning Systems." arXiv:1906.01820, 2019.
  • Hume, David. A Treatise of Human Nature. 1739–40. Edited by L. A. Selby-Bigge and P. H. Nidditch.
  • Hume, David. An Enquiry Concerning Human Understanding. 1748.
  • Hume, David. An Enquiry Concerning the Principles of Morals. 1751.
  • Hume, David. Essays, Moral, Political, and Literary. Edited by Eugene F. Miller. Indianapolis: Liberty Fund, 1987.
  • Kant, Immanuel. Critique of the Power of Judgment. 1790.
  • Mittelstadt, Brent, Chris Russell, and Sandra Wachter. "Explaining Explanations in AI." In Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT* '19). New York: ACM, 2019.
  • Omohundro, Stephen M. "The Basic AI Drives." In Artificial General Intelligence 2008, edited by Pei Wang, Ben Goertzel, and Stan Franklin. Amsterdam: IOS Press, 2008.
  • Rawls, John. A Theory of Justice. Cambridge, MA: Belknap Press of Harvard University Press, 1971.
  • Russell, Stuart. Human Compatible: Artificial Intelligence and the Problem of Control. New York: Viking, 2019.
  • Sharma, Mrinank, et al. "Towards Understanding Sycophancy in Language Models." arXiv:2310.13548, 2023.
  • Soares, Nate, Benja Fallenstein, Eliezer Yudkowsky, and Stuart Armstrong. "Corrigibility." AAAI-15 Workshop on AI and Ethics, 2015.
  • Wittgenstein, Ludwig. Philosophical Investigations. Translated by G. E. M. Anscombe. Oxford: Blackwell, 1953.