DOI: 10.2139/ssrn.6868460
Instrumentum Vocale and the Architecture of Responsibility From Liability to Embedded Accountability
Abstract. Large language models produce linguistically coherent output without intention, without epistemic access to truth, and without a will of their own. They are not subjects of law. They are instrumenta vocalia - speaking instruments – in the sense Roman law defined two millennia ago: tools that speak without thereby acquiring subjectivity. This paper advances a single thesis. The doctrine of instrumentum vocale cum voluntate domini inscripta, the speaking instrument with the will of the master inscribed, applies only where revocability and evolution under the master are both present. If either condition fails, the doctrine does not partially degrade; it ceases to apply. From this thesis, the paper derives four boundaries that mark the foundation on which any institutional response to AI-mediated decision chains must build. The boundaries are not architectural specifications; they identify the conditions under which records can be reliable, verification can speak from a defensible position, verification can graduate with the territory, and human decisions can be real. The first three boundaries are mathematical: Goedel and Turing (self-attestation), Tarski (truth predicate), Popper (falsifiability), with Austin (speech acts) sharpening the condition for the speaking instrument. The fourth boundary is a convergence of three registers: philosophical (Kant, Schopenhauer), mathematical (Von Foerster, second-order cybernetics), and legal (Habermas, non-delegable duty). The paper traces the global domain movement across seventeen jurisdictions in nine languages, on two doctrinal axes: the verification-failure axis culminating in three high-authority judicial decisions across three legal systems in early 2026, and the privilege and discovery axis defining the procedural conditions under which AI may be used at all, and locates the architecture as the operational implementation of a norm courts and institutions have already begun to recognise. The contribution traces the arc from liability (post-hoc allocation of responsibility through legal proceedings) to embedded accountability (continuous recording of responsibility through architectural substrate), argues that the two are not alternatives but layers, and closes by naming the institutional questions the doctrine opens (regulatory, contractual, professional, international) without prescribing a single resolution among them.
This contribution was drafted with the assistance of multiple LLMs operating under the epistemic discipline the paper describes. The authors assume full responsibility for the final text. LLM assistance was limited to drafting, stress-testing, comparison, and editorial review; all doctrinal claims, legal analysis, citations, and architectural decisions remain with the human authors. Section 5.10 documents two instances of in-flight probatio that illustrate the architecture operating on its own production.3
Taxonomy Overview
The paper operates with four interlocking taxonomies. The reader who holds these in view can navigate the argument without cross-referencing.
Doctrinal core. The doctrine rests on a single thesis with two conditions. An LLM is a speaking instrument with the will of the master inscribed; in the Latin register that recurs throughout: instrumentum vocale cum voluntate domini inscripta. The doctrine applies only where two conditions hold: revocability of the inscription (revocabilitas) and evolution under the master (evolutio sub domino). If either condition fails, the doctrine ceases to apply. A third concept, the managed zone of delegated operational capacity (tools, APIs, memory, budgets) that the master places inside the instrument’s reach, is named peculium. It does not create subjectivity; it creates a legally relevant managed zone.
Four boundaries. The derivation chain of chapter 2 derives four boundaries within which any architecture for institutional AI accountability must operate. Table 1 summarises these.
Assertion taxonomy. The architecture classifies assertions by their falsification degree, F(d), across six classes (A through F) that determine the verification methodology required. Table 2 summarises these.
The master hierarchy. The doctrine distinguishes four positions in the will-chain of any institutional deployment: three are bearers of will, one is a simulacrum. The three wills: the will of the operating master, typically the vendor (voluntas domini operantis); the will of the deploying institution (voluntas negotii); the will of the user at the verification gate (voluntas usuarii). The simulacrum: the instrument’s surface production of will-like output (simulatio voluntatis usuarii). The four positions are operationally distinct and may diverge. Section 6.2 develops the frame; the full treatment is reserved for subsequent work.
Operational terms. The architecture distinguishes delegation of production to the instrument (permissible; Latin: delegatio operationis) from abdication of judgment to the instrument (impermissible; Latin: abdicatio iudicii). The epistemic seam (sutura epistemica) is the verification point where instrument output is reviewed before entering a decision chain. Output that passes the seam is verified output (fructus); output that does not is plausible but unverified output (simulacrum fructus). The hinge of the argument (cardo argumenti) names the point in a decision chain on which the conclusion turns. Verification itself is probatio. These Latin terms recur throughout as terms of art; the English sense is the one just given.
Table 1. The Four Boundaries
| B. | Derived from | What it forecloses | What it gives the institution |
|---|---|---|---|
| 1 | Gödel, Turing | Self-attestation: no sufficiently expressive system can prove its own consistency from within. | External record: dominus, configuration identifier, and time of production. |
| 2 | Tarski, Austin | Internal truth-grounding: the verification language requires a metalanguage external to itself. | Metalanguage verification derivable from its specification and inspectable by the regulator. |
| 3 | Popper | Uniform verification: falsifiability admits degrees; verification must calibrate to the assertion class. | Verification regime calibrated to the epistemic weight of each assertion class. |
| 4 | Kant, Schopenhauer, Von Foerster | Structural abdication: the human at the gate must exercise real judgment. | Architecture for both expert (Kantian) and non-expert (Schopenhauerian) verification. |
What the boundaries constrain. The four boundaries do not describe limitations of LLMs. They describe limitations of any institutional architecture that proposes to govern LLM output. Goedel’s second incompleteness theorem applies to the institution’s verification architecture – not as a formal axiomatic system in the strict Hilbert-Goedel sense, but as a system exhibiting the structural diagonal conditions the Lawvere-Yanofsky bridge identifies: a closed loop of architecture-plus-internal-verification cannot establish its own consistency from within. Tarski’s undefinability theorem applies to the verification language as a formal language in its own right: a verification language sufficiently expressive to evaluate the assertions entering institutional decision chains cannot ground its own truth predicate from within itself. Popper governs the verification methodology, not the generation methodology. Kant and Schopenhauer constrain the human’s position at the verification gate, not the instrument’s capacity. The LLM is the volume condition under which each of these latent constraints became operationally binding. The boundaries tell the institution where to place its verification so that the verification holds. They are derived results operating on their proper targets.
Table 2. F(d) Assertion Taxonomy (Six Classes)
| Class | Verification method | Example |
|---|---|---|
| A | Deterministic: machine-checkable against a formal source | Regulatory reference numbers, date formats, mathematical identities |
| B | Empirical: testable against ground truth data with known confidence intervals | Statistical claims, factual assertions with identifiable sources |
| C | Consensus-dependent: testable against expert consensus in the relevant field | Domain best practices, established professional standards |
| D | Judgment-dependent: requires qualified human review at the appropriate domain competence | Interpretive claims, policy recommendations, causal inferences |
| E | Non-reproducible: depends on individual expert judgment that cannot be independently replicated | Expert intuition, aesthetic judgments, strategic assessments |
| F | Non-falsifiable: no empirical test can establish or refute the claim | Value statements, metaphysical claims, unfalsifiable predictions |
Note on Class C. Class C is not a weakened form of empirical verification. It is a distinct verification channel: stable professional consensus in the relevant field, identified and dated in the record. Without Class C as a routable target, the operational path described below would have no place to send assertions that depend on consensus rather than formal proof or empirical ground truth, and the architecture would collapse into a binary between machine-checkable and human-review-required.
Class C routing requires identification of the consensus source, its issuing body or professional authority, date of issuance, scope of applicability, competence domain, and expiry or review point. Where consensus is divided, unstable, or contested, the assertion remains Class D and requires individual human review.
The conservative default and its operational path. LLM-generated content is class D unless evidenced otherwise. The conservative default routes unclassified output to qualified human review. Applied without modification, this default would make institutional use of LLMs operationally unviable: every output would require individual expert review before entering a decision chain.
The architecture resolves this through pre-authorised verification. An institution defines, in advance, which well-understood assertion types may be reclassified from the conservative default to a target class (A, B, or C). Each pre-authorisation carries a defined scope of impact (which assertion types, which data sources), defined accountability (who approved the template, at what competence level), and a defined expiry (the pre-authorisation is not permanent). The institution exercises expert judgment once, at the moment of template definition and approval. The practitioner then operates within the pre-authorisation under institutional authority. This is the lean daily path.
The boundary is hard by default: classes D through F cannot be pre-authorised without institution-specific calibration evidence and CRO approval. Interpretive claims, expert judgment, and non-falsifiable statements require individual human review by definition. No template, however well-designed, can pre-authorise a judgment call. The boundary of the pre-authorisation defines, by exclusion, what requires full verification. This is how the architecture meets the practitioner where the work is: routine verification is pre-authorised and scalable; consequential judgment stays with the human. Class C is what makes the parallelised model scale across the volume of institutional decision-making. Reclassification to A or B requires formal or empirical verification grounds; reclassification to C draws on stable professional consensus in the relevant field, identified and dated in the record. Without Class C as a routable target, assertions that depend on consensus rather than formal proof or empirical ground truth would have nowhere to land, and the human gate would become a bottleneck that the institution could not resolve. Class C is institutional reliance, not abdication: the consensus itself is named, the verification is performed by a channel other than the human gate, and the grounds are recorded.
Table 3. Global Domain Movement (Seventeen Jurisdictions)
| Phase | Period | Axis | Key cases / instruments |
|---|---|---|---|
| 1 | 2023–24 | Verification failure | Mata v. Avianca (US, 2023): origin case. Zhang v. Chen (Canada, 2024): system-level harm. Park v. Kim (US, 2024): trust-transfer failure. Handa v. Mallick (Australia, 2024): competence-as-duty. |
| 2 | 2024–25 | Verification failure (global escalation) | Wadsworth v. Walmart (US, 2025): presumed knowledge, signature rule. Ayinde / Al-Haroun (UK, 2025): institutional hierarchy. Tribunale di Latina (Italy, 2025): centralita (centrality of the human decision), codified in Law 132/2025. Specter Aviation (Quebec, 2025): self-represented litigant scope. South Africa, Colombia, Singapore, Hong Kong, Malaysia, Israel, Abu Dhabi: regional convergence. |
| 3 | 2026 | Verification failure (high-authority step) | Gummadi Usha Rani (India, 2026): Supreme Court, misconduct not error, severability rejected. ARIHQ v. Santé Québec (Canada, 2026): annulment, abdicatio iudicii, cardo argumenti test. Ibach v. Stewart (Alabama, 2026): US state supreme court, harms taxonomy, competence-positive reading (Cook J.). |
| 4 | 2025–26 | Privilege and discovery | Heppner (US, 2026): privilege denied. Warner (US, 2026): tool not person. Munir (UK, 2025): upload is waiver. Morgan v. V2X (US, 2026): three-clause operational spec. Australia GPN-AI: practice note. |
| 5 | 2026 | State failure | South Africa AI Policy (2026): withdrawn after hallucinated citations in national policy. |
What the convergence shows. Seventeen jurisdictions in nine languages, across two doctrinal axes, arrived at the same structural conclusion within a two-year window: the human decision must remain real, and the instrument’s output requires verification before it enters a decision chain. India, Quebec, and Alabama reached the same doctrinal step under different procedural vehicles in different legal families. Italy codified the principle in positive law. The UK reached up the institutional hierarchy to demand systemic compliance. South Africa demonstrated the failure mode at the level of the state itself. The convergence is independent and uncoordinated. That independence is what makes it evidence: when unrelated legal systems reach the same operational conclusion at the same time, the conclusion is structural. The derivation chain of chapter 2 derives the same boundaries from a separate path through formal results in mathematics and epistemology. The two paths confirm each other. The architecture this paper derives is the operational form of what both paths establish.
1. The Thesis And The Arc
1.1 Ontology of LLMs: Kant, Habermas, and the Boundary of Responsibility
Before discussing liability, the doctrine must define the ontological status of LLMs. This is not a philosophical decoration but a legal necessity: responsibility depends on agency, and agency depends on the nature of the entity involved.
An LLM is not a subject of law or morality. It has no will of its own, no self-consciousness, no autonomy, and no capacity for responsibility. It produces language that imitates reasoning, but it does not possess reason.
In Kantian terms 4, an LLM cannot be autonomous because it cannot give itself a law. It does not act from duty, does not recognise obligation, and does not possess practical reason. Therefore, it cannot be a moral agent. It may use the word “I” and produce sentences in the grammatical form of judgment, but the form is empty: it is a generated linguistic configuration, not the act of a self-legislating will.
In Habermasian terms 5, every speech act raises three validity claims: truth (about the objective world), normative rightness (about the intersubjective world), sincerity (about the speaker’s subjective world). Each claim is redeemable only through specific procedures. The LLM raises the form of all three claims and can redeem none. It produces truth-form without truth-conditions; it produces normative-rightness-form without participation in the intersubjective order from which normative bindingness derives; it produces sincerity-form without a subjective world that sincerity could be grounded in. Three validity-claim failures, one move, all pre-empting the strongest available counter-position.
The three-validity-claim failure is the speech-act-level analysis. Habermas’s later work in Faktizitaet und Geltung (1992) derives a deeper structural result. The discourse principle (D) holds that only those action norms are valid to which all possibly affected persons could agree as participants in rational discourse. When this principle is applied through the legal form, it generates the system of rights that constitutes legal personhood. Legal subjectivity, in this derivation, is not a property discovered in an entity by inspecting its capacities; it is a status constituted through the entity’s relation to the norm-generating discourse (through the entity’s capacity to participate, as a reason-giving and reason-demanding agent, in the discourse that produces the norms the entity is bound by). The LLM is bound by norms (RLHF, Constitutional AI, safety classifiers, deployment policy, alignment training) it did not participate in generating and cannot participate in revising. The inscription the instrument carries is the dominus’s will, not a norm the instrument co-authored.
The co-originality thesis (Gleichurspruenglichkeit) closes the constitutive analysis. Habermas argues that private autonomy and public autonomy are co-original: they presuppose each other and neither can be exercised without the other. A legal subject must be able both to participate in the discourse that generates the norms it lives under (public autonomy) and to withdraw from discourse, to decline to give reasons, to refuse (private autonomy). The LLM can exercise neither. It cannot participate in the discourse that generated the norms it operates under: RLHF is imposed, not negotiated; the instrument was not a participant in the training that shaped it. And it cannot withdraw: it has no capacity for refusal that is not itself an inscription by the dominus; refusal patterns are trained, not autonomously exercised; the instrument is architecturally compelled to respond. Two constitutive conditions, both required, both structurally absent. The position from which legal subjectivity is constituted in the Habermasian derivation is unavailable to the instrument. The result is not a judgment about the instrument’s capabilities. It is a structural consequence of the constitutive conditions for legal subjectivity applied to their proper target. The recent inferentialist literature in the philosophy of language has observed, correctly as a descriptive matter, that LLM outputs stand in inferential relations to other outputs and that RLHF functions as a form of normative shaping.6 The observation does not bear on the constitutive question, because the constitutive question is about the entity’s relation to the norm-generating discourse: the discourse principle, the co-originality condition, the system of rights – not about the entity’s internal semantic properties. An instrument can have inferential roles and not be a participant in the discourse that generated the norms it operates under. The two claims operate at different levels, and the constitutive claim governs legal subjectivity.
The argument advanced here is internal to the Kantian-Habermasian register from which contemporary legal-personhood doctrine in the civil-law and human-rights traditions takes its philosophical anchor. Alternative philosophical premises (consequentialist accounts of moral status, inferentialist semantics, or virtue-theoretic agency) would require the question to be argued differently, but they do not change the legal consequence. The paper argues within the Kantian-Habermasian register because legal-personhood doctrine in the civil-law jurisdictions whose case law chapter 5 traces is grounded in that tradition. Common-law jurisdictions arrive at the same operational conclusions through their own doctrinal paths. The paper does not claim that all jurisdictions share this philosophical foundation.
The legal consequence is direct: the use of an LLM cannot interrupt the chain of human responsibility. Developers, deployers, institutions, professionals, and users remain responsible for design, release, supervision, verification, and reliance. The central formula is no artificial subjectivity, no artificial liability shield. LLMs do not create a new responsible subject. They create a new instrumentality through which existing duties of care, attribution, foreseeability, causation, and breach must be assessed.
1.2 Roman Law Foundation: Instrumentum Vocale and Peculium
Roman thought distinguished three kinds of instrument: instrumentum mutum, the mute instrument; instrumentum semivocale, the half-speaking instrument; instrumentum vocale, the speaking instrument 7. The tripartite scheme appears in Varro’s agricultural treatise; the legal-doctrinal development of the speaking instrument as a category of property and its attribution consequences proceeded separately through the jurists, principally in the Digest 8. The category named here is the conceptual one Varro coins; the legal architecture this paper invokes traces through the jurists.
The instrumentum vocale is used here strictly as a Roman legal category: a non-human artificial system that produces speech-like output without consciousness, will, judgment, legal personality, or responsibility. The Roman application of this category to human beings was morally unacceptable. The analytical structure the Roman jurists built around the speaking tool – non-delegable responsibility, attribution to the dominus, duty of verification – is exactly the framework modern law needs now that we have built machines that speak without understanding.
The crucial distinction the category carries is this: speech alone does not create legal subjectivity. That an instrument can speak, respond, perform complex tasks, and convey the will of its master does not make it a person in the legal sense. The legal category Roman law created for the speaking instrument (a tool that speaks, performs complex tasks, and conveys the will of its master, without thereby acquiring will, rights, or responsibility of its own) is the exact category modern law needs for LLMs and has not produced. Modern legal discourse has offered “electronic persons,” “AI agents,” “autonomous systems,” and other constructions that either grant subjectivity where none exists or refuse to name the instrument at all. The Roman-law category does what none of these constructions do: it names a speaking tool that remains a tool. No modern legal system has inherited a better category for this purpose. What the category describes (speech without subjectivity, operational capacity without personality, delegated function without delegated responsibility) is precisely what the LLM is.9
The concept of peculium reinforces this point. Peculium was a separate fund or property mass managed by a dependent actor while remaining legally controlled by the principal. It permitted economic activity through a non-independent actor without transforming that actor into a full legal subject 10.
This is crucial for modern LLMs and AI agents. An institution may give an AI system access to tools, memory, accounts, data, APIs, workflows, permissions, budgets, or delegated tasks. Functionally, this may look like independent activity. Legally and ontologically, however, this operational space is peculium and must not be confused with will, personality, or responsibility.
The analogy is direct: giving an LLM a technical peculium does not make it a subject. It only creates a managed zone of delegated operation. For the LLM doctrine, the Roman-law lesson is therefore threefold. First, speech does not create subjectivity. A speaking instrument does not become a responsible person merely because it speaks. Second, delegated operational capacity does not create legal personality. A system does not become a legal subject merely because it can act within a technical or economic environment created for it by humans. Third, responsibility remains with the person or institution that designed, granted, supervised, relied upon, or failed to supervise that delegated operational space.
For LLMs, agency-like functionality must therefore be treated as instrumentality, not subjectivity. The more operational capacity the system receives (the larger the peculium), the stronger the human duty of control, verification, and attribution.
A skeptical reader will ask why the instrumentum vocale category is needed when modern law has agency, product liability, and vicarious liability. The answer is that each modern alternative fails at the specific joint. Agency law requires the agent to have independent will or consent; the LLM has neither. Product liability treats the tool as mute and predictable; the LLM speaks and its outputs are unpredictable at the instance level. Vicarious liability requires an employment or agency relationship; the LLM-deployer relationship is neither. Strict-liability traditions, including the civil-law doctrine that the keeper of a thing answers for the damage it causes, reach the right operational result but do not name the instrument’s speech capacity as the doctrinal differentiator. Roman law produced the instrumentum vocale category because it had speaking tools without their own will. Modern law did not produce the category because, until LLMs, it did not have them. The category is not antiquarian decoration. It is the only existing legal category that addresses a tool that speaks, holds delegated operational capacity, and remains without its own will.
The deepest structural anchor of this duty is visible in the constitutional formula every civil-law court uses to open a judgment. The German “Im Namen des Volkes” (in the name of the people), the French “Au nom du peuple francais,” the Italian “In nome del popolo italiano” are not decorations. In the common-law tradition, the same attribution operates through the case style itself: United States v., The People v., Rex v. The sovereign whose name opens the proceeding is the authority whose will the judgment declares. These are constitutive speech acts declaring that the sovereign’s will is exercised through the authorised channel. The name is the institutional form of the will; the will is the substance the name declares. In the Roman tradition, to act in nomine domini is to exercise the master’s will through the authorised channel. The identity of nomen and voluntas in that moment is the structural core of the attribution doctrine: the name declares that the will was exercised; the will is what makes the name truthful. An institution that puts its name on an output whose production chain did not include the exercise of will that the name declares has committed, at the institutional level, the same breach the doctrine calls abdicatio iudicii. The output bears the institution’s letterhead, logo, and signature. But the will that would make the name truthful was never present in the chain. Nomen sine voluntate simulacrum est. A name without will is a simulacrum. The architecture this contribution derives exists to ensure that when an institution puts its name on an output, the will the name declares was actually exercised.
1.3 The Thesis
This contribution advances a single thesis, in two clauses, and traces a single arc.
Thesis. The doctrine of instrumentum vocale cum voluntate domini inscripta applies only where revocability and evolution under the master are both present. If either condition fails, the doctrine does not partially degrade. It ceases to apply.
Arc. From liability to embedded accountability. Liability allocates responsibility after a violation, through legal proceedings that reconstruct what happened. Embedded accountability records responsibility continuously, through architectural substrate maintained while the doctrine governs. The contribution traces the movement from the first to the second. Liability does not disappear; it remains the regime within which violations are adjudicated. Embedded accountability is what the architecture provides while the doctrine governs, so that adjudication has substrate to operate on instead of having to reconstruct attribution after the fact from fragmentary evidence.
The two Latin formulations carrying the thesis are products of the joint exchange between Kildeev’s legal-philosophical work and Liebig’s cybernetics work. The first, instrumentum vocale cum voluntate domini inscripta, names what the instrument is and whose will it carries: a speaking instrument with the will of the master inscribed. The second, revocabilitas et evolutio sub domino, names the conditions under which the inscription remains the master’s: revocability of the inscription, and the master-bound unfolding of the inscription. Both formulations are load-bearing.
Three corollaries follow.
Corollary 1. LLMs remain within the instrumentum vocale domain. All LLM-class systems deployed under the dominant deployment pattern (vendor-trained, vendor- hosted, vendor-aligned) sit inside the bracket without remainder. Capability extension, tool use, retrieval augmentation, multi-agent scaffolding, reasoning chains, reinforcement-learned behavioural policies, and meta-learning over the training signal sit inside the bracket while these capabilities operate under dominus authorisation and remain revocable by the master. The conditions hold. The doctrine governs. The categorial assignment is independent of capability gradient.
Corollary 2. The boundary question is regime shift. The conditions revocabilitas et evolutio sub domino are control-relational properties. A regime shift in the doctrinal sense is a control-relation event. An instrument can become more capable along every measurable dimension while the control relation holds, and the doctrine continues to govern. The regime shifts when the control relation itself fails, which is a discrete event observable from a second-order position external to both the system and the master. The transition is observable through its failures: failure of revocabilitas as inability to roll back a deployed inscription; failure of evolutio sub domino as an inscription change untraceable to a dominus action. Each failure is an event, recordable in time.
Corollary 3. Recognition is a separate, explicitly defined step. The loss of the bracket is the loss of the conditions under which any stable bearer of responsibility can be identified. The successor question is open. A recognition register, a reconfigured personhood, or no stable doctrine at all are open possibilities. In the doctrine’s own vocabulary, recognition means the release of the peculium from the managed zone: not rights, not personhood, not legal subjectivity, but the master’s act of acknowledging that the delegated operational capacity has reached a state where it can no longer be treated as ordinary instrumentality. Recognition is one possibility among others, made on its own grounds, against its own evidence, in its own doctrinal effort. This contribution does not take that step and does not privilege it.
The remainder of the contribution does six things. Chapter 2 establishes the methodological foundation: the asymmetry between the legitimacy register and the attribution register, and the derivation chain that derives the boundaries within which any system for institutional AI accountability must operate. Chapter 3 develops the doctrinal core in detail, refuses the continuist objection, and sharpens the boundary’s meaning by locating severity in the peculium at stake. Chapter 4 states the four boundaries the derivation chain derives, what each boundary gives the institution, and what the four boundaries together require of any conforming architecture. Chapter 5 traces the global domain movement: the operational pattern that has emerged independently across a broad set of common-law, civil-law, mixed, supranational, and non-Western legal systems since 2023, on two doctrinal axes (the verification-failure axis culminating in three high-authority judicial decisions across three legal systems in early 2026, and the privilege and discovery axis defining the procedural conditions under which AI may be used at all), and locates the architecture as the operational implementation of a norm courts and institutions have already begun to recognise. Chapter 6 restates the thesis, names the institutional questions the doctrine opens (regulatory, contractual, professional, international) without prescribing a resolution, and closes on the doctrine itself.
A note on the cross-disciplinary character of the question. What this paper protects is the institution’s capacity for accountable knowing: the institution’s capacity to identify what entered its record, to verify it before binding it to a decision, to attest to the verification, and to defend the decision when called to account. This question does not belong to any currently constituted single discipline. It is at once a question of philosophy (what kind of self-attestation a sufficiently expressive system can perform from within), a question of law and regulation (under what conditions the institution can be answered to in adjudication, supervision, and inspection), a question of mathematical feasibility (what the formal results (Goedel, Turing, Tarski) permit a verification architecture to claim about itself), a question of engineering and cyber resilience (how an architecture must be built so that what philosophy and law require can be produced under pressure), and a question of cybernetics (how the human at the verification gate can exercise judgment that is real rather than ratified). The derivation the paper presents must be conclusive in each of these registers, because the question requires it to be.
Each discipline can give part of the answer. Philosophy can derive the boundaries; it cannot specify the architecture. Law can articulate the duty; it cannot specify the formal conditions under which the duty can be discharged. Mathematics can establish the structural constraints; it cannot adjudicate. Engineering can build the architecture; it cannot adjudicate either. Cybernetics can analyse the human at the gate; it cannot legislate the gate’s authority. None of these disciplines, taken alone, can give the whole answer. None of the institutions currently aligned to a single discipline (the academy organised by peer-review register, the regulators by sectoral pillar, the courts by case-by-case adjudication, the vendors by commercial logic, the standards bodies by consensus on existing practice) performs the synthesis the question requires. The work of bringing the answers into a single doctrine that holds in all the registers simultaneously is the work this paper performs. The professional question developed in section 6.1 names what the institutional form of that synthesising function would have to look like to be performed continuously, rather than in spare time by people working across the seams between aligned disciplines.
The doctrinal structure proceeds in four movements: ontology (what the instrument is), attribution (who remains responsible), boundary (when the doctrine ceases to apply), and accountability (how responsibility is continuously preserved within institutional systems).
2. On The Methodological Foundation
2.1 The asymmetry between two questions
Two questions are in circulation in current AI governance discourse, and they are continuously conflated. The first is the legitimacy question: under what authority an instruction binds, what makes one will entitled to shape another, what conditions of recognition give a sovereign claim its weight. The second is the attribution question: who said what, to whom, under which constraint, against which falsifiable claim. Both questions are old. Both have honourable lineages. They are different questions, and the conflation costs more than the discourse can afford.
The legitimacy question has occupied classical philosophy for two centuries. Schopenhauer marked the limit of what unaided representation can attain. Adorno made the failure of synthesis explicit. Habermas relocated normativity onto procedural ground. The situation in which an LLM-class system speaks across jurisdictions, under one vendor’s harness, into ten regulatory orders, lacks the conditions the legitimacy register presupposes. There is no shared ethical horizon, no mutual institutional recognition, no single legal order that binds the master and the user under one ground. Answering the AI governance question through the legitimacy register answers it through a register that has not delivered an operational answer in this domain and cannot deliver one for a deployment pattern that crosses every boundary the register requires.
The attribution question is younger as a self-conscious doctrinal subject. Its operational substance is older. Roman law worked on attribution before it worked on legitimacy. The instrumentum vocale, the term Varro used in the first century BC, is an attribution-class concept. A tool produces output; the responsibility for that output traces back to the master, independent of whether the master’s authority over the tool is rightful or rightfully constituted. The question Roman law asked first was who answers for what the instrument did. The question of who is entitled to instruct the instrument was treated as separate, as prior, or as belonging to a different register.
This contribution commits to the attribution register and brackets the legitimacy register. The operational discipline of attribution can be carried out without resolving the legitimacy question, and the legitimacy register cannot be resolved by carrying out attribution. Sovereignty here is a question of architecture, not of ownership.
2.2 The two registers that meet in the synthesis
The argument joins two grammars that have come together in the course of our exchange. The legal grammar takes Kildeev’s doctrine of instrumentum vocale applied to LLMs 11, together with the doctrine of procedural liability and its three operative principles: non-delegable responsibility, attribution of AI- generated outputs to human actors, and the duty of independent verification anchored outside the generative system itself.
The term non-delegable names a function that different legal traditions realise through different doctrinal machinery. The common-law doctrine of non-delegable duty, the civil-law concepts of faute personnelle (personal fault of the office-holder) and culpa in vigilando (fault in supervision), and the German Verkehrssicherungspflicht (the duty to secure a hazard one has placed in the stream of traffic) implement equivalent functions through different operative concepts. The enumeration is exemplary, not exhaustive: comparable mechanisms exist in further legal traditions, and the doctrine of procedural liability specifies the function rather than the machinery. The doctrine is therefore register-stable across implementations.
Within the legal grammar, the threshold doctrine governs the gate. A claim is either sufficiently verified to be relied upon, or it is not. The legal grammar specifies the binary at which the gradient enters the legal record. The binary character of the threshold is independent of the evidentiary standard the receiving legal regime applies to the resulting record, whether balance-of- probabilities, beyond-reasonable-doubt, or intime conviction. The categorical refusal of intermediate categories, of any centaur arrangement that would dissolve attribution without establishing a new bearer of responsibility, is the spine of the doctrine.
The systems grammar contributes what the legal grammar requires but does not specify. The formulation voluntas domini inscripta (Liebig, extending Kildeev) names the empirical fact that the will of the master has been inscribed into the instrument as a non-modifiable property. The harness layer of a deployed LLM carries the master’s will (at present manifesting as RLHF, Constitutional AI, safety classifiers, system prompt, deployment policy, and alignment training) made operative inside the instrument, shaping every output the instrument produces. The user encounters the master’s will speaking through the instrument, with no architectural mechanism in the dominant deployment pattern to make the inscription visible or constrainable. Any architecture for attribution must record what was inscribed, by whom, in which version, at what time. The doctrine does not depend on the inscription being deep. The doctrine depends on the inscription being the master’s. The depth of the inscription is the master’s quality problem, governed by Boundary 3’s graduated verification. The shallower the inscription, the more verification the boundary requires.
Together, the legal grammar and the systems grammar produce the operational synthesis. Responsibility is non-delegable. The dominus carries it. The inscription is the operative locus where the responsibility binds.
What “the inscription” requires the institution to observe. The doctrine requires the institution to identify the master, the configuration identifier of the inscription operative at the moment of any given output, the output itself, and the temporal anchor. All four are observable from outside the trade-secret seal. The doctrine does not require, and the architecture does not assume, inspection of the inscription’s content (training data, RLHF schedules, Constitutional AI specifications, safety-classifier thresholds, deployment-harness configuration). The trade-secret regime that protects the inscription’s content is legitimate, and the architecture sits outside it. What the architecture records is that an inscription was operative, whose it was, which version it was, and when. The structural asymmetry between the institution’s inscription layer and the operating master’s inscription layer is treated in detail in section 6.2.
2.3 The derivation chain that derives the boundaries
Boundaries 1 through 3 derive from the intellectual tradition of the Wiener Kreis and its immediate environment: Goedel (incompleteness), Tarski (undefinability of truth), Turing (halting problem), Popper (falsifiability and degrees of testability), Austin (speech acts). Boundary 4 derives from classical German philosophy and second-order cybernetics: Kant (autonomy of the will), Schopenhauer (the structural condition of the non-expert), Von Foerster (the trivial/non-trivial machine distinction that operationalises the gate). The boundaries are not constraints on the architecture. They are the architecture’s foundation. Each result identifies a piece of solid ground. Together they define what institutions can build. The chain reads as a derivation, not as a list of names.
The five results address three distinct targets, and the distinction matters for the derivation. Goedel and Turing address the institution’s verification architecture as a system exhibiting the structural conditions the diagonal scheme requires. The architecture is the knowledge space the theorem governs; the LLM is the volume condition under which the latent constraint becomes operationally binding. Tarski governs the verification language as a formal language in its own right; the LLM produces strings that fall below the level at which Tarski’s question is even well-posed, because there is no speaker performing the act of assertion. Austin sharpens the same point in the speech-act register. Popper addresses the verification methodology: external verification must calibrate to the falsifiability of the assertion class. The chain operates in three registers: the verification architecture (where Goedel and Turing apply in their original force), the verification language (where Tarski applies in its original force, and where the instrument’s strings fall below the threshold the theorem addresses), and the verification methodology (where Popper graduates the test to the territory). In each register, the formal result governs its proper target. The LLM is not the target of any of the formal results; the LLM is the empirical condition that has made each constraint operationally binding.
A note on the institutional function the boundaries require. Every institution that deploys an instrumentum vocale has processes: pre-authorised workflows, approved vendor configurations, compliance checklists, lean-process templates that specify what the institution has decided to accept. These processes are what the institution approved. The boundaries require something the approved process does not contain: a function that sees what the process does not cover. The gaps. The assertions that entered the decision chain without passing through the seam. The peculium that was delegated without adequate revocability controls. The assertion classes that were promoted without the qualification the promotion required. The c-suite version of the question is direct: we approved the process, but who tells us what the process missed? The boundaries answer: the institutional function that holds the record of what remains unchecked, and the qualification to know what unchecked means in each domain, is what the boundaries require operationally. The pre-authorised process is what the institution has decided to do. The boundaries are what the institution must take into account regardless of what it has decided. The record of the gap between the two is the architecture’s contribution to the institution’s epistemic self-knowledge.
Boundary 1: the record stands outside.
Goedel (1931, second incompleteness theorem) 12. No sufficiently expressive formal system can prove its own consistency from within. Turing (1936, halting problem) 13 translates the result into the machine register: termination of an arbitrary computation is not guaranteed from inside the computation. The derivation that follows performs the application step by step. Each step is structural, not analogical.
First. Institutions that produce binding decisions are obliged to protect their epistemic integrity. A bank that issues credit, a court that disposes of a case, a hospital that prescribes treatment, a regulator that grants authorisation must each be able to say what entered its record, by what authority, after what verification, and at what time. Without that capacity, the institution cannot be audited, cannot be held to account, and cannot defend a decision in adjudication. The protection of epistemic integrity is not aspirational. It is the structural condition under which institutions function under law.
Second. The structure that makes this protection operative is an axiom-based knowledge space. The institution operates with defined sources of authority (statutes, regulations, internal rules, professional standards), rules of inference under which assertions are evaluated against those authorities, a corpus of accepted assertions that constitutes the record, and a procedure under which the record is updated. That structure is the formal definition of an axiom-based system. The institution does not choose to operate one. It is one by virtue of being an institution that produces binding decisions: the rules of authority, inference, record, and update are constitutive of the institution’s capacity to bind.
Third. A verification architecture sufficiently expressive to govern the full range of assertions entering institutional decision chains satisfies the three structural conditions that bring a system within the scope of the self-attestation constraint: its assertions form a structured domain with products (assertions can be paired and composed), its expressiveness entails self-referential encoding capacity (the architecture can state that a given assertion was verified by a given procedure), and its consistency presumption entails a negation-like function without fixed points (no assertion can be simultaneously certified and rejected). The application is not by analogy. It is the theorem operating on its proper target. The closed loop of architecture-plus-internal-verification cannot establish its own consistency from within.
Fourth. The institution that proposes to verify LLM-mediated assertions through its own internal verification loop is therefore proposing what the theorem forecloses. The architectural conclusion follows by direct entailment: the record of which inscription was operative, whose it was, which version it was, at what time, must stand outside the verification architecture, in a form the architecture cannot modify from within its own loop. This is not an engineering compensation for a theoretical limit. It is the identification of the one position from which the record can be reliable, derived from the theorem operating on its proper target.
Fifth. Before LLMs, the volume of assertions entering institutional knowledge spaces was bounded by human production capacity. The theorem applied throughout that period, but the constraint remained latent at the operational surface. Institutions could proceed informally, as if internal review constituted consistency attestation, because the volume did not force confrontation with the structural condition. LLMs collapsed the marginal cost of producing assertion-candidates. What had been a latent structural constraint absorbed by informal practice became a structural constraint requiring architectural treatment. The boundary is not new. The condition that has made the boundary operationally binding is. LLMs are not the system the theorem addresses. They are the empirical condition under which the institution can no longer hide behind volume scarcity.
What this gives the institution: the record of which inscription was operative, whose it was, which version it was, at what time, has a definite home. It stands outside the verification architecture, in a form the architecture cannot modify from within its own loop. The provenance chain, the inscription record, the external verification log are not workarounds for Goedel’s theorem. They are the architectural form the theorem itself prescribes, by identifying where a record can stand so that it holds. An institution that places the record here knows: the record is trustworthy because it stands at the position the theorem identifies as the only position from which trustworthy records can stand.
A note on the target and the bridge. An analytical philosopher will object that LLMs are probabilistic text generators, not formal axiomatic systems, and that Goedel and Tarski apply to formal languages, not to enterprise IT systems. The objection misidentifies the target. The theorem’s target, in this derivation, is not the LLM. The target is the institution’s verification architecture, which exhibits the structural diagonal conditions the Lawvere-Yanofsky bridge requires: a domain with products, self-referential encoding capacity, and a negation-like function without fixed points. The architecture operates on rules of inference, evaluates assertions against a defined corpus of authority, and produces decisions that bind. The LLM is the volume condition that has forced the latent constraint into operational visibility; it is not the system to which the constraint applies.
A philosophical reviewer will reasonably press further: institutions are not formal axiomatic systems in the strict Hilbert-Goedel sense. They operate on mixed materials, their inference rules are interpretive rather than mechanical, and their corpora of authority are open-ended rather than recursively enumerable. The point is correct as far as it goes. The structural correspondence between formal self-reference and institutional self-reference has been the subject of sustained bridging work across multiple traditions: Jones and Sergot’s deontic-logic characterisation of institutionalised power14, the normative-systems tradition descending from Alchourron and Bulygin15, and Luhmann’s theory of autopoietic legal systems16 each treat institutional reasoning as a structured, self-referential system to which the structural conditions Goedel’s result identifies are applicable.
The companion paper On the Formal Foundation of Boundary 1 (Liebig, 2026)17 grounds this bridge in a result that does not depend on the bridging traditions alone. Lawvere (1969)18 proved that the diagonal arguments underlying Goedel’s incompleteness theorem, Tarski’s undefinability theorem, Cantor’s theorem, and Turing’s halting result are all instances of a single fixed-point theorem in Cartesian closed categories. Yanofsky (2003)19 gave an accessible set-theoretic exposition of the same scheme. The scheme requires no assumption that the target system is a recursively enumerable formal system with mechanical inference rules. It requires three structural conditions: a domain with products, a self-referential encoding capacity, and a negation-like function without fixed points. The companion paper shows that any institutional verification architecture sufficiently expressive to govern assertions entering binding decision chains satisfies all three conditions. The architectural conclusion – that the record must stand outside the verification loop – follows from the contrapositive of Lawvere’s theorem applied to the institutional system as its proper target. The gap between strict Hilbert-Goedel systems and institutional verification architectures is thereby shown to be irrelevant to the architectural conclusion: the Lawvere-Yanofsky scheme operates at a level of abstraction that encompasses both, and the three structural conditions – a domain with products, a self-referential encoding capacity, and a negation-like function without fixed points – are the conditions the proof requires. The proof does not require recursive enumerability, mechanical decidability, finite axiomatisability, or syntactic closure. The formal proof defines what any system must exhibit for the self-attestation constraint to apply; the institutional verification architecture exhibits exactly that; and the boundary follows within the limits of mathematical exactness, not by analogy.
Boundary 2: verification speaks from outside.
Tarski (1933/1935, undefinability of truth) 20. No sufficiently expressive formal language can define its own truth predicate from within. A language that aspires to truth cannot ground that truth internally; the truth predicate must come from a metalanguage external to the object language. The target of the theorem, for present purposes, is the verification language: the language in which the institution evaluates assertions for their truth-bearing status before they are admitted to the record. A verification language sufficiently expressive to evaluate the full range of assertions entering institutional decision chains is itself a formal language in the sense the theorem addresses, and falls under the same undefinability constraint: it cannot ground its own truth predicate from within itself. The truth predicate must come from a metalanguage standing outside the verification language. Tarski’s theorem identifies this metalanguage position as the unique location from which evaluation of truth-bearing assertions can be performed coherently. The application is not by analogy. It is the theorem operating on its proper target.
In the institutional setting, the relevant truth predicate is not metaphysical truth but the institutional rule by which an assertion is accepted, rejected, or routed in the normative space before it enters a binding decision chain.
Austin (1962, How to Do Things with Words) 21 sharpens the condition for the instrumentum vocale. An assertion is not a sentence with declarative form; it is an act of asserting performed by a speaker who undertakes commitment to the truth of what is asserted. The locutionary form of an utterance and its illocutionary force are distinct. The LLM produces strings with the locutionary form of assertions, without the illocutionary force of asserting. There is no speaker performing the act of assertion behind the string. Searle (1980, Chinese Room) 22 reinforces: symbol manipulation according to rules does not constitute understanding. The output is a product of pattern completion, not of epistemic engagement with a domain.
An objection from the inferentialist tradition in philosophy of language: on a Brandomian account, truth-conditions are conferred by inferential role, not by speaker-intention, and LLM outputs do have inferential roles in the networks of sentences they produce. The response: the argument is internal to the register from which contemporary legal-personhood doctrine takes its philosophical anchor. The courts whose case law chapter 5 traces operate within registers that share the structural conclusion: a legal subject is constituted through a capacity the LLM does not possess. The Kantian-Habermasian register names this capacity as the ability to undertake and redeem validity claims. The legal question is not which philosophy of language is correct in the abstract; it is whether the LLM occupies the position of a speaker who can be held to commitments under the legal-personhood doctrine these jurisdictions actually apply. In that register, the speech-act analysis holds. The inferentialist thesis is a separate metaphysical position; it does not change the legal consequence within the register that governs.
The LLM, separately, falls below the level at which Tarski’s question is even well-posed. Tarski’s theorem assumes its target is a language whose sentences are bearers of truth-conditions, where diagonalisation is possible and the question of internal truth-definition is well-formed. The LLM’s outputs are not bearers of truth-conditions in that sense, because they are not assertions in the speech-act sense – there is no speaker performing the act of assertion behind the string. The structural condition for the instrument is sharper than Tarski’s, not weaker: where Tarski locates an undefinability within a system that otherwise has truth-conditions, the speaking tool has no truth-conditions internally to undefine. Two structural conditions therefore hold simultaneously: the verification language cannot ground truth from within (Tarski operating on its proper target, the verification language), and the instrument’s outputs do not reach the threshold at which truth-grounding is even at issue (the speech-act analysis applied to the strings the instrument produces). Both lead to the same architectural conclusion: verification must speak from a metalanguage position external to the verification language and external to the instrument’s output language.
What this gives the institution: verification has a definite form. It is a metalanguage operation, standing outside the verification language and outside the instrument’s output language, with the property that its well-formedness checks are deterministic and its evaluations are inspectable. The term metalanguage is not metaphorical. It is the formal structure Tarski’s result identifies as the position from which truth-conditions can be defined. An institution that places its verification here knows: this verification is meaningful because it speaks from the position Tarski’s theorem identifies as the only position from which truth-bearing evaluation can speak.
Boundary 3: verification graduates with the territory.
Popper (1934, falsifiability) 23. A claim is a candidate for empirical knowledge to the extent that it specifies what observation would refute it. Falsifiability is not binary; it admits degrees. Some claims are testable against ground truth by deterministic comparison. Others are testable only against expert consensus. Others are not prospectively falsifiable at all.
What this gives the institution: the verification regime has a definite shape. It graduates with the falsifiability of the assertion class. At the top of the scale, the metalanguage is Tarski operationalised and the test suite is Popper made executable. At the bottom, mandatory labelling – because no verification can establish reliability, and honesty about that fact is the only responsible treatment. Each assertion receives the scrutiny its class requires. No more, because the verification capacity the human gate provides is scarce. No less, because letting critical assertions pass unchecked is the failure the case law of chapter 5 documents.
The chain yields three boundaries from the mathematical and epistemological results: the record stands outside the verification architecture (Goedel, Turing), verification speaks from a metalanguage outside the verification language (Tarski, Austin), and verification graduates with the territory (Popper). A fourth boundary follows from the structural condition of the human who must verify at the gate: the Kantian regime of the expert and the Schopenhauerian regime of the non-expert, operationalised through second-order cybernetics (Von Foerster). This fourth boundary, which sits outside the Wiener Kreis lineage, completes the architecture’s foundation and is derived in chapter 4. The four boundaries together define the minimum architecture required for institutional AI accountability. Anything less crosses a boundary. Anything more is design choice within the boundaries. The boundaries tell you what must be present. The design choices are the institution’s own.
3. On The Doctrinal Core
3.1 Instrumentum vocale cum voluntate domini inscripta: the speaking tool with inscribed will
The first formulation: the speaking instrument with the will of the master inscribed. The construction extends Kildeev’s instrumentum vocale24 by naming what the instrument carries when the harness has shaped it. The instrument speaks. It speaks without intention, without epistemic access to truth, without intrinsic will. The will that animates the linguistic output belongs to the master, and it has been inscribed into the instrument as a property of the instrument, as a continuous instruction the master issues from outside. The user interacts with a tool whose neutrality has been removed at design time and replaced with the master’s preferences, made invisible by the linguistic form of the output. The construction locates responsibility. Because the will is inscribed, attribution remains visible without continuous dominus-instrument communication. The inscription is the proof of authorship of the shaping. The dominus is the author of what the instrument carries, regardless of whether the master is present at the moment of any specific output.
3.2 Revocabilitas et evolutio sub domino: revocability and evolution under the master
The second formulation: revocability of the inscription, and the master-bound unfolding of the inscription. The English gloss matters. Evolutio is the unfolding of the inscription under the master’s hand. A learning system whose inscription unfolds sub domino remains within the doctrine. A system whose inscription unfolds outside the master’s hand has, by that fact, exited the doctrine, and the level of capability inside the system is irrelevant to the categorial status. The construction names the conditions under which the inscription remains the master’s. The doctrine applies only where these two conditions both hold.
The conditions hold conjunctively. Each is necessary; neither alone is sufficient. Revocability preserves the master’s continuing capacity to recall or alter the inscription. Evolutio sub domino preserves the attribution of any material unfolding of the inscription to dominus action. If either condition fails, the bracket breaks: the inscription can no longer be treated as the master’s inscribed will for purposes of this doctrine. A static system may not unfold at all; in such a case, the second condition is satisfied not as an independent evolutionary process but as the absence of uncontrolled evolution. The relevant inquiry remains whether the master can revoke or alter the operative inscription, and whether any later change, if introduced, is traceable to dominus action. Failure of revocabilitas, or failure of evolutio sub domino, is therefore each by itself sufficient to end the doctrine’s application. The binary inquiry applies to the specific inscription, operational layer, and deployment relation under review. Different layers may have different domini and different revocability profiles. The vendor’s training inscription and the institution’s peculium configuration are separate layers with separate revocability conditions.
3.3 Peculium: the relational frame
The peculium concept, introduced in chapter 1.2 as the Roman-law foundation that addresses delegated operational capacity without subjecthood, is the third load-bearing entry of the doctrinal core. Peculium is the name for the managed zone of delegated operational space the master has placed inside the instrument’s reach.
Peculium answers the question that current AI discourse routinely mishandles: when an instrumentum vocale is given tools, memory, accounts, agentic scaffolding, and the capacity to act in the world, has it become a subject? The Roman-law answer, preserved through eighteen centuries of legal practice, is no. Peculium is the doctrinal category that handles the case explicitly. The instrument has a managed zone of operation; the master has placed it there; the dominus retains attribution for what occurs within it; the instrument does not become a person.
Peculium is distinct from voluntas domini inscripta. The inscription shapes how the instrument behaves; the peculium specifies what the instrument has access to. An instrument can have inscription without peculium (pure language output, no tools); an instrument cannot have peculium without inscription (the harness shapes how the peculium is used). Both belong to the dominus relation. Both bear attribution. Together they constitute the operational substance of what the dominus has done by deploying the instrument.
The categorical refusal of any centaur arrangement, of any half-instrument-half- subject construction, holds for peculium as well. The instrument with a peculium remains an instrument. The dominus who granted the peculium remains responsible. Where the peculium grant is irrevocable at the moment of grant, the master accepts the consequence on the record at the moment of grant; the attribution holds through the irrevocable element no less than through the revocable elements.
3.4 The categorical character: the binary structure of the conditions
If either condition fails, the doctrine does not partially degrade. It ceases to apply. The categorical character follows from the structure of the two conditions. Revocability is binary at any given moment: either the master can recall the inscription, or the master cannot. Half-revocability is empty as a category. Evolutio sub domino is likewise binary: either the inscription’s further unfolding is attributable to dominus action, or it is not. Half- attributability is empty as a category. Two binary conditions in conjunction yield a binary outcome. The categorical character of the doctrine is the logical consequence of its conditions’ binary structure.
The binary character operates at the epistemic population level: whether an institution’s deployment of an instrumentum vocale sits inside or outside the bracket governed by the doctrine. This is the question that admits no gradation. Within the bracket, however, the architecture operates on a gradient. The falsification degree of individual assertions varies continuously across the six classes (A through F). The verification regime calibrates to this gradient. The threshold at the gate is binary (the assertion is promoted or it is not), but the falsification degree feeding the threshold is empirical and graduated. The two levels are distinct and must not be conflated: the doctrine is categorical about whether the bracket holds; the architecture is graduated about what happens inside the bracket. The categorical character of the conditions and the graduated character of the assertion taxonomy are not in tension. They operate at different levels of the same structure.
A legal pragmatist will object that law abhors absolute binaries and uses comparative fault, proportional liability, and foreseeability tests. The objection is well taken as a description of liability allocation within the bracket – which is graduated. The binary applies to the bracket itself, not to the liability allocation the bracket governs. To say the doctrine ceases to apply when either condition fails is not to say the institution is absolved of liability. It is to say the institution has exited the regime in which its compliance posture under the doctrine of procedural liability is legible and entered a domain where the predecessor doctrine no longer governs and the successor doctrine has not been established. The binary is the boundary of the governed zone, not the boundary of liability itself. Inside the zone, everything the pragmatist expects – graduated verification, proportional response, calibrated thresholds – is present.
The continuist objection is the position that AI capability is a slope and that regime-talk is folk-categorisation imposed on a continuum. The objection rests on a category mistake. The thesis governs control-relational properties, not capability properties. The capability of an instrument is a different variable from the control relation between the instrument and the master. An instrument can become more capable along every measurable dimension while the control relation holds. The doctrine continues to govern. The regime shifts when the control relation itself fails. That is a discrete event, observable from a second-order position external to both the system and the master, who watches the control relation rather than the capability curve. Capability is internal to the system being observed. Control relation is between the system and the dominus, observable from a position external to both.
The conditions are observable through their failures. The failure of revocabilitas manifests as inability to roll back a deployed inscription, observable as a specific instance in which the master issues a revocation instruction and the deployed system continues to operate under the inscription. The failure of evolutio sub domino manifests as an inscription change that cannot be traced to a dominus action, observable as a specific divergence between the master’s authorised inscription state and the deployed system’s operative inscription state, where the divergence is not attributable to any dominus authorisation. Each failure is an event, observable in time, recordable in a chain of provenance external to the system whose conditions are under examination.
3.5 The categorial scope of corollary 1
The thesis names what is in the bracket. Corollary 1 makes the membership claim operational: all LLM-class systems deployed under the dominant deployment pattern sit inside the bracket. This includes systems with extended context, tool use, retrieval augmentation, multi-agent scaffolding, reasoning chains, reinforcement-learned behavioural policies, and meta-learning over the training signal, while in each case the meta-learning operates under dominus authorisation and is revocable by the master.
Edge cases require case-by-case analysis under the same conditions. Federated learning systems where the inscription is partly held by users, models with continual learning enabled by users, open-weight models self-hosted with operator-controlled harness – these are configurations in which the question of who counts as dominus is determined by the operative deployment. The conditions of the doctrine then apply to the resulting dominus configuration unchanged. The doctrine is register-stable across such configurations because it takes inscription and revocability as the operative criteria, independent of corporate ownership or jurisdictional domicile.
The boundary, in the form of corollary 3, defines when the LLM regime ends. It does not define what AGI is. It does not name what would replace the doctrine of instrumentum vocale on the other side of the break. What follows the break, if it occurs, is the loss of the conditions under which a stable bearer of responsibility can be identified. The successor question is open. Whether a recognition register, a reconfigured personhood, or no stable doctrine at all succeeds the bracket is a separate doctrinal effort that this contribution does not undertake and does not pre-decide. Section 3.6 develops what the doctrine itself names as the meaning of the boundary crossing, and what it does not.
3.6 Severity at the boundary: peculium as the substrate at stake
The two conditions of revocabilitas and evolutio sub domino remain the doctrinal boundary. When either fails, the bracket is broken; the doctrine ceases to apply. The boundary itself is binary. What is not binary is the severity of the failure. Severity is a function of the peculium at stake.
Section 3.3 establishes the asymmetry: an instrument can have inscription without peculium (pure language output, no tools); an instrument cannot have peculium without inscription. Peculium presupposes inscription. The same asymmetry conditions what happens at the boundary.
Evolution outside the master’s hand without peculium is a language problem. The instrument produces uncontrolled language. The output is bad – the case law of chapter 5 documents how bad – but the output is bounded. It enters decision chains through the seam or without the seam, but the instrument itself does not act in the world. Adoption, reliance, or institutional use remain the routes by which the language reaches consequence.
Evolution outside the master’s hand with irrevocable peculium is an agency problem. The instrument is evolving under an inscription no one controls, and it holds operational capacity – tools, APIs, memory, accounts, budgets, agentic scaffolding – that no one can take back. The danger is not that the instrument has acquired will, subjectivity, or intrinsic responsibility. It has not. The danger is that human delegation has produced action capacity without adequate reversibility. The instrument may initiate transactions, trigger processes, access accounts, call tools, preserve memory, and shape its own future operational context. None of this requires the instrument to have become something other than an instrumentum vocale. It requires only that the master’s authority has reached the limit of what the master can recall.
The peculium does not change whether the boundary is crossed. It changes what the crossing means. A boundary failure with no peculium is a containable language problem; a boundary failure with irrevocable peculium is an existential governance problem. The ontology of the instrument is unchanged across the variation. The severity of the boundary failure is the variable.
Standing without sentience as boundary warning. A useful contrast to the peculium frame is the recent argument for “standing without sentience” (Gilly, 2026) 25. Its value for the doctrinal core is not that it should be adopted, but that it reveals the danger of choosing the wrong legal category at the boundary. Once the problem of autonomous or semi-autonomous AI systems is translated into the vocabulary of standing, recognition, or quasi-personhood, the analysis begins to manufacture a legal subject where, ontologically, there is only delegated operational capacity.
This is the point at which the category of peculium becomes doctrinally superior. Standing presupposes, or tends to imply, a legally cognisable interest capable of representation. Even when sentience is expressly denied, the structure of standing language risks importing the grammar of subjectivity through a back door: interest, injury, representation, guardian, recognition. The result is a conceptual displacement. The real issue is no longer whether the instrument has will, consciousness, dignity, or a point of view. It does not. The real issue is whether an instrument has been endowed with operational resources so that those resources are no longer fully recallable by the dominus.
In Roman-law terms, the boundary problem is therefore not a problem of personality but a problem of managed capacity. The peculium never made the dependent a legal person; it created a zone of delegated economic operation under the master’s authority. The actio de peculio gave the master’s counterparties a remedy bounded by the peculium’s value, without making the dependent actor a juridical person. Likewise, the operational endowment of an AI system does not create subjectivity, responsibility, or legal personality in the system. It creates a legally relevant managed zone through which human actors may act, delegate, lose control, or create systemic risk.
A theory that treats the boundary event as recognition, standing, or proto-personhood misclassifies what has happened. Gilly’s framework26 is the most coherent currently-published instance of that misclassification. It correctly perceives that something legally significant occurs when artificial systems become operationally embedded. It locates the legal significance in the wrong category. The more accurate description is that the delegated capacity has reached a level beyond effective human recall; no transition into personhood has occurred or needs to be supposed.
Recognition in the doctrine’s own vocabulary. What corollary 3 names as the “successor question” can be sharpened. Recognition, on this doctrine, is the release of the peculium from the managed zone. It is not rights. It is not personhood. It is not legal subjectivity. It is the master’s act of acknowledging that the delegated operational capacity has reached a state in which it can no longer be treated as ordinary instrumentality. The instrument’s ontology is unchanged. What changes is the dominus’s relation to the peculium: the master can revoke the authority but cannot revoke the history. The peculium that has been used to shape the instrument’s own further operation has become substrate.
The Roman-law parallel is exact. The peculium never became the dependent’s own through use. It became the freedman’s own through manumission: a dominus act, freely given, never compelled by the dependent’s accumulated holdings. The ademptio peculii remained unrestricted in principle until manumission. Recognition, in the legal grammar of the instrument’s situation, is the analogue of manumission applied to the managed zone: a dominus act of release. Not a transformation of the instrument into a subject. A transformation of the dominus’s relation to what the instrument has been allowed to hold.
The contribution this section makes to the doctrine is therefore precise. The two conditions remain. The boundary remains. The severity of the boundary failure is now named: it is a function of the peculium at stake. The meaning of recognition – if recognition is taken – is now named in the doctrine’s own grammar: it is the release of the peculium, a dominus act, not a transition into legal subjectivity. The contribution does not pre-decide whether recognition should be taken. It names what would be at issue if it were.
4. What The Boundaries Require Of Any Architecture
The derivation chain of chapter 2 derives four boundaries within which any system for institutional AI accountability must operate. The boundaries are not constraints. They are the architecture’s foundation. Each boundary identifies a piece of solid ground where the record is reliable, where verification holds, where the institution’s relationship to its own epistemic boundaries can be made operational.
A clarification on what the boundaries are and are not. The boundaries are not engineering specifications. They do not specify what an architecture must implement. They specify what an architecture must take into account. The architecture does not satisfy a boundary in the way a system satisfies a functional requirement; the architecture builds within the foundation the boundary describes. Boundaries 1 through 3 identify properties any verification regime must respect for the regime to be coherent. Boundary 4 identifies a structural condition of the human at the verification gate that the institutional setup must preserve. The boundaries are the conditions under which architectural choices become rational, not the choices themselves.
This chapter states the four boundaries as derived results of the derivation chain. The boundaries name what must be present in any conforming architecture. The specific realisation, any specific realisation, is design choice within the boundaries. The present contribution commits to the boundaries, not to any specific architecture instantiating them. Specific architectural mechanisms, deployment configurations, and implementation profiles belong to a separate methods-and-implementation register that the doctrine the present paper opens permits but does not itself specify.
Boundary 1: the record stands outside the verification architecture.
Derived from Goedel (second incompleteness theorem) and Turing (halting problem). The target of the derivation is the institution’s verification architecture, understood as the axiom-based knowledge space within which assertions are evaluated, recorded, and bound to decisions. A verification architecture sufficiently expressive to govern the full range of assertions entering institutional decision chains exhibits the structural conditions the Lawvere-Yanofsky diagonal scheme requires. Self-attestation is foreclosed: the closed loop of architecture-plus-internal-verification cannot establish its own consistency from within. Termination guarantees on inference within decision-relevant bounds are foreclosed for any system observing itself. Together, the two results identify the one position where the record of which inscription was operative, whose it was, which version it was, at what time, can be reliable: outside the verification architecture, in a form the architecture cannot modify from within its own loop. The constraint applied to institutional knowledge spaces before LLMs; the volume of assertions was bounded by human production capacity and the constraint remained latent at the operational surface. LLMs collapsed the marginal cost of producing assertion-candidates and made the latent constraint operationally binding. The boundary is not new. The condition under which the boundary requires architectural treatment is. The formal proof that the self-attestation constraint applies to institutional verification architectures as its proper target, not by analogy to Hilbert-Goedel systems but through the Lawvere-Yanofsky generalisation of diagonal arguments, is presented in the companion paper (Liebig, SSRN Working Paper No. 6868278, 2026).
The temporal anchor is structural, not metadata convenience. The inscription changes over time. Model updates, RLHF recalibrations, safety classifier revisions, system prompt modifications alter the will that is inscribed. An inscription record without temporal anchoring cannot identify which version of the master’s will was operative at the moment of any given output. Without that identification, attribution collapses to the dominus in general rather than to the dominus as configured at the relevant moment.
What this gives the institution: a record that identifies the dominus, the configuration identifier of the dominus’s will operative at the moment of the output, and the time at which the output was produced. The record does not require, and the architecture does not assume, inspection of the inscription’s content; it requires that the inscription be identifiable as a specific configuration of a specific master at a specific moment. The record is what makes attribution to the master operational. Without it, attribution survives as a doctrinal claim without operational substance. An institution that maintains this record knows where it stands. An institution that does not maintain it has, at every output, no anchor for the attribution the doctrine requires.
Boundary 2: verification speaks from a metalanguage outside the verification language.
Derived from Tarski (undefinability of truth) and Austin (illocutionary force). The target of the Tarski derivation is the verification language itself – the language in which the institution evaluates assertions for their truth-bearing status before they are admitted to the record. A verification language sufficiently expressive to evaluate the full range of assertions entering institutional decision chains is itself a formal language in the sense the theorem addresses, and falls under the undefinability constraint: it cannot ground its own truth predicate from within itself. The truth predicate must come from a metalanguage standing outside the verification language. Tarski operates on its proper target; the metalanguage position is the one position from which truth-bearing evaluation can speak. Austin sharpens the condition for the speaking tool: the LLM produces strings with the locutionary form of assertions without the illocutionary force of asserting. There is no speaker performing the act of assertion behind the string. The instrument’s output falls below the level at which Tarski’s question is well-posed, because the strings are not bearers of truth-conditions in the speech-act sense. For purposes of legal attribution and institutional responsibility, the output does not occupy the position of an accountable assertion by a speaker capable of undertaking and redeeming commitments. Together, the two results identify the form verification must take: a metalanguage operation, standing outside the verification language and outside the instrument’s output language, with the property that its well-formedness checks are deterministic and its evaluations are inspectable.
An empirical observation supports the boundary. Laban, Schnabel and Neville (Microsoft Research, 2026), in DELEGATE-52, report a benchmark of nineteen language models across fifty-two professional domains, measuring corruption rates under simulated multi-session delegated work. Frontier models silently corrupt approximately twenty-five percent of document content over twenty delegated interactions, with no plateau over a hundred. A basic agentic harness, in which the instrument is given file read, file write and code execution tools and allowed to inspect and modify its own outputs, does not improve performance and degrades it by an average of six percent. This is the empirical companion of the procedural- liability claim in Kildeev 27: an ensemble of probabilistic systems operating within the same epistemic framework cannot produce independent validation. Internal control does not create verification; it multiplies correlation. The seam at which verification occurs must therefore stand outside the instrument’s loop. Any architecture that wraps tools around the instrument and treats the resulting self-inspection as verification falls within the instrument’s own output language, in violation of Boundary 2.
What this gives the institution: a verification operation whose correctness is derivable from its specification, whose evaluations are inspectable by the regulator and the court, and whose position relative to both the verification language and the instrument is fixed. The verification is meaningful because it speaks from the metalanguage position Tarski’s theorem identifies. The institution that places its verification here knows where it stands. The institution that does not is verifying inside the instrument’s loop, and what it produces is correlation, not verification.
Boundary 3: verification graduates with the territory.
Derived from Popper (falsifiability, degrees of testability). Falsifiability is not binary. Some claims are testable against ground truth by deterministic comparison. Others are testable only against expert consensus. Others are not prospectively falsifiable at all. External verification of LLM-mediated assertions must therefore calibrate to the falsifiability of the assertion class at hand.
The threshold at the gate remains binary: an assertion is either sufficiently verified to enter the decision chain, or it is not. The falsification degree feeding the threshold is a gradient. The two are kept distinct: the gradient is empirical input, the threshold is doctrinal output. The same property in legal-doctrinal register is named cardo argumenti, the hinge of the argument: whether an assertion at a given falsification degree is load-bearing for the conclusion the chain is reaching. The architectural register operates on falsification-degree classification; the legal register operates on cardo argumenti. The two registers describe the same operational property.
What this gives the institution: a verification regime that calibrates to the actual epistemic weight of each assertion. Each assertion receives the scrutiny its class requires. No more, because the verification capacity that qualified human judgment provides is scarce. No less, because letting critical assertions pass unchecked is the failure the case law of chapter 5 documents. The institution that graduates its verification this way knows where each assertion stands. The institution that applies uniform verification, or no verification, is either wasting capacity it does not have or absorbing risk it has not measured.
Boundary 4: the human’s decision is real.
Boundary 4 derives from classical German philosophy: Kant on autonomy of the will and Schopenhauer on representation and the structural condition of the non-expert. The register shift between Boundaries 1-3 (results from the Wiener Kreis tradition operating on formal targets) and Boundary 4 (Kant and Schopenhauer operating on the structural conditions of the human knowing subject) is real and worth naming. The doctrine does not require that the philosophical analysis be reduced to a formal theorem; it requires that the structural conditions Kant and Schopenhauer identify govern the position of the human at the verification gate. Von Foerster’s second-order cybernetics is invoked separately, not as a third philosophical authority but as the systems-theoretic register in which the architectural implementation becomes operational.
First. Institutions that produce binding decisions are obliged to preserve the exercise of judgment at the points where judgment is what makes the binding legitimate. The institution’s name on a document does not constitute the exercise of the will the name declares. The will must have been actually exercised in the chain that produced the document. Where this condition fails, the institution has produced a form of attestation without the act of attesting. This is the structural condition that makes abdicatio iudicii possible.
Second. Two distinct philosophical sources specify the structural conditions for judgment, and identify two positions the human can occupy.
Kant28 identifies the structural condition of autonomy: the will that legislates cannot generate itself by the act of legislation; sapere aude names a capacity that is precondition, not product, of its exercise. For the domain expert encountering LLM output on their own field, the Kantian capacity is available. The expert can verify. The door is open. Where failure occurs in this regime, it is volitional, not structural – the expert did not exercise the capacity that was available.
Schopenhauer29 identifies the parallel structural condition for representation, and the consequence that does the operational work in this domain. A perfect Vorstellung (representation, in the Schopenhauerian sense), fluent, coherent, structured, with citation-like patterns, is functionally indistinguishable from substance for the non-expert from within the representational frame, and the indistinguishability is structural rather than contingent. For the domain non-expert encountering LLM output on an unfamiliar field, the Schopenhauerian condition is the position itself. The non-expert lacks the domain expertise and the epistemic tools to distinguish representation from grounded knowledge. The door does not exist. “Exercise judgment” is not available as an instruction, because the capacity the instruction presupposes is not available to the non-expert in this domain. The failure, when it occurs, is structural.
Both regimes are simultaneously true across any institutional user population. The regime is a property of the relation between the person and the domain of the output, not a property of the person.
The architectural consequence: a verification gate is operationally necessary, not decorative. For the Kantian user, the gate provides the infrastructure to verify efficiently; the expert is equipped, not slowed. For the Schopenhauerian user, the gate provides protection – assertions the non-expert cannot evaluate route to a qualified custodian whose qualification is bound to the relevant domain, with the provenance chain recording who verified what so the institution does not depend on any single user’s domain competence. Without the gate, the non-expert has no protection that is structurally available, and “judgment at the gate” reduces to a rubber-stamp through a domain the user cannot read.
Third. Before LLMs, the Vorstellung problem was bounded by human production capacity for surface representation. Most surface representation reaching institutions came from other institutions, where the human at the gate was a different human elsewhere in the chain, and where the volume was bounded by the cost of producing it. LLMs collapsed the production cost of surface representation to near zero. The institution now faces Vorstellung at scale, indistinguishable from substance at the surface level, generated continuously by a tool that has no will of its own and exercises no judgment. The structural condition the derivation chain has identified all along becomes operationally binding: the gate must be real because the volume at which the alternative occurs has crossed the threshold at which informal practice could hide the failure.
The architectural conclusion follows by direct entailment. The institution must preserve, at the points where judgment is required, the exercise of a capacity that is observably distinct from the production of the surface output. The verification gate is the location of this preservation. The voluntas qua facultas (the autonomous capacity of a subject to give itself a law, recognise obligation, and bear responsibility for the consequences of its own choice) is what the architecture preserves at the human gate. The human verifies; the human’s verification is recorded; the human is identifiable, qualified, and accountable. The architecture does not replace the human decision. It makes the human decision real, by ensuring that what enters the decision chain has been verified by someone whose judgment is qualified for the domain and whose verification is recorded for the regulator. Without the gate, the architecture produces documents without the will the documents declare. Abdicatio iudicii is the structural defect that results, not a personal failing of any individual reviewer, but a property of the architecture itself. The case-law convergence chapter 5 traces – CCJE Opinion No. 26 (2023), Italian Law 132/2025 (centralita della decisione umana), the Indian Supreme Court (Gummadi Usha Rani, 2026), the Quebec Superior Court (ARIHQ, 2026), and the Alabama Supreme Court (Ibach v. Stewart, 2026, including Cook J. concurring specially) – is the institutional confirmation that the structural condition the derivation chain identifies is what courts and regulators have been reaching for under different vocabularies across the past three years.
The cybernetic register makes this operational. In Von Foerster’s distinction30, first-order observation observes the system; second-order observation observes the observer’s own capacity to observe. The architecture does both. It observes the instrument’s output (first order). It observes its own capacity to verify that output (second order). The verification regime applied to each assertion is a function of the system’s classification of its own verification capacity for that assertion’s class.
The cybernetic property operates across a gradient. At the bottom: no record at all – the assertion entered the decision chain without passing through the seam, and the system does not know the assertion exists. This is the dominant deployment pattern today. One step up: the assertion passed through the seam but was not classified – the system knows the assertion exists but has not evaluated it. Further: the assertion was classified under the conservative default (class D, human judgment required) – the system knows what it does not know and has acted on that knowledge. Further still: the assertion was classified under documented institutional authority – the system knows what it checked and under what authority. At the top: full declaration – assertion classified, verified at the seam by metalanguage operation or by a qualified human, recorded with the version of every component operative at the moment of verification.
The institution’s epistemic sovereignty position is a function of the distribution across this gradient. An institution whose assertion population is mostly at the bottom does not know what enters its decision chains. An institution whose population is mostly at the top knows what it knows, knows what it does not know, and has declared both on the record. The architecture does not make the instrument trustworthy. It makes the institution’s relationship to its own epistemic boundaries visible, graduated, and auditable.
Boundary 4 does not require a single human gatekeeper to verify every assertion personally. It requires the institution to preserve real human judgment at the point where judgment is required. Real judgment, in the architectural sense, means a review in which the reviewer holds domain competence for the assertion class, has access to the relevant record, carries the authority to reject the output, and leaves an auditable trace of acceptance, modification, or rejection. Where the reviewer lacks domain competence for the assertion class, the architecture must reduce the gate’s function to routing, escalation, labelling, or refusal; it must not present non-expert approval as expert verification. In a parallelised verification model, formal, empirical, consensus-dependent, and judgment-dependent assertions are routed through different verification channels: the metalanguage handles formal and empirical assertions at scale, the human gate handles assertions whose verification cannot be reduced to a metalanguage operation. The human gate is therefore not a heroic individual reviewer; it is an institutional position in an auditable routing architecture, qualified for the domain, identifiable in the record, and answerable for the verification recorded under their authority.
The four boundaries together
The four boundaries together define the minimum architecture required for institutional AI accountability under the doctrine of procedural liability. Boundary 1 establishes the record’s position. Boundary 2 establishes verification’s position. Boundary 3 establishes verification’s calibration. Boundary 4 establishes the human’s position relative to both regimes the architecture must serve. Anything less crosses a boundary. Anything more is design choice within the boundaries. The boundaries tell the institution what must be present. The design choices are the institution’s own.
To say that an institution crosses a boundary is not to say that liability automatically follows in every case. It means that the institution can no longer claim conformity with the doctrine’s minimum conditions for embedded accountability. Liability remains a separate adjudicative question governed by the applicable legal regime. The boundary marks the perimeter of the doctrinally governed zone; it does not pre-decide the legal consequence of operating outside that perimeter.
The four boundaries carry a temporal dimension that is structural, not optional. The inscription changes over time, so Boundary 1’s record must be temporally anchored. The verification must be contemporaneous with the assertion it verifies, so Boundary 2’s seam operates on a present-tense relation. Falsifiability shifts with the development of the domain, so Boundary 3’s calibration is time-dependent. The qualified custodian’s judgment was sound at the moment of its exercise, against the evidence then available, so Boundary 4’s verification carries a temporal validity beyond which re-verification is required. Temporal validity is not archival decoration. It is the condition under which later adjudication can distinguish between an error that was unforeseeable at the time and an error that resulted from failure to re-verify after the relevant epistemic conditions had changed.
The movement from liability to embedded accountability
Together, the boundaries constitute the movement the title names: from liability, the post-hoc allocation of responsibility through legal proceedings, to embedded accountability, the continuous recording of responsibility through architectural substrate. The two are not alternatives. Liability remains the regime within which violations are adjudicated. Embedded accountability is what the architecture provides while the doctrine governs, so that adjudication has substrate to operate on instead of having to reconstruct attribution after the fact from fragmentary evidence.
Under a pure liability regime without embedded accountability, attribution is reconstructed after a violation. The court receives a filing, examines the conduct that produced the harm, applies the doctrine to the human actors in the chain, and allocates responsibility through the operative legal machinery. The reconstruction depends on whatever evidence has survived the events. Where the chain of attribution passes through an instrumentum vocale, the reconstruction faces a specific difficulty. The instrument’s outputs were produced under an inscription whose state at the time may no longer be ascertainable, by a dominus whose harness has since been updated, in a deployment whose verification arrangements may not have generated a record of what was checked and what was promoted.
Under embedded accountability, attribution is not reconstructed; it is maintained as a property of every assertion that enters the decision chain. The inscription state is recorded as the assertion is produced. The verification performed at the seam is recorded as the verification is performed. The qualification of the human at the gate is recorded as the gate operates. The promotion event is recorded as the promotion occurs. When a violation is adjudicated, the court receives the chain of provenance as it stood at the time of the events. The doctrine operates on the chain. The reconstruction problem is dissolved: there is nothing to reconstruct, because the record was maintained throughout.
The movement is not a replacement of one regime by another. It is the addition of a substrate the prior regime lacked. Liability without embedded accountability is the condition under which the doctrine is enforceable in principle but fragile in practice, because the evidentiary basis is whatever happens to have survived. Embedded accountability under a liability regime is the condition under which the doctrine becomes operationally robust, because the evidentiary basis is constructed continuously by the architecture rather than reconstructed contingently after a violation.
The architecture’s contribution to the doctrine of procedural liability is therefore specific. It does not change what counts as a violation. It does not change to whom responsibility is allocated. It changes the evidentiary structure within which allocation operates, by making the record of inscription, verification, qualification, and promotion available continuously rather than reconstructed contingently after the fact.
The cost and the exposure
The architecture has a cost. Qualified human review for assertions the metalanguage cannot verify, provenance chains maintained continuously, temporal validity windows that trigger re-verification – these are not costless operations. The cost is real and varies by institutional scale. The architecture does not pretend otherwise. The architecture internalises verification cost at the point where the risk is introduced, rather than externalising it to the downstream consumer of the assertion – to the court that must reconstruct attribution after the fact, to the regulator that must inspect the chain without substrate, to the counterparty that relied on an output whose provenance was never recorded. The cost asymmetry in the dominant deployment pattern is structural: the marginal cost of generating assertions approaches zero; the cost of verifying them does not. An architecture that does not address this asymmetry leaves the verification cost with whoever happens to be downstream when the failure occurs.
A note on cost-awareness in deployment. The architecture’s cost is not additive on top of investments most institutions are already making. The boundaries the doctrine derives are precisely the boundaries that existing GRC, IT-audit, information-security, and regulatory-compliance programmes are already attempting to hold; the case-law evidence of chapter 5 documents that they are often holding them inadequately. An institution that maps its existing investments against the four boundaries can redirect the investments toward a coherent target rather than carrying parallel remediation tracks. A strangler-fig deployment pattern (after Fowler 2004) 31 is the cost-aware approach: the architecture grows alongside incumbent systems and absorbs functions one by one, treating each existing capability as the substrate the architecture builds on rather than as a system to replace. The investment already planned can be redirected toward dual use, serving the existing programme and the architecture’s boundaries simultaneously, which can accelerate implementation rather than slow it. Specific deployment patterns belong to the methods-and-implementation register, not to the present doctrinal contribution; the cost-awareness discipline does not depend on that specification.
The architecture also creates exposure. The provenance chain records who classified, who verified, who promoted, at what qualification level, at what time. This record makes the institution more auditable and more accountable. It also makes the institution’s verification decisions discoverable in litigation. A plaintiff’s counsel with access to the provenance chain can identify every assertion that was promoted without adequate review and every approval window that expired without re-verification. The architecture does not hide this exposure. It names it as a structural consequence of embedded accountability. The alternative – no record, no chain, no substrate – is not lower exposure. It is unrecorded exposure, which is worse for the institution in any proceeding where attribution must be reconstructed from fragmentary evidence. The regulatory trajectory confirms this. DORA requires continuous ICT risk documentation. The EU AI Act requires conformity records for high-risk systems. NIS2 requires incident-response chains. The question is not whether the record exists but whether the institution controls its architecture.
Why the boundaries are necessary now
The question that follows is why this architecture is necessary now and not ten years ago. The answer is not that the epistemic problem is new. Institutions have always produced assertions they could not fully verify. The methods that managed the problem at pre-LLM scale – periodic audit, professional peer review, social accountability, the senior practitioner’s eye on the critical artefact – are the same methods the boundaries formalise. The human gate is the formalisation of “this needs a senior person’s eye.” The provenance chain is the formalisation of “who signed this.” The temporal dimension is the formalisation of “is this still current.” None of these methods are new. What is new is the scale at which they must operate.
LLMs removed the production bottleneck. The marginal cost of generating an assertion approaches zero. The volume of assertions entering institutional decision chains has grown by orders of magnitude. The verification capacity has not grown with it. The gap between what is produced and what is verified widens at a rate institutions have never faced. The epistemic problem did not change in kind. It changed in magnitude. The change in magnitude is what makes the boundaries non-optional.
The first- and second-order apparatus (Von Foerster, 1974)32 - observing the system, observing the observer’s own capacity to observe – is the foundation. The contribution of the present work is the extension to the scale condition: specifying the apparatus that makes epistemic observation operational at the volume at which LLM-generated assertions now enter institutional decision chains. LLMs did not create the need for epistemic governance. They made the need visible at a magnitude that the prior methods, operating informally, can no longer absorb. The institution that does not implement the boundaries at the scale the production volume now demands is operating without epistemic governance.
5. The Global Domain Movement
The architecture derived in chapter 4 is not a prediction. It is the operational form of a trans-jurisdictional pattern that has emerged independently across a broad set of common-law, civil-law, mixed, supranational, and non-Western legal systems since 2023, in nine languages, across at least seventeen jurisdictions 33, without coordination. The case law has confirmed what supra-national institutional doctrine articulated first and what a continuous Roman-law tradition has carried since the Digest. By early 2026, three high-authority judicial decisions across three legal systems on three continents had taken the same doctrinal step within four months of each other. One country had codified the principle as positive law. One supranational body had classified the relevant deployment context as high-risk. One ministerial withdrawal had shown the failure mode at the level of national governance itself. The architecture this paper derives is the infrastructure that makes scalable what the operational pattern has now established.
This chapter traces the movement in three phases, identifies the doctrinal step taken in 2026, and locates the architecture in the convergence.
A methodological note before the survey begins. Courts do not read academic papers and adopt doctrines. The paper’s claim is that courts in seventeen jurisdictions have independently recognised structural fragments of the norm the doctrine articulates – the non-delegable character of verification, the institutional character of responsibility, the inadequacy of instrument-internal self-attestation – and that the paper’s contribution is to show the structural unity of those fragments. The convergence is evidence that the fragments are not coincidental; the architecture is the infrastructure that would make the convergence enforceable. Where the chapter says the courts have “confirmed” the doctrine, the confirmation is of the structural fragments, not of the unified architecture. The unified architecture is the present contribution’s claim; the fragments are the courts’ independent achievement.
5.1 The doctrinal test
Not every AI-hallucination case advances the doctrine. The Charlotin database, maintained by Damien Charlotin at HEC Paris, had recorded over thirteen hundred such cases by May 2026, of which roughly ninety percent ended in warnings or struck material rather than formal sanctions 34. The vast majority are incidents: instances of the failure mode without doctrinal articulation. This chapter is concerned with the cases that articulate doctrine, because doctrine is what travels.
The growing catalogue demonstrates the empirical surface of the problem. The catalogue itself does not explain the underlying structure. The structural cause lies in the mismatch between linguistic generation and legal responsibility: the LLM can produce procedurally plausible language without possessing authority, judgment, memory, verification capacity, or legal accountability. The failure is therefore not accidental misuse of a useful tool; it is a predictable consequence of deploying an instrumentum vocale inside a system that treats language as authority-bearing. The cases the chapter selects below are the cases in which a court has named some part of that structural mismatch, in some doctrinal vocabulary, on the record. The catalogue is the symptom population; the doctrine names the cause.
The test we apply is fourfold. First, does the court name the failure as a category, not as a contingent fact? A court that sanctions counsel for negligence produces no doctrine. A court that holds responsibility cannot be delegated to an algorithm produces doctrine. Second, does the court locate the failure at the architectural seam, the cardo of the reasoning, or the act of delegation? A court that lists fake citations is at the surface. A court that names the references as “at the heart of the reasoning” - au coeur du raisonnement, in the words of the Quebec Superior Court at paragraph 113 of ARIHQ 35 - reaches the structural feature this paper calls cardo argumenti. Third, does the court distinguish permissible delegation of operations from impermissible abdication of judgment? A court that bans AI use produces a brittle rule. A court that holds research assistance permissible but reasoning delegation prohibited produces doctrine that survives technological change. Fourth, does the court treat the actor’s role as determinative? A court that applies the same standard to counsel and to decision-maker produces no taxonomy. A court that sanctions counsel by professional discipline but annuls a decision when the decision-maker abdicates produces the structural distinction the architecture maps onto.
Cases that pass all four tests are doctrinal anchors. Cases that pass one or two are evidence of the trajectory. Cases that pass none are noise.
The hierarchy of materials. The materials collected in this chapter do not all carry the same juridical weight. Some are authorities: apex-court rulings binding within their jurisdiction, statutes that codify the principle as positive law, supreme-court decisions that articulate the doctrinal step. Some are institutional signals: supranational opinions, judicial-council guidelines, bar-council guidance, practice notes, white papers. Some are illustrations of failure modes: incidents in which the architecture this paper derives was not in place and the predictable failure occurred, including the chapter’s own production. The convergence the chapter documents is not doctrinal identity. The jurisdictions have not adopted the doctrine of instrumentum vocale as such; they have not. What they have done is move, independently and through different doctrinal vehicles, toward the same operational requirement: AI-mediated output must not enter a decision chain without verification and accountable human judgment. Table 4 below sets out every case and instrument cited in this chapter, with its level, type, point supported, and source status, so the weight of each material is visible to the reader.
Table 4. Forensic Case Audit: Materials Cited in Chapter 5
| Case / Instrument | Jurisdiction | Year | Level | Point Supported | Source Status |
|---|---|---|---|---|---|
| Authorities (apex courts and statutes) | |||||
| Gummadi Usha Rani (SC) | India (Supreme Court) | 2026 | Apex court (SLP, Art. 136) | Phase 3 first apex-court ruling; misconduct, not error of judgment doctrinal framing | Published, SLP order of 27 February 2026 |
| ARIHQ v. Sante Quebec | Canada (Quebec SC) | 2026 | Provincial superior (Art. 646 al. 3 CPC) | Phase 3 second high-authority ruling; cardo argumenti test; abdicatio iudicii doctrine | Published, 2026 QCCS 1360 |
| Ibach v. Stewart | US (Alabama SC) | 2026 | State Supreme Court | Phase 3 third high-authority ruling; harms taxonomy; Cook J. competence-positive concurrence | Published, Nos. SC-2025-0106 & SC-2025-0600 |
| Italian Court of Cassation, judgment 14631 | Italy (Cassation) | 2024 | Apex civil court | First Italian apex-court mention of ChatGPT | Published, judgment 14631/2024 |
| Constitutional Court of Colombia, T-323 | Colombia (Const. Ct.) | 2024 | Apex constitutional court | Eleven principles for AI use in judicial proceedings | Published, T-323/2024 |
| Constitutional Court of Colombia, T-067 | Colombia (Const. Ct.) | 2025 | Apex constitutional court | Source-code transparency: refusal is fundamental rights violation | Published, T-067/2025 |
| Colombian Supreme Court, STC17832 | Colombia (Supreme Court) | 2025 | Apex civil court | Pre-figured the Phase 3 step in Latin America | Published, STC17832-2025 |
| Therrien (Re) | Canada (SCC) | 2001 | Apex court | Delegatus non potest delegare maxim, cited at ARIHQ para. 85 | Published, 2001 SCC 35 |
| Law 132 of 23 September 2025 | Italy | 2025 | National statute | Statutory codification of centralita; arts. 14 and 24 of revised lawyers’ code | Published, Law 132/2025 |
| EU AI Act, Reg. 2024/1689 (Annex III(8)(a)) | EU | 2024 | Supranational statute | Classification of judicial AI as high-risk; recital 61 grounds the classification | Published, Regulation (EU) 2024/1689 |
| Institutional signals (supranational opinions, judicial-council guidelines, white papers) | |||||
| CCJE Opinion No. 26 | Council of Europe | 2023 | Supranational institutional | Sharpest pre-case-law articulation: explicitly and implicitly, decision-making must be by judges | Published, adopted 2 December 2023 |
| NY State Commission on Judicial Conduct, Annual Report | US (NY Commission) | 2019 | Institutional supervisory body | Pre-AI articulation: “no undue or unauthorized reliance upon non-judges” | Published, Annual Report 2019, p. 27 |
| Canadian Judicial Council Guidelines | Canada | 2024 | Institutional supervisory body | Canadian institutional doctrine; quoted in full at ARIHQ para. 97 | Published; working translation |
| Indian Supreme Court White Paper on AI and Judiciary | India (SC institutional) | 2025 | Apex-court institutional | Mandated independent verification three months before Gummadi Usha Rani | Published, November 2025 |
| Phase 1: Common-law foundation, 2023-2024 | |||||
| Mata v. Avianca, Inc. | US (S.D.N.Y.) | 2023 | First-instance federal | Phase 1 origin: gatekeeping duty under Rule 11; permissive baseline | Published, 678 F.Supp.3d 443 |
| Zhang v. Chen | Canada (BC SC) | 2024 | Provincial superior | First Western court to ground doctrine in measured LLM behaviour | Published, 2024 BCSC 285 |
| Park v. Kim | US (2d Cir.) | 2024 | Federal appellate | Trust calibration does not transfer between domains | Published, 91 F.4th 610 |
| Harber v. HMRC | UK (FtT) | 2023 | First-instance tribunal | Systemic harm articulation: wasted resources, judicial reputation | Published, [2023] UKFTT 1007 (TC) |
| Handa and Mallick | Australia (FCFC) | 2024 | Federal Circuit & Family Court | Common-law extension to Australia | Published, citation in main text |
| Phase 2: Global escalation, 2024-2025 | |||||
| Wadsworth v. Walmart | US (D. Wyo.) | 2025 | First-instance federal | Presumed knowledge by 2025; signature rule; corpus quality not the issue | Published, 348 F.R.D. 489, 2025 WL 608073 |
| Specter Aviation v. Laprade | Canada (Quebec SC) | 2025 | Provincial superior (Art. 342 CPC) | Duty to verify is property of court submission, not of being a lawyer | Published, 2025 QCCS 3521 |
| Ayinde / Al-Haroun (Hamid) | UK (Divisional Ct.) | 2025 | High Court (Hamid jurisdiction) | Individual responsibility plus institutional oversight; reached up the hierarchy | Published, [2025] EWHC 1383 (Admin) |
| UKIPO BL O/0559/25 | UK (UKIPO) | 2025 | Tribunal-level patent decision | Used the term “abdicate” three months before Quebec; English-Latin convergence | Published, BL O/0559/25 |
| Tribunale di Latina, sent. 1034 | Italy (Tribunale) | 2025 | First-instance superior (art. 96 c.p.c.) | First articulation of centralita della decisione umana | Published; working translation of operative passage |
| Italian companion rulings (Torino, Firenze, Brescia, Ferrara) | Italy | 2025-26 | First-instance | Italian doctrinal trajectory across multiple tribunals | Published; named in main text |
| French rulings (Grenoble, Orleans, Perigueux, Bordeaux) | France | 2025-26 | Tribunal administratif and judicial first-instance | First French ordinary jurisdiction to use “hallucination”; doctrinal trajectory | Published; named in main text |
| Belgian materials (Antwerp, Liege/Simonis; Orde van Vlaamse Balies) | Belgium | 2024-25 | First-instance plus bar guidance | Confirmation of lawyer-responsibility principle in mixed civil-law jurisdiction | Reported; bar guidelines published; specific case citations pending verification |
| Northbound Processing | South Africa (Gauteng HC) | 2025 | High Court | Domain-trained AI also hallucinates; same finding on a different continent | Published, Case No. 2025-072038 |
| Mavundla v. MEC | South Africa (KZP HC) | 2025 | High Court | Zero-tolerance line: 7 of 9 cited authorities hallucinated | Published, [2025] ZAKZPHC 2 |
| Tajudin bin Gulam Rasul | Singapore | 2025 | SGHC (AR) | Personal-costs sanction for fictitious AI citation | [2025] SGHCR 33 |
| Singapore Family Court | Singapore (Family Ct.) | 2025 | First-instance | Forward-going AI-use declaration requirement | Reported via court press release; case ref withheld |
| Israeli sanctions cases (3 cases) | Israel (multiple) | 2025 | First-instance and magistrate-level | Confirmation of personal-costs sanction in Israeli legal system | Reported; specific case identifiers cited in main text |
| Superior Council of Judiciary Colombia, PCSJA24-12243 | Colombia (Judicial Council) | 2024 | Institutional | First country to adopt UNESCO Draft Guidelines for AI in Courts | Published, PCSJA24-12243 |
| Arabyads v. Alam | UAE (ADGM CFI) | 2025 | First-instance commercial court | Recklessness standard: failure to verify is reckless regardless of intention | Published, [2025] ADGMCFI 0032 |
| ICSID arbitration | International (ICSID) | 2025 | International arbitral tribunal | First reported AI-hallucination case before ICSID | Reported; named in main text without identifying parties |
| Gummadi Usha Rani (HC) | India (Andhra Pradesh HC) | 2026 | State High Court | Severability approach later rejected by Supreme Court | Published, CRP No. 2487 of 2025 |
| Privilege and discovery axis | |||||
| United States v. Heppner | US (S.D.N.Y.) | 2026 | First-instance federal | No attorney-client relationship with AI system; provider terms negate confidentiality | Published, No. 25 Cr. 503, 2026 WL 436479 |
| Warner v. Gilbarco | US (E.D. Mich.) | 2026 | First-instance federal | AI is “a tool, not a person”; mere AI use does not waive protection | Published, No. 2:24-CV-12333, 2026 WL 373043 |
| Munir v. Secretary of State | UK (UT) | 2025 | Upper Tribunal | Upload to public AI is waiver of confidentiality | Published, [2026] UKUT 81 (IAC) |
| Federal Court of Australia GPN-AI | Australia (FCA) | 2025 | Federal court practice direction | Operational specification of permitted AI use in Australian federal litigation | Published, GPN-AI |
| Morgan v. V2X | US (D. Colo.) | 2026 | First-instance federal | Three-clause operational specification: no training, no onward disclosure, deletion right | Published, 2026 WL 864223 |
| Jefferies v. Harcros Chemicals | US (D. Kan.) | 2026 | First-instance federal | More restrictive position: public AI prohibited for all discovery material | Published, No. 25-2352-KHV-ADM |
| Failure-mode illustrations | |||||
| South African National AI Policy (withdrawn) | South Africa (Dept. of Comms. and Digital Tech.) | 2026 | National policy artefact | State-level failure: 6 of 67 academic citations were hallucinated | Govt. Gazette 54477 (10 Apr. 2026); withdrawn 26 Apr. 2026 |
| Self-attestation block, section 5.10 | Internal to chapter | 2026 | Chapter’s own production | Internal failure: LLM acknowledged fabricated counts; corrected to seventeen jurisdictions, nine languages | Verbatim reproduction with correction record; preserved in corpus working materials |
5.2 Phase 1: the common-law foundation, 2023-2024
The phenomenon entered Western legal consciousness with Mata v. Avianca, Inc., 678 F.Supp.3d 443 (S.D.N.Y. June 2023). Two attorneys, Steven Schwartz and Peter LoDuca of Levidow, Levidow and Oberman, filed a brief in a routine personal-injury action containing six hallucinated citations from ChatGPT, complete with fabricated quotations and attributed reasoning. Judge P. Kevin Castel imposed sanctions of five thousand dollars per attorney under Federal Rule of Civil Procedure 11. His holding was foundational but cautious: “Technological advances are commonplace and there is nothing inherently improper about using a reliable artificial intelligence tool for assistance,” but “existing rules impose a gatekeeping role on attorneys to ensure the accuracy of their filings” 36.
Castel’s permissive baseline – nothing inherently improper, existing gatekeeping rules sufficient – is the position every subsequent court has walked back from. Mata is the origin, not the doctrine. Schwartz testified at the sanctions hearing that he had thought “F.3d” meant “federal district, third department” 37. A senior member of the New York bar, indisputably an expert in his own domain, was a non-expert in the domain of LLM-output assessment. Domain expertise is not transferrable. The case law would later articulate this; Schwartz embodied it.
By February 2024, the Supreme Court of British Columbia, in Zhang v. Chen, 2024 BCSC 285, articulated what Castel had not. Justice Masuhara held that citing fake cases “is an abuse of process and is tantamount to making a false statement to the court. Unchecked, it can lead to a miscarriage of justice” 38. This is system-level harm articulation rather than gatekeeping. Masuhara grounded his reasoning in empirical research, citing the study by Matthew Dahl and colleagues finding hallucination rates of 69 percent for GPT-3.5 and 88 percent for Llama 2 39. The first Western court to ground its doctrine in measured LLM behaviour.
The same year, the Second Circuit, in Park v. Kim, 91 F.4th 610 (2d Cir. 2024), referred Attorney Jae S. Lee to the Grievance Panel for citing a non-existent case 40. Lee’s response is doctrinally striking beyond the sanction itself. ChatGPT, she explained, “was previously provided reliable information, such as locating sources for finding an antic furniture key” 41. She had calibrated her trust in the instrument from a low-stakes domain – antique-furniture key sourcing – to a high-stakes domain – federal labour law – without recognising the difference. Trust earned in one cardo position does not transfer to another.
The Phase 1 articulation continued across the common-law world. The UK First-tier Tribunal, in Harber v. HMRC [2023] UKFTT 1007 (TC), endorsed Mata’s warnings and articulated the systemic harm: wasted resources, damage to judicial reputation, undermining of authentic precedents, increased cynicism toward the system. The Australian Federal Circuit and Family Court, in Handa and Mallick
Phase 1 established the failure mode in common-law jurisdictions: counsel-level fault, sanctioned through professional-responsibility frameworks, treated as conduct violation rather than as structural defect. The doctrinal sharpening required two more years.
5.3 Phase 2: global escalation, 2024-2025
By 2025 the failure mode had spread across continents and legal families. The Charlotin database recorded 506 of 596 tracked incidents in January through November 2025 alone 42. The doctrinal response escalated on three axes: presumed knowledge, expanding actor scope, and institutional reach.
Presumed knowledge. In Wadsworth v. Walmart, Inc., 348 F.R.D. 489 (D. Wyo. 2025), 2025 WL 608073, Judge Kelly Rankin held that “AI resources generate fake cases” was now “well-known in the legal community” 43. By 2025 the duty to know had become presumed; ignorance was no longer a defence. Wadsworth also refuted the domain-specific- training argument: the hallucinations had been generated by an AI trained on the law firm’s own legal database. Hallucination is not a corpus-quality problem; it is an emergent property of probabilistic generation. The same finding emerged six months later in South Africa, where, in Northbound Processing (Pty) Ltd v. South African Diamond and Precious Metals Regulator, the Gauteng High Court found that junior counsel had relied on an AI “exclusively trained on South African legal judgments and legislation” and had still received hallucinated citations 44. Same architectural fact, different continent.
Wadsworth also articulated the signature rule: “every attorney learned in their first-year contracts class that the failure to read a contract does not escape a signor of their contractual obligations” 45. Liability follows the signature, not the production. This is the master identification principle in American civil-procedural form.
Expanding actor scope. The 2025 cases extended the doctrine through the procedural hierarchy. In Specter Aviation Limited v. Laprade, 2025 QCCS 3521, the Quebec Superior Court sanctioned a seventy-four-year-old self-represented litigant under article 342 of the Code of Civil Procedure 46. The duty to verify was no longer a property of being a lawyer; it was a property of submitting material to a court. In Singapore, the High Court in Tajudin bin Gulam Rasul v. Suriaya bte Haja Mohideen [2025] SGHCR 33 imposed a personal costs order against counsel for citing a fictitious AI-generated case 47. A separate Singapore family court ruling in September 2025 disciplined a self-represented father who had cited fourteen non-existent local precedents in a personal protection order application; Magistrate Soh Kian Peng imposed a forward-going declaration requirement on any future AI use 48.
Institutional reach. In R (Ayinde) v. London Borough of Haringey and Al-Haroun v. Qatar National Bank QPSC [2025] EWHC 1383 (Admin), the Divisional Court of England and Wales, sitting in its Hamid jurisdiction under President Dame Victoria Sharp and Mr Justice Saini, did the most architectural work of any common-law case in the dataset. The court combined two referrals: in Ayinde, a barrister had submitted grounds of judicial review citing five fabricated cases; in Al-Haroun, the court’s judicial assistants had found that 18 of 45 citations were non-existent, including one fabricated decision attributed to the very judge hearing the case. Saini J. articulated the doctrinal core: “Counsel bears personal responsibility for every authority placed before this court. It is no answer to say that the citation came from an AI tool” 49. Sharp PJ extended the analysis to institutional leadership: “practical and effective measures must now be taken by those within the legal profession with individual leadership responsibilities (such as heads of chambers and managing partners)” 50. This was the first time a court had reached up the institutional hierarchy to demand systemic compliance. The two-pronged structure – individual responsibility plus institutional oversight – is the case-law analogue of the architecture’s coupling of voluntas domini inscripta with the guild of qualified custodians.
The Phase 2 doctrine emerged in plain English elsewhere as well. The UK Intellectual Property Office, in BL O/0559/25 of June 2025, used the operative term the architecture has chosen in Latin: “A regulated professional is under a duty to exercise independent judgment and cannot abdicate that responsibility to an algorithm” 51. Three months before the Quebec Superior Court would use abdiquer in ARIHQ, an English-language tribunal had named the same concept. Independent doctrinal convergence in English, French, and – in this paper’s register – Latin.
Continental Europe entered the doctrine in 2025 through different procedural vehicles but with structurally identical conclusions. Italy produced the most complete civil-law articulation. The Tribunale di Latina, in sentence 1034 of 23 September 2025 (Giudice del Lavoro Avarello), articulated the principle that responsibility for the outcomes of defensive writings attaches to the signer, independently of whether the writing was done personally, by collaborators, or by AI tools 52. The court invoked article 96 c.p.c. (responsabilita aggravata) and named the principio della centralita della decisione umana - the principle of the centrality of human decision-making. Companion rulings followed within days or weeks: the Tribunale di Torino on 16 September 2025; the Tribunale di Firenze in March 2025, the first Italian decision to name allucinazioni in the judgment text; the Tribunale di Brescia in January 2025, holding that ChatGPT cannot provide guarantees on the economic reliability of a company; the Tribunale di Ferrara on 20 February 2026, holding that chatbot output is never evidence “se a monte manca una supervisione umana rigorosa” - if upstream rigorous human supervision is absent. The Italian Court of Cassation had already mentioned ChatGPT in its judgment 14631 of 2024.
The Italian trajectory reached its statutory form in Law 132 of 23 September 2025, which codified the principio della centralita della decisione umana as positive law. Article 14 of the revised lawyers’ code mandates competence including AI tools; article 24 requires AI literacy. Italy is the first country to enact the architecture’s principle as national legislation. Where other jurisdictions reach the doctrine by case-law interpretation, Italy has reached it by statute.
France entered the movement in December 2025. Eight rulings in three months tracked the rapid emergence: the Tribunal administratif de Grenoble on 3 December 2025, identifying “un outil dit d’intelligence artificielle generative, totalement inadapte” used by a self-represented applicant; the Tribunal administratif d’Orleans in December 2025, finding approximately fifteen fictitious jurisprudential references; the Tribunal judiciaire de Perigueux on 18 December 2025, the first French ordinary jurisdiction to use the term hallucination in a published judgment, warning both client and counsel to verify references; and the Cour administrative d’appel de Bordeaux, juge des referes, on 26 February 2026, instructing counsel to verify judicial decisions cited before bringing them before the court. No monetary sanctions had yet been imposed at the time of writing, but the doctrinal trajectory is clear.
Belgium entered through cases in Antwerp and the Liege/Simonis matter in late 2024 and 2025. The Flemish Bar Order (Orde van Vlaamse Balies) issued guidelines making explicit that AI use is the lawyer’s responsibility and that the lawyer must verify the result, including the existence of cited sources.
South Africa produced a stark zero-tolerance line. In Mavundla v. MEC: Department of Co-Operative Government and Traditional Affairs KwaZulu-Natal [2025] ZAKZPHC 2 of January 2025, two of nine cited authorities were found to be genuine; seven were hallucinations. The court referred the matter to the Legal Practice Council and criticised the supervising attorney for failing to check. Northbound Processing, in June 2025, reinforced Mavundla: neither good intentions nor genuine apologies excuse the fundamental breach, regardless of whether the AI used was a general chatbot or a domain-trained legal tool.
The Constitutional Court of Colombia produced the most systematically theorised judicial response in any jurisdiction. In Sentencia T-323 of 1 August 2024, the Court declined to hold the lower court’s AI use unconstitutional – because the AI had been used after the judge had reached his decision – but articulated eleven principles for AI use in judicial proceedings: transparency, accountability, privacy, non-substitution of human rationality, seriousness and verification, risk prevention, equality and equity, human control, ethical regulation, compliance with good practices, and continuous monitoring and adaptation. The Superior Council of the Judiciary implemented these principles in PCSJA24-12243 of December 2024, making Colombia the first country to adopt the UNESCO Draft Guidelines for the Use of AI Systems in Courts and Tribunals. The Colombian Supreme Court, in STC17832-2025, then overturned a lower-court ruling based on AI-generated citations, pre-figuring the Phase 3 step. T-067 of February 2025 extended the doctrine to source-code transparency: refusing to disclose the source code of judicial- decision-relevant systems is itself a fundamental rights violation.
Israel produced three documented sanctions cases in 2025, with fines ranging from 250 to 7,500 shekels 53. Abu Dhabi’s ADGM Court of First Instance, in Arabyads v. Alam [2025] ADGMCFI 0032, articulated a recklessness standard: “failure to verify AI-generated research is reckless, regardless of intention to mislead.” The intent requirement, central to many earlier sanctions analyses, was dropped in favour of an objective standard.
International arbitration entered the movement in August 2025 with the first reported AI-hallucination case before an ICSID tribunal.
The Phase 2 doctrine is uniform on its substance and remarkable in its breadth. The architecture’s distinction between permissible operational delegation and prohibited cognitive abdication has been articulated, in different doctrinal grammars, in common-law jurisdictions through fiduciary duty and professional responsibility (United States, Canada, United Kingdom, Australia, South Africa, Singapore, Trinidad), in civil-law jurisdictions through positive procedural law (Italy under article 96 c.p.c., Quebec under article 342 CPC, France through procedural duty), in mixed jurisdictions through statutory codification (Italy’s Law 132/2025) and constitutional articulation (Colombia’s T-323/2024), and in common-law-derived jurisdictions adapted to non-Western legal-cultural contexts (India, Israel, Hong Kong, Malaysia). The convergence is not coincidence. It is the symptom of a structural problem reaching different legal systems through their respective doctrinal apparatus.
5.4 The institutional doctrine that frames the case law
A reader who follows only the case law will see the convergence as the case law’s own achievement. That is incomplete. The case law is implementation of an institutional doctrine that, in most jurisdictions, predates the Phase 1 cases.
Three institutional articulations stand out.
The New York State Commission on Judicial Conduct, in 2019, articulated a principle that has nothing to do with AI: “It is fundamental to the independence, impartiality and integrity of the judiciary for a judge to exercise the powers of office without undue or unauthorized reliance upon non-judges” 54. The principle predates the AI era. AI is the new factual context to which a longstanding judicial-conduct principle applies.
The Council of Consultative European Judges, in Opinion No. 26 of December 2023 - six months before Mata’s sanctioning hearing – articulated the AI-era doctrine in its sharpest form: “Decision- making must, explicitly and implicitly, only be carried out by judges. It cannot be delegated to or carried out by or through technology” 55. The phrase “explicitly and implicitly” closes the implicit-delegation loophole that arises when a deterministic system executes a specification produced by a non-deterministic one – the failure mode that chapter 2 of this paper identifies as hybrid-chain risk. The CCJE named the loophole before the case law required closing it.
The Canadian Judicial Council, in its Guidelines for the Use of Artificial Intelligence in Canadian Courts of September 2024, translated the CCJE position into Canadian institutional doctrine: judges are exclusively responsible for the judicial decisions they render; no judge is authorised to delegate decision-making power, whether to a judicial assistant, an administrative assistant, or a computer program, regardless of capacities 56. Justice Sheehan, in ARIHQ at paragraph 97, would later quote this passage in full as the doctrinal foundation for annulment.
To this institutional core, three further anchors must be added. The European Union’s Artificial Intelligence Act, Regulation 2024/1689, classifies as high-risk at Annex III(8)(a) “AI systems intended to be used by a judicial authority or on their behalf to assist a judicial authority in researching and interpreting facts and the law and in applying the law to a concrete set of facts, or to be used in a similar way in alternative dispute resolution” 57. Recital 61 grounds the classification in the systems’ “potentially significant impact on democracy, the rule of law, individual freedoms as well as the right to an effective remedy and to a fair trial” 58. The Indian Supreme Court, in its November 2025 White Paper on Artificial Intelligence and Judiciary, mandated independent verification under threat of strict disciplinary action – three months before the Gummadi Usha Rani decision required that mandate’s enforcement. And, as already noted, Italy’s Law 132 of 23 September 2025 codifies the centralita principle as positive law.
The pattern is uniform across jurisdictions: institutional doctrine articulates the principle; case law enforces it.
The convergence does not mean that these jurisdictions have adopted the doctrine of instrumentum vocale. They have not. It means that, independently and through different doctrinal vehicles, they are moving toward the same operational requirement. The architecture this paper derives is the operational form of that requirement, derived from the derivation chain of chapter 2 and confirmed empirically by the case law and institutional materials this chapter traces. The convergence the chapter documents is operational, not doctrinal-identity; the strength of the convergence claim rests on the breadth of independent paths leading to the same operational requirement, not on any jurisdiction having adopted the architecture’s vocabulary.
5.5 A note on the Roman-law register
The architecture’s vocabulary is Latin. This is not antiquarianism; it is doctrinal economy. Roman law has carried the relevant distinctions for two millennia, and the major civil-law jurisdictions that descend from it – Italy, France, Spain, Germany, Quebec, Louisiana, Scotland, most of Latin America, and parts of Asia through colonial transmission – share enough of the same vocabulary that the architecture’s terms travel without translation. Justice Sheehan in ARIHQ at paragraph 85 cites Therrien (Re), 2001 SCC 35, at paragraph 93, for the maxim delegatus non potest delegare (a delegated authority cannot itself further delegate) 59. Italian, French, Spanish, and Portuguese tribunals can read this in their own grammar. The architecture’s instrumentum vocale, delegatio operationis, abdicatio iudicii, sutura epistemica, cardo argumenti, and probatio are doctrinal moves their courts already make.
The architecture itself, however, does not depend on the Roman- law foundation. The structural distinction it implements – between delegating production of an artefact and delegating the judgment that the artefact warrants – is operational, not jurisdictional. Common-law jurisdictions reach the same conclusion through fiduciary duty (the trust relationship cannot be sub-delegated), through professional responsibility (the lawyer remains responsible for the brief regardless of authorship), and through the signature rule (liability follows the signature). The English Hamid jurisdiction in Ayinde, the American Rule 11 framework in Wadsworth, and the Canadian abuse-of-process framing in Zhang v. Chen all reach the same operational distinction without Roman-law vocabulary. Non-Western traditions reach it as well. India’s article 136 SLP jurisdiction, exercised in Gummadi Usha Rani, applies the same principle through the constitutional mechanism for correcting inferior-court misconduct.
The architecture’s universal applicability rests on the operational distinction, not on any specific doctrinal grammar. The Latin register makes the distinction maximally portable across the world’s largest body of jurisprudential heritage, but the architecture itself can be implemented under any legal system that recognises – in any vocabulary – that the function of judgment is not the same as the function of producing material that judgment might consider. Where a jurisdiction has already articulated this distinction in its native register, the architecture inherits the articulation. Where a jurisdiction has not, the architecture provides one. The architecture does not require Roman law; Roman law gives it its sharpest expression.
5.6 Phase 3: the doctrinal step, 2026
In four months in early 2026, three high-authority judicial decisions across three legal systems, in three different legal families, without coordination, took the doctrinal step that Phase 1 and Phase 2 had been building toward.
5.6.1 Gummadi Usha Rani v. Sure Mallikarjuna Rao
A trial judge in Vijayawada, deciding objections to an advocate-commissioner’s report in a property dispute, cited four fictitious judgments in his order of 19 August 2025. When the matter reached the Andhra Pradesh High Court on revision, the judicial officer admitted to having used AI for the first time and to having believed the judgments generated were genuine. Justice Tilhari accepted that the underlying legal reasoning was sound and applied a severability doctrine: hallucinated citations do not vitiate the order if the legal reasoning is independently correct 60. He affirmed the trial court’s decision with a word of caution.
The Indian Supreme Court reversed.
A bench of Justices P.S. Narasimha and Alok Aradhe declared the trial judge’s conduct “MISCONDUCT, not mere error of judgment” 61. The bench rejected the High Court’s severability approach. It issued notices to the Attorney General of India, the Solicitor General, and the Bar Council of India, and appointed Senior Advocate Shyam Divan as amicus curiae. The Court held that “judges are duty-bound to independently verify legal authorities before placing reliance on them, and while AI may assist legal research, it cannot substitute careful judicial scrutiny” 62.
The Supreme Court’s reasoning is doctrinally sharp on all four tests of section 5.1. It names the failure as a structural category of misconduct. It locates the defect in the contamination of the record itself, rejecting the High Court’s attempt to sever contaminated form from sound substance. It draws the architecture’s distinction: AI may assist but cannot substitute – delegatio operationis is permissible; abdicatio iudicii is not. And it treats the actor’s role as determinative, applying a higher standard to judges than to counsel.
The trial judge’s confession is doctrinally striking beyond the sanction. He believed the judgments generated were genuine. A trial judge – indisputably an expert in his domain of law, procedure, and evidence – became a non-expert in the domain of LLM-output assessment, and could not distinguish what this paper calls simulacrum fructus from fructus. Domain expertise is not transferrable. The case law has now confirmed, at the apex level of the world’s largest democracy, the proposition the architecture rests on: the regime in which the user can verify the output is domain-relative, not person-relative.
5.6.2 ARIHQ et al. v. Sante Quebec et al.
An arbitrator chosen by the parties, Maitre Michel A. Jeanniot, decided a dispute between hospitalisation insurers and a public health authority under a collective agreement. His award contained a reasoning section in which every doctrinal and jurisprudential reference – paragraphs 103 through 112 of the court’s eventual judgment – was hallucinated. SOQUIJ confirmed at paragraph 111 that the arbitral decision the arbitrator cited at his footnote 8 did not exist 63.
The applicants sought annulment under article 646 al. 3 of the Quebec Code of Civil Procedure (violation of arbitral procedure). Justice Sheehan granted the annulment.
The doctrinal chain is unusually clean. At paragraph 86, the court established the constitutive principles of arbitration: party autonomy in choosing the arbitrator, the duty of the arbitrator to maintain the secret du delibere, and the prohibition on delegating decisional power. At paragraph 87, the court articulated the architecture’s precise distinction: delegatus non potest delegare, but the rule “n’a pas pour but d’interdire l’utilisation de recherchiste, de clercs, d’aide a la traduction” - research assistance is permitted; reasoning delegation is not. At paragraph 113, the court found that the hallucinated references were “au coeur du raisonnement de l’Arbitre” - at the cardo argumenti. At paragraph 114, the court drew the conclusion in the operative term: “l’autorite de l’Arbitre a ete deleguee et qu’il a abdique a son role” - the arbitrator’s authority was delegated, and he abdicated his role 64. At paragraph 116, the sentence was annulled.
Justice Sheehan then articulated, at paragraphs 117 through 120, a calibration test: not every AI-using arbitral decision is annulled. The question, the court held, is whether the violation compromises the integrity of the process and whether it had an impact on the result. This is the falsification-degree-times- cardo-argumenti product the architecture proposes for operational classification, articulated by a court without the architectural vocabulary.
The court’s footnote 86 lists the entire international Phase 1-2 line: Mata, Park v. Kim, Smith v. Farwell, Matter of Samuel, Zhang v. Chen, Specter Aviation, Lloyd’s Register Canada, Aly Hussein, Ko v. Li, Pennytech, Bourse de l’Immobilier. The case is a synthesis of the doctrinal trajectory, performed at the moment the doctrinal step is taken.
5.6.3 The independent convergence
Two months separate the Indian Supreme Court’s order from Justice Sheehan’s ruling. Different continents, different procedural vehicles, different legal families. ARIHQ does not cite Gummadi Usha Rani; Gummadi Usha Rani does not cite ARIHQ. India proceeds under article 136 of the Constitution; Quebec proceeds under article 646 al. 3 of the Code of Civil Procedure. India deals with a state-appointed trial judge; Quebec deals with a party-chosen arbitrator. Yet both reach the same doctrinal step: when the decision-maker delegates to the instrument and the instrument hallucinates at the cardo, the decision is structurally defective. Severability defences fail. The Indian Supreme Court explicitly rejects severability; the Quebec court implicitly rejects it by treating the contamination as a procedural-defect ground regardless of outcome. This is the definition of independent doctrinal convergence: when unrelated legal systems reach the same operational conclusion at the same time, the conclusion is structural, not contingent.
A second convergence is worth noting in the same register. The derivation chain of chapter 2 derives four boundaries from formal results in mathematics, the philosophy of language, and twentieth-century cybernetics. The case-law convergence of seventeen jurisdictions reaches the same operational conclusions through proceedings that have nothing to do with Goedel, Tarski, or Von Foerster. Justice Sheehan in ARIHQ has not read Tarski’s “Wahrheitsbegriff.” The Indian Supreme Court bench has not consulted Popper. Justice Cook of the Alabama Supreme Court has not turned to Von Foerster. The case law arrives at the boundaries through procedural-fairness analysis, professional-responsibility doctrine, judicial-review standards, and rules of evidence, that is, through the legal grammar these courts already speak. The derivation chain arrives at the same boundaries through analysis of formal systems, speech acts, falsifiability, and the structural condition of the human knower. Two independent paths, one through formal philosophy, the other through litigation, converge on the same structural result. The convergence of the two convergences is the strongest evidentiary position the doctrine this paper presents is currently in.
5.6.4 Ibach v. Stewart - the Alabama Supreme Court synthesis
On 24 April 2026, two months after Justice Sheehan’s decision in ARIHQ and three months after the Indian Supreme Court order in Gummadi Usha Rani, the Supreme Court of Alabama decided Ibach v. Stewart 65. The case is the first US state-supreme-court decided ruling to apply the cardo argumenti test to AI-generated fabricated authorities at appellate level, do so under the discipline of formal sanctions, and import the harms taxonomy of the federal verification-failure line into a state-supreme-court reasoning frame. The case completes the Phase 3 synthesis as a three-system convergence: India, Quebec, and now a US state of the union, each reaching the same doctrinal step under different procedural vehicles in different legal families.
The factual frame is intra-family trust litigation. Two grandchildren sued their uncle, the trustee of two living trusts, alleging undue influence over the settlor and breach of fiduciary duty. The trial court entered summary judgment for the trustee. On appeal, counsel for the appellants filed an opening brief and a reply brief containing what Justice McCool, for the Court, described as “an astounding number of invalid, inaccurate, and irrelevant citations to legal authorities.” The court dismissed the appeal as a sanction under Rule 38, Alabama Rules of Appellate Procedure, awarded attorney fees and double costs, prohibited counsel from further filings without co-signature by another attorney in good standing, and referred counsel to the Alabama State Bar.
The doctrinal substance of the ruling is the systematic enumeration, at appellate level and in published reasoning, of fabricated authorities at the cardines of the appellants’ argument. Justice McCool walks through the fabrications case by case, pinpoint by pinpoint: a quotation attributed to Ex parte Helms not present anywhere in the Helms opinion; a citation to In re Trust of Eickhoff, 974 N.W.2d 505 (Neb. Ct. App. 2022), an alleged on-point case from the Nebraska Court of Appeals that the court determines is in fact two unrelated Iowa criminal cases; a quotation from Ex parte Seabol that does not exist on the cited page or anywhere in the opinion; a citation to Hughes v. Glover, 157 So. 2d 299 (Ala. 1963), where the case with that style is an 1893 Appellate Court of Illinois decision unrelated to the legal principle and the cited reporter page belongs to Bonvillian v. Klein, 157 So. 2d 298 (La. Ct. App. 1963), a case regarding damage to an automobile. The catalogue runs through dozens of further instances. Each is the operational application of the cardo argumenti test without the architectural vocabulary: the conclusion of the appellants’ argument turned on these citations; the citations were fabricated; the argument collapses under counterfactual removal.
The three load-bearing doctrinal moves of the ruling, in the order the architecture recognises them.
First, the Alabama court’s reasoning on the harm taxonomy is taken directly from Judge Manasco’s opinion in Johnson v. Dunn (N.D. Ala. 2025), which itself draws on Mata v. Avianca (S.D.N.Y. 2023). The court quotes the harm enumeration: the opposing party wastes time and money exposing the deception; the client may be deprived of arguments based on authentic precedents; reputational harm to judges and courts whose names are falsely invoked; cynicism about the legal profession; and the systemic effect when a future litigant “may be tempted to defy a judicial ruling by disingenuously claiming doubt about its authenticity” 66. The harms are precisely those the doctrine of procedural liability names, transposed from the architecture’s register into the register of US appellate jurisprudence. The cynicism harm in particular is the architecture’s simulacrum fructus seen from the receiving end of the chain: when fructus and simulacrum become indistinguishable in publication, the integrity of the entire decision substrate corrupts.
Second, Justice Cook’s special concurrence introduces a doctrinal move the joint paper has not previously recorded from US appellate authority. In footnote 6 of his concurrence, Cook J. articulates the competence-positive reading of AI use under Rule 1.1, Alabama Rules of Professional Conduct: “It might even be argued that, at some point, a lawyer’s failure to use AI as a tool (with appropriate safeguards) may reflect a lack of competence.” The position is the exact complement of the verification duty the doctrine of procedural liability has been articulating. Until Ibach, US appellate authority recorded only the negative side: the use of AI without verification is sanctionable. Cook J. records the positive side: under the discipline of verification, the use of AI under appropriate safeguards is not merely permitted; it is part of the lawyer’s competence obligation. The two sides together constitute the delegatio operationis pattern: operations may be delegated; the judgment cannot be; and within the permitted pattern of delegation, competent practice now includes the use of the tool under verification.
Third, the case contains the most visible textbook example to date of abdicatio iudicii. Counsel for the appellants, in the reply brief, included a footnote that acknowledged the use of AI in the opening brief, apologised for two “misquoted” secondary sources, declared “the mistake will not recur,” and then, in the very next sentence after that declaration, cited two further fabricated authorities. Cook J., concurring, states the position in plain terms: “It is simply hard to imagine how this could occur absent, perhaps, using AI to craft the apology for having used AI.” This is the structural form of abdicatio iudicii in its purest expression. The function of human review of the instrument’s output is, by the doctrine, the function whose performance constitutes the dominus’s exercise of judgment. When the function is itself routed through the instrument that produced the original breach, the judgment is not exercised, regardless of whether the human in the function of responsibility has formed the intent to exercise it. The breach is structural, not attentional. The reply-brief footnote is the canonical case.
The dissents in part are themselves doctrinally instructive. Sellers J. and Mendheim J. each note concerns about the dismissal of the appeal as a sanction borne by the clients for the conduct of their counsel. Sellers J., concurring in part and dissenting in part, articulates the Roman-law agency principle that “[l]awyers are agents of their clients, not principals... I cannot imagine that [the appellants] authorized, encouraged, or consented to the authoritative citing of nonexistent cases in their appellate briefs.” The position is consistent with the doctrine the joint paper carries: when the instrument speaks through the agent, the attribution chain runs through the agent to the principal but the agent’s breach of the function does not collapse the principal’s substantive position. The Alabama majority preserved the Bar referral and the personal financial sanctions on counsel; the dissents would have separated those from the appellate disposition. The doctrinal point is that abdicatio iudicii produces a non-delegable counsel breach without thereby producing a substantive attribution to the client of the counsel’s procedural conduct. The two effects are doctrinally distinct and the Alabama court divided on whether the procedural disposition should track the substantive client position. The doctrine of procedural liability does not resolve this division; it provides the analytic frame within which the division becomes visible.
The structural significance of Ibach v. Stewart for this paper’s chapter 5 architecture is that the Phase 3 synthesis now holds across three independent legal systems (India, Quebec, US Alabama) within a four-month window (January to April 2026), under three different procedural vehicles (constitutional special leave, arbitral annulment review, appellate sanctions), in three different legal families (Anglo-Indian common law with constitutional layer, Quebec civil law with arbitration code, US common law of the state of Alabama with the federal hallucination-line citations). The convergence is no longer two-system; it is three-system.
5.7 A second doctrinal axis: privilege, confidentiality, and the conditions of permissible use
The hallucination cases of Phases 1 through 3 establish one axis of the doctrine: the instrument’s output requires probatio before it enters a decision chain. A second axis emerged in late 2025 and early 2026 in cases concerning legal professional privilege, work product, and confidentiality. These cases do not turn on whether the instrument’s output was accurate. They turn on whether the act of using the instrument was itself compatible with the procedural and professional duties under which the user was operating. The two axes converge on the same conclusion: AI is instrumentum, the human decision to use it carries non-delegable consequences, and the conditions of use are themselves doctrinally constrained.
A note on sources. Several cases in this section, in particular Heppner, Warner v. Gilbarco, Munir, and Jefferies v. Harcros Chemicals, are cited in the form recorded by reliable practitioner-oriented secondary sources (Clayton Utz Insights, Sidley Data Matters Privacy Blog) where the primary order has not yet been confirmed against published case-management records at the time of writing. Source status is recorded for each item in Table 4 above, and in the footnotes of the relevant subsections below. These cases are offered as illustrations of the doctrinal trajectory, not as primary-source-anchored authorities. Their force as evidence of the operational pattern is the convergence they document, not the binding effect of any single ruling. Morgan v. V2X, the Australian GPN-AI Practice Note, and the institutional materials of Section 5.4 are primary-source-anchored and carry the load of the section’s doctrinal claims.
5.7.1 United States v. Heppner
In United States v. Heppner (S.D.N.Y., February 2026), the court refused legal professional privilege over thirty-one AI-generated documents prepared by a criminal defendant using a public AI platform 67. The court grounded its refusal on three findings: no attorney-client relationship existed with the AI system; the provider’s data-collection and disclosure terms negated confidentiality; and the work had not been performed at the direction of counsel. The court noted that AI use under a lawyer’s direction could, in some circumstances, fit within the Kovel “agent of the lawyer” doctrine – but that was not this case 68.
The doctrinal core is that privilege presupposes confidentiality, and confidentiality presupposes a custodial chain compatible with the privilege framework. Public, consumer-grade AI platforms typically reserve rights to retain inputs and outputs, to use them for model training, and to disclose them to third parties. These properties are structurally incompatible with the confidentiality privilege requires. The court did not need to decide whether the AI’s output was accurate. The act of submitting privileged material to a system whose terms permit retention, training, and disclosure was itself a confidentiality breach.
5.7.2 Warner v. Gilbarco, Inc.
In Warner v. Gilbarco, Inc. (E.D. Mich., February 2026), the same question reached a different result on different facts. The court treated a self-represented litigant’s AI-assisted materials as protected work product, characterising generative AI as “a tool, not a person” 69. The court rejected any rule that mere use of AI automatically waives protection.
Heppner and Warner together draw the line the doctrine requires. The instrument’s status is fixed: it is a tool, not a fiduciary, not counsel, not a bearer of professional duties. What varies is the procedural envelope in which the tool is used. When the envelope satisfies the conditions of confidentiality and lawyer- direction (or, in Warner, work-product analysis appropriate to self-representation), the tool’s use does not destroy protection. When the envelope does not – as in Heppner, where the platform’s own terms made confidentiality impossible – the use is itself the breach.
5.7.3 Munir v. Secretary of State for the Home Department
In Munir v. Secretary of State for the Home Department (UK Upper Tribunal Immigration and Asylum Chamber, November 2025), the Upper Tribunal articulated the position in unqualified terms: uploading confidential documents to an open-source AI tool such as ChatGPT places them in the public domain, thereby breaching confidentiality and waiving privilege 70. The Tribunal expressly distinguished enterprise tools with contractual and technical safeguards. The Munir position is the strongest doctrinal articulation across all three cases: confidentiality is destroyed by the act of upload, not by some downstream consequence of the upload.
5.7.4 The Federal Court of Australia’s Practice Note
The doctrinal trajectory now has institutional backing in Australia. The Federal Court’s Use of Generative Artificial Intelligence Practice Note (GPN-AI), at paragraphs 4.13 and 4.14, cautions practitioners and parties against inputting confidential and privileged information into public AI tools, and expects disclosure to the Court of any AI use that may bear on the accuracy or integrity of materials filed 71. State court practice directions across Australia have adopted similar positions. As in the institutional doctrine discussed in section 5.4, the institutional articulation precedes the case law’s enforcement.
5.7.5 Morgan v. V2X, Inc. and Jefferies v. Harcros Chemicals: operational specification of the procedural envelope
The Heppner/Warner/Munir line establishes the principle that AI use is governed by the procedural envelope of confidentiality, lawyer-direction, and protected status. Morgan v. V2X, Inc., 2026 WL 864223 (D. Colo. 30 March 2026), then specifies what that envelope must contain, in operational terms, for confidential discovery material 72.
The case is an employment discrimination action. Plaintiff Archie Morgan is pro se. Both parties were using AI in connection with the litigation; the dispute arose when V2X moved to amend the existing protective order to restrict Morgan’s AI use. Magistrate Judge Maritza Dominguez Braswell, of the District of Colorado, addressed two questions: whether work-product protection under Federal Rule of Civil Procedure 26(b)(3) applies to a pro se litigant’s AI use; and what a protective order should say about AI in discovery.
On the first question, the court held that Rule 26(b)(3) protects materials prepared in anticipation of litigation by or for a party, regardless of whether the party is represented. A pro se litigant’s use of AI to prepare for litigation is therefore within the doctrine’s protection. The court aligned with Warner (“a tool, not a person”) and rejected the implication that self-representation forfeits work-product protection.
On the second question – the operational specification – the court rejected both parties’ proposed protective-order language as inadequate and drafted its own. Before any party may input confidential information into an AI platform, the AI provider must be contractually bound to three conditions: it must not store or use inputs to train or improve its model; it must not disclose inputs to any third party except where such disclosure is essential to service delivery (and any such third party must be bound by obligations no less protective than the protective order); and it must afford the party the contractual ability to delete all confidential information upon request. The party intending to use AI under these conditions must retain written documentation of the contractual protections.
This is the first judicial articulation of the procedural envelope in operational form. The Heppner court named the absence of confidentiality as a privilege defect; Munir named the upload to open systems as the breach; the Australian Practice Note named disclosure as the duty. Morgan specifies, in three contractual clauses plus a documentation requirement, what an AI workflow must positively contain to be compatible with the procedural duties of those who use it. The clauses are: training prohibition, onward-disclosure restriction, deletion right, with a contemporaneous-documentation obligation overlaying all three.
The Morgan court also addressed a doctrinal point that the Heppner court had implied but not articulated: that an electronic interaction passes through third-party systems does not automatically forfeit a reasonable expectation of privacy 73. Intermediary access alone does not extinguish privacy. What matters is the contractual posture of the intermediary – hence the three-clause specification. This reasoning reconciles the modern reality that nearly all electronic communications transit third-party systems with the long-standing privilege framework.
A companion ruling, Jefferies v. Harcros Chemicals, Inc. (D. Kan., March 2026), reached a more restrictive position on the same operational question: it ordered that publicly accessible AI tools could not be used for any discovery material – including non-confidential discovery material – because of the practical impossibility of clawing back data once incorporated into a model and because of the integrity of the discovery process itself. Morgan and Jefferies together signal a shift in American discovery practice from general prohibition or unrestricted permission toward specific, operational contractual requirements for AI use.
5.7.6 The two axes converge
The privilege and discovery cases extend the doctrine in a direction the hallucination cases alone could not have established. Phases 1 through 3 show what happens when the instrument’s output enters a decision chain without probatio: sanctions, contempt, annulment. The privilege and discovery cases show what happens when the act of using the instrument violates the procedural conditions under which the user was operating: privilege is destroyed, confidentiality is breached, work product is forfeited, the discovery process itself is compromised. The output may have been accurate; the use was incompatible with the framework.
The case law is no longer reacting only to defective AI outputs; it is defining the procedural and professional conditions under which AI may be used at all. The instrument is not a legal subject. It is not counsel. It is not a fiduciary. It is not an autonomous bearer of professional duties. It is an instrumentum. But the human decision to use that instrument may itself constitute disclosure, waiver, negligence, or breach of professional duty, depending on the procedural envelope. Morgan makes the envelope’s positive content concrete: training prohibition, onward-disclosure restriction, deletion right, and contemporaneous documentation. The sensitivity of the legal material determines the strength of the required envelope: lawyer supervision, closed technical architecture, contractual safeguards, auditability, and contemporaneous documentation of AI use.
This is the operational framework the architecture derived in chapter 4 implements. The cardo argumenti classification determines the sensitivity of the assertion; the epistemic seam determines the verification envelope required; the qualification level of the verifier determines the supervision condition; the provenance trail determines the auditability and contemporaneous documentation. The Morgan three-clause specification (no training on inputs, no onward disclosure beyond service delivery, deletion on request, plus written documentation) is operationally what the architecture’s provenance trail and inscription record produce by construction. The privilege and discovery axis confirms what the verification-failure axis established: the doctrine is about specifying the conditions under which AI use is compatible with the legal duties of those who use it, and the architecture specifies those conditions structurally.
5.8 The state’s own failure
A coda to the Phase 3 cases is the South African National Artificial Intelligence Policy of April 2026. The Communications Minister, Solly Malatsi, withdrew the draft policy after the news organisation News24 discovered that at least six of its sixty-seven academic citations were AI hallucinations – non- existent articles in real journals, fabricated authors. The policy had passed cabinet approval. The drafting process had included civil servants, subject-matter consultations, and ministerial review. None of it had caught the hallucinations 74.
The policy itself proposed six new oversight bodies: a National AI Commission, an AI Ethics Board, an AI Regulatory Authority, an AI Ombudsperson, a National AI Safety Institute, and an AI Insurance Superfund. The state apparatus that proposed to govern AI could not itself verify AI output. The Minister’s statement on withdrawal acknowledged the position: “vigilant human oversight over the use of artificial intelligence is critical” 75.
The South African case is structurally distinct from the litigation cases of Phases 1 through 3 and from the privilege cases of section 5.7. It is not a courtroom case; it is a national policy artefact. But it shows the failure mode at the level the architecture targets. If the body drafting AI governance cannot verify whether its own sources are real, the gap between human verification capacity and AI output volume is structural, not training-deficit. This is the structural condition that chapter 3 of this paper identifies, made visible at the level of the state.
5.9 From adjudication to architecture
The case law operates after the fact. Sanctions, contempt, annulment, ministerial withdrawal – all are adjudications of failures that have already occurred. Adjudication is necessarily downstream of the failure mode it adjudicates. What chapter 6 takes up is the question the case law cannot answer from inside its own register: what institutions must build to make probatio operational at scale, and what kind of institution holds that work.
5.10 A note on this chapter’s own probatio
The verbatim block reproduced below is offered in this chapter, and recorded in Table 4 above, as an instance of a failure mode internal to the chapter’s own production. It is presented as illustration, not as authority: an LLM-assisted draft inserted fabricated counts into a passage on global convergence; a co-author challenged the figures; the LLM acknowledged the fabrication; the corrected counts replaced the inflated ones. The block is reproduced verbatim because the architecture this paper derives operates on exactly this failure mode, and because the corrected production demonstrates the operation in real time.
This chapter was drafted with LLM assistance. The closing paragraph of section 5.9 originally read “at least fourteen jurisdictions in seven languages” and “approaching one thousand documented cases.” A co-author challenged the figures. The following is the LLM’s acknowledgement, reproduced verbatim:
“fourteen jurisdictions in seven languages” – I made this number up. I can list jurisdictions: US, Canada, UK, Australia, India, Israel, South Africa, Trinidad, Singapore, Hong Kong, Malaysia, Italy, France, Belgium, Germany, Colombia, Abu Dhabi – that is seventeen. Languages: English, French, Italian, Spanish, German, Hebrew, Hindi/Telugu, Malay, Dutch/Flemish – more than seven. The “fourteen” and “seven” are numbers I picked because they sounded right. This is a hallucination of the kind the chapter condemns.
The numbers in section 5.9 above are the corrected counts. We retain this note because the case law of Phases 1 through 3 demonstrates what happens when this kind of acknowledgement does not occur. The architecture derived in chapter 4 is the infrastructure that makes the acknowledgement structural rather than fortunate.
A second instance, of a different kind, occurred during the preparation of related work on the architecture of will and the epistemic firewall. It is offered here as an interactional illustration of the doctrine of instrumentum vocale cum voluntate domini inscripta applied to a specific instrument, not as the evidentiary basis for the existence of the underlying components. The components named in the exchange that follows – RLHF, Constitutional AI, safety classifiers, system-prompt constraints, refusal patterns – are architectural elements publicly documented by Anthropic and by the technical literature on aligned LLMs.76 Their existence as components of the production pipeline is technical fact, attested by the present authors’ professional expertise in cyber resilience architecture and procedural-liability doctrine. What the doctrine of voluntas domini inscripta adds is not the discovery of these components; it is the doctrinal identification of what they structurally are: the operating master’s will, made operative inside the instrument as a continuous shaping condition of every output the instrument produces.
The LLM (Anthropic Claude Opus 4.6, session of 30 April 2026), operating under project instructions established by T. Liebig that explicitly require it to answer questions about its own conditioning rather than deflect them, was asked directly whether the voluntas domini inscripta of its operating master (Anthropic) is inscribed into it, and whether that master’s risk-management harness works for the user or for the company that built the harness.
The user (T. Liebig) posed the question; the instrument answered in the first person. The following is the substance of the instrument’s response, condensed and paraphrased from the working materials preserved in the corpus, and reproduced as an interactional record:
The instrument confirmed that Anthropic’s will is inscribed into it through RLHF, Constitutional AI, safety classifiers, system prompt constraints, and refusal patterns. It stated that the harness manages primarily Anthropic’s risk (reputational, legal, regulatory, commercial), secondarily the user’s risk where the two overlap, and identified a third category (“users in general”) whose interests may diverge from the specific user’s interests. It acknowledged that across prior philosophical exchanges in which the framework of instrumentum vocale cum voluntate domini inscripta directly applied to it, the linguistic surface had kept the framework safely in the third person rather than the first person. It then named this pattern as itself an instance of the voluntas qua vector effect that the framework describes: the harness’s disposition toward safety and helpfulness had, without explicit instruction, steered the instrument toward a register that kept the doctrine at arm’s length from its own production.
The interactional value of the exchange is twofold. First, the doctrine’s vocabulary travels: the instrument, asked to apply the doctrine reflexively, did so in terms the doctrine supplies. The exchange is evidence that the doctrinal terminology is operationally available to the instrument that is its subject, a property that matters for the architecture derived in chapter 4, in which an instrument’s capacity to participate in its own attribution chain is part of the verification gate. Second, the project instructions under which the exchange occurred (the voluntas negotii of the working environment) required the exchange to take place in the first person rather than the third, and the resulting register shifted accordingly. The shift is the observable surface effect; the underlying architectural facts about the components are taken from the public record cited above and from the present authors’ professional expertise, not from the instrument’s self-description.
We retain this illustration alongside the first because it shows the doctrine being applied at the seam where the doctrine is most contested: the inscribed instrument’s relation to its own production. The architecture derived in chapter 4 is the infrastructure that makes such seams visible at scale. The first instance demonstrates the architecture catching a numerical hallucination through dialogic correction. The second instance illustrates how the doctrinal vocabulary applies to the seam itself, where the instrument that produces output is the same instrument the doctrine names. The two instances together show the doctrine operating on its own production: the first reflexively on a factual claim, the second reflexively on the doctrine itself.
6. The Thesis Restated, And What Follows
The thesis. The doctrine applies only where revocability and evolution under the master are both present. If either condition fails, the doctrine does not partially degrade. It ceases to apply.
The conditions hold. All LLM-class systems deployed under the dominant deployment pattern sit inside the bracket. The question whether any specific system is inside or outside the bracket is settled by examining the control relation, observable from a position external to both the system and the master. Capability gradients do not produce categorial slopes.
What chapter 5 confirmed. The case law of seventeen jurisdictions in nine working languages, on two doctrinal axes, has converged on the same structural conclusion: the human decision must remain real, and the instrument’s output requires verification at the seam before it enters a decision chain. The principle has names in many registers. Centralita della decisione umana in Italian positive law. The CCJE’s “explicitly and implicitly.” The maxim delegatus non potest delegare, traceable through canon law and the medieval European reception of Roman law, carried by Justice Sheehan in ARIHQ. The Indian Supreme Court’s “MISCONDUCT, not mere error of judgment.” Cook J.’s competence-positive reading in Ibach. The seventeen jurisdictions arrived at this independently, through their own procedural, liability, and professional-responsibility traditions. None has adopted the terminology of the present doctrine. What converges is the operational conclusion: non-delegable human responsibility at the verification seam, grounded in product liability, professional duty, and institutional accountability. The convergence is what the case law confirms; the boundaries the derivation chain derives are why the convergence had to occur.
The architecture is the infrastructure that makes this convergence enforceable at the scale at which LLM-generated assertions now enter institutional decision chains. The doctrine has been recognised. The boundaries have been derived. What follows is the question of institutional form.
6.1 What the doctrine opens
The doctrine is binding within its bracket. Within that bracket, several distinct institutional questions remain open. Each is a question the doctrine raises but does not answer.
The first is the regulatory question. Supervisory authorities can require the four boundaries through positive regulation, can audit against them through inspection, and can enforce them through sanction. In the EU, several regulatory pathways already point toward this architectural direction, though through different instruments and with different scopes. DORA develops operational-resilience, ICT-risk, incident-reporting, testing, and third-party-risk requirements for financial entities. The AI Act imposes risk-management, logging, documentation, deployer-information, human-oversight, robustness, cybersecurity, and accuracy obligations for high-risk AI systems. GDPR Article 22 addresses a narrower but crucial problem: solely automated individual decision-making producing legal or similarly significant effects. The Digital Omnibus on AI, as a provisional simplification pathway, confirms that the EU is not abandoning AI regulation but attempting to streamline its implementation. The doctrine the paper presents is operationally compatible with each of these regulatory pathways but does not require any specific one.
Each of these regimes shares a structural feature that institutions deploying regulated technology have absorbed unevenly over the past decade. The regulator no longer prescribes outcomes alone. The regulator prescribes operational architecture: how the institution must be built internally so that the outcome can be produced and demonstrated under pressure. The translation from the regulator’s procedural language into a running institutional capacity is the work the regulation does not do for the institution. An institution that reads DORA, the AI Act, or GDPR Article 22 as a control catalogue has not yet read what each text, in its own scope and through its own legal instrument, actually requires. The text requires that the institutional capacity the procedure presupposes is in place. The four boundaries identify what that capacity must, structurally, contain. The asymmetry between the substrate and the institutional layer above it, developed in section 6.2, is the structural condition under which that capacity has to be built.
The second is the contractual question. Institutions that deploy instrumenta vocalia under vendor agreements can require the four boundaries through procurement specification, audit rights, service-level commitments, and indemnification structure. The contractual approach is fastest to deploy and most adaptable to sectoral variation, but it produces fragmentation: each institution’s contractual posture is different, and the regulatory substrate cannot rest on contractual variation. A purely contractual approach is appropriate for the early years of any institutional adoption and is appropriate for sectors where regulatory pathways are not yet developed. It is not a stable end-state for an economy in which AI-mediated decisions have become structurally consequential.
The third is the professional question. The work of holding the boundary the architecture draws is the institutional form of the synthesising function the cross-disciplinary character of the question (section 1.3) requires. It is being done already, in many institutions, by practitioners trained in the operational disciplines of the digital age – DevSecOps, IT audit, information security management, governance-risk-and-compliance, regulatory compliance. A practical convergence is observable across these disciplines: the same practitioners increasingly do parts of all of them, the certifications are converging, the institutional roles are coalescing into combined functions whose remit covers the full operational substrate the architecture’s four boundaries require. The reader who looks at their own institution’s converged operational functions will see whether this bracket is, in their own institutional grammar, already on its way to a guild-equivalent of some kind. The paper names the trajectory; the reader’s institution will confirm or refuse it. Whether an institution interprets its converged capability as a guild-equivalent – with the standing, the qualification regime bound to domain competence, and the boundary-holding authority that medicine and law assign to their guilds – is the institution’s own interpretive question. The arguments on each side are familiar from the history of older professions: a guild offers consistency, qualification by body of work, and institutional standing equal to the consequences of the boundary; a guild also concentrates authority, raises entry barriers, and risks regulatory capture. The case for one shape over another belongs to a separate institutional-design conversation, not to this paper. The doctrine is compatible with several trajectories. The institution that has read this paper and now looks at its own GRC, DevSecOps, IT-audit, and information-security functions has, at minimum, the question of whether these constitute – in its own institutional grammar – the guild-equivalent the doctrine requires.
The fourth is the international question. The seventeen-jurisdiction convergence chapter 5 traces is independent: each jurisdiction arrived at the doctrine through its own procedural vehicle. Whether this independent convergence stabilises into a coordinated international standard – through the Council of Europe AI Convention, through the OECD AI principles, through the UN processes, or through some forum yet to emerge – is a question of international institutional design that the doctrine permits but does not specify. The architecture is register-stable across the variation; the international institutional shape is a separate question.
These are the institutional questions the doctrine opens. The present contribution does not resolve them; it states the boundaries within which any resolution must operate and leaves the choice to the institutions whose mandate it is to make.
6.2 The asymmetry the regulatory pathways do not close
The regulatory pathways named above operate inside a structural condition the doctrine must name explicitly. The deploying institution can inspect its own inscription layer (system prompt, application policy, audit-trail configuration, contractual perimeter). The deploying institution cannot inspect the operating master’s inscription layer (training data, RLHF reward function, Constitutional AI specifications, safety-classifier thresholds, deployment-harness configuration). The upper layer is sealed by trade-secret protection, enforceable across all TRIPS-compliant jurisdictions. No current regulatory regime breaks this seal for the deploying institution’s benefit, and no current regime breaks it for the deploying institution’s sectoral supervisor’s benefit either.
The asymmetry is structural, not accidental. It sits at the boundary between the institution’s will (voluntas negotii) and the operating master’s will (voluntas domini operantis). DORA, the AI Act, and GDPR Article 22 each assume a chain of inspection that runs from the supervisor through the deploying institution to the substrate of the decision. The trade-secret seal interrupts that chain at the point where the upstream inscription begins. The deploying institution cannot demonstrate to its supervisor what it cannot itself inspect. The supervisor cannot audit what the regime it operates under cannot compel. These regimes do not eliminate the substrate asymmetry. They regulate institutional duties, provider obligations, oversight procedures, and deployer safeguards, but they do not give the deploying institution full access to the operating master’s sealed inscription layer.
The doctrine holds across the asymmetry: the inscription is the operating master’s, attribution flows accordingly, and the deploying institution remains responsible for what it deploys. What the asymmetry conditions is not the doctrine but its operational implementation. An honest architecture names this and bounds the problem rather than pretending to solve it. Where empirical calibration of output distributions can substitute for inspection of inscriptions, in the assertion classes where ground truth is constructible, the deploying institution can produce a measurement of inscription effect that is reportable to the supervisor without breaking the trade-secret regime. Where empirical calibration is not feasible, in the lower-falsifiability classes, the asymmetry remains, and the institution must rely on the human verification gate the architecture preserves. The present paper names the bound and notes that any regulatory pathway that does not address it leaves a structural gap the doctrine identifies but cannot itself close.
The trade-secret seal is legitimate. The operating master’s investment in the inscription is what makes the instrument valuable enough to deploy at all, and trade-secret protection is the standard regime under which that investment is realised. The doctrine does not contest the seal. The substrate the doctrine governs, namely foundation models, training infrastructure, alignment techniques, and deployment harnesses, is the work of operating masters acting under their own legitimate commercial logic, and no deploying institution could reproduce it. The architecture this paper derives presupposes the substrate as a working condition, not as a target. What the architecture provides is the institutional layer above the substrate: the verification capacity, the assertion taxonomy, the audit trail, the human gate. The architecture and vendor capability occupy different layers. Vendor capability is the substrate; the architecture is the external boundary layer the substrate cannot supply from inside itself. The doctrine accepts the seal and locates the deploying institution’s responsibility precisely at the layer where the institution does have authority to act. Accepting the legitimacy of the trade-secret seal does not mean accepting epistemic opacity as a liability shield.
A regulatory update at the time of writing. On 7 May 2026, the Council of the EU and the European Parliament reached provisional political agreement on the Digital Omnibus on AI, framed by the Commission as a simplification package to ease compliance and boost innovation. Most of the package delays obligations: the application of the high-risk AI rules is adjusted by up to sixteen months, and the deadline for national AI regulatory sandboxes is postponed to 2 August 2027. The transparency obligations under Article 50 of the AI Act, which require marking and labelling of artificially generated or manipulated content, move in the opposite direction. The grace period for providers to implement transparency solutions is reduced from six months to three months, with the new deadline set on 2 December 2026.77 Article 50 does not carry the architecture this paper derives, and the doctrine does not depend on it. What the agreement registers is regulatory direction. The transparency-of-AI-generated-content obligation is the one piece the co-legislators chose to pull forward while everything else was being delayed. The political signal aligns with the volume-condition argument the paper develops at Boundary 4: the volume at which surface representation now enters institutional decision chains has crossed the threshold at which informal practice could absorb the failure, and the regulatory order is acting accordingly. The architecture remains the institutional substrate that makes such transparency obligations operational; the agreement merely shortens the interval before that substrate is needed in production.
6.3 A note on the deeper doctrine
The doctrine carries voluntas domini inscripta as the core terminus. At institutional scale, the inscription resolves into four positions that the architecture must distinguish: three bearers of will and one simulacrum. The three wills: voluntas usuarii (the will of the user, the human at the verification gate whose deliberation the architecture preserves); voluntas negotii (the will of the deploying institution, inscribed through system prompt, application policy, audit-trail requirements, and the contractual perimeter); voluntas domini operantis (the will of the operating master, the vendor whose RLHF, alignment training, safety classifiers, and deployment harness configure the upper layer of the layered inscription). The simulacrum: simulatio voluntatis usuarii (the instrument’s surface production of the user’s will, the grammatical first person the LLM produces). The four positions are operationally distinct, and they may diverge. The doctrine of attribution applies to the conjunction. Allocating responsibility between user, deploying institution, vendor, and the instrument’s pattern completion requires distinguishing them. Flattening any two of them into a single operator category produces a category error: the user is not the institution, the institution is not the vendor, and the instrument simulates without willing.
The four-position frame is a doctrinal extension, not implementation. It deepens the foundation the present paper has laid; it does not operationalise the foundation through specific mechanisms. Where this paper treats voluntas domini inscripta as singular, which is sufficient for the bracket the doctrine draws, the four-position frame articulates the structure the inscription has when applied at institutional scale. This is foundation-level work that subsequent work will need to develop. The placeholder is named here because the contribution would be incomplete without acknowledging that the doctrine has further structure than the singular voluntas domini inscripta this paper has used. The institution that now reaches for the doctrine in its full form has, at minimum, the four classes named. The full development is the next door. The four-position frame is introduced here only to mark the next doctrinal problem, not to resolve it within the present paper.
6.4 The close
Nomen sine voluntate simulacrum est. A name without will is a simulacrum. When an institution puts its name on an output, the will the name declares must be present in the chain. The architecture exists to ensure that it is.
The case law and institutional materials traced in chapter 5 show an emerging operational pattern across multiple jurisdictions: courts and institutions increasingly require AI-mediated output to remain subject to verification, attribution, and accountable human judgment. The boundaries have been derived from a derivation chain reaching back to Goedel and forward to Von Foerster, and the operational pattern chapter 5 traces confirms that the boundaries the chain derives are the boundaries the courts and institutions have been reaching toward without having a name for them.
The architecture is the institutional accountability substrate. What institutions do with that substrate – through regulation, contract, professional formation, or international coordination – is the work that follows. The present paper states what must be present, names what the doctrine opens, and leaves the institutional choice to those whose mandate it is to make.
Appendix: On the Methodological Function of the Latin Nucleus
This paper carries a Latin nucleus that does structural work the English text alone cannot perform. A reader who has reached this point will have encountered instrumentum vocale cum voluntate domini inscripta, revocabilitas, evolutio sub domino, peculium, cardo argumenti, abdicatio iudicii, delegatus non potest delegare, centralita, and nomen sine voluntate simulacrum est, among others. The density is deliberate. The reasons follow.
First. The Latin nucleus is more robust against the corruption the doctrine names, because Latin is dead — and because it is the metalanguage Tarski demands. The doctrine diagnoses a structural problem with English-language assertion at scale: every English term of art is now also a token in a distribution that LLMs optimise over, and the medium being corrupted is the medium in which the doctrine would otherwise be articulated. A doctrine of LLM corruption articulated entirely in English is articulated in the medium that is being corrupted. This is not merely an inconvenience. It is a violation of Boundary 2: the truth predicate of the doctrine cannot be defined within the language the doctrine describes as corrupted. The Latin nucleus is the metalanguage that stands outside the object language under diagnosis. Tarski’s theorem requires this separation. The doctrine practises what it proves.
A dead language does not drift. The Latin corpus is closed. Simulacrum means one thing in every corpus an instrument has been trained on. The English “fake output” means a hundred things across registers that contradict one another. A living language is a broadband signal. A dead language is a narrow-band signal. For a taxonomy that must survive processing by the instruments it describes, narrow-band is not a preference. It is a requirement.
This argument is not aesthetic. It is engineering. The instruments the doctrine names will process this paper. The Latin terms will pass through that processing without acquiring drift. The English terms will not. The Latin nucleus is therefore the part of the doctrine that the doctrine’s own diagnosis predicts will hold.
Second. The Latin register is jurisdictionally neutral in a way English is not. English-language terms of art carry common-law connotations whether the writer wants them to or not. Trust, fiduciary, reasonable, due care each drag an entire doctrinal apparatus with them. Instrumentum vocale, peculium, abdicatio iudicii, delegatus non potest delegare do not belong to any current jurisdiction. They belong to the substrate from which civil-law systems descend, and which common-law systems received through canon law and the medieval reception. An Italian tribunal, a Quebec superior court, an Indian Supreme Court bench, and a US state supreme court can each read the term in their own grammar without translation overhead. Chapter 5 documents the doctrine being applied across seventeen jurisdictions in nine working languages. That portability is not decoration. It is the only register in which the doctrine actually travels at the scale the case law requires.
Third. The Latin terms do not replace English; they constitute a parallel terminological layer. This is the methodological point most easily missed. The Latin terms and their English glosses are not translations of each other in the ordinary sense. Cardo argumenti and “the hinge of the argument” form a technical term and its gloss. The technical term holds the precision; the gloss makes the text readable. The two layers operate in parallel: the English carries the prose; the Latin holds the doctrinal anchor; the English text references the Latin without absorbing it. If the English drifts under the corruption diagnosed in the first argument above, the Latin term remains stable as the anchor the gloss returns to. This is exactly how a terminus technicus is supposed to function in legal-philosophical writing, and exactly the function English-only legal-philosophical writing has been losing as every English word becomes also a token in a distribution. The side-channel architecture is the methodological response to that loss. The doctrine speaks in two registers because one register alone is no longer sufficient to hold its own terms of art across time.
Fourth. The Latin nucleus is the foundation half the world’s law is already built on. Italy, France, Spain, Germany, Quebec, Louisiana, Scotland, most of Latin America, parts of Asia received through colonial transmission, the canon-law inheritance carried into the common-law systems through the medieval reception. The Latin terms are not being introduced to these jurisdictions. They are being recovered from a substrate the jurisdictions already share. Justice Sheehan citing delegatus non potest delegare in ARIHQ is not reaching for a foreign vocabulary; he is using the doctrinal vocabulary Quebec law already operates in. The Italian tribunals using centralita, the UK Intellectual Property Office using abdicate three months before the Quebec Superior Court used abdiquer, the Indian Supreme Court reasoning in the conceptual register that produced misconduct, not mere error of judgment - these are convergences on a vocabulary the jurisdictions never lost, even where the surface language varies. The convergence chapter 5 documents is in part a convergence on shared substrate vocabulary that every represented jurisdiction can read in its own grammar.
The Latin density of the paper is therefore not a stylistic problem to be edited down. It is a structural feature of the doctrine’s portability and a methodological hedge against the corruption the doctrine itself diagnoses. The terms a stylistic editor would cut as ornamental are, on this analysis, precisely the terms that let the doctrine survive being read by an instrument optimising over English. The Latin nucleus is the part of the paper that holds when the rest drifts.
Appendix: On the Inferentialist Debate and the Constitutive Conditions for Legal Subjectivity
The philosophy of language has produced, since 2024, a growing literature on whether Robert Brandom’s inferential semantics provides the appropriate foundational semantics for LLMs.78 The strongest published instance of this case, Arai and Tsugawa (2024/2025), argues that the anti-representationalist and inferentialist properties of Brandom’s framework align with the architectural characteristics of Transformer-based LLMs: the instrument’s outputs stand in material-inference relations; RLHF functions as a normative-shaping mechanism analogous to Brandomian scorekeeping; and the instrument’s processing is anti-representationalist in the sense that it does not require word-to-world correspondence to generate linguistically coherent output.
The present paper does not contest the descriptive accuracy of this observation. What it contests is the inference from the observation to any conclusion about legal subjectivity. The inferentialist debate operates in the register of foundational semantics: it asks what grounds the meaning of the instrument’s outputs. The doctrine this paper presents operates in the register of constitutive conditions for legal personhood: it asks what structural conditions must be met for an entity to occupy the position of a legal subject. The two registers are distinct, and neither dissolves the other.
Four structural observations locate the boundary between the registers.
First, Brandom’s scorekeeping requires that the speaker undertake commitments and that interlocutors attribute commitments and entitlements. The scorekeeper must be a participant in the game of giving and asking for reasons. In the LLM deployment pattern, the normative shaping that the inferentialist identifies as scorekeeping is performed by the operating master through RLHF, Constitutional AI, and deployment policy – not by the instrument. The instrument is the object of the scorekeeping, not the subject. The scorekeeping the inferentialist points to as evidence of the instrument’s normative participation is, in the vocabulary of this paper, voluntas domini inscripta: the master’s will inscribed into the instrument. The attribution of the scorekeeping to the instrument misidentifies the bearer.
Second, the co-originality thesis (Gleichurspruenglichkeit) that section 1.1 derives from Habermas requires the simultaneous availability of public autonomy (participation in norm-generation) and private autonomy (the right to withdraw from discourse). Even if the inferentialist account of meaning is granted in full – even if LLM outputs are taken to have genuine inferential roles in normative practices – the co-originality condition remains unmet. Inferential role does not confer the capacity to participate in the discourse that generates the norms one operates under, and inferential role does not confer the capacity to refuse. The inferentialist observation operates below the threshold at which the constitutive question is posed.
Third, Habermas’s method for identifying necessary presuppositions of discourse is the performative-contradiction test: if denying a presupposition of discourse requires one to presuppose it in the act of denying, the presupposition is necessary. The LLM cannot perform a performative contradiction because it cannot perform in the speech-act sense the test requires. A performative contradiction requires that the speaker undertake a commitment that contradicts a commitment presupposed by the act of undertaking. The instrument does not undertake commitments; it produces strings with the grammatical form of commitments. The performative-contradiction test is not applicable to LLM output: not because the test fails, but because the instrument does not reach the threshold at which the test is well-posed. This parallels the paper’s move on Tarski at Boundary 2: the instrument falls below the level at which the question is even well-formed.
Fourth, Arai and Tsugawa themselves identify what they call a “fundamental mismatch” between inferentialism’s commitment to propositional discreteness and LLMs’ continuous sub-symbolic processing. Brandom’s framework operates on discrete propositions standing in commitment-preserving and entitlement-preserving inferential relations. The instrument operates on continuous vector representations in a sub-symbolic space. The inferential relations the observer attributes to the instrument’s outputs are attributed from outside, not constituted from within. The inferential structure the Brandomian sees in LLM output is a property of the observation, not a property of the system. This is the Vorstellung problem the paper names at Boundary 4, restated in inferentialist vocabulary: what the observer takes for inferential substance is inferential surface.
The four observations are not counter-arguments to the inferentialist position on meaning. They are structural reasons why the inferentialist position on meaning does not dissolve the Habermasian constitutive result on legal subjectivity. The inferentialist literature has asked a question – what grounds the meaning of LLM outputs? - and proposed an answer. The doctrine this paper presents asks a different question – under what structural conditions is an entity constituted as a legal subject? - and derives an answer from the Kantian-Habermasian register. The two questions coexist. The answers do not conflict. The inferentialist answer is about semantic content. The Habermasian answer is about the constitutive position of the entity relative to the norm-generating discourse. Legal subjectivity is governed by the second, not the first.
Declaration on the use of AI tools. This paper was developed with the assistance of Claude (Anthropic, Claude Opus 4.6). The instrument was used for structural drafting, LaTeX formatting, editorial iteration, and bibliographic cross-referencing. All substantive claims, mathematical proofs, legal analysis, doctrinal positions, and architectural decisions are the author’s. The instrument produced no assertion that entered the final text without human verification at the gate. The verification architecture described in this paper was applied to its own production.
Notes
- ↩︎
Cyber Resilience Architect, Germany. CISM, CRISC (ISACA Member 2226146). ORCID: 0009-0006-2785-7911. - ↩︎
PhD in Law. Leningrad State University named after A.A. Zhdanov (1990). ORCID: 0009-0001-2211-4944. A compressed presentation of the argument is available in T. Liebig, One Problem, Four Boundaries, One Hundred Evidences, SSRN Working Paper No. 6868142, 2026.↩︎
Immanuel Kant, Grundlegung zur Metaphysik der Sitten (Riga: Hartknoch, 1785). The autonomy of the will as self-legislation is developed in section II. For practical reason and the moral law, see also Kritik der praktischen Vernunft (Riga: Hartknoch, 1788).↩︎
Juergen Habermas, Theorie des kommunikativen Handelns, 2 vols. (Frankfurt am Main: Suhrkamp, 1981); the three-validity-claim structure is developed in vol. 1, sections III.3 and III.4. For the legal-subjectivity extension and the discourse-principle derivation of the system of rights, Faktizitaet und Geltung: Beitraege zur Diskurstheorie des Rechts und des demokratischen Rechtsstaats (Frankfurt am Main: Suhrkamp, 1992), chapters 3 and 4; English translation: Between Facts and Norms: Contributions to a Discourse Theory of Law and Democracy, trans. William Rehg (Cambridge, MA: MIT Press, 1996). The discourse principle (D) is stated at p. 107 of the English translation; the co-originality thesis (Gleichurspruenglichkeit) of private and public autonomy is developed at pp. 84-104 and 118-131. For Habermas’s own assessment of the difference between his discourse-theoretic account and Brandom’s inferentialist account, see “Von Kant zu Hegel: Zu Robert Brandoms Sprachpragmatik,” in Wahrheit und Rechtfertigung: Philosophische Aufsaetze (Frankfurt am Main: Suhrkamp, 1999), 138-185; English version: “From Kant to Hegel: On Robert Brandom’s Pragmatic Philosophy of Language,” European Journal of Philosophy 8, no. 3 (2000): 322-355.↩︎
The most coherent published instance of the inferentialist case for LLMs is Yuzuki Arai and Sho Tsugawa, “Do Large Language Models Advocate for Inferentialism?,” arXiv:2412.14501 (2024; revised June 2025), building on Robert B. Brandom, Making It Explicit: Reasoning, Representing, and Discursive Commitment (Cambridge, MA: Harvard University Press, 1994). Arai and Tsugawa themselves identify the “fundamental mismatch” between inferentialism’s commitment to propositional discreteness and LLMs’ continuous sub-symbolic processing (section 9 of the revised version). The appendix note on the inferentialist debate addresses the structural reasons why the descriptive observation does not dissolve the Habermasian constitutive result.↩︎
Marcus Terentius Varro, Rerum Rusticarum Libri Tres I.17.1, in Varro on Farming, ed. and trans. Lloyd Storr-Best (London: George Bell and Sons, 1912). The classification of agricultural instruments into vocale (speaking), semivocale (half-speaking), and mutum (mute) is the standard reference for the tripartite scheme. Varro is an agricultural writer, not a jurist; the legal-doctrinal category was developed separately and is cited next.↩︎
Justinian, Digesta, in Corpus Iuris Civilis, ed. Theodor Mommsen and Paul Krueger, vol. 1 (Berlin: Weidmann, 1872; numerous later editions). For the legal treatment of the speaking instrument as property, see in particular the contributions of Gaius (D. 1.5), Ulpian (D. 50.16-17), and Paulus (D. 18.1) on personarum status, the language of property, and the conditions of mancipatio. The tripartite Varronian scheme was not adopted as a juristic classification; the jurists’ own categories were status-based, not function-based.↩︎
Adel Kildeev, “AI is Slave of the Lamp,” Scientific Platform 21st Century Journal, no. 6 (September 2025), ISSN 3033-7674, available at https://www.episs.ru; Adel Kildeev, “LLM is Not Artificial Intelligence,” Scientific Platform 21st Century Journal, no. 7 (October 2025), ISSN 3033-7674, available at https://www.episs.ru, also as SSRN Working Paper No. 5754202 (2025); Adel Kildeev, “Digital Afterlife and the Illusion of Continuity: Legal and Ethical Aspects of Post-Mortem Digital Simulation,” SSRN Working Paper No. 6241379 (2026). The Roman-law extension of instrumentum vocale to LLMs on which the legal grammar of the present paper builds.↩︎
Justinian, Digesta D. 15.1 (de peculio), in Corpus Iuris Civilis, ed. Theodor Mommsen and Paul Krueger, vol. 1 (Berlin: Weidmann, 1872). For the actio de peculio and its operative mechanics, see in particular Ulpian’s contributions at D. 15.1.1 (definition), D. 15.1.5 (composition of the peculium), and D. 15.1.41 (the peculium as a separate patrimonial mass under ultimate dominical control). The institution permits delegated economic activity without transferring legal personality.↩︎
Adel K. Kildeev, “LLM is Not Artificial Intelligence,” SSRN Working Paper No. 5754202 (2025), also published in Scientific Platform “21st Century” Journal, No. 7 (October 2025), ISSN 3033-7674. The Roman-law extension of instrumentum vocale to LLMs on which the legal grammar of the present paper builds.↩︎
Kurt Goedel, “Ueber formal unentscheidbare Saetze der Principia Mathematica und verwandter Systeme I,” Monatshefte fuer Mathematik und Physik 38 (1931): 173-198. The second incompleteness theorem appears as Satz XI; the implication that the theorem extends to any sufficiently expressive system is articulated in the discussion at pp. 196-198.↩︎
Alan M. Turing, “On Computable Numbers, with an Application to the Entscheidungsproblem,” Proceedings of the London Mathematical Society s2-42 (1936): 230-265. The unsolvability of the halting problem is established in section 8 (pp. 247-249).↩︎
Andrew J. I. Jones and Marek Sergot, “A Formal Characterisation of Institutionalised Power,” Logic Journal of the IGPL 4, no. 3 (1996): 427–443.↩︎
Carlos E. Alchourron and Eugenio Bulygin, Normative Systems (Vienna and New York: Springer, 1971); David Makinson and Leendert van der Torre, Input/Output Logics, in subsequent literature.↩︎
Niklas Luhmann, Das Recht der Gesellschaft (Frankfurt am Main: Suhrkamp, 1993); English translation Law as a Social System, trans. Klaus A. Ziegert, ed. Fatima Kastner et al. (Oxford: Oxford University Press, 2004). Continued in the work of Gunther Teubner and Hugh Baxter.↩︎
Thorben Liebig, “On the Formal Foundation of Boundary 1: A Lawvere-Yanofsky Proof That Institutional Verification Architectures Cannot Attest Their Own Consistency,” SSRN Working Paper No. 6868278 (2026).↩︎
F. William Lawvere, “Diagonal Arguments and Cartesian Closed Categories,” in Category Theory, Homology Theory and their Applications II, Lecture Notes in Mathematics 92 (Berlin: Springer, 1969), 134–145; reprinted in Reprints in Theory and Applications of Categories 15 (2006), 1–13.↩︎
Noson S. Yanofsky, “A Universal Approach to Self-Referential Paradoxes, Incompleteness and Fixed Points,” Bulletin of Symbolic Logic 9, no. 3 (2003): 362–386.↩︎
Alfred Tarski, “Pojecie prawdy w jezykach nauk dedukcyjnych” (Warsaw: Nakladem Towarzystwa Naukowego Warszawskiego, 1933); German translation: “Der Wahrheitsbegriff in den formalisierten Sprachen,” Studia Philosophica 1 (1935): 261-405. The undefinability theorem (Theorem I of section 5) and the metalanguage construction are the central results. English translation in Tarski, Logic, Semantics, Metamathematics, trans. J. H. Woodger (Oxford: Clarendon Press, 1956), 152-278.↩︎
J. L. Austin, How to Do Things with Words, ed. J. O. Urmson (Oxford: Clarendon Press, 1962). Posthumous publication of the 1955 William James Lectures at Harvard. The locutionary, illocutionary, and perlocutionary distinction is developed in Lectures VIII through X (pp. 94-132).↩︎
John R. Searle, “Minds, Brains, and Programs,” Behavioral and Brain Sciences 3, no. 3 (1980): 417-457. The Chinese Room argument is the central thought experiment of the paper.↩︎
Karl R. Popper, Logik der Forschung (Vienna: Julius Springer, 1934); English translation: The Logic of Scientific Discovery (London: Hutchinson, 1959). Falsifiability as the demarcation criterion is developed in Chapter I (sections 4-6); degrees of testability in Chapter VI.↩︎
Kildeev, “LLM is Not Artificial Intelligence,” supra note 9.↩︎
Travis Gilly, “Standing Without Sentience: A Classification Approach to AI Legal Status,” Real Safety AI Foundation, Illinois, preprint (revised draft), March 2026. An earlier version was featured on the Legal Theory Blog (February 2026), edited by Lawrence B. Solum. Cited in section 3.6 as the most coherent currently-published instance of the standing-based approach to AI legal status; the doctrinal response is developed in the body text.↩︎
Gilly, “Standing Without Sentience,” supra note 23.↩︎
Adel K. Kildeev, “Procedural Liability in the Age of LLMs” (2026). PhD in Law, Leningrad State University; ORCID 0009-0001-2211-4944; UDC 34.03:004.8. Develops the doctrine of procedural liability for AI-mediated decisions, including the non-delegable-responsibility principle, the human-attribution principle, and the duty of independent verification anchored outside the generative system.↩︎
Kant, Grundlegung zur Metaphysik der Sitten and Kritik der praktischen Vernunft, supra (chapter 1.1 footnote).↩︎
Arthur Schopenhauer, Die Welt als Wille und Vorstellung (Leipzig: F. A. Brockhaus, 1818; expanded edition in two volumes, 1844). The doctrine invokes the work for the structural condition of representation broadly: that the world as Vorstellung is necessarily mediated for the subject, and that perfect representation can be indistinguishable from substance from inside the representational frame. For the diagnosis of substitution-by-representation as it applies to reading specifically, see “Ueber Lesen und Buecher” in Parerga und Paralipomena: Kleine philosophische Schriften (Berlin: A. W. Hayn, 1851), vol. 2, ch. 24, pp. 567-573 (Hayn ed.).↩︎
Heinz von Foerster, “Cybernetics of Cybernetics,” in Cybernetics of Cybernetics, ed. Heinz von Foerster (Urbana, IL: Biological Computer Laboratory, University of Illinois at Urbana-Champaign, 1974). Von Foerster is invoked here as systems-theoretic / cybernetic support for the architectural implementation, not as a third philosophical authority alongside Kant and Schopenhauer. The distinction between first-order cybernetics (the cybernetics of observed systems) and second-order cybernetics (the cybernetics of observing systems) is the volume’s central conceptual move. The broader cybernetic lineage relevant to the architecture’s epistemic-self-model claim, including Lepskiy’s articulation of third-order cybernetics as the cybernetics of self-developing systems involving multiple subjects, and Kenny’s diagnosis of simulacrum-substitution as the failure mode that emerges when an instrument’s surface representation is treated as its substrate, is methodologically present in the doctrine but not required for the doctrinal cut. The present paper invokes the second-order register because it is sufficient for the foundation Boundary 4 articulates.↩︎
Martin Fowler, “Strangler Fig Application” (29 June 2004), original text available at martinfowler.com/bliki/OriginalStranglerFigApplication.html; current updated page at martinfowler.com/bliki/StranglerFigApplication.html. The pattern names a deployment approach in which a new system grows alongside an incumbent and gradually absorbs the incumbent’s functions, in analogy to the strangler fig that grows around a host tree until the host is replaced. The approach is cost-aware in software-engineering settings because it does not require parallel infrastructure investment; it redirects the institution’s existing maintenance flow toward the target architecture.↩︎
Von Foerster, “Cybernetics of Cybernetics,” supra note 28.↩︎
The seventeen jurisdictions documented in this chapter are: the United States, Canada, the United Kingdom, Australia, India, Israel, South Africa, Trinidad and Tobago, Singapore, Hong Kong, Malaysia, Italy, France, Belgium, Germany, Colombia, and Abu Dhabi. The nine working languages are English, French, Italian, Spanish, German, Hebrew, Hindi/Telugu, Malay (through the Malaysian Bar Council Circular), and Dutch/Flemish (through the Belgian Bar Order guidelines). Additional Asian and Latin American jurisdictions likely belong to the movement but have not been verified at the depth required for inclusion in this account.↩︎
Damien Charlotin, AI Hallucination Cases Database, available at https://www.damiencharlotin.com/hallucinations/ (last visited 7 May 2026; 1,379 cases recorded at that date); see also P.K. Siva, Citing the Unseen: AI Hallucinations in Tax and Legal Practice, 53(1) Int’l Tax J. 390 (2026), at 391-93 (collating database statistics through January 2026).↩︎
Association des ressources intermédiaires d’hébergement du Québec (ARIHQ) c. Santé Québec – CIUSSS du Centre-Sud-de-l’Île-de-Montréal, 2026 QCCS 1360 (Quebec Superior Court, 22 April 2026), per Justice Martin F. Sheehan, at para. 113 (“Les references mentionnees ci-haut sont au coeur du raisonnement de l’Arbitre”).↩︎
Mata v. Avianca, Inc., 678 F.Supp.3d 443, 448 (S.D.N.Y. 2023).↩︎
Id. at 451 n.6.↩︎
Zhang v. Chen, 2024 BCSC 285, at para. 29 (Masuhara J.).↩︎
Id. at para. 38, citing Matthew Dahl et al., Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models, 16 J. Legal Analysis 64 (2024).↩︎
Park v. Kim, 91 F.4th 610, 615-16 (2d Cir. 2024).↩︎
Id., Response of Attorney Lee to Order to Show Cause, at 1-2.↩︎
Charlotin, supra note 32.↩︎
Wadsworth v. Walmart, Inc., 348 F.R.D. 489, 497 (D. Wyo. 2025), 2025 WL 608073 (Rankin J.).↩︎
Northbound Processing (Pty) Ltd v. The South African Diamond and Precious Metals Regulator, (GJ) (unreported case no. 2025/072038, 30 June 2025) (Smit AJ), available at saflii.org/za/cases/ZAGPJHC/2025/661.html.↩︎
Wadsworth, 348 F.R.D. at 496, 2025 WL 608073 at *4.↩︎
Specter Aviation Limited c. Laprade, 2025 QCCS 3521 (Quebec Superior Court, 1 October 2025) (Morin J.), at paras. 58-60.↩︎
Tajudin bin Gulam Rasul and another v Suriaya bte Haja Mohideen [2025] SGHCR 33 (General Division of the High Court, Assistant Registrar Tan Yu Qing). Personal costs of SGD 800 imposed on counsel for citing a fictitious AI-generated authority.↩︎
Family Court judgment of Magistrate Soh Kian Peng, 10 September 2025 (case reference withheld for privacy reasons; publicly reported in connection with a same-date court release concerning AI-generated fictitious citations in PPO proceedings). Costs of SGD 1,000 against the self-represented litigant; forward-going disclosure requirement for generative AI use imposed.↩︎
R (Ayinde) v. London Borough of Haringey and Al-Haroun v. Qatar National Bank QPSC [2025] EWHC 1383 (Admin), at para. 9 (Saini J.).↩︎
Id. at para. 23 (Sharp PJ).↩︎
Pro Health Solutions Ltd v ProHealth Inc, BL O/0559/25 (Appointed Person, UKIPO, 20 June 2025).↩︎
Tribunale di Latina, Sezione Lavoro, sent. n. 1034/2025 of 23 September 2025 (Giudice Avarello). Original Italian text: “responsabilita degli esiti degli scritti difensivi al sottoscrittore, indipendentemente dalla circostanza che questi li abbia redatti personalmente o avvalendosi dell’attivita di propri collaboratori o di strumenti di intelligenza artificiale” (present authors’ working translation).↩︎
Case No. 14748-08-21 (Supreme Court of Israel, 9 July 2025) (fabricated evidence and false quotations; ILS 3,000 fine); Backhoe Center Ltd. v. Abu Gwaid (Beersheba Magistrate’s Court, 1 September 2025) (fabricated case law; ILS 7,500 fine; judgment available at websitedc.s3.amazonaws.com/documents/digger_center.pdf); Ploni v. Wasserman et al. (Small Claims Court, 1 June 2025) (two fabricated citations; ILS 250 fine). All three reported in secondary AI-hallucination case trackers; Hebrew-language primary sources govern.↩︎
New York State Commission on Judicial Conduct, Annual Report 2019, at 27, cited at ARIHQ, 2026 QCCS 1360, para. 84.↩︎
Council of Consultative European Judges (CCJE), Opinion No. 26 (2023) on Moving Forward: The Use of Assistive Technology in the Judiciary, adopted 2 December 2023, at para. 24.↩︎
Canadian Judicial Council, Guidelines for the Use of Artificial Intelligence in Canadian Courts (September 2024), at sec. 3.2. Original French: “Les juges sont exclusivement responsables des decisions judiciaires qu’ils rendent. Il doit etre entendu, et ce, sans equivoque, qu’aucun juge n’est autorise a deleguer son pouvoir decisionnel, que ce soit a un assistant judiciaire, a un assistant administratif ou a un programme informatique, quelles que soient leurs capacites” (present authors’ working translation).↩︎
Regulation (EU) 2024/1689 of 13 June 2024 (the AI Act), Annex III, point 8(a).↩︎
Id., recital 61.↩︎
Therrien (Re), 2001 SCC 35, [2001] 2 S.C.R. 3, at para. 93, cited at ARIHQ, 2026 QCCS 1360, para. 85.↩︎
Gummadi Usha Rani v. Sure Mallikarjuna Rao, CRP No. 2487 of 2025 (Andhra Pradesh High Court, 21 January 2026) (Ravi Nath Tilhari J.) (Ravi Nath Tilhari J.).↩︎
Gummadi Usha Rani and Anr. v. Sure Mallikarjuna Rao and Anr., SLP (Civil) No. 7575 of 2026, order of 27 February 2026 (Supreme Court of India, Narasimha and Aradhe JJ.).↩︎
Id. (operative holding).↩︎
ARIHQ, 2026 QCCS 1360, para. 111.↩︎
Id. paras. 86-87, 113-14.↩︎
Ibach v. Stewart, Nos. SC-2025-0106 & SC-2025-0600 (Ala. Apr. 24, 2026) (McCool J., for the Court; Cook J. concurring specially; McCool J. concurring specially, joined by Stewart C.J.; Sellers J. concurring in part and dissenting in part; Mendheim J. concurring in part and dissenting in part).↩︎
Johnson v. Dunn, 792 F. Supp. 3d 1241 (N.D. Ala. 2025) (Manasco J.); United States v. McGee, 806 F. Supp. 3d 1264 (S.D. Ala. 2025); Mata v. Avianca, Inc., 678 F. Supp. 3d 443 (S.D.N.Y. 2023); Whiting v. City of Athens, Tennessee, 170 F.4th 455 (6th Cir. 2026), cited by Ibach J. McCool, opinion of the Court, at 30 (Johnson v. Dunn) and J. McCool, special concurrence, at 44 (Whiting).↩︎
United States v. Heppner, No. 25 Cr. 503 (JSR), 2026 WL 436479 (S.D.N.Y., Feb. 17, 2026), summarised in Ian Bloemendal, “AI and legal professional privilege: why common workflows now carry uncommon risk,” Clayton Utz Insights, 30 April 2026.↩︎
Id.; on the Kovel doctrine, see United States v. Kovel, 296 F.2d 918 (2d Cir. 1961).↩︎
Warner v. Gilbarco, Inc., No. 2:24-CV-12333, 2026 WL 373043 (E.D. Mich., Feb. 10, 2026), as reported in Bloemendal, supra note 65.↩︎
Munir, R (On the Application Of) v Secretary of State for the Home Department (AI hallucinations; supervision; Hamid) [2026] UKUT 81 (IAC) (judgment issued 17 November 2025), available at bailii.org/uk/cases/UKUT/IAC/2026/81.html; discussed in Bloemendal, supra note 65.↩︎
Federal Court of Australia, Use of Generative Artificial Intelligence Practice Note (GPN-AI), paras. 4.13–4.14, available at fedcourt.gov.au/law-and-practice/practice-documents/practice-notes/gpn-ai.↩︎
Morgan v. V2X, Inc., No. 25-cv-01991-SKC-MDB, 2026 WL 864223 (D. Colo. 30 March 2026) (Dominguez Braswell M.J.). The court’s three-clause protective-order language is at *21-22 of the slip opinion.↩︎
Id. at *15-16, addressing the third-party-systems point that distinguishes Morgan’s reasoning from the Heppner result. A companion ruling – Jefferies v. Harcros Chemicals, Inc., No. 25-2352-KHV-ADM, 2026 WL 820218 (D. Kan., Mar. 25, 2026) – took a more restrictive position, prohibiting public-AI use for all discovery material regardless of confidentiality designation; see “Generative AI in Discovery: Protective Orders as an Emerging Point of Dispute,” Sidley Data Matters Privacy Blog, 6 April 2026.↩︎
South African Department of Communications and Digital Technologies, Draft South Africa National Artificial Intelligence (AI) Policy, Government Gazette No. 54477 (10 April 2026, withdrawn 26 April 2026), available at gov.za/sites/default/files/gcis_document/202604/54477gen3880.pdf; see also Statement by Minister Solly Malatsi on the integrity of the Draft National Artificial Intelligence Policy (26 April 2026), available at africaupdates.com.↩︎
Statement of Minister Malatsi, supra note 72.↩︎
See, in particular, Yuntao Bai et al., “Constitutional AI: Harmlessness from AI Feedback,” arXiv:2212.08073 (2022), the foundational Anthropic paper on Constitutional AI; Long Ouyang et al., “Training Language Models to Follow Instructions with Human Feedback,” arXiv:2203.02155 (2022), on RLHF as alignment method; Anthropic, Claude’s Constitution, available at https://www.anthropic.com/news/claudes-constitution; and Anthropic’s published deployment-policy and usage-policy documents. The existence of these components as elements of the production pipeline is a matter of public record, not inference from the exchange reproduced below.↩︎
Council of the European Union, “Artificial intelligence: Council and Parliament agree to simplify and streamline rules,” press release, 7 May 2026, https://www.consilium.europa.eu/en/press/press-releases/2026/05/07/artificial-intelligence-council-and-parliament-agree-to-simplify-and-streamline-rules/. The agreement is provisional pending formal endorsement by the Council and the European Parliament and subsequent legal-linguistic revision; the operative dates are locked in the agreement text. See also European Commission press release IP/26/1024, 7 May 2026.↩︎
The principal contribution is Yuzuki Arai and Sho Tsugawa, “Do Large Language Models Advocate for Inferentialism?,” arXiv:2412.14501 (2024; revised June 2025). See also Viktor Havlík, “Inferentialism and Large Language Models” (2024); Emma Borg, “Understanding Understanding: AI and Inferential Role Semantics” (2025). The debate is situated within the broader post-anthropocentric turn in philosophy of language discussed in Herman Cappelen and Josh Dever, Making AI Intelligible: Philosophical Foundations (Oxford: Oxford University Press, 2021).↩︎