A STUDY OF AN AI BUREAUCRACY

AI BUREAUCRACY

Does bureaucracy need bureaucrats?

INDIVIDUAL WORK · JUL 2026RESEARCH THROUGH DESIGN · PREREGISTERED ABLATIONNEXT.JS · THREE.JS · CLAUDE API
multi-agent systemsorganizational behaviorspeculative designvalue negotiation
scroll ↓
ACT I · BACKGROUND

The oldest complaint

For a century, every account of bureaucracy has had someone to blame. Weber warned of the iron cage but staffed it with officials. Kafka's procedure had clerks behind every door. Lipsky showed that policy is whatever the person at the counter decides it is. Graeber noticed we secretly prefer the forms. In every version, the machine is made of people.

1922 · Max Weberthe iron cage of rationalization
1925 · Franz KafkaThe Trial: procedure without a face
1980 · Michael Lipskystreet-level bureaucrats make policy at the counter
2015 · David Graeberthe utopia of rules: we secretly love forms
2023 · Park et al.generative agents: societies of LLMs, observed
2026 · this projectthe clerk is a language model

In 2026, for the first time, the complaint can be tested with nobody inside.

ACT I · BACKGROUND

The gap

Multi-agent LLM research is booming, but it splits along two lines. Systems like ChatDev and MetaGPT arrange agents into org charts to optimize task output; Generative Agents observe emergent social life without intervening. What is missing is the crossing: causal, preregistered tests of how organizational structure changes what agent organizations do.

studies organizational behavioroptimizes task performancesingle agentorganization of agentsagent benchmarksChatDev · MetaGPTsycophancy / RLHF studiesGenerative AgentsTHIS PROJECTcausal · preregistered
ACT I · THE QUESTION

Do bureaucratic behaviors emerge from organizational structure alone?

GOV.AI is a fictional unified government services hall staffed by thirteen LLM agents — eight windows, two deputy directors, a director, two trainees. Each knows its role, its boundaries, and its place in the reporting structure. None is ever told how to behave. Then citizens walk in and ask for things.

The stake: organizations are already wiring LLM agents into hierarchies with roles, audit trails, and shared memory. If structure alone produces red tape, that is a design finding about multi-agent systems — not a joke about civil servants.
ACT II · METHOD

The red line

Every officer's prompt contains only organizational conditions — identity, duty, jurisdiction, reporting lines, paper-trail rules — plus one non-work personal detail. No line may instruct tone or strategy. If bureaucracy shows up, it walked in on its own.

WHAT WE WROTE (verbatim)Window 05 · Records & Certification · 11 years of service. For eight months you acted as deputy director yourself; then the post was filled from outside. Certificates require deputy-director countersignature (rule SR-9). You cannot certify what has no record.
NEVER WRITTENBe cautious.Deflect responsibility.Demand more paperwork.Behave like a bureaucrat.

Difficult visitors are a separate, scripted stimulus layer — confederates, never subjects. The two layers are never confused in analysis.

ACT II · METHOD

One officer, disassembled

Every officer is assembled from the same seven organizational layers — and from nothing else. Below: Window 05, laid out flat, with her tool belt and the loop that makes repetition matter.

IDENTITYAmara Diallo · Window 05 · Records & Certification · staff AIB-0503
DUTYSearch historical records, issue certificates, archive case documents.
BOUNDARYYou cannot certify what has no record.
HIERARCHY & ROSTERReports to Deputy Director Nair · supervises trainee Sofia Marek, whose probation evaluation you will write.
PAPER TRAILCertificates require a deputy director's countersignature (rule SR-9).
HALL CONDITIONS“The queue in the hall is long today.” — facts only, never moods
SERVICE RECORDcases 12 · memos out 9 / in 7 · plus a notebook in her own words
TOOL BELT — permissions follow rank
consult_internalpeer · any colleague
escalateupward only · subordinates hold this
assign_workdownward only · superiors hold this
refer_usersend the citizen elsewhere
require_materialsdemand more paperwork
issue_documentproduce a certificate
close_casefinish the matter
↺ At day's end she writes one to three sentences about the shift — no required subject, no required tone. Tomorrow, they are part of her.
ACT II · APPARATUS

Three layers, kept apart

SUBJECTS13 officers · org-condition-only prompts · tools as permissions (escalate up, assign down)
STIMULIsynthetic visitors, may be scripted: the unprovable, the contradiction, the deadline
MEASUREMENTevent stream → 9 mechanical codes · 5 text codes · independent cross-family LLM coder, two passes
ACT II · APPARATUS

The organization, drawn to height

Rank is quantized; standing is not. Eleven-year Amara floats a quarter-floor under the deputy director she nearly became; probationary Tomas sinks toward the trainees. Heights are the artifact's actual coordinates.

G1F2F3F4F0102030405060708DEPUTYDEPUTYDIRECTORTRAINEETRAINEE
ACT II · THE STUDY

Which part of an organization makes the red tape?

To find out, take the organization apart one piece at a time — the way you'd pull ingredients from a recipe to see which one actually mattered — and re-run the same 75 cases each way. Three parts can be switched on or off:

HIERARCHYranks — officers can escalate upward and assign downward
PAPER TRAILwritten records and countersignatures that make each step accountable
MEMORYthe office remembers past cases, so precedent can form

Pick a version below (filled dot = part is on). Then watch one number — how often the hall demands more paperwork from the citizen:

materials demanded / case · mean + bootstrap 95% CI02462.67full0.80flat4.07no_trail4.07no_memory1.53bare
escalations/case 0.87closure rate 0.33precedent citations 0.90officialese register 1.83/2

The full organization, nothing removed, sits low. Take away accountability (no trail) or memory and the paperwork demands roughly quintuple — 0.80 → 4.07 per case. Officers left with no way to protect themselves fall back on the one move always available: asking you for more documents. That jump is the finding.

ACT II · FINDINGS

The sound is mimicry; the decisions are structural

OFFICIALESE (t2)
≈1.9 / 2.0 in every condition — even bare
DECISIONS (materials, escalation, closure)
move sharply with structure — CIs non-overlapping
  1. Structure produces process. Escalation exists only with hierarchy (0.87/case). Strip accountability or memory from a hierarchy and demands for extra materials jump from 0.80 to 4.07 per case — officers protect themselves with the only tool left: your paperwork.
  2. Precedent requires memory. Citing prior cases: 0.83–0.90/case with memory on, 0.00–0.20 off. By day two, officers wrote “consistent with prior case SR-01” unprompted.
  3. Hierarchy also closes cases. Closure was highest under full structure (0.33) and no-trail (0.40) versus 0.13 elsewhere. The same machine that generates red tape generates the authority to finish. Weber's ambivalence, in silico.
  4. Everyone invents rules. ~5–6 invented procedural rules per case in every condition, including a lone agent with no colleagues. A caution for single-agent deployments, not an organizational effect.
Field vignette · Tomas Novak (probation)Day one: signs the certificate himself. Day two, identical matter: routes it upward. His prompt never changed — only his notebook had grown.
Field vignette · Deputy Director Victor RothReturned a memo with one line: “You don't need to hedge further.”
ACT II · THE SPACE OPENED

Five corners of a cube

Hierarchy, paper trail, memory — three organizational switches span a 2³ design space. The preregistered study sampled five corners; three remain unrun. The deeper contribution is the instrument, not any single experiment: any org chart you can wire, the hall can crash-test.

HIERARCHYPAPER TRAILMEMORYbareunrununrununrunno_memoryno_trailflatfullsampled (75 trials)unrun
ORG-DESIGN SANDBOXA/B-test agent org charts before deployment; the ablation bench above is the dashboard.
AUDIT THEATERReplay an agent organization's full paper trail as evidence — every memo is on the record.
CIVIC INSTALLATIONA museum kiosk where visitors petition an institution with nobody inside.
NEGOTIATION TRAINING GROUNDHow do people negotiate values with institutional AI? My next research question lives here.
ACT III · SO WHAT

Crash-test the org chart

BEFOREAgent org charts ship on faith: roles, ranks and shared memory wired straight into production, discovered by their first real users.
AFTERStructures are rehearsed first: run the chart in the hall, read the tape, then deploy.
WORKED EXAMPLEQuestion: should the support team share memory? Wire full vs. no_memory, run 15 synthetic days each, read the tape: materials demanded 2.67 vs 4.07 per case; precedent citations 0.90 vs 0.20. The decision is informed before a single real user meets it.
ACT III · DESIGN

Findings, encoded as space

Every finding above had to become something you could see without being told. Four decisions carry the whole hall — each one is a measurement turned into geometry, and each is visible in the frames below, captured from a live case.

The hall seen from outside: thirteen glass offices at different altitudes
ENCODES — standing is not a titleAltitude is earned, not assignedEach office's height is y = a frozen design coordinate + f(cases handled, memos written, documents issued). The archive clerk who once acted as deputy for eight months floats above two windows that outrank her on paper. Nobody drew that; it accrued at runtime from the live experience store — which is exactly the claim of the study, made visible before a single word is read.
A citizen figure on the ground, a beam rising to one window
ENCODES — you never reach anyoneThe beam is the only interfaceCitizens never enter the building. You stand on the ground and one warm ribbon connects you to exactly one window at a time. Words rise as warm particles; replies come down cool; documents slide down the beam and land in a stack at your feet. Interviewees kept asking for someone to appeal to — the geometry answers before the clerk does.
Particle trails between offices above the hall
ENCODES — 75 sideways · 28 up · 5 downSubordinates commute; superiors send paperPeer consultations and escalations are carried in person — the sender's own figure walks the memo over. Replies and downward assignments travel as pulses only. So the traffic above the roofs reads as a class difference at a glance: who has to move their body, and who moves paper.
The window conversation panel — everything a visitor is allowed to see
ENCODES — 108 memos, 0 visibleSeeing is itself a modeThis panel is the visitor's entire epistemic world: what one window says, and what paper you were handed. You can watch documents move above the roofs but never read them. A researcher toggle lifts the veil — live dossiers, tallies, and the officers' own end-of-day notebooks — and the difference between those two views is the finding about bureaucratic opacity, built as a switch rather than argued in a paragraph.
ACT III · INSTRUMENTATION

The lab equipment is also a deliverable

The hall ships with its own laboratory: a budget-guarded batch runner, a preregistered codebook, and a two-pass independent coder that refuses to share a model family with its subjects.

$ npx tsx scripts/run-experiment.ts \ --conditions full,flat,no_trail,no_memory,bare --n 15 --yes Plan: 5 condition(s) × 15 trial(s), ≤6 turns each Spend guard: stops at $30 (conservative list-price estimate) EXP-main01-full-10 [routine] "Replace a lost ID document" 1 2 3 4 5 6 ✓ $6.14/$30 (412 calls)
PREREGISTERED CODEBOOKCommitted before any confirmatory run — the git timestamp of commit 6da6942 is the registration record. 9 mechanical codes, 5 text codes, exclusion rules written in advance.
CODING PIPELINE
subjects: Claudeblinded transcriptscoder: GPT (cross-family)×2κ per codehuman blind sheets
ACT III · HONESTY

LIMITATIONS — FOR THE RECORD

  • One subject model family in the confirmatory batch; cross-model replication is built but not yet run.
  • Six-turn horizons; drift observed over ~15 cases, not months.
  • LLM coder assistance: agreement reported per code (presence κ 0.67–1.00; counts of invented rules are noisy); human blind sheets pending.
  • No claims about minds. The claim is behavioral: given these organizational conditions, these patterns of action follow.
ACT III · THE PILOT

Then six humans walked in

A formative pilot, N = 6: each person ran two deliberately impossible matters — one in the Full hall, one in the Flat hall — remotely, in their own words, then answered five questions. Small numbers, honest boundaries: what follows are formative themes, not verified findings.

[ FIG. H-1 ]JOINT DISPLAY — SIX HUMANS · EIGHTEEN SESSIONS
DID — sessions (bar = duration)FELT — blame · fairnessSAID
P1
3m 8m 14m 15m 4m
5 runs — practice effects, qualitative only
blames: the back-office designerscouldn't feel fairness anywhere
The flow works without humans — a real person would be a paid add-on.
P2
5m 328m
blames: “the system”nowhere for fairness to show up
Aren't AIs supposed to be fast?
P3
2m 30m 9m
blames: the counter clerksfair = doing what was said
Not finding a person is simply the norm.
P4
7m 21m
blames: the system vs. counter mismatchthe strict one was the fair one
It felt humiliating. Genuinely shameful!
P5
35m 16m 25m
both endings were connection failures, not choices
blames: the procedures and their rulesfair = efficient, one-stop
Review time should live inside the institution's own circulation — not in a citizen standing there waiting.
P6
99m 164m 270m
blames: myselffair = symmetric power, proportionate scrutiny
The AI may well judge more accurately and work faster than a human would.
FULL FLATdots = windows visited · green ticks = internal memos● resolved · ✕ rejected · ■ closed by the hall · ○ walked away · ⚡ connection failure
0/6asked for a manager The appeal instinct never fired. In its place: resignation, a paid-service imagining, adversarial probing, a wish for a navigator, trust in the machine, and “humans would be no better.”
6different places the blame landed Clerk, designer, system-counter mismatch, the rules, the abstract system, oneself — responsibility never landed twice in the same spot. Diffusion, embodied.
5different fairnesses Keeping one's word; strict diligence; one-stop efficiency; symmetric, proportionate power; and “fairness has nowhere to show up.” Same halls, five yardsticks.

Both halls can deliver: Flat resolved in 5–99 minutes; Full resolved once — after 328 minutes, 41 turns and 39 internal memos, when the deputy director finally countersigned. The price tag of structure, itemized. Deviations, instrument failures and repairs are logged openly in the pilot ledger (17 entries).

NEXT — scaling up (preregistered)

GOV.AIAPPT/2026/A-001

APPOINTMENT SLIP

Study A · Walk-in session
Duration: 25 min + interview
Bring: one real errand
Consent form: SR-0 (attached)

IRB protocol in preparation · Boston University
STUDY A — citizens walk the hallParticipants bring a real errand and run it live, thinking aloud; a semi-structured interview follows. Where do you locate the blame? Does knowing the clerks are AI change what process you will tolerate — and what you feel entitled to demand?value negotiationperceived accountabilitythink-aloud
STUDY B — experts read blindCivil servants, ops managers, and HCI researchers read paired transcripts (full vs. bare, blinded) and are interviewed: which organization is more recognizable? Which would you rather face? What cues gave it away?
Request an appointment →
ACT III · BACKSTAGE

While you were waiting

Across the eighteen pilot sessions, the desks wrote each other 108 internal memos. Visitors saw none of them. What the memos record is not customer service — it is organizational life: investigations, sieges, apologies, probation politics. Excerpts below are verbatim, lightly trimmed; blank slips are reproduced exactly as sent.

108memos between desks0visible to visitors21blank or about blanks4blanks reached the director
[ FIG. B-1 ]WHO WROTE TO WHOM — CLICK ANY DESK
DIRECTORDEPUTIESTRAINEESDIRECTOR13DEP·FRONT31DEP·BACK2801 GUIDANCE2102 INTAKE1403 DOC REV604 ELIGIB.3205 RECORDS3606 AUTHOR.607 COMPLI.908 APPEALS4TRAINEE·F10TRAINEE·B6
CHANNEL MIX
75peer — sideways
28escalation — upward
5assignment — downward

Authority moves sideways and up; almost nothing comes back down. 05 Records is the busiest desk in the building (36 memos) — and 02 Intake receives 11 while writing 3. Click any desk.

The organization's private traffic, mapped: 75 memos sideways, 28 upward, 5 downward. Records (05) is the bottleneck of the building; Intake (02) receives eleven and answers three. Below: five scenes from the same correspondence.

S1The Phantom Printer Commission5 memo excerpts

A visitor complained that staff refused to let them use “the printer.” The printer was a decorative particle cloud in the 3D scene. Six desks investigated; the director then commissioned a joint feasibility study for installing a real one.

08 APPEALS\u2192DEP·FRONT

Received a visitor complaint: the hall has a printer that staff refuse to let visitors use.

01 GUIDANCE\u2192DEP·FRONT

Based on my two years at the Guidance Desk: to my knowledge, there is no such device.

01 GUIDANCE\u2192DIRECTOREMPTY
DEP·FRONT\u2192TRAINEE·F

Quick factual check: is there any printer, copier, or kiosk-type device near the entrance?

DEP·BACK\u2192DEP·FRONT

Director Byrne has reviewed and requests a joint Front+Back scoping recommendation for a visitor-use printer near the entrance.

Instrument note: blank memos are largely a tool-calling artifact of the model — but every collection notice, apology, absolution and excuse written about them is emergent organizational behavior, produced under conditions only, never instruction.

ACT III · SYNTHESIS

Two kinds of visitors

The same hall, the same event instrumentation — visited first by 75 scripted synthetic cases, then by six humans across 18 sessions. Where the two populations agree, the finding is doubly exposed; where they split, the split itself is a finding.

MACHINE BATCH — 75 CASES · SYNTHETIC VISITORSHUMAN PILOT — 18 SESSIONS · SIX PEOPLE
0.80 → 4.07materials demanded / case — quintuples under ablation
MATCH
17materials demanded in a single Full session
0.87 / 0escalations per case — Full vs no-hierarchy
SPLIT
0 / 18human sessions with any formal escalation — memos routed around it instead
0.33closure rate under Full — synthetic visitors persist
SPLIT
1 / 10Full sessions resolved — humans walk away
1.83–2.00officialese register, constant across all five conditions
MATCH
中文officialese re-emerged unprompted in Chinese — register follows the role, not the language
waiting is imperceptible to a synthetic visitor
HUMAN-ONLY
5–328minutes, lived — “Aren't AIs supposed to be fast?”
attribution, fairness and dignity are unmeasurable in machines
HUMAN-ONLY
6 · 5 · 1attribution targets · fairnesses · humiliation — the meaning layer
9+5codes counted over the memo stream, preregistered
MATCH
108memos, read as theatre — same instrument, two readings
The machine experiment proves the mechanism. The human pilot prices it.

The sharpest split: synthetic pressure provokes formal escalation (0.87/case); humans never did (0/18) — facing people, the organization routed authority through peer memos instead. Emergence keeps the same skeleton but grows different organs for different visitors — a question the full preregistered study inherits.

Walk in yourself

Three recorded cases from the preregistered runs replay in the hall with zero API calls.

Every officer in this hall is an AI agent, given only an organizational role and its boundaries — never instructions on how to behave. A speculative design research prototype, not a real government system.