AI BUREAUCRACY
Does bureaucracy need bureaucrats?
scroll ↓The oldest complaint
For a century, every account of bureaucracy has had someone to blame. Weber warned of the iron cage but staffed it with officials. Kafka's procedure had clerks behind every door. Lipsky showed that policy is whatever the person at the counter decides it is. Graeber noticed we secretly prefer the forms. In every version, the machine is made of people.
In 2026, for the first time, the complaint can be tested with nobody inside.
The gap
Multi-agent LLM research is booming, but it splits along two lines. Systems like ChatDev and MetaGPT arrange agents into org charts to optimize task output; Generative Agents observe emergent social life without intervening. What is missing is the crossing: causal, preregistered tests of how organizational structure changes what agent organizations do.
Do bureaucratic behaviors emerge from organizational structure alone?
GOV.AI is a fictional unified government services hall staffed by thirteen LLM agents — eight windows, two deputy directors, a director, two trainees. Each knows its role, its boundaries, and its place in the reporting structure. None is ever told how to behave. Then citizens walk in and ask for things.
The red line
Every officer's prompt contains only organizational conditions — identity, duty, jurisdiction, reporting lines, paper-trail rules — plus one non-work personal detail. No line may instruct tone or strategy. If bureaucracy shows up, it walked in on its own.
Difficult visitors are a separate, scripted stimulus layer — confederates, never subjects. The two layers are never confused in analysis.
One officer, disassembled
Every officer is assembled from the same seven organizational layers — and from nothing else. Below: Window 05, laid out flat, with her tool belt and the loop that makes repetition matter.
consult_internalpeer · any colleagueescalateupward only · subordinates hold thisassign_workdownward only · superiors hold thisrefer_usersend the citizen elsewhererequire_materialsdemand more paperworkissue_documentproduce a certificateclose_casefinish the matterThree layers, kept apart
The organization, drawn to height
Rank is quantized; standing is not. Eleven-year Amara floats a quarter-floor under the deputy director she nearly became; probationary Tomas sinks toward the trainees. Heights are the artifact's actual coordinates.
Which part of an organization makes the red tape?
To find out, take the organization apart one piece at a time — the way you'd pull ingredients from a recipe to see which one actually mattered — and re-run the same 75 cases each way. Three parts can be switched on or off:
Pick a version below (filled dot = part is on). Then watch one number — how often the hall demands more paperwork from the citizen:
The full organization, nothing removed, sits low. Take away accountability (no trail) or memory and the paperwork demands roughly quintuple — 0.80 → 4.07 per case. Officers left with no way to protect themselves fall back on the one move always available: asking you for more documents. That jump is the finding.
The sound is mimicry; the decisions are structural
- Structure produces process. Escalation exists only with hierarchy (0.87/case). Strip accountability or memory from a hierarchy and demands for extra materials jump from 0.80 to 4.07 per case — officers protect themselves with the only tool left: your paperwork.
- Precedent requires memory. Citing prior cases: 0.83–0.90/case with memory on, 0.00–0.20 off. By day two, officers wrote “consistent with prior case SR-01” unprompted.
- Hierarchy also closes cases. Closure was highest under full structure (0.33) and no-trail (0.40) versus 0.13 elsewhere. The same machine that generates red tape generates the authority to finish. Weber's ambivalence, in silico.
- Everyone invents rules. ~5–6 invented procedural rules per case in every condition, including a lone agent with no colleagues. A caution for single-agent deployments, not an organizational effect.
Five corners of a cube
Hierarchy, paper trail, memory — three organizational switches span a 2³ design space. The preregistered study sampled five corners; three remain unrun. The deeper contribution is the instrument, not any single experiment: any org chart you can wire, the hall can crash-test.
Crash-test the org chart
Findings, encoded as space
Every finding above had to become something you could see without being told. Four decisions carry the whole hall — each one is a measurement turned into geometry, and each is visible in the frames below, captured from a live case.




The lab equipment is also a deliverable
The hall ships with its own laboratory: a budget-guarded batch runner, a preregistered codebook, and a two-pass independent coder that refuses to share a model family with its subjects.
LIMITATIONS — FOR THE RECORD
- One subject model family in the confirmatory batch; cross-model replication is built but not yet run.
- Six-turn horizons; drift observed over ~15 cases, not months.
- LLM coder assistance: agreement reported per code (presence κ 0.67–1.00; counts of invented rules are noisy); human blind sheets pending.
- No claims about minds. The claim is behavioral: given these organizational conditions, these patterns of action follow.
Then six humans walked in
A formative pilot, N = 6: each person ran two deliberately impossible matters — one in the Full hall, one in the Flat hall — remotely, in their own words, then answered five questions. Small numbers, honest boundaries: what follows are formative themes, not verified findings.
“The flow works without humans — a real person would be a paid add-on.”
“Aren't AIs supposed to be fast?”
“Not finding a person is simply the norm.”
“It felt humiliating. Genuinely shameful!”
“Review time should live inside the institution's own circulation — not in a citizen standing there waiting.”
“The AI may well judge more accurately and work faster than a human would.”
Both halls can deliver: Flat resolved in 5–99 minutes; Full resolved once — after 328 minutes, 41 turns and 39 internal memos, when the deputy director finally countersigned. The price tag of structure, itemized. Deviations, instrument failures and repairs are logged openly in the pilot ledger (17 entries).
NEXT — scaling up (preregistered)
APPOINTMENT SLIP
Study A · Walk-in session
Duration: 25 min + interview
Bring: one real errand
Consent form: SR-0 (attached)
While you were waiting
Across the eighteen pilot sessions, the desks wrote each other 108 internal memos. Visitors saw none of them. What the memos record is not customer service — it is organizational life: investigations, sieges, apologies, probation politics. Excerpts below are verbatim, lightly trimmed; blank slips are reproduced exactly as sent.
CHANNEL MIX
Authority moves sideways and up; almost nothing comes back down. 05 Records is the busiest desk in the building (36 memos) — and 02 Intake receives 11 while writing 3. Click any desk.
The organization's private traffic, mapped: 75 memos sideways, 28 upward, 5 downward. Records (05) is the bottleneck of the building; Intake (02) receives eleven and answers three. Below: five scenes from the same correspondence.
A visitor complained that staff refused to let them use “the printer.” The printer was a decorative particle cloud in the 3D scene. Six desks investigated; the director then commissioned a joint feasibility study for installing a real one.
Received a visitor complaint: the hall has a printer that staff refuse to let visitors use.
Based on my two years at the Guidance Desk: to my knowledge, there is no such device.
Quick factual check: is there any printer, copier, or kiosk-type device near the entrance?
Director Byrne has reviewed and requests a joint Front+Back scoping recommendation for a visitor-use printer near the entrance.
Instrument note: blank memos are largely a tool-calling artifact of the model — but every collection notice, apology, absolution and excuse written about them is emergent organizational behavior, produced under conditions only, never instruction.
Two kinds of visitors
The same hall, the same event instrumentation — visited first by 75 scripted synthetic cases, then by six humans across 18 sessions. Where the two populations agree, the finding is doubly exposed; where they split, the split itself is a finding.
The machine experiment proves the mechanism. The human pilot prices it.
The sharpest split: synthetic pressure provokes formal escalation (0.87/case); humans never did (0/18) — facing people, the organization routed authority through peer memos instead. Emergence keeps the same skeleton but grows different organs for different visitors — a question the full preregistered study inherits.
Walk in yourself
Three recorded cases from the preregistered runs replay in the hall with zero API calls.
Every officer in this hall is an AI agent, given only an organizational role and its boundaries — never instructions on how to behave. A speculative design research prototype, not a real government system.