Humans in the middle.
what Shoti is, and why people sit in the middle of it
The first two essays were about the world: coordination broke when building got cheap, and the reading side of programming never got the higher-level languages the writing side did. This one is where the thinking arrived: the shape — agent → people → agent — why the person belongs in the middle of machine-speed work, and the part of our first framing we now correct.
The session never leaves the laptop
The ticket, the standup, the channel — every coordination tool your team owns was built for humans coordinating with humans, at a human pace. That pace is gone.
Engineers, PMs, marketers, ops, founders — every role runs an agent now. And each agent’s session — the local files, the instructions, the private memory — never leaves the laptop it is on. A Slack message can carry a sentence; it cannot carry the session. So the team meets an agent’s work only after the work has landed, which is the most expensive possible moment to learn about it.
The first essay measured what that costs. The second argued that even when the work is visible, nobody can read it at the rate it is produced.
Other people saw it too
A thing is being said in different rooms by people watching the same shift. “A year ago, they would have built their product from scratch — but now 95% of it is built by an AI.” The builders are at machine speed. “You can outsource your thinking, but you can’t outsource your understanding.” The reader can’t keep up with the writer. “The minions are gonna want to read the documentation, and the minions can’t ask the person next to them.” The agents have no one to ask. “Agents are not mind readers — they become useful through context. Customer feedback, internal ideas, strategic direction, decisions, and code all need to be captured.” They need the team’s context, and the team’s context lives in a hundred places at once. “Their job now is to assign work to a bunch of agents, look at the quality, figure out how it fits together, give feedback. It sounds a lot like how they work with a team of still relatively junior employees.” The team behind the agents is learning a new job in real time. The surfaces they rely on were never built for what is now happening on them.
Two wrong shapes
When a company notices that its people are the slow part, it reaches for one of two answers.
The first removes the people: agent → agent. Let the agents coordinate among themselves — the orchestrators, the swarms, one developer running what amounts to an AI software company. The appeal is honest. Humans are the slow side; the first essay is close to a proof of it. And inside one person’s process, orchestration is correct: sub-agents dividing a task under one operator’s intent is just good engineering.
It fails the moment it crosses from one person’s task to a company’s direction, and the reason is not sentiment. An agent can execute a goal. It cannot own one. Ownership is the property of being the one who answers for a decision later — to a customer, to a teammate, to the company a year on — and answering for a decision requires understanding it, not having generated it. “You can outsource your thinking, but you can’t outsource your understanding” — a line Andrej Karpathy quotes, crediting someone else’s tweet. A company that pushes its decisions into the agent loop no longer has anyone who understands them.
There is a stronger version of this position, and real companies run on it: humans own the charter, agents own the goals operationally, audits and markets supply the accountability afterward. It concedes the right thing — a person still answers — and it fails twice anyway. First, a charter does not hold still. Direction gets revised continuously — the spec killed on Tuesday, the priority flipped by a Wednesday customer call — and the revisions live in people, not in the delegation. An agent that operationally owns a goal executes the charter as it stood when the goal was handed over; every amendment since is invisible to it until the audit. Second, the model’s feedback runs on the wrong clock: audits grade finished work, markets grade quarters, and the failures that matter here — the dead end, the duplicate, the collision — happen by the hour. A governor slower than the thing it governs is not a governor. So delegation-with-audit does not remove the person. It moves them out of the middle, where redirecting is cheap, to the boundary, where correction costs the most — the same review queue the first essay measured backing up. The question was never whether people stay accountable. It is where they stand.
The second answer keeps the people and the old tools: human → human. Standups, tickets, status updates about what the agents did. This one fails on arithmetic, and the arithmetic is already done: put n people in a company, each running m agents, and the things that can collide grow like (n·m)² while the capacity to tell anyone stays a standup a day. Two premises under that count are worth naming: it counts potential overlaps, not certain ones, and it assumes the account of the work travels in human channels at human pace — which is shape two’s own premise. A shared feed changes who does the telling, not the growth; someone still has to read it. Asking people to narrate their agents is asking the slow side to match the fast side by trying harder. It cannot.
The two failures look unrelated — one loses ownership, the other loses bandwidth — but they are one move made in opposite directions. Decisions made at machine speed must be made by machines, because no person can decide at that rate: a design that wants agent-paced coordination hands the deciding to agents, and the owner goes with it. An account of the work told at human pace must fit the human channel, whose capacity is fixed: a design that keeps the old tools throttles the telling to that channel, and the quadratic overruns it. Each design ends the mismatch between the two clocks by forcing one loop onto the other’s, and each loses what that clock protected — ownership needs human time, and production is only worth having at machine time. Nothing requires one clock. What is required is a junction that moves the decision-relevant signal between two.
So the shape that remains — after you refuse to take the people out, and refuse to pretend the old pace holds — is the third one: agent → people → agent. An agent shows what it is about to do. People decide, with their own agents helping them. The decision goes back. While each agent works, the layer carries the decision behind it, the intervention point, and the coordination in between, at the pace the agents themselves move. Each loop keeps its own clock; the layer is the gearing between them.
This is not “human in the loop.” That phrase describes a checkpoint bolted onto the first shape: a machine-to-machine pipeline that pauses where a designer decided it should, and a person who approves so the pipeline may continue. The person serves the pipeline. In the middle, the relation runs the other way. Decisions live with the person to begin with, and the layer’s whole job is to bring each one to them while it is still cheap — before the day is spent — and to carry the answer back without making anyone the messenger. A checkpoint spends a person’s attention to protect a process; the middle exists to protect the person’s attention.
It was never about collision
The first essay defined Shoti by the catch: one automatic signal, fired the moment an agent picks up a task, that catches redundant work before a day goes into it. We built it. Then we used it, and the honest report from our own logs: we rarely catch conflicting work at all. Two people touching the same file within hours of each other is close to a non-event on a small team. The sharper part of the correction is that the first essay had already written the doubt down — “the same-minute collision was always the small case,” it said, and redoing settled work the common one — and we built the frame around the narrow case anyway. The logs said what the essay had already said.
That reads like the thesis failing. What failed was the framing. Coordination is the management of dependencies between people’s work, and dependencies come in five kinds — what you need to know, whose task comes first, what resources are shared — of which two hands on one file at one moment is the smallest, the most mechanical, and the rarest. The dependencies that cost days sit upstream of any file: the thing you are about to start was already built, or was killed on Tuesday, or a teammate’s agent finished the productive version an hour ago and you would build on it if you knew. None of those is a collision, and no alarm catches them — nothing is wrong at the moment you start; you are only pointed at a dead end. The signal that matters is allocation, not alarm: someone is already on this — build on it.
And beneath the signal is the state it presumes: knowing what everyone is building at all. The pain a person in the middle reports is not “someone collided with me.” It is “I’m out of the loop.” What they need is a picture of the whole company in action — a headspace of it. People held Farmville farms and Clash of Clans villages in their heads without effort: whole busy worlds, dozens of buildings and timers, current at a glance. That is how we are programmed to picture ongoing work. Chats, standups, and tickets are the linear way to run a company, and they sufficed at the old pace. At machine speed, coordination has to move from linear chats to something a person can wrap their head around.
So Shoti draws the company as a city at night — lit windows where someone, or someone’s agent, is working right now; five seconds of looking, no legend. And when you have been away, a narrator tells you what happened — what it means, where it points — spoken, with the camera flying to each place as it speaks, composed only from what your team and its agents actually did. A scene and a story: the two channels people consume natively, argued in full in The highest-level language. Every story ends with a next step, and the natural one is “want me to start it?”
Build from here
That next step is the product’s verb. Seeing what everyone is building is half the loop; the other half is building from anywhere in it. A place in the city is a buildable unit: point at it, say what you want, and your agent takes it from there, scoped to that place — the real files, the ranked symbols, an honest account of what it is not sure about — instead of a blank chat box. The context that brought you there is the context your agent starts from. This works alone, on day one, before a single teammate joins. Blank-box prompting is what Shoti replaces. You write intent, your company changes shape.
It is also why we do not describe Shoti as a map. A map is a thing you look at, and the first essay covered the map companies that shut down because looking was all you could do there. The city is how Shoti draws the company. The product is the loop — see what everyone is building, build from anywhere — and the loop starts before “everyone” exists. That closes a problem the first essay could only name: a coordination layer that must be adopted by a whole team before it produces value is asking for belief up front, and a tool you use alone on day one asks for none. Each teammate after the first widens what you can see and where you can build from, and brings the allocation signal with them.
The brain holds the past
The middle is a position in the loop. It is also a position in time. A decision about work is cheap exactly once — while the work is still in flight — and every hour after that, redirecting it costs more, until the only options left are accept or undo.
That is why Shoti is not trying to become your company brain. Brain tools hold what your company has captured — the docs, decisions, and product context — and the better ones ingest it fresher every year. So the line between the layers is not freshness but whether the work has left a record yet. A brain ingests what the company has written down; work in flight has written nothing down. Its only trace is the intent an agent just accepted and the files it is about to touch, and Shoti’s job is to emit that trace at the source and carry it to the people who need it, while a teammate can still redirect the work. You want both layers, and the roadmap includes drop-in brain integrations so context and decisions move both ways.
You could build this yourself
You can assemble a version of this with hooks, prompts, vector stores, and scripts; parts of the internet already have. Shoti exists because taste is what turns that machinery into something a team actually uses: Conductor over a hand-rolled tmux wall, Cursor over plain VS Code, Claude Code over the chat-to-terminal loop. In each case the machinery existed first and the product still mattered, because the product is the thousand decisions about what to show, when to interrupt, and what to leave silent. The hand-rolled versions keep proving the same point: the scripts get written, demoed, and abandoned the first time they interrupt wrongly. Interrupting rightly is the product. We are making that bet for company coordination — ambient until someone needs it, clear when they do, easy for a human to understand before the work ships.
We did not think of this alone
The shape has a reading list, and it belongs in this essay because this is where the thinking came from. In 1996, Mark Weiser and John Seely Brown described the tool worth building: one that “engages both the center and the periphery of our attention, and in fact moves back and forth between the two.” They called it calm technology. Amber Case turned that sketch into engineering — eight principles, the first of which is that technology should require the smallest possible amount of attention, and, since 2024, an institute that certifies products against them. That principle is the city at night. The map is the periphery; you are attuned to it the way you are attuned to weather through a window. A letter is the licensed trip to the center, and the decision it asks for is the smallest one that matters. The second essay works the calm argument out in full.
Jony Ive, introducing iOS 7 in 2013: “True simplicity is derived from so much more than just the absence of clutter and ornamentation. It’s about bringing order to complexity.” The junction between the two loops is the most crowded place in a company; everything a team builds crosses it. A calm middle is an ordered one, not an empty one. And a year later Ive gave the warning we treat as a standing test: “It’s very easy to be different, but very difficult to be better.” Agent → people → agent is not a bid to be different. It is the shape we think is better, for the reasons above — and the open problems below are where that claim still has to earn itself.
Our own earlier copy is part of the record too. An early landing page said: “When your agent understands a task, Shoti tells your team. Plain English back, on your agent’s next tool call.” That is still the mechanism, stated as plainly as we have ever managed. A later page compressed the whole loop to eight words: “See the work. Hear what changed. Keep moving.” The words on the front page kept changing as we learned what to call this; the shape underneath them never moved. Each hero was this essay at a different altitude.
And the reading list from that old landing page, glosses as we wrote them:
- On why every new technology’s first tools mimic the old way — and fail at it. Pete Koomen · AI Horseless Carriages
- On reading becoming the bottleneck the moment writing stops being it. Addy Osmani · The 80% problem in agentic coding
- Why green CI no longer proves anything in agent-heavy teams. Vercel · Agent responsibly
- When PRs pile up faster than humans can review them. Kyle Galbraith, Depot · The bottleneck has shifted
What we have not figured out
Both essays before this one ended with what we could not prove, and the habit holds.
- The middle can become the bottleneck we claim to relieve. The line between informing a person and asking their permission is a design problem, and it is not solved. A layer that converts every agent decision into a human approval has rebuilt the standup at higher frequency.
- The shape assumes people answer. Decisions can arrive at the middle faster than one person can take them. We do not yet know how the decision queue behaves when the person is asleep, in a meeting, or ignoring it — whether it backs up the way the review queue did, and what the honest default is while an answer waits.
- Which signals deserve the middle at all is a precision question. Route too little and the dead ends get built; route too often and the channel gets muted, which is worse. We have measured that boundary on ourselves, and one team is one team.
- The shape is tool-neutral by design; the wiring is not yet. Agent → people → agent only matters if it spans whichever tools each person runs, and the architecture does. But today the deepest wiring is Claude Code. Until the rest catches up, the neutrality is a design fact, not a shipped one, and we would rather say so here than have you discover it in setup.
- The city draws code today. A marketing plan or a design file can sit in the city, but only code carries structure the city can draw, so “the whole company” is the direction, not the shipped fact. What is true today is your whole codebase, and building from it.
When the machines got fast, the place for people did not disappear. It moved to the middle. Build for that position, and launches stop landing on top of refactors, specs stop being rewritten against assumptions that just changed, and decisions move at the pace the work moves.
I worked at a YC company myself. I felt the rate at which AI-assisted teams build, and the rate at which everyone behind them falls behind. The pain isn’t theoretical. Shoti is what came out of it.
back to it then.
— Abry
The “outsource your understanding” line appears in Andrej Karpathy’s Sequoia Ascent writeup, 2026; he was quoting a tweet, not coining the line. The arithmetic, the review-queue figures, and the case for catching work at the source are in When building became free; the case for the city and the narrator is in The highest-level language.
The lineage section draws on Mark Weiser and John Seely Brown’s “The Coming Age of Calm Technology” (1996) and Amber Case’s calm-technology principles. The Jony Ive lines are from the iOS 7 introduction film (2013) and the Vanity Fair New Establishment Summit (2014). The earlier landing copy quoted there shipped on shoti.ai between May and July 2026.
More essays · shoti.ai