Paddock20
Wireframe DigestAgentic Thought LeadershipJuly 11, 2026 · 7 min read

I Managed 5,000 People.
Now I Manage AI Agents.
The Playbook Did Not Change.

MIT found 95% of AI pilots return zero measurable ROI. The missing skill is management — and operators already own it.

Crew One

5,000 people
1,586 stores

AT&T retail through Spring Mobile. $1.2B in attributable revenue. Brooks University cut new-hire ramp 40%.

Crew Two

2 voice agents
9 pipelines

Live business lines screened and routed. DRX Lot Assistant™ runs $125M+ monthly inventory at duPont REGISTRY.

Twenty-six years sit between those two crews. The playbook does not.


Why 95% of Pilots Die

MIT's NANDA initiative studied enterprise generative AI and found 95% of pilots return zero measurable ROI. Fortune covered the report in August 2025 and the number traveled fast. Most takes blamed the models.

I read the finding as an operator. I watched pilots die in retail for two decades before AI was a line item. A pilot dies the same way every time: nobody owns it, nobody inspects it, and nobody measures it after launch week. That is not a technology failure. That is a management failure.

Google Cloud published its own guidance: grade production agents on operational KPIs, not demos. That is a management sentence. Every district manager I trained knew it by month three.

The industry keeps hiring engineers and teaching them management. There is a shorter path. Take a person who already manages and hand them agents.


The 3D Build System, on a Live Build

Every build runs through three steps. Here is how it shipped DRX Lot Assistant™.

Digest

Watch the work. Count it. Find the stall. The recon standard is a 6-day turn per car. Walk the lot the way you once walked stores. Trace cars from transporter drop to front line and write down every handoff. The stall was status: it lived in text threads and memory instead of a system, so the 6-day clock had no referee.

Develop

Write the spec before anything gets built. A spec is an SOP with a compiler waiting. Name who owns each step, what data gets written at that step, and by whom. One writer per data record — because two writers per record is how counts go bad, on a stock ledger or in a database. The acceptance gates get written here too: the exact checks the build must pass before it touches live inventory. Pass or fail. Nothing graded on effort.

Deliver

The coding agents build against the spec. Inspect output against the written standard, not against mood that day. Nothing ships until 100% of the gates pass. After ship, the scoreboard takes over, and a watchdog service pages when anything breaks. The same three steps shipped both voice agents and all nine pipelines. The system does not care what it is building. That is the point of a system.


Five Habits That Transfer Straight Across

1.

Written standards

Brooks University cut ramp time 40% for one reason: the standard lived on paper instead of in a trainer's head. Agents are the most literal staff you will ever direct. Hand one a vague brief and it builds your vagueness at machine speed. Hand it a written spec and it builds the spec.

2.

One owner per task

The Amazon delivery company ran 46 routes and 100 employees at peak with two management layers. Every route had one owner every day. The agent stack keeps the rule: one writer per data record, one agent per job. Shared ownership is zero ownership, for drivers and for pipelines.

3.

Inspection before done

You get what you inspect, not what you expect. In stores, inspection meant walking the floor with a checklist. In agent work it means acceptance gates that run before anything goes live. "The demo looked great" is not a gate. A gate is binary and written in advance.

4.

Scoreboard after ship

The delivery company moved 40,000 packages a week at 99.8% on-time, and the number held because every driver saw a weekly scorecard with their name on it. Agents get the same treatment. Each pipeline reports its counts, and the watchdog pages on a miss. Launch day is not the finish line. It is the first day the scoreboard counts.

5.

Fire-drill discipline

Peak season taught me that failures arrive on schedule, so you rehearse them in October, not December. Every agent has a written failure path before launch: the watchdog pages, the pipeline halts, the fallback takes over. A plan is only cheap before the outage.

Specs are SOPs. Gates are inspections. Guardrails are house rules. Scoreboards are scoreboards. I did not learn a new discipline. I pointed an old one at a new crew.


What This Means If You Run a Business

The loudest voices in agentic AI deserve their audience. Simon Willison writes from an engineer's seat. Steve Yegge writes from an engineer's seat. Ethan Mollick writes from a professor's seat, and his One Useful Thing essay calling management the AI superpower sits closest to this position. One line to add: none of those seats run inventory. This one does — 280 cars a month of it.

Meanwhile the tools moved toward you. Gartner projects citizen developers will outnumber professional developers 4 to 1. In Y Combinator's winter 2025 batch, TechCrunch reported 25% of startups shipped codebases that were 95% AI-generated. Founders with every engineering option available are renting the syntax. What they cannot rent is the judgment: what to build, who owns it, how to know it still works on a Tuesday in month six.

You built that judgment across years of Monday scoreboards. If you have run a store, a route network, a service lane, or a kitchen, you have written standards, assigned owners, inspected work before sign-off, and read numbers with names attached. Managing agents is that job. The typing is the rented part now.


Build With the Playbook

The scoping session maps one workflow end to end, produces a written spec, and sets the acceptance gates before any code is written. It is the first step in the 3D Build System — and it starts with a free 15-minute Fit Call.

Book a Free Fit CallRun the Free Diagnostic

Sources

Fortune (MIT NANDA)fortune.com/2025/08/18

Google Cloudcloud.google.com/transform (production agent KPIs)

One Useful Thingoneusefulthing.org (management as AI superpower)

Gartner via Quixyquixy.com (citizen developer projection)

TechCrunchtechcrunch.com/2025/03/06 (YC W25 AI-generated codebases)


Related Reading

The 6-Day Turn — DRX Lot Assistant as Proof CaseThe Tools Businesses Used to RentWhat Is Agentic Engineering — The Full DefinitionWhy 69% of AI Pilots Never Reach ProductionCase Studies — Production Software in the Wild

Build With RAIL: Free

90 Seconds to
Build Your App
With RAIL.

Answer three questions—type them or say them out loud—and watch RAIL scope your app in real time. Named app, working interface, and a build blueprint you keep. Before you ever talk to anyone.

A named app. A working interface. A blueprint PDF you keep. Whether we ever talk or not.

AI-powered90 secondsType or speakPDF blueprint
Build Your App Blueprint →Name + email to receive your PDF blueprint.