Imagine a factory floor where software agents are the workers, pull requests are the conveyor belt, and a cost dashboard blinks politely in Slack. At Uber, more than 70% of pull requests now involve local or cloud agents. That is less like hiring an assistant and more like adopting a hyper-efficient squirrel that can also debug your CI.
I remember asking a chatbot to explain a failing test. It gave a five-minute lecture on testing philosophy, then missed the obvious: I had left a debug print inside the assertion. My heart did a tiny drumroll. If I can forget to delete one print, an entire company of robots can definitely wander into token-costed mischief.
Uber’s story is practical. The team treats AI like a supply chain. Some models are the espresso machine: brilliant, expensive, best for hard decisions. Others are the paperclip dispenser: not glamorous, but useful. Their secret is a cost equation that measures price per token, tokens per request, and requests per turn. It is the software equivalent of noticing you kept paying for a parking meter after leaving the lot.
One delightful anecdote is the code-review agent. Uber built a benchmark from real pull requests that actually had bugs, grading them easy, medium, and hard like a chef scoring souffles. When they switched models, accuracy improved and cost dropped. The lesson is delicious: measure real work, not toy work. Benchmarking on ‘Hello, world’ is a love letter to your own ignorance.
Another relatable moment is context bloat. Have you ever opened a browser with forty-seven tabs, convinced one contains the answer to your life? Agents can do the same with tool schemas. Uber had thousands of MCP tools, and loading every schema added tens of thousands of tokens before anyone typed a sentence. That is like carrying a telephone book into a conversation where you only need your neighbor’s number.
So they made tools callable through a shell, letting agents fetch what they need only when needed. They also introduced code-mode, where a loop can run quietly instead of making the model narrate every step. If an agent once polled a query five times while sighing theatrically, now it runs a script and returns the summary. Less chit-chat, more receipts.
The human part is funny. There is a live cost counter in the terminal, like a fitness tracker for money. Slack nudges at 50, 80, and 100 percent of expected spend remind engineers that budgets are not enemies, merely disappointed relatives. No one is locked in a room with a stopwatch. Instead, the system says, gently, ‘You are spending more than a small coffee shop, and we do not know if this is a breakthrough or a typo.’
Uber’s optimization is that instinct scaled to millions of requests. Compact conversations, use caching, route subagents to cheaper models, and keep the expensive brain for hard questions. The result is dignity. Engineers stop shouting at machines, machines stop shouting at invoices, and the software factory hums like a dishwasher.
The Uber Software Factory, or How to Make Robots Tidy Their Own Desks










