Uber cuts AI token costs as engineer adoption quadruples
Uber Technologies Inc. has reduced its cost per token for AI tools while quadrupling employee adoption since the start of the year. CTO Praveen Neppalli attributes this efficiency to engineering optimizations like prompt caching and model selection, signaling an end to the 'tokenmaxxing' era. This shift aligns with broader industry trends where leaders like Chamath Palihapitiya and Dario Amodei predict recursive self-improvement and increased automation in software engineering.

*this image is generated using AI for illustrative purposes only.
Uber Technologies Inc. has successfully lowered its cost per token for artificial intelligence usage while simultaneously expanding adoption across its engineering workforce, marking a strategic shift from volume consumption to operational efficiency. On Wednesday, Chief Technology Officer Praveen Neppalli disclosed that the company has more than quadrupled the number of people using frontier AI tools since the start of the year, with thousands of engineers now utilizing these systems daily. This expansion in usage occurred alongside a decline in per-token costs, a trend Neppalli described as a signal that the industry is exiting the so-called 'tokenmaxxing' era, where competitive advantage was driven by sheer token consumption.
The improvement in cost efficiency stems from targeted engineering optimizations rather than restrictions on tool access. Uber has enhanced its prompt caching and reuse systems to reduce input token spending, adjusted default model settings and context sizes, and provided engineers with real-time visibility into their AI usage and associated costs. Additionally, the company is testing open-weight AI models and selecting specific models based on distinct use cases to maximize value. "The next phase, whatever we call it, will not be characterized by who spends the most tokens, but about how people use them as efficiently as possible," Neppalli said, echoing comments made by Chief Financial Officer Balaji Krishnamurthy during the company's recent earnings call.
Strategic Shift in AI Usage
The company’s approach reflects a broader industry maturation where the focus is shifting from raw compute power to intelligent resource allocation. By providing engineers with data on their usage patterns, Uber enables teams to optimize their prompts and model selections, thereby reducing unnecessary expenditure without hindering productivity. The implementation of prompt caching allows the system to reuse previously processed inputs, significantly cutting down on redundant token generation. These technical adjustments have allowed Uber to scale its AI infrastructure sustainably, ensuring that increased adoption does not lead to proportional increases in operational costs.
| Metric | Status / Change |
|---|---|
| AI Tool Adoption | More than quadrupled since beginning of year |
| Cost Per Token | Declined |
| Primary Driver | Engineering optimizations (caching, model selection) |
| User Base | Thousands of engineers using AI daily |
Industry Context and Expert Views
Uber’s internal developments align with broader predictions from technology leaders regarding the evolution of artificial intelligence. Venture capitalist Chamath Palihapitiya recently suggested that AI may have entered a recursive self-improvement cycle, where systems help build more advanced models, accelerate breakthroughs, and reduce costs. This trend, he noted, could represent a step toward the AI singularity. Similarly, Anthropic CEO Dario Amodei predicted that AI could soon handle most software engineering tasks, observing that some engineers are already using AI to generate and edit code. AI pioneer Andrew Ng warned that engineers who fail to adopt AI tools risk falling behind, emphasizing that developers who combine technical experience with AI skills will be better positioned as the industry evolves.
What the Numbers Show
The divergence between rising adoption rates and falling unit costs at Uber highlights a critical efficiency gain in enterprise AI deployment. Typically, increased usage of cloud-based AI services leads to linear or exponential cost growth; however, Uber’s ability to decouple these variables suggests that initial inefficiencies in prompt engineering and model selection have been largely resolved. This optimization indicates that the marginal cost of adding new users to an AI ecosystem can be minimized through architectural improvements such as caching and context management. For investors, this signals that AI integration may become a scalable operational lever rather than a purely additive expense, potentially improving long-term margins as productivity gains outpace incremental costs.
How might Uber's success in decoupling AI adoption from cost growth influence the competitive landscape among other ride-hailing and logistics platforms?
What specific metrics will Uber use to quantify the productivity gains from AI-assisted coding to justify continued investment in these tools?
Could the industry-wide shift away from 'tokenmaxxing' pressure AI model providers to alter their pricing structures or feature offerings?
































