Uber's AI Budget Blowout—and Recovery—Reveals How Software Will Eat Inference Costs

By
Lakshmi Reddy
1 min read

Uber Technologies burned through its entire annual AI coding budget by April. Four months later, its weekly agent requests have grown 9.4 times over February levels, more than 70% of code-change submissions involve autonomous agents, and engineers run upward of 30,000 agent tasks daily. Aggregate AI spending, meanwhile, has barely moved. Cost per 1,000 requests has dropped nearly 34% from April's peak; cost per session is down 52% from its June high.

The sequence matters: explosion, crisis, engineering discipline, and then sustained scaling at flat cost. Uber's CTO Praveen Neppalli Naga has called it the end of the "tokenmaxxing era." The company's internal AI Gateway now routes calls across OpenAI, Anthropic, and open-weight models like Llama and Mixtral, while prompt caching, developer-facing cost dashboards, per-tool spending caps of $1,500 per month, and continuous model benchmarking compress the bill from multiple directions at once.

The Software Tape Tells a Bigger Story

Uber reported on a day when enterprise software broadly rallied. Okta surged roughly 20–25% after posting GAAP gross margins of 80%, up from 77% a year ago, with GAAP operating margin doubling from 6% to 13%. Salesforce rose approximately 17%—though for different reasons. Salesforce's Q2 FY27 GAAP gross margin actually contracted about 145 basis points year over year, to roughly 76.6%, and GAAP operating margin fell from 22.8% to 20.5%.

Investors bid up Salesforce shares because of demand signals: constant-currency cRPO accelerated to 14%, Agentforce plus Data 360 ARR approached $3.9 billion (up more than 210%), and Agentforce ARR alone topped $1.5 billion (up more than 240%). Adjusted EPS of $5.90 got a large boost from an Anthropic investment gain; strip that out and the figure was closer to $3.37. The BVP Emerging Cloud Index edged up about 0.5% on the session while the Nasdaq was roughly flat.

Two conclusions follow from this tape. First, the market rewarded AI monetization credibility and booking acceleration, and it did so independently of any near-term margin improvement. Second, aggregate cloud-software margins have not collapsed: the BVP constituent median gross margin sits around 75%, and a Benchmarkit dataset covering hundreds of B2B software companies puts 2025 median software gross margin at roughly 80%. The damage is concentrated in pure usage-based models, where median gross margin runs closer to 62%.

The Cost Curve Has Moved by Orders of Magnitude

Stanford's 2025 AI Index documented the cost of reaching GPT-3.5-level performance falling from about $20 per million tokens in late 2022 to $0.07 by October 2024—a 280-fold decline. Epoch AI estimates a 5–10× annual drop in the cost required to hit a fixed capability level, which compounds to roughly 33–44% deflation per quarter.

Current model pricing makes the routing opportunity concrete. Anthropic’s frontier Claude Opus 5 costs $5 per million input tokens and $25 per million output tokens; OpenAI’s cost-optimized GPT-5.6 Luna costs $0.20 input, $0.02 cached input, and $1.20 output per million tokens; and Google’s Gemini 3.5 Flash-Lite costs $0.30 input, $0.03 cached input, and $2.50 output per million tokens. On uncached input alone, that puts the spread between a frontier model and current economy tiers at roughly 17–25×. For a simplified action using 2,000 input and 500 output tokens, Opus 5 costs about $0.0225, GPT-5.6 Luna about $0.0010—or roughly $0.00064 if all input tokens are cache reads—and Gemini 3.5 Flash-Lite about $0.00185, or roughly $0.00131 with cached input, excluding cache-storage charges. Salesforce’s Agentforce remains $0.10 per standard action through Flex Credits ($500 per 100,000 credits, with 20 credits consumed per action), $2 per conversation, alongside per-user and flat-fee pricing structures. At the $0.10 action price, a single-call GPT-5.6 Luna route leaves roughly 99% of revenue after model-token cost alone; actual contribution margin will be lower once multi-call reasoning, retrieval, tool execution, observability, security, and other workflow overhead are included.

The Real Bear Case Is Jevons, and It Got Stronger

If unit costs halve but agent intensity rises 9.4× again, total COGS still explodes. A workflow that goes from three model calls at $0.01 each to fifty calls at $0.005 has cut the per-call price in half—and increased total inference cost more than eightfold. Uber's CTO has acknowledged this arithmetic, and the company's COO has said it remains hard to draw a direct line from token consumption to proportional gains in customer-facing product output. Gross-margin expansion requires unit-cost deflation to outrun inference intensity per dollar of revenue. That is the governing equation, and nothing in today's data guarantees it holds permanently.

Who Owns the Spread

Salesforce disclosed something worth more than any single margin datapoint: for Agentforce Vibes, the customer's Flex Credit or flat-fee price does not change depending on which underlying AI model Salesforce selects. The customer pays for the business outcome. Salesforce procures intelligence underneath an abstraction layer and pockets the difference when it routes to a cheaper model.

This is the mechanism that cloud platforms used against hardware vendors and payment networks used against bank rails: sell the outcome at a price anchored to business value, then continuously reprice the supply chain. The company that can classify an incoming request—hard reasoning to frontier, routine extraction to a nano or open-weight tier—captures the spread between a fixed customer price and a falling input cost. Capital markets are pricing this insight aggressively: Stripe's reported $8 billion acquisition of OpenRouter, at roughly 57× annualized revenue, puts an extraordinary premium on the model-selection layer itself. Fireworks AI, now at over $1 billion annualized revenue with 95% of its 40 trillion daily tokens running specialized models, raised at a $17.5 billion valuation. Baseten and Together AI have raised at $13 billion and $8.3 billion respectively, on similar inference-arbitrage logic.

The highest-quality position in this value chain, then, belongs to whoever controls workflow, proprietary customer data, identity, authorization, and the customer-facing price—while treating model invocation as a substitutable commodity input. Investors tracking this thesis should ignore aggregate GAAP gross margin, which is polluted by services mix, acquisitions, and amortization. Track instead cost per successful AI action, premium-model traffic share, model calls per completed workflow, and the widening spread between customer price and underlying compute cost. Those metrics will separate the companies that own the spread from those still paying retail for every token.

not investment advice

You May Also Like

This article is submitted by our user under the News Submission Rules and Guidelines. The cover photo is computer generated art for illustrative purposes only; not indicative of factual content. If you believe this article infringes upon copyright rights, please do not hesitate to report it by sending an email to us. Your vigilance and cooperation are invaluable in helping us maintain a respectful and legally compliant community.

Subscribe to our Newsletter

Get the latest in enterprise business and tech with exclusive peeks at our new offerings

We use cookies on our website to enable certain functions, to provide more relevant information to you and to optimize your experience on our website. Further information can be found in our Privacy Policy and our Terms of Service . Mandatory information can be found in the legal notice