
Anthropic Hires Google’s TPU Chief as Frontier Labs Race to Break Sole-Source GPU Pricing
Anthropic hired Amir Salek on August 21 to lead its internal semiconductor effort. Salek ran Google's custom-silicon organization and shepherded seven generations of TPUs from internal experiment to production platform. He reports to compute lead James Bradbury. The hire caps a rapid three-step sequence: in July, The Information reported Anthropic had opened early foundry discussions with Samsung around a 2nm chip with advanced packaging; on August 5, Anthropic publicly confirmed it was building an internal silicon team; and now, the executive who built Google's TPU franchise sits inside Anthropic's compute organization.
OpenAI, meanwhile, is further along. Its Jalapeño inference processor, designed with Broadcom and integrated by Celestica, already runs ML workloads — including GPT-5.3-Codex-Spark — at target production frequency and power. OpenAI claims substantially better performance per watt in early lab testing. Initial deployment is slated for late 2026, with Broadcom disclosing a contractual commitment covering 1.3 GW for 2027. Meaningful volume, by OpenAI's own admission, arrives next year.
The Reservation-Price Mechanism
Custom silicon does not need majority workload share to alter Nvidia's economics. If a lab can move 20–30% of predictable, high-volume inference onto captive chips, that capacity becomes its walk-away price in GPU negotiations. Nvidia can still win the workload — but must justify its premium through flexibility, throughput, latency, or software. The ASIC acts as a price ceiling on the marginal inference token even at modest penetration.
And power amplifies the incentive beyond chip cost alone. At a power-constrained 1 GW campus, a 25% gain in useful tokens per watt creates the economic equivalent of 250 MW of additional compute without another grid connection. When electricity is the binding constraint, efficiency gains become revenue capacity.
Nvidia's Financial Record Says Otherwise — For Now
Nvidia's most recent quarter produced $81.6 billion in total revenue, $75.2 billion from Data Center, and 74.9% GAAP gross margin. Independent estimates suggest Nvidia's inference-chip revenue share actually rose from roughly 66% to 74% over the past year. Custom silicon and Nvidia growth are, at the moment, happening simultaneously inside an exploding market. There is no consolidated financial evidence of an ASIC-driven margin break.
The bear case requires ASIC substitution to outrun market growth. That has not happened yet.
The Dual-Sourcing Chain Reaction
Google's expanded Marvell agreement this week — covering TPUs and infrastructure silicon, with warrants worth up to roughly $12.2 billion — sent a pointed message. Marvell surged ~10%; Broadcom fell 4.6%. The hyperscalers have no interest in swapping dependence on Nvidia for dependence on Broadcom. They want competitive tension at every silicon layer: Broadcom plus Marvell plus internal architecture plus foundry choice.
Broadcom's AI semiconductor revenue hit $10.8 billion in fiscal Q2 (+143% YoY), with Q3 guided to $16 billion and FY2027 targeted above $100 billion. The revenue trajectory is enormous. The margin trajectory, under deliberate customer-driven dual sourcing, is a separate question.
Who Absorbs the Damage First
The initial pressure likely lands on leveraged GPU lessors, not on Nvidia's income statement. When stable inference migrates to captive silicon, released GPU-hours become marginal capacity. Hourly clearing rates and utilization decline. Debt service and depreciation do not. CoreWeave's Q2 illustrates the structure: $2.575 billion revenue, $1.51 billion adjusted EBITDA, but $640 million interest expense, $9.4 billion capex, and a $626 million net loss — with roughly 72% of revenue concentrated in three customers.
A 10–20 point utilization decline on older GPU fleets can damage neocloud equity long before it registers in Nvidia's consolidated gross margin.
Own the Scheduler, Auction Every Token
The highest-information buyers — Anthropic, OpenAI, Google — are each constructing heterogeneous compute portfolios paired with a proprietary control plane. Anthropic already runs on TPU, Trainium, Nvidia, and reportedly Fractile, with its own ASIC years away. OpenAI runs merchant GPUs alongside Jalapeño. Google balances internal architecture against two external design partners.
The winning position is the scheduling layer that routes each token request to the cheapest accelerator meeting the latency and quality target. That layer produces three distinct advantages at once: it exploits stable workloads on specialized silicon while holding GPUs for volatile architectures; it grounds every procurement negotiation in a real alternative; and it allocates work by tokens per megawatt when power is scarce.
Qualcomm's July acquisition of Modular and d-Matrix's purchase of Wallaroo.ai confirm that this abstraction layer is becoming the contested strategic asset. The chip matters. The compiler and workload router that decides which chip serves which token may matter more.
not investment advice