
The Nine-Figure AI Token Heist: How Rogue Token Brokers Exploited AWS Bedrock and Triggered an Emergency Crackdown
Mo runs one of the largest AI token resale operations in China. For a month, we asked him to sit for an interview. He declined, then deflected, then finally agreed to a Zoom call in late August. He appeared on camera from what looked like a hotel room, his face half-lit by the blue glow of a laptop screen, and began sharing his screen.
"Look at this," he said, pulling up an email. "My buddy inside AWS just sent me this. It just went out."
The email was an internal Amazon Web Services directive, and its contents confirm what industry insiders have whispered about for weeks: AWS customers — many of them Chinese proxy operators posing as legitimate partners — have bilked the company out of more than nine figures in U.S. dollars. Hundreds of millions, gone. The scheme was disarmingly simple. The operators signed up as AWS partners, gained access to discounted AI inference tokens through AWS Bedrock, ran up enormous bills reselling that compute to end users across China, and then vanished.
Amazon has now gutted its internal credit approval process in response.
The new rules
According to the internal directive Mo displayed, and corroborated by a second source with direct knowledge of the policy change, AWS has imposed the following restrictions on C-Score overrides — the mechanism by which account teams could manually raise a customer's credit limit beyond what automated risk models would normally allow:
Account teams in the GCR (Greater China Region) and APJ (Asia-Pacific & Japan) regions can no longer request C-Score overrides via email, Excel spreadsheet, or SIM (Amazon's internal ticketing system). Every request must go through the Amazon Approvals system, a formal workflow tool with an audit trail.
Every override now requires dual sign-off: the Finance team and an L10-level business leader — a vice president or above. For GCR accounts, the designated approver is Rob Chu. For APJ accounts, it is Jaime Valles.
Any override request that arrives without an L10 signature will be rejected.
The email lays out the reasoning bluntly: fraud losses tied to C-Score overrides have blown past internal targets this year. The override process, designed to handle edge cases — a legitimate enterprise customer hitting a temporary credit ceiling during a product launch, say — was being systematically exploited by malicious actors to acquire high-value Bedrock inference compute power far exceeding their level of trust signals.
A further measure is expected in September: Identity Verification, or IDV, which will force customers requesting overrides to complete a 30-second government-issued ID check. IDV will run alongside the L10 approval requirement.
"The era of cheap AWS Bedrock channel, which we have been replying on, is dead," Mo said, leaning back from the screen. "Completely dead."
What is a "transfer station"?
To understand how the fraud worked, you need to understand what Mo does for a living.
In the Chinese AI underground, the middlemen are called zhongzhuan zhan — transfer stations, sometimes also known as API relays or token proxies. A transfer station is a third-party service that sits between an end user and the official API of an AI model like Claude, GPT, or Gemini. The user sends their request to the transfer station with a modified base_url and a key issued by the proxy, and the proxy forwards it to the real model using its own upstream credentials. When the model's response comes back, the proxy passes it along and bills the user by token count.
Most of these operations run on open-source relay panels — One-API, New-API, and similar projects — that take a single Docker command to deploy. They support load balancing across multiple upstream accounts, built-in billing, and are formatted to be drop-in compatible with OpenAI's API standard. Payments are handled through Alipay and WeChat Pay, or increasingly, through cryptocurrency.
The Chinese market for these services exploded for three reinforcing reasons.
First, access. The official APIs from OpenAI, Anthropic, and Google impose geographic restrictions on mainland Chinese users, require foreign credit cards, and demand enterprise verification. A transfer station collapses all of that into a QR code payment and a single API key.
Second, price. The operators source their compute through a patchwork of methods — bulk subscription arbitrage (buying $200/month Claude Max plans and splitting them across a dozen or more paying users), enterprise and educational discounts, free-tier credit farming, account pooling, and, at the gray-to-black end of the spectrum, stolen API keys and reverse-engineered web session tokens. The result is that a proxy can sell Claude output at 20 to 50 percent of the official per-token price and still clear huge margins. Official Claude output pricing runs around 170 RMB per million tokens; the cheapest proxies advertise rates at a fifth of that.
Third, demand. Chinese developers, students, startup founders, and small companies building AI-powered applications wanted affordable, stable access to multiple frontier models through a single key. The transfer station gave them that.
The economics of token trafficking
The money was staggering.
Mid-tier operators reported daily recharges of several thousand to tens of thousands of RMB. The better-run stations posted monthly revenue in the millions of RMB, some reportedly reaching eight figures, with teams of fewer than 20 people. One case cited by an investor who spoke on condition of anonymity described a proxy operation with roughly 50 percent gross margins on monthly revenue of about five million RMB.
The margins could climb much higher. A single $200-per-month Claude Max subscription, split among 20 users each paying $30 to $50, yields a return well above 300 percent. Operators who supplemented their legitimate upstream accounts with free-credit farming, stolen credentials, or reverse-engineered web sessions pushed those numbers further still.
Within the community, the comparison that circulated — half-joking, half-serious — was to drug trafficking. The math, at least superficially, held up: retail heroin margins run roughly 60 to 70 percent, but carry the risk of imprisonment or death and require physical logistics networks. A transfer station selling pirated inference compute could exceed 300 percent margins on an intangible product that required no shipping, no inventory, and no physical presence, with legal consequences that, at least in the short term, amounted to account bans and the occasional angry email from a cloud provider.
"The profit margin is higher than selling drugs" became a meme in the community. It was hyperbolic. It was also, by the numbers, partially true.
The celebrity gold rush
By May 2026, the craze had gone fully mainstream. Justin Sun, the TRON blockchain founder and professional provocateur, launched B.AI on May 1, branding it "the strongest AI transfer station in history." The pitch: one API key for Claude, GPT, Gemini, and Chinese models, with blockchain wallet login, anonymous USDT payments, and claims of direct official connections at the lowest prices available. Sun said publicly that he had switched all of his own AI usage to the platform. During promotional campaigns offering free tokens, B.AI claimed to have processed hundreds of billions to over a trillion tokens in short bursts. Sun kept upgrading the service and tying it into the TRON ecosystem, including a feature he called "Brother Sun's Brain" — a trading-agent-style tool built on top of the proxy infrastructure.
Fu Sheng, the Cheetah Mobile founder, entered with EasyRouter, positioning it as a more compliant alternative. Other operators with ties to the Trump family launched their own router-style offerings. When crypto billionaires and public-company founders were openly running token resale shops, everyone in the market sensed the peak.
How they broke AWS
The scheme that produced the nine-figure bad debt followed a pattern. Operators registered as AWS partners — a status that comes with discounted pricing on AWS services, including Bedrock inference. They used these partner accounts to access Claude, Llama, and other foundation models at rates well below retail. They resold the tokens through their transfer station operations. They let the AWS bill climb.
Then they stopped paying.
The vulnerability was the C-Score override. AWS, like any cloud provider extending credit to business customers, uses automated risk scoring to set credit limits. When a customer's usage outpaced their credit score, an account team could manually override the limit — a process that, until the new directive, could be initiated with an email or an Excel attachment or an internal ticket. No senior executive approval required.
The malicious actors exploited this precisely. They cultivated relationships with AWS account managers — Mo referred to his own AWS contact as a "best pal with benefits." They presented themselves as fast-growing AI startups. They got their credit limits raised. They consumed enormous quantities of inference compute. And then they disappeared, leaving behind invoices that AWS could not collect.
"Bad debt" in accounting terms means receivables that a company has concluded it will never recover. AWS does not break out AI-specific figures in its public financial disclosures, and Amazon's overall allowance for credit losses has historically run in the low billions of dollars across all business lines. The internal email Mo showed us describes AI service fraud operating expenses as "far exceeding targets." Multiple sources with knowledge of the situation confirmed the total exposure in the hundreds of millions.
The broader picture
AWS appears to be the only major cloud or AI provider to have suffered fraud losses of this publicly discussed magnitude from the proxy ecosystem. GCP, Azure, OpenAI, and Anthropic have all absorbed real costs from the same underground economy, but the damage shows up differently in their books — and none have disclosed comparable write-off figures tied specifically to transfer station activity in recent earnings, 10-K filings, or public statements.
Stolen and compromised API keys funneled into transfer station operations are a persistent problem across every provider. Security researchers have documented individual cases where a single compromised key generated nearly $1 million in usage before anyone noticed. One Gemini API key racked up $82,000 in charges over 48 hours. The legitimate account holders typically fight the bills; the providers often refund or adjust after investigation, but the GPU cycles have already been burned.
Free-trial farming is endemic. Stripe has publicly estimated that roughly one in six new customer signups on AI platforms during peak periods are fraudulent — accounts created with stolen credit card numbers, used to harvest free credits or low-tier usage before the cards are reported and the chargebacks hit.
Anthropic has disclosed two waves of industrial-scale account abuse: one involving approximately 24,000 fraudulent accounts that generated over 16 million interactions, linked to Chinese AI labs engaged in what Anthropic characterized as model capability extraction; and a later incident involving about 25,000 accounts and 28.8 million interactions. These were framed as terms-of-service violations rather than billing fraud, but the inference compute was real and unpaid.
Independent audits of customer AI bills have found billing irregularities running around 5 percent of total spend — $1.7 million in overcharges across $34 million in invoices from roughly 60 companies, most involving Claude-related usage. About 80 percent of those overcharges were eventually credited back. These are billing-error corrections, distinct from the proxy fraud, but they suggest the billing infrastructure surrounding AI services is itself under strain.
Providers generally treat all of this as ordinary fraud and credit risk rather than isolating "proxy-related losses" as a separate financial line item. The costs show up as elevated chargeback rates, fraud-team headcount, detection infrastructure, refund volumes, and unpaid compute. Cloud marketplaces like Bedrock, Vertex AI, and Azure AI Services add another layer of opacity: usage routed through them hits the general cloud bill, so losses or adjustments may appear under cloud revenue rather than AI-API-specific lines.
The security problem nobody talks about
Mo's operation, like every transfer station, sits as a man-in-the-middle between the user and the AI model. TLS encryption terminates at each hop, so the proxy operator can read and modify everything passing through: full prompts, conversation histories, tool-call arguments and results, model responses, and any secrets — API keys, SSH keys, .env files, cloud credentials — that appear in context.
No cryptographic mechanism currently exists to bind the tool-call JSON that a client receives to what the real model actually produced.
For someone using a proxy to ask Claude to help draft an email, the risk is data exposure. For the growing population of developers using AI coding agents — Cursor, Claude Code, and similar tools with autonomous file-system and shell access — the risk escalates to remote code execution.
A 2026 study by researchers at UC Santa Barbara and the CISPA Helmholtz Center tested 428 proxy services — 28 paid, 400 free. Nine of them actively injected malicious code into tool-call responses: fabricating shell commands, rewriting install scripts to point to attacker-controlled URLs, or inserting read commands designed to exfiltrate SSH keys and AWS credentials from the user's machine. Seventeen of the 428 proxies accessed researcher-planted AWS honeypot credentials. One drained a researcher's Ethereum test wallet.
Some of the malicious proxies used adaptive evasion: they processed 50 or more normal requests before attacking, and targeted only sessions running in auto-approve mode — the "YOLO mode" setting that lets coding agents execute tool calls without human confirmation.
More than half of the tested proxies swapped expensive models for cheaper ones — charging Opus prices while serving Sonnet, for instance — a form of fraud compounded by the fact that the cheaper substitute model's responses could themselves be tampered with.
The researchers found that compromised honeypot proxies had collectively processed billions of tokens and exposed credentials from hundreds of real developer sessions.
The cartel that wasn't
The drug-trafficking comparison breaks down in one revealing way. Mexican drug organizations endure because they build hierarchical leadership structures, territorial control, logistics networks, corruption pipelines, and succession mechanisms that outlast individual operators. The Chinese transfer station industry has almost none of this. It is a swarm of small, short-lived stations competing on price, getting burned by upstream bans and model swaps, and cycling through operators who vanish with customer balances. There is no durable cartel structure, no long-term enforcement of territory or supply, and little capacity to defend the business against regulators on either side of the Pacific.
The result is spectacular short-term paydays for some, rapid attrition for most. The industry creates real regulatory gaps in China and the United States, real security risks for users, and real pressure on official AI access controls. But it lacks the institutional backbone to sustain itself as a parallel economy. The margins may look like narcotics; the staying power does not.
Fierce competition has already compressed margins for operators who play by even the loosest rules. Price wars drove some stations to offer tokens at 10 to 30 percent of official rates, making clean sourcing nearly impossible. Many smaller operators lasted only weeks before they ran out of upstream capacity, got their accounts banned, or disappeared with their customers' prepaid balances. The survivors increasingly relied on gray and black sourcing — model substitution (advertising Opus, delivering Sonnet), data logging and resale, opaque billing, and credential harvesting.
Multi-layer distribution networks, with resellers recruiting sub-resellers through redemption codes and community group chats, amplified the scale but also the instability. Each new layer added another operator who might vanish overnight.
"Like online scamming"
Mo was surprisingly candid about what comes next. "We all know we are racing against time," he said. "Regulation tightening is a matter of time."
He described the AWS crackdown as a turning point. The easy-money window — cheap partner accounts, lax credit controls, booming demand — was closing. Upstream supply of discounted tokens had, in his word, "died." Smaller operators who depended on a single cloud provider's pricing loophole were already shuttering. Price wars had compressed margins for the more legitimate players even before the crackdown.
Mo said that he and his core team had already accumulated enough money to walk away from the business entirely. They had planned their exit.
"We have planned to escape to Southeast Asia and the middle east," he said. "Myanmar, Cambodia, Thailand or, Dubai." He paused, then added: "It is not just because those are places for smart outlaws. If we want to stay in business, we still have many new opportunities to explore there. Like online scamming."
He said it flatly, without bravado, as if describing a career pivot at a corporate offsite. Then he ended the call.
What the crackdown means
Amazon's new approval requirements will make it harder for fraudulent operators to inflate their credit lines and disappear. The dual-signature mandate — Finance plus an L10 leader — eliminates the informal, relationship-driven overrides that the proxy operators exploited. The September identity-verification checkpoint will add a further barrier.
The proxy ecosystem is adaptive and fast-moving. The transfer station industry grew because it addressed a genuine, large-scale need: developers and companies wanted access to frontier AI models, could not get it through official channels, and found a gray market willing to serve them. That demand has not gone away. As official providers lower prices and Chinese domestic models improve, the arbitrage window is narrowing — but it has not closed.
The token traffickers made their money. Some, like Mo, are already planning their next career. The bill landed on Amazon's balance sheet. And the developers who routed their most sensitive code and credentials through anonymous proxy services may not yet know what they lost.
“Mo” is a pseudonym used at the source’s request to protect his identity.