gpudiff the public record of change in GPU cloud pricing
Follow the diffs: RSS · Free API · Sponsor this site

Cheapest LLM API gateway, ranked

Six gateways resell the same models at different prices. We rank them by input price across the models they share, updated hourly. Frontier models are usually identical everywhere (gateways pass list pricing through), so the ranking is decided by open-weight models — where the gap can be large.

Cheapest LLM gateway right now: Requesty — the lowest input price on 116 of 162 shared models (72%).

RankProviderTypical priceCheapest on
🥇Requestyusually cheapest116 of 162 (72%)visit →
🥈OpenRouterusually cheapest192 of 296 (65%)visit →
🥉Glamausually cheapest185 of 290 (64%)visit →
4DeepInfra+11% typical35 of 86 (41%)visit →
5Ramp Router+11% typical4 of 57 (7%)visit →
6Novita+21% typical20 of 96 (21%)visit →

How this ranking works

For every model sold by two or more gateways we find the lowest input price, then measure each gateway's premium over it. "Typical price" is the median of those premiums across the models a gateway carries — median, so a handful of wildly-priced models can't skew the result, which is why a gateway that is cheapest most of the time reads as "usually cheapest" even if its average is dragged up by outliers. "Cheapest on" counts models where a gateway matches the lowest price (identical frontier prices are shared wins). We rank on published per-token input price only — not latency, throughput, or uptime (methodology). Gateway links may carry a referral code; the ranking is computed from prices and is unaffected.

Where to buy tokens by tier → · Cheapest GPU cloud →