On August 22, OpenAI reduced the price of GPT-5.6 Luna — its lowest-cost API offering — by 80 percent, cutting input costs to $0.20 per million tokens and output to $1.20. Mid-tier GPT-5.6 Terra fell roughly 20 percent. Within 24 hours, Google had cut Gemini 3.7 Flash by approximately 50 percent. Both companies framed the cuts as competitive positioning. Neither framing fully captures what the announcements reveal: that the cost of running a transformer inference workload has been falling faster than either company’s earlier pricing implied, and that the gap between actual cost and market price had become wide enough to attract competitive entry from below.
What is actually driving the cuts
The proximate driver is the intersection of three forces arriving simultaneously. First, model architecture improvements have produced meaningful efficiency gains — the ability to run GPT-5.6-class inference on less compute per token than GPT-4-class models required is a measurable consequence of three years of architectural refinement, not a marketing claim. Second, hardware costs have continued to fall as manufacturing yields have improved and as AMD and custom silicon (Google TPUs, Amazon Trainium) have introduced credible alternatives at the margin. Third, and most structurally significant, open-source competition from Chinese labs has established a price ceiling below which proprietary models cannot sustain premium positioning. Z.ai’s GLM-5.3, released in August and reported to have identified over a thousand genuine security vulnerabilities in widely-used codebases, is not a toy model. It runs at a fraction of OpenAI’s API cost. When capable open-source alternatives exist, the proprietary pricing premium collapses toward the cost of the infrastructure layer alone.
Structural winners
The structural winners of this dynamic are enterprises and developers who have built workloads on API access and have been absorbing above-competitive pricing for the past two years. A workflow costing $10,000 per month on GPT-5.6 Luna in July costs approximately $2,000 after the August cut — illustrative arithmetic, but the direction is exact. At that price point, use cases that were previously marginal — bulk document processing, always-on conversational applications, inference at consumer scale — become economically viable without a fundamental change in the underlying application. Startups that assumed AI inference costs were a structurally limiting factor in their unit economics will need to revise those models upward.
Structural losers
The structural losers are harder to identify from a price announcement alone, but the direction is legible. GPU suppliers have benefited from an inference-scarcity narrative in which data centre demand was assumed to grow faster than supply indefinitely. If inference cost is falling faster than demand is growing, the utilisation rates that justify continued aggressive capital expenditure on GPU capacity come into question. Hyperscalers that have committed to multi-year infrastructure buildouts on the assumption that AI inference would remain high-margin are now running the same arithmetic. The wave of AI-infrastructure bond issuance reported this week — tens of billions in debt financing against projected AI compute revenues — carries a different risk profile if the revenue-per-unit-of-compute is contracting at the pace the August cuts imply.
The trajectory and what is not yet resolved
OpenAI stated that the Luna cuts were achieved through model-level efficiency improvements, not infrastructure subsidies — which means the cost reduction is real, not an accounting artefact. If that efficiency improvement rate continues, inference pricing has further to fall. The relevant floor is not zero — data centre power and cooling impose real physical minimums — but it is substantially below current API pricing levels.
What is not yet resolved is whether the price war produces sustainable businesses or a subsidy race. Both OpenAI and Google are running inference infrastructure that is not yet profitable at current prices; the question is whether the cuts reflect genuine margin normalisation or a competitive burn-rate contest. History of platform markets suggests that at sufficient scale, one player’s efficiency advantage becomes self-reinforcing. Which player that is in LLM inference — and whether it is a US company, a Chinese open-source ecosystem, or a hardware-layer provider — is not established.