Insight

Inference Needs a Forwards Market

ReportVishwa Naik

Introduction

There is a particular kind of conversation happening inside every company that has put a large language model into production. It does not happen on stage at developer conferences. It happens in the finance org, on the second Tuesday of the month, when someone pulls up the inference bill.

The bill is large, it is growing, and — this is the part that matters — it is unpredictable. Usage spikes when a feature goes viral or an enterprise customer turns on a workflow. Per-token prices move underneath you as providers reprice frontier models, deprecate old ones, and route traffic across a shifting fleet of accelerators. The team that promised the board a 70% gross margin on the AI product is now explaining why it came in at 55%, and why next quarter is anyone’s guess.

Thanks for reading Insights by Anera! Subscribe for free to receive new posts and support my work.

This is not a billing problem. It is a market structure problem. And it is the same problem electricity solved forty years ago.

Inference is a utility, and the token is its kilowatt-hour

Start with the unit. A token is the atomic quantum of inference — the smallest billable increment of intelligence production. Every model interaction, no matter how it is dressed up at the application layer, decomposes into input tokens consumed and output tokens generated, metered and priced per million. The token is to inference what the kilowatt-hour is to electricity: a standardized, fungible-enough unit of an otherwise abstract industrial output.

And inference behaves like a utility in the ways that matter for market design:

It is produced industrially. Tokens come out of GPU clusters running at high utilization, transforming compute and energy into model outputs. The marginal cost is real — power, depreciation, networking — and the fixed cost (the GPUs themselves, the data center, the multi-year power contracts) is enormous and sunk well before the first token is sold. Semianalysis extensively cover the breakdown of costs in the AI Cloud business.

It is delivered with time sensitivity. A token you needed at 2pm during a traffic spike is not interchangeable with a token available at 3am. Latency and availability are part of the product. Inference cannot be meaningfully stored — like electricity, it is consumed at the moment it is produced, and idle capacity is value that evaporates.

It is demand-variable against fixed supply. GPU capacity is set by procurement and buildout decisions made quarters or years in advance. Demand arrives in unpredictable bursts. The gap between a fixed, capital-intensive supply curve and a volatile demand curve is precisely the gap that every commodity and utility market has had to engineer around.

Once you see the token as a utility unit rather than an API line item, the analogy stops being a metaphor and starts being a blueprint.

Capacity planning is now a gross-margin problem

Here is the uncomfortable part for the application layer. As more of a software product’s cost of goods sold migrates into inference, the discipline of buying inference becomes the discipline of defending margin.

Consider what the AI-native company’s P&L actually looks like. Traditional SaaS had near-zero marginal cost — once the software was written, serving an additional user cost almost nothing, and gross margins of 80%+ were the norm. AI-enabled products break this. Every query has a real, non-trivial marginal cost denominated in tokens. The product’s gross margin is now a function of two things the company does not control: how much inference each user consumes, and what that inference costs at the moment of consumption.

That second variable is the killer. If you are buying inference purely on the spot market — paying the prevailing per-token rate, request by request — then your cost of goods sold inherits the full volatility of the inference market. This is not just price volatility, but it is also access and quality of service volatility. Your margin is short volatility. When frontier prices spike because a new model launches and demand outruns capacity, your COGS spikes with it, and there is nothing in your stack to absorb the shock. Many will try to convince you that reserving infrastructure can solve this problem, completely eliminating the cost volatility. However, it opens a new, and far more expensive cost head- underutilisation.

This is the same position a load-serving entity occupies in an electricity market. A retail electricity supplier sells power to households at a fixed monthly rate but buys a part of it from a wholesale market. If it bought everything on the spot market, a single heat wave could bankrupt it. So it does not. It hedges — it locks in forward contracts for the bulk of its expected load and uses the spot market only for the residual. The fixed retail price it offers consumers is, in effect, manufactured out of a carefully layered portfolio of forward purchases. Many modern electricity retailers only buy ahead a part of their inventory, freeing up cashflow and drawing from the grid via real time and day ahead markets.

AI applications are LSEs for intelligence. They sell intelligence to end users at some posted or bundled price and procure it from the real time spot market (on demand end points) OR via reserved instances. Right now, almost none of them have the procurement toolkit that every electricity retailer takes for granted.

OpenAI already built the supply side of this market

The most telling development is that the largest inference producer in the world has already started building the contracts — from its own side of the table.

In late 2025, OpenAI introduced Guaranteed Capacity. The mechanics are worth stating plainly. Customers choose one-, two-, or three-year commitments, with discounts that increase based on the annual commitment, and certainty of access to compute based on spend levels, drawing down from that commitment across the portfolio of OpenAI products. Alongside it sits the older Reserved Capacity offering, where reserved instances are a static allocation of capacity dedicated to a customer, rented on three-month or one-year commitments, with roughly 15% savings on the longer term.

Request OpenAI Guaranteed Capacity | OpenAI

Read these as what they are: forward contracts on inference, sold over the counter, originated by the producer. The customer trades a long-dated spend commitment for two things — a discount, and the right not to be throttled when it matters. It is, in effect, a reservation system: you trade a multi-year commitment for the right to not be throttled when your business needs the throughput most.

Now ask why OpenAI is doing this, because the producer’s motive is the whole point.

A GPU fleet is a brutal fixed-cost business. The accelerators depreciate whether or not they are processing tokens; the data center draws power and rent regardless; the capital was raised against an assumption of utilization. Reports suggest OpenAI is targeting roughly $600 billion in total compute spending by 2030. When you have committed to spending on that scale, the single most valuable thing you can buy is certainty of utilization. Idle GPUs are the inference equivalent of a power plant spinning with no one drawing from the grid — pure value destruction.

Guaranteed Capacity is how a producer converts uncertain future demand into contracted, bankable revenue. It smooths cash flow. It lets the producer underwrite the next buildout against signed commitments rather than hope. This is exactly what a generator does in a capacity market. For longer-term grid resource stability, system operators run a capacity market that auctions the commitment of a resource to provide energy, with revenues paid regardless of whether energy is produced or not. The generator gets paid for standing ready. The buyer gets assurance. The fixed-cost asset gets financed against contracted availability rather than merchant hope.

OpenAI has, quietly, built the generator’s side of a power market. What it has not built — what no one has built — is the rest of the market.

What’s actually missing: the middle of the stack

A mature commodity market is not one market. It is a stack of layered timeframes, each solving a different problem, and electricity is the cleanest example. Strip it down and you get four layers:

The spot and real-time market. The spot market deals with short-term supply and demand fluctuations through real-time and day-ahead markets, where prices move with real-time conditions. This is where the residual gets cleared and where the true marginal price is discovered, minute by minute. It is also notoriously volatile, in terms of costs and capacity!

The day-ahead market. A scheduling layer. Generators and load-serving entities submit bids the day before delivery, the operator forecasts demand and clears a price, and the bulk of physical delivery gets committed in advance. This is the coordination mechanism that lets a system with fixed supply and variable demand plan one cycle ahead.

The forward market. This is the layer that does the financial work. Market participants hedge the risk of price fluctuation on spot markets with financial trades taking place up to several years before physical delivery, and price formation is based on the expected average price of the day-ahead timeframe. Forwards are how risk gets transferred from the people who cannot bear it (the LSE defending a margin) to the people who will (generators locking in revenue, and speculators warehousing the rest). Forward market hedging lets both generators and load-serving entities smooth irregular scarcity conditions into average expected values, transforming a boom-and-bust revenue environment into one that is more predictable and bankable.

The capacity market. The reliability backstop, paying assets to stand ready years in advance — PJM’s capacity market, for instance, is based on three-year forward-looking annual obligations for locational capacity needs.

Map this onto inference and the asymmetry jumps out.

The spot and real-time layer exists. It is the per-token pay-as-you-go API every provider runs by default — and, increasingly, the routing layer above it. OpenRouter and platforms like it already function as the inference market’s real-time balancer: they take a request, survey live prices and availability across providers, and clear it against whatever capacity is cheapest and fastest right now. That is genuinely useful, and it is genuinely a spot market — with all the volatility that implies. The price you pay is the price at the instant of consumption, and you wear every spike.

Andrej Karpathy YC startup school talk

The capacity layer is being built — that is precisely what OpenAI’s Guaranteed and Reserved Capacity is: bilateral, producer-originated reservations of standing capacity. Anthropic is looking to augment its reselling efforts too and I suspect every lab will look to go in the same direction. The sheer capex on their books demands this!

But the two transformative layers in the middle — day-ahead and forwards — do not exist yet for inference. And these are the layers that do the coordination and the risk transfer. A real-time balancer optimizes the present; it gives you the best execution available at the moment you ask. It does nothing for the problem of the future — it cannot tell you what a million tokens will cost tomorrow or next week or event next month. It cannot let you lock that cost in, and it cannot move the risk of that cost off your P&L and onto someone who wants it.

Critically, these middle layers are the one part of the stack that cannot be supplied by any single producer, because their entire function is to sit between producers and consumers as neutral infrastructure. OpenAI can sell you a forward commitment on OpenAI compute. A router can sell you best execution today. Neither can give you a market — a place where inference demand and supply meet across providers, where a price for future inference is discovered openly, where risk is transferred to whoever will price it best, and where a buyer can hedge their token COGS the way an electricity retailer hedges power.

The middle layers are a risk-transfer layer

It helps to be precise about what the day-ahead and forward layers actually do, because “hedging” undersells it. They are a risk-transfer mechanism between two parties with exactly opposite problems.

The application has demand it cannot perfectly forecast and a cost it cannot control, and it desperately wants certainty on both. The provider has capacity it has already paid for and cashflow it cannot predict, and it desperately wants certainty on utilization. Today these two parties meet either at the spot price — which means neither gets certainty, and the volatility just sits in the open between them OR at the reserved infrastructure price, where regardless of whether the infrastructure is managed or not, the application will have to eat the cost of underutilisation.

The issue with reserved infrastructure is that it does not solve the idle GPU problem. It merely changes who is willing to pay for it. Ultimately, this will leave the application unhappy and drag on their margins.

A day-ahead market lets them coordinate one cycle out: applications signal expected demand, providers commit capacity against it, and delivery gets scheduled rather than scrambled. A forward market lets the price risk move further out and across providers — the app locks tomorrow’s cost today, and the provider best positioned to absorb that obligation, whether by virtue of cheap capital, idle fleet, or a software stack that delivers below market cost, wins the trade. That movement is the product. It is the single most important thing a commodity market does, and it is the thing inference has no mechanism for.

The payoff for the application is direct: a token COGS line you can actually underwrite, a gross margin that survives a model launch, and a procurement function that looks like an electricity retailer’s instead of a gambler’s. The payoff for the provider is the same one OpenAI is already chasing with Guaranteed Capacity — contracted utilization and smoother cashflow — except sourced from a liquid market rather than negotiated one enterprise at a time.

Forwards turn tokens into collateral — and capital into velocity

The buildout has happened. The capex is committed, the GPUs are racked, the data centers are lit, the upstream contracts are signed. The financing question for the inference supply side is no longer “how do we raise the capital to build it.” It is the harder, more interesting question every mature industrial sector eventually has to answer: how does the capital tied up in the deployed base turn faster.

I remember listening to Neil Tiwari from Magnetar on No Priors with Sarah Guo actively talking about how they are looking for way to build a finance business for Inference. While their focus is probably on converting inference clouds into GPU owners, we believe that the finance business for inference is actually working capital, not GPU loans.

This is what forwards enable. They turn tokens into collateral, and a collateralizable token is the difference between capital sitting still and capital in motion. The mechanism is a flywheel:

A provider signs a forward. The forward, on a cleared venue at a known price with a defined delivery window, is a receivable — the cleanest possible quality of short-dated contracted cashflow. The provider discounts that receivable, takes cash today against delivery in the future (1 hour to 90 days), and redeploys the cash into more operational capacity: more rented hours from upstream, more leased fleet, more inventory of compute against the next set of forward sales. The new capacity gets sold forward. The new forwards get financed. The wheel turns again. Same underlying balance sheet, materially larger book of business — because the capital is recycling through the working-capital cycle instead of waiting on the back end of spot delivery to come back as cash.

The token claim is the asset that makes this possible. A pay-as-you-go API booking is unfinanceable — it is a promise of unspecified volume at unspecified price at some unspecified future moment. A forward token claim is the opposite: standardized quantity, locked price, defined window, cleared counterparty. That is what a lender, a trade-finance desk, an asset-based lender, or another participant in the market can accept as collateral. Tokens stop being a perishable, unbillable abstraction and become a unit of account that the financial system can lend against.

None of this is exotic. It is the same trade-finance plumbing that has run physical commodity markets for a century. Oil and gas producers fold their near-dated hedge book into the borrowing base of a revolving credit facility. Metals merchants finance 30-90 day forward sales through trade banks. Freight operators borrow against locked charter receipts. Exporters monetize letters of credit on day-of-shipment terms. In every case, a short-dated contracted receivable is treated as fundamentally different collateral from merchant exposure, and priced accordingly. This plumbing does not exist for inference yet because the contracts to finance don’t exist yet.

The contrast makes the point sharper. Orders filled on the real-time spot market are not financeable, and they never recycle capital. When OpenRouter routes a query to whichever provider is cheapest right now, the provider that fills it earns a few cents of revenue at a price they could not have predicted yesterday, collected on standard billing terms after delivery. No lender underwrites working-capital lines against “spot fills, in volumes we cannot forecast, at prices that move every minute.” Spot revenue is real, but it is merchant revenue — the lowest possible quality of cashflow from a financing standpoint, and the slowest to convert into deployable capital. A provider running on spot is collecting cash after delivery, in unpredictable quantities, at unpredictable prices, with no instrument to discount against. The capital tied up in their fleet sits still. That is the trap most of the inference supply side is in today, and it is the same trap shale producers were in before the near-dated hedge market matured.

A forward market changes which cashflows providers generate, and therefore changes how fast capital can turn through the inference supply side. That is the deep capital-efficiency story. It is not about financing the buildout — that is already done. It is about taking the deployed base and making the dollars move through it faster: forward, finance, more compute, more forward, more finance. Tokens-as-collateral is what makes the wheel turn.

What’s missing, concretely

So lay out the gap directly. Inference has a spot-and-real-time layer (the APIs, and routers like OpenRouter balancing across them). It is starting to get a producer-originated capacity layer (Guaranteed Capacity and its kin). What it does not have:

  • A price for future inference. A forward curve — what a million tokens of frontier capability will cost next month, next quarter, next year — discovered by a market rather than dictated by a single seller’s discount schedule.

  • A scheduling layer. A day-ahead mechanism where applications signal forward demand and providers commit capacity against it, so the system plans ahead instead of colliding at the spot price.

  • Standardized, transferable contracts. A capacity reservation you cannot resell is a liability if your forecast was wrong. A forward you can transfer is a hedge with an exit — and, as above, collateral. Transferability is what converts a procurement contract into a market instrument.

  • A neutral venue and a clearing layer. Someone has to stand between counterparties, net exposures, hold collateral, and guarantee performance. This is the clearinghouse role in every derivatives market, and it is what lets the structure scale past bilateral relationships.

Introducing Anera

Anera is the forward capacity market for AI inference. A scheduling layer where applications signal forward demand and providers commit capacity against it. A forwards layer where the price of future inference is discovered, locked, and transferred — opening provider competition across the curve, turning token claims into collateral, and routing price risk to the balance sheet best equipped to hold it. Ultimately, it becomes the risk transfer layer for the future of work.

The starting point is the open-weight frontier. The same model served by many independent providers — fungible underlying, multi-provider supply, no single seller in control of the reference price. Those are the conditions a forward market needs to clear, and they exist today on every open-weight frontier model. That is where Anera begins.

Electricity figured this out. Inference is next.


Anera is the forward capacity market for AI inference — the coordination and risk-transfer layer between token buyers and providers. Find out more on our website

Thanks for reading Insights by Anera! Subscribe for free to receive new posts and support my work.

Read and subscribe on Substack.

Comments, archives and new posts straight to your inbox.

View on Substack