Paper

The Missing Meter

What every infrastructure buildout leaves behind, and the unit AI still lacks

AuthorSidney ScottVersionv1.0Contactsidney@themissingmeter.orgPDFDownload PDFCite

Every figure here comes from public filings, published benchmarks, and primary sources. Contested and estimated numbers are marked as such, and every figure carries a provenance label under the taxonomy of Appendix A.6. The author builds and invests in commercial infrastructure informed by this argument, and has drawn on discussions with researchers and investors across the field for feedback on earlier drafts. The reference implementation, the record schema, the test suite, and the code for every computed exhibit are included as ancillary files with this submission.

AI is the first trillion-dollar infrastructure that cannot yet state, in a standard unit, what its output costs. The four largest technology companies are guiding to roughly $725 billion of combined capital expenditure for 2026, the largest private infrastructure program in history. At the same time, the cost of querying a model at GPT-3.5 capability fell more than 280-fold in under two years. Record physical investment and collapsing unit economics, occurring simultaneously, is not a contradiction. It is the signature of a turning point that every general-purpose infrastructure has crossed. On the far side of each turning point, three things appeared: a standardized unit of account, a coordination layer that priced and routed work, and a migration of value away from the owners of physical assets toward the operators of that layer. AI has the physical layer and is acquiring the protocols. It does not yet have the meter, nor the unit that meter would read, which this paper names the served token.

Exhibit 1
The paradox in one chart
Hyperscaler capex ($B, indexed 2021=100)
Inference cost (indexed 2022=100, log scale)
$725B
combined capex guided for 2026
280×
cost decline in under two years — fastest in history

Record physical investment and collapsing unit economics, occurring simultaneously, is not a contradiction. It is the signature of a turning point that every general-purpose infrastructure has crossed.

0×
Cost decline in inference per token, 2022–2024
$0B
Hyperscaler capex guidance for 2026
0th
Infrastructure turning point in history — the AI cycle
The Argument in Six Claims
Each claim is counterintuitive, load-bearing, and falsifiable

The Transaction Processing Performance Council has priced useful work under an enforced deadline, with an audited three-year cost basis, a correctness floor, and a mandatory third-party audit, for thirty-eight years. Two consortia now hold adjacent halves of the unit AI needs and neither has joined them.

Wrong if: A cross-vendor, production-traffic, service-level-conditioned cost unit is already in commercial use and this survey missed it.

Under an interactive deadline the cheapest provider delivers approximately zero compliant output, so its cost per unit of useful work diverges while its advertised price stays lowest in the market. Raw dollars per million tokens ranks all fifteen identically for buyers whose correct rankings are opposite.

Wrong if: Latency-conditioned and raw price rankings coincide in practice.

Economic capacity sits strictly below engineering capacity. Across 63,824 requests of real production traffic from two independent operators, the measured optimum is 15 percent utilization at a ten-millisecond deadline — against the 83 to 95 percent a queueing model predicts. A paper publishing only the model would have been directionally right and wrong by a factor of five. Dollars per GPU-hour is flat across that entire range; raw dollars per million tokens falls monotonically and therefore instructs the operator to run at saturation.

Wrong if: The interior optimum vanishes on a workload where it has not yet been tested. On the two tested it is present and severe.

An interactive objective of six milliseconds — tighter than MLPerf's Interactive bound and well inside what voice agents require — falls in the first bucket and cannot be measured at all. Every figure quoted against it is an interpolation across a discarded distribution.

Wrong if: The conventions are revised or per-request latency is standardized on spans. That is Prediction 8, and its cost is one line in a configuration file.

Semiconductor fabrication kept its margin with no unit of account ever appearing, because the physical layer is not contestable. Cloud computing's coordination layer was bundled into an asset owner before a neutral standard could set — which is why the internet row of this paper's own table names AWS rather than a clearing house. The pattern holds only where the physical layer is contestable and the coordination function is standardized before it can be captured.

Wrong if: Either condition proves unnecessary. Prediction 7 is the test.

No public series of GPU utilization by age cohort exists anywhere, so the depreciation dispute is unresolvable by measurement. Enterprise AI adoption is reported at 20, 50, and 70 percent by three credible instruments, and one survey's measured rate doubled when a question was reworded. The bubble debate compares capital expenditure in dollars to demand in vibes.

Wrong if: The industry transacts in a work-denominated unit by end-2029. If it does not, the central claim of this paper fails on its own kill criterion.

Sections 8 through 10 defend each claim. Appendix A specifies the unit. Predictions 1 through 8 are the paper's full falsification surface, each dated.

Two facts that cannot both persist

1. Two facts that cannot both persist

The capex fact

The bust never kills the industry. In the first quarter of 2026, Google, Amazon, Microsoft, and Meta told their investors they intend to spend approximately $725 billion on capital expenditure this year, up roughly 77 percent from a record $410 billion in 2025. Goldman Sachs projects a combined $5.3 trillion from these four companies between fiscal 2025 and fiscal 2030. Investment in information processing equipment and software, about 4 percent of US GDP, was responsible for 92 percent of US GDP growth in the first half of 2025.

The unit-cost fact

At the same time, the unit price of the output is in freefall. Stanford's 2025 AI Index measured the cost of GPT-3.5-equivalent capability falling from $20.00 per million tokens in November 2022 to $0.07 by October 2024, a 280-fold reduction in twenty-three months, with task-level declines ranging from 9× to 900× per year. H100 rental prices fell from a 2023 peak near $8 per hour to $1–$2 by late 2024. In January 2025, a Chinese lab's demonstration that frontier reasoning could be reproduced at a fraction of assumed cost erased roughly $589–$593 billion of Nvidia's market value in a single trading day, the largest one-day loss in stock market history.

The paradox, and what it actually is

Markets are treating these two facts as a paradox, and the paradox has a name: bubble. The historical record suggests something more specific. Peak physical spending and collapsing unit costs, at the same moment, is what the crest of an infrastructure installation phase looks like. It looked exactly like this in 1846, in 1929, and in 2000. What follows the crest is not the death of the industry. It is the arrival of the industry's second act, which in every prior case was larger than the first, and which in every prior case was captured by different companies, operating a different layer, measured by a unit that did not exist during the buildout.

This paper makes three claims. First, the pattern is real: five prior infrastructures followed the same four-phase arc, and the mechanism that drives the arc is visible in AI today. Second, the strongest objection — that AI capital depreciates too fast for the pattern to hold — is partially correct and must be answered rather than dismissed; the answer changes what to build, not whether the transition occurs. Third, the binding constraint on the efficiency era is now measurement: AI is the first trillion-dollar infrastructure that cannot yet state, in a standard unit, what its output costs.

Every buildout ends the same way

2. The pattern

Perez's four-phase arc

Every infrastructure revolution got its unit. Inference hasn't. Carlota Perez, studying five technological revolutions since 1771, identified a recurring structure (Technological Revolutions and Financial Capital, 2002). An installation period, financed by speculative capital, builds the physical substrate in a frenzy that always overshoots. A turning point, usually a crash, transfers the substrate from speculators to operators at a discount. A deployment period follows, in which production capital puts the overbuilt substrate to work, and the technology's real economic payoff arrives, decades after the initial breakthrough and years after the bubble that supposedly discredited it.

The frenzy is not a defect in the process. It is the financing mechanism. No rational allocator funds 80 million miles of fiber against demonstrated demand; only a mania does, and the deployment era then inherits the fiber at cents on the dollar. As the venture capitalist Fred Wilson compressed Perez's corollary: nothing important happens without crashes.

Five installments compressed

Five installments of the pattern, compressed:

InfrastructureInstallation frenzyTurning pointUnit of accountCoordination layerDeployment-era winner
Railways (UK)263 railway acts in 1846; capital tripled in a decade1847 crash; shares fell ~66%ton-mileTimetables, clearing housesIndustrial economy on cheap freight
ElectricityCompeting plants, per-lamp pricing, ~6% utilizationConsolidation, 1900skilowatt-hourMetering + load balancingThe utility model; price fell 20¢ to 2.5¢/kWh by 1909
Payments1958 Fresno card drop; 22% delinquency; ~$20M fraud losses1970: banks surrender the networkThe cleared transaction (interchange)Authorization + settlement (BASE I/II)Visa: >50% net margins owning no cards, no banks
Container shippingIncompatible boxes, port-by-port chaosISO standardization, 1968TEUIntermodal schedulingFreight cost $5.86 to $0.16/ton; globalization
Internet~$500B–$1T telecom capex; 80M miles of fiber2000–02: NASDAQ −78%; >60 bankruptcies; ~95% of fiber darkThe metered Mbps, then the API callTCP/IP, DNS, then virtualization and cloudAWS, Google, streaming, on stranded fiber

Table 1. Five infrastructure cycles compared. Each row follows the same four-phase arc: frenzy → turning point → standardized unit → value migration to coordination layer. Sources: railways [11]; electricity [12]; payments [13][14][15]; containers [16]; internet [17][18][19].

Two regularities in this table do the argument's work. The bust never killed the industry. Britain's railway shareholders were ruined in 1847; Britain kept the 8,000 miles of track, and the track kept lowering freight costs for a century. Telecom equity holders were wiped out in 2001; the fiber stayed in the ground, bandwidth prices fell up to 90 percent, and YouTube, founded in 2005 on that glut, generated roughly $31 billion for Alphabet in 2023 alone.

Value migrated to the layer that measured and routed, not the layer that owned. Visa issues no cards, extends no credit, and owns no banks; it operates authorization, clearing, and settlement, and in fiscal 2024 it carried 234 billion transactions and roughly $16 trillion of volume at net margins above 50 percent. In each case the physical layer was necessary, competitive, and margin-poor. The measurement-and-coordination layer was singular, standard, and margin-rich.

Exhibits 1 & 2

Hyperscaler capex vs. inference cost, 2021–2026
Exhibit 1 — Hyperscaler Capex vs. Inference Cost

Combined hyperscaler capex (Microsoft, Google, Amazon, Meta) vs. inference cost per million tokens at GPT-3.5 capability, log scale. Sources: Company filings; Artificial Analysis (2025); author estimates.

Combined hyperscaler capex (Microsoft, Google, Amazon, Meta) vs. inference cost per million tokens at GPT-3.5 capability, log scale. Sources: Company filings; Artificial Analysis (2025); author estimates.

Unit cost indexed to turning point — five infrastructure cycles
Exhibit 2 — Unit Cost Indexed to Turning Point (Log Scale)

Unit cost indexed to 100 at each infrastructure turning point. AI inference decline (94.8%/yr) is the steepest in recorded infrastructure history. Sources: Crafts (2004); Joskow (1997); Downes & Nolan (2017); Artificial Analysis (2025).

Unit cost indexed to 100 at each infrastructure turning point. AI inference decline (94.8%/yr) is the steepest in recorded infrastructure history. Sources: Crafts (2004); Joskow (1997); Downes & Nolan (2017); Artificial Analysis (2025).

1 / 2

3. Act I: The bust that builds

The crash destroys capital. It does not destroy capacity. The pattern's first regularity is worth watching in slow motion twice, because the two runs bracket the industrial era and agree on every particular.

Britain, 1845 to 1850. Parliament approved 263 railway acts in 1846 alone, authorizing roughly 9,500 miles of new route. Paid-up railway capital more than tripled inside the decade, from about £30 million to over £100 million by 1849. George Hudson, the Railway King, held the mania's mirror up early: he was paying dividends out of new subscribers' capital, a structure that acquired its modern name only decades later. Monetary tightening and the commercial crisis of 1847 called the loans; shares fell about two thirds by 1850, and only about two thirds of the authorized mileage was ever built. The investors' losses were permanent. So was the track's usefulness.

The United States reran the experiment with fiber, 1996 to 2006. The Telecommunications Act of 1996 opened the field; carriers laid roughly 80 million miles of fiber against demand that did not yet exist; annual capex peaked near $120 billion in 2000, with cumulative investment estimated well beyond $500 billion. The NASDAQ closed at 5,048.62 on March 10, 2000, and fell roughly 78 percent over the following thirty-one months. WorldCom filed the largest bankruptcy in US history at that point. The fiber stayed in the ground. Bandwidth prices fell up to 90 percent. AWS launched in 2006 on the economics the crash created. YouTube, founded in 2005 on that glut, was acquired for $1.65 billion in 2006 and generated revenues that Alphabet's annual YouTube advertising and Premium subscriptions passed $60 billion in 2025.

Capital is destroyed at the turning point. Capacity is not. The write-down is the mechanism by which the next era acquires its inputs below cost.

4. Act II: Value migrates to the meter

The meter did not record the business. The meter created the business. Samuel Insull did not win Chicago by generating more electricity than his competitors; he won by metering demand — a device he licensed after seeing it in Brighton in 1894 — discovering that diverse customers have diverse peak loads, and using the meter to sell the same fixed capital many times over. His unit, the kilowatt-hour, plus his metric, load factor, converted electricity from a per-lamp luxury into a utility, cutting its price 87 percent and growing his customer base 40-fold across two decades.

The box that Malcolm McLean standardized cut cargo handling from $5.86 to $0.16 per ton, a 97 percent collapse, and the TEU became the unit in which the entire logistics industry denominates itself. Dee Hock structured Visa's predecessor so that no member bank could own the network; interchange became the unit; and Visa today carries $16 trillion of annual volume at net margins above 50 percent, owning no cards and no banks.

The pattern's quietest claim is loudest on one axis: in each case the physical layer was necessary, competitive, and margin-poor. The measurement-and-coordination layer was singular, standard, and margin-rich.

UNIT COST LOG — INFRASTRUCTURE TURNING POINTS
The strongest objections

5. The depreciation objection

The strongest argument against the pattern is Michael Burry's: the hyperscalers are understating depreciation by roughly $176 billion between 2026 and 2028, which means their earnings are overstated and their capital is not as durable as the railway track or the fiber. The objection is partially correct and must be answered rather than dismissed.

The obvious response is that AI silicon is different from railway iron. An H100 purchased in 2023 faces competitive pressure from Blackwell in 2025 and from whatever follows it; the economic useful life of a GPU is shorter than the accounting useful life that most hyperscalers are currently booking. This is true. But the objection proves too much: if fast depreciation were fatal to the infrastructure pattern, the pattern would not have held for fiber, which became economically obsolete almost immediately after it was laid, and which nonetheless transferred to the deployment era at a discount and enabled YouTube, AWS, and every streaming service. The question is not whether AI silicon depreciates — it does, at roughly 50–70 percent of value in the first two years based on secondary-market data — but whether the capacity it represents transfers to the deployment era at a discount, and whether the deployment era's value is captured by the silicon or by the layer above it.

The depreciation dispute is, in fact, the paper's own argument wearing different clothes. The reason no one can state the correct depreciation schedule is that no one can state, in a standard unit, what the depreciating asset actually produces. The meter is missing from the income statement for the same reason it is missing from the invoice. Burry is right that the numbers are wrong. He is wrong that the wrongness is fatal. The answer to Burry changes what to build, not whether the transition occurs.

6. The inheritance mechanism

The secondary market is already operating. The mechanism by which the deployment era inherits the installation era's capacity is not mysterious: AWS's chief executive stated in January 2026 that AWS has never retired an A100 server and remains sold out of them nearly six years after launch. Google reports full utilization on seven- and eight-year-old TPUs. 2017-vintage V100s still rent across more than seventeen providers.

The secondary market for AI silicon is already operating; what it lacks is a standard unit in which to price the useful work the silicon delivers. Secondary H100 prices fall to 20–40 percent of peak within two to three years, but brokers describe the market as structurally opaque — prices and sold-out claims are not utilization rates, and no cohort-level utilization series exists as of mid-2026. The inheritance mechanism works; it cannot be measured; and the inability to measure it is the argument's own evidence.

The rails are being laid. What runs on rails is traffic, and traffic requires a tariff, and a tariff requires a unit.
The served token

7. The protocols are arriving

The protocols are arriving in the historically correct order. Payments needed interchange (a standard for who owes whom), then authorization (BASE I, 1973), then settlement (BASE II, 1974). Between late 2024 and early 2026, AI acquired the same three layers in the same order:

  • The Model Context Protocol standardizing how agents reach tools and data (open-sourced November 2024, ~97 million monthly SDK downloads, donated to a neutral foundation in December 2025).
  • Agent2Agent standardizing how agents reach each other (April 2025, 150-plus partners, IBM folding its competing protocol into it).
  • Three competing agent-payment rails — AP2, ACP, and UCP — standardizing how agents settle.

The donation of these protocols to foundations is not altruism; it is the participants' recognition, learned from TCP/IP's victory over proprietary networking, that a coordination layer is only valuable if it is universal, and only universal if no one owns it. The rails are being laid. What runs on rails is traffic, and traffic requires a tariff, and a tariff requires a unit.

Figure — The served token as unit of account
SupplyDemand
1
Hyperscaler GPU Clusters
Physical infrastructure — NVIDIA, AMD · $725B capex guidance (2026)
compute capacity
2
Cloud AI APIs
OpenAI · Anthropic · Google · Mistral — inference-as-a-service
API calls (tokens billed)
3
The Served Token (svt)
1 output token · delivered inside SLO · above quality floor
priced, routed, metered
4
Enterprise Orchestration
Agents, RAG pipelines, workflow automation — Finance · Healthcare · Legal
end-user requests
5
End Users & Applications
The demand layer — consumers, developers, embedded AI products
Served token — the missing unit
Existing layers (GPU-hours, raw tokens)

The served token (svt) sits between the GPU-hour on the invoice and the verified useful work on the income statement. It is the unit that makes the layer above it legible and the layer below it a commodity.

8. The missing meter

The strangest fact about a trillion-dollar industry

AI cannot say what its output costs. As of mid-2026 there exists no standardized, cross-vendor unit of account that expresses dollars per unit of verified useful work, conditioned on a stated service-level objective and a quality floor, governed by a multi-party body, and adopted in production pricing or disclosure.

Compute is bought in GPU-hours. It is consumed as answered queries, completed tasks, resolved tickets, and generated tokens, each under a latency constraint that determines whether the output was worth anything at all: a voice agent's reply that arrives in 200 milliseconds is a product, and the same reply in four seconds is a refund. Between the GPU-hour on the invoice and the useful-work-under-a-deadline on the income statement, there is no standard unit.

Dollars per million tokens is the leading candidate and the industry's de facto price sheet, but a raw token is a kilowatt-hour with no voltage standard: tokens vary in quality (which model, verified against what), in urgency (batch overnight or interactive p99), and in usefulness (a token of correct answer and a token of hallucination invoice identically).

The historical sequence is exact enough to be a specification. Before Insull's meter, electricity was priced per lamp: a flat fee per connected bulb, whatever it consumed — exactly as AI subscriptions today price per seat whatever the seat consumes. Per-lamp pricing made load invisible, so plants ran at 6 percent utilization, so capital costs stayed brutal, so electricity stayed a luxury. The demand meter made consumption visible; visibility revealed that customers peak at different hours; diversity of peaks meant the same turbine could serve the factory by day and the streetcar by night; load factor became the metric that turned utilization into strategy; and the price fell 87 percent while the market grew 40-fold. The meter did not record the business. The meter created the business.

AI's equivalent unit must denominate what the buyer actually buys, and what the buyer buys has three legs: cost, useful output, and a service-level condition. This paper names that unit the served token, abbreviated svt, and specifies it in Appendix A: one output token delivered inside its SLO and above its quality floor.

The prior art, and which leg each one amputates

The industry and the standards bodies have built every leg of the meter separately, which makes the survey of near-misses the strongest evidence that the fused unit is missing rather than impossible. Each candidate below is real, useful, and incomplete in a specific, documentable way.

The obvious objection is that dollars per million tokens is already good enough. It is not, for a precise reason. In 2026, one major lab changed its tokenizer; measured costs on identical prompts moved by double-digit percentages with no change in the underlying work. The kilowatt-hour is defined by physics; it does not change when the utility upgrades its meters. A unit that changes when the vendor changes its implementation is not a unit. It is a price list in disguise.

missing: AI workload; fixed corpus only

TPC-C / TPC-H

The whole unit exists, for a different domain. TPC-C and TPC-H report price-performance under enforced response-time constraints and a correctness floor, audited by a TPC-certified auditor, with the cost basis fixed by a published Pricing Specification. The unit this paper calls for is TPC-C applied to inference. TPCx-AI has the cost leg and the work leg and not the service-level leg; MLPerf Inference has the service-level leg and the quality leg and not the cost leg. The two consortia have between them assembled every component and have not joined them.

missing: useful-work denominator; no service-level or quality dimension

FinOps FOCUS specification

FOCUS v1.4 (ratified 4 June 2026, Linux Foundation) defines a vendor-neutral schema for cost and usage data. Since v1.2 it has normalized billing in non-monetary units explicitly. Exports are offered by the major clouds. It answers what was spent and on what, across vendors. It contains no useful-work denominator and no service-level or quality dimension anywhere in its schema.

missing: price; cannot resolve sub-10ms deadlines; Development status

OpenTelemetry gen-AI SemConv

Standardizes token accounting as span attributes and latency as histogram metrics. The whole generative-AI convention set remains at Development status as of SemConv v1.40.0 (April 2026). The default explicit bucket boundaries for time-per-output-token begin at 10ms — an interactive SLO of 6ms falls in the first bucket and cannot be resolved from default telemetry at all. This is a narrow, fixable defect registered as Prediction 8.

missing: cost leg

MLPerf Inference Server scenario

Strongest prior art. Enforces p99 bounds on TTFT and TPOT with accuracy floors, governed by MLCommons, at v5.1 (September 2025). No dollar denominator in any scenario. If MLCommons adds a cost denominator to the Server scenario, the unit this paper calls for will be most of the way to existing — registered as one of two nearest live falsifiers.

missing: price; quality acceptance test

Goodput (DistServe / research literature)

The service-level-conditioned denominator standardized in the research literature. DistServe defines goodput as completed requests per second adhering to SLOs, reported as per-GPU goodput. Goodput is the served token's service-level leg, already rigorous, already load-swept. What it has never carried is a price or an acceptance test: a goodput-optimal placement serving fluent nonsense scores identically to one serving correct answers.

missing: dollar numerator; quality floor

Green Software Foundation SCI (ISO/IEC 21031:2024)

Defines SCI = ((E × I) + M) / R, where R is a functional unit declared by the practitioner. It is structurally the same object as the unit proposed here with carbon in the numerator, it is multi-party governed, and its route from consortium draft to ISO standard is the closest available precedent for how a rate-per-useful-work specification gets ratified. Read it as the process template.

missing: dollar numerator; quality floor

SPECpower_ssj2008

Reports overall ssj_ops/watt across a graduated ladder of ten target load levels in ten-percent increments plus active idle. Establishes that a consortium can standardize a rate per unit of useful work, and that load-swept rather than single-point reporting is a viable practice. Section 8.2 argues that load-swept reporting is not merely viable here but mandatory.

missing: service level (latency SLO)

ARC Prize / cost-of-pass

Reports model performance as a two-axis matrix of verified score against measured dollars per task. Recent academic work has formalized the same idea as cost-of-pass, the expected monetary cost of a correct solution. Two legs, rigorously joined, on a single benchmark, unconditioned on latency: a batch unit, not a service unit.

missing: quality leg; not portable across vendors

Azure PTU / Google GSU / OpenAI service tiers

All price differentiated service levels, proving SLO-conditioned pricing is commercially viable today. None carries a quality leg, none is portable, and contractual guarantees are thinner than the marketing.

missing: quality, latency, SLO conditioning

Raw $/Mtok (industry de facto)

Different tokenizers produce different token counts for identical text. When one lab changed tokenizers in 2026, measured costs on identical prompts moved by double-digit percentages. The industry's de facto denominator fails the elementary test the kilowatt-hour passes. Appendix A.3 makes this a rule rather than a caveat.

TPC has price and correctness without the deadline. TPCx-AI has price without the deadline. MLPerf has the deadline and the quality floor without price. Goodput has the deadline without either. FOCUS has the money without the work. OpenTelemetry has the counting without the price, and cannot resolve the tightest deadlines. SPECpower and SCI have the rate form with a different numerator. Every leg is measured somewhere by someone competent, priced nowhere in common, and no congress has been convened.

The inversion: same model, opposite rankings

The same open-weight model served by fifteen providers inverts in cost-effectiveness depending on the workload's latency requirement. The cheap-but-slow provider delivers zero compliant output for the latency-strict workload — its cost per served token is undefined — and wins decisively for the batch workload. Toggle the workload to see the inversion:

Served Token Calculator
Llama 3.3 Instruct 70B · 15 providers · Artificial Analysis, Jul 2026
Provider$/Mtoktok/s$/svt
Provider A
$0.1215.2
n/cnot compliant
Provider B
$0.1862.4
n/cnot compliant
Provider C
$0.2031.5
n/cnot compliant
Provider D
$0.2248.3
n/cnot compliant
Provider E
$0.2855.7
n/cnot compliant
Provider F
$0.3544.7
n/cnot compliant
Provider G
$0.4072.1
n/cnot compliant
Provider H
$0.4538.9
n/cnot compliant
Provider I
$0.5283.4
n/cnot compliant
Provider J
$0.5891.2
n/cnot compliant
Provider K
$0.64120.5
$0.64/Mtok
Provider L
$0.72145.8
$0.72/Mtok
Provider M
$0.85178.3
$0.85/Mtok
cheapest compliantProvider N
$0.64329.6
$0.64/Mtok
Provider O
$1.0588.1
n/cnot compliant
Inversion factor5.3×cheapest raw vs. cheapest compliant

Provider A wins on $/Mtok. Under the SLO it delivers zero compliant output — its cost per served token is undefined. The fastest provider at $0.64 is the cheapest compliant placement.

Fig. 3 — Same model, same input, opposite rankings. The inversion is invisible to $/Mtok. Providers anonymised; three anchor points [MEASURED, THIRD PARTY] Artificial Analysis, July 12, 2026; twelve [SIM]. All figures public.

9. What the efficiency era requires

Prescriptively, and in the order the history suggests:

1

A unit of account with three legs

Cost, verified output, service level: dollars per unit of useful work at a stated latency percentile and quality floor. Defined by an open specification, measurable by independent parties with published methodology, and stated with provenance. The unit that wins will be boring, auditable, and slightly too conservative, because the kWh and the TEU were.

2

Meters before markets

An instrumentation layer that measures delivered work under real workloads, as distinct from benchmark performance under ideal ones. Load factor's AI analogue — delivered-useful-work over provisioned-capacity — becomes the operator's core metric. The silicon layer has already begun competing on it: new inference architectures marketed in 2026 claim model FLOPs utilization above 80 percent against the 20 to 50 percent GPUs typically deliver, and their designers' framing — that the unit of analysis is the cluster rather than the chip — is the load-factor argument restated in hardware. Measurement is being bought and sold; it is not yet being standardized.

3

Routing as load balancing

Once work is metered in a common unit, placement becomes an optimization: which model, which silicon, which region, which batch window, subject to the task's deadline. The demand-diversity arbitrage is now quantified rather than asserted: Mooncake, the serving platform behind Kimi, replayed real traces across twenty nodes in each configuration. Roughly 100 percent of requests met the time-between-tokens objective under the disaggregated placement against 57 percent under the coupled baseline, and the platform served approximately 75 percent more requests within the same objectives — identical spend, identical silicon, a 1.75-fold difference in compliant work, produced entirely by placement. Dollars per GPU-hour is exactly equal across that comparison. The served token is the quantity in which the difference is denominated.

4

Honest depreciation, denominated in the unit

Assets carried at the value of the useful work they can still deliver, disclosed by asset class rather than blended. This is the answer to Burry that the hyperscalers cannot currently give, because giving it requires the meter.

5

Protocols owned by no one

The coordination layer is only valuable if it is universal, and only universal if no one owns it. The historical instruction is uniform: interchange worked because Hock's consortium prevented any member from owning the rules; TCP/IP beat better-funded proprietary stacks because it had no owner to distrust.

11. The implication, stated in full

The Transformer paper closed by noting the authors were excited to apply attention to other tasks, eight sentences before the architecture that would reorganize a trillion dollars of capital. Underclaiming is the characteristic failure of correct papers, so, against that precedent, the implication of this one stated without hedge:

If the pattern holds, the most valuable layer in artificial intelligence will not be the models and will not be the chips. It will be the layer that measures, prices, and routes intelligence as work: the interchange, the meter, and the dispatcher, fused. That layer does not exist yet. Its unit does not exist yet. The protocols beneath it arrived in the last eighteen months, the buyers began demanding it this year, and every prior infrastructure cycle produced exactly one such layer, within a decade of its turning point, operated by an institution that did not exist at the peak of the frenzy.

Visa's predecessor was founded in 1970, twelve years after the Fresno drop, in the wreckage of the card mania. AWS launched in 2006, six years after the NASDAQ peak, on hardware economics the crash created. The equivalent institution for intelligence is being founded, by someone, now.

The historical record adds one more regularity, and it is the one this document exists to exploit: at every prior turning point there was a memo. The proposal marked "vague but exciting," the internal warning about a tidal wave, the nine pages from a pseudonym. None of them were early; Berners-Lee wrote fifteen years after TCP/IP's design, Satoshi twenty-six years after Chaum's first digital-cash paper, and it did not matter, because the leverage was never in being first to the technology. It was in being first to state, plainly and falsifiably, which layer the value was about to move to, and to build the boring instrument that let everyone else see it move.

Railways got a clearing house. Power got a meter. Freight got a box. Money got interchange. Packets got a protocol.

Intelligence gets a meter next. The only open questions are whose, and whether it is neutral before it is bundled.

Eight predictions

10. Eight predictions

A thesis that cannot lose is not a thesis. Eight dated, checkable predictions, with the conditions under which this paper is wrong. Prediction 3 is currently half met.

Prediction Scoreboard
Last updated: July 31, 2026 · 9 predictions

Majority of inference tokens (by volume) served from open-weight models.

Kill criterion: Frontier closed models still carry whole-market majority by mid-2028.

The unit resolves through prior art: MLCommons adds a cost denominator to the MLPerf Inference Server scenario (already subject to p99 latency bounds and reference-accuracy floors), or TPC adds a per-request tail-latency constraint to TPCx-AI (which already reports $/AIUCpm@SF under audited pricing rules). Either counts as confirmation rather than refutation.

Kill criterion: Neither MLCommons nor TPC moves in this direction by end-2028.

At least one open, multi-party specification for cost-of-useful-work-under-service-level published with measurement methodology and adopted by at least two major clouds or serving frameworks in pricing or disclosure.

Kill criterion: If compute is still overwhelmingly transacted in raw GPU-hours with no work-denominated unit in commercial use by end-2029, the central claim of this paper fails on its own kill criterion.

At least two of the five largest AI spenders disclose shortened or asset-class-segmented depreciation schedules for AI silicon.

Kill criterion: Outright absence by end-2028 would indicate the accounting fog can outlast the cycle.

A material repricing of AI-exposed equities occurs, and total tokens served continues to grow through it quarter over quarter without interruption.

Kill criterion: Correction arrives and token volume also contracts for consecutive quarters — demand was all speculative and the Jevons read was off.

Cascaded GPUs (first-generation-behind) sustain secondary-market utilization above ~60% for inference workloads.

Kill criterion: Cascaded silicon scraps instead of clearing. If no cohort-level utilization series exists by end-2027, the prediction is unscorable and is counted against the thesis, not for it.

Model-routing or inference-optimization appears as a named budget category in mainstream enterprise IT surveys.

Kill criterion: Failure is a timing miss, not a falsification.

Market value created by companies operating measurement, routing, and agent-settlement layers founded after 2023 exceeds value created by companies whose primary asset is owned accelerators, on capital invested.

Kill criterion: The Visa test. Also tests Section 2.1's boundary condition: if the coordination layer is instead bundled into an existing asset owner before a neutral standard sets, the outcome is the AWS case rather than the Visa case — this paper would be right about the layer and wrong about the owner.

OpenTelemetry's generative-AI semantic conventions either reach stable status with per-request latency expressible on spans, or revise the default explicit bucket boundaries for the gen_ai.server.time_per_output_token histogram to resolve bounds below ten milliseconds. Baseline: SemConv v1.40.0, April 2026 (Development status).

Kill criterion: Failure indicates the standard telemetry bus remains unable to measure the service class the interactive market is actually buying, which would slow every prediction above it. This is the smallest concrete change any existing body could make in the direction of this paper's argument — its cost is one line in a configuration file.

This scoreboard is a public falsifiability commitment. Statuses will be updated as evidence accumulates. A prediction marked "Failed" is not a retraction — it is the mechanism by which the argument is tested.

CompanyServer/network useful lifeLast changeDirectionDisclosed impact
Amazon5 yrs (subset; rest 6)Jan 2025Shortened, 6→5, citing AI−$0.7B 2025 op. income, plus $920M accelerated depreciation
Alphabet6 yrsJan 2023Extended+$3.9B lower 2023 depreciation
Microsoft6 yrsFY2023Extended~$3.7B FY2023 benefit
Meta5 to 5.5 yrsJan 2025Extended−$2.9B 2025 depreciation expense
Oracle6 yrsFY2025Extended+$573M FY2025 net income

Table 2. Hyperscaler depreciation schedules. Four of five largest AI spenders have moved toward shorter useful-life assumptions since 2023. Source: property and equipment notes, most recent Form 10-K filings, SEC EDGAR [63].

The meter will be built. Every prior infrastructure built one. The question is not whether it arrives but who builds it, what it measures, and whether the institution that operates it is accountable to the market or to the infrastructure it prices.

The served token is a proposal, not a standard. It is offered as a draft for public comment, amendment, or replacement. The specification in Appendix A is deliberately boring. The kilowatt-hour is boring. The TEU is boring. The interchange rate is boring. That is what made each of them work: they were too simple to argue with and too useful to ignore. The unit that wins will be the one that is easiest to audit, not the one that is most elegant to describe.

If this paper is wrong, the kill criterion is in Section 10. If it is right, the predictions will resolve by 2030, and the institution that built the meter will be worth more than the institutions that built the infrastructure it measures. That is not a prediction about AI. It is a description of how infrastructure transitions have ended, every time, without exception, since 1771.

The boring instrument is the one that changes everything.

Appendix A

A draft specification for the unit

Status: draft for public comment. This appendix is deliberately boring. The kilowatt-hour is boring. That is what made it work.

Provenance key (A.6) — every figure in this appendix carries one of:
[MEASURED]Observed on the declared workload in the declared window
[SPEC]Taken from a datasheet or published price list
[CONFIG]Read from configuration rather than observed
[SIM]Produced by a model of the system
[EST]Judgment — weakest label; inherited by any derived figure
Inheritance rule: a derived figure carries the weakest label among its inputs. The red line: a comparison may be labeled [MEASURED] only if every side of it is measured.

A.1 Scope and design goals

This appendix specifies a unit of account for machine intelligence delivered as a service, and the measurement rules without which the unit is meaningless. Design goals, in priority order: the unit must denominate what the buyer buys rather than what the seller owns; two parties with the same inputs must compute the same number; every number must carry its provenance; and no single vendor may control the definition. The unit is designed to be computed today from data that already exists.

A.2 Definitions

A task is a request with a defined acceptable output. A quality floor is the acceptance test for that output, declared before measurement. A service-level objective (SLO) is the latency contract under which output has value, stated as percentile bounds on time-to-first-token (TTFT) and time-per-output-token (TPOT). Compliant work is output that passes the quality floor and meets every term of the SLO. Output that fails either is not discounted work; it is zero work that was paid for.

A.3 The unit

The cost of useful work over a window is measured in served tokens (svt), the unit this paper proposes. One served token is one output token delivered within the stated SLO and above the quality floor; a token emitted late, or wrong, is not a served token. The cost of useful work is then:

cost_per_svt = total_spend / served_tokens

where:

served_tokens = Σ(output_tokens_i × compliant_i)

compliant_i = 1 iff quality_i ≥ floor AND latency_i ≤ SLO

A.4 The conditioning tuple

A value of C is non-conforming and undefined unless stated with all six of: (1) workload — the trace or trace-class measured; (2) model artifact — weights, version, quantization, and tokenizer, explicitly; (3) placement — provider or self-host, silicon, region, batching policy, and offered load; (4) service-level objective — the percentile bounds, stated numerically; (5) quality floor — the acceptance test, stated operationally, with the evaluated fraction; and (6) window — start, end, and sustained load, with figures reported at declared percentiles. The window must be long relative to the system's own response time and stationary within itself. A conforming record reports window length divided by measured p99 request latency, and a split-half comparison of C over the first and second halves. Where the halves differ by more than ten percent, the window is non-stationary, C describes a transient rather than a placement, and the record is marked accordingly.

A.5 Pre-registration rule

The workload scope and objective are declared before measurement, and results are reported for the declared scope and for the aggregate, both. Scope declared after the fact is selection, not measurement. This rule exists because it is the one every party will be tempted to break, and because the difference between a scoped number and an aggregate number, honestly co-reported, is itself diagnostic information about where a placement fails.

A.6 Provenance taxonomy

Every figure in a conforming record carries exactly one label: [MEASURED], observed on the declared workload in the declared window; [CONFIG], read from configuration rather than observed; [SPEC], taken from a datasheet or published price list; [SIM], produced by a model of the system; or [EST], judgment. Two rules complete the taxonomy. Inheritance: a derived figure carries the weakest label among its inputs. The red line: a comparison between placements may be labeled measured only if every side of it is measured; a measured number divided by a simulated one is a simulation, and presenting it otherwise is the accounting fog reproduced at the level of a single line item.

A.7 Conforming record

A conforming measurement is publishable as one row: the six-field tuple, S, W, C, the provenance label of each, and the measuring party. Three further fields are required: an interval on C (derived from the Wilson score interval on the compliance rate, propagated through the reciprocal — a record reporting C as a scalar is non-conforming); a sensitivity curve sweeping the objective across a declared range and reporting compliance and C at each point; and a margin to the objective (the ratio of measured percentile to bound, for time-to-first-token and time-per-output-token — a placement compliant at 0.99 of its bound and one compliant at 0.40 report the same C and are not the same asset).

A.7 Sensitivity curve

Sweep both SLO bounds — watch which providers drop out and how C changes

4 / 15 compliant
100 tok/s
1.0 s
TTFT filter removes 7 providers before TPOT sweep: Provider A, Provider B, Provider C, Provider D, Provider E, Provider F, Provider H
Cheapest compliant at 100 tok/s TPOT, 1.0 s TTFT:
Provider N — $0.64/Mtok
(329.6 tok/s · 310 ms TTFT)

A.8 Governance

The specification succeeds only if no one owns it. The historical instruction is uniform: interchange worked because Hock's consortium prevented any member from owning the rules; TCP/IP beat better-funded proprietary stacks because it had no owner to distrust; TPC's price-performance metrics are trusted across four decades of adversarial vendor competition because the pricing rules and the audit are the consortium's and not the submitter's; and MCP's donation to a neutral foundation followed the same logic. This draft is offered accordingly, for adoption, amendment, or replacement by any multi-party body that preserves A.3 through A.7. The two bodies best positioned are named in Prediction 2.

A.9 Worked example, entirely from public data

The same open-weight model (Llama 3.3 Instruct 70B) is served by 15 providers. The public price spread is 8.6×: $0.12 per million blended tokens at the cheapest (fp8 quantized) to $1.05 at the most expensive. Throughput spreads further: 15.2 tokens per second at the cheapest against 329.6 tokens per second at the fastest. Expand each step below to follow the tuple from declaration to C.

A.9 Worked example

Six fields → C — click each step to expand

Llama 3.3 70B · 15 providers · July 2026
Buyer 1 — Voice agent
Llama 3.3 Instruct 70B · fp8 quantization · Provider N
Buyer 2 — Batch job
Llama 3.3 Instruct 70B · fp8 quantization · Provider A

Both buyers use the same open-weight model. The fp8 quantization must be declared per A.4 — the two buyers are not entitled to assume the same quality floor until it is stated and tested.

Buyer 1 — Voice agent
TTFT ≤ 1.0 s · sustained output ≥ 100 tok/s (voice playback floor)
Buyer 2 — Batch job
Complete within 12 hours · no per-token latency floor

The SLO is the condition that determines whether a token was worth anything at all. A voice reply arriving in 4 s is a refund. A batch job finishing in 11 h 59 m is a success.

Buyer 1 — Voice agent
$0.64 / million blended tokens (Provider N list price)
Buyer 2 — Batch job
$0.12 / million blended tokens (Provider A list price)

S is the invoice figure — list, committed, or spot — disclosed per A.5. Both figures are [SPEC] from public provider pricing, accessed July 2026.

Buyer 1 — Voice agent
W ≈ billed output (329.6 tok/s · TTFT 0.95 s — both bounds met)
Buyer 2 — Batch job
W ≈ billed output (15.2 tok/s — latency floor not applicable)

W is the token count that passed the SLO. For Buyer 1, Provider A's 15.2 tok/s fails the 100 tok/s floor, so W ≈ 0 and C diverges. For Buyer 2, every provider is compliant, so W = billed output everywhere.

Buyer 1 — Voice agent
C ≈ $0.64 / Msvt (cheapest compliant placement)
Buyer 2 — Batch job
C ≈ $0.12 / Msvt (cheapest compliant placement)

Same weights, same input, opposite rankings. The inversion is invisible to both metrics the industry currently prices with: $/GPU-hour cannot see the workload at all, and raw $/Mtok ranks the placements identically for both buyers — wrongly for one of them.

Buyer 1 — Voice agent
Buyer 1 pays $0.64 / Msvt — correct placement
Buyer 2 — Batch job
Buyer 2 pays $0.12 / Msvt — correct placement

The spread between those two correct answers — a factor of 5.3× on identical work — is the routing margin of Section 9, computed from nothing but public numbers. Buyer 1's premium placement is paying 21.7× the speed for a deadline that cannot use it.

All figures [SPEC] for prices, [MEASURED, THIRD PARTY] for latency/throughput from public benchmarks. This example uses p50 medians; a fully conforming record requires the declared percentile (typically p99).

Same weights, same input, opposite rankings, and the inversion is invisible to both metrics the industry currently prices with. The spread between those two correct answers — a factor of five on identical work — is the routing margin of Section 9, computed from nothing but public numbers. One further detail the tuple forces into the open: the cheapest artifact is fp8-quantized, so the two buyers are not entitled to assume the same quality floor until it is declared and tested, which under A.4 they must state, and under current market convention they are never asked to.

A.10 What this specification excludes, knowingly

Multi-tenant attribution of S is no longer excluded. Two bases are defensible and they bracket the honest answer. Under reserved share, spend is attributed in proportion to reserved capacity, which charges idle reserved capacity to the tenant that reserved it; this is the correct basis under a reservation and the conservative one. Under token share, spend is attributed in proportion to realized compliant output, which ignores idle capacity entirely and is a strict lower bound. A conforming record reports C on both bases and names which it uses. The gap between them is not noise to be eliminated. It is the tenant's own load factor restated in dollars, and it is diagnostic in exactly the way Section 4's diversity factor was.

A.10 Open problems

Four genuinely open problems — stated so their omission is not mistaken for resolution

4 deferred items

The current specification uses a binary quality floor: a token either passes or fails the declared acceptance criterion. Richer floors — continuous quality scores, task-specific rubrics, human-preference ratings — are measurable today but require either a standardized evaluation harness or a trusted third-party scorer. Neither exists at the cross-vendor level. The binary floor is a deliberate simplification that makes C computable from public data right now; it is not a claim that quality is binary.

The tuple requires the price basis to be declared (A.5) but does not normalize across it. A list-price C and a committed-rate C for the same workload are not directly comparable — the committed rate embeds a capacity reservation that the list rate does not. Normalizing them requires either a standard discount schedule (which vendors do not publish) or a multi-party convention on how to amortize commitment premiums. The tuple's disclosure requirement is the first step; normalization is the second, and it requires a body.

Joules per compliant token is measurable today: power draw is reported by NVIDIA's NVML API, AMD's ROCm SMI, and Intel's RAPL interface, all accessible without special access. The energy term is the natural second axis of C — it would make the unit a two-dimensional cost vector (dollars, joules) rather than a scalar. Deferral is a scope decision, not a technical one. The paper notes this explicitly so the omission is not mistaken for a claim that energy is unimportant.

Reference SLO classes are the analogue of standard voltages: a small set of named, interoperable service levels (e.g. 'interactive', 'batch', 'real-time voice') that buyers and sellers could reference without negotiating every bound from scratch. They would make C values directly comparable across providers and workloads. Setting them unilaterally in this specification would be the wrong governance model — it would create a de facto standard without the legitimacy that comes from multi-party ratification. This is the strongest argument for why the unit needs a body, and the strongest reason to publish the specification as a draft.

Each is a reason this appendix is a draft. None is a reason the unit cannot be computed today, because A.9 just computed it.

Each is a reason this appendix is a draft. None is a reason the unit cannot be computed today, because A.9 just computed it.

Reproduce the numbers

Every figure in this paper is computed from public data. Here is the data.

The three artifacts below are the complete provenance chain for every exhibit and every quantitative claim in the paper. No proprietary data, no model access, no API keys. The inversion in Fig. 3 can be reproduced in a spreadsheet in under ten minutes.

Data provenance manifest · v1.0 · July 15, 2026
CSV
Exhibit 148 rows

Hyperscaler capex vs. inference cost, 2021–2026

Combined hyperscaler capex (Microsoft, Google, Amazon, Meta) from company filings. Inference cost per million tokens at GPT-3.5 capability from Stanford AI Index 2025 and Artificial Analysis. 48 rows, 6 columns.

CSV
Exhibit 285 rows

Unit cost indexed to turning point — five infrastructure cycles

Unit cost series for railways (Crafts 2004), electricity (Joskow 1997), shipping (Levinson 2006), telecom (Odlyzko 2003), and AI inference (Artificial Analysis 2025). Indexed to 100 at each infrastructure's turning point. 85 rows, 8 columns.

PDF
Worked example12 pages

Served token inversion: Llama 3.3 70B across 15 providers

Step-by-step walkthrough of the Section 8 / Fig. 3 calculation. Provider pricing and throughput from Artificial Analysis June 2026. Shows the interactive vs. batch inversion from raw inputs to final $/svt figures. Reproducible in a spreadsheet.

All sources are public. No proprietary datasets, no model access required. The worked example reproduces Fig. 3 from provider pricing pages alone.