
By Matt Toussain, CIO, Open Security Research
We reverse-engineered the compute hidden inside the biggest frontier-AI subscriptions. The lowest subscription price per unit of API-list-equivalent compute came from an unexpected place. Then actual developer behavior broke the spreadsheet.
The subscription that provides the most compute for the money is not necessarily the subscription that creates the most value for the developer.
I did not set out to build an AI-subscription benchmark. I set out to buy more Cursor.
My Cursor Ultra dashboard was behaving like a warning light. The $400 Other Models allowance, the pool spent at third-party API prices, was disappearing faster than the separate Cursor Models pool used by Grok 4.6, Grok 4.5, and Composer 2.5. At my observed pace, one Ultra subscription was no longer enough. The obvious move was a second one.
Then my COO Joshua Christman mentioned I should consider SuperGrok Heavy. At first, I scoffed, but then I took a deeper look and the math does not lie.
The spreadsheet said Heavy. My working model said that a $300 Heavy subscription represented roughly 4.622 billion monthly fresh-input-equivalent frontier tokens, compared with roughly 4.245 billion for two $200 Cursor Ultra subscriptions. Heavy was not only cheaper. On the benchmark’s central estimate, it offered more Grok compute.
But two Ultras bought something Heavy did not: another $400 Other Models pool, another first-party Cursor pool, and the ability to spend that capacity inside the development environment I was already using. The economic winner and the behavioral winner were not automatically the same product.
That was the real investigation.
The quota is the product, and the quota is mostly hidden
Frontier AI subscriptions are strange financial instruments. You pay a fixed monthly fee for access to expensive, variable-cost infrastructure. The vendor usually does not tell you how many tokens you bought.
Instead, you get abstractions:
- Cursor publishes a dollar-denominated Other Models allowance but calls the first-party Cursor Models allocation “generous included usage.” Its documentation does not publish the first-party token pool.
- xAI moved paid Grok products to one weekly, compute-weighted percentage pool. Different actions consume different shares; no official SuperGrok or Heavy token total appears on the page.
- Anthropic sells Max as 5x or 20x more usage per five-hour session than Pro, while still applying weekly limits.
- Google says usage now depends on prompt complexity, features, and chat length; the meter refreshes every five hours until a weekly cap.
- OpenAI says the number of Codex messages varies with the model, context, reasoning, tool use, retrieval, and caching. Prompt length alone is not a reliable estimate.
The provocative version is “they don’t want you to know.” The defensible version is more useful: the observable product is opaque enough that a customer cannot infer capacity from the plan page. Motive is not required. The measurement problem exists either way.
Traditional subscription framing
Fixed monthly fee, abstract usage meter, unclear underlying quota.
What this report tries to recover
List-equivalent compute, comparable economics, and actual decision value.
That opacity matters because a subscription is perishable. Unused capacity expires. The best allowance is not the largest allowance; it is the largest amount of capacity you will actually consume on work you actually want to route to that model.
How incomparable subscription models were normalized
We wanted frontier-class, high-reasoning models, not each vendor’s cheapest token generator. The normalized set was:
| Comparison family | Normalized model | Reasoning | Speed tier | Fresh / cached / output API price per 1M |
|---|---|---|---|---|
| xAI / Cursor | Grok 4.6 | xhigh | Standard, non-Fast | $2.00 / $0.50 / $6.00 |
| Anthropic | Claude Opus 5 | xhigh | Standard, non-Fast | $5.00 / $0.50 / $25.00 |
| Gemini 3.1 Pro Preview | High | Standard, non-Fast | $2.00 / $0.20 / $12.00 | |
| OpenAI | GPT-5.6 Sol | xhigh | Standard, non-Fast | $5.00 / $0.50 / $30.00 |
| Z.AI | GLM-5.3 | max | Standard, non-Fast | $1.40 / $0.26 / $4.40 |
Those are short-context standard prices. Long-context requests can step up: Grok 4.6 doubles its rates at 200,000 prompt tokens, Gemini 3.1 Pro moves to $4 / $0.40 / $18 above 200,000, and GPT-5.6 Sol applies a long-context uplift above 272,000. Cursor also documented Fast as its paid-plan default during this research; Fast is priced above Standard. The benchmark normalizes to Standard so speed-tier premiums do not masquerade as quota.
Why raw token counts are misleading
Agentic coding workloads reuse enormous contexts. In the earlier experiment, one measured SuperGrok workload was about 95.85% cached input. A billion cache reads are not economically equivalent to a billion fresh-input tokens, and neither is equivalent to a billion output tokens.
So we convert the observed or estimated mix into retail API-list-equivalent value, then divide by that model’s fresh-input price:
API-list-equivalent value
= fresh input × fresh-input list price
+ cached input × cached-input list price
+ output/reasoning × output list price
fresh-input-equivalent frontier compute (FIE tokens)
= API-list-equivalent value / fresh-input list price
If a workload contains $200 of list-priced cached input and output, and the model’s fresh-input price is $2 per million, it represents 100 million FIE tokens. The unit is deliberately input-token-shaped, but it is not a transcript token count.
Three caveats belong in every serious reading of this report
- FIE tokens are list-equivalent compute, not official quotas. Most absolute capacity points are reverse-engineered central estimates.
- FIE is not quality-adjusted. One billion Grok FIE tokens and one billion Opus FIE tokens do not guarantee equal work.
- API list price is not serving cost. It is a catalog used for conversion, not a claim about what a vendor spends.
The decision that started it: Heavy or a second Ultra?
Here is the quantitative case.
| Choice | Monthly price | Monthly FIE compute | Subscription $ / 1M FIE | What else the money buys |
|---|---|---|---|---|
| Cursor Ultra | $200 | 2.122B | $0.094 | $400 Other Models pool plus unpublished Cursor Models pool |
| SuperGrok Heavy | $300 | 4.622B | $0.065 | xAI’s highest consumer usage tier, Grok products, Build, and Grok Bot |
| 2x Cursor Ultra (modeled) | $400 | 4.245B | $0.094 | Two $400 Other Models pools plus two Cursor Models pools |
On the central FIE estimate, Heavy beats two Ultras on both price and Grok capacity. It supplies about 8.9% more FIE compute for 25% less money. Its unit cost is roughly 31% lower.
Current official pages do not support the earlier wording that “Heavy includes Cursor Ultra.” xAI lists Heavy at $300 and Cursor Ultra at $200 as separate products; both include Grok Bot. Cursor sells Ultra. xAI sells Heavy. Their quotas are not documented as one pool.
Communication with the Cursor AI support bot indicates a rapidly unifying set of subscription tiers. It is likely that this promotion will eventually be converted into a single product.
Raw-token sensitivity
The raw-token sensitivity tells a related but different story. The earlier work placed one Ultra around 6–7 billion raw first-party Grok-equivalent tokens, two Ultras around 13 billion, and Heavy around 14.2–16.5 billion. In that model, paying another $100 for two Ultras buys another $400 Other Models allowance while giving up roughly 1.2–3.5 billion raw Grok-equivalent tokens.
That trade is not absurd. It is a portfolio decision.
If I would voluntarily choose Opus or Sol for half my serious tasks, then Heavy’s unused Grok capacity expires worthless. If I would happily route almost everything to Grok, buying expensive third-party optionality is waste. Capacity has no utility independent of model preference.
The spreadsheet says Grok; the developer may say Opus.
The broader subscription benchmark
Every 2x, 4x, or 8x row below is a linear stack, not a vendor SKU. It assumes separate eligible subscriptions or accounts, linear quota scaling, no volume discount, and no shared enforcement ceiling.
| Family | Plan or scenario | Type | $ / mo | Monthly FIE compute | $ / 1M FIE | Evidence confidence |
|---|---|---|---|---|---|---|
| xAI / Cursor | X Premium | Native | $8 | 83.3M | $0.096 | Low |
| xAI / Cursor | SuperGrok | Native | $30 | 250.0M | $0.120 | Medium |
| xAI / Cursor | X Premium+ | Native | $40 | 250.0M | $0.160 | Low–medium |
| xAI / Cursor | SuperGrok Plus | Native | $100 | 833.3M | $0.120 | Low |
| xAI / Cursor | Cursor Ultra | Native | $200 | 2.122B | $0.094 | Medium–low |
| xAI / Cursor | SuperGrok Heavy | Native | $300 | 4.622B | $0.065 | Medium–low |
| xAI / Cursor | 2x Cursor Ultra | Modeled | $400 | 4.245B | $0.094 | Scenario |
| Anthropic | Claude Pro | Native | $20 | 52.3M | $0.382 | Low |
| Anthropic | Claude Max 5x | Native | $100 | 261.7M | $0.382 | Low–medium |
| Anthropic | Claude Max 20x | Native | $200 | 550.4M | $0.363 | Low–medium |
| Anthropic | 2x Claude Max 20x | Modeled | $400 | 1.101B | $0.363 | Scenario |
| Anthropic | 8x Claude Max 20x | Modeled | $1,600 | 4.403B | $0.363 | Scenario |
| Google AI Plus | Native | $4.99 | 17.9M | $0.279 | Low–medium | |
| Google AI Pro | Native | $19.99 | 35.8M | $0.558 | Low–medium | |
| Google AI Ultra 5x | Native | $99.99 | 178.8M | $0.559 | Low–medium | |
| Google AI Ultra 20x | Native | $199.99 | 715.0M | $0.280 | Low–medium | |
| 2x Google AI Ultra 20x | Modeled | $399.98 | 1.430B | $0.280 | Scenario | |
| 4x Google AI Ultra 20x | Modeled | $799.96 | 2.860B | $0.280 | Scenario | |
| OpenAI | ChatGPT Plus | Native | $20 | 29.5M | $0.678 | Low |
| OpenAI | ChatGPT Pro 5x | Native | $100 | 147.4M | $0.678 | Low–medium |
| OpenAI | ChatGPT Pro 20x | Native | $200 | 589.7M | $0.339 | Low–medium |
| OpenAI | 2x ChatGPT Pro 20x | Modeled | $400 | 1.180B | $0.339 | Scenario |
| OpenAI | 8x ChatGPT Pro 20x | Modeled | $1,600 | 4.718B | $0.339 | Scenario |
| Z.AI | GLM Coding Lite | Native | $18 | 93.2M | $0.193 | Medium–high |
| Z.AI | GLM Coding Pro | Native | $72 | 570.2M | $0.126 | Medium–high |
| Z.AI | GLM Coding Max | Native | $160 | 1.329B | $0.120 | Medium–high |
| Z.AI | 2x GLM Coding Max | Modeled | $320 | 2.658B | $0.120 | Scenario |
| Z.AI | 4x GLM Coding Max | Modeled | $640 | 5.316B | $0.120 | Scenario |
Many of the estimates are based on low-confidence reverse mathematics. While imprecise, they should be generally proximal to reality. A four-decimal calculation built on a low-confidence quota remains low-confidence.
The X Premium and Premium+ prices are retained as historical comparison points from an archived March 18, 2025 official X page because the live US price page was inaccessible; they are not verified as current 2026 prices.
What roughly $400 buys across providers
Figure 1 asks a practical question: what fits into roughly $400?
Heavy is the only native plan in the extreme upper-left. Cursor Ultra and the modeled two-Ultra stack also enter the high-value zone. Z.AI Max sits just below its value threshold; the modeled 2x Max crosses the capacity threshold at $320.
The Google and OpenAI curves contain genuine-looking pricing cliffs. Google’s $19.99 Pro and $99.99 Ultra 5x points are less efficient than both the $4.99 Plus and $199.99 Ultra 20x points under the literal public multipliers. OpenAI’s $200 Pro 20x tier is roughly twice as efficient as its $20 Plus and $100 Pro 5x tiers in this central series.
We did not smooth those discontinuities. If the plan schedule is nonlinear, the chart should look nonlinear.
The most uncomfortable point is the archived $8 X Premium row, marked with a dagger in the charts. It looks spectacular on unit value and carries low confidence, but its price is historical rather than verified current. It is a reminder that the left edge of this chart is not the recommendation. A tiny cheap quota can be efficient and still be useless to a heavy developer. Google’s AI Plus subscription shares similar characteristics.
What does Heavy-scale capacity cost elsewhere?
Figure 2 changes the question from “what can I buy for about $400?” to “what would it cost to approach this much compute?”
- SuperGrok Heavy: $300 → 4.622B FIE
- Z.AI 4x Max: $640 → 5.316B FIE
- Google 4x Ultra 20x: $799.96 → 2.860B FIE
- Anthropic 8x Max 20x: $1,600 → 4.403B FIE
- OpenAI 8x Pro 20x: $1,600 → 4.718B FIE
The words “would cost” matter. Eight consumer accounts are not an elegant purchasing strategy, and a vendor may restrict or aggregate them. These points isolate the price curve. They do not promise that stacking works operationally. It is particularly worth noting that fraud prevention systems may flag subscription stacking as potentially unwanted or malicious activity.
The Z.AI comparison
Z.AI is the real second story.
Unlike most vendors in this report, Z.AI publishes enough machinery to audit the quota model. Its current Coding Plan documentation lists:
- 10,000 / 60,000 / 140,000 weekly credits for Lite / Pro / Max
- GLM-5.3 credit multipliers of 6.9 fresh input, 1.7 cached input, and 24 output per 10,000 tokens
- token ranges under a published 90.9% coding cache-hit assumption
- a defined weekday peak window and 50% off-peak credit charging
- multiple cache-hit sensitivity rows
Its GLM-5.3 documentation supports low, high, and max, defaults to max, and recommends max for complex coding. Its API list price is approximately $1.40 fresh input, $0.26 cached input, and $4.40 output per million.
That does not make the central point official. We still convert weekly ranges and a workload mix into a monthly FIE estimate. But the evidence chain is stronger. As a result, we assigned higher confidence, medium–high, to the Z.AI native points rather than the lower confidence applied to unpublished xAI, Cursor, Claude, Google, and OpenAI absolute limits.
The indexed official-checkout result used here is $18 / $72 / $160. During final verification, other client-rendered Z.AI surfaces exposed inconsistent promotional or product-namespace prices. We preserved the indexed result and its evidence limits, named the access date, and logged the discrepancy rather than guessing which unlabeled card would win tomorrow. While this is a set of assumptions, we are confident that it details the best-available numbers on Z.AI at the time of writing.
At 5.316B FIE for a modeled $640, 4x Z.AI Max subscriptions enter the gold region. That is a compelling second subscription-economics story. It is not Heavy’s unit economics: Z.AI remains near $0.120 per million FIE, about 1.85x Heavy’s roughly $0.065. Depending on the future of promotions and subscription subsidies, there is a significant likelihood of this relationship inverting in the future.
Why developer behavior changes the economics
The benchmark asks a deliberately narrow question: how much frontier compute does the subscription appear to contain at retail API-list equivalence?
A developer is buying on output: What can I build with this today?
Claude may be preferred for long-running code changes, editorial judgment, or document-heavy work. Codex may be preferred for its harness, local tooling, review behavior, or the ability to move between Sol, Terra, and Luna. Google packages storage and consumer surfaces. Cursor integrates routing, codebase context, review, and third-party models. Grok offers build, bot, media, and an unusually deep first-party allowance. Z.AI can plug into coding harnesses while preserving a different cost and openness posture.
Realized subscription value
realized subscription value
= included capacity
× probability you choose that model
× probability you consume it before reset
× workflow fit
Effectively: a plan with 5B estimated FIE can underperform a 500M plan if you avoid the model or cannot leverage its product surface. A smaller Claude or OpenAI subscription may be rational insurance for tasks where the harness or output is worth more than the token efficiency.
The subscription becomes a season pass; the objective is to ski until the resort regrets selling it. But a season pass to the wrong mountain is not a bargain. Like a time-share, it is a burden.
Stop hiding spend inside a ratio
Ratios are useful, but they compress two different realities into one number.
The absolute map makes three things hard to miss:
- Heavy is the only native point above 4B FIE.
- The cheapest native paths to at least 1B in this model are Z.AI Max at $160 and Cursor Ultra at $200.
- The Anthropic and OpenAI capacity-match scenarios land near Heavy only after spend reaches $1,600.
This is why a ratio alone is dangerous. Google AI Plus looks efficient at $4.99, but its central capacity is 17.9M. Heavy looks efficient and enormous. Those are different economic objects.
What users can measure, and what they cannot
Public discussion contains valuable evidence, but much of it is contaminated by promotions, silent speed modes, client bugs, changing limits, and incomparable workloads.
The strongest examples we found illustrate the problem more than they solve it:
- Cursor staff and forum replies confirm separate first-party and Other Models behavior, while declining to translate the first-party pool into a fixed token count. A temporary Grok 4.6 launch multiplier further distorted dashboard observations.
- A detailed OpenAI Codex issue reported four user turns expanding into 1,031 Sol requests over a long tool loop, about 98% cache-read, while the weekly indicator moved from 59% to exhausted. That is powerful evidence that “tokens processed” and “subscription percentage” are not interchangeable.
- Another Codex Pro 20x report later found a Fast/Priority configuration. OpenAI’s official documentation says GPT-5.6 Fast consumes credits at 2.5x Standard.
- Anthropic and Google community reports include dramatic depletion incidents, but some look like metering bugs or account issues.
Token efficiency and actual throughput
Quota is only half of throughput. The other half is tokens spent by a model to finish a task.
Artificial Analysis’ public model pages report that Grok 4.6 and GPT-5.6 Sol reached the same Intelligence Index score with similar total output-token volume in its suite, while Opus 5 scored higher with more output and GLM-5.3 used substantially more output at a nearby score. The source audit also included Grok 4.5 as a historical comparator. Grok 4.5 is particularly interesting because of its extreme token efficiency to task completion and its extremely low direct API pricing.
That does not settle the matter. Token efficiency may amplify output-price dominance for both xAI and OpenAI. On a different workload, quality or shorter iteration count may invert that hypothesis.
Whether current subscription economics are sustainable
The chart cannot answer the sustainability question by itself.
At Grok 4.6’s $2 per million standard fresh-input list price, the Heavy central estimate converts to:
4.622B FIE × $2/M = about $9,244 of API-list-equivalent compute
$9,244 / $300 = about 30.8x the subscription sticker
That is the defensible dramatic number, not the oft-touted $21,000+.
It is also easy to misconstrue. xAI is not necessarily spending $9,244 to service the subscriber. Cached inference, spare capacity, batching, internal hardware economics, model routing, and the strategic value of user acquisition can all make marginal cost radically different from the retail API menu. The calculation measures a retail API-list-equivalent value-to-price ratio, not COGS or economic viability.
Using open-weight as a control moves the bar closer to commodity compute, in some cases. The llm-d project served GLM-5.2-FP8 on 64 H200s using production Claude Code traces and calculated GPU-rental-only costs of $0.05–$0.08 per million input tokens and $1.92–$3.55 per million output tokens from measured throughput and a $2.71 per H200-hour assumption under sustained load.
An August 18 Vast H200 snapshot found roughly $0.29–$0.48 per million FIE: about $386–$644 for 1.329B FIE, or 2.4–4.0x the $160 Z.AI Max subscription, although still far below the $1.40 per million FIE API baseline. The calculation excludes storage, bandwidth, engineering, orchestration, and idle capacity.
Even a saturated rented-GPU deployment therefore fails to undercut Max, making the subscription discount even harder to dismiss.
We can responsibly say the subscription price is far below the benchmark’s retail API-list equivalence.
The potential impact of Grok 4.7
The economically interesting scenario is conditional.
If Grok 4.7 reaches Opus / Sol-class capability on the work developers care about, and if xAI retains roughly similar subscription economics and token efficiency, the behavioral reason to maintain large competing frontier subscriptions could shrink dramatically.
Every variable is meaningful. Capability can miss. Quotas can tighten. Pricing can change. A new model can consume more compute per task. Cursor, OpenAI, and Anthropic can retain workflow advantages that raw capability does not erase. Documents, code review, ecosystem integrations, safety behavior, and harness quality can justify a smaller subscription even when the token spreadsheet points elsewhere.
Grok may not win, but the pressure an economically credible Grok 4.7 would create represents a titanic shift in the model provider space.
The final buying framework
On this benchmark’s central estimates, SuperGrok Heavy is the strongest native subscription for high-volume frontier compute. It lands at roughly 4.622B monthly FIE, $0.065 per million FIE, and a $300 sticker. Z.AI is the most compelling second story: it publishes materially better quota mechanics and reaches 5.316B FIE at a modeled $640.
But the benchmark did not make the purchase decision automatic.
Two Cursor Ultras cost $100 more than Heavy and offer slightly less estimated Grok FIE. They also provide a second $400 Other Models pool and another integrated Cursor first-party pool. That optionality is valuable only if it matches actual routing behavior.
The right buying sequence is therefore:
- Identify the model and product surface that currently blocks your work.
- Estimate how much of each allowance you will consume before reset.
- Compare the marginal subscription against on-demand overflow.
- Use FIE economics to expose expensive capacity, not to pretend all tokens are equal.
- Re-run the measurement after major model, quota, speed-tier, or pricing changes.
The best AI subscription is not necessarily attached to the best AI model. It is attached to the model you will use, inside the workflow you will keep, at a quota you can exhaust, for a price below your next-best source of capacity.
That is a harder question than “which model is best?” It is also the one that drives the economics of your bill.
Methodology and assumptions
Benchmark date and scope
- Price and official-product verification date: 2026-08-18
- Scope: five comparison families, xAI / Cursor, Anthropic, Google, OpenAI, and Z.AI
- Consumer or individual subscription prices are US web prices where available
- Teams, enterprise, API-key-only, promotional annual equivalents, taxes, and app-store markups are excluded
Capacity evidence hierarchy
- Published vendor credits, token ranges, and price tables
- Official relative multipliers, windows, and pool descriptions
- Reproducible empirical usage measurements with cache/input/output mix
- High-signal user reports with screenshots or local logs
- Editorial extrapolation
Z.AI has the strongest native capacity evidence because it publishes weekly credits, cache assumptions, token ranges, and charging windows. Arithmetic confidence is high across all rows; realization confidence is not.
Monthly conversion
The central estimates preserve the calendar-month convention established in the prior benchmark. A four-week month versus 365.25 / 12 / 7 weeks changes a weekly extrapolation by roughly 8.7%, which is larger than display rounding. Do not read the third decimal place as measurement precision.
Native versus modeled
- Filled chart points are currently sold retail plans.
- Open chart points are linear 2x / 4x / 8x stacks.
- A modeled stack assumes independent capacity and spend scale linearly.
- “2x Cursor Ultra” means two subscriptions or accounts, not a $400 Cursor SKU.
- Capacity-match stacks are analytical counterfactuals, not operational recommendations.
Confidence labels
- Medium–high: published allowance mechanics plus normalization; current Z.AI native points
- Medium / medium–low: useful empirical anchor but absolute vendor allowance remains unpublished
- Low / low–medium: relative multiplier or anecdotal anchor with substantial extrapolation
- Scenario: arithmetic is deterministic if the parent point and linear-stacking assumption hold; operational feasibility is unverified
Exclusions and rejected claims
- The distrusted outlier series from the earlier draft is excluded completely.
- No claim that a plan is “unlimited.”
- No claim that FIE equals transcript tokens, quality-adjusted work, serving cost, or vendor subsidy.
- No claim that retail H200 rental equals a hyperscale provider’s marginal hardware cost.
- No claim that Heavy includes Cursor Ultra; current official pages show separate plans.
- No unverified screenshot or inaccessible social post is presented as captured evidence.
- No proprietary Artificial Analysis chart is copied.
This article is part of Open Security Research: technical, evidence-driven work on the operational realities behind fast-moving security and AI markets.

