Research

The Best Sub Isn't the Best Model

Matt Toussain

Cursor Ultra Plan & Usage dashboard showing Cursor Models at 86% used and Other Models at 100% used

By Matt Toussain, CIO, Open Security Research

We reverse-engineered the compute hidden inside the biggest frontier-AI subscriptions. The lowest subscription price per unit of API-list-equivalent compute came from an unexpected place. Then actual developer behavior broke the spreadsheet.

The subscription that provides the most compute for the money is not necessarily the subscription that creates the most value for the developer.

I did not set out to build an AI-subscription benchmark. I set out to buy more Cursor.

My Cursor Ultra dashboard was behaving like a warning light. The $400 Other Models allowance, the pool spent at third-party API prices, was disappearing faster than the separate Cursor Models pool used by Grok 4.6, Grok 4.5, and Composer 2.5. At my observed pace, one Ultra subscription was no longer enough. The obvious move was a second one.

Then my COO Joshua Christman mentioned I should consider SuperGrok Heavy. At first, I scoffed, but then I took a deeper look and the math does not lie.

The spreadsheet said Heavy. My working model said that a $300 Heavy subscription represented roughly 4.622 billion monthly fresh-input-equivalent frontier tokens, compared with roughly 4.245 billion for two $200 Cursor Ultra subscriptions. Heavy was not only cheaper. On the benchmark’s central estimate, it offered more Grok compute.

But two Ultras bought something Heavy did not: another $400 Other Models pool, another first-party Cursor pool, and the ability to spend that capacity inside the development environment I was already using. The economic winner and the behavioral winner were not automatically the same product.

That was the real investigation.

The quota is the product, and the quota is mostly hidden

Frontier AI subscriptions are strange financial instruments. You pay a fixed monthly fee for access to expensive, variable-cost infrastructure. The vendor usually does not tell you how many tokens you bought.

Instead, you get abstractions:

  • Cursor publishes a dollar-denominated Other Models allowance but calls the first-party Cursor Models allocation “generous included usage.” Its documentation does not publish the first-party token pool.
  • xAI moved paid Grok products to one weekly, compute-weighted percentage pool. Different actions consume different shares; no official SuperGrok or Heavy token total appears on the page.
  • Anthropic sells Max as 5x or 20x more usage per five-hour session than Pro, while still applying weekly limits.
  • Google says usage now depends on prompt complexity, features, and chat length; the meter refreshes every five hours until a weekly cap.
  • OpenAI says the number of Codex messages varies with the model, context, reasoning, tool use, retrieval, and caching. Prompt length alone is not a reliable estimate.

The provocative version is “they don’t want you to know.” The defensible version is more useful: the observable product is opaque enough that a customer cannot infer capacity from the plan page. Motive is not required. The measurement problem exists either way.

Cursor Ultra usage dashboard showing the separate Cursor Models and Other Models pools
Cursor Ultra usage dashboard

Traditional subscription framing

Fixed monthly fee, abstract usage meter, unclear underlying quota.

What this report tries to recover

List-equivalent compute, comparable economics, and actual decision value.

That opacity matters because a subscription is perishable. Unused capacity expires. The best allowance is not the largest allowance; it is the largest amount of capacity you will actually consume on work you actually want to route to that model.

How incomparable subscription models were normalized

We wanted frontier-class, high-reasoning models, not each vendor’s cheapest token generator. The normalized set was:

Comparison family Normalized model Reasoning Speed tier Fresh / cached / output API price per 1M
xAI / Cursor Grok 4.6 xhigh Standard, non-Fast $2.00 / $0.50 / $6.00
Anthropic Claude Opus 5 xhigh Standard, non-Fast $5.00 / $0.50 / $25.00
Google Gemini 3.1 Pro Preview High Standard, non-Fast $2.00 / $0.20 / $12.00
OpenAI GPT-5.6 Sol xhigh Standard, non-Fast $5.00 / $0.50 / $30.00
Z.AI GLM-5.3 max Standard, non-Fast $1.40 / $0.26 / $4.40

Those are short-context standard prices. Long-context requests can step up: Grok 4.6 doubles its rates at 200,000 prompt tokens, Gemini 3.1 Pro moves to $4 / $0.40 / $18 above 200,000, and GPT-5.6 Sol applies a long-context uplift above 272,000. Cursor also documented Fast as its paid-plan default during this research; Fast is priced above Standard. The benchmark normalizes to Standard so speed-tier premiums do not masquerade as quota.

Why raw token counts are misleading

Agentic coding workloads reuse enormous contexts. In the earlier experiment, one measured SuperGrok workload was about 95.85% cached input. A billion cache reads are not economically equivalent to a billion fresh-input tokens, and neither is equivalent to a billion output tokens.

So we convert the observed or estimated mix into retail API-list-equivalent value, then divide by that model’s fresh-input price:

API-list-equivalent value
 = fresh input × fresh-input list price
 + cached input × cached-input list price
 + output/reasoning × output list price

fresh-input-equivalent frontier compute (FIE tokens)
 = API-list-equivalent value / fresh-input list price

If a workload contains $200 of list-priced cached input and output, and the model’s fresh-input price is $2 per million, it represents 100 million FIE tokens. The unit is deliberately input-token-shaped, but it is not a transcript token count.

Three caveats belong in every serious reading of this report

  1. FIE tokens are list-equivalent compute, not official quotas. Most absolute capacity points are reverse-engineered central estimates.
  2. FIE is not quality-adjusted. One billion Grok FIE tokens and one billion Opus FIE tokens do not guarantee equal work.
  3. API list price is not serving cost. It is a catalog used for conversion, not a claim about what a vendor spends.

The decision that started it: Heavy or a second Ultra?

Here is the quantitative case.

Choice Monthly price Monthly FIE compute Subscription $ / 1M FIE What else the money buys
Cursor Ultra $200 2.122B $0.094 $400 Other Models pool plus unpublished Cursor Models pool
SuperGrok Heavy $300 4.622B $0.065 xAI’s highest consumer usage tier, Grok products, Build, and Grok Bot
2x Cursor Ultra (modeled) $400 4.245B $0.094 Two $400 Other Models pools plus two Cursor Models pools

On the central FIE estimate, Heavy beats two Ultras on both price and Grok capacity. It supplies about 8.9% more FIE compute for 25% less money. Its unit cost is roughly 31% lower.

Current official pages do not support the earlier wording that “Heavy includes Cursor Ultra.” xAI lists Heavy at $300 and Cursor Ultra at $200 as separate products; both include Grok Bot. Cursor sells Ultra. xAI sells Heavy. Their quotas are not documented as one pool.

Cursor support exchange clarifying the relationship between Cursor Ultra and SuperGrok Heavy
Plan-separation support exchange

Communication with the Cursor AI support bot indicates a rapidly unifying set of subscription tiers. It is likely that this promotion will eventually be converted into a single product.

Raw-token sensitivity

The raw-token sensitivity tells a related but different story. The earlier work placed one Ultra around 6–7 billion raw first-party Grok-equivalent tokens, two Ultras around 13 billion, and Heavy around 14.2–16.5 billion. In that model, paying another $100 for two Ultras buys another $400 Other Models allowance while giving up roughly 1.2–3.5 billion raw Grok-equivalent tokens.

That trade is not absurd. It is a portfolio decision.

If I would voluntarily choose Opus or Sol for half my serious tasks, then Heavy’s unused Grok capacity expires worthless. If I would happily route almost everything to Grok, buying expensive third-party optionality is waste. Capacity has no utility independent of model preference.

The spreadsheet says Grok; the developer may say Opus.

The broader subscription benchmark

Every 2x, 4x, or 8x row below is a linear stack, not a vendor SKU. It assumes separate eligible subscriptions or accounts, linear quota scaling, no volume discount, and no shared enforcement ceiling.

Family Plan or scenario Type $ / mo Monthly FIE compute $ / 1M FIE Evidence confidence
xAI / Cursor X Premium Native $8 83.3M $0.096 Low
xAI / Cursor SuperGrok Native $30 250.0M $0.120 Medium
xAI / Cursor X Premium+ Native $40 250.0M $0.160 Low–medium
xAI / Cursor SuperGrok Plus Native $100 833.3M $0.120 Low
xAI / Cursor Cursor Ultra Native $200 2.122B $0.094 Medium–low
xAI / Cursor SuperGrok Heavy Native $300 4.622B $0.065 Medium–low
xAI / Cursor 2x Cursor Ultra Modeled $400 4.245B $0.094 Scenario
Anthropic Claude Pro Native $20 52.3M $0.382 Low
Anthropic Claude Max 5x Native $100 261.7M $0.382 Low–medium
Anthropic Claude Max 20x Native $200 550.4M $0.363 Low–medium
Anthropic 2x Claude Max 20x Modeled $400 1.101B $0.363 Scenario
Anthropic 8x Claude Max 20x Modeled $1,600 4.403B $0.363 Scenario
Google Google AI Plus Native $4.99 17.9M $0.279 Low–medium
Google Google AI Pro Native $19.99 35.8M $0.558 Low–medium
Google Google AI Ultra 5x Native $99.99 178.8M $0.559 Low–medium
Google Google AI Ultra 20x Native $199.99 715.0M $0.280 Low–medium
Google 2x Google AI Ultra 20x Modeled $399.98 1.430B $0.280 Scenario
Google 4x Google AI Ultra 20x Modeled $799.96 2.860B $0.280 Scenario
OpenAI ChatGPT Plus Native $20 29.5M $0.678 Low
OpenAI ChatGPT Pro 5x Native $100 147.4M $0.678 Low–medium
OpenAI ChatGPT Pro 20x Native $200 589.7M $0.339 Low–medium
OpenAI 2x ChatGPT Pro 20x Modeled $400 1.180B $0.339 Scenario
OpenAI 8x ChatGPT Pro 20x Modeled $1,600 4.718B $0.339 Scenario
Z.AI GLM Coding Lite Native $18 93.2M $0.193 Medium–high
Z.AI GLM Coding Pro Native $72 570.2M $0.126 Medium–high
Z.AI GLM Coding Max Native $160 1.329B $0.120 Medium–high
Z.AI 2x GLM Coding Max Modeled $320 2.658B $0.120 Scenario
Z.AI 4x GLM Coding Max Modeled $640 5.316B $0.120 Scenario

Many of the estimates are based on low-confidence reverse mathematics. While imprecise, they should be generally proximal to reality. A four-decimal calculation built on a low-confidence quota remains low-confidence.

The X Premium and Premium+ prices are retained as historical comparison points from an archived March 18, 2025 official X page because the live US price page was inaccessible; they are not verified as current 2026 prices.

What roughly $400 buys across providers

Figure 1 asks a practical question: what fits into roughly $400?

Figure 1: Subscription Efficiency Frontier comparing native and modeled AI subscriptions by monthly FIE compute and unit cost
Figure 1 · Subscription Efficiency Frontier

Heavy is the only native plan in the extreme upper-left. Cursor Ultra and the modeled two-Ultra stack also enter the high-value zone. Z.AI Max sits just below its value threshold; the modeled 2x Max crosses the capacity threshold at $320.

The Google and OpenAI curves contain genuine-looking pricing cliffs. Google’s $19.99 Pro and $99.99 Ultra 5x points are less efficient than both the $4.99 Plus and $199.99 Ultra 20x points under the literal public multipliers. OpenAI’s $200 Pro 20x tier is roughly twice as efficient as its $20 Plus and $100 Pro 5x tiers in this central series.

We did not smooth those discontinuities. If the plan schedule is nonlinear, the chart should look nonlinear.

The most uncomfortable point is the archived $8 X Premium row, marked with a dagger in the charts. It looks spectacular on unit value and carries low confidence, but its price is historical rather than verified current. It is a reminder that the left edge of this chart is not the recommendation. A tiny cheap quota can be efficient and still be useless to a heavy developer. Google’s AI Plus subscription shares similar characteristics.

What does Heavy-scale capacity cost elsewhere?

Figure 2 changes the question from “what can I buy for about $400?” to “what would it cost to approach this much compute?”

Figure 2: Heavy-scale capacity cost comparison showing spend required to approach SuperGrok Heavy absolute capacity
Figure 2 · Heavy-scale capacity cost comparison
  • SuperGrok Heavy: $300 → 4.622B FIE
  • Z.AI 4x Max: $640 → 5.316B FIE
  • Google 4x Ultra 20x: $799.96 → 2.860B FIE
  • Anthropic 8x Max 20x: $1,600 → 4.403B FIE
  • OpenAI 8x Pro 20x: $1,600 → 4.718B FIE

The words “would cost” matter. Eight consumer accounts are not an elegant purchasing strategy, and a vendor may restrict or aggregate them. These points isolate the price curve. They do not promise that stacking works operationally. It is particularly worth noting that fraud prevention systems may flag subscription stacking as potentially unwanted or malicious activity.

Statement explaining subscription enforcement for third-party clients and shared API traffic
Stacking and enforcement caveat

The Z.AI comparison

Z.AI is the real second story.

Unlike most vendors in this report, Z.AI publishes enough machinery to audit the quota model. Its current Coding Plan documentation lists:

  • 10,000 / 60,000 / 140,000 weekly credits for Lite / Pro / Max
  • GLM-5.3 credit multipliers of 6.9 fresh input, 1.7 cached input, and 24 output per 10,000 tokens
  • token ranges under a published 90.9% coding cache-hit assumption
  • a defined weekday peak window and 50% off-peak credit charging
  • multiple cache-hit sensitivity rows

Its GLM-5.3 documentation supports low, high, and max, defaults to max, and recommends max for complex coding. Its API list price is approximately $1.40 fresh input, $0.26 cached input, and $4.40 output per million.

That does not make the central point official. We still convert weekly ranges and a workload mix into a monthly FIE estimate. But the evidence chain is stronger. As a result, we assigned higher confidence, medium–high, to the Z.AI native points rather than the lower confidence applied to unpublished xAI, Cursor, Claude, Google, and OpenAI absolute limits.

The indexed official-checkout result used here is $18 / $72 / $160. During final verification, other client-rendered Z.AI surfaces exposed inconsistent promotional or product-namespace prices. We preserved the indexed result and its evidence limits, named the access date, and logged the discrepancy rather than guessing which unlabeled card would win tomorrow. While this is a set of assumptions, we are confident that it details the best-available numbers on Z.AI at the time of writing.

At 5.316B FIE for a modeled $640, 4x Z.AI Max subscriptions enter the gold region. That is a compelling second subscription-economics story. It is not Heavy’s unit economics: Z.AI remains near $0.120 per million FIE, about 1.85x Heavy’s roughly $0.065. Depending on the future of promotions and subscription subsidies, there is a significant likelihood of this relationship inverting in the future.

Why developer behavior changes the economics

The benchmark asks a deliberately narrow question: how much frontier compute does the subscription appear to contain at retail API-list equivalence?

A developer is buying on output: What can I build with this today?

Claude may be preferred for long-running code changes, editorial judgment, or document-heavy work. Codex may be preferred for its harness, local tooling, review behavior, or the ability to move between Sol, Terra, and Luna. Google packages storage and consumer surfaces. Cursor integrates routing, codebase context, review, and third-party models. Grok offers build, bot, media, and an unusually deep first-party allowance. Z.AI can plug into coding harnesses while preserving a different cost and openness posture.

Realized subscription value

realized subscription value
 = included capacity
 × probability you choose that model
 × probability you consume it before reset
 × workflow fit

Effectively: a plan with 5B estimated FIE can underperform a 500M plan if you avoid the model or cannot leverage its product surface. A smaller Claude or OpenAI subscription may be rational insurance for tasks where the harness or output is worth more than the token efficiency.

The subscription becomes a season pass; the objective is to ski until the resort regrets selling it. But a season pass to the wrong mountain is not a bargain. Like a time-share, it is a burden.

Stop hiding spend inside a ratio

Ratios are useful, but they compress two different realities into one number.

Figure 3: Subscription absolute scale map showing monthly FIE compute and monthly spend as separate series
Figure 3 · Subscription absolute scale map

The absolute map makes three things hard to miss:

  1. Heavy is the only native point above 4B FIE.
  2. The cheapest native paths to at least 1B in this model are Z.AI Max at $160 and Cursor Ultra at $200.
  3. The Anthropic and OpenAI capacity-match scenarios land near Heavy only after spend reaches $1,600.

This is why a ratio alone is dangerous. Google AI Plus looks efficient at $4.99, but its central capacity is 17.9M. Heavy looks efficient and enormous. Those are different economic objects.

What users can measure, and what they cannot

Public discussion contains valuable evidence, but much of it is contaminated by promotions, silent speed modes, client bugs, changing limits, and incomparable workloads.

The strongest examples we found illustrate the problem more than they solve it:

  • Cursor staff and forum replies confirm separate first-party and Other Models behavior, while declining to translate the first-party pool into a fixed token count. A temporary Grok 4.6 launch multiplier further distorted dashboard observations.
  • A detailed OpenAI Codex issue reported four user turns expanding into 1,031 Sol requests over a long tool loop, about 98% cache-read, while the weekly indicator moved from 59% to exhausted. That is powerful evidence that “tokens processed” and “subscription percentage” are not interchangeable.
  • Another Codex Pro 20x report later found a Fast/Priority configuration. OpenAI’s official documentation says GPT-5.6 Fast consumes credits at 2.5x Standard.
  • Anthropic and Google community reports include dramatic depletion incidents, but some look like metering bugs or account issues.

Token efficiency and actual throughput

Quota is only half of throughput. The other half is tokens spent by a model to finish a task.

Artificial Analysis’ public model pages report that Grok 4.6 and GPT-5.6 Sol reached the same Intelligence Index score with similar total output-token volume in its suite, while Opus 5 scored higher with more output and GLM-5.3 used substantially more output at a nearby score. The source audit also included Grok 4.5 as a historical comparator. Grok 4.5 is particularly interesting because of its extreme token efficiency to task completion and its extremely low direct API pricing.

That does not settle the matter. Token efficiency may amplify output-price dominance for both xAI and OpenAI. On a different workload, quality or shorter iteration count may invert that hypothesis.

Figure 4: Output tokens per Intelligence Index task for GPT-5.6 Sol, Gemini 3.1 Pro, Grok 4.5, Claude Opus 5, Grok 4.6, Gemini 3.7 Flash, and GLM-5.3
Figure 4 · Output tokens per Intelligence Index task

Whether current subscription economics are sustainable

The chart cannot answer the sustainability question by itself.

At Grok 4.6’s $2 per million standard fresh-input list price, the Heavy central estimate converts to:

4.622B FIE × $2/M = about $9,244 of API-list-equivalent compute
$9,244 / $300 = about 30.8x the subscription sticker

That is the defensible dramatic number, not the oft-touted $21,000+.

It is also easy to misconstrue. xAI is not necessarily spending $9,244 to service the subscriber. Cached inference, spare capacity, batching, internal hardware economics, model routing, and the strategic value of user acquisition can all make marginal cost radically different from the retail API menu. The calculation measures a retail API-list-equivalent value-to-price ratio, not COGS or economic viability.

Using open-weight as a control moves the bar closer to commodity compute, in some cases. The llm-d project served GLM-5.2-FP8 on 64 H200s using production Claude Code traces and calculated GPU-rental-only costs of $0.05–$0.08 per million input tokens and $1.92–$3.55 per million output tokens from measured throughput and a $2.71 per H200-hour assumption under sustained load.

An August 18 Vast H200 snapshot found roughly $0.29–$0.48 per million FIE: about $386–$644 for 1.329B FIE, or 2.4–4.0x the $160 Z.AI Max subscription, although still far below the $1.40 per million FIE API baseline. The calculation excludes storage, bandwidth, engineering, orchestration, and idle capacity.

Even a saturated rented-GPU deployment therefore fails to undercut Max, making the subscription discount even harder to dismiss.

We can responsibly say the subscription price is far below the benchmark’s retail API-list equivalence.

The potential impact of Grok 4.7

Post stating that Grok 4.7 should be ready in three to four weeks
Grok 4.7 timing signal

The economically interesting scenario is conditional.

If Grok 4.7 reaches Opus / Sol-class capability on the work developers care about, and if xAI retains roughly similar subscription economics and token efficiency, the behavioral reason to maintain large competing frontier subscriptions could shrink dramatically.

Every variable is meaningful. Capability can miss. Quotas can tighten. Pricing can change. A new model can consume more compute per task. Cursor, OpenAI, and Anthropic can retain workflow advantages that raw capability does not erase. Documents, code review, ecosystem integrations, safety behavior, and harness quality can justify a smaller subscription even when the token spreadsheet points elsewhere.

Grok may not win, but the pressure an economically credible Grok 4.7 would create represents a titanic shift in the model provider space.

The final buying framework

On this benchmark’s central estimates, SuperGrok Heavy is the strongest native subscription for high-volume frontier compute. It lands at roughly 4.622B monthly FIE, $0.065 per million FIE, and a $300 sticker. Z.AI is the most compelling second story: it publishes materially better quota mechanics and reaches 5.316B FIE at a modeled $640.

But the benchmark did not make the purchase decision automatic.

Two Cursor Ultras cost $100 more than Heavy and offer slightly less estimated Grok FIE. They also provide a second $400 Other Models pool and another integrated Cursor first-party pool. That optionality is valuable only if it matches actual routing behavior.

The right buying sequence is therefore:

  1. Identify the model and product surface that currently blocks your work.
  2. Estimate how much of each allowance you will consume before reset.
  3. Compare the marginal subscription against on-demand overflow.
  4. Use FIE economics to expose expensive capacity, not to pretend all tokens are equal.
  5. Re-run the measurement after major model, quota, speed-tier, or pricing changes.

The best AI subscription is not necessarily attached to the best AI model. It is attached to the model you will use, inside the workflow you will keep, at a quota you can exhaust, for a price below your next-best source of capacity.

That is a harder question than “which model is best?” It is also the one that drives the economics of your bill.

Methodology and assumptions

Benchmark date and scope

  • Price and official-product verification date: 2026-08-18
  • Scope: five comparison families, xAI / Cursor, Anthropic, Google, OpenAI, and Z.AI
  • Consumer or individual subscription prices are US web prices where available
  • Teams, enterprise, API-key-only, promotional annual equivalents, taxes, and app-store markups are excluded

Capacity evidence hierarchy

  1. Published vendor credits, token ranges, and price tables
  2. Official relative multipliers, windows, and pool descriptions
  3. Reproducible empirical usage measurements with cache/input/output mix
  4. High-signal user reports with screenshots or local logs
  5. Editorial extrapolation

Z.AI has the strongest native capacity evidence because it publishes weekly credits, cache assumptions, token ranges, and charging windows. Arithmetic confidence is high across all rows; realization confidence is not.

Monthly conversion

The central estimates preserve the calendar-month convention established in the prior benchmark. A four-week month versus 365.25 / 12 / 7 weeks changes a weekly extrapolation by roughly 8.7%, which is larger than display rounding. Do not read the third decimal place as measurement precision.

Native versus modeled

  • Filled chart points are currently sold retail plans.
  • Open chart points are linear 2x / 4x / 8x stacks.
  • A modeled stack assumes independent capacity and spend scale linearly.
  • “2x Cursor Ultra” means two subscriptions or accounts, not a $400 Cursor SKU.
  • Capacity-match stacks are analytical counterfactuals, not operational recommendations.

Confidence labels

  • Medium–high: published allowance mechanics plus normalization; current Z.AI native points
  • Medium / medium–low: useful empirical anchor but absolute vendor allowance remains unpublished
  • Low / low–medium: relative multiplier or anecdotal anchor with substantial extrapolation
  • Scenario: arithmetic is deterministic if the parent point and linear-stacking assumption hold; operational feasibility is unverified

Exclusions and rejected claims

  • The distrusted outlier series from the earlier draft is excluded completely.
  • No claim that a plan is “unlimited.”
  • No claim that FIE equals transcript tokens, quality-adjusted work, serving cost, or vendor subsidy.
  • No claim that retail H200 rental equals a hyperscale provider’s marginal hardware cost.
  • No claim that Heavy includes Cursor Ultra; current official pages show separate plans.
  • No unverified screenshot or inaccessible social post is presented as captured evidence.
  • No proprietary Artificial Analysis chart is copied.

This article is part of Open Security Research: technical, evidence-driven work on the operational realities behind fast-moving security and AI markets.

Talk through what you are seeing.

Operators can help you decide what to test next.