Hey, Paweł here. I’ve been sick, so please expect the next standard post on Sunday. Today, I have a few interesting findings to share.
Let’s dive in.
1. My Benchmark: GPT-6.1 Sol, Sonnet 5.5, and Opus 5.5
First off, new models. Just recently Opus 5.5, Sonnet 5.5, and GPT-6.1 Sol dropped.
I don’t trust the benchmarks, so I tested those models on real work: 2 repos, 105 bugs frontier models missed at the beginning of 2026, find and fix what you can.
The results:

What drew my attention:
Opus 5.5 is significantly stronger than Opus 5 and nearly matches Fable 5.1 (41.7 vs 43) at two-thirds of the cost ($58.53 vs $87.18). I recommend you stop using Fable 5.1 unless you have an unlimited budget.
Sonnet 5.5 is 2x cheaper than Opus 5.5 on input and output tokens ($2/$10 vs $4/$20 per MTok). But in long agentic sessions, most of your costs are cache reads, and those are priced identically: $0.20/MTok.
GPT-6.1 Sol is slightly stronger while being over 10x cheaper than GPT-5.6 Sol on complex tasks.
GPT-6.1 Sol’s input and output tokens cost half as much
Its cache reads are 4x cheaper (framed as a 95% discount by OpenAI). This may be temporary. Soon, we may get back to the standard “90% discount.”
The 10x difference is best explained by the number of turns (485 for GPT-5.6 Sol xhigh vs 200 for GPT-6.1 xhigh on the same task).
On the release day (yesterday), GPT-6.1 Sol was almost 2x slower than Opus 5.5. According to Tibo from OpenAI, it should become almost 2x faster than yesterday over the next hours.
Sonnet 5.5 (max) surprise
Sonnet 5.5 (max) won the benchmark with 51.3/105 by running 1,497 turns over 287 min compared to 270 turns over 90 min of GPT-6 Astra. It’s not smarter, just more motivated and able to think for longer.
I’m not emphasizing this score, as it was too slow and too expensive to use even in agentic coding. You won’t see those gains in your standard tasks.
2. The Real Cost of The AI Subscription
We’ve all been hearing that subscriptions are cheaper than APIs, but I’d never seen the exact number.
So yesterday, I took several popular plans and burned a slice of their weekly allowance on purpose. I logged every model call and priced those calls at the vendor’s public API list price.
The results:
While I couldn't test all subscription levels, according to OpenAI their monthly allowance is now directly proportional to the subscription price:
Plus ($20/mo) = the reference OpenAI monthly allowance
Pro 100 ($100/mo) = 5x monthly allowance - tested
Pro 200 ($200/mo) = 10x monthly allowance (just downgraded from 20x)
For Claude, their Max 20x is subsidized twice as much:
Pro ($20/mo) = the reference Claude monthly allowance
Max 5x ($100/mo) = up to 5x monthly allowance - we don’t know the exact number, multipliers apply to the 5-hour window
Max 20x ($200/mo) = up to 20x monthly allowance, though we don’t know the exact number, multipliers apply to the 5-hour window - tested
Note: I’ve removed the Anthropic Max 5x row after being corrected on X. The only plan we can reason about without a measurement is Max 20x ($200/mo). I’m going to test lower Anthropic plans separately.
According to Meta, for Muse Code:
Everyday Usage ($5/mo) = the reference Muse monthly allowance
High Usage ($15/mo) = 5x monthly allowance (not 3x) - tested
Power Usage ($50/mo) = 20x monthly allowance (not 10x)
What does all this mean?
If you want the cheapest western subscription, it’s Grok in the SuperGrok subscription.
Just after Grok, Muse Code if you are willing to share your data with Meta (Contributor).
Just after Muse Code + Contributor, OpenAI:
GPT-6.1 Sol (medium) is 13x cheaper than Grok 4.7 (xhigh) while getting similar scores. SuperGrok subscription multiplier is 18.5x bigger than OpenAI’s.
Despite Anthropic’s 2-4x bigger multipliers (depending on the plan), GPT-6.1 Sol will still cost you less than Opus 5.5. A linear scale from my tests:
How to use Grok Build or Muse Code without the CLI?
Neither Grok Build nor Muse Code has a native GUI. As many of the readers know, I created one. It currently supports Grok Build, Muse Code, Codex, and Claude Code:
Extension for VS Code (32K installs)
Extension for Cursor, Antigravity (134K downloads)
Desktop app for Windows, macOS, and Linux: https://afkpilot.com/desktop
Unlike OpenAI or Anthropic, small Grok and Muse plans ($30 SuperGrok, $15 Muse) are enough to do real agentic work. SuperGrok $30/mo also gives you access to the Grok Bot.
We will discuss Bots and Dots another time ;)
Side note: Here, you can see the recent episode with Moe Ali from Product Faculty, in which we discussed hardening production apps without coding.
The ultimate From Prototype to Production guide:
Thanks for Reading The Product Compass
It’s amazing to learn and grow together.
Have an amazing rest of the week,
Paweł





