Qaf · AI cost · October 2026

AI costs $54k a month.
In January it doubles.

Tier 2 for weak-payment markets saves 18%. Making every question leaner saves about half.

AI cost today
$54.5k
per month · 4¢ a question
From Jan 2027
$102k
Gemini Flash intro price ends
Revenue
$7.8k
MRR · AI costs 7× this

The question

Tier 2 saves $10k now, $20k in 2027

Free Android users in 14 countries where almost nobody can pay.

Saving today
−$9.6k
$54.5k → $44.9k · −18%
Saving in 2027
−$19.7k
$102k → $83k · −19%
Revenue at risk
≈ $0
All payers keep premium
YemenIranAfghanistan AlgeriaEthiopiaSyriaMauritaniaPalestineLibyaSudanNigerTajikistanMaliChad Egypt — later

Why it stays high

Free users are 89% of the bill

The 14 countries are only a fifth of it. Most free usage is in markets that pay.

Paying, trials, gifts
Tier 2 as recommended
Signed-out guests, apps <1.3.0
Free, outside top-18 markets
Free, in markets making 90% of revenue

How to cut further

Six packages, from safe to aggressive

Monthly cost. Each step includes the ones above it. Ranges are conservative → optimistic.

No change
Baseline
Now
$54.5k
2027
$102k
Tier 2 · 14 countries
The plan above
Now
$44.9k
2027
$82.7k
Quick wins START HERE
+ guests and old apps on tier 2, leaner prompts, cheaper reranker
Now
$28–31k
2027
$52–57k
Deeper trimming
+ fewer passages and searches, shorter answers — after quality tests
Now
$15–21k
2027
$27–38k
Payer-market focus
+ tier 2 for free users outside the 18 markets that make 90% of revenue
Now
$13–18k
2027
$21–30k
Aggressive
+ premium for a free user's first 50 questions only, Google's half-price tier, 10 welcome questions
Now
$8–12k
2027
$12–19k
All free users on tier 2
Only payers get premium
Now
$7–8k
2027
$8–11k
The biggest lever isn't tier 2. Leaner questions alone — no one downgraded — cut the bill roughly in half: $20–29k now, $37–54k in 2027.

The levers

What each one is worth

Monthly saving today, on its own. 2027 is roughly double for model-cost levers.

LeverSavesRiskEffort
Leaner prompts: short citations, trim history, dedupe passages$9–13kLow1 week
Fewer passages and searches, shorter answers+$14–20kMedium2–4 weeks
Tier 2: 14 countries + guests + old apps$15–18kLow2 days + eval
Tier 2: free users outside top-18 markets+$8–9kMediumLow
Tier 2: every free user$42–50kHigh2 days + eval
Google half-price Flex tier for free traffic$13–20kLatencyLow
Google savings plan, credits, price extension$5–10k+UncertainLow
Force update to app 1.3.0+$2kLowHours
Welcome questions 20 → 10$3–4kConversionA/B test
Sign-in after 3 guest questions$2–4kFunnel1–2 days
Cheaper reranker, titles off GPT$2–3kLowTrivial

Recommendation

Do the quick wins before January

Cuts the 2027 bill from $102k to about $55k.

  1. October
    Leaner prompts, cheaper reranker, force app 1.3.0. No quality trade-off. Ask Google for credits and a price extension.
  2. November
    Pick a tier-2 model that passes a quality test in Arabic, Farsi, Dari and Pashto. Then launch it for the 14 countries, guests and old apps — Yemen and Iran first, as an A/B test.
  3. Q1 2027
    Test deeper trimming and Google's Flex tier. Decide on wider tier 2 once retention and conversion data are in.

Fine print

Tier 2 won't really cost 5%

In our model benchmark, a tuned DeepSeek cost 16–30% of a current question; GLM lost 88% of head-to-heads. With a realistic model, tier-2 savings shrink 20–25%. No candidate has passed quality yet.

Where the numbers come from

30 days, 9 Sep – 8 Oct 2026. Cost from Google Cloud token counts and our Cloudflare AI gateway at list price (they agree within 5%). Langfuse is not used for totals — it records almost only Android and misses thinking tokens. Usage and country from Mixpanel; payers from Autumn and Stripe; logic from the product code.

Assumptions
  • Savings use the conservative reading: only the model gets 95% cheaper, search stays the same.
  • 2027 figures assume today's volume and Gemini's announced price ($1.50 / $7.50 per million tokens).
  • List prices. Google or Azure credits may make cash cost lower.
  • Earlier versions said $59.9k: reranking was double-counted.