Qaf · internal analysis · data window 9 Sep – 8 Oct 2026 (30 UTC days) · prepared 9 Oct 2026

Tier 2 cost analysis

What happens to Qaf's monthly AI bill when a model at about 5% of today's per-question cost answers non-paying users in the countries where Qaf earns almost nothing, and which country × platform segments should get it.

Headline answer

Today Qaf spends about $59.9k a month on AI. Moving non-paying Android users in the 14 recommended countries (scenario C) to tier 2 brings that to $48.2k–$50.1k: a saving of $9.7k–$11.6k a month (16–19%). In year 2, when Gemini Flash prices double, the saving is about $19k–$21k a month.

AI cost today
$59.9k
per month · $58.2k answering questions + $1.7k other AI · 4.6¢ per question · 1.25M questions in 30 days
After tier 2 · scenario C
$50.1k
−$9.7k · −16.3%
Interpretation B. Interpretation A: $48.2k (−$11.6k, −19.4%). Low case: −$8.7k to −$10.6k.
Year 2 · 2027 Gemini prices
$89.5k
−$19.4k · −17.9%
from $108.9k at today's volume. Interpretation A: $87.6k (−$21.3k, −19.6%).
Year 2 · month 12, base growth
$144.5k
−$39.1k · −21.3%
from $183.6k. Interpretation A: $140.7k (−$42.9k). High growth: $305.2k → $229.4k.
Two readings of "5% of the cost per question". Interpretation A: a tier-2 question costs 5% of everything a question costs today ($46 per 1,000 questions → $2 per 1,000). Interpretation B (conservative, used for headlines): only the language model gets 95% cheaper; retrieval (search reranking, embeddings), which is about 16% of a question's cost today, stays the same. Under B a tier-2 question costs about 20% of today's, not 5%, unless retrieval is cheapened too (see retrieval).

Revenue at risk is about $0 directly: every payer, trial, Play grace-period, gift, comp and lifetime user keeps the premium model. The only cost is possible lost future conversion, at most about $1.3k to $2.2k of MRR after a year even if Android billing were fixed, against a saving of $9.7k+ every month.

Update: corrected for a rerank double count, the bill is $54.5k today and $102.3k at 2027 prices. Scenario C alone leaves most of it; see why it is still high and how to reach $12.6k–$21.8k.

Why it's still high, and how to cut further

Corrected, Qaf's AI bill is $54.5k a month today and $102.3k at 2027 prices. The earlier $59.9k / $108.9k counted every rerank call twice. Scenario C only reaches 21% of it. The cost per question has to fall for everyone, and tier 2 then has to reach more of the free traffic.

Where the $54.5k goes today

By who asks
  • Paid, trial, grace, comp, gift · $5.6k (10%)
  • Scenario C (14 weak-payment countries, Android, free) · $11.1k (20%)
  • Guests + unmetered legacy apps (outside C) · $6.0k (11%)
  • Other free users outside the 18 payer markets · $8.7k (16%)
  • Free users in the 18 payer markets · $21.3k (39%)
  • Non-chat AI (agents, TTS) · $1.7k (3%)
By what it pays for
  • Answer generation (Gemini 3.7 Flash) · $47.9k (88%)
  • Rerank (Cohere) · $4.6k (8%)
  • Titles · $0.2k (0%)
  • Non-chat AI · $1.7k (3%)

1 · Generation is the bill, and it doubles

Answer generation on Gemini 3.7 Flash is 88% of cost ($47.9k). Google's intro price ends on 31 Dec 2026, which takes this line to about $95.8k. Rerank is only ~$4.6k.

2 · Each question is heavy

About 47k prompt tokens and 2.5k output tokens over ~3 Gemini calls. 66% of the prompt is ~42 retrieved passages (20 per search, sent with diacritics, footnotes and long IDs); 23% is chat history, a third of it old citation tags. Only ~21% is cached. That is $0.042 per question today and $0.079 in 2027.

3 · Free users carry 89%, and it is not abuse

~90k people asked questions in the window; the median free user asked 5 a month, the free tier is already 20 welcome questions plus 3 a day, and nobody averages 50 a day. Most free cost sits in markets that pay: Saudi Arabia alone is ~$12.7k a month and 41% of MRR. Two leaks: apps below 1.3.0 skip the free cap (~$5.3k), and guests cost ~$5.2k (up to ~$8.7k counting signed-out questions of people who sign in later).

4 · Payers do not cover their own usage in 2027

A payer asks ~162 questions a month: $6.73 of AI today against $9.17–11.13 after store fees. Break-even needs 11–18% of users paying; the actual rate is 0.71%. At 2027 prices a payer costs $12.84, more than they pay after fees, so no conversion rate breaks even until the cost per question falls.

Options (each alone, against the corrected baseline)

Monthly savings. Tier-2 ranges run from a realistic tier-2 model (tuned DeepSeek from the bench) to the 5%-of-generation target. Options overlap, so they cannot be added up; the packages below stack them properly.

OptionLeverSaving todaySaving 2027 pricesRiskEffortConfidence
C0 Fix the rerank double count in the cost model
Langfuse logs every rerank call twice; counted once, rerank is ~$4.6k, not $9.1k.
measurement— (baseline $59.9k → $54.5k)— ($108.9k → $102.3k)none: it is an overstatement, not a savingnonehigh
E-A Efficiency package A (core)
Short citation refs instead of long IDs, strip old citations from history, de-duplicate passages, cap history at 16k tokens, minimal thinking on the policy check.
efficiency$8.4–11.7k$16.9–23.3klow: citation rendering, expand-chunk mapping, stream resume~1–1.5 weeksmedium-high
E-A+ Efficiency A, gated items
Policy check on Flash-Lite; step cap 6. Each needs a targeted check (missed prohibited queries; truncated 3+-step answers).
efficiency+$0.9–1.4k (A total $9.3–13.1k)+$2.0–2.9k (A total $18.9–26.2k)medium: safety false negatives, answer truncationdays + checksmedium
E-B Efficiency package B
Top-12 instead of 20 passages, 1–2 searches instead of 2–4, drop diacritics/footnotes from passages, router in the policy call, shorter answers, cache work, response thinking low.
efficiency / retrievalA+B $23.5–32.6k; top-12 alone $7.9–9.4k (lower if passages are a smaller share of the prompt: $4.8–7.2k)A+B $46.3–63.6k; top-12 $15.7–18.9kmedium: completeness, coverage of differing madhab views; thinking-low was reverted once2–4 weeks incl. evalsmedium-low
T2-C Tier 2: scenario C
Non-paying Android users in the 14 weak-payment countries.
audience$7.7–9.6k$17.8–19.7kvery low (~0.5% of converters)1–2 dev days + evalmedium-high
T2-S1 Tier 2: C + all guests + legacy apps
Adds signed-out users and signed-in users on unmetered app versions below 1.3.0.
audience$11.9–14.7k$27.3–30.3klow; about a fifth of payers asked as a guest first+ hoursmedium-high
T2-S2 Tier 2: S1 + free users outside the 18 payer markets
The 18 countries that produce 90% of MRR keep premium for free users.
audience$18.0–22.3k$41.3–45.7k9.6% of MRR and 22% of trials start outside the list (~$0.4–3.1k first-year MRR); VPN, fairnesslow–mediummedium
T2-S4 Tier 2: S2 + payer markets after 50 lifetime questions
Free users in payer markets get premium for their first 50 lifetime questions.
audience$22.7–28.1k$52.1–57.7k38–50% of converters pass 50 questions before paying; $1.8–7k first-year MRRmedium (lifetime counter)medium
T2-ALL Tier 2: every non-paying user
Only paid, trial, grace, gift, comp and lifetime users keep premium.
audience$32.8–40.6k$75.3–83.4khigh: every converter's first impression; $4.7–14k first-year MRR, more on lifetime value1–2 dev days + evalhigh on cost, low on conversion
T2-M Which tier-2 model
Tuned DeepSeek ≈30% of a question today / 16% in 2027 (searches 2.4× more). Gemini 3.1 Flash-Lite ≈35% / 17.5% of Flash generation. Claude Haiku 5.5 ≈25–30% / 13–15% once its 5× price on prompts over 100k tokens is counted. Only GLM no-think reaches ~5%, and it lost 88% of comparisons.
vendordecides where each tier-2 range lands: low end ≈ DeepSeek, high end needs ~5%sameno candidate passes quality yet (DeepSeek: serious errors in 66% of answers vs 33%)2–3 days benchmedium
L1 Force update to app 1.3.0+
Meters the legacy apps that skip the free cap today.
limits$1.8–2.1k (up to ~$3.4k with drop-off); overlaps S1$3.4–4.0klocks out users who cannot update (Iran sideloads); needs an APK/web fallbackhourshigh
L3 Fewer starter credits (A/B)
Starter 20 → 10; or 2/day + starter 10.
limits$3.1–4.0k; $5.7–7.4k (upper bounds)$5.8–7.6k; $11.0–14.3kwalls precede 31.6% of first paid starts; causal effect unknown<1 day + 1–2 week testmedium on cost, low on revenue
L4 Sign-in after 3 guest questions; 50/day ceiling; meter reference routeslimits$1.7–3.3k; $0.5–0.7k; $0.1–0.3k$3.2–6.3k; $0.8–1.4k; $0.2–0.5ktop-of-funnel loss (sign-in wall)<1–2 days eachlow–medium
R1 Cheaper reranker
rerank-v4.0-fast instead of pro; or a token-priced reranker.
retrieval$0.9k (token-priced $2.4–3.8k)samesmall relevance loss; new vendor for token-pricedtrivial / mediumhigh on price, unknown quality
V1 Vertex Flex PayGo for free traffic
50% off the same model, with Standard fallback.
vendor$12.8–20.3k$25.6–40.6klatency and throttling on streaming chat; test on 1–5% firstlow–mediummedium on price, low on latency
V2 Savings plan, price negotiation, startup credits
Savings plan 10–20%; ask Google to extend the intro price; startup credits up to $350k one-off over 2 years (≤$250k in year 1, ≈3.6 months of 2027 generation; VC-funded only).
vendorplan $4.8–9.6kplan $9.6–19.2k; negotiation $0–48klock-in while prices fall; eligibility uncertainlowlow
V5 Titles and non-chat agents off gpt-5.6-terravendor$1.2–1.3ksamelowlowmedium
X Ads / rewarded ads
One rewarded video earns far less than a question costs outside the Gulf.
other~$0 net~$0 netbrand risk in a religious app1–2 weekslow

Stacked packages: monthly AI cost

The higher number uses a realistic tier-2 model and each option's low saving; the lower number assumes tier 2 at ~5% of generation and each option's high saving.

PackageTodaySaving2027 pricesSavingRevenue risk
P0 Baseline (corrected)
No change.
$54.5k–$102.3k––
P1 Scenario C only
Tier 2 for non-paying Android users in the 14 weak-payment countries.
$44.9k–$46.7k14–18%$82.7k–$84.6k17–19%very low (~0.5% of converters)
P2 Quick wins
Tier 2 for C + all guests + unmetered legacy (<1.3.0) apps; efficiency package A; rerank-fast; titles/agents off terra.
$28.4k–$33.7k38–48%$51.8k–$59.7k42–49%low
P3 P2 + full efficiency
P2 + efficiency B after evals (top-12 passages, fewer searches, payload slimming, router, shorter answers, cache work).
$15.1k–$24.2k55–72%$26.6k–$41.3k60–74%low on conversion; medium on answer quality
P4 Payer-market focus
P3 + tier 2 for every free user outside the 18 countries that make 90% of MRR.
$12.6k–$21.8k60–77%$21.3k–$34.7k66–79%low–medium: 9.6% of MRR and 22% of trials start outside the list
P5 Aggressive
P4 + payer-market free users premium only for their first 50 lifetime questions; Vertex Flex for remaining free traffic; starter credits 20 → 10.
$8.2k–$17.4k68–85%$12.2k–$24.8k76–88%medium: first-50 rule touches 38–50% of converters, mostly in Saudi Arabia
P6 All non-paying users on tier 2
Everyone without a paid, trial, grace, gift or comp plan on tier 2, plus full efficiency.
$6.6k–$15.8k71–88%$8.4k–$18.8k82–92%high: every converter starts on tier 2 ($4.7–14k first-year MRR at risk)
R Reference: no tier 2
Efficiency A+B, rerank-fast and the terra move only; everyone keeps the premium model.
$20.0k–$29.1k47–63%$36.9k–$54.1k47–64%low on conversion; medium on answer quality
P0 · Baseline (corrected)
$54.5k
$102.3k
P1 · Scenario C only
$44.9k–$46.7k
$82.7k–$84.6k
P2 · Quick wins
$28.4k–$33.7k
$51.8k–$59.7k
P3 · P2 + full efficiency
$15.1k–$24.2k
$26.6k–$41.3k
P4 · Payer-market focus
$12.6k–$21.8k
$21.3k–$34.7k
P5 · Aggressive
$8.2k–$17.4k
$12.2k–$24.8k
P6 · All non-paying users on tier 2
$6.6k–$15.8k
$8.4k–$18.8k
R · Reference: no tier 2
$20.0k–$29.1k
$36.9k–$54.1k

P6 is cheaper than P4 even with a realistic tier 2: about $6.0k a month less today and $16.0k less at 2027 prices. It is held back for revenue risk, not because the gain is small.

Recommended sequence: aim for P4 by 31 December

Target $12.6k–$21.8k a month today and $21.3k–$34.7k at 2027 prices, against $54.5k / $102.3k. Treat Flex and starter 10 as experiments.

  1. October, weeks 1–2. Efficiency A core (short citation refs, strip old citations, passage dedupe, 16k history cap, minimal policy thinking). Switch to rerank-v4.0-fast. Move titles and non-chat agents off terra. Force-update apps to 1.3.0 with an APK/web fallback. Fix the rerank double count in reporting. Apply for Google startup credits and open a price conversation with Google. Start a 1–5% Flex latency test on free traffic. In parallel, check the Flash-Lite policy check for missed prohibited queries and the step cap for truncated answers.
  2. Weeks 2–4. Bench Gemini 3.1 Flash-Lite and Claude Haiku 5.5 as tier 2 (at most 2 searches, citation resolver, 16k history cap). No candidate passes today. When one does, ship tier 2 to C + guests + legacy apps: P2, $28.4k–$33.7k today, $51.8k–$59.7k in 2027.
  3. November. Evaluate and ship efficiency B (top-12 passages, fewer searches, payload slimming, router, shorter answers, cache work): P3, $15.1k–$24.2k today, $26.6k–$41.3k in 2027.
  4. December, before prices double. Extend tier 2 to free users outside the 18 payer markets (P4). A/B test starter 10 and the first-50 rule in payer markets; both land mostly on Saudi free users. Keep Flex only if time to first token holds.
  5. Hold P6 (all non-paying users on tier 2) until a conversion A/B shows tier 2 does not hurt sign-ups to Pro.
What could move these numbers. No tier-2 model passes quality yet. Efficiency B factors need evals. iOS and web token use is not observed (Langfuse covers Android). Rerank may bill 2 units for long passages (unchecked). Cloud credits may make cash savings smaller than list-price savings. Revenue-at-risk figures are first-year, order-of-magnitude estimates.

Recommendation

Adopt scenario C: tier 2 for non-paying Android users in 14 countries

YemenIranAfghanistanAlgeriaEthiopiaMauritaniaSyriaPalestineLibyaNigerSudanTajikistanMaliChad

These are exactly the Android segments an objective rule flags: at least 50 active users, conversion and revenue per user both under 25% of Qaf's average, and payment feasibility of "none" or "poor" (sanctions, no card rails, or weak Play billing). Together they ask 22% of all questions and cost 21% of the AI bill, but bring in almost no revenue: across all their Android users there are fewer than 5 payers.

  • Yemen, Iran and Afghanistan (the founder's list) are the core: $7.3k–$8.7k of the saving. Iran and Afghanistan have no working payment rails at all.
  • Algeria, Ethiopia, Mauritania, Syria, Palestine, Libya, Niger, Sudan, Tajikistan, Mali, Chad add another ~$2.5k. Algeria and Libya have zero payers and zero trials among 1,474 and 403 Android users.

Extensions worth taking

  • Scenario E: also put iOS and desktop web on tier 2 in Iran, Afghanistan, Syria and Sudan, where no payment is possible. Adds ~$0.5k/month and closes the "switch to iOS or desktop" loophole. Require a timezone or IP in the country (Farsi locale alone does not qualify).
  • Cheaper retrieval for tier 2: with no rerank, C saves $11.7k/month instead of $9.7k; with rerank-v4.0-fast, $10.1k.

Decide separately or exclude

  • conditional Egypt Android (+$2.1k/month): Egypt is payable. iOS and desktop convert at about the average, and Android has a pipeline of trials and grace-period subscribers. Its weak Android conversion looks like Qaf's own Play billing problems. Fix billing first, or A/B test tier 2 with a "Pro = premium model" upsell measured on trial starts.
  • exclude India, Bangladesh, Morocco, Indonesia, Iraq, Pakistan, Jordan: weak on the data but with good or fair payment rails and live trials. Revisit after measuring C (scenario D).
  • exclude Somalia: not weak on the data (above-average revenue per user) and tiny (~$0.2k/month).
  • exclude All signed-out guests everywhere: saves $4.5k–$5.4k/month but gives every new user a cheaper first impression. Not before the quality evaluation.

How to target it in production

  1. Country = device timezone OR Cloudflare CF-IPCountry. GeoIP alone misses 97.7% of Iranian Android users, who are on VPNs (they appear as Germany or the US); the x-qaf-client-timezone header the apps already send catches 96.6% of them. This rule saves $10.1k (B) for C, within ~4% of the modelled figure; GeoIP alone saves only $7.3k.
  2. GeoIP only for Palestine, Syria and Sudan. Syrian and Sudanese timezones are shared by diaspora users in Turkey, Brazil, the UK and the Netherlands (payable countries); Palestine shares Asia/Jerusalem.
  3. Android app first (header x-qaf-client-platform=android). Android web needs new user-agent parsing; it is a phase-2 item worth about $0.3–0.4k/month.
  4. Keep premium for anyone with a non-free plan in status active, scheduled, trialing or past_due, plus gifts, comps and lifetime. Do not reuse hasPaidAccess as it is: it excludes trialing and past_due, so trial and Play grace-period users would be downgraded.
  5. Gate the launch on quality. Run an offline evaluation of tier 2 against the premium model on real questions in Arabic, Farsi/Dari, Pashto, Amharic, Somali and French, scoring citation and hadith accuracy, rulings and the policy check. Label the model in the app and show the Pro upsell.
  6. Roll out to Yemen and Iran first as an A/B test. Measure retention (D7/D30), questions per user, answer-quality complaints and feedback-portal reports before widening to all 14.

Saving by country (non-paying Android users, interpretation B, today's prices)

Yemen C
−$2,962
Iran C
−$2,440
Egypt C+EG
−$2,051
Afghanistan C
−$1,891
India not C
−$872
Morocco not C
−$817
Indonesia not C
−$679
Algeria C
−$662
Iraq not C
−$607
Bangladesh not C
−$598
Ethiopia C
−$482
Pakistan not C
−$381
Uzbekistan not C
−$329
Syria C
−$315
Mauritania C
−$313
Jordan not C
−$275
Nigeria not C
−$227
Palestine C
−$222
Libya C
−$213
Somalia not C
−$197
Tunisia not C
−$168
Senegal not C
−$137
Country (Android)DecisionPayment railsActive usersPayersTrials + graceQuestions / moAI cost / moSaving / mo (B)Saving / mo 2027 (B)
Yemen YE C Poor 4,164 <5<5 79,118$3,687 −$2,962−$5,909
Iran IR C None 3,170 00 70,467$3,075 −$2,440−$4,867
Egypt EG conditional Good 3,367 <520 51,099$2,527 −$2,051−$4,092
Afghanistan AF C None 2,524 0<5 54,723$2,384 −$1,891−$3,772
India IN exclude Good 1,584 05 23,504$1,087 −$872−$1,740
Morocco MA exclude Good 1,577 <56 22,542$1,022 −$817−$1,630
Indonesia ID exclude Good 1,183 513 17,826$843 −$679−$1,354
Algeria DZ C Poor 1,546 0<5 19,958$840 −$662−$1,321
Iraq IQ exclude Fair 1,230 <511 18,016$769 −$607−$1,212
Bangladesh BD exclude Fair 1,562 0<5 18,035$759 −$598−$1,193
Ethiopia ET C Poor 668 0<5 14,592$612 −$482−$961
Pakistan PK exclude Fair 678 <5<5 9,229$467 −$381−$760
Syria SY C None 573 00 9,227$398 −$315−$628
Mauritania MR C Poor 287 0<5 8,999$394 −$313−$625
Palestine PS C Poor 421 0<5 6,817$283 −$222−$443
Libya LY C Poor 425 00 6,295$269 −$213−$424
Somalia SO exclude Poor 178 <50 2,912$228 −$197−$393
Sudan SD C None 150 0<5 1,790$80 −$64−$127
Niger NE C Poor 70 00 1,542$70 −$56−$112
Tajikistan TJ C Poor 89 00 1,381$56 −$43−$87
Chad TD C Poor 69 00 1,211$54 −$43−$85
Mali ML C Poor 73 00 1,409$55 −$42−$85

Questions and cost include Android users seen only in Langfuse (mostly signed-out guests). Payer and trial counts under 5 are shown as "<5".

Scenarios

Every scenario moves only non-paying users (payers, trials, grace, gifts, comps and lifetime keep premium). Bars show the monthly saving today: dark = interpretation B, light = interpretation A, tick = low case (B).

A Founder's list
−$7.3k / −$8.7k
B A + Egypt
−$9.3k / −$11.1k
C Recommended core
−$9.7k / −$11.6k
C+EG Conditional add-on
−$11.8k / −$14.0k
E C + all platforms in no-payment countries
−$10.3k / −$12.2k
D Broad: 28 countries
−$17.2k / −$20.4k
S1 Sensitivity: C + India, Bangladesh
−$11.2k / −$13.4k
S2 Sensitivity: all signed-out guests
−$4.5k / −$5.4k
ScenarioUsers / moShare of questionsToday, B: afterToday, A: afterLow case (B)2027 prices, BMonth 12, base growth (B)Deployable rule (B)GeoIP only (B)"Everyone" variant: payers / MRR exposedConversion risk after 12 mo (MRR)
A Founder's list
Android app + web, non-paying users in YE, AF, IR
12,74716.1% $52.6k
−$7.3k · −12.2%
$51.2k
−$8.7k · −14.5%
−$6.5k $94.4k
−$14.5k (A: −$15.9k)
$152.0k
from $181.3k · −$29.3k
−$7.3k −$4.8k <5 / <$100 $180–$300
billing fixed: $697–$1.2k
B A + Egypt
A plus Egypt Android
16,72820.2% $50.5k
−$9.3k · −15.6%
$48.8k
−$11.1k · −18.5%
−$8.3k $90.3k
−$18.6k (A: −$20.4k)
$145.7k
from $183.2k · −$37.5k
−$9.4k −$6.9k <5 / <$100 $587–$978
billing fixed: $1.3k–$2.1k
C Recommended core
Android app + web, non-paying users in the 14 countries flagged by the weak-monetization rule
18,38921.9% $50.1k
−$9.7k · −16.3%
$48.2k
−$11.6k · −19.4%
−$8.7k $89.5k
−$19.4k (A: −$21.3k)
$144.5k
from $183.6k · −$39.1k
−$10.1k −$7.3k <5 / <$100 $360–$600
billing fixed: $1.3k–$2.2k
C+EG Conditional add-on
C plus Egypt Android, after a Play billing fix or an A/B test
22,37025.9% $48.1k
−$11.8k · −19.7%
$45.8k
−$14.0k · −23.5%
−$10.5k $85.4k
−$23.5k (A: −$25.8k)
$138.2k
from $185.5k · −$47.4k
−$12.2k −$9.3k <5 / <$100 $691–$1.2k
billing fixed: $1.9k–$3.1k
E C + all platforms in no-payment countries
C plus iOS and desktop web in IR, AF, SY, SD (timezone or IP in-country required)
19,18223.0% $49.6k
−$10.3k · −17.1%
$47.6k
−$12.2k · −20.5%
−$9.3k $88.4k
−$20.5k (A: −$22.5k)
$142.9k
from $184.1k · −$41.2k
−$10.8k −$7.7k <5 / <$100 $401–$668
billing fixed: $1.3k–$2.2k
D Broad: 28 countries
Android in all 28 low/middle-income candidate countries (includes payable EG, MA, ID, IQ, PK, JO)
35,29637.7% $42.7k
−$17.2k · −28.7%
$39.4k
−$20.4k · −34.2%
−$15.3k $74.6k
−$34.3k (A: −$37.5k)
$121.6k
from $190.6k · −$69.0k
−$17.7k −$14.7k 22 / $240 $1.9k–$3.2k
billing fixed: $3.7k–$6.1k
S1 Sensitivity: C + India, Bangladesh
C plus India and Bangladesh Android (no payers yet, but payable markets)
22,08025.2% $48.6k
−$11.2k · −18.7%
$46.5k
−$13.4k · −22.4%
−$10.0k $86.5k
−$22.4k (A: −$24.6k)
$140.0k
from $185.0k · −$45.0k
−$11.6k −$8.7k <5 / <$100 $648–$1.1k
billing fixed: $1.8k–$3.1k
S2 Sensitivity: all signed-out guests
Every signed-out guest, every country and platform
24,62010.3% $55.3k
−$4.5k · −7.5%
$54.5k
−$5.4k · −9.0%
−$4.4k $99.9k
−$9.0k (A: −$9.9k)
$160.5k
from $178.6k · −$18.1k
– – 10 / $104 $0–$0
billing fixed: $713–$1.2k

Baseline $59.9k/month today, $108.9k at 2027 Gemini prices with today's volume. "Low case" assumes Android-app questions cost 0.883× the average (see Method). "Deployable rule" = timezone OR CF-IPCountry instead of the modelled inferred country. "Everyone" variant also moves paying users (not recommended). Conversion risk = MRR lost after 12 months if tier 2 cuts free-to-paid conversion by 30–50%, counting trial and grace pipeline; the second line is an upper bound if Android billing were fixed and these markets converted at Qaf's average.

Where the cost goes

Monthly AI cost by component

  • Answer generation (Gemini 3.7 Flash) $48,799 (82%)
  • Retrieval (rerank, embeddings) $9,107 (15%)
  • Chat titles $250 (0%)
  • Non-question AI $1,700 (3%)

Cost per question today: 4.59¢ (3.87¢ generation + 0.72¢ retrieval). At 2027 prices: 8.44¢. Gemini 3.7 Flash list price doubles on 1 Jan 2027, when introductory pricing ends.

Cost vs revenue by platform

PlatformAI cost / moShareMRRMRR ÷ AI cost
Android app$32,83654.9%$2,0390.06
iOS app$17,86129.8%$4,1220.23
Desktop web$5,3028.9%$1,3730.26
Non-question AI$1,7002.8%$00.00
Android web$1,4712.5%$260.02
iOS web$6851.1%$80.01

iOS and web cost per question is not measured directly (see Method); they are assumed to cost the average.

Top 15 segments by AI cost, with revenue

Saudi Arabia · iOS app
$7,774 / $1,840
Saudi Arabia · Android app
$4,514 / $793
Yemen · Android app C
$3,499 / <$100
Iran · Android app C
$2,848 / $0
Egypt · Android app
$2,367 / <$100
Afghanistan · Android app C
$2,068 / $0
United States · iOS app
$1,136 / $367
Saudi Arabia · Desktop web
$1,057 / $389
United Kingdom · iOS app
$1,022 / $333
India · Android app
$956 / $0
Morocco · Android app
$916 / <$100
France · iOS app
$832 / $83
Indonesia · Android app
$782 / $54
France · Android app
$766 / $67
Kuwait · iOS app
$761 / $303
Segment (inferred country · main platform)Active usersQuestions / moAI cost / moShare of costPayersMRRMRR ÷ AI costWeak monetization
Saudi Arabia · iOS app 10,731169,361$7,77413.4% 117$1,8400.24
Saudi Arabia · Android app 6,45899,192$4,5147.8% 55$7930.18
Yemen · Android app 4,06674,441$3,4996.0% <5<$100<0.03Yes
Iran · Android app 2,77764,888$2,8484.9% 0$00Yes
Egypt · Android app 3,06447,292$2,3674.1% <5<$100<0.04
Afghanistan · Android app 2,42446,219$2,0683.6% 0$00Yes
United States · iOS app 1,61724,751$1,1362.0% 37$3670.32
Saudi Arabia · Desktop web 1,16723,016$1,0571.8% 32$3890.37
United Kingdom · iOS app 1,24222,272$1,0221.8% 22$3330.33
India · Android app 1,35020,480$9561.6% 0$00
Morocco · Android app 1,38719,878$9161.6% <5<$100<0.11
France · iOS app 1,11118,114$8321.4% 7$830.10
Indonesia · Android app 95816,357$7821.3% 5$540.07
France · Android app 88816,330$7661.3% 7$670.09
Kuwait · iOS app 1,25816,578$7611.3% 21$3030.40
Algeria · Android app 1,47416,617$7001.2% 0$00Yes
Iraq · Android app 1,18816,077$6881.2% <5<$100<0.15
United Kingdom · Android app 62412,287$6741.2% 14$1860.28
Bangladesh · Android app 1,35414,851$6261.1% 0$00
Germany · iOS app 68811,785$5410.9% 10$1460.27
Ethiopia · Android app 62812,761$5410.9% 0$00Yes
Egypt · iOS app 64810,440$4790.8% 6$650.14
Netherlands · Android app 1884,407$4560.8% <5<$100<0.22
Egypt · Desktop web 5109,850$4520.8% 5$530.12
Pakistan · Android app 5578,418$4380.8% <5<$100<0.23
Saudi Arabia · Android app (guests seen only in Langfuse) 1,51710,349$4370.8% 5$610.14
Turkey · Android app 6119,431$4270.7% <5<$100<0.23
Uzbekistan · Android app 57910,135$4090.7% <5<$100<0.24
Canada · iOS app 5098,698$3990.7% 11$1190.30
UAE · iOS app 6088,250$3790.7% 14$1760.47

Highlighted rows are in scenario C. Saudi Arabia is Qaf's largest cost and revenue market and stays on premium. Yemen's Android app costs about $3.5k a month against roughly $20 of MRR. MRR is gross (mobile includes VAT); segments with 1–4 payers show MRR as "<$100". Total MRR in the window $7,768.

Trend: the target markets are growing fastest

ScenarioIn-scope questions, AugIn-scope, SepGrowth in scopeGrowth elsewhere
A128,813189,547+47%+9%
B165,999238,582+44%+8%
C179,628253,543+41%+8%
C+EG216,814302,578+40%+7%
E179,628253,543+41%+8%
D325,948421,557+29%+7%

Mixpanel chat_message_sent, calendar months. In-scope usage grew ~40–47% month over month vs ~8% elsewhere, so the saving grows over time. One month of growth can't be compounded, so the model uses assumed growth cases: base = +6%/month in scope, +4% elsewhere; high = +12% / +8%.

Method and data sources

Sources

  • Google Cloud (Vertex AI) monitoring: Gemini 3.7 Flash token and call counts for the window (58.4B input, 3.11B output tokens, 18.5% cached). This is the ground truth for generation cost, priced at list.
  • Cloudflare AI Gateway analytics: cross-check of requests and tokens (agrees with Vertex within ~5%).
  • Langfuse (LLM tracing): cost per traced question by user, rerank volume, title cost. Covers the Android app almost completely and iOS/web barely.
  • Mixpanel: questions (chat_message_sent) per user, platform, country signals and monthly trend.
  • Autumn and Stripe (billing): plan and status per customer (paying, trialing, grace, gift, comp, lifetime) and MRR.
  • Production database (read replica): timezone, platform and guest status to join users across sources.
  • Qaf codebase: answer pipeline, models per step, available request-time targeting signals.

Definitions

  • Window: 9 Sep – 8 Oct 2026, 30 full UTC days. Monthly = window × 30.4/30.
  • Question: one sent chat message. 1.25M in the window, including ~68k from Android guests seen only in Langfuse.
  • Segment: inferred country × dominant platform per user. Country is inferred from SIM carrier, or Farsi locale overriding GeoIP for Iranian VPN users. Platforms: Android app, Android web, iOS app, iOS web, desktop web.
  • Non-paying: no non-free Autumn plan in any status. Guests count as non-paying.
  • Weak-monetization rule: ≥50 users, conversion and MRR per user both <25% of average (0.71% and $0.093), payment feasibility none or poor.
  • Central vs low case: Android-app questions cost 1.0× (input-token anchor, measured 0.997) or 0.883× (call-count anchor) the average question.
  • Tier-2 cost: A = 5% of total per-question cost; B = 5% of generation, retrieval unchanged.

Monthly change = Σ over moved segments of questions × cost per question × (1 − tier-2 ratio). Per-segment cost per question varies only within the Android app (where Langfuse measures it; target countries sit at 0.92–1.09× the Android mean). Every other platform gets one flat index.

Caveats and open questions

Implementation notes

  • Today every user gets the same model (Gemini 3.7 Flash through a single Cloudflare AI Gateway route). There is no plan-based model selection yet. Estimated effort: ~1–2 dev days plus the evaluation.
  • Model switch: add a second gateway route (e.g. qaf-t2) or branch inside the dynamic route on request metadata, so no model id lives in code.
  • Decision point: the chat handler already has headers, session (timezone, platforms, guest flag) and the billing snapshot per turn, so the paid check costs no extra latency. The same decision must also feed the per-question policy-check call, which uses the same model.
  • Signals: x-qaf-client-platform (native), x-qaf-client-timezone (all clients, also persisted on the user), CF-IPCountry (api.qaf.ai is Cloudflare-proxied; confirm with one log line). Android web needs User-Agent or client-hint parsing, which nothing does today.
  • Edge cases: legacy native clients (<1.3.0) and requests without a platform header are unmetered and have no billing context; treat them as unpaid or load the customer. A paying user travelling to a target country stays premium because the check runs per request.
  • Retrieval: under interpretation B, reranking (0.72¢/question) becomes most of a tier-2 question's cost. Use rerank-v4.0-fast or skip rerank for tier 2.
  • Measurement: add tier, country and platform to the Langfuse trace metadata, and fix Langfuse coverage of iOS and web, so the real saving can be tracked after launch.