PagesPilot · Cost Estimate · Verified March 8, 2026
Infrastructure Cost Expected Monthly
No brand-count assumptions. Worst case = maximum plausible monthly bill if every service runs at peak load simultaneously. Every figure sourced from official provider docs.
🖥️ Compute
$37
normal $37 · worst $69
🗄️ Database
$19
normal $19 · worst $60
☁️ Storage
$0
normal $0 · worst $8
⚡ AI Models
$73
normal $73 · worst $380
Monthly expected — normal operating load
$129/mo
2-month estimate
$258
Service
Pricing Unit
Normal
Worst
🪰
Fly.io — Node.js Backend
shared-cpu-2x · 1GB RAM · always-on
$0.00000292/sec
$7
$19
🪰
Fly.io — Python Scraper
shared-cpu-2x · 2GB RAM · job-based
$0.00000584/sec for 1GB machine
$10
$30
▲
Vercel — Frontend
Pro plan (required for commercial use)
$20/user/month (Pro plan)
$20
$20
🗄️
Neon — PostgreSQL + pgvector
Launch plan · usage-based
$0.106/CU-hour
$19
$60
☁️
Cloudflare R2
Pay-as-you-go · Standard storage
$0.015/GB-month
—
$8
⚡
Gemini 2.5 Flash
Paid tier · primary model
Input: $0.30/M tokens
$25
$120
🧠
Gemini 2.5 Pro
Paid tier · use sparingly
Input: $1.25/M tokens
$35
$200
🖼️
Gemini 2.5 Flash Image
Paid tier · per image
$30/M output tokens
$12
$55
🔢
Gemini text-embedding-004
Paid tier · negligible cost
Free tier generous
$1
$5
Monthly Total
All services combined
$129
$517
⚠ Cost Risk Register — What Can Blow Up the Budget
🔴 HIGH
Gemini 2.5 Pro — long context pricing
Prompts exceeding 200K tokens are billed at 2× rate on BOTH input and output. A brand system prompt + conversation history + RAG chunks can easily breach this. Always chunk context, never send full brand memory in one call.
Mitigation
Enforce max context budget per API call. Use sliding window for chat history.
🔴 HIGH
Gemini 2.5 Flash — thinking tokens
Gemini 2.5 Flash's thinking tokens count as output tokens and are billed at $2.50/M. A single request can silently generate thousands of thinking tokens, multiplying your expected output cost.
Mitigation
Set thinking budget explicitly: thinkingConfig: {thinkingBudget: 0} for simple tasks. Only enable thinking for complex reasoning tasks.
🟡 MEDIUM
Fly.io — no autostop on Python scraper
If autostop is not configured, Playwright machines run 24/7 even between jobs. A forgotten always-on 2GB machine = $15+/mo wasted.
Mitigation
Set min_machines_running = 0 in fly.toml. Machines spin up on job trigger, stop after completion.
🟡 MEDIUM
Neon — concurrent query spikes
RAG similarity searches, scraper writes, and chatbot queries hitting simultaneously during peak hours can spike CU usage dramatically beyond estimate.
Mitigation
Enable connection pooling (PgBouncer built into Neon). Set autoscaling max CU limit to cap runaway costs.
🟢 LOW
Cloudflare R2 — operations cost
Free tier covers 1M Class A and 10M Class B ops/month. Only becomes a cost above this threshold. At early scale, very unlikely to exceed.
Mitigation
Monitor via Cloudflare dashboard. Only becomes material at very high image serving volume.
💡 How to Stay Close to Normal — Not Worst Case
💾
Context caching on brand system prompts
Up to −90% on repeated input tokens
Cache each brand's voice/product prompt. Cache reads = 10% of input price. Every chatbot message reuses the same brand system prompt.
⏳
Batch API for scheduled content
−50% on Gemini Pro + Flash
All non-real-time jobs (weekly posts, campaign drafts) go through Batch API. Only live chatbot replies need standard API.
🧠
Flash-first routing rule
−70% vs routing everything to Pro
Classify request complexity before calling API. Simple Q&A → Flash ($0.30/$2.50). Complex campaign → Pro ($1.25/$10).
🪰
Fly.io autostop on scraper
−40–60% on Python machine cost
Scraper only runs on scheduled jobs. Machine stops between runs. You only pay for actual scraping time.
💡
Disable thinking tokens for chatbot
−30–50% on Flash output cost
Chatbot replies don't need deep reasoning. Set thinkingBudget: 0 for all conversational tasks. Save thinking for content generation only.
🖼️
Default to 1024×1024 images
−60% vs 2K resolution
1K image = $0.039. 2K image = $0.101. For social media posts, 1024×1024 is more than sufficient quality.
Sources verified March 8, 2026: fly.io/docs/about/pricing · neon.com/pricing · vercel.com/pricing · developers.cloudflare.com/r2/pricing · ai.google.dev/gemini-api/docs/pricing Worst case = all services at peak load simultaneously. Normal = realistic operating load. Click any row for full breakdown + source link.