LLM API Pricing for a Micro-SaaS: Price the Workload
Price LLM APIs by successful workflow, using dated token rates, a 100,000-completion cost fixture, batch discounts, and explicit self-hosting limits.
Filed under Decision Guides· see every report on this topic
Start with one direct provider. Use a stronger model where a failure costs the customer. Use a cheap model only where a machine checks the result. Add routing after a second provider fixes a measured reliability, capability, or spend problem. Self-host after measured API spend can pay for capacity and its ops work.
Customers pay for a finished reply, a valid extracted record, or a completed agent task. They do not pay for a low token rate.
Proof status. Token prices came from vendor pages fetched on 21 August 2026 and are linked at every figure. The monthly example is arithmetic, not a benchmark. The self-hosted comparison is blocked / unverified. Needs a measured benchmark. No GPU benchmark was run. Check sources before setting a customer price.
The operator and the leak
A solo builder has an AI feature inside a SaaS and about 100,000 user-facing completions a month. The work includes chat, upload-to-fields extraction, multi-tool loops, and a nightly back-catalogue job.
One default model handles every request. A 200-token extraction, a long agent loop, and asynchronous bulk work get the same price, latency, and failure plan. A rate-limit spike or model retirement then becomes a product incident.
Split the work before choosing a model
| Workload | Customer result | Stack decision |
|---|---|---|
| Chat / completions | Response while they wait | First token, output speed, context, capacity fallback |
| Agentic multi-tool loops | Task completed across tools | Tool-call reliability, retries, context growth, idempotency |
| Structured extraction | Valid JSON or database fields | Schema adherence, validation failures, retry cost |
| Bulk batch jobs | Queue completes later | Token rates, batch discount, throughput, queue back-pressure |
Chat needs low perceived latency. A pricier agent model can make sense when one bad tool call causes five retries. Extraction needs a validator. Batch work needs an asynchronous rate.
Customer records add a data-handling check. Read the direct provider’s retention and regional terms, then the router’s policy for the exact model. A low token price does not override an agreement that forbids that data transfer.
Compare listed token prices
These standard listed prices are per 1M input/output tokens, as of 21 Aug 2026. They do not prove that a model fits your task. Cached input, tool charges, reasoning tokens, regional processing, and batch rates can change the bill. OpenAI says its Batch API cuts input and output prices by 50%; Anthropic says batch processing saves 50%.12
| API path | Model | Input / 1M | Output / 1M |
|---|---|---|---|
| OpenAI direct | GPT-5.6 Sol | $5.00 | $30.00 |
| OpenAI direct | GPT-5.6 Terra | $2.00 | $12.00 |
| OpenAI direct | GPT-5.6 Luna | $0.20 | $1.20 |
| Anthropic direct | Fable 5 | $10.00 | $50.00 |
| Anthropic direct | Sonnet 5 | $2.00 | $10.00 |
| Anthropic direct | Haiku 4.5 | $1.00 | $5.00 |
| OpenRouter | Gemini 3.7 Flash | $0.375 | $1.875 |
| OpenRouter | Qwen3.7 Flash | $0.03 | $0.13 |
| OpenRouter | gpt-oss-120b (CoreWeave) | $0.03 | $0.17 |
Receipts, as of 21 Aug 2026: OpenAI pricing, Anthropic API pricing, OpenRouter models · Qwen3.7 Flash · gpt-oss-120b.
OpenRouter is a routing and billing layer, not a model vendor. Its pay-as-you-go plan lists a 5.5% platform fee plus model-based pricing.3 The gpt-oss-120b number is CoreWeave’s listed provider price on that model page; other hosts list different prices. Treat it as an endpoint choice.
A lower token rate can still cost more per finished job. Schema misses, bad tool calls, and larger prompts erase savings through retries. Put every model behind a named workload and a test set.
Price 100,000 completions
This is a pricing fixture, not observed production traffic. The monthly split is:
| Workload | Completions | Tokens each | Monthly tokens |
|---|---|---|---|
| Chat | 40,000 | 800 input + 400 output | 32M input / 16M output |
| Agentic | 20,000 | 5,000 input + 1,500 output | 100M / 30M |
| Extraction | 30,000 | 1,000 input + 200 output | 30M / 6M |
| Bulk | 10,000 | 2,000 input + 300 output | 20M / 3M |
The bulk row applies the vendor-stated 50% Batch discount to OpenAI Terra.
| Stack fixture | Workload routing | Monthly token cost | Meaning |
|---|---|---|---|
| Single-provider direct | Terra for chat, agents, extraction; Terra Batch bulk | $986.00 | $256 + $560 + $132 + $38 |
| Routed by workload | Terra chat; Sonnet 5 agents; Luna extraction; Terra Batch bulk | $807.20 | $256 + $500 + $13.20 + $38; direct calls only |
| Self-hosted fixture | Open-weight model on own capacity | Blocked / unverified | No measured throughput, utilization, hardware, electricity, serving, reliability data |
The routed fixture saves $178.80/month under its assumptions. It excludes evaluation, fallback testing, observability, and whether Sonnet is best for the tool calls. Its useful point is narrower: my take is to challenge one default model once traffic is real.
For self-hosting, calculate:
monthly self-host cost = fixed GPU + electricity + serving/monitoring + labour + variable costs
Then compare the measured API bill with:
successful workflows/month = monthly fixed self-host cost ÷ (direct API cost/workflow − self-host variable cost/workflow)
Measure model, hardware, context, concurrency, and reliability target first. A numeric crossover without them is unverified. Needs a measured benchmark. Free weights do not make serving free.
Expect operational failures
A model can disappear. TubeSpark’s builder said Google deprecated gemini-1.5-flash; requests returned 404 until the check moved to gemini-2.5-flash.4 This is one builder’s account, not an SLA. Keep model IDs configurable and run a canary. A fallback only helps after a test.
Track 429s, queue delay, retry count, and fallback activation separately from model failures. With multiple providers, token accounting, caches, tool charges, and invoice periods also need a request record for customer, workflow, provider, model, and retry outcome.
Cheap models can produce valid JSON with the wrong field. Track schema-valid rate and sampled task correctness, not only HTTP 200s and token cost.
When to keep it simple
Do not add a router before more than one provider is in production for availability, capability, compliance, or measured cost/quality. “Optionality” adds an abstraction layer, tests, logging, model-version policy, and failure states.
Do not self-host before tokens are a real P&L line item and you can name the box operator. You need sustained utilization and an answer for upgrades, cold starts, incidents, and regressions. Direct API can cost less operationally even with a higher token rate.
For version one, keep one direct provider, model IDs in configuration, a request-level usage log, and a manual fallback procedure that you have run.
Bottom line
OpenAI direct, Anthropic direct, OpenRouter, and self-hosting are different purchases with different capability, routing, billing, data-handling, and operations trade-offs. Price a successful workflow first. Use strong models for hard work, cheaper models after validation, batch for work that can wait, and routing after measured evidence appears.
A leaderboard supplies candidates. It cannot price your SaaS.
More on this decision, three ways to look at it:
Footnotes
-
OpenAI API Pricing, fetched 21 August 2026. ↩
-
Anthropic API Pricing, fetched 21 August 2026. ↩
-
OpenRouter Pricing, fetched 21 August 2026. ↩
-
“Why I chose 4 AI providers instead of just OpenAI, and what happened when Google killed a model without warning,” Indie Hackers, posted 2 March 2026 and fetched 21 August 2026. ↩
Get the next verdict before it's everywhere.
One email when a new lab post or cost table ships. No spam, no confirmation step — unsubscribe anytime.