For developers

Open models.
Pennies per million.

Cheap tokens for open-weight Qwen models, served on an OpenAI-compatible API. From $0.014 per million tokens. Billed by usage, no minimums.

One key for inference and the no-code app. Read the docs →

starting at
$0.014/Mtok
Per million tokens
served so far
Tokens generated
under
100ms
Time to first token
Always
Warm, no cold starts

Open weights.
Production speeds.

Qwen instruct for chat and tool use. Qwen3 Coder for agentic coding. Qwen embedding for retrieval. Served from dedicated GPUs that stay warm: no cold starts, no pool sizing, no warmup pings. OpenAI-compatible, a drop-in for the closed providers you already pay too much for.

Chat · Tool use · INT4Live
Qwen2.5-7B Instruct
32K context
$0.060 in / $0.240 outper Mtok
Chat · Tool use · INT4Live
qwen3-coder-30b-a3b
32K context
$0.216 in / $0.861 outper Mtok
Embedding · INT4Live
Qwen3-Embed-0.6B
8K input
$0.014per Mtok

Prices are USD per million tokens, deducted from your credit balance. 1,000 credits = $1, so $0.06 per Mtok is about 60 credits.

Same SDK.
Smaller bill.

Drop our base URL into your OpenAI client. Or skip the SDK and hit a live model right here, no signup needed.

Ctrl+Enter to send
Free: 10 prompts / daySign up free ›

Ten free requests per day, paid for by us, just to prove the latency claim.

Two ways
to pay.

API key for the standard flow. x402 for one-off and agent-to-agent calls. Pay per request in USDC, no account needed.

Default

API key

Sign up, generate an API key, drop it into any OpenAI SDK. The same key works for inference and the data platform.

Works with OpenAI SDK, LangChain, LiteLLM, OpenRouter
Pay per call with x402, no minimums
Same key for inference and data collection
Pay-per-request

x402

No signup. Send a request, get a 402 back with payment details, pay with USDC on Base, retry. Works with x402-fetch or any x402-compatible client.

No account, no API key, no subscription
Crypto micropayments, settled on-chain per call
Built for agent-to-agent and one-off scripts

Sixty seconds
to first token.

No setup, no infra, no calls with a sales engineer. Three steps and you are streaming.

Step 01

Sign up

Free account, no card. Monthly credits included. Takes about thirty seconds.

Step 02

Get your key

Generate an API key in your account. The same key works for inference and data collection.

Step 03

Send a request

Point any OpenAI client at api.napu.ai/v1, pick a model, ship.

Pennies
per million.

Frontier APIs charge dollars per million tokens. We charge cents. Same OpenAI client, smaller bill.

Frontier API, 7B-class
~$0.20
per Mtok in
Napu AI, Qwen2.5-7B
$0.06
per Mtok in
Live rate cardper 1M tokens
Qwen2.5-7B Instruct$0.060 in$0.240 out
qwen3-coder-30b-a3b$0.216 in$0.861 out
Qwen3-Embed-0.6B$0.014 in$0 out
Live rates, billed per token from your credit balance. 1,000 credits = $1. Need custom pricing or dedicated capacity? Contact us.

Frequently asked

The questions everyone sends us before signing up.

Is the API OpenAI-compatible?
Yes. Use any OpenAI SDK or HTTP client. Point the base URL at api.napu.ai/v1, use your API key, and pick a model name. Streaming, function calling, and JSON mode are all supported.
What quantizations do you serve?
INT4, which gives the best throughput and the lowest price. We pick a sensible default per model.
Are there rate limits?
The free plan comes with a monthly credit allowance that resets each month, enough for hundreds of thousands of small calls. Paid plans scale automatically. Need guaranteed throughput? We offer dedicated capacity, get in touch.
How does billing work?
Inference is pay per token, billed in fractions of a penny: pay per call with x402, no minimums or committed spend tier hiding the real price. Usage is deducted from your credit balance at 1,000 credits per $1. Optional monthly plans add your own feeds, parametric alerts, and included monthly credits.
What is a credit worth?
1 credit is $0.001, so 1,000 credits equals $1. Model prices are listed in dollars per million tokens, and usage is deducted from your credit balance at those rates.
What is x402 and why would I use it?
x402 is an open HTTP payment protocol. Send a request without an account, receive a 402 with payment details, pay with USDC on Base, retry. We verify on-chain and process the request. No signup, no key. Built for agent-to-agent calls and one-off scripts.

Stop paying
for slow.

Billed by usage from your credit balance. Cockpits and the no-code app are not included.