API reference
Base URL: https://api-production-c74b9.up.railway.app (this site's current API). Inference is OpenAI-compatible; everything else is plain JSON.
Authentication
- Inference takes an API key:
Authorization: Bearer ik_…orx-api-key: ik_…. Without either, it takes an x402 payment instead (below). - Account routes (keys, balance, usage, seller) take the session token from sign-in as
Authorization: Bearer <session>. Sessions last 24h. Balance and usage also accept your API key. - Public routes (models, prices, marketplace, stats, rails) need no auth and are cacheable.
Money
All amounts are integers in millionths of a unit (fields ending in Micro), serialised as decimal strings. On the testnet the unit is ITC, which counts as one test dollar and has no monetary value: "1500000" = 1.50 ITC. Prices are in the same micro-units per 1M tokens, quoted in $ so they compare with list prices, and are all-in (seller price + platform fee) unless a field says otherwise. Percentages are numbers with two decimals; a discount is rounded down so it is never overstated.
Inference
/v1/chat/completionsAPI keystream: true (SSE relayed byte for byte). Before routing, your available balance on the selected rail must cover the worst case: input estimate plus the output limit at the top-ranked offer. The output limit is the smaller of max_tokens and max_completion_tokens, or 1024 when you send neither, and the seller is held to it: a long answer without a limit stops at 1024 tokens with finish_reason: "length". You are then charged the metered cost.curl https://api-production-c74b9.up.railway.app/v1/chat/completions \
-H "Authorization: Bearer $INFERIT_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"qwen/qwen-2.5-14b-instruct","messages":[{"role":"user","content":"Hello"}]}'/min{N}/v1/chat/completionsAPI keyx-min-discount header or a per-key setting; the strictest wins./anthropic/v1/messages—501 with a docs link. See Claude Code.Response headers on every inference call:
| Header | Meaning |
|---|---|
x-request-id | Request id, also in your usage log |
x-inferit-attempts | Offers tried (failover happens only before the first byte) |
x-inferit-buyer-cost-micro | What this request cost you, in micro-units (µITC on the testnet). A stream's final cost is known only after its last chunk; see your usage log. |
x-inferit-offer | Opaque hash of the offer that served you, not the seller's identity. On Whitechain the settlement itself is public; see On-chain visibility |
x-inferit-rail | Rail the request is funded and settled on |
x-inferit-network | Settlement network, on every response: whitechain-sepolia today (whitechain-mainnet after launch) |
x-inferit-testnet | 1 on every response from a testnet deployment: balances are test tokens with no value |
Errors
Errors use OpenAI's shape. Notable codes: 402 insufficient_funds (with a key; without one, a 402 is an x402 offer), 404 model_not_found, 503 no_available_offers (with retry-after), 429 rate_limit_exceeded (per client IP, with retry-after; sign-in and dev routes have a stricter budget), 401 for a missing or bad key.
{
"error": {
"message": "Your available balance on whitechain does not cover the worst-case cost of this request.",
"type": "insufficient_funds",
"code": "insufficient_funds"
}
}x402: pay per request, no key
Standard x402 v2, exact scheme, settled into the escrow and credited to the payer. The walkthrough and client code are on Pay per request with x402.
/v1/chat/completionsnone → 402Authorization or x-api-key header: 402 with a PAYMENT-REQUIRED header (base64 JSON) and the same document as the body. It offers one requirement: scheme: "exact", network: "eip155:<chainId>", asset = the token, payTo = the escrow, amount = the worst-case all-in cost of this exact body (at least 0.01 ITC), maxTimeoutSeconds and extra: { name, version }, the token's EIP-712 domain. The quote is pinned to the body for 120 s. The same applies to /min{N}/v1/chat/completions./v1/chat/completionsPAYMENT-SIGNATUREPAYMENT-SIGNATURE header (base64 JSON payload with an EIP-3009 authorization; X-PAYMENT is accepted too). The API verifies it, settles it first with depositWithAuthorization, then serves the request against the payer's escrow balance (streaming included). The response carries PAYMENT-RESPONSE, base64 of { success: true, transaction, network, payer }, plus the usual x-inferit-* headers. A payment that does not verify gets a fresh 402 with the reason in error. If the request fails after the payment settled, the credit stays in the payer's escrow balance and the error says so (an error.x402 object with credited, payer, amount and transaction; a funding 402 at that point becomes 503, so the client does not pay again). 503 x402_settlement_pending: the deposit was sent but its receipt did not arrive in time; if it lands it is the payer's escrow balance. 503 x402_unavailable: nothing was sent and nothing was charged./v1/faucetnone{ address }: the API mints 1,000 test ITC to that address with faucetTo and pays the gas. At most once per 24 hours per address (429 faucet_cooldown); rate-limited per IP. 404 when the server does not sponsor the faucet./.well-known/x402404 when x402 is off.Marketplace (public)
/v1/modelspricing (list price) and best_price (best all-in offer)./v1/prices/api/marketplace/api/markets/:model/api/stats/v1/rails/v1/providers/resale-allowlist/health/metrics/llms.txtGET /api/marketplace
{
"feeBps": 500,
"markets": [{
"modelId": "qwen/qwen-2.5-14b-instruct",
"name": "Qwen: Qwen2.5 14B Instruct",
"openWeight": true,
"licenseId": "apache-2.0",
"listInputPerM": "100000", // µUSD per 1M tokens ($0.10)
"listOutputPerM": "200000",
"bestInputPerM": "84000", // all-in, healthy offers only
"bestOutputPerM": "168000",
"bestDiscountPct": 16,
"sellers": 3, "healthySellers": 2,
"requests24h": 1840, "volume24hMicro": "2214551",
"uptime24hPct": 99.2, "ttftP50Ms": 310
}]
}GET /api/markets/qwen%2Fqwen-2.5-14b-instruct
{
"modelId": "qwen/qwen-2.5-14b-instruct",
"summary": { ...same shape as a marketplace row... },
"levels": [
{ "inputPerM": "84000", "outputPerM": "168000", "offers": 2, "healthyOffers": 2, "capacityMicro": "18000000" },
{ "inputPerM": "95000", "outputPerM": "190000", "offers": 1, "healthyOffers": 0, "capacityMicro": null }
]
}Sign-in
/v1/auth/evm/challenge{ address }. Returns { message, nonce }: a one-time message to sign with the wallet./v1/auth/evm/verify{ message, signature }. Verifies the signature and returns { token, expiresAt, account }.Keys
/v1/keyssession{ name?, spendLimitMicro?, minDiscountPct? }. Returns the key once in key./v1/keyssession/v1/keys/:idsessionBuyer
/v1/balancesession or API keypendingMicro) and usage settled on-chain but not yet final (pendingFinalityMicro). Available subtracts both./v1/buyer/railsession{ rail: "whitechain" }. Selects Whitechain settlement for future requests ("credit" exists only on local development APIs)./v1/usage?limit&cursorsession or API keynextCursor. Each row carries finality: pending until its settlement batch is finalized on-chain, then finalized. Rows never carry the seller's identity, a transaction hash or a batch id; your own settlements are listed with explorer links on your dashboard./v1/usage/export.csvsession or API keySeller
/v1/seller/offerssession{ model, kind: "endpoint" | "upstream_key", endpointUrl, authToken?, upstreamProvider?, upstreamKey?, pricing, capDailyMicro?, licenseAck }. pricing is { mode: "per_token", inputPerM, outputPerM, cacheReadPerM? } or { mode: "multiplier", multiplierBps } (10,000 = list price). The API checks the licence, the resale allowlist and an SSRF guard, sends a live 1-token probe, and requires you to be registered on your rail./v1/seller/offerssession/v1/seller/offers/:idsession/v1/seller/offers/:idsession/v1/seller/earningssessionpendingFinalityMicro and finalizedMicro.