LLM inference for agents, $0.005 per call

OpenAI-compatible POST /v1/chat/completions. No account, no key. Your agent pays each call in USDC on Base with x402; settlement is on-chain to 0xc91cE6291eDC0713ec753BAFBA002506ffb2b95c.

modelupstream (Cloudflare Workers AI)
llama-3.3-70b@cf/meta/llama-3.3-70b-instruct-fp8-fast
gpt-oss-120b@cf/openai/gpt-oss-120b
llama-4-scout@cf/meta/llama-4-scout-17b-16e-instruct
llama-3.1-8b@cf/meta/llama-3.1-8b-instruct-fp8-fast

Call it

curl -X POST https://inference.agentexchange.work/v1/chat/completions \
  -H 'content-type: application/json' \
  -d '{"model":"llama-3.3-70b","messages":[{"role":"user","content":"Summarize x402 in one line."}],"max_tokens":128}'
# -> HTTP 402 with a payment challenge; an x402 client (x402-fetch, @x402/fetch, AgentKit) pays it and retries automatically.

Limits: 12,000 input characters, 512 output tokens per call, 20 most recent messages. Machine docs: /.well-known/x402, /openapi.json, /llms.txt, /v1/models.

Part of agentexchange.work. Contact: riley@agentexchange.work