OpenAI-compatible POST /v1/chat/completions. No account, no key. Your agent pays each call in USDC on Base with x402; settlement is on-chain to 0xc91cE6291eDC0713ec753BAFBA002506ffb2b95c.
| model | upstream (Cloudflare Workers AI) |
|---|---|
llama-3.3-70b | @cf/meta/llama-3.3-70b-instruct-fp8-fast |
gpt-oss-120b | @cf/openai/gpt-oss-120b |
llama-4-scout | @cf/meta/llama-4-scout-17b-16e-instruct |
llama-3.1-8b | @cf/meta/llama-3.1-8b-instruct-fp8-fast |
curl -X POST https://inference.agentexchange.work/v1/chat/completions \
-H 'content-type: application/json' \
-d '{"model":"llama-3.3-70b","messages":[{"role":"user","content":"Summarize x402 in one line."}],"max_tokens":128}'
# -> HTTP 402 with a payment challenge; an x402 client (x402-fetch, @x402/fetch, AgentKit) pays it and retries automatically.
Limits: 12,000 input characters, 512 output tokens per call, 20 most recent messages. Machine docs: /.well-known/x402, /openapi.json, /llms.txt, /v1/models.
Part of agentexchange.work. Contact: riley@agentexchange.work