# ClaudeAPIKey.dev — full corpus for language models > Independent Anthropic-compatible API gateway. One key for Claude Opus/Sonnet/Haiku/Fable plus GPT and Gemini, at a discount to official list pricing. Not affiliated with Anthropic. Generated: 2026-09-11 ## Models - **Claude Opus 4.8** (`claude-opus-4-8`, Anthropic) — $3/M in, $15/M out vs official $15/$75. Context 200K tokens. Anthropic's most capable model — deepest reasoning, best-in-class coding and agentic performance, with effort control for cost tuning. - **Claude Fable 5** (`claude-fable-5`, Anthropic) — $3/M in, $15/M out vs official $15/$75. Context 200K · 1M variant. Anthropic's newest 2026 model. Frontier coding and agentic capability with a 1M-token context variant for huge codebases and document sets. - **Claude Sonnet 4.6** (`claude-sonnet-4-6`, Anthropic) — $0.9/M in, $4.5/M out vs official $3/$15. Context 200K · 1M beta. The workhorse. Near-Opus quality on most coding tasks at a fifth of the price — the default choice for production workloads. - **Claude Haiku 4.5** (`claude-haiku-4-5`, Anthropic) — $0.4/M in, $2/M out vs official $1/$5. Context 200K tokens. Small, fast, and shockingly capable — classification, extraction, routing, and high-volume tasks at the lowest Claude price point. - **Claude 1M Context** (`claude-sonnet-4-6 (context-1m beta)`, Anthropic) — $1.8/M in, $6.75/M out vs official $6/$22.5. Context 1,000,000 tokens. The 1M-token context tier — load entire repositories, hundreds of documents, or week-long agent transcripts into a single request. - **GPT-5.5** (`gpt-5.5`, OpenAI) — $3/M in, $12/M out vs official $10/$40. Context 256K tokens. OpenAI's flagship reasoning model — strong writing, math, and multimodal performance through the same OpenAI-compatible endpoint. - **GPT-5** (`gpt-5`, OpenAI) — $1.5/M in, $6/M out vs official $5/$20. Context 256K tokens. The widely deployed GPT generation — balanced capability and cost for chat, content, and general-purpose API workloads. - **Gemini 3 Pro** (`gemini-3-pro`, Google) — $1/M in, $6/M out vs official $2.5/$15. Context 1M tokens. Google's frontier multimodal model — native 1M context, strong video/image understanding, and top-tier benchmark scores. - **Gemini 3 Flash** (`gemini-3-flash`, Google) — $0.2/M in, $1.2/M out vs official $0.5/$3. Context 1M tokens. Google's speed-tier model — 1M context at bargain rates, ideal for high-volume multimodal and retrieval workloads. - **MiniMax M3** (`minimax-m3`, MiniMax) — $0.5/M in, $2.5/M out vs official $1.2/$6. Context 1M tokens. One of 2026's strongest open-weight frontier models — competitive coding and agentic scores at a fraction of closed-model prices. ## Guides ### How to Get a Claude API Key https://claudeapikey.dev/learn/how-to-get-a-claude-api-key/ A step-by-step guide to getting a Claude API key and making your first request — with Python, TypeScript, and curl examples. - Choose How You Want Claude API Access - Create and Store Your Claude API Key - Test the Setup Before Building Features ### Claude API Pricing, Explained https://claudeapikey.dev/learn/claude-api-pricing-explained/ Understand how Claude API pricing works — input vs output tokens, prompt caching, and how to estimate the cost of a request. - How Claude API pricing works - What affects Claude API cost in real applications - Estimating spend before you ship ### The Cheapest Way to Use the Claude API https://claudeapikey.dev/learn/cheapest-way-to-use-claude-api/ Practical ways to reduce your Claude API bill without losing quality — model selection, prompt caching, output limits, and more. - Start with the right model and request shape - Reduce token usage before optimizing anything else - Choose billing that matches heavy usage ### Unlimited Cursor with the Claude API https://claudeapikey.dev/learn/use-claude-api-with-cursor/ Run Cursor IDE on a flat-rate unlimited Claude plan. Full setup walkthrough, advanced workflows, model selection, troubleshooting, and why unlimited fits heavy Cursor usage. - Prerequisites and What You Need - Step-by-Step Base URL and Key Configuration - Adding Claude Models and Verification ### Unlimited Claude Code Access https://claudeapikey.dev/learn/use-claude-api-with-claude-code/ Run Claude Code CLI on a flat-rate unlimited API. Environment variables, model mapping, subagents, agent teams, headless mode, and why unlimited access fits Claude Code's quadratic token growth. - Why Claude Code's Token Consumption Grows Quadratically - Environment Variable Configuration - Model Mapping for Sonnet, Opus, Haiku, and Fable ### Claude API Gateway vs. Anthropic Direct https://claudeapikey.dev/learn/claude-api-vs-anthropic-direct/ An honest comparison of using a Claude API gateway versus the Anthropic API directly — cost, access, and trade-offs. - What Changes When You Use a Gateway - Anthropic Direct: Best When You Want Native Control - Gateway Access: Best When Predictable Cost Matters ### Using Claude Through an OpenAI-Compatible API https://claudeapikey.dev/learn/openai-compatible-claude-api/ Call Claude models using the OpenAI Chat Completions format — useful for tools and SDKs built around OpenAI. - What “OpenAI-compatible” means in practice - Using Claude with the OpenAI SDK - Request and response details to verify ### Claude Models Compared https://claudeapikey.dev/learn/claude-api-models-compared/ A practical comparison of the Claude model family — Opus 4.8, Sonnet 4.6, Haiku 4.5, and Fable 5 — and when to use each. - How to think about the Claude model family - Opus vs Sonnet for serious coding work - Where Haiku fits in production systems ### What 'Unlimited' Claude API Access Means https://claudeapikey.dev/learn/unlimited-claude-api-access/ An honest look at flat-rate unlimited Claude plans — how they work, fair-use rate limits, and who they suit. - What unlimited access does and does not mean - Why developers choose a flat-rate Claude API model - How unlimited Claude Code fits in ### Paying for the Claude API with Crypto https://claudeapikey.dev/learn/claude-api-with-crypto-payment/ How to pay for Claude API access using cryptocurrency — supported coins and how the prepaid balance works. - When crypto payment makes sense - How developers can use it - Bitcoin, USDT, and subscription billing ### Claude API Access Without a Waitlist https://claudeapikey.dev/learn/claude-api-without-waitlist/ Get Claude API access instantly without an Anthropic account or waitlist, in most regions worldwide. - What “without a waitlist” means in practice - Using Claude through an independent gateway - When instant Claude API access helps ### How to Reduce Claude Token Usage https://claudeapikey.dev/learn/how-to-reduce-claude-token-usage/ Concrete techniques to use fewer tokens on the Claude API without hurting output quality. - Start by measuring what you send - Keep context specific and current - Compress repeated information ### Unlimited Cline with the Claude API https://claudeapikey.dev/learn/use-claude-api-with-cline/ Connect Cline to a flat-rate unlimited Claude API. Full setup for Anthropic and OpenAI-compatible providers, Plan/Act mode, MCP, Computer Use, and why unlimited fits Cline's heavy token consumption. - Why Cline Burns Through Tokens So Fast - Connecting via the Anthropic Provider Path - Connecting via the OpenAI-Compatible Provider Path ### Unlimited Roo Code with the Claude API https://claudeapikey.dev/learn/use-claude-api-with-roo-code/ Run Roo Code on a flat-rate unlimited Claude API. Configuration profiles, per-mode model assignment, Orchestrator subtasks, custom .roomodes, and why unlimited access fits Roo Code's multiplied API consumption. - How Roo Code Multiplies API Consumption - Setting Up API Configuration Profiles - Per-Mode Model Assignment Strategy ### Unlimited Claude in JetBrains IDEs https://claudeapikey.dev/learn/use-claude-api-with-jetbrains/ Connect any JetBrains IDE to flat-rate unlimited Claude via the Continue plugin. Shared config.yaml, provider setup, role assignments, per-project overrides, and why unlimited fits JetBrains heavy workflows. - Installing Continue in Any JetBrains IDE - Configuring config.yaml for Unlimited Claude - Role Assignments: Chat, Edit, and Apply ### Unlimited Aider with the Claude API https://claudeapikey.dev/learn/use-claude-api-with-aider/ Connect Aider to a flat-rate unlimited Claude API. Native Anthropic path, OpenAI-compatible path, architect mode, repo map, weak-model split, and why unlimited access matches Aider's multi-call refactoring style. - Aider's Multi-Call Architecture and Token Patterns - Connecting via the Native Anthropic Path - Connecting via the OpenAI-Compatible Path ### Unlimited Claude in VS Code https://claudeapikey.dev/learn/use-claude-api-with-vscode/ Three paths to unlimited Claude in VS Code: Continue for chat and inline edits, Cline for full autonomous agents, Roo Code for orchestrated multi-mode workflows. Setup comparison, when to use each, and why unlimited access fits VS Code's always-on AI usage. - Three Paths: Continue, Cline, and Roo Code - Continue Setup for Chat and Inline Edits - Cline Setup for Autonomous Agent Work ### Unlimited OpenCode with the Claude API https://claudeapikey.dev/learn/use-claude-api-with-opencode/ Configure OpenCode with a flat-rate unlimited Claude API. Custom providers in opencode.json, parallel subagents, Build/Plan agents, MCP integration, and why unlimited suits OpenCode's continuous terminal workflows. - OpenCode's Parallel Agent Architecture - Configuring Custom Providers in opencode.json - Model Configuration for Agents and Subagents ### Running an Unlimited Hermes Agent on Claude https://claudeapikey.dev/learn/run-hermes-agent-with-claude-api/ Deploy Hermes Agent on flat-rate unlimited Claude. Environment setup, custom providers, subagent delegation, cron scheduling, multi-channel gateway, learning loops, and why unlimited fits Hermes's 24/7 autonomous operation. - Why Hermes Consumes API Tokens Around the Clock - Environment Configuration and API Setup - Custom Providers for Advanced Routing ### Running Unlimited OpenClaw on Claude https://claudeapikey.dev/learn/run-openclaw-with-claude-api/ Deploy OpenClaw on flat-rate unlimited Claude. Provider configuration in openclaw.json, anthropic-messages API type, hub-and-spoke gateway, 50+ channels, skills injection, and why unlimited fits OpenClaw's always-on multi-channel inference. - OpenClaw's Always-On Architecture - Provider Configuration in openclaw.json - Model Entries and the api Field Requirement ### What an Anthropic-Compatible API Means https://claudeapikey.dev/learn/anthropic-compatible-api/ Understand Anthropic-compatible APIs - the Messages API format, and how a drop-in gateway works with existing Claude code. - What compatibility means in practice - How the Messages API shape works - When a drop-in Claude API helps ### Claude Code API Key — How to Get and Set Up https://claudeapikey.dev/learn/claude-code-api-key-setup/ Step-by-step guide to getting a Claude Code API key, setting ANTHROPIC_BASE_URL, verifying with /status, and claiming $5 free credit with WELCOME5. - What Is a Claude Code API Key - Step 1 — Register and Generate Your Key - Step 2 — Set Environment Variables ### Claude Code API Costs — Pricing and How to Cut Them https://claudeapikey.dev/learn/claude-code-api-costs/ Understand real Claude Code API costs, why token usage grows quadratically, and how unlimited flat-rate plans eliminate unpredictable billing. - Real Anthropic API Pricing for Claude Models - Why Claude Code Burns Tokens Faster Than Chat - Cost Comparison: Pay-As-You-Go vs Max Subscription vs Unlimited Gateway ### Claude API Errors — Rate Limits, 400, 500, 529 Fixes https://claudeapikey.dev/learn/claude-api-errors-troubleshooting/ Fix Claude API errors: 400 bad request, 429 rate limit reached, 500 server error, and 529 overloaded_error. Real causes, retry logic, and how a gateway reduces them. - Error 400 — Bad Request - Error 429 — Rate Limit Reached - Error 500 — Internal Server Error ### How to Buy Claude API Credits https://claudeapikey.dev/learn/buy-claude-api-credits/ Step-by-step guide to buying Claude API credits — registration, payment methods (card and crypto), WELCOME5 promo, and when to choose unlimited instead. - Step 1 — Create Your Account - Step 2 — Add Balance (Pay-As-You-Go Credits) - Step 4 — Generate an API Key ### Unlimited Claude Usage — Plans and Pricing https://claudeapikey.dev/learn/unlimited-claude-usage-plan/ Unlimited Claude usage plans explained — 1-day and 1-week flat-rate pricing, what unlimited means, included models, how to extend, and comparison to Anthropic limits. - Available Plans and Pricing - What 'Unlimited' Actually Means - What Models Are Included ### Unlimited Claude Opus 4.8 Access https://claudeapikey.dev/learn/unlimited-claude-opus/ Get unlimited access to Claude Opus 4.8 — the strongest reasoning model with 1M context. Why Opus is expensive, who needs it, and how flat-rate access removes the cost barrier. - Why Opus Is Expensive — And Worth It - Who Needs Unlimited Opus Specifically - How to Access Unlimited Opus ### Claude API Free Trial — Get $5 Free Credit https://claudeapikey.dev/learn/claude-api-free-trial/ Get $5 free Claude API credit with promo code WELCOME5 — enough for ~1M Sonnet tokens. Honest answer on free API access and the upgrade path to unlimited. - The Honest Answer About Free Claude API Access - How to Claim the $5 Free Credit - What You Can Do With $5 of Claude API Credit ### Claude API vs Claude Pro/Max Subscription — Which to Choose https://claudeapikey.dev/learn/claude-api-vs-subscription/ Honest comparison of Claude API access vs Claude Pro and Max subscriptions — features, limits, pricing, and which fits developers vs casual users. - What Claude Pro and Max Subscriptions Include - What Raw API Access Provides Instead - Can You Use Claude Code on Both? ### Unlimited Claude Code Usage — No Token Limits https://claudeapikey.dev/learn/unlimited-claude-code-usage/ Remove token limits from Claude Code with flat-rate unlimited access. Why official plans hit limits fast, setup in 2 env vars, and workflows that become possible without caps. - Why Claude Code Hits Limits Faster Than Anything Else - How Official Plans Fall Short for Heavy Claude Code - Setup: Two Environment Variables ### How to Get Unlimited Claude Usage https://claudeapikey.dev/learn/get-unlimited-claude-usage/ The honest breakdown of all ways to get unlimited Claude usage — Max subscription limits, flat-rate gateways, why free unlimited does not exist, and step-by-step setup. - All the Ways to Get 'Unlimited' Claude — Ranked by Honesty - Why Free Unlimited Claude Does Not Exist - Common Reddit Questions Answered Honestly ### How to Get an Unlimited Claude API Key https://claudeapikey.dev/learn/unlimited-claude-api-key/ Get a flat-rate unlimited Claude API key with no per-token billing. How it works, what unlimited means in practice, supported models, fair-use rate limits, setup steps, and common questions about unlimited Claude access. - What Unlimited Actually Means - Fair-Use Rate Limits - Supported Models and Capabilities ### Using Claude in Codex https://claudeapikey.dev/learn/use-claude-api-with-codex/ Buy a Claude API key and run OpenAI's Codex CLI on Claude in minutes — config.toml provider, wire protocol, model id, no Anthropic account needed. - Get a key, then add the provider - Base URL and wire_api - Why buy a key for Codex ### Connecting Claude to n8n https://claudeapikey.dev/learn/use-claude-api-with-n8n/ Buy a Claude API key and wire it into n8n's Anthropic node in minutes — Base URL field, the '/v1/models' fix, AI Agent workflows, no Anthropic account. - Buy a key and build the credential - The '/v1/models' credential fix - Why a bought key suits n8n ### Using LiteLLM with a Claude Key https://claudeapikey.dev/learn/use-claude-api-with-litellm/ Buy a Claude API key and route LiteLLM to it — model_list prefix, api_base, the suffix gotcha, one OpenAI-compatible endpoint for all your apps. - Buy a key, configure model_list - The suffix gotcha - Why route a bought key through LiteLLM ### Using Claude in Kilo Code https://claudeapikey.dev/learn/use-claude-api-with-kilo-code/ Buy a Claude API key and drive the Kilo Code VS Code extension on Claude — provider choice, custom base URL, model id, no Anthropic account or waitlist. - Get a key and pick the provider - Model id and context - Why buy a key for Kilo ### Claude in GitHub Copilot — No Anthropic Account Needed https://claudeapikey.dev/learn/use-claude-api-with-github-copilot/ Drop Claude into GitHub Copilot using a purchased key from claudeapikey.dev — no Anthropic signup or waitlist. Custom Endpoint, Messages API type, and WELCOME5 for $5 of free credit to test it. - Get a key first - Open Copilot's Custom Endpoint - Set the type, URL and key ### Claude in Zed — Instant Key, No Waitlist https://claudeapikey.dev/learn/use-claude-api-with-zed/ Set up the Zed editor with a purchased Claude key from claudeapikey.dev — no Anthropic account required. Use the openai_compatible block (Zed's native provider has no api_url), and start with WELCOME5 for $5 free credit. - Buy the key, skip the account - Use openai_compatible, not the native provider - Point it and list a model ### Claude in Cherry Studio — Buy a Key, Start Chatting https://claudeapikey.dev/learn/use-claude-api-with-cherry-studio/ Add Claude to Cherry Studio with a purchased key from claudeapikey.dev — no Anthropic account or waitlist. Anthropic-type provider, API Host at the host root, the # lock, and WELCOME5 for $5 free credit. - Start with a bought key - Add an Anthropic-type provider - Set the API Host right ### Claude in LangChain — Just Add a Key https://claudeapikey.dev/learn/use-claude-api-with-langchain/ Build LangChain apps on Claude with a purchased key from claudeapikey.dev — no Anthropic account or waitlist. ChatAnthropic base_url at the root, the env route, the ChatOpenAI /v1 variant, and WELCOME5 for $5 free credit. - Buy a key, start building - ChatAnthropic with your key - Or configure from the environment ### Factory Droid on Claude — No Anthropic Account https://claudeapikey.dev/learn/use-claude-api-with-factory-droid/ Run Factory's Droid CLI on Claude with a purchased key from claudeapikey.dev — no Anthropic signup or waitlist. customModels in settings.json, provider anthropic, ${VAR} apiKey, and WELCOME5 for $5 free credit. - Buy the key, skip onboarding - Add it to customModels - Reference the key safely # ClaudeAPIKey.dev — Documentation # Chat Completions API > POST /v1/chat/completions — the OpenAI-format endpoint, for Cursor, LangChain, LiteLLM and the OpenAI SDKs. _Source: https://claudeapikey.dev/docs/api-reference/chat-completions/ · Home > Docs > API reference_ The OpenAI-compatible surface. Same models, same prices, different envelope. - **Endpoint:** `POST https://claudeapikey.dev/v1/chat/completions` - **Base URL for SDKs:** `https://claudeapikey.dev/v1` — note the /v1 - **Auth:** `Authorization: Bearer sk-...` ## Request body | Field | Type | Required | Description | |---|---|---|---| | `model` | string | yes | Any id from `/v1/models`, including Claude ids | | `messages` | array | yes | Supports system, user and assistant roles | | `max_tokens` | integer | no | Optional here, unlike the Messages API | | `temperature` | number | no | 0–2 in the OpenAI convention | | `stream` | boolean | no | true for SSE chunks | | `tools` | array | no | OpenAI function-calling shape | > Unlike Messages, `system` **is** a role here — put it as the first entry of `messages`. Do not send a top-level `system` field to this endpoint. ## Example ```bash curl https://claudeapikey.dev/v1/chat/completions \ -H "Authorization: Bearer $CLAUDEAPIKEY" \ -H "content-type: application/json" \ -d '{ "model": "claude-sonnet-4-6", "messages": [{"role": "user", "content": "Hello"}] }' ``` ## Response ```json { "id": "chatcmpl-...", "object": "chat.completion", "model": "claude-sonnet-4-6", "choices": [{ "index": 0, "message": {"role": "assistant", "content": "..."}, "finish_reason": "stop" }], "usage": {"prompt_tokens": 9, "completion_tokens": 12, "total_tokens": 21} } ``` ## Using the OpenAI SDK ```python from openai import OpenAI client = OpenAI( api_key="sk-your-key", base_url="https://claudeapikey.dev/v1", # the /v1 matters ) resp = client.chat.completions.create( model="claude-sonnet-4-6", messages=[{"role": "user", "content": "Hello"}], ) print(resp.choices[0].message.content) ``` ## Field mapping between the two formats | Concept | Messages | Chat Completions | |---|---|---| | System prompt | top-level `system` | `messages[0]` with role system | | Reply text | `content[0].text` | `choices[0].message.content` | | Why it stopped | `stop_reason` | `finish_reason` | | Input tokens | `usage.input_tokens` | `usage.prompt_tokens` | | Output tokens | `usage.output_tokens` | `usage.completion_tokens` | | Stream delta | `content_block_delta` | `choices[0].delta.content` | - [Cursor](https://claudeapikey.dev/docs/guides/cursor/) — Configure Cursor against this endpoint - [Migrating from OpenAI](https://claudeapikey.dev/docs/guides/migration-from-openai/) — Switch an existing codebase - [Streaming](https://claudeapikey.dev/docs/api-reference/streaming/) — SSE for both formats --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Errors > Every status code the gateway returns, what actually causes it, and the fix. _Source: https://claudeapikey.dev/docs/api-reference/errors/ · Home > Docs > API reference_ Errors return JSON with a `type` and a human-readable `message`. Status codes follow the Anthropic and OpenAI conventions. ```json {"type": "error", "error": {"type": "authentication_error", "message": "invalid api key"}} ``` ## Status codes | Code | Meaning | Most common real cause | Fix | |---|---|---|---| | `400` | Bad request | Missing `max_tokens` on Messages, or a malformed `messages` array | Check the body against [Messages](/docs/api-reference/messages/) | | `401` | Unauthenticated | Key sent in the wrong header for the format | `x-api-key` for Messages, Bearer for Chat Completions | | `402` | Payment required | Balance exhausted or plan expired | [Top up](/topup/) or [buy a plan](/plans/) | | `403` | Forbidden | Key revoked, or model not enabled for the account | Create a new key; check `/v1/models` | | `404` | Not found | Base-URL mistake — usually a doubled or missing `/v1` | See [Base URLs](/docs/getting-started/base-urls/) | | `429` | Rate limited | Too many concurrent requests | Back off exponentially; see [Rate limits](/docs/api-reference/rate-limits/) | | `500` | Gateway error | Unexpected internal failure | Retry once; if it persists, contact support with the request id | | `529` | Upstream overloaded | Provider capacity, not your account | Retry with backoff — usually clears in seconds | ## Error types | type | Meaning | |---|---| | `invalid_request_error` | The body is malformed or a required field is missing | | `authentication_error` | Key missing, malformed, or revoked | | `permission_error` | Key is valid but not allowed to use that model | | `not_found_error` | Unknown path or unknown model id | | `rate_limit_error` | Throttled | | `api_error` | Internal failure | | `overloaded_error` | Upstream capacity exhausted | ## Retrying correctly Retry `429`, `500`, `502`, `503` and `529`. Never blind-retry `400`, `401`, `402` or `403` — the same request fails identically, and for partially generated responses you may pay twice. ```python import time, httpx def call(body, key, tries=4): for i in range(tries): r = httpx.post("https://claudeapikey.dev/v1/messages", json=body, timeout=120, headers={"x-api-key": key, "anthropic-version": "2023-06-01"}) if r.status_code in (429, 500, 502, 503, 529): time.sleep(2 ** i) # 1s, 2s, 4s, 8s continue r.raise_for_status() return r.json() raise RuntimeError("upstream unavailable after retries") ``` > Add jitter to the sleep in production. Synchronised retries from many workers turn one blip into a thundering herd against the same upstream. - [Rate limits](https://claudeapikey.dev/docs/api-reference/rate-limits/) — Concurrency and throughput - [Troubleshooting](https://claudeapikey.dev/docs/guides/troubleshooting/) — Symptom-first index --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Messages API > POST /v1/messages — the Anthropic-format endpoint: parameters, response shape and examples. _Source: https://claudeapikey.dev/docs/api-reference/messages/ · Home > Docs > API reference_ The endpoint every Anthropic-compatible client uses, including Claude Code and the official SDKs. - **Endpoint:** `POST https://claudeapikey.dev/v1/messages` - **Auth:** `x-api-key: sk-...` - **Version header:** `anthropic-version: 2023-06-01` - **Content type:** `application/json` ## Request body | Field | Type | Required | Description | |---|---|---|---| | `model` | string | yes | Model id, exactly as listed by `/v1/models` | | `messages` | array | yes | Alternating user / assistant turns, oldest first | | `max_tokens` | integer | yes | Maximum output tokens. Generation stops here. | | `system` | string or array | no | Top-level system prompt. Not a message role. | | `temperature` | number | no | 0–1. Omit for the model default. | | `top_p` | number | no | Nucleus sampling. Use this or temperature, not both. | | `stop_sequences` | array | no | Strings that halt generation when produced | | `stream` | boolean | no | true for server-sent events | | `tools` | array | no | Tool definitions — see [Tool use](/docs/api-reference/tool-use/) | | `metadata` | object | no | Passed through; `user_id` is useful for your own attribution | ## Example ```bash curl https://claudeapikey.dev/v1/messages \ -H "x-api-key: $CLAUDEAPIKEY" \ -H "anthropic-version: 2023-06-01" \ -H "content-type: application/json" \ -d '{ "model": "claude-sonnet-4-6", "max_tokens": 1024, "messages": [{"role": "user", "content": "Hello"}] }' ``` ## Response ```json { "id": "msg_01ABC...", "type": "message", "role": "assistant", "model": "claude-sonnet-4-6", "content": [{"type": "text", "text": "..."}], "stop_reason": "end_turn", "stop_sequence": null, "usage": {"input_tokens": 9, "output_tokens": 12} } ``` | Field | Description | |---|---| | `content` | Array of blocks. Text blocks carry `type: text` and a `text` field; tool calls carry `type: tool_use`. | | `stop_reason` | `end_turn`, `max_tokens`, `stop_sequence` or `tool_use` | | `usage.input_tokens` | Billed input, including the full resent history | | `usage.output_tokens` | Billed output — typically 4–5× the input rate | ## System prompts ```json { "model": "claude-sonnet-4-6", "max_tokens": 1024, "system": "You are a terse assistant. Answer in at most two sentences.", "messages": [{"role": "user", "content": "Explain HTTP caching."}] } ``` > A stable `system` block is the ideal target for [prompt caching](/docs/guides/prompt-caching/) — it is identical on every call, so it can be read from cache at a fraction of the input price. ## Multimodal input Image blocks use the standard Anthropic shape — see [Vision](/docs/api-reference/vision/). - [Streaming](https://claudeapikey.dev/docs/api-reference/streaming/) — Token-by-token output - [Tool use](https://claudeapikey.dev/docs/api-reference/tool-use/) — Function calling - [Errors](https://claudeapikey.dev/docs/api-reference/errors/) — Failure modes --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Models endpoint > GET /v1/models — discover exactly which model ids your key can call. _Source: https://claudeapikey.dev/docs/api-reference/models-endpoint/ · Home > Docs > API reference_ The authoritative list. If an id is not here, your key cannot call it — do not guess ids from documentation or blog posts. ```bash curl https://claudeapikey.dev/v1/models \ -H "Authorization: Bearer $CLAUDEAPIKEY" ``` ## Response ```json { "object": "list", "data": [ {"id": "claude-opus-4-8", "object": "model", "owned_by": "anthropic"}, {"id": "claude-sonnet-4-6", "object": "model", "owned_by": "anthropic"}, {"id": "gpt-5.5", "object": "model", "owned_by": "openai"} ] } ``` ## Practical uses - **Health check.** A 200 proves base URL and key are both correct before you debug anything else. - **Populating a model picker.** Read ids at startup instead of hardcoding a list that goes stale. - **Catching typos.** Validate a configured model id against this list and fail loudly at boot rather than on the first user request. ```python import os, httpx r = httpx.get( "https://claudeapikey.dev/v1/models", headers={"Authorization": f"Bearer {os.environ['CLAUDEAPIKEY']}"}, timeout=20, ) r.raise_for_status() ids = [m["id"] for m in r.json()["data"]] assert "claude-sonnet-4-6" in ids, "configured model not available on this key" ``` - [Models](https://claudeapikey.dev/docs/getting-started/models/) — Prices and context windows - [Errors](https://claudeapikey.dev/docs/api-reference/errors/) — What model_not_found means --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # API overview > Every endpoint the gateway exposes, which format it speaks, and what it returns. _Source: https://claudeapikey.dev/docs/api-reference/overview/ · Home > Docs > API reference_ Three endpoints. Two request formats. One key. | Method | Path | Format | Purpose | |---|---|---|---| | POST | `/v1/messages` | Anthropic Messages | Chat completion, Anthropic envelope | | POST | `/v1/chat/completions` | OpenAI Chat Completions | Chat completion, OpenAI envelope | | GET | `/v1/models` | OpenAI-style list | Model ids your key can call | A machine-readable description is published at [/openapi.json](/openapi.json), and a plain-text summary of the whole site at [/llms.txt](/llms.txt) and [/llms-full.txt](/llms-full.txt). ## Choosing an endpoint They front the **same models at the same price**. Pick by what your client already emits — do not translate formats yourself. | Endpoint | Auth header | Response shape | Streaming | |---|---|---|---| | `/v1/messages` | `x-api-key` or Bearer | `content[]` array + `usage` | SSE, `content_block_delta` events | | `/v1/chat/completions` | Bearer | `choices[].message.content` | SSE, `choices[].delta.content` chunks | > The streaming shapes differ and are not interchangeable. A client written for `choices[].delta.content` will read nothing at all from a Messages stream. See [Streaming](/docs/api-reference/streaming/). ## Common headers | Header | Applies to | Notes | |---|---|---| | `content-type: application/json` | both POSTs | Required | | `x-api-key` | Messages | Anthropic style | | `Authorization: Bearer` | both | OpenAI style; also accepted by Messages | | `anthropic-version: 2023-06-01` | Messages | Set automatically by Anthropic SDKs | | `accept: text/event-stream` | both | Set by SDKs when `stream` is true | - [Messages](https://claudeapikey.dev/docs/api-reference/messages/) — Anthropic format reference - [Chat Completions](https://claudeapikey.dev/docs/api-reference/chat-completions/) — OpenAI format reference - [Models endpoint](https://claudeapikey.dev/docs/api-reference/models-endpoint/) — Discovering ids - [Errors](https://claudeapikey.dev/docs/api-reference/errors/) — Status codes --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Rate limits > How throttling works, what a 429 means here, and how to build a client that does not trigger one. _Source: https://claudeapikey.dev/docs/api-reference/rate-limits/ · Home > Docs > API reference_ Limits keep shared upstream capacity fair. They apply per account, not per key, so adding keys does not raise your ceiling. ## What is limited | Dimension | Behaviour | |---|---| | Concurrent requests | In-flight requests per account. The usual cause of a 429. | | Requests per minute | Sustained call rate. | | Tokens per minute | Combined input and output throughput. | Unlimited plans apply fair-use limits rather than a credit balance — they remove the per-token charge, not the concurrency ceiling. ## Handling 429 properly 1. Back off exponentially with jitter — 1s, 2s, 4s, 8s, each plus or minus a random fraction. 2. Cap concurrency client-side with a semaphore instead of firing everything and retrying the rejects. 3. Queue batch work rather than parallelising it maximally; throughput is bounded by tokens per minute, not by how many sockets you open. 4. Treat `529` differently from `429` — that is upstream capacity, not you, and usually clears within seconds. ```python import asyncio sem = asyncio.Semaphore(8) # ceiling on concurrency, tune to your plan async def ask(client, body): async with sem: return await client.post("/v1/messages", json=body) ``` > Agentic tools such as Claude Code fan out aggressively by design. If you see 429s during heavy agent use, lower the tool's parallelism before raising your plan. ## Reducing pressure instead of retrying - Route cheap work to `claude-haiku-4-5` — smaller models finish sooner and hold a slot for less time. - Turn on [prompt caching](/docs/guides/prompt-caching/): cached prefixes cut input tokens and therefore token-per-minute pressure. - Trim conversation history; see [Context management](/docs/guides/context-management/). - Batch offline work into off-peak windows. - [Errors](https://claudeapikey.dev/docs/api-reference/errors/) — Status code reference - [Cost control](https://claudeapikey.dev/docs/guides/cost-control/) — Spend less per task - [Plans](https://claudeapikey.dev/docs/billing/plans/) — Flat-rate options --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Streaming > Server-sent events in both formats — the event shapes differ, and mixing them up reads as silence. _Source: https://claudeapikey.dev/docs/api-reference/streaming/ · Home > Docs > API reference_ Set `stream` to true and the response becomes `text/event-stream`. Both endpoints stream, but they emit **different event shapes**. ## Anthropic format ```bash curl -N https://claudeapikey.dev/v1/messages \ -H "x-api-key: $CLAUDEAPIKEY" \ -H "anthropic-version: 2023-06-01" \ -H "content-type: application/json" \ -d '{"model":"claude-sonnet-4-6","max_tokens":512,"stream":true, "messages":[{"role":"user","content":"Count to five"}]}' ``` ```json event: message_start data: {"type":"message_start","message":{"id":"msg_...","usage":{"input_tokens":9}}} event: content_block_delta data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"One"}} event: message_delta data: {"type":"message_delta","delta":{"stop_reason":"end_turn"},"usage":{"output_tokens":14}} event: message_stop data: {"type":"message_stop"} ``` Text arrives in `content_block_delta` events at `delta.text`. Final output token counts arrive in `message_delta`, not `message_start`. ## OpenAI format ```json data: {"choices":[{"delta":{"content":"One"},"index":0}]} data: {"choices":[{"delta":{},"finish_reason":"stop","index":0}]} data: [DONE] ``` Text arrives at `choices[0].delta.content`, and the stream ends with a literal `data: [DONE]` sentinel that is not JSON. > This mismatch produces an empty response with no error: a client parsing `choices[].delta.content` against a Messages stream finds nothing, every time, and reports success. If streaming appears to work but returns nothing, check which shape you are parsing. ## Python, Anthropic SDK ```python from anthropic import Anthropic client = Anthropic(api_key="sk-your-key", base_url="https://claudeapikey.dev") with client.messages.stream( model="claude-sonnet-4-6", max_tokens=512, messages=[{"role": "user", "content": "Count to five"}], ) as stream: for text in stream.text_stream: print(text, end="", flush=True) ``` ## Operational notes - Use `curl -N` when testing by hand, or curl buffers the whole stream and it looks like nothing is happening. - Disable proxy buffering in front of your own app (`proxy_buffering off` in nginx) or clients see the reply arrive in one lump at the end. - A dropped connection mid-stream still bills the tokens already generated. - Streaming does not change the price — only when the bytes arrive. - [Messages API](https://claudeapikey.dev/docs/api-reference/messages/) — Non-streaming reference - [Errors](https://claudeapikey.dev/docs/api-reference/errors/) — Failures mid-stream - [Troubleshooting](https://claudeapikey.dev/docs/guides/troubleshooting/) — Empty responses --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Tool use and function calling > Let a model call your functions — the request loop, both formats, and the mistakes that cost tokens. _Source: https://claudeapikey.dev/docs/api-reference/tool-use/ · Home > Docs > API reference_ Tool use lets the model ask you to run a function and then continue with the result. It is the mechanism behind every coding agent. ## The loop 1. You send `messages` plus `tools` describing what is available. 2. The model replies with `stop_reason: tool_use` and a `tool_use` block naming the tool and its arguments. 3. **You** execute it — the gateway never runs your code. 4. You send the whole history back plus a `tool_result` block. 5. The model produces its final answer, or asks for another tool. ## Defining tools (Messages format) ```json { "model": "claude-sonnet-4-6", "max_tokens": 1024, "tools": [{ "name": "get_weather", "description": "Current weather for a city. Use when the user asks about weather.", "input_schema": { "type": "object", "properties": {"city": {"type": "string", "description": "City name"}}, "required": ["city"] } }], "messages": [{"role": "user", "content": "Weather in Paris?"}] } ``` ## Returning a result ```json { "role": "user", "content": [{ "type": "tool_result", "tool_use_id": "toolu_01ABC...", "content": "18C, light rain" }] } ``` > `tool_result` goes in a message with role `user`, and `tool_use_id` must match the id from the model's request exactly. Mismatched ids are rejected as an invalid request. ## OpenAI format The `/v1/chat/completions` endpoint takes the OpenAI `tools` / `tool_calls` shape instead, with results returned as messages with role `tool`. ## What tool use costs - **Definitions are resent every turn.** Twenty verbose tool schemas can dominate your input tokens before the conversation even starts. - **Results become context.** A tool returning 200 rows of JSON puts all of it in the next request, and every request after that. - **Each round trip is a billed call.** A five-tool task is at least six requests. - Cap tool output length and summarise before returning. Truncating a result to what the model actually needs is the highest-leverage change in most agent loops. > Tool schemas are stable across turns, which makes them a strong [prompt caching](/docs/guides/prompt-caching/) candidate alongside the system prompt. - [Messages API](https://claudeapikey.dev/docs/api-reference/messages/) — Base reference - [Cost control](https://claudeapikey.dev/docs/guides/cost-control/) — Where agent tokens go - [MCP servers](https://claudeapikey.dev/docs/guides/mcp-servers/) — Ready-made tools for Claude Code --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Vision and image input > Send images to multimodal models — encoding, size limits, and what it costs. _Source: https://claudeapikey.dev/docs/api-reference/vision/ · Home > Docs > API reference_ Claude and Gemini models accept images alongside text. Images are content blocks inside an ordinary message. ## Base64 image ```json { "model": "claude-sonnet-4-6", "max_tokens": 1024, "messages": [{ "role": "user", "content": [ {"type": "image", "source": { "type": "base64", "media_type": "image/png", "data": "iVBORw0KGgoAAAANSUhEUg..." }}, {"type": "text", "text": "What does this chart show?"} ] }] } ``` ## Python ```python import base64, anthropic img = base64.standard_b64encode(open("chart.png", "rb").read()).decode() client = anthropic.Anthropic(api_key="sk-your-key", base_url="https://claudeapikey.dev") msg = client.messages.create( model="claude-sonnet-4-6", max_tokens=1024, messages=[{"role": "user", "content": [ {"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": img}}, {"type": "text", "text": "What does this chart show?"}, ]}], ) print(msg.content[0].text) ``` ## Practical limits | Constraint | Guidance | |---|---| | Formats | PNG, JPEG, GIF and WebP | | Resolution | Very large images are downscaled upstream. Resizing to roughly 1500px on the long edge before sending saves tokens with no quality loss. | | Cost | Images bill as input tokens, roughly proportional to area. A full-page screenshot can cost more than a page of text. | | Multiple images | Several image blocks per message are allowed; each one is billed. | > In an agent loop, images are resent with the history on every turn like any other content. Drop them from the transcript once they have been described, or one screenshot gets paid for repeatedly. - [Messages API](https://claudeapikey.dev/docs/api-reference/messages/) — Content block reference - [Models](https://claudeapikey.dev/docs/getting-started/models/) — Which models are multimodal --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Credits > How the prepaid balance works, what draws it down, and how to avoid running dry mid-run. _Source: https://claudeapikey.dev/docs/billing/credits/ · Home > Docs > Billing_ Credits are a prepaid balance spent per token. Every request deducts input and output tokens at that model's rate. - **Rate:** **$10 = $100** in credits - **Expiry:** None - **Shared across models:** Yes — one balance for Claude, GPT, Gemini and MiniMax - **Runs out:** Requests return `402` until you top up ## What draws the balance down - **Input tokens** — the whole request: system prompt, full history, tool schemas, images. - **Output tokens** — what the model generates, at 4–5× the input rate. - Nothing else. There is no per-request fee and no monthly minimum. ## Reading your usage Every response carries a `usage` object with exact token counts — that is the billing basis, and it is worth logging alongside your own request ids. The dashboard aggregates the same data per model and per key. ## Not running out mid-task - Long agent runs consume credits faster than chat because history is resent every turn. - A `402` mid-session stops work immediately — top up before a long run rather than during one. - If usage is steady and heavy, a [plan](/docs/billing/plans/) is usually cheaper than credits. > Crypto top-ups currently carry +30%, and a new account's first crypto top-up carries +100%. Both are applied automatically at checkout. - [Pricing](https://claudeapikey.dev/docs/billing/pricing/) — Rates and conversion - [Top up](https://claudeapikey.dev/topup/) — Add credits - [Cost control](https://claudeapikey.dev/docs/guides/cost-control/) — Make credits last --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Payment methods > Card and crypto — which to use, and why the bonus differs. _Source: https://claudeapikey.dev/docs/billing/payment-methods/ · Home > Docs > Billing_ | Method | Bonus | Notes | |---|---|---| | Bank card | standard rate | Instant. Some issuers decline cross-border AI purchases. | | Crypto (USDT, BTC, ETH) | +30% | Works everywhere; credited on confirmation. | New accounts additionally receive **+100%** on their first crypto top-up. ## Why crypto carries a bonus Card processing carries fees and chargeback exposure that crypto does not. The bonus passes part of that difference back rather than pricing it into everyone's rate. ## If a card is declined 1. The issuer may have blocked it as an international or high-risk merchant transaction — a bank app approval often clears it. 2. Try another card, or switch to crypto. 3. Contact [support@claudeapikey.dev](mailto:support@claudeapikey.dev) with the approximate time. Never send full card details by email. - [Top up](https://claudeapikey.dev/topup/) — Add credits now - [Pricing](https://claudeapikey.dev/docs/billing/pricing/) — Conversion and rates - [Availability](https://claudeapikey.dev/docs/resources/supported-countries/) — Regional notes --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Unlimited plans > Flat-rate access for a fixed window — what they include and when they beat credits. _Source: https://claudeapikey.dev/docs/billing/plans/ · Home > Docs > Billing_ A plan replaces per-token billing for a fixed period. No token arithmetic, no balance to watch — under fair-use limits. | Plan | Price | Duration | Per day | What it is | |---|---|---|---|---| | Unlimited 1 Hour | $1 | 1 hour | $24.00 | One hour of unlimited access | | Unlimited 24 Hours | $10 | 1 day | $10.00 | A full day of unlimited access | | Unlimited 1 Week | $49 | 7 days | $7.00 | A full week of unlimited access | | Unlimited 15 Days | $89 | 15 days | $5.93 | 15 days of unlimited access | ## What is included - Unlimited requests to the Claude family — Fable 5, Opus, Sonnet and Haiku — for the window. - No per-token charge and no balance drawdown while the plan is active. - Works with every client: Claude Code, Cursor, Cline and the SDKs. - Instant activation, no auto-renewal. Access ends when the plan expires. ## Plan or credits? | Your usage | Cheaper option | Why | |---|---|---| | A few requests a day | Credits | You pay only for what you use, and the balance does not expire | | Bursty — heavy some days, idle others | Credits | A plan's clock runs whether you use it or not | | Agent running for hours daily | Plan | Per-token billing on long sessions outruns the flat rate quickly | | Predictable heavy load | Plan | Fixed, known cost | > Fair-use limits still apply on a plan — it removes the per-token charge, not the concurrency ceiling. See [Rate limits](/docs/api-reference/rate-limits/). ## Buying Plans are listed on the public [plans page](/plans/) and purchased from the dashboard. Activation is immediate. - [Plans page](https://claudeapikey.dev/plans/) — Buy a plan - [Pricing](https://claudeapikey.dev/docs/billing/pricing/) — Compare with credits - [Rate limits](https://claudeapikey.dev/docs/api-reference/rate-limits/) — Fair-use ceilings --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Pricing > How credits convert to tokens, what the bonuses are, and how to work out the cost of a task. _Source: https://claudeapikey.dev/docs/billing/pricing/ · Home > Docs > Billing_ Two ways to pay: prepaid **credits** spent at per-token rates, or a flat-rate **unlimited plan** for a fixed window. ## Credits - **Conversion:** **$10 = $100** in API credits (10× multiplier) - **Crypto bonus:** +30% extra credits on every crypto top-up - **First top-up:** +100% on a new account's first crypto top-up - **Expiry:** Credits do not expire | You pay | Credits | With the +30% crypto bonus | |---|---|---| | $5 | $50 | $65 | | $10 | $100 | $130 | | $25 | $250 | $325 | | $50 | $500 | $650 | | $100 | $1000 | $1300 | > Bank card top-ups are credited at the standard rate; the bonus percentages above apply to crypto. Live figures are always shown on the [top-up page](/topup/) — this table is generated from the same configuration. ## Per-token rates Credits are spent at these rates, per million tokens: | Model | API id | Context | $/M in | $/M out | Saving | |---|---|---|---|---|---| | Claude Opus 4.8 | `claude-opus-4-8` | 200K | $3 | $15 | 80% | | Claude Fable 5 | `claude-fable-5` | 200K / 1M | $3 | $15 | 80% | | Claude Sonnet 4.6 | `claude-sonnet-4-6` | 200K / 1M | $0.9 | $4.5 | 70% | | Claude Haiku 4.5 | `claude-haiku-4-5` | 200K | $0.4 | $2 | 60% | | GPT-5.5 | `gpt-5.5` | 256K | $3 | $12 | 70% | | GPT-5 | `gpt-5` | 256K | $1.5 | $6 | 70% | | Gemini 3 Pro | `gemini-3-pro` | 1M | $1 | $6 | 60% | | Gemini 3 Flash | `gemini-3-flash` | 1M | $0.2 | $1.2 | 60% | | MiniMax M3 | `minimax-m3` | 1M | $0.5 | $2.5 | 58% | ## Working out a task's cost Cost is `(input ÷ 1M × in-rate) + (output ÷ 1M × out-rate)`. A 20K-in / 2K-out request: | Model | Input | Output | Total | |---|---|---|---| | `claude-opus-4-8` | $0.060 | $0.030 | **$0.090** | | `claude-sonnet-4-6` | $0.018 | $0.009 | **$0.027** | | `claude-haiku-4-5` | $0.008 | $0.004 | **$0.012** | Same task, 7.5× the price on Opus. If it is classification, that multiple buys nothing — see [Cost control](/docs/guides/cost-control/). ## Unlimited plans | Plan | Price | Duration | Per day | What it is | |---|---|---|---|---| | Unlimited 1 Hour | $1 | 1 hour | $24.00 | One hour of unlimited access | | Unlimited 24 Hours | $10 | 1 day | $10.00 | A full day of unlimited access | | Unlimited 1 Week | $49 | 7 days | $7.00 | A full week of unlimited access | | Unlimited 15 Days | $89 | 15 days | $5.93 | 15 days of unlimited access | Plans remove per-token billing for their window under fair-use limits. See [Plans](/docs/billing/plans/) for which model of billing suits your usage. - [Credits](https://claudeapikey.dev/docs/billing/credits/) — Balance mechanics - [Plans](https://claudeapikey.dev/docs/billing/plans/) — Flat-rate access - [Cost control](https://claudeapikey.dev/docs/guides/cost-control/) — Spend less per task --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Refunds and support > What can be refunded, what cannot, and how to raise a billing issue. _Source: https://claudeapikey.dev/docs/billing/refunds/ · Home > Docs > Billing_ ## Refundable - A duplicate charge. - A failed top-up where credits were never delivered. - A documented service fault that prevented use of a plan for a material part of its window. ## Not refundable - Credits already spent on completed API calls — the tokens were generated and paid upstream. - A plan window that has elapsed while unused. - Output quality not matching expectations. Test with a small top-up or the shortest plan before committing. > The shortest plan exists precisely so you can evaluate the service for a small amount before buying a longer window. ## Raising an issue Email [support@claudeapikey.dev](mailto:support@claudeapikey.dev) with the account email, approximate date and time, amount and payment method. Do not include your API key or full card number. For a technical fault rather than a billing one, include the endpoint, model id, status code and error body — see [Troubleshooting](/docs/guides/troubleshooting/). - [Refund policy](https://claudeapikey.dev/refund) — Full terms - [FAQ](https://claudeapikey.dev/docs/resources/faq/) — Common questions --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Authentication > How to send your key, which header each format expects, and how to keep keys safe. _Source: https://claudeapikey.dev/docs/getting-started/authentication/ · Home > Docs > Getting started_ Every request needs an API key created in your dashboard. Keys start with `sk-` and carry the permissions and balance of the account that made them. ## Two headers, one key The same key works in both formats — only the header name changes. | Format | Endpoint | Header | |---|---|---| | Anthropic | `/v1/messages` | `x-api-key: sk-...` | | Anthropic (alt) | `/v1/messages` | `Authorization: Bearer sk-...` | | OpenAI | `/v1/chat/completions` | `Authorization: Bearer sk-...` | The Anthropic format additionally expects a version header. Any Anthropic SDK sets it automatically: ```bash -H "anthropic-version: 2023-06-01" ``` ## Environment variables Most tooling reads these rather than taking a key as an argument: | Variable | Used by | Value | |---|---|---| | `ANTHROPIC_BASE_URL` | Claude Code, Anthropic SDK | `https://claudeapikey.dev` | | `ANTHROPIC_AUTH_TOKEN` | Claude Code | your `sk-` key | | `ANTHROPIC_API_KEY` | Anthropic SDK (Python/TS) | your `sk-` key | | `OPENAI_BASE_URL` | OpenAI SDK, LangChain, LiteLLM | `https://claudeapikey.dev/v1` | | `OPENAI_API_KEY` | OpenAI SDK | your `sk-` key | > Note the `/v1` suffix on `OPENAI_BASE_URL` but not on `ANTHROPIC_BASE_URL`. That asymmetry is in the official SDKs, not something we introduced — the OpenAI client appends paths to the base you give it, the Anthropic client appends `/v1` itself. Getting this wrong is the single most common setup error. ## Key hygiene - Create a **separate key per machine or project** so one can be revoked without disturbing the others. - Never commit keys. Use environment variables or a secret manager; add `.env` to `.gitignore`. - Keys are shown in full **once**. Rotate rather than trying to recover a lost key. - A leaked key spends your balance. Revoke it in the dashboard immediately — revocation is instant. - Prefer server-side calls. A key shipped in a browser bundle or mobile app is a public key. ## Revoking and rotating 1. Dashboard → **API Keys**. 2. Create the replacement key first and deploy it. 3. Delete the old key once traffic has moved. In-flight requests using it fail immediately after deletion. - [Quickstart](https://claudeapikey.dev/docs/getting-started/quickstart/) — Make the first call - [API key configuration](https://claudeapikey.dev/docs/guides/api-key-configuration/) — Per-tool placement of the key - [Errors](https://claudeapikey.dev/docs/api-reference/errors/) — What a 401 vs 402 means --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Base URLs > The one setting that trips everyone up: which URL each client expects, with and without the /v1 suffix. _Source: https://claudeapikey.dev/docs/getting-started/base-urls/ · Home > Docs > Getting started_ Almost every failed integration is a base-URL mistake. This page is the reference for exactly what to set. ## The rule | Client | Set base URL to | Then it calls | |---|---|---| | Anthropic SDK (Python / TS) | `https://claudeapikey.dev` | `https://claudeapikey.dev/v1/messages` | | Claude Code | `https://claudeapikey.dev` | `https://claudeapikey.dev/v1/messages` | | OpenAI SDK (Python / TS) | `https://claudeapikey.dev/v1` | `https://claudeapikey.dev/v1/chat/completions` | | LangChain / LiteLLM / Cursor | `https://claudeapikey.dev/v1` | `https://claudeapikey.dev/v1/chat/completions` | > **Anthropic clients: no `/v1`. OpenAI clients: with `/v1`.** The Anthropic SDK appends `/v1/messages` to whatever base you give it; the OpenAI SDK appends only `/chat/completions`. Adding `/v1` to an Anthropic base produces requests to `/v1/v1/messages`, which 404s. ## How to tell which mistake you made | You see | Meaning | |---|---| | `404` and the path contains `/v1/v1/` | You added `/v1` to an Anthropic-style base URL. Remove it. | | `404` on `/chat/completions` with no `/v1` | You omitted `/v1` from an OpenAI-style base URL. Add it. | | `404` on `/v1/chat/completions` from an Anthropic tool | The tool speaks Messages, not Chat Completions. Use the Anthropic base and endpoint. | | Connection works but every call 401s | Right URL, wrong auth header for that format. See [Authentication](/docs/getting-started/authentication/). | ## Verifying without any SDK This is the fastest way to prove the URL and key are right before blaming your framework: ```bash # should return JSON with a "data" array of model ids curl -s -o /dev/null -w "%{http_code}\n" https://claudeapikey.dev/v1/models \ -H "Authorization: Bearer $CLAUDEAPIKEY" ``` A `200` means URL and key are both correct, and any remaining problem is in your client configuration. - [Authentication](https://claudeapikey.dev/docs/getting-started/authentication/) — Headers and env vars - [Troubleshooting](https://claudeapikey.dev/docs/guides/troubleshooting/) — Symptom-first debugging - [Overview](https://claudeapikey.dev/docs/api-reference/overview/) — Endpoint reference --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Your first request explained > A line-by-line walkthrough of one Messages call — what every field does and what comes back. _Source: https://claudeapikey.dev/docs/getting-started/first-request/ · Home > Docs > Getting started_ The Quickstart gets you a response. This page explains what each part of it means, so the next request is one you write yourself. ## The request ```bash curl https://claudeapikey.dev/v1/messages \ -H "x-api-key: $CLAUDEAPIKEY" \ -H "anthropic-version: 2023-06-01" \ -H "content-type: application/json" \ -d '{ "model": "claude-sonnet-4-6", "max_tokens": 1024, "messages": [{"role": "user", "content": "Hello"}] }' ``` | Field | Required | What it does | |---|---|---| | `model` | yes | Which model to run. Must match an id from `/v1/models` exactly. | | `max_tokens` | yes | Hard ceiling on the **output**. The model stops here even mid-sentence. This is a cost control, not a target. | | `messages` | yes | The conversation so far, oldest first. Each entry has a `role` (`user` or `assistant`) and `content`. | | `system` | no | Instructions that apply to the whole conversation. A top-level field, not a message with `role: system`. | | `temperature` | no | 0–1. Lower is more deterministic. Leave unset unless you have a reason. | | `stream` | no | `true` returns server-sent events instead of one JSON body. See [Streaming](/docs/api-reference/streaming/). | ## The response ```json { "id": "msg_01ABC...", "type": "message", "role": "assistant", "model": "claude-sonnet-4-6", "content": [{"type": "text", "text": "Hello! How can I help?"}], "stop_reason": "end_turn", "usage": {"input_tokens": 9, "output_tokens": 12} } ``` - **`content` is an array**, not a string. Text lives at `content[0].text`. Treating it as a string is the most common client bug. - **`stop_reason`** tells you why generation ended: `end_turn` (finished), `max_tokens` (hit your ceiling — the reply is truncated), `stop_sequence`, or `tool_use`. - **`usage`** is what you are billed on. Log it; it is the only reliable basis for cost attribution. > If `stop_reason` is `max_tokens`, your answer was cut off. Raise `max_tokens` or ask for a shorter reply — do not retry blindly, you pay for the truncated output too. ## Multi-turn conversations The API is stateless. To continue a conversation you resend the whole history, including the model's previous replies: ```json { "model": "claude-sonnet-4-6", "max_tokens": 1024, "messages": [ {"role": "user", "content": "What is the capital of France?"}, {"role": "assistant", "content": "Paris."}, {"role": "user", "content": "And its population?"} ] } ``` > Because history is resent every turn, input tokens grow with the conversation and so does the bill. This is why long agent sessions get expensive. [Prompt caching](/docs/guides/prompt-caching/) and [context management](/docs/guides/context-management/) both attack this directly. - [Messages API](https://claudeapikey.dev/docs/api-reference/messages/) — Complete field reference - [Streaming](https://claudeapikey.dev/docs/api-reference/streaming/) — Token-by-token responses - [Python SDK](https://claudeapikey.dev/docs/sdks/python/) — The same call in Python --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Introduction > What ClaudeAPIKey.dev is, which endpoints it speaks, and how it differs from calling Anthropic directly. _Source: https://claudeapikey.dev/docs/getting-started/introduction/ · Home > Docs > Getting started_ ClaudeAPIKey.dev is an **Anthropic-compatible API gateway**. You point an existing SDK at one base URL, use one key, and reach Claude, GPT, Gemini and MiniMax models through the request format you already use. - **Base URL:** `https://claudeapikey.dev` - **API formats:** Anthropic Messages **and** OpenAI Chat Completions - **Auth:** `x-api-key` (Anthropic style) or `Authorization: Bearer` (OpenAI style) - **Models:** 9, across Anthropic, OpenAI, Google and MiniMax - **Billing:** Prepaid credits, or a flat-rate unlimited plan - **Setup time:** Two environment variables ## What you get | Area | What it means | |---|---| | Drop-in compatibility | The same request bodies, the same response shapes, the same streaming events as the official APIs. You change the base URL, not your code. | | One key, many vendors | A single credential reaches Claude, GPT, Gemini and MiniMax. No separate accounts or dashboards per provider. | | Both wire formats | `/v1/messages` for Anthropic-style clients, `/v1/chat/completions` for OpenAI-style clients — against the same models. | | Tool compatibility | Claude Code, Cursor, Cline, Aider, Zed, Continue, OpenCode and anything else that honours a custom base URL. | | Pay-as-you-go | Prepaid credits that do not expire, or a flat-rate unlimited plan for a fixed window. | ## How it differs from calling Anthropic directly Functionally it should not differ at all — that is the design goal. The differences that matter are commercial and operational: - **Price.** Published rates run 60–80% below list for the Claude family. See [Pricing](/docs/billing/pricing/). - **No waitlist or org approval.** An account and a top-up are all that is needed. - **Payment options.** Bank card or crypto, including from regions where Anthropic billing is unavailable. - **Routing.** Requests are distributed across upstream capacity with automatic failover, so a single upstream error does not surface as a hard failure. > We are a reseller of API capacity, not Anthropic. If you need a contractual relationship with Anthropic itself, an enterprise SLA, or data-processing terms signed by Anthropic, use the official API. ClaudeAPIKey.dev is an independently operated API gateway. "Claude" and "Anthropic" are trademarks of Anthropic; we are not affiliated with, endorsed by, or sponsored by Anthropic, OpenAI or Google. ## Which format should I use? Use whichever your tool already speaks — both hit the same models and cost the same. | If your tool is… | Use | Auth header | |---|---|---| | Claude Code, Anthropic SDK, Cline, any Claude client | `POST /v1/messages` | `x-api-key` | | Cursor, Continue, OpenAI SDK, LangChain, LiteLLM | `POST /v1/chat/completions` | `Authorization: Bearer` | ## Next steps - [Quickstart](https://claudeapikey.dev/docs/getting-started/quickstart/) — First working request in about a minute - [Authentication](https://claudeapikey.dev/docs/getting-started/authentication/) — Keys, headers and env vars - [Models](https://claudeapikey.dev/docs/getting-started/models/) — Every model id and what it costs - [Claude Code](https://claudeapikey.dev/docs/guides/claude-code/) — The most common setup - [Pricing](https://claudeapikey.dev/docs/billing/pricing/) — Credits, rates and bonuses --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Models > Every model id you can call, its context window, and what it costs per million tokens. _Source: https://claudeapikey.dev/docs/getting-started/models/ · Home > Docs > Getting started_ One key reaches all of these. Use the **API id** column verbatim in the `model` field of a request — display names are not accepted. | Model | API id | Context | $/M in | $/M out | Saving | |---|---|---|---|---|---| | Claude Opus 4.8 | `claude-opus-4-8` | 200K | $3 | $15 | 80% | | Claude Fable 5 | `claude-fable-5` | 200K / 1M | $3 | $15 | 80% | | Claude Sonnet 4.6 | `claude-sonnet-4-6` | 200K / 1M | $0.9 | $4.5 | 70% | | Claude Haiku 4.5 | `claude-haiku-4-5` | 200K | $0.4 | $2 | 60% | | GPT-5.5 | `gpt-5.5` | 256K | $3 | $12 | 70% | | GPT-5 | `gpt-5` | 256K | $1.5 | $6 | 70% | | Gemini 3 Pro | `gemini-3-pro` | 1M | $1 | $6 | 60% | | Gemini 3 Flash | `gemini-3-flash` | 1M | $0.2 | $1.2 | 60% | | MiniMax M3 | `minimax-m3` | 1M | $0.5 | $2.5 | 58% | Prices are per million tokens in credits. Live ids are always available from the models endpoint: ```bash curl https://claudeapikey.dev/v1/models -H "Authorization: Bearer $CLAUDEAPIKEY" ``` ## Choosing a model | Task | Recommended | Why | |---|---|---| | Agentic coding, refactors, architecture | `claude-opus-4-8` | Strongest multi-step reasoning; worth the output price on hard work | | Everyday coding, chat, general work | `claude-sonnet-4-6` | The default — roughly a fifth of Opus's rate at close quality on most tasks | | Classification, extraction, routing | `claude-haiku-4-5` | Cheapest Claude; these tasks do not need a frontier model | | Very large documents | `gemini-3-pro` | 1M context at a low input rate | | High-volume batch | `gemini-3-flash` or `minimax-m3` | Lowest per-token cost in the catalog | > The biggest single cost lever is **routing by task**. Output tokens cost 4–5× input on every frontier model, so sending classification work to Opus wastes money at roughly 7× the Haiku rate. See [Cost control](/docs/guides/cost-control/). ## Long context `claude-sonnet-4-6` also serves a 1M-token variant, billed at a higher rate than the 200K version because long context is more expensive upstream. Gemini models take 1M natively. Sending 500K tokens of context on every turn is rarely the cheapest way to solve a problem — see [Context management](/docs/guides/context-management/). ## Model aliases Ids are pinned, not floating. `claude-sonnet-4-6` always means that version — it will not silently become a different model under you. When we add a successor it gets a new id, and the old one keeps working until formally retired via the [changelog](/docs/resources/changelog/). - [Model catalog](https://claudeapikey.dev/models/) — Per-model pages with benchmarks - [Pricing](https://claudeapikey.dev/docs/billing/pricing/) — How credits convert to tokens - [Cost control](https://claudeapikey.dev/docs/guides/cost-control/) — Cut spend without losing quality --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Quickstart > Get a key, set two environment variables, and make your first request in about a minute. _Source: https://claudeapikey.dev/docs/getting-started/quickstart/ · Home > Docs > Getting started_ This page takes you from nothing to a working response. It assumes only `curl`; SDK versions are in [SDKs](/docs/sdks/). ## Step 1 — Create a key 1. [Register an account](/register) and confirm your email. 2. [Top up credits](/topup/) or [buy an unlimited plan](/plans/). New accounts get a bonus on the first crypto top-up — see [Pricing](/docs/billing/pricing/). 3. Open the dashboard, go to **API Keys**, and create a key. It begins with `sk-`. 4. Copy it immediately — the full value is shown once. ## Step 2 — Export it ```bash export CLAUDEAPIKEY="sk-your-key-here" ``` ## Step 3 — Call the API ### Anthropic format ```bash curl https://claudeapikey.dev/v1/messages \ -H "x-api-key: $CLAUDEAPIKEY" \ -H "anthropic-version: 2023-06-01" \ -H "content-type: application/json" \ -d '{ "model": "claude-sonnet-4-6", "max_tokens": 1024, "messages": [{"role": "user", "content": "Hello"}] }' ``` ### OpenAI format Identical models, different envelope — pick whichever matches your client: ```bash curl https://claudeapikey.dev/v1/chat/completions \ -H "Authorization: Bearer $CLAUDEAPIKEY" \ -H "content-type: application/json" \ -d '{ "model": "claude-sonnet-4-6", "messages": [{"role": "user", "content": "Hello"}] }' ``` ## Step 4 — Confirm what you can reach ```bash curl https://claudeapikey.dev/v1/models \ -H "Authorization: Bearer $CLAUDEAPIKEY" ``` That returns every model id your key can address. Use those ids verbatim in the `model` field. ## Step 5 — Point a real tool at it For Claude Code, two variables are the whole setup: ```bash export ANTHROPIC_BASE_URL="https://claudeapikey.dev" export ANTHROPIC_AUTH_TOKEN="$CLAUDEAPIKEY" claude ``` > If `ANTHROPIC_API_KEY` is already set in your shell, Claude Code may bill that key instead and silently ignore your subscription. Unset it when switching. See [Claude Code](/docs/guides/claude-code/). ## Troubleshooting the first call | Symptom | Cause | Fix | |---|---|---| | `401` | Key missing, mistyped, or sent in the wrong header | Anthropic format uses `x-api-key`; OpenAI format uses `Authorization: Bearer` | | `402` | No credits and no active plan | [Top up](/topup/) or [buy a plan](/plans/) | | `404` on `/v1/chat/completions` | Client is calling the Anthropic path with an OpenAI body, or vice versa | Match the format to the endpoint — see [Overview](/docs/api-reference/overview/) | | `model_not_found` | Model id not spelled exactly as returned by `/v1/models` | Copy the id from the models endpoint | - [Authentication](https://claudeapikey.dev/docs/getting-started/authentication/) — Header and env-var reference - [Messages API](https://claudeapikey.dev/docs/api-reference/messages/) — Full request and response schema - [Errors](https://claudeapikey.dev/docs/api-reference/errors/) — Every status code and what it means --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Aider > Run the terminal pair-programmer against the gateway in either wire format. _Source: https://claudeapikey.dev/docs/guides/aider/ · Home > Docs > Integration guides_ Aider routes model calls through LiteLLM, so it accepts either format depending on how you name the model. ## OpenAI-compatible route ```bash export OPENAI_API_BASE="https://claudeapikey.dev/v1" export OPENAI_API_KEY="sk-your-key" aider --model openai/claude-sonnet-4-6 ``` ## Anthropic route ```bash export ANTHROPIC_API_BASE="https://claudeapikey.dev" export ANTHROPIC_API_KEY="sk-your-key" aider --model anthropic/claude-sonnet-4-6 ``` > The `openai/` or `anthropic/` prefix tells LiteLLM which format to speak. Get it wrong and you will see a 404 as the request goes to the endpoint that does not match the body. ## Recommended flags ```bash # a cheaper model for commit messages and summaries aider --model anthropic/claude-sonnet-4-6 \ --weak-model anthropic/claude-haiku-4-5 ``` Aider makes frequent small auxiliary calls. Sending those to Haiku instead of the main model is a meaningful saving on a long session at no quality cost. - [Models](https://claudeapikey.dev/docs/getting-started/models/) — Ids and prices - [Cost control](https://claudeapikey.dev/docs/guides/cost-control/) — Route by task --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # API key configuration by tool > One table: where each client expects the key and the base URL, with the /v1 suffix spelled out. _Source: https://claudeapikey.dev/docs/guides/api-key-configuration/ · Home > Docs > Integration guides_ A single reference so you do not have to open five guides to remember one suffix. | Tool | Key goes in | Base URL | |---|---|---| | Claude Code | `ANTHROPIC_AUTH_TOKEN` | `https://claudeapikey.dev` | | Anthropic SDK (Python/TS) | `api_key=` or `ANTHROPIC_API_KEY` | `https://claudeapikey.dev` | | OpenAI SDK (Python/TS) | `api_key=` or `OPENAI_API_KEY` | `https://claudeapikey.dev/v1` | | Cursor | Settings → Models → OpenAI key | `https://claudeapikey.dev/v1` | | Cline | Extension settings, Anthropic provider | `https://claudeapikey.dev` | | Continue | `apiKey` in `config.json` | `https://claudeapikey.dev` (anthropic) / `https://claudeapikey.dev/v1` (openai) | | Aider | `ANTHROPIC_API_KEY` / `OPENAI_API_KEY` | `ANTHROPIC_API_BASE` / `OPENAI_API_BASE` | | Codex CLI | `OPENAI_API_KEY` | `https://claudeapikey.dev/v1` | | Zed | Assistant provider settings | `https://claudeapikey.dev/v1` | | LangChain / LiteLLM | `OPENAI_API_KEY` | `https://claudeapikey.dev/v1` | | curl (Messages) | `x-api-key` header | `https://claudeapikey.dev/v1/messages` | | curl (Chat Completions) | `Authorization: Bearer` | `https://claudeapikey.dev/v1/chat/completions` | > Anthropic-format clients take the bare origin. OpenAI-format clients take the origin plus `/v1`. That one line resolves most support questions we get. - [Authentication](https://claudeapikey.dev/docs/getting-started/authentication/) — Header reference - [Base URLs](https://claudeapikey.dev/docs/getting-started/base-urls/) — Why the suffix differs - [Troubleshooting](https://claudeapikey.dev/docs/guides/troubleshooting/) — Symptom-first index --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Production best practices > What to get right before an integration carries real traffic. _Source: https://claudeapikey.dev/docs/guides/best-practices/ · Home > Docs > Working with the API_ ## Reliability - **Retry the right codes.** `429`, `500`, `502`, `503`, `529` with exponential backoff and jitter. Never blind-retry `400`, `401`, `402` or `403`. - **Set timeouts.** A long generation can legitimately take minutes; a hung socket should not take your worker with it. 120s is a reasonable ceiling for non-streaming calls. - **Cap concurrency** with a semaphore rather than firing everything and retrying rejections. - **Have a fallback model.** If your primary is overloaded, degrading to a smaller model beats returning an error. ## Cost - Set `max_tokens` on every call. It is a spending cap, not a hint. - Route by task — see [Cost control](/docs/guides/cost-control/). - Turn on [prompt caching](/docs/guides/prompt-caching/) for stable prefixes. - Log `usage` per request with a task label so you can attribute spend later. ## Security - Keys live in environment variables or a secret manager — never in the repo, never in a client bundle. - One key per service, so revocation is surgical. - Never build prompts by string-concatenating untrusted user input with instructions; keep user content in a `user` message and instructions in `system`. - Treat model output as untrusted. Do not `eval` it, do not run generated SQL against a writable role. ## Observability - Log model, latency, `usage`, `stop_reason` and your own task type for every call. - Alert on the `stop_reason: max_tokens` rate — a rise means answers are being truncated. - Track cost per task type, not just total. Totals hide which workload regressed. > The highest-value alert is on truncation, not errors. A truncated answer returns HTTP 200 and looks fine to your monitoring while being wrong to the user. - [Errors](https://claudeapikey.dev/docs/api-reference/errors/) — Retry semantics - [Rate limits](https://claudeapikey.dev/docs/api-reference/rate-limits/) — Concurrency - [Cost control](https://claudeapikey.dev/docs/guides/cost-control/) — Spend management --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Claude Code > Point Anthropic's CLI agent at the gateway with two environment variables — and avoid the one trap that silently bills the wrong key. _Source: https://claudeapikey.dev/docs/guides/claude-code/ · Home > Docs > Integration guides_ Claude Code reads its endpoint from the environment, so no patching is needed. This is the most common setup on the gateway. ## Setup ```bash export ANTHROPIC_BASE_URL="https://claudeapikey.dev" export ANTHROPIC_AUTH_TOKEN="sk-your-key" claude ``` That is the whole integration. Add both lines to `~/.zshrc` or `~/.bashrc` to make it permanent. > **The trap:** if `ANTHROPIC_API_KEY` is set in your shell, Claude Code may prefer it and bill that key, silently ignoring both your gateway key and any Pro/Max subscription. Run `echo $ANTHROPIC_API_KEY` — if it prints anything, `unset ANTHROPIC_API_KEY` before starting. ## Per-project configuration To scope the gateway to one repo rather than your whole shell, use a settings file: ```json { "env": { "ANTHROPIC_BASE_URL": "https://claudeapikey.dev", "ANTHROPIC_AUTH_TOKEN": "sk-your-key" } } ``` Place it at `.claude/settings.json` in the project, or `~/.claude/settings.json` for a user-wide default. Add `.claude/settings.json` to `.gitignore` — it contains a key. ## Verifying it worked ```bash # 1. the gateway answers and the key is valid curl -s -o /dev/null -w "%{http_code}\n" https://claudeapikey.dev/v1/models \ -H "Authorization: Bearer $ANTHROPIC_AUTH_TOKEN" # 2. start Claude Code and confirm usage appears in your dashboard claude ``` If requests do not appear in your dashboard, they are going somewhere else — almost always the `ANTHROPIC_API_KEY` trap above. ## Choosing a model ```bash claude --model claude-opus-4-8 ``` Any id from [Models](/docs/getting-started/models/) works. Opus for hard refactors, Sonnet for everyday work — see [Cost control](/docs/guides/cost-control/) for why that choice dominates your bill. ## Keeping the cost sane - Claude Code resends file contents and tool results every turn — long sessions grow expensive non-linearly. Use `/clear` between unrelated tasks. - Large tool outputs (a 500-row query, a whole build log) become permanent context. Pipe through `head` where you can. - Prompt caching helps most here because the system prompt and tool definitions are stable across turns. - [MCP servers](https://claudeapikey.dev/docs/guides/mcp-servers/) — Give Claude Code real tools - [Cost control](https://claudeapikey.dev/docs/guides/cost-control/) — Where agent tokens go - [Troubleshooting](https://claudeapikey.dev/docs/guides/troubleshooting/) — When nothing happens --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Cline > Configure the Cline VS Code extension against the Anthropic-format endpoint. _Source: https://claudeapikey.dev/docs/guides/cline/ · Home > Docs > Integration guides_ Cline speaks the Anthropic Messages format natively, so it uses `/v1/messages` and the base URL takes **no** `/v1` suffix. ## Setup 1. Install Cline from the VS Code marketplace and open its settings panel. 2. Set **API Provider** to **Anthropic**. 3. Paste your `sk-` key into the API key field. 4. Enable the custom base URL option and set it to `https://claudeapikey.dev`. 5. Pick a model id, for example `claude-sonnet-4-6`. - **Provider:** Anthropic - **Base URL:** `https://claudeapikey.dev` — no `/v1` - **Endpoint used:** `/v1/messages` > If Cline's settings offer an OpenAI-Compatible provider instead, that path works too — but then the base URL becomes `https://claudeapikey.dev/v1`. Pick one and match the suffix to it. ## Cost notes Cline is an agent: it reads files, runs commands and feeds results back as context. The same economics as [Claude Code](/docs/guides/claude-code/) apply — output tokens dominate, and tool results accumulate. Start new tasks rather than continuing one long thread. - [Messages API](https://claudeapikey.dev/docs/api-reference/messages/) — The format Cline uses - [Cost control](https://claudeapikey.dev/docs/guides/cost-control/) — Keeping agent runs affordable --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Codex CLI > Run OpenAI's Codex CLI and the Codex extension against the gateway. _Source: https://claudeapikey.dev/docs/guides/codex-cli/ · Home > Docs > Integration guides_ Codex CLI reads the standard OpenAI environment variables, so redirecting it needs no patch. ```bash export OPENAI_BASE_URL="https://claudeapikey.dev/v1" export OPENAI_API_KEY="sk-your-key" codex ``` ## Config file ```bash # ~/.codex/config.toml model = "gpt-5.5" model_provider = "claudeapikey" [model_providers.claudeapikey] name = "ClaudeAPIKey.dev" base_url = "https://claudeapikey.dev/v1" env_key = "OPENAI_API_KEY" ``` Both GPT and Claude ids work here — the gateway translates, so you can run Codex CLI against `claude-sonnet-4-6` if you prefer it. > Codex is an agent and streams heavily. If output arrives in one lump rather than progressively, a proxy between you and the gateway is buffering — see [Streaming](/docs/api-reference/streaming/). - [Chat Completions](https://claudeapikey.dev/docs/api-reference/chat-completions/) — Endpoint reference - [Models](https://claudeapikey.dev/docs/getting-started/models/) — GPT and Claude ids --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Context management > Why long sessions get expensive non-linearly, and the patterns that keep history bounded. _Source: https://claudeapikey.dev/docs/guides/context-management/ · Home > Docs > Working with the API_ The API is stateless. Every turn resends the entire conversation, so input tokens grow with the transcript — and cost grows with them, on every subsequent call. ## The shape of the problem A conversation of N turns, each adding roughly T tokens, costs on the order of N²T/2 input tokens in total rather than NT. Doubling the session length roughly quadruples the input bill. This is why an agent that felt cheap for ten minutes is alarming after an hour. ## Four patterns that work | Pattern | How | Trade-off | |---|---|---| | Sliding window | Keep the last K turns, drop the rest | Loses early context entirely | | Running summary | Periodically replace old turns with a model-written summary | Costs one extra call; keeps the gist | | Fresh session per task | Start clean when the topic changes | Free, and usually correct | | Externalise | Keep state in files or a database, load only what the current step needs | Most work, best result | ## Summarising older turns ```python def compact(history, keep=6): """Replace everything but the last `keep` turns with one summary turn.""" if len(history) <= keep: return history old, recent = history[:-keep], history[-keep:] summary = summarise(old) # one cheap Haiku call return [{"role": "user", "content": f"Summary of earlier discussion: {summary}"}] + recent ``` Run the summarisation on a cheap model. It is a compression task, not a reasoning task. ## Tool results are the usual culprit - A tool returning 500 rows puts 500 rows in every later request. Cap the rows, or return a summary plus a handle to fetch detail on demand. - Full file contents read by an agent stay in context. Read the section you need, not the whole file. - Screenshots and images are resent like any other content — drop them once described. - Build logs and stack traces: keep the first and last 50 lines, discard the middle. > In agentic tools, starting a new session between unrelated tasks is the single cheapest optimisation available. In Claude Code that is `/clear`. - [Cost control](https://claudeapikey.dev/docs/guides/cost-control/) — The full picture - [Prompt caching](https://claudeapikey.dev/docs/guides/prompt-caching/) — Making the stable part cheap - [Models](https://claudeapikey.dev/docs/getting-started/models/) — 1M-context options --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Continue > Configure the Continue extension for VS Code and JetBrains. _Source: https://claudeapikey.dev/docs/guides/continue/ · Home > Docs > Integration guides_ Continue is configured by a single JSON or YAML file and supports custom providers directly. ## config.json ```json { "models": [ { "title": "Claude Sonnet 4.6", "provider": "anthropic", "model": "claude-sonnet-4-6", "apiKey": "sk-your-key", "apiBase": "https://claudeapikey.dev" }, { "title": "Claude Opus 4.8", "provider": "anthropic", "model": "claude-opus-4-8", "apiKey": "sk-your-key", "apiBase": "https://claudeapikey.dev" } ] } ``` The file lives at `~/.continue/config.json`. With `provider: anthropic`, `apiBase` takes **no** `/v1`; if you use `provider: openai` instead, append `/v1`. ## A cheaper autocomplete model ```json { "tabAutocompleteModel": { "title": "Haiku", "provider": "anthropic", "model": "claude-haiku-4-5", "apiKey": "sk-your-key", "apiBase": "https://claudeapikey.dev" } } ``` > Autocomplete fires constantly. Pointing it at a frontier model is the fastest way to burn credits for no benefit — Haiku is the right tool for that job. - [VS Code extensions](https://claudeapikey.dev/docs/guides/vscode-extensions/) — Other extension options - [JetBrains](https://claudeapikey.dev/docs/guides/jetbrains/) — IntelliJ, PyCharm, GoLand --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Cost control > Where the tokens actually go in an agentic workload, and the four levers that move the bill. _Source: https://claudeapikey.dev/docs/guides/cost-control/ · Home > Docs > Working with the API_ If you have run a coding agent for a full working day you have seen the bill. Agentic loops are token-hungry in a way ordinary chat never is: every file read, every tool result and every diff lands in context and is resent on the next turn. ## Output tokens are the expensive ones On every frontier model, output costs 4–5× input. A model that thinks out loud in its final answer burns money faster than one that answers tightly. Cap `max_tokens` deliberately and prefer concise models for routine work. | Model | $/M in | $/M out | Ratio | |---|---|---|---| | `claude-opus-4-8` | $3 | $15 | 5× | | `claude-sonnet-4-6` | $0.90 | $4.50 | 5× | | `claude-haiku-4-5` | $0.40 | $2 | 5× | | `gemini-3-flash` | $0.20 | $1.20 | 6× | ## The four levers 1. **Route by task.** Classification, extraction and routing do not need Opus — Haiku does them at roughly a seventh of the price. Reserve the flagship for reasoning-heavy work. 2. **Prompt caching.** Stable system prompts and tool definitions are re-read from cache at a fraction of the input price. For agents this is the single biggest lever — see [Prompt caching](/docs/guides/prompt-caching/). 3. **Trim context each turn.** Do not resend an entire transcript; keep recent turns plus a running summary. Unbounded context growth is what turns a $2 session into a $20 one. 4. **Cap the loop.** Set a hard step limit on autonomous runs so a stuck agent cannot spin forever. ## Estimating before you run ```python IN_RATE = {"claude-opus-4-8": 3.0, "claude-sonnet-4-6": 0.9, "claude-haiku-4-5": 0.4} OUT_RATE = {"claude-opus-4-8": 15.0, "claude-sonnet-4-6": 4.5, "claude-haiku-4-5": 2.0} def cost(model, tok_in, tok_out): return tok_in / 1e6 * IN_RATE[model] + tok_out / 1e6 * OUT_RATE[model] print(cost("claude-opus-4-8", 20_000, 2_000)) # 0.09 print(cost("claude-sonnet-4-6", 20_000, 2_000)) # 0.027 print(cost("claude-haiku-4-5", 20_000, 2_000)) # 0.012 ``` The same 20K-in / 2K-out task costs 7.5× more on Opus than on Haiku. If the task is classification, that multiple buys nothing. ## Measuring after Every response carries a `usage` object. Log it against a request id and a task type — without that, cost optimisation is guesswork. The gateway dashboard shows per-model usage, but only your own logs know *why* a request was made. > The most common expensive mistake is not the model choice — it is resending a growing transcript on every turn of a long agent session. Fix context growth first, then tune models. - [Prompt caching](https://claudeapikey.dev/docs/guides/prompt-caching/) — The biggest single lever - [Context management](https://claudeapikey.dev/docs/guides/context-management/) — Keeping history bounded - [Models](https://claudeapikey.dev/docs/getting-started/models/) — Rates for every model --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Cursor > Use Claude models inside Cursor through the OpenAI-compatible endpoint, without an Anthropic account. _Source: https://claudeapikey.dev/docs/guides/cursor/ · Home > Docs > Integration guides_ Cursor talks the OpenAI Chat Completions format, so it connects through `/v1/chat/completions` — which means the **`/v1` suffix on the base URL is required**. ## Setup 1. Open **Settings → Models** (or **Cursor Settings → Models** depending on version). 2. Find the OpenAI API section and paste your `sk-` key. 3. Enable **Override OpenAI Base URL** and set it to `https://claudeapikey.dev/v1`. 4. Add a custom model name — use an exact id such as `claude-sonnet-4-6`. 5. Click **Verify** / **Save**. Cursor makes a test call; a green result means it is working. - **Base URL:** `https://claudeapikey.dev/v1` - **API key:** your `sk-` key - **Model ids:** `claude-sonnet-4-6`, `claude-opus-4-8`, `gpt-5.5`, `gemini-3-pro` - **Endpoint used:** `/v1/chat/completions` > Omitting `/v1` is the number-one Cursor setup failure — verification fails with a 404 because Cursor appends only `/chat/completions` to whatever base you give it. See [Base URLs](/docs/getting-started/base-urls/). ## What works and what does not | Feature | Status | |---|---| | Chat with a custom model | Works | | Inline edit (Ctrl/Cmd+K) | Works | | Composer / agent mode | Works with models that support tool use | | Cursor Tab (autocomplete) | Uses Cursor's own in-house model and is unaffected by this setting | Cursor Tab is not routed through any custom base URL — that is a Cursor product decision, not a gateway limitation. ## If verification fails - `404` — the `/v1` suffix is missing from the base URL. - `401` — the key is wrong or has a stray space. Re-copy it from the dashboard. - Model not found — the custom model name must be an exact id from [Models](/docs/getting-started/models/), not a display name like "Claude Sonnet". - Verified but replies are empty — see [Troubleshooting](/docs/guides/troubleshooting/). - [Chat Completions](https://claudeapikey.dev/docs/api-reference/chat-completions/) — The endpoint Cursor uses - [Base URLs](https://claudeapikey.dev/docs/getting-started/base-urls/) — The /v1 rule - [Cline](https://claudeapikey.dev/docs/guides/cline/) — The Anthropic-format alternative --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # JetBrains IDEs > IntelliJ IDEA, PyCharm, WebStorm, GoLand and the rest — via Continue or an OpenAI-compatible plugin. _Source: https://claudeapikey.dev/docs/guides/jetbrains/ · Home > Docs > Integration guides_ JetBrains' built-in AI Assistant does not currently accept a custom endpoint. The reliable route is a plugin that does. ## Continue plugin (recommended) 1. Install **Continue** from the JetBrains marketplace. 2. Open its config (the gear icon in the Continue panel) — it is the same `~/.continue/config.json` used by VS Code. 3. Add a model block with `apiBase` set to `https://claudeapikey.dev`. 4. Reload the plugin and select the model. ```json { "models": [{ "title": "Claude Sonnet 4.6", "provider": "anthropic", "model": "claude-sonnet-4-6", "apiKey": "sk-your-key", "apiBase": "https://claudeapikey.dev" }] } ``` ## Terminal alternative The JetBrains terminal runs [Claude Code](/docs/guides/claude-code/) and [Aider](/docs/guides/aider/) perfectly well, which is often a better agentic experience than any plugin. - [Continue](https://claudeapikey.dev/docs/guides/continue/) — Shared configuration - [Claude Code](https://claudeapikey.dev/docs/guides/claude-code/) — Terminal agent --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # MCP servers for Claude Code > A Model Context Protocol config that earns its place, plus the traps that cost real time. _Source: https://claudeapikey.dev/docs/guides/mcp-servers/ · Home > Docs > Integration guides_ MCP turns Claude Code from a code generator into something that can read your repo, query a database, drive a browser and open a PR. MCP wiring is independent of the gateway — it changes what tools exist, not where model calls go. ## A .mcp.json worth running ```json { "mcpServers": { "filesystem": {"command": "npx", "args": ["-y", "@modelcontextprotocol/server-filesystem", "."]}, "git": {"command": "uvx", "args": ["mcp-server-git", "--repository", "."]}, "github": {"command": "npx", "args": ["-y", "@modelcontextprotocol/server-github"], "env": {"GITHUB_TOKEN": "..."}}, "fetch": {"command": "uvx", "args": ["mcp-server-fetch"]}, "playwright": {"command": "npx", "args": ["-y", "@playwright/mcp@latest"]} } } ``` Drop it in the repo root; Claude Code detects it and asks you to trust it once. ## The traps - **The trust prompt is a feature.** Claude Code refuses to load project MCP servers until you approve them, per project. That is the security boundary, not a bug. - **Give a database server a read-only role.** An MCP server executes whatever SQL the model emits; a read-only role makes a hallucinated `DROP` physically impossible. - **Scope tokens narrowly.** For GitHub, `repo` plus `read:org` is enough. Do not hand it an all-scopes classic PAT. - **MCP output is context.** A 200-row result or a full page fetch is resent on every subsequent turn. Large tool outputs are the main reason agentic sessions get expensive. > Start with filesystem, git and github. Add browser and database servers only when a specific task needs them — every loaded server adds its schema to the token bill on every request. - [Claude Code](https://claudeapikey.dev/docs/guides/claude-code/) — Base setup - [Tool use](https://claudeapikey.dev/docs/api-reference/tool-use/) — How tool calling works - [Cost control](https://claudeapikey.dev/docs/guides/cost-control/) — Why tool output matters --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Migrating from the Anthropic API > Move an existing Anthropic-SDK codebase over by changing two values. _Source: https://claudeapikey.dev/docs/guides/migration-from-anthropic/ · Home > Docs > Working with the API_ Because the gateway speaks the Messages format, migration is a configuration change. Request bodies, response parsing, streaming handlers and tool definitions all stay as they are. ## Python ```python from anthropic import Anthropic client = Anthropic( api_key="sk-your-gateway-key", # was your Anthropic key base_url="https://claudeapikey.dev", # new line ) # everything below is unchanged ``` ## TypeScript ```typescript import Anthropic from "@anthropic-ai/sdk"; const client = new Anthropic({ apiKey: process.env.CLAUDEAPIKEY, baseURL: "https://claudeapikey.dev", }); ``` ## Environment-only migration ```bash export ANTHROPIC_BASE_URL="https://claudeapikey.dev" export ANTHROPIC_API_KEY="sk-your-gateway-key" ``` ## Migration checklist 1. Change the base URL and the key. No `/v1` suffix for Anthropic clients. 2. Map model ids — see [Models](/docs/getting-started/models/). Ids are pinned, so `claude-3-5-sonnet-20241022` style names from an older integration need updating to a current id. 3. Re-run your test suite. Response shapes are identical, so anything that breaks is a model-id or base-URL issue. 4. Check streaming if you use it: the event shape is unchanged, but any hardcoded hostname in a proxy or allowlist needs updating. 5. Move one service first and compare outputs before switching everything. ## What does not carry over - Anthropic Console features — workbench, org billing, per-workspace limits — are Anthropic products and have no equivalent here. - Batch API and any beta feature not exposed by this gateway. Check [Overview](/docs/api-reference/overview/) for the current endpoint list. - Contractual terms. If you need a DPA or SLA signed by Anthropic, stay on the official API. - [Base URLs](https://claudeapikey.dev/docs/getting-started/base-urls/) — The suffix rule - [Models](https://claudeapikey.dev/docs/getting-started/models/) — Id mapping - [Python SDK](https://claudeapikey.dev/docs/sdks/python/) — Full example --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Migrating from the OpenAI API > Keep the OpenAI SDK and reach Claude models by changing the base URL. _Source: https://claudeapikey.dev/docs/guides/migration-from-openai/ · Home > Docs > Working with the API_ The gateway exposes an OpenAI-compatible endpoint, so an existing OpenAI codebase can reach Claude models without rewriting request handling. ## Python ```python from openai import OpenAI client = OpenAI( api_key="sk-your-gateway-key", base_url="https://claudeapikey.dev/v1", # the /v1 is required here ) resp = client.chat.completions.create( model="claude-sonnet-4-6", # a Claude model through the OpenAI SDK messages=[{"role": "user", "content": "Hello"}], ) ``` ## LangChain ```python from langchain_openai import ChatOpenAI llm = ChatOpenAI( model="claude-sonnet-4-6", api_key="sk-your-gateway-key", base_url="https://claudeapikey.dev/v1", ) ``` ## Checklist 1. Set `base_url` to `https://claudeapikey.dev/v1` — with the suffix. 2. Swap model ids to ones from [Models](/docs/getting-started/models/). 3. Response parsing is unchanged: `choices[0].message.content`. 4. Streaming is unchanged: `choices[0].delta.content` plus the `[DONE]` sentinel. 5. Function calling uses the OpenAI `tools` shape on this endpoint. > Behaviour differs even when the interface does not. Claude and GPT respond differently to the same prompt, particularly on instruction-following and output formatting. Re-check prompts that depend on a specific style, and re-run evaluations rather than assuming parity. - [Chat Completions](https://claudeapikey.dev/docs/api-reference/chat-completions/) — Endpoint reference - [Models](https://claudeapikey.dev/docs/getting-started/models/) — Available ids - [Base URLs](https://claudeapikey.dev/docs/getting-started/base-urls/) — Why /v1 here --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # OpenClaw > Configure the OpenClaw agent framework to use gateway models. _Source: https://claudeapikey.dev/docs/guides/openclaw/ · Home > Docs > Integration guides_ OpenClaw accepts Anthropic-compatible endpoints, so it needs the base URL and key and nothing else. ```bash export ANTHROPIC_BASE_URL="https://claudeapikey.dev" export ANTHROPIC_API_KEY="sk-your-key" ``` If your build uses a config file rather than the environment, set the provider to Anthropic, the base URL to `https://claudeapikey.dev` (no `/v1`), and the model to an id from [Models](/docs/getting-started/models/). ## Agent economics Autonomous frameworks can run for a long time without supervision. Two safeguards are worth setting before the first long run: - A hard step limit, so a stuck loop cannot spin indefinitely. - A cheaper model for routine sub-tasks, reserving the frontier model for planning. - [Cost control](https://claudeapikey.dev/docs/guides/cost-control/) — Bounding an autonomous run - [Rate limits](https://claudeapikey.dev/docs/api-reference/rate-limits/) — Concurrency ceilings --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # OpenCode > Point the open-source terminal coding agent at the gateway. _Source: https://claudeapikey.dev/docs/guides/opencode/ · Home > Docs > Integration guides_ OpenCode supports custom providers through its config file, using the OpenAI-compatible surface. ```json { "provider": { "claudeapikey": { "npm": "@ai-sdk/openai-compatible", "options": { "baseURL": "https://claudeapikey.dev/v1", "apiKey": "sk-your-key" }, "models": { "claude-sonnet-4-6": {"name": "Claude Sonnet 4.6"}, "claude-opus-4-8": {"name": "Claude Opus 4.8"} } } } } ``` Config lives at `~/.config/opencode/opencode.json` or `opencode.json` in the project root. Restart OpenCode and select the provider. > OpenCode is under fast development and its config schema moves. If the block above is rejected, check the provider section of OpenCode's own docs — the values you need are always base URL `https://claudeapikey.dev/v1`, your key, and exact model ids. - [Chat Completions](https://claudeapikey.dev/docs/api-reference/chat-completions/) — The endpoint used - [Claude Code](https://claudeapikey.dev/docs/guides/claude-code/) — The Anthropic-native equivalent --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Prompt caching > Re-read a stable prefix at a fraction of the input price — what to cache, what not to, and how to verify it worked. _Source: https://claudeapikey.dev/docs/guides/prompt-caching/ · Home > Docs > Working with the API_ Prompt caching stores a prefix of your request so later calls that share it are billed at a reduced input rate. For anything with a large stable preamble — a system prompt, tool schemas, a document being questioned — it is the highest-return change available. ## What makes a good cache target | Good | Why | |---|---| | System prompts | Byte-identical on every call in a session | | Tool / function definitions | Stable across an entire agent run, and often large | | A document being asked about repeatedly | One upload, many questions | | Few-shot examples | Fixed, and frequently long | | Bad | Why | |---|---| | The user's latest message | Different every time — never a cache hit | | Anything with a timestamp or request id near the top | One changing byte early in the prefix invalidates everything after it | | Short prompts | Below the minimum cacheable length, so there is nothing to gain | ## Marking a prefix ```json { "model": "claude-sonnet-4-6", "max_tokens": 1024, "system": [{ "type": "text", "text": "", "cache_control": {"type": "ephemeral"} }], "messages": [{"role": "user", "content": "Question that changes every call"}] } ``` Everything up to and including the marked block becomes the cached prefix. Order matters: put stable content first and volatile content last. ## Verifying it works Check the `usage` object. A cache hit reports tokens under cache-read rather than ordinary input. If cache-read stays at zero across repeated calls, your prefix is not actually identical — a common cause is a timestamp, a session id or non-deterministic JSON key ordering in the preamble. > Serialise cached content deterministically. `json.dumps(obj)` without `sort_keys=True` can reorder keys between runs, which changes the bytes and silently costs you every cache hit. ## Where it matters most Agentic tools resend the same system prompt and the same tool schemas on every single turn. In a 50-turn session that preamble is paid for 50 times without caching, and roughly once with it. - [Cost control](https://claudeapikey.dev/docs/guides/cost-control/) — The other three levers - [Tool use](https://claudeapikey.dev/docs/api-reference/tool-use/) — Why schemas dominate input - [Messages API](https://claudeapikey.dev/docs/api-reference/messages/) — Where cache_control goes --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Troubleshooting > Symptom-first index — find what you are seeing and jump to the cause. _Source: https://claudeapikey.dev/docs/guides/troubleshooting/ · Home > Docs > Working with the API_ ## Nothing works at all Establish the baseline before debugging anything else: ```bash curl -s -o /dev/null -w "%{http_code}\n" https://claudeapikey.dev/v1/models \ -H "Authorization: Bearer $CLAUDEAPIKEY" ``` `200` means the URL and key are both fine and the problem is in your client. Anything else is covered below. ## By symptom | Symptom | Likely cause | Fix | |---|---|---| | `404` on every call | Base URL has the wrong `/v1` handling | [Base URLs](/docs/getting-started/base-urls/) | | `401` with a key you just created | Wrong header for the format, or a trailing space | [Authentication](/docs/getting-started/authentication/) | | `402` | No credits, or plan expired | [Top up](/topup/) · [Plans](/plans/) | | `400 max_tokens` | Messages requires `max_tokens`; Chat Completions does not | [Messages](/docs/api-reference/messages/) | | `model_not_found` | Display name used instead of an api id | [Models endpoint](/docs/api-reference/models-endpoint/) | | Streaming connects but yields nothing | Parsing the other format's event shape | [Streaming](/docs/api-reference/streaming/) | | Reply cut off mid-sentence | `stop_reason: max_tokens` | Raise `max_tokens` | | Works in curl, fails in the SDK | SDK appends its own path to your base URL | [Base URLs](/docs/getting-started/base-urls/) | | Usage not showing in dashboard | Requests going to a different key or endpoint | Check `ANTHROPIC_API_KEY` is unset — [Claude Code](/docs/guides/claude-code/) | | Frequent `429` | Agent parallelism too high | [Rate limits](/docs/api-reference/rate-limits/) | | Bill higher than expected | Context growth across turns | [Context management](/docs/guides/context-management/) | ## Reading the error body The status code tells you the class of problem; the `error.type` and `error.message` fields tell you which one. Log the whole body — a `400` with `"missing max_tokens"` and a `400` with `"invalid role"` need completely different fixes. ## Still stuck Contact [support@claudeapikey.dev](mailto:support@claudeapikey.dev) with the endpoint, the model id, the status code and the error body. Never include your API key — if you already pasted it somewhere, revoke it first. - [Errors](https://claudeapikey.dev/docs/api-reference/errors/) — Status code reference - [Base URLs](https://claudeapikey.dev/docs/getting-started/base-urls/) — The most common fault - [API key configuration](https://claudeapikey.dev/docs/guides/api-key-configuration/) — Per-tool settings --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # VS Code extensions > Which VS Code AI extensions accept a custom base URL, and how to point each at the gateway. _Source: https://claudeapikey.dev/docs/guides/vscode-extensions/ · Home > Docs > Integration guides_ Any extension that exposes a base-URL or custom-endpoint setting will work. Those that hardcode a vendor endpoint will not, and no gateway can change that. | Extension | Setting to change | Base URL | |---|---|---| | Cline | API Provider: Anthropic + custom base URL | `https://claudeapikey.dev` | | Continue | `apiBase` in `~/.continue/config.json` | `https://claudeapikey.dev` (anthropic) or `https://claudeapikey.dev/v1` (openai) | | Roo Code | Anthropic provider + custom base URL | `https://claudeapikey.dev` | | Kilo Code | OpenAI-compatible provider | `https://claudeapikey.dev/v1` | | Any OpenAI-compatible extension | Base URL / endpoint field | `https://claudeapikey.dev/v1` | > GitHub Copilot does not support custom endpoints and cannot be redirected. Use Cline or Continue alongside it if you want Claude models in VS Code. ## The suffix rule, once more - Extension configured as an **Anthropic** provider — base URL without `/v1`. - Extension configured as **OpenAI-compatible** — base URL with `/v1`. - A `404` almost always means you got this backwards. - [Cline](https://claudeapikey.dev/docs/guides/cline/) — Full setup - [Continue](https://claudeapikey.dev/docs/guides/continue/) — Full setup - [Base URLs](https://claudeapikey.dev/docs/getting-started/base-urls/) — The rule explained --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Zed > Add the gateway as a custom assistant provider in the Zed editor. _Source: https://claudeapikey.dev/docs/guides/zed/ · Home > Docs > Integration guides_ Zed's assistant supports custom OpenAI-compatible endpoints, configured in `settings.json`. ## Configuration ```json { "language_models": { "openai": { "api_url": "https://claudeapikey.dev/v1", "available_models": [ {"name": "claude-sonnet-4-6", "display_name": "Claude Sonnet 4.6", "max_tokens": 200000}, {"name": "claude-opus-4-8", "display_name": "Claude Opus 4.8", "max_tokens": 200000} ] } } } ``` 1. Open the command palette and choose **zed: open settings**. 2. Paste the block above, merging it with your existing settings. 3. Restart Zed, open the assistant panel, and enter your `sk-` key when prompted. 4. Select one of the models you declared. > Zed's settings schema changes between releases. If `language_models` is rejected, check Zed's own assistant configuration docs for the current key name — the value to supply is always the same: base URL `https://claudeapikey.dev/v1` plus your key. - [Chat Completions](https://claudeapikey.dev/docs/api-reference/chat-completions/) — The format Zed uses - [Base URLs](https://claudeapikey.dev/docs/getting-started/base-urls/) — Why /v1 is needed here --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Changelog > Notable changes to the API surface, models and documentation. _Source: https://claudeapikey.dev/docs/resources/changelog/ · Home > Docs > Resources_ > Model ids are pinned. A published id keeps its behaviour; new versions arrive as new ids, and retirements are announced here before they take effect. ## Documentation - Full docs site published: getting started, API reference, tool and integration guides, SDK examples, billing and resources. - Every page now serves a Markdown mirror at the same path with a `.md` suffix, plus copy-for-LLM and ask-in-an-assistant actions. - [llms.txt](/llms.txt) and [llms-full.txt](/llms-full.txt) extended to cover the docs tree. ## API - `/v1/messages` — Anthropic Messages format, streaming and tool use. - `/v1/chat/completions` — OpenAI Chat Completions format against the same models. - `/v1/models` — model discovery. - Machine-readable schema at [/openapi.json](/openapi.json). ## Models The current catalog is always listed on [Models](/docs/getting-started/models/) and served live from [/v1/models](/docs/api-reference/models-endpoint/). Treat the endpoint as authoritative over any documentation page. - [Models](https://claudeapikey.dev/docs/getting-started/models/) — Current catalog - [API overview](https://claudeapikey.dev/docs/api-reference/overview/) — Endpoint list --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # FAQ > Direct answers to the questions we are actually asked. _Source: https://claudeapikey.dev/docs/resources/faq/ · Home > Docs > Resources_ ## Is this the official Anthropic API? No. ClaudeAPIKey.dev is an independently operated gateway that resells API capacity and exposes an Anthropic-compatible interface. We are not affiliated with, endorsed by or sponsored by Anthropic. If you need a contract with Anthropic itself, use the official API. ## Will my existing code work? If it uses the Anthropic or OpenAI SDK, yes — change the base URL and the key. See [Migrating from Anthropic](/docs/guides/migration-from-anthropic/) or [from OpenAI](/docs/guides/migration-from-openai/). ## Do credits expire? No. Prepaid credits stay on the account until spent. Unlimited plans are different — they run for a fixed window and end when it expires. ## Which is cheaper for me, credits or a plan? Credits suit bursty or low-volume use because you pay only for tokens. A flat-rate plan wins when you are running an agent for hours a day, where per-token billing on a long session outruns the plan price. See [Plans](/docs/billing/plans/). ## Can I use this with Claude Code? Yes — two environment variables. It is the most common setup. See [Claude Code](/docs/guides/claude-code/). ## Why is my bill higher than I estimated? Almost always context growth. Every turn resends the whole conversation, so a long agent session pays for its history repeatedly. See [Context management](/docs/guides/context-management/). ## Do you support tool use / function calling? Yes, in both formats. See [Tool use](/docs/api-reference/tool-use/). ## Do you store my prompts? Request metadata needed for billing and abuse prevention is retained. Do not send data you are contractually barred from sharing with a third-party processor — see [Security and privacy](/docs/resources/security/). ## What happens if a model is deprecated? Model ids are pinned and do not silently change under you. Retirements are announced in the [changelog](/docs/resources/changelog/). ## How do I get support? [support@claudeapikey.dev](mailto:support@claudeapikey.dev). Include the endpoint, model id, status code and error body — never your API key. - [Troubleshooting](https://claudeapikey.dev/docs/guides/troubleshooting/) — Symptom-first index - [Pricing](https://claudeapikey.dev/docs/billing/pricing/) — How billing works --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Glossary > The terms used throughout these docs, defined precisely. _Source: https://claudeapikey.dev/docs/resources/glossary/ · Home > Docs > Resources_ | Term | Meaning | |---|---| | **Token** | The unit models read and write, roughly ¾ of an English word. Billing is per million tokens. | | **Input tokens** | Everything you send: system prompt, full message history, tool schemas, images. Grows every turn of a conversation. | | **Output tokens** | What the model generates. Costs 4–5× input on every frontier model. | | **Context window** | The maximum tokens a model can consider at once, input plus output. Exceeding it is an error, not silent truncation. | | **Base URL** | The origin a client sends requests to. The single most common misconfiguration — see [Base URLs](/docs/getting-started/base-urls/). | | **Messages API** | Anthropic's request format. `POST /v1/messages`, replies in a `content[]` array. | | **Chat Completions** | OpenAI's request format. `POST /v1/chat/completions`, replies in `choices[]`. | | **Streaming / SSE** | Server-sent events delivering the reply token by token instead of in one body. | | **stop_reason** | Why generation ended: `end_turn`, `max_tokens`, `stop_sequence` or `tool_use`. | | **Prompt caching** | Re-reading a stable prefix at a reduced input rate. See [Prompt caching](/docs/guides/prompt-caching/). | | **Tool use** | The model requesting that you run a function, then continuing with the result. | | **MCP** | Model Context Protocol — a standard way to expose tools to an agent. | | **Credits** | Prepaid balance spent at per-token rates. Do not expire. | | **Unlimited plan** | Flat-rate access for a fixed window, under fair-use limits, instead of per-token billing. | | **Gateway** | A service that fronts one or more model providers behind a single API and key. | - [Models](https://claudeapikey.dev/docs/getting-started/models/) — Ids, context windows, rates - [Pricing](https://claudeapikey.dev/docs/billing/pricing/) — How credits work --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Security and privacy > How to handle keys, what to avoid sending, and the boundaries of this service. _Source: https://claudeapikey.dev/docs/resources/security/ · Home > Docs > Resources_ ## Key handling - Keys are shown in full once. Store them in a secret manager or environment variable. - One key per service or machine so revocation is surgical. - Revocation takes effect immediately. - Never put a key in client-side JavaScript, a mobile binary, or a public repository. Anything shipped to a user is public. - If a key leaks, revoke first and investigate afterwards. ## What to avoid sending This is a third-party gateway sitting between you and the upstream provider. Treat it as an external processor: - Do not send data you are contractually barred from sharing with a subprocessor. - Do not send regulated data — health records, payment card numbers, government identifiers — without a legal basis for doing so. - Redact secrets before sending logs or config files to a model. Agents that read whole files will happily read your `.env`. ## Prompt injection Any content the model reads can attempt to instruct it — web pages, files, tool results, issue comments. Consequences follow from what the model can *do*, so constrain that: - Give database tools a read-only role. - Scope API tokens to the minimum needed. - Require human approval for irreversible actions. - Never `eval` model output or execute generated SQL against a writable connection. ## Transport All traffic is HTTPS. Reject plain HTTP in your own client configuration; a base URL of `http://` would send your key in cleartext. > Certificate pinning to our hostname is not recommended — infrastructure changes would break your client without warning. Standard CA validation is the right level. - [Authentication](https://claudeapikey.dev/docs/getting-started/authentication/) — Key mechanics - [Best practices](https://claudeapikey.dev/docs/guides/best-practices/) — Production hardening - [MCP servers](https://claudeapikey.dev/docs/guides/mcp-servers/) — Least-privilege tools --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Status and reliability > How the gateway handles upstream failures, and what to do when something is degraded. _Source: https://claudeapikey.dev/docs/resources/status/ · Home > Docs > Resources_ ## How failover works Requests are distributed across upstream capacity. When one upstream returns an error or is overloaded, traffic moves to another rather than surfacing a hard failure — which is why a `529` here is rarer than calling a single provider directly. ## What each failure means for you | You see | Where the fault is | What to do | |---|---|---| | `529` overloaded | Upstream capacity | Retry with backoff; usually clears in seconds | | `500` | Gateway | Retry once, then report with the request id | | `429` | Your account's concurrency | Reduce parallelism — [Rate limits](/docs/api-reference/rate-limits/) | | Slow first token | Upstream queueing under load | Expected during peaks; streaming makes it visible sooner | ## Building for degradation - Retry `429`, `500`, `502`, `503` and `529` with exponential backoff and jitter. - Keep a fallback model configured — degrading from Opus to Sonnet beats returning an error. - Set timeouts so a stalled request cannot hold a worker indefinitely. - Run the `/v1/models` health check in CI so a broken credential is caught before deploy. For an incident affecting your account, contact [support@claudeapikey.dev](mailto:support@claudeapikey.dev) with timestamps and a request id. - [Errors](https://claudeapikey.dev/docs/api-reference/errors/) — Full status reference - [Best practices](https://claudeapikey.dev/docs/guides/best-practices/) — Resilient clients --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Availability and payment methods > Where the service can be used, and how to pay when card networks are unavailable. _Source: https://claudeapikey.dev/docs/resources/supported-countries/ · Home > Docs > Resources_ The API itself has no geographic restriction — it is a standard HTTPS endpoint reachable anywhere with internet access. What varies is payment. ## Paying | Method | Availability | Notes | |---|---|---| | Bank card | Most countries | Processed by our payment provider; some issuers decline cross-border AI purchases | | Crypto (USDT, BTC, ETH) | Everywhere | The reliable route where card payments are blocked, and it carries the credit bonus | Crypto is the common path for users whose region cannot complete card checkout with the upstream providers directly. Current bonus rates are on [Pricing](/docs/billing/pricing/). ## If a card is declined 1. Check the issuer did not block it as an international or high-risk merchant transaction. 2. Try a different card, or switch to crypto. 3. Contact [support@claudeapikey.dev](mailto:support@claudeapikey.dev) with the approximate time of the attempt — never send full card details. > You are responsible for complying with the laws applicable to you, and with the acceptable-use policies of the underlying model providers. - [Pricing](https://claudeapikey.dev/docs/billing/pricing/) — Rates and bonuses - [Credits](https://claudeapikey.dev/docs/billing/credits/) — How the balance works - [Plans](https://claudeapikey.dev/docs/billing/plans/) — Flat-rate options --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # SDKs overview > Official Anthropic and OpenAI SDKs both work unmodified — pick by which format you prefer. _Source: https://claudeapikey.dev/docs/sdks/ · Home > Docs > SDKs_ There is no bespoke SDK to install. The gateway is wire-compatible with the official clients, so you use the ones you already know. | Language | Anthropic format | OpenAI format | |---|---|---| | Python | `pip install anthropic` | `pip install openai` | | TypeScript | `npm i @anthropic-ai/sdk` | `npm i openai` | | Go | community clients, or plain HTTP | community clients, or plain HTTP | | Anything else | HTTP + JSON | HTTP + JSON | > Base URL differs by format: `https://claudeapikey.dev` for Anthropic clients, `https://claudeapikey.dev/v1` for OpenAI clients. See [Base URLs](/docs/getting-started/base-urls/). - [Python](https://claudeapikey.dev/docs/sdks/python/) — Both SDKs, working examples - [TypeScript](https://claudeapikey.dev/docs/sdks/typescript/) — Node and browser-server - [cURL](https://claudeapikey.dev/docs/sdks/curl/) — No dependencies - [Go](https://claudeapikey.dev/docs/sdks/go/) — Plain net/http --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # cURL > Dependency-free examples for every endpoint — useful for debugging and CI checks. _Source: https://claudeapikey.dev/docs/sdks/curl/ · Home > Docs > SDKs_ ## List models ```bash curl https://claudeapikey.dev/v1/models \ -H "Authorization: Bearer $CLAUDEAPIKEY" ``` ## Messages ```bash curl https://claudeapikey.dev/v1/messages \ -H "x-api-key: $CLAUDEAPIKEY" \ -H "anthropic-version: 2023-06-01" \ -H "content-type: application/json" \ -d '{ "model": "claude-sonnet-4-6", "max_tokens": 1024, "system": "Be concise.", "messages": [{"role": "user", "content": "What is a CDN?"}] }' ``` ## Chat Completions ```bash curl https://claudeapikey.dev/v1/chat/completions \ -H "Authorization: Bearer $CLAUDEAPIKEY" \ -H "content-type: application/json" \ -d '{ "model": "claude-sonnet-4-6", "messages": [ {"role": "system", "content": "Be concise."}, {"role": "user", "content": "What is a CDN?"} ] }' ``` ## Streaming ```bash curl -N https://claudeapikey.dev/v1/messages \ -H "x-api-key: $CLAUDEAPIKEY" \ -H "anthropic-version: 2023-06-01" \ -H "content-type: application/json" \ -d '{"model":"claude-sonnet-4-6","max_tokens":512,"stream":true, "messages":[{"role":"user","content":"Count to five"}]}' ``` `-N` disables curl's own buffering. Without it the whole stream appears at once and streaming looks broken. ## A CI health check ```bash #!/usr/bin/env bash set -euo pipefail code=$(curl -s -o /dev/null -w "%{http_code}" https://claudeapikey.dev/v1/models \ -H "Authorization: Bearer $CLAUDEAPIKEY") [ "$code" = "200" ] || { echo "gateway check failed: $code" >&2; exit 1; } echo "gateway ok" ``` - [Errors](https://claudeapikey.dev/docs/api-reference/errors/) — Interpreting the status - [Troubleshooting](https://claudeapikey.dev/docs/guides/troubleshooting/) — When curl works but the SDK does not --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Go > Call the gateway from Go with the standard library — no third-party client needed. _Source: https://claudeapikey.dev/docs/sdks/go/ · Home > Docs > SDKs_ Both endpoints are plain JSON over HTTP, so `net/http` is sufficient. ```go package main import ( "bytes" "encoding/json" "fmt" "net/http" "os" "time" ) type msg struct { Role string `json:"role"` Content string `json:"content"` } type req struct { Model string `json:"model"` MaxTokens int `json:"max_tokens"` Messages []msg `json:"messages"` } type resp struct { Content []struct { Text string `json:"text"` } `json:"content"` Usage struct { InputTokens int `json:"input_tokens"` OutputTokens int `json:"output_tokens"` } `json:"usage"` } func main() { body, _ := json.Marshal(req{ Model: "claude-sonnet-4-6", MaxTokens: 1024, Messages: []msg{{Role: "user", Content: "Hello"}}, }) r, _ := http.NewRequest("POST", "https://claudeapikey.dev/v1/messages", bytes.NewReader(body)) r.Header.Set("x-api-key", os.Getenv("CLAUDEAPIKEY")) r.Header.Set("anthropic-version", "2023-06-01") r.Header.Set("content-type", "application/json") client := &http.Client{Timeout: 120 * time.Second} res, err := client.Do(r) if err != nil { panic(err) } defer res.Body.Close() var out resp json.NewDecoder(res.Body).Decode(&out) fmt.Println(out.Content[0].Text) fmt.Println(out.Usage.InputTokens, out.Usage.OutputTokens) } ``` > Always set an explicit `Timeout` on the client. Go's default `http.Client` has none, so a stalled connection blocks a goroutine indefinitely. - [Messages API](https://claudeapikey.dev/docs/api-reference/messages/) — Field reference - [Best practices](https://claudeapikey.dev/docs/guides/best-practices/) — Retries and timeouts --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # Python > Working examples with the anthropic and openai packages, including streaming and async. _Source: https://claudeapikey.dev/docs/sdks/python/ · Home > Docs > SDKs_ ## Anthropic SDK ```bash pip install anthropic ``` ```python import os from anthropic import Anthropic client = Anthropic( api_key=os.environ["CLAUDEAPIKEY"], base_url="https://claudeapikey.dev", # no /v1 ) msg = client.messages.create( model="claude-sonnet-4-6", max_tokens=1024, system="You are a concise assistant.", messages=[{"role": "user", "content": "Explain HTTP caching in three sentences."}], ) print(msg.content[0].text) print(msg.usage.input_tokens, msg.usage.output_tokens) ``` ## Streaming ```python with client.messages.stream( model="claude-sonnet-4-6", max_tokens=1024, messages=[{"role": "user", "content": "Write a haiku about latency."}], ) as stream: for text in stream.text_stream: print(text, end="", flush=True) final = stream.get_final_message() print("\n", final.usage.output_tokens, "output tokens") ``` ## Async ```python import asyncio from anthropic import AsyncAnthropic client = AsyncAnthropic(api_key="sk-your-key", base_url="https://claudeapikey.dev") async def ask(q): m = await client.messages.create( model="claude-haiku-4-5", max_tokens=256, messages=[{"role": "user", "content": q}], ) return m.content[0].text async def main(): sem = asyncio.Semaphore(8) # respect rate limits async def guarded(q): async with sem: return await ask(q) print(await asyncio.gather(*(guarded(q) for q in ["1+1?", "2+2?", "3+3?"]))) asyncio.run(main()) ``` ## OpenAI SDK ```python from openai import OpenAI client = OpenAI(api_key="sk-your-key", base_url="https://claudeapikey.dev/v1") # with /v1 resp = client.chat.completions.create( model="claude-sonnet-4-6", messages=[{"role": "user", "content": "Hello"}], ) print(resp.choices[0].message.content) ``` > Both snippets call the same model at the same price. Use whichever library your project already depends on. - [Messages API](https://claudeapikey.dev/docs/api-reference/messages/) — Field reference - [Streaming](https://claudeapikey.dev/docs/api-reference/streaming/) — Event shapes - [Best practices](https://claudeapikey.dev/docs/guides/best-practices/) — Timeouts and retries --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._ # TypeScript > Node examples with @anthropic-ai/sdk and openai, including streaming. _Source: https://claudeapikey.dev/docs/sdks/typescript/ · Home > Docs > SDKs_ ## Anthropic SDK ```bash npm install @anthropic-ai/sdk ``` ```typescript import Anthropic from "@anthropic-ai/sdk"; const client = new Anthropic({ apiKey: process.env.CLAUDEAPIKEY!, baseURL: "https://claudeapikey.dev", // no /v1 }); const msg = await client.messages.create({ model: "claude-sonnet-4-6", max_tokens: 1024, messages: [{ role: "user", content: "Explain event loops briefly." }], }); const block = msg.content[0]; if (block.type === "text") console.log(block.text); ``` > `content` is a discriminated union. Narrow on `block.type === "text"` rather than indexing blindly — a tool-use reply has no `.text` and will be `undefined` at runtime. ## Streaming ```typescript const stream = await client.messages.stream({ model: "claude-sonnet-4-6", max_tokens: 1024, messages: [{ role: "user", content: "Count to five." }], }); for await (const event of stream) { if (event.type === "content_block_delta" && event.delta.type === "text_delta") { process.stdout.write(event.delta.text); } } const final = await stream.finalMessage(); console.log(final.usage); ``` ## OpenAI SDK ```typescript import OpenAI from "openai"; const client = new OpenAI({ apiKey: process.env.CLAUDEAPIKEY!, baseURL: "https://claudeapikey.dev/v1", // with /v1 }); const r = await client.chat.completions.create({ model: "claude-sonnet-4-6", messages: [{ role: "user", content: "Hello" }], }); console.log(r.choices[0].message.content); ``` > Never ship a key in browser-side code. Call the gateway from your server and expose your own endpoint to the front end — anything in a bundle is public. - [Chat Completions](https://claudeapikey.dev/docs/api-reference/chat-completions/) — OpenAI-format reference - [Best practices](https://claudeapikey.dev/docs/guides/best-practices/) — Production concerns --- _ClaudeAPIKey.dev is an independently operated, Anthropic-compatible API gateway. Not affiliated with Anthropic._