AI Gateway
Send your own code's model requests through g1t in Anthropic's or OpenAI's format, with a workspace access token, to Claude, open models or your own providers, with a log of every request.
The AI Gateway takes model requests from your own code and sends them to the model. It speaks two formats, and any model works in either:
| Format | Base URL | For |
|---|---|---|
| Anthropic’s Messages API | https://models.g1t.sh/anthropic |
Anthropic’s SDKs, Claude Code, and anything else that speaks that format |
| OpenAI’s Chat Completions API | https://models.g1t.sh/openai/v1 |
OpenAI’s SDKs, and any tool that lets you set an OpenAI-compatible base URL |
In both, the API key is a workspace access token (g1t_…) with the
models:write scope.
The model a request names decides where it goes:
- g1t’s models: Claude on Anthropic, and open models on Workers AI. Each request is charged to the workspace at the model’s price and paid from the plan’s included usage and AI credit. While the gateway is in beta there is no markup.
- Your own providers: an Anthropic key, an OpenAI key, or any endpoint that speaks either API, connected under Integrations. You choose which models go to each. Those requests are counted and never charged on g1t.
Every request is logged with its format, who served it, its model, tokens, cost and status. Prompts and answers are never kept.
Before you start
Section titled “Before you start”You need one of these:
- The workspace on the g1t plan, with AI credit or this month’s included usage left. See AI credit.
- One of the workspace’s own model providers, connected under Integrations. Requests for the models it takes need no plan and cost nothing on g1t.
Make a token
Section titled “Make a token”The gateway takes a workspace’s own token, so its usage is the workspace’s and keeps working when the person who set it up leaves. Only owners make them.
- Open the workspace’s Settings → Access tokens.
- Under New token, give it a name, such as
release-notes. The log shows each request’s token by this name. - The scopes start on the CI preset. Untick what the code does not need,
and under AI Gateway tick
models:write. A token with onlymodels:writecan send model requests and nothing else. - Choose an expiry and select Create token. Copy the token now: it is not shown again.
A personal access token is refused, even with models:write: the gateway
has to know which workspace to charge. A token with full access has every
scope, models:write included. See scopes.
Send a request
Section titled “Send a request”The token goes in x-api-key or in Authorization: Bearer, in either
format. The gateway answers these routes:
| Route | What it does |
|---|---|
POST /anthropic/v1/messages |
A message, streamed ("stream": true) or whole. Logged and charged. |
POST /anthropic/v1/messages/count_tokens |
Counts a request’s input tokens. Not logged, and costs nothing. For a model that does not speak Anthropic’s API, an estimate. |
POST /openai/v1/chat/completions |
A chat completion, streamed or whole. Logged and charged. |
POST /openai/v1/embeddings |
Embeddings, from an embeddings model. Logged and charged by their input tokens. |
GET /openai/v1/models |
The models this workspace can use, with g1t’s prices. |
Anthropic’s format
Section titled “Anthropic’s format”With curl:
curl https://models.g1t.sh/anthropic/v1/messages \ -H "x-api-key: $G1T_TOKEN" \ -H "anthropic-version: 2023-06-01" \ -H "content-type: application/json" \ -d '{ "model": "claude-haiku-5-5", "max_tokens": 1024, "messages": [{ "role": "user", "content": "Write a commit message for: fix the login redirect" }] }'With Anthropic’s TypeScript SDK:
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({ baseURL: "https://models.g1t.sh/anthropic", apiKey: process.env.G1T_TOKEN,});
const message = await client.messages.create({ model: "claude-haiku-5-5", max_tokens: 1024, messages: [{ role: "user", content: "Write a commit message for: fix the login redirect" }],});With Anthropic’s Python SDK:
import os
import anthropic
client = anthropic.Anthropic( base_url="https://models.g1t.sh/anthropic", api_key=os.environ["G1T_TOKEN"],)
message = client.messages.create( model="claude-haiku-5-5", max_tokens=1024, messages=[{"role": "user", "content": "Write a commit message for: fix the login redirect"}],)Request and answer bodies are Anthropic’s, and so are streamed events. To
an open model, such as workers-ai/@cf/openai/gpt-oss-120b, the request is
translated: messages, system prompt, images, tools and tool results,
tool_choice, stop sequences, output_config.effort (as
reasoning_effort, xhigh and max as high) and output_config.format
(as a JSON schema). The answer comes back as an Anthropic message, tool
calls included. Server tools have no counterpart there and are refused on
g1t’s models.
OpenAI’s format
Section titled “OpenAI’s format”With curl:
curl https://models.g1t.sh/openai/v1/chat/completions \ -H "Authorization: Bearer $G1T_TOKEN" \ -H "content-type: application/json" \ -d '{ "model": "anthropic/claude-haiku-5-5", "messages": [{ "role": "user", "content": "Write a commit message for: fix the login redirect" }] }'With OpenAI’s TypeScript SDK:
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://models.g1t.sh/openai/v1", apiKey: process.env.G1T_TOKEN,});
const completion = await client.chat.completions.create({ model: "workers-ai/@cf/openai/gpt-oss-120b", messages: [{ role: "user", content: "Label this issue: the login page is blank on Safari" }],});With OpenAI’s Python SDK:
import os
from openai import OpenAI
client = OpenAI( base_url="https://models.g1t.sh/openai/v1", api_key=os.environ["G1T_TOKEN"],)
completion = client.chat.completions.create( model="anthropic/claude-sonnet-5-5", messages=[{"role": "user", "content": "Summarize this diff in one sentence."}], stream=True, stream_options={"include_usage": True},)for chunk in completion: print(chunk.choices[0].delta.content or "" if chunk.choices else "", end="")Embeddings:
curl https://models.g1t.sh/openai/v1/embeddings \ -H "Authorization: Bearer $G1T_TOKEN" \ -H "content-type: application/json" \ -d '{ "model": "workers-ai/@cf/baai/bge-m3", "input": ["fix the login redirect"] }'What OpenAI’s format supports, to any model:
| In the request | |
|---|---|
messages |
system, developer, user, assistant and tool messages. User content can be text, image_url (a URL or a data: URL) and, to Claude, file with file_data (a PDF as a data: URL). |
tools, tool_choice, parallel_tool_calls |
Function tools. To Claude, required is any and a named function is that tool. |
stream, stream_options.include_usage |
Server-sent chunks, ending with data: [DONE]. With include_usage, a last chunk carries the usage. |
max_tokens, max_completion_tokens |
To Claude, 8,192 when neither is given. |
temperature, top_p, stop, user |
As given. Some models refuse sampling settings. |
reasoning_effort |
To Claude, output_config.effort; minimal and none are low. |
response_format |
json_schema is a JSON schema the answer follows. json_object asks Claude for one JSON object. |
thinking |
To Claude, passed as it is, for a caller that sets Anthropic’s thinking. |
To Claude, the answer’s usage counts cached tokens in prompt_tokens,
with prompt_tokens_details.cached_tokens; Claude’s thinking comes back as
reasoning_content. Claude’s thinking blocks must go back with the tool
calls they led to, so the gateway carries them in the first tool call’s
id: send the id back unchanged in the assistant message and the tool
message, as OpenAI’s SDKs do. n above 1 is refused for Claude.
Claude Code and other tools
Section titled “Claude Code and other tools”Claude Code speaks Anthropic’s format. Set two environment variables before you start it:
export ANTHROPIC_BASE_URL=https://models.g1t.sh/anthropicexport ANTHROPIC_AUTH_TOKEN=g1t_…claudeClaude Code’s own small requests go to Claude Haiku 4.5, which the gateway
offers. To choose the main model, also set ANTHROPIC_MODEL, such as
claude-opus-5-5, or ANTHROPIC_SMALL_FAST_MODEL, such as
claude-haiku-5-5.
Any other tool that lets you set an OpenAI-compatible base URL and key
works the same way: give it https://models.g1t.sh/openai/v1 and the
token, and a model id from the models list.
Models
Section titled “Models”Model ids
Section titled “Model ids”A request names a model:
| Id | Goes to |
|---|---|
anthropic/claude-sonnet-5-5 |
Claude on g1t’s account, in either format |
claude-sonnet-5-5 |
The same, as Anthropic’s API names it |
workers-ai/@cf/openai/gpt-oss-120b |
An open model on g1t’s account, in either format |
@cf/openai/gpt-oss-120b |
The same |
Any id one of your own providers takes, such as gpt-5.5 or ollama/llama3.3 |
That provider, with its key. See your own providers. |
Your own providers come first: when one of them takes a model, the request
goes there, even one that names a model g1t offers. An Anthropic key takes
claude-* unless you choose otherwise, so with one connected, Claude goes to
your key.
GET /openai/v1/models lists what the workspace can use: its own
providers’ models first, then g1t’s, starting with the Claude g1t suggests
starting with (Claude Haiku 5.5 today). g1t adds models as providers
release them, once their prices are confirmed, so the list and the tables
below grow over time; a model a provider stops offering is listed until it
is retired. Each has
billed_to (workspace or g1t), connection (your provider’s name) and,
on g1t’s models, pricing in dollars per million tokens:
{ "object": "list", "data": [ { "id": "anthropic/claude-haiku-5-5", "object": "model", "created": 0, "owned_by": "anthropic", "name": "Claude Haiku 5.5", "kind": "chat", "billed_to": "g1t", "connection": null, "pricing": { "currency": "usd", "input": 0.1, "output": 0.5, "cache_read": 0.01, "cache_write": 0.125, "cache_write_1h": 0.2, "long_prompt": { "above_tokens": 100000, "input": 0.5, "output": 2.5, "cache_read": 0.05, "cache_write": 0.625, "cache_write_1h": 1 } } } ]}Claude, on Anthropic
Section titled “Claude, on Anthropic”Prices are per million tokens, Anthropic’s list price. Cache writes are
five-minute ones; one-hour cache writes ("ttl": "1h") cost twice the
input price.
| Model | model |
Input | Output | Cache reads | Cache writes | One-hour cache writes |
|---|---|---|---|---|---|---|
| Claude Haiku 5.5 | claude-haiku-5-5 |
$0.10 | $0.50 | $0.01 | $0.125 | $0.20 |
| Claude Haiku 5.5, prompts over 100,000 tokens | claude-haiku-5-5 |
$0.50 | $2.50 | $0.05 | $0.625 | $1.00 |
| Claude Sonnet 5.5 | claude-sonnet-5-5 |
$2.00 | $10.00 | $0.10 | $2.50 | $4.00 |
| Claude Opus 5.5 | claude-opus-5-5 |
$4.00 | $20.00 | $0.20 | $5.00 | $8.00 |
| Claude Haiku 4.5 | claude-haiku-4-5, claude-haiku-4-5-20251001 |
$1.00 | $5.00 | $0.10 | $1.25 | $2.00 |
Claude Haiku 5.5 is the cheapest Claude and the one to start with. It is
priced by the prompt’s length: a request whose prompt (its input, cache
read and cache write tokens) is longer than 100,000 tokens is charged
entirely at the higher prices. It takes effort, like Opus: set
output_config.effort in Anthropic’s format, or reasoning_effort in
OpenAI’s.
Open models, on Workers AI
Section titled “Open models, on Workers AI”Prices are per million tokens, Cloudflare’s list price. Workers AI has no prompt-cache price: cached tokens, where a model reports them, cost what input does.
| Model | model |
Input | Output |
|---|---|---|---|
| GLM-5.3 Flash | workers-ai/@cf/zai-org/glm-5.3-flash |
$0.15 | $0.50 |
| gpt-oss-20b | workers-ai/@cf/openai/gpt-oss-20b |
$0.20 | $0.30 |
| Llama 4 Scout | workers-ai/@cf/meta/llama-4-scout-17b-16e-instruct |
$0.27 | $0.85 |
| gpt-oss-120b | workers-ai/@cf/openai/gpt-oss-120b |
$0.35 | $0.75 |
| Mistral Small 3.1 | workers-ai/@cf/mistralai/mistral-small-3.1-24b-instruct |
$0.351 | $0.555 |
| DeepSeek V4 Flash | workers-ai/@cf/deepseek-ai/deepseek-v4-flash-0731 |
$0.44 | $1.32 |
| Nemotron 3 120B | workers-ai/@cf/nvidia/nemotron-3-120b-a12b |
$0.50 | $1.50 |
| Kimi K2.6 | workers-ai/@cf/moonshotai/kimi-k2.6 |
$0.95 | $4.00 |
| DeepSeek V4 Pro | workers-ai/@cf/deepseek-ai/deepseek-v4-pro-0813 |
$1.32 | $3.96 |
| GLM-5.3 | workers-ai/@cf/zai-org/glm-5.3 |
$1.40 | $4.40 |
Embeddings, through POST /openai/v1/embeddings only:
| Model | model |
Input |
|---|---|---|
| BGE M3 | workers-ai/@cf/baai/bge-m3 |
$0.012 |
| BGE Base (English) | workers-ai/@cf/baai/bge-base-en-v1.5 |
$0.067 |
Open models cost much less per call than Claude, and suit one-shot work: titles, summaries, labels, triage, embeddings. In a long loop that sends the same context every turn, Claude’s cache reads close most of that gap.
What is refused on g1t’s models
Section titled “What is refused on g1t’s models”A request for a model nobody offers is refused with 404 before it
reaches a provider, and the error names the models offered. On g1t’s
models a request is charged only by its tokens, so what a provider bills
some other way is refused with 400 for now:
| Not offered on g1t’s models yet | In the request |
|---|---|
| Fast mode | speed other than standard |
| Inference in one region | inference_geo other than global |
| Server-side fallbacks | fallbacks |
| Server tools, such as web search, web fetch and code execution | In Anthropic’s format, a tool whose type is not your own (custom or none) or a client tool (bash_…, text_editor_…, computer_…, memory_…). In OpenAI’s, a tool that is not a function, or web_search_options. |
| Containers and skills | container |
All of them work on your own provider, which bills them. In Claude Code on g1t’s models, its web search fails for this reason; the rest of Claude Code works.
Your own providers
Section titled “Your own providers”Connect a model provider under Integrations, and choose which models your own code’s gateway requests send to it. Requests there use its key, are counted in the log, and are never charged on g1t. Any provider works:
| Provider | Takes, unless you choose |
|---|---|
| An Anthropic key, or an Anthropic-compatible endpoint | claude-* |
| An OpenAI key, or any other provider | Nothing until you choose |
| An OpenAI-compatible endpoint: a self-hosted vLLM or Ollama, LiteLLM, another provider | Nothing until you choose |
-
Open the workspace’s Integrations and choose a provider under Model providers. For your own server, choose OpenAI-compatible endpoint or Anthropic-compatible endpoint and give its base URL.
-
Paste its key. It is sealed when saved and never shown again: the page, the API and MCP show only its last four characters.
-
Under AI Gateway models, list the models to send there, separated by spaces:
Write Takes gpt-5.5That model only gpt-*Every model whose id starts with gpt-ollama/*Every model named ollama/…, sent without the prefix:ollama/llama3.3arrives asllama3.3*Every model Nothing No gateway requests -
Select Connect. To change the list or replace the key later, open Change its AI Gateway models or key under the provider.
The first provider, in the order they were connected, that takes a model gets its requests. Either format reaches either kind of provider: a Claude key answers OpenAI-format requests, and an OpenAI-compatible endpoint answers Claude Code. Agent runs choose their models under routing, apart from this list.
From code, connect one with
POST /workspaces/{workspace}/integrations
and change it with
PATCH /workspaces/{workspace}/integrations/{id},
with config.gateway_models, or the workspace MCP tool’s
connect_integration and update_integration actions. Both need
workspace:admin and act for an owner. The key is write-only: neither
returns it.
curl -X PATCH https://api.g1t.sh/workspaces/acme/integrations/con_01kpx5c2d8e4f6g0h2j4k6m8n0 \ -H "Authorization: Bearer $G1T_ADMIN_TOKEN" \ -H "content-type: application/json" \ -d '{ "config": { "base_url": "https://gpu.acme.dev/v1", "gateway_models": ["ollama/*"] }, "secret": "…" }'config replaces the provider’s settings whole, so send the ones it has
with the change.
What it costs
Section titled “What it costs”| Where it goes | You pay |
|---|---|
| g1t’s models | Its tokens at the model’s price above, with no markup while the gateway is in beta |
| Your own providers | Nothing on g1t. The provider bills you for the model. |
On g1t’s models:
- Each request that used tokens is one line on the statement, under AI Gateway, such as AI Gateway: Claude Sonnet 5.5, 14,352 tokens, token release-notes. A Claude Haiku 5.5 request over 100,000 prompt tokens says long-prompt price.
- The plan’s included usage pays first, then AI credit. Trial credit and g1t’s open-source pool never pay for gateway requests.
- It is not an agent run, so the agent rate does not apply.
- It counts toward the workspace’s spend limit like any other usage, and shows on Usage under the AI Gateway product.
- A workspace with a 100% discount gets it free through the discount; an enterprise is invoiced for it after use.
Limits and errors
Section titled “Limits and errors”A request on g1t’s models is refused before it reaches the model when:
- The workspace is over its spend limit.
- It is on the plan, and has no AI credit and none of this month’s included usage left. If auto-reload is on, g1t tries it first.
- It is not on the g1t plan.
Errors are in the format of the route, so SDKs raise their usual errors. Anthropic’s:
{ "type": "error", "error": { "type": "billing_error", "message": "The acme workspace is out of AI credit …" } }OpenAI’s:
{ "error": { "message": "The acme workspace is out of AI credit …", "type": "insufficient_quota", "param": null, "code": "insufficient_quota" } }| Status | Anthropic’s error.type |
OpenAI’s error.type (code) |
Why |
|---|---|---|---|
400 |
invalid_request_error |
invalid_request_error |
The body is not JSON, the model is of the wrong kind, the request asks for something not offered on g1t’s models yet, or it cannot be said to the model (such as n above 1 to Claude). |
401 |
authentication_error |
authentication_error (invalid_api_key) |
The token is unknown, expired or deleted, or your provider refused its key. |
402 |
billing_error |
insufficient_quota (insufficient_quota) |
Out of AI credit, over the spend limit, or not on the plan. The message says what an owner can do. |
403 |
permission_error |
permission_error |
Not a workspace’s token, or it lacks models:write. |
404 |
not_found_error |
invalid_request_error (model_not_found) |
No provider offers the model, or a route the gateway does not answer. |
An error from the model provider, such as 429 or 529, comes back with
its status and message, in the route’s format. A provider’s key never
appears in an error or the log, even when the provider quotes it. Refused
and failed requests are logged with their status and why, and cost
nothing. Every answer carries x-g1t-request-id, the request’s id in the
log.
A deleted token, a provider added or changed under Integrations, or AI credit just bought takes effect within about ten seconds.
See every request
Section titled “See every request”The AI Gateway page lists the workspace’s requests, newest first. Open
it from the link under Usage, or at g1t.sh/<workspace>/-/gateway.
Every member can see it.
| Column | |
|---|---|
| Time | When it was sent. Hover for the exact time, how long it took and whether it streamed. |
| Model | The model it named on g1t’s models, or the one that answered on your own provider; below it, the format it was sent in. |
| Served by | g1t’s account and the provider (g1t · Anthropic, g1t · Workers AI), or your provider by name. None when it was refused first. |
| Input, Output | Its tokens by kind. Hover Input for all of them. |
| Cache | Cache reads, then cache writes. Hover for how many writes were to the one-hour cache. |
| Cost | What it was charged, before included usage and AI credit paid for it, or Not charged on your own provider. |
| Status | The status it was answered with. Hover a refusal or failure for why. |
| Token | The name of the token that sent it. |
Requests are kept 30 days.
From code, list them with
GET /workspaces/{workspace}/gateway/requests,
or the billing MCP tool’s
gateway_requests action. Each request has
format, provider, connection, model and its tokens, with
cache_write_hour for one-hour cache writes. Both need models:read,
which the Read only and Agent presets include.