Skip to main content
A Claude apps gateway deployment is configured by one YAML file, conventionally gateway.yaml. The file defines everything the gateway does: where it listens, how developers sign in, where inference goes, and which policies and telemetry apply. This page is the reference for every option in that file. To write your first one, start from the quickstart, which builds a minimal working config and runs it. Once you have a config you’re happy with, the deployment guide covers containerizing and hosting it on Kubernetes, Cloud Run, or your own platform. The gateway reads the file once, at startup, with claude gateway --config /path/to/gateway.yaml. Every option is validated against a schema at boot, so a malformed config fails at start with a field-level error rather than at first use. The complete example at the end of this page exercises every section.

File structure

Five sections are required. Every other section is optional, and an omitted section takes its defaults. Unknown keys fail boot, so a typo surfaces as a named error rather than a silently ignored setting. Required sections:
  • listen: bind address, public URL, TLS termination
  • oidc: your identity provider (IdP), including issuer, client, claim mapping, and who may sign in
  • session: the bearer tokens the gateway mints, with secret and lifetime
  • store: PostgreSQL, for device grants and rate-limit counters
  • upstreams: where inference goes, whether Anthropic, Amazon Bedrock, Claude Platform on AWS, Google Cloud’s Agent Platform, or Microsoft Foundry
Optional sections:
  • admin: Admin API auth and retention for spend limits
  • enforcement: spend-limit fail-open or fail-closed behavior
  • pricing: contracted rates and a multiplier for the spend meter and for the cost figures developers see
  • models and auto_include_builtin_models: admin-curated model list and per-upstream IDs
  • managed: managed settings policies by IdP group
  • telemetry: OTLP forwarding to your observability stack
  • access_control, limits, timeouts, rate_limits: IP allow/deny, request size caps, upstream time-to-first-byte, and per-IP sign-in limits
  • load_test_mode: load testing the gateway without calling a model provider

Secret expansion

Don’t write secrets such as client_secret, jwt_secret, or postgres_url directly in gateway.yaml. Reference them with one of the forms below, and the gateway resolves the value at boot from an environment variable or a file:

Required sections

listen

The listen block controls where the gateway serves: the bind address and port, the externally visible origin, and optional TLS termination.

oidc

The oidc block connects the gateway to your identity provider and decides who can sign in. It names the issuer and OAuth client, maps the claims that carry email and groups, and restricts sign-in by email domain or group. OpenID Connect (OIDC) is the SSO protocol the gateway uses with your identity provider; see Identity provider setup for what to register on the IdP side.

IdP requests through a forward proxy

The inference upstreams honor HTTPS_PROXY and HTTP_PROXY on every version. The gateway’s own requests to the IdP, discovery, JWKS, token, and userinfo, go direct unless you set oidc.use_proxy: true, which requires v2.1.227 or later. When a proxy variable is set, use_proxy is unset, and the issuer isn’t covered by NO_PROXY, the gateway keeps those requests direct and logs a notice at boot asking you to choose; use_proxy: false keeps them direct and silences the notice. With use_proxy: true, the pod resolves each IdP endpoint’s hostname itself and asks the proxy to CONNECT to the resolved IP address, so the proxy must accept CONNECT to the IP address of every host the discovery document names, not only the issuer. Use an http:// proxy URL. ca_cert_pem and the SSRF guard apply on the proxied path as well. Proxy-only egress changes both of these: while it’s active, IdP requests follow the proxy unless you set use_proxy: false, and the gateway hands the proxy each IdP hostname without resolving it first.

Proxy-only egress

Set CLAUDE_GATEWAY_PROXY_IS_EGRESS_BOUNDARY=1 in the gateway’s environment, next to HTTPS_PROXY, when the pod reaches other hosts only through that forward proxy and can’t resolve public DNS names itself, or when the proxy refuses CONNECT to an IP address. Requires v2.1.277 or later. It’s an environment variable rather than a gateway.yaml key so that nothing in the config file can relax the gateway’s address check.
The gateway logs one network: line at boot while proxy-only egress is active. Each row below is one class of outbound request on a gateway with HTTPS_PROXY set, by default and while proxy-only egress is active. Proxy-only egress stays off unless the gateway’s environment meets all three of these conditions:
  • HTTPS_PROXY or HTTP_PROXY is set.
  • NO_PROXY and no_proxy are empty. If your platform injects either into pods, set both to an empty value on the gateway container. Listing a telemetry collector in NO_PROXY keeps proxy-only egress off.
  • CLAUDE_GATEWAY_ALLOW_LOOPBACK isn’t turned on. A collector or IdP on the pod’s own loopback can’t be combined with proxy-only egress, because a loopback address handed to the proxy would be the proxy host’s own, so give those services an address the proxy can reach instead. For the same reason the gateway refuses localhost-style names outright while proxy-only egress is active.
When one of those conditions isn’t met, the gateway logs a warning at boot naming the variable that stopped it and keeps the default behavior. Once proxy-only egress is active, allow every destination in the proxy, including an internal collector and any host configured by IP address. You can still keep an internal IdP direct with oidc.use_proxy: false.
Turn this on only when the proxy’s allowlist is at least as strict as the gateway’s own check. The proxy must refuse cloud metadata endpoints such as 169.254.169.254 and metadata.google.internal, link-local addresses, and the proxy host’s own loopback, and it must refuse them by the address a name resolves to, not only by name, because the gateway no longer catches a hostname that resolves to one of them. A proxy that connects anywhere it’s asked removes the gateway’s SSRF guard for these requests.

session

The session block shapes the bearer tokens the gateway mints after sign-in: the secret that signs them and how long they live.

store

The store block points the gateway at its PostgreSQL database, which holds device grants and rate-limit counters. For local development, point postgres_url at a throwaway Postgres container, for example docker run --rm -p 5432:5432 -e POSTGRES_HOST_AUTH_METHOD=trust postgres.

upstreams

upstreams is an ordered list. The gateway forwards inference to the first upstream that resolves the requested model. On 5xx, 429, 401, 403, 404, or timeout the gateway fails over to the next upstream; other 4xx doesn’t, because those errors are attributable to the request rather than the upstream. A 401 or 403 means the gateway’s own credential failed against that upstream. A 404 means that upstream doesn’t serve the requested model, so a later upstream in the list still can. If you set forward_user_identity: true on an upstream, a 429 it returns to a request that carried the developer’s email doesn’t fail over. See how a per-user limit denial reaches the developer. Failover on 404 requires gateway v2.1.198 or later. Earlier releases returned the first 404 to the client even when a later upstream in the list served the model. Multiple upstreams of the same provider must set a distinct name:. Amazon Bedrock, Claude Platform on AWS, Google Cloud’s Agent Platform, and Microsoft Foundry clients are built once at startup, and their SDKs refresh credentials internally, so rotating cloud credentials doesn’t require a restart. Static Anthropic API keys and bearers are read at startup; see Anthropic API.

Upstream error messages

The gateway returns one upstream’s error response, or its own 502, depending on how the upstreams answered:
  • An upstream returned a status the gateway doesn’t fail over on: that upstream’s response. The gateway tries no further upstreams.
  • Every upstream the gateway tried failed in a way it fails over on: the last 429. When none returned a 429, the gateway prefers, in order, the last 401 or 403, the last 404, and the last 501. When none returned any of those, the gateway’s own 502, all upstreams failed (N attempted), where N counts every entry in upstreams, including entries the gateway skipped because they don’t serve the requested model.
When the gateway returns an upstream’s response, it keeps the upstream’s status code. Whether it keeps the upstream’s message depends on the provider. An Anthropic API upstream’s error body reaches the developer unchanged. The Amazon Bedrock, Claude Platform on AWS, Google Cloud’s Agent Platform, and Microsoft Foundry upstreams can name your account IDs, role ARNs, and project IDs in their error text. The gateway records that full text in the operational log. What the developer sees from those upstreams depends on the rejection:
  • 400 or 413 in Anthropic’s standard error envelope: the upstream’s own message, such as prompt is too long. Claude Platform on AWS, Agent Platform, and Microsoft Foundry return this envelope for model API rejections.
  • 400 or 413 in the provider’s own shape: a capability_rejected: token. When the gateway can’t classify the rejection, upstream rejected the request on a 400 or request too large for this upstream on a 413.
  • Any other status: generic per-status copy, such as upstream rate limit exceeded on a 429.
For example, the gateway replaces Amazon Bedrock’s Input is too long for requested model. with capability_rejected: prompt_too_long. Claude Code compacts automatically on that token, as it does on prompt is too long. Keeping a cloud upstream’s 400 or 413 message, or replacing it with a capability_rejected: token, requires gateway v2.1.233 or later.

Anthropic API

The minimal Anthropic upstream is an API key from the Claude Console:
The two credential forms differ in the header they send:
  • api_key: sends x-api-key. Rotate it in the Claude Console and update the env var.
  • oauth_token: sends Authorization: Bearer. Use the bearer form when your org issues short-lived tokens instead of long-lived API keys. The bearer is read once at startup, so refresh by remounting the secret and restarting.
Instead of a static key or bearer, you can use Workload Identity Federation. Create a federation rule by following the Workload Identity Federation guide, then mount your workload’s OIDC JWT as a file, such as a Kubernetes projected service-account token or a CI platform’s id-token. The gateway exchanges the JWT for a short-lived bearer and refreshes it automatically. The token file is re-read on every exchange, so rotated projected tokens are picked up without a restart.
Per-user identity headers for a proxy you run
You can point a provider: anthropic upstream’s base_url at a proxy you run instead of at the Anthropic API. To tell that proxy which developer sent each request, set forward_user_identity: true on that upstream. The proxy can then attribute spend per developer. Requires a gateway running Claude Code v2.1.233 or later. For example, for a proxy at upstream-gateway.internal.example.com:
The gateway adds these headers to every request it forwards to that upstream. When the IdP token carries no email, the gateway sends only x-claude-gateway-user-id and omits the two email headers. If your IdP puts the email in a different claim, set oidc.email_claim to that claim. When your proxy answers 429 to a request that carried the developer’s email, the gateway returns that response to the developer as-is instead of failing over to the next upstream, so your proxy’s per-user budget or rate limit holds. The proxy’s other responses follow the ordinary failover rules. If a developer’s IdP token carries no email, the gateway forwards their requests without the email headers, so a 429 to one of those requests counts as upstream capacity and fails over. Before v2.1.267 on the gateway server, every 429 failed over. Set forward_user_identity only on an upstream whose base_url is a proxy you operate. The gateway sends developer emails to whatever server that base_url names. If the base_url is the Anthropic API, which is the default, the gateway refuses to start.

Amazon Bedrock

For the client-side Amazon Bedrock deployment that the gateway replaces or fronts, see Claude Code on Amazon Bedrock. The gateway-side upstream:
An empty auth block uses the AWS SDK’s default credential chain: env vars, ~/.aws/credentials, ECS task role, EC2 instance metadata, or IRSA on EKS. In production, give the gateway pod an IAM role instead of embedding static keys in a container image. Explicit credentials must be complete: the gateway fails at boot when aws_access_key_id and aws_secret_access_key aren’t set together, or when aws_session_token is set without them. Before v2.1.207, a partial auth: block passed validation.

Claude Platform on AWS

Claude Platform on AWS serves the first-party Anthropic API on AWS infrastructure at aws-external-anthropic.<region>.api.aws. It uses first-party model IDs, honors anthropic-beta headers as sent, and serves count_tokens, so none of the Bedrock-specific translation applies. The anthropicAws provider requires Claude Code v2.1.198 or later; earlier gateway releases reject it at boot. For the client-side deployment of the same platform, see Claude Code on Claude Platform on AWS. The gateway-side upstream:
The platform runs in a separate AWS account from Amazon Bedrock and signs SigV4 requests for its own service name, aws-external-anthropic, so a Bedrock-scoped IAM role doesn’t authorize it. An API key in auth.api_key takes precedence when SigV4 credentials are also set. An empty auth block uses the AWS SDK’s default credential chain, the same chain the Amazon Bedrock upstream uses. Because the platform resolves first-party model IDs, the built-in catalog routes to it with no models: block. When you curate a models: list, key the entry anthropicAws: with the first-party ID.

Google Cloud Agent Platform

For the equivalent client-side setup, see Claude Code on Google Cloud. The gateway-side upstream:
An empty auth block uses Application Default Credentials: GOOGLE_APPLICATION_CREDENTIALS, GCE metadata, or GKE Workload Identity. Service-account JSON key files are supported but discouraged; use Workload Identity or attach a service account to the GCE or Cloud Run instance. Set region: global to use the global endpoint for Google Cloud’s Agent Platform instead of a regional one. Google then routes each request to an available region, so you don’t track per-region model availability. Setting a specific region pins every request to it.

Microsoft Foundry

For the client-side Microsoft Foundry deployment, see Claude Code on Microsoft Foundry. The gateway-side upstream:
use_azure_ad: true resolves through DefaultAzureCredential: Managed Identity on AKS, ACI, or App Service; the Azure CLI; or environment credentials. API keys work but are project-wide and don’t rotate automatically. Microsoft Foundry’s endpoint is derived from resource:; set the optional base_url to override it for sovereign clouds such as Azure Government.

Static headers on upstream requests

To add fixed headers to the requests the gateway sends to one upstream, set headers: on that upstream. Use it when a proxy you run in front of the provider routes or attributes traffic by a header. headers: requires Claude Code v2.1.277 or later on the gateway server. An earlier gateway refuses to start when it finds the key. Upgrade every replica before you add the key, and remove the key before you roll back to an earlier version. The headers go to the server that base_url names, or to the provider’s own endpoint when base_url is unset. The provider receives them too unless your proxy removes them. This example reaches a provider: vertex upstream through a proxy at upstream-proxy.internal.example.com. It sets the x-source header the proxy reads, and sends a token from the PROXY_TOKEN environment variable as x-proxy-token:
Values are printable ASCII text with no space at either end. Quote a number, true, or false so YAML reads it as text. To keep a secret out of the config file, use secret expansion to load the value from an environment variable with ${VAR} or from a file with ${file:/path}. A ${VAR} that resolves to an empty value stops the gateway from starting. headers: works on every provider, and each upstream sends only its own. Not every request that the gateway sends to an upstream carries them: On an Amazon Bedrock or Claude Platform on AWS upstream that signs requests with AWS SigV4, these headers are part of the signature, so your proxy must pass them through unchanged. If you use a name the gateway reserves, it refuses to start, and the startup error names the header. Reserved names include:
  • authorization and x-api-key
  • host, content-type, and user-agent
  • Any name starting with anthropic-, x-goog-, x-amz-, or x-amzn-

Multiple upstreams

The same provider can appear more than once with a distinct name:. This covers different regions, different accounts via different credential chains, provisioned throughput versus on-demand, and cross-provider fallback. The gateway tries upstreams in order. 5xx, 429, 401, 403, 404, timeouts, and missing-endpoint (501) fail over; other 4xx doesn’t. 429 is per-upstream capacity, so provisioned-throughput (PT) exhaustion fails over to on-demand. If you set forward_user_identity: true on an upstream, a 429 to a request that carried the developer’s email is a per-user denial instead and doesn’t fail over. Every request starts at the first upstream. A request reaches a later upstream only when every upstream ahead of it has failed or doesn’t serve the requested model. The gateway keeps no record of failed upstreams, so while an upstream is down, every request that reaches it still tries it and waits for it to fail before moving on. For an Anthropic API upstream, timeouts.upstream_ttfb_ms bounds the wait on a down upstream. That setting doesn’t apply to the other providers, where the gateway waits up to one hour for an upstream to start responding. 404 is per-upstream model availability, so an upstream that hasn’t enabled a model doesn’t block a later upstream that serves it. An upstream that can’t resolve the requested model is skipped without a network round-trip. This example routes a provisioned-throughput Amazon Bedrock allotment first, overflows to on-demand and a second account, and falls back to the Anthropic API last:
Failing over between cloud providers, or to the direct Anthropic API, changes which agreement, geography, and other terms govern the request. The CLI applies the same feature gating to gateways regardless of which upstream serves a given request, so failover doesn’t send a body field an upstream would reject.

Optional sections

admin

Optional. Enables /v1/organizations/spend_limits, which mirrors Anthropic’s public Admin API, and per-developer spend enforcement on /v1/messages. See Spend limits for how caps are set and enforced; this section covers the gateway.yaml keys that turn the feature on and tune it.

enforcement

The enforcement block controls how spend-limit checks behave when the store is unavailable.

pricing

The pricing block tells the spend meter what to charge instead of USD list price, so caps and /effective reflect your contracted rates. Amounts stay in USD and remain an estimate, not an invoice. Two prerequisites:
  • Claude Code v2.1.227 or later on the gateway server. Earlier versions reject the unknown key at boot.
  • An admin: block or, in v2.1.268 or later, a managed: block with at least one policy. The gateway refuses to start with pricing set and neither block, because nothing would read it.
How the meter matches an override row:
  • A row replaces list price for requests that upstream, an upstreams[].name, serves for model. That includes the higher fast mode rate, so fast and standard requests meter at the same four rates.
  • A built-in ID such as claude-sonnet-4-6, matched like models[].id, covers every dated form, regional Amazon Bedrock form, or Google Cloud’s Agent Platform form the meter prices as that model. Any other string, such as an alias or an inference-profile ARN, matches the ID the client sent or the string sent upstream, case-insensitively.
  • Where rows overlap, the meter picks the most specific row rather than the first row: a row whose model is the exact model string sent upstream, then a row matching the exact ID the client sent, then a row naming the built-in model.
  • An unknown upstream name fails boot, and so do two rows for one upstream that name the same model, including two spellings of one built-in model. The gateway warns at boot about a row no requestable model can use.
  • Web-search requests stay at the $0.01 list price; the multiplier still applies to them.
For per-region rates, give each region its own named upstream and one row per upstream.

Mark prices up

With v2.1.271 or later on the gateway server, you can set multiplier above 1, up to 10, to meter more than the provider charges, for example an internal chargeback rate. This example meters every request at 120% of the price:
With an admin: block, the markup also applies to spend limits. The meter counts 120% of the price, so developers reach their caps sooner. The gateway logs a warning at boot that says so. The multiplier doesn’t change what the upstream provider charges for the requests. If the gateway also sends the rates to signed-in clients, developers need Claude Code v2.1.271 or later to see the markup. Earlier clients ignore a multiplier above 1 and show costs without it. A gateway server earlier than v2.1.271 refuses to start if you set a multiplier above 1.

Send the rates to signed-in clients

With v2.1.268 or later on the gateway server, the gateway also puts the rates from pricing into the managed policies it serves, as the modelPricing managed setting. Developers matched by a policy then see the pricing rates for the first upstream that serves each model ID in /usage, the status line, and OpenTelemetry. A developer who matches no policy receives no managed settings, so their figures stay at list price. Clients apply the setting in Claude Code v2.1.242 or later.
  • What the gateway adds: unless a policy’s cli block already sets modelPricing, the gateway adds the multiplier and, for every model ID a client can request, the override row of the first upstream that serves that ID. A rate that only a failover upstream charges stays on the gateway.
  • Opt one policy out: set modelPricing to {} in that policy’s cli block, and its developers stay at list price.
  • Keep a policy’s own rates: a policy whose cli block sets modelPricing with its own multiplier or overrides keeps that modelPricing whole, and the gateway adds no rates of its own to it.

models

The models block is an optional admin-curated model list, served at /v1/models and used to translate model IDs per upstream. It is required for non-US Amazon Bedrock regions, Amazon Bedrock provisioned-throughput ARNs, and Microsoft Foundry deployment names.
Each key under upstream_model must match the name of a configured upstream, which defaults to the provider name. A key that matches no upstream fails boot, so omit the lines for providers you don’t use.

managed

The managed block defines role-based access policies keyed on IdP groups or email domain. Policies are evaluated in order; the first match is selected, then merged onto the match: {} catch-all base. They are served per-user at GET /managed/settings with ETag/304 caching.
A match: {} catch-all, conventionally listed last, is treated as a base layer. Every other policy inherits any key it doesn’t set from the catch-all, so per-role entries only need to list what differs from the org default. The merge rules depend on the key type:
  • Allow-lists: availableModels and permissions.allow. A specific policy’s list fully replaces the base’s.
  • Deny-lists and hook arrays: permissions.deny, permissions.ask, disabledMcpjsonServers, deniedMcpServers, blockedMarketplaces, and every hooks event-type array. These take the union of base and policy, so an org-wide deny or audit hook can’t be accidentally dropped by a per-role override.
  • Record-typed keys: env, modelOverrides, and skillOverrides. These shallow-merge, so a per-role env block overrides keys it sets and inherits the rest from the base.
availableModels is also enforced server-side at /v1/messages, so a denied model returns 400 regardless of what the client sends. The gateway validates the model value itself before it relays a request, so a malformed value never reaches an upstream. It rejects the request with a 400 in two cases:
  • When the value is missing or empty, the gateway rejects the request with the message model is required. That check requires a gateway running Claude Code v2.1.228 or later.
  • When the value is present but isn’t a string, the gateway rejects the request with the message model must be a string. Requires a gateway running Claude Code v2.1.221 or later.
An authenticated user who matches no policy gets the gateway’s defaults, which means every model in the catalog and no managed settings. Add a match: {} catch-all last if you want a guaranteed default policy.
The gateway keeps no user directory of its own. It authorizes each request from the user’s IdP token, reading group membership from the token’s groups claim and evaluating policies against it. There is no roster to enumerate and no accounts to pre-create, and therefore no SCIM endpoint, because there is nothing for SCIM to sync into.Run user and group lifecycle management at the source of truth, which is your IdP’s native SCIM provisioning or a dedicated identity-governance platform. Membership and deprovisioning governed there flow into the gateway automatically through the token. If you want SCIM provisioning of Claude accounts themselves, that is a Claude for Enterprise capability.Two propagation clocks apply:
  • Policy contents: editing a policy and redeploying reaches connected clients on their next managed-settings poll, within an hour, apart from the changes that apply only at the next launch
  • Group membership: changing a user’s group membership changes which policy matches them. This takes effect on the next session re-mint, meaning the next silent refresh, bounded by session.ttl_hours.

Matcher values that stop the gateway at boot

At boot, the gateway checks the match block of every policy and the admin_groups list. Any of these values stops the gateway with an error that names the field:
  • An empty groups list
  • An empty entry in groups or in admin_groups
  • An empty email_domain
  • An email_domain that contains @, whitespace, or a comma. The gateway trims the value and strips one leading @ before this check. Write one bare domain, such as example.com.
Before v2.1.232, the gateway started with these values. Each value had this effect:
  • An empty email_domain: the gateway skipped the domain check, so a policy with an empty email_domain and no groups list matched every authenticated user
  • An empty groups list: the policy matched no one
  • An email_domain containing @, whitespace, or a comma: the policy matched no one
  • An empty entry in groups or in admin_groups: the entry matched a user only when that user’s IdP groups claim also contained an empty entry. In admin_groups, that match granted admin access. If your admin_groups list never contained an empty entry, no one gained admin access this way.

What goes in cli

Each cli value is a complete Claude Code managed-settings.json document, the same schema you would deploy via MDM or /etc/claude-code/managed-settings.json, expressed here as YAML. The CLI applies the delivered document at the managed tier, above user and project settings, in place of server-managed settings. It therefore ignores the settings restricted to OS-level policy sources, such as policyHelper and wslInheritsWindowsSettings. The gateway validates each document against the CLI’s settings schema at boot, so an unrecognized top-level key fails boot with an error naming every offending key. Deliberately open parts of the schema still accept arbitrary values, because newer clients may recognize entries the gateway’s schema doesn’t. These open keys include env, pluginConfigs, and keys nested under permissions. Because validation uses the schema bundled with the gateway’s installed version, putting a top-level settings key introduced by a newer Claude Code release into managed config requires upgrading the gateway first. Smoke-test a new policy on one client before rolling it out. The full key reference is in Claude Code settings. The keys most operators reach for first:
Because these settings arrive over the network, the CLI shows each developer a security approval dialog before applying the settings listed below:
  • hooks
  • env variables that require the developer’s approval, such as proxy and base-URL variables
  • shell-execution settings such as apiKeyHelper and statusLine
  • the sandbox binary settings sandbox.bwrapPath, sandbox.socatPath, and sandbox.ripgrep
  • Sandbox settings that intercept traffic, inject credentials, or weaken isolation, such as sandbox.network.tlsTerminate and the proxy port settings. Security approval dialogs lists them all.
Approval memory covers how long an approval lasts and when the dialog appears again. Claude Code applies some delivered env variables without showing the developer the approval dialog, such as model selection settings and numeric limits. Other delivered variables can require the developer’s approval before they take effect; a non-empty proxy, base-URL, or OTEL_EXPORTER_OTLP_ENDPOINT value always does. When a delivered variable needs approval, the dialog names it. Environment variables and the approval dialog has the details, including four privacy toggles whose delivered value decides whether they need approval. Before v2.1.218, Claude Code applied fewer variables without asking the developer, so more delivered variables triggered the dialog. The gateway’s telemetry configuration pushes OTEL_EXPORTER_OTLP_ENDPOINT, so setting telemetry.forward_to triggers the dialog on each interactive client. The dialog protects the developer’s machine from a compromised or hostile gateway, not the organization from the developer. A non-interactive run with the -p flag can’t show the dialog. It applies the pushed settings for that run only and doesn’t record them as approved, so the developer’s next interactive session still shows the dialog. Before v2.1.207, a non-interactive run saved the settings as approved and no later interactive session showed the dialog for them. If a developer declines, Claude Code exits that session rather than applying the policy. When you push a new hook, or any env var that triggers the dialog, to a broad policy, Claude Code therefore shows the dialog to every matching developer. It shows the dialog in a running session on the next hourly poll, and otherwise at the developer’s next startup. The cli key was named settings in earlier releases. That spelling is still accepted as an alias, but new deployments should use cli.

MCP servers in a policy

To provide MCP servers to the Claude Code clients a policy matches, set managedMcpServers in that policy’s cli block. You need Claude Code v2.1.259 or later on the gateway server and on clients. The gateway checks each entry at boot with the same rules Claude Code applies on the client, and if an entry fails a check, the gateway refuses to start and names the entry. If you write a ${VAR} reference in gateway.yaml, the gateway resolves it from its environment at boot through secret expansion before it runs the entry checks, so every matching client receives the literal value and can read it. The header guidance for provided servers applies to the expanded value. The gateway rejects the .mcp.json spelling mcpServers in a cli block, and its boot error names managedMcpServers as the key to use. Before v2.1.259, the gateway rejected any MCP server definition in a cli block.

Claude Desktop overlay

If your organization also deploys Claude Desktop, the same gateway serves both clients. Point bootstrapUrl, in Claude Desktop’s managed configuration, at <listen.public_url>/user/bootstrap. Claude Desktop derives the OAuth issuer from that URL, runs the same device-code sign-in against this gateway, and fetches its configuration from the response.
Requires Claude Code v2.1.203 or later on the gateway server, and an explicit opt-in: /user/bootstrap returns 404 unless the policy matching the user carries a desktop key. An empty desktop: {} opts a policy in, and a desktop key on the match: {} base layer opts in every policy that inherits it. The audit log records each request as desktop_bootstrap.serve or desktop_bootstrap.denied.
The gateway derives much of the response from the matched policy’s cli block and from top-level gateway config:
  • The model list, from availableModels
  • Disabled tools, from bare tool-name permissions.deny entries. If you set disabledBuiltinTools in the policy’s desktop block, the gateway serves the union of your value and the derived list, so you can disable more tools this way but can’t re-enable one you disabled through permissions.deny
  • The egress allowlist, from sandbox.network.allowedDomains. If you set coworkEgressAllowedHosts in the policy’s desktop block, the gateway uses that value instead of the derived list
  • An OTLP endpoint that points at the gateway itself, and the signed-in user’s identity attributes. The gateway relays the exports it receives at that endpoint to your forward_to destinations. It includes the endpoint and the attributes when you set both telemetry.forward_to and listen.public_url. Claude Desktop exports every signal with one encoding: http/protobuf, or http/json when you set OTEL_EXPORTER_OTLP_PROTOCOL or one of its per-signal variants to http/json in the policy’s env. Before Claude Code v2.1.261 on the gateway server, the response set http/json regardless, so a collector that accepts only protobuf rejected Claude Desktop’s exports
To set disabledBuiltinTools, coworkEgressAllowedHosts, or Claude Desktop’s own managedMcpServers setting in a policy’s desktop block, you need Claude Code v2.1.232 or later on the gateway server. Claude Desktop’s managedMcpServers takes an array value rather than an object. The gateway omits keys with no Claude Desktop equivalent, such as hooks and scoped permission rules like Bash(npm *), from the bootstrap response. Add the optional desktop block alongside cli to set Claude Desktop settings directly. Write settings from Claude Desktop’s managed configuration reference as flat key names. Leave out keys Claude Desktop reads only from MDM or local files, such as bootstrapUrl; the gateway rejects them at boot. Before v2.1.232, the gateway accepted a fixed list of 11 feature-gate keys, such as chatTabEnabled and disableAutoUpdates, and rejected every other key at boot. Before v2.1.227, the gateway also rejected chatTabEnabled and chatAdvancedFileAnalysisEnabled at boot.
Every key is optional; Claude Desktop applies its own default for any key you omit. The gateway validates each desktop block at boot against the configuration schema Claude Desktop itself uses, so a mistake surfaces at gateway start as an error naming the key rather than reaching every connected desktop. The gateway fails at boot when a block contains:
  • An unknown key
  • A recognized key whose value Claude Desktop would reject or silently drop, such as an empty value or a misspelled sub-key inside a nested entry. Before v2.1.260, the gateway silently dropped a misspelled field inside a nested object of a managedMcpServers or orgPluginSettings entry instead of failing at boot.
  • A key the gateway computes itself: the inference connection, the model list, and the OTLP relay. Configure those through upstreams, models, and the telemetry section’s forward_to.
  • A legacy alias of a current key. In the boot error, the gateway names the canonical key to write.
If you use a deprecated value or entry shape, such as a managedMcpServers entry without transport, the gateway starts and logs a warning that names the replacement. The gateway validates a desktop block against the schema bundled with its installed version, as it does the cli block. To deliver a setting introduced by a newer Claude Desktop release, upgrade the gateway first. For example, userPluginMarketplacesEnabled and userPluginUploadsEnabled need Claude Code v2.1.260 or later on the gateway server and Claude Desktop 1.37937.0 or later on members’ machines. If you set orgPluginSettings in a policy’s desktop block, the gateway serves it in the array form that Claude Desktop 1.15200.0 and later reads. Older desktops ignore the array and enforce no plugin tool policy, so update members to 1.15200.0 or later before you rely on it. The gateway fills in keys a policy’s desktop block doesn’t set from the match: {} catch-all’s desktop block, the same way it fills in a policy’s cli block from the base. If you set disabledBuiltinTools or builtinToolPolicy in both the base and a role policy, the gateway keeps the base’s restriction:
  • disabledBuiltinTools: the gateway uses the union of the base’s list and the policy’s list
  • builtinToolPolicy: if you set a tool to a value other than allow in the base, the gateway keeps that value even if you set allow for the same tool in a role policy
For every other key, if you set it in the role policy, the gateway uses the role policy’s value. The gateway replaces an array or a nested object such as banner whole, so if you set banner.text in a role policy, the gateway drops the base’s banner.backgroundColor. If you don’t deploy Claude Desktop, leave desktop out of your policies entirely; the gateway then returns 404 from /user/bootstrap for every user.

Precedence with other managed sources

If a device also has an MDM-delivered policy or a local managed-settings.json, gateway-delivered settings rank first. Precedence within the managed tier on the managed settings page says when the local sources apply, and has the keys Claude Code reads from every admin source regardless of which source it selected, such as the sandbox lock keys, forceRemoteSettingsRefresh, and the per-variable env merge. A policyHelper configured in an MDM profile or the managed settings file runs only when the gateway delivers no settings; the entry says what its output replaces. Embedding hosts such as Claude Desktop can supply policy through the SDK managedSettings option. Parent settings from embedding hosts says when Claude Code applies it, and Restrict parent settings lists which allow-direction settings still apply without the allowManaged*Only locks. Gateway policies apply to every Claude Code invocation on the machine, including non-interactive claude -p runs and sessions spawned by the Agent SDK. If the gateway is unreachable at startup, signed-in sessions exit with an error rather than running without their policy.

telemetry

The CLI sends metrics, logs, and, when enabled, traces to the gateway, which relays them verbatim to each configured destination. The exports use OpenTelemetry Protocol (OTLP) over HTTP. To skip the relay and have sessions export straight to your collector, name the collector in a policy. See Monitoring usage for the metrics and events the CLI emits. The CLI stamps each export with the authenticated user’s identity, read from the gateway-issued JWT: the user.id, user.email, and user.groups attributes. Per-developer cost and usage attribution therefore works with no developer-side configuration. Claude Desktop and Cowork sessions signed in through the gateway stamp their telemetry with user.email and user.groups alongside enduser.id, so you can cover terminal, Desktop, and Cowork usage with one query on user.email or user.groups. user.groups is the comma-separated IdP group list. Desktop and Cowork telemetry also carries enduser.sub, the sub claim your identity provider issues for the user, which stays the same when a user’s email changes. Terminal sessions stamp the same value under user.id, so a query that matches enduser.sub against terminal user.id covers one user’s terminal, Desktop, and Cowork usage together. On Desktop and Cowork exports, user.id is an anonymous identifier, not the subject. Like all OpenTelemetry data from Claude Code, these attributes go only to destinations your organization configures, never to Anthropic. If a user’s group list is longer than 255 characters once percent-encoded, or a group name contains a comma or equals sign, the gateway leaves user.groups off that user’s Desktop and Cowork telemetry rather than truncating it. That user’s terminal sessions still carry the full list. The gateway leaves enduser.sub off when the subject is longer than 255 characters once percent-encoded, or contains a space, a character outside printable ASCII, or one of , ; = \ " %. That user’s Desktop and Cowork telemetry keeps its other attributes. You need Claude Code v2.1.265 or later on the gateway server for user.email and user.groups on Desktop and Cowork telemetry, and Claude Desktop 1.24012 or later on each developer’s machine for user.groups. You need Claude Code v2.1.274 or later on the gateway server for enduser.sub.
Each destination opts into metrics, logs, and traces independently, and the default is metrics only. The signals differ in sensitivity:
  • Metrics: aggregate counters such as token counts, request counts, and latency
  • Logs and traces: can carry full Bash commands, tool inputs, and file paths, covering anything Claude Code does on a developer’s machine
Enable logs and traces only on destinations with the access controls and retention policy that data warrants.
Each forward_to URL must use https://, with one exception for a collector on the gateway’s own loopback interface:
  • http://localhost:<port> passes config validation, but the SSRF guard blocks every export with ECONNREFUSED_SSRF unless you set CLAUDE_GATEWAY_ALLOW_LOOPBACK=1 in the gateway’s environment
  • http://127.0.0.1:<port> or http://[::1]:<port> fails boot unless that variable is set
For an in-cluster collector, expose it over HTTPS at its own internal address, or run it as a sidecar with the variable set. When HTTPS_PROXY is set, the gateway sends exports through that proxy. To reach an internal collector directly, add it to NO_PROXY by hostname or by a domain with a leading dot such as .internal.example.com, which requires Claude Code v2.1.277 or later on the gateway server. Make sure the gateway can reach the collector without the proxy. An entry without a leading dot matches only that exact name, not names under it. CIDR ranges don’t match. With proxy-only egress turned on, allow the collector in the proxy instead, since any NO_PROXY entry keeps proxy-only egress off. Telemetry is off in the CLI by default. When you set both telemetry.forward_to and listen.public_url, the gateway turns it on for connected clients by pushing six environment variables through /managed/settings:
  • CLAUDE_CODE_ENABLE_TELEMETRY=1
  • OTEL_METRICS_EXPORTER, OTEL_LOGS_EXPORTER, and OTEL_TRACES_EXPORTER, each set to otlp if at least one forward_to destination enables that signal and to none otherwise
  • OTEL_EXPORTER_OTLP_ENDPOINT=<public_url>
  • OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf
When you add your own labels, the gateway also pushes OTEL_RESOURCE_ATTRIBUTES. Before Claude Code v2.1.265 on the gateway server, the gateway pushed all three exporter selectors as otlp, including for signals no destination opted into. The pushed endpoint is built from the public URL, so metrics and logs need no OTEL configuration from developers or policies. Developers signed in through /login can’t redirect exports with their own OTEL configuration:
  • Locally set variables: Claude Code applies the pushed variables at the managed tier, so each one overrides the value a developer sets for it locally.
  • Locally configured endpoints: with OTLP/HTTP export enabled, the CLI ignores any locally configured endpoint, whether or not the gateway pushed the telemetry variables. Its exports go to the gateway unless a policy names your collector as the endpoint.
Without a forward_to destination for a signal, the gateway accepts and discards it. If developers already export Claude Code telemetry to one of your collectors, add it as a forward_to destination, with logs or traces enabled if they export those, so it keeps receiving their data after they sign in. To skip the relay instead, name the collector in a policy. Traces also require CLAUDE_CODE_ENHANCED_TELEMETRY_BETA=1 on each client. Set it in a managed policy’s env block, since the gateway doesn’t push it. Developers approve it in the same security approval dialog that the pushed endpoint already triggers. Set it to 1 only in the policies whose groups you want traced. A policy that doesn’t set it inherits the value from your match: {} catch-all policy if that policy sets one, per the merge rules. To keep a group’s clients from sending traces even when a developer sets the variable locally, set it to 0 in that group’s policy. Both protobuf and JSON OTLP encodings are relayed, and any OpenTelemetry-compatible backend works as a destination.

Add your own labels

To put fixed labels such as service.namespace or deployment.environment.name on the telemetry of sessions signed in through the gateway, set telemetry.resource_attributes. Each label is an OpenTelemetry resource attribute, and every destination receives the same labels. Sessions get the labels only when you also set telemetry.forward_to and listen.public_url. This example adds two labels:
The gateway refuses to start when a label breaks one of these rules, and the startup error names the label:
  • Names use only letters, digits, ., _, and -
  • Names aren’t reserved. Compared in any letter case, the reserved names are everything that starts with user., enduser., or identity., plus service.name, service.version, claude.deployment_mode, host.arch, os.type, os.version, and wsl.version
  • Values are non-empty printable ASCII with no space and none of , ; = \ " %
  • Values are at most 255 characters as the gateway counts them after percent-encoding, so /, :, and @ each count as three
  • Values are text, so quote a number, true, or false
You need Claude Code v2.1.281 or later on the gateway server to set telemetry.resource_attributes. An earlier gateway refuses to start when it finds the key. Upgrade every replica before you add the key, and remove the key before you roll back to an earlier version. Terminal sessions signed in through /login receive the labels as OTEL_RESOURCE_ATTRIBUTES, pushed with the other telemetry variables. If you set OTEL_RESOURCE_ATTRIBUTES in a policy’s env block, terminal sessions that policy matches get that value instead of the labels. Claude Desktop receives the labels from the gateway alongside user.email and the other identity attributes. Claude Code also copies each label onto every metric data point, so you can filter metrics by it in a backend that doesn’t index resource attributes. To turn that copy off, see Metrics cardinality control.

Export directly to your collector

To have sessions signed in through /login send telemetry straight to your collector instead of through the relay, set OTEL_EXPORTER_OTLP_ENDPOINT to the collector’s https:// base URL in the env block of a managed policy. Claude Code appends /v1/metrics, /v1/logs, or /v1/traces to the URL you set, such as https://otel-collector.example.com:4318, and exports each signal there over OTLP/HTTP. Requires Claude Code v2.1.265 or later on each developer’s machine. Earlier clients export through the relay. To authenticate to the collector, set OTEL_EXPORTER_OTLP_HEADERS in the same env block. Sessions never send the developer’s gateway session token to a collector named this way. When you add or change this endpoint in a policy, Claude Code asks each developer to approve it in the security approval dialog before applying it in an interactive session. Claude Code checks the endpoint before it exports a signal directly, and keeps that signal on the relay when a check fails. The checks include:
  • The endpoint comes from the gateway itself. If you set the same variable in an MDM profile or a local managed-settings.json, exports stay on the relay.
  • The URL uses https://, or http:// to a loopback address
  • The URL resolves to a path ending in /v1/<signal>, with no query or fragment. Claude Code builds that path itself from the generic variable. It uses a per-signal variable such as OTEL_EXPORTER_OTLP_METRICS_ENDPOINT as written, so include the full path there.
  • The URL isn’t the gateway’s own host. An endpoint addressed to the gateway keeps the relay path and its session token.
  • Neither you nor the developer has configured otelHeadersHelper in any settings source. With a helper configured, every signal stays on the relay.
The endpoint you name changes only where exports go. You still choose which signals export at all with the OTEL_*_EXPORTER selectors. The endpoint alone doesn’t turn export on, so also set the variables that do, unless the gateway already pushes them:
  • If the gateway already pushes the telemetry variables, they cover enablement, selectors, and protocol, and your explicit endpoint overrides the pushed <public_url> value. Set an OTEL_*_EXPORTER selector to otlp yourself only for a signal that no forward_to destination enables.
  • If it doesn’t, also set CLAUDE_CODE_ENABLE_TELEMETRY=1, the OTEL_*_EXPORTER selectors, and OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf.
When the developer signs out, or signs in to a different gateway, exports to the collector stop and Claude Code drops each remaining batch rather than sending it.

When a destination fails

The gateway doesn’t buffer, retry, or store telemetry, so it drops an export that doesn’t reach a destination rather than delivering it late. Each destination succeeds or fails on its own, and the exporting client receives a success response either way, so a failed delivery appears only in the gateway’s log. After five consecutive failed deliveries to a destination, the gateway pauses forwarding to it in 30-second stretches, logging each pause, until a delivery succeeds. Any error response, timeout, or connection error counts as a failed delivery, except 400, 413, 415, 422, and 431, which mean the collector refused that export’s payload as malformed or too large. A refused payload neither advances nor resets the failure count: the gateway keeps forwarding to the destination and logs a warning naming it and the status, on the destination’s first refusal and every hundredth after.

HTTP tuning

Four optional top-level blocks, access_control, limits, timeouts, and rate_limits, tune the HTTP surface. The defaults suit most deployments. If you leave both access_control lists empty, which is the default, the gateway serves any client address, so only your network restricts who can reach it. That matters because a gateway can push managed settings that run commands on developer machines. While allow_cidrs is empty, the gateway warns in two places, without changing how it answers any request:
  • At boot: a warning in the operational log recommends allowing only the private ranges 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16, 100.64.0.0/10, 127.0.0.0/8, ::1/128, and fc00::/7, plus any other internal ranges your developers connect from. If you bind the gateway to a loopback address and set neither trusted_proxies nor public_url, as in local development, the warning doesn’t appear.
  • At runtime: the first time a request arrives from an address outside those private ranges, the gateway logs a warning and emits an access.public_client audit event carrying the client IP. Both fire once per process. Link-local addresses, 169.254.0.0/16 and fe80::/10, don’t count as public. The gateway answers /healthz and /readyz before this check runs, so health probes from public ranges don’t trigger it.
Both signals use the client address as the gateway resolves it. If a load balancer, port-forward, or tunnel relays traffic and isn’t listed in listen.trusted_proxies, the gateway sees the relay’s address, which is usually private, so neither the runtime warning nor a private allow list catches traffic relayed through it. Behind such a front end, set listen.trusted_proxies first so the gateway sees real client addresses, and keep the gateway and everything in front of it unreachable from the public internet regardless.

load_test_mode

The load_test_mode block lets you load test a gateway without calling a model provider. While it’s on, the gateway builds and signs each provider request as usual, discards it instead of sending it, and streams a canned reply back through its normal response path. The reply is filler text that begins with a sentence saying it is canned. Requires Claude Code v2.1.282 or later on the gateway server. An earlier gateway refuses to start when it finds the key. Upgrade every replica before you add the block, and remove the block before you roll back. The example below turns the mode on with the defaults, a reply of roughly 750 tokens of text streamed over about 10 seconds:
A load test in this mode covers the gateway, your Postgres, and everything in front of the gateway. It doesn’t cover the provider’s limits, speed, or network path. No model request is sent to the provider, so a replica’s CPU per request is an estimate and reads lower than production, which also encrypts its traffic to the provider. Confirm a replica count with a small pilot against the real provider. Before v2.1.283, the estimate reads much lower. While the mode is on, a request can carry an x-load-test-user header holding a whole number of up to seven digits. The gateway counts each number as a separate developer, with the email and groups of the developer whose token came with the request. Give the load-test deployment its own empty database, because the gateway refuses to start with the mode on against a database in which any developer has already spent anything.
Never turn this on for a gateway that developers use. Every request gets the canned reply and no model is called. The gateway logs a load_test_mode is on warning at boot and marks each inference audit event with load_test: true while the mode is on.

Complete example

This full reference config exercises every core section; the HTTP tuning blocks keep their defaults. Copy it, delete what you don’t need, and fill in your values. The config in the Quickstart is a minimal version of this.
gateway.yaml

Client-side managed settings

Everything above configures the gateway server. You point developer machines at the gateway separately, on each device, through Claude Code’s managed settings. The gateway can’t push the login keys itself, because they’re what tell the client where the gateway is. For the CLI, set these keys in the per-OS managed-settings.json. The two login keys route each developer’s /login to your gateway:
parentSettingsBehavior: "merge" keeps Claude Desktop’s delivery of the egress allowlist to its embedded Claude Code sessions working; Deliver policy to Claude Desktop sessions explains the mechanism and where the opt-in must sit. Deploy the managed-settings.json file to each device, typically via your MDM platform. The file path differs by platform. See where each mechanism stores the policy. By default, a registry policy on Windows or a managed-preferences plist on macOS replaces the managed-settings.json file rather than merging with it, apart from the exception keys and cross-source checks above. All three keys in this snippet follow the highest-priority-source rule, so fleets that deliver policy through Group Policy or configuration profiles must put all three in that mechanism instead. For Claude Desktop, set the bootstrapUrl key in Claude Desktop’s own managed configuration to <listen.public_url>/user/bootstrap. The sign-in flow and per-group policy then match the CLI’s once a policy opts in server-side with a desktop key; without the opt-in, /user/bootstrap returns 404. See Claude Desktop overlay for the server-side half. Claude Code honors forceLoginGatewayUrl, gatewayInternalNetworks, and the "gateway" value of forceLoginMethod only from a managed source on the machine: managed-settings.json, the macOS plist or Windows HKLM registry, or a policy helper. Setting them in a developer’s own ~/.claude/settings.json or in the gateway payload doesn’t configure the gateway sign-in. Leave forceLoginMethod and forceLoginOrgUUID out of the payload. Claude Code still reads both keys from the payload for its startup credential check, so a developer who keeps an Anthropic-issued credential on the machine gets the startup exit described under Administrator policy requires a Cloud gateway sign-in even after they sign in.