Overview
Databricks serves foundation models through two surfaces, both reachable under your workspace host and both OpenAI-compatible on the wire. Bifrost’sdatabricks provider covers both behind one provider, so you configure a workspace once and address models by name.
Key characteristics:
- One provider, two surfaces - the surface is chosen per request, by model name or by an explicit setting
- OpenAI-compatible - Bifrost uses its shared OpenAI converters, so tool calling, structured outputs and streaming work unchanged
- Responses API - works on every endpoint, so coding agents (Claude Code, Cursor, Codex CLI) work through Bifrost; served natively where the endpoint supports it and emulated through Chat Completions elsewhere
- Reasoning on Claude - a
reasoning_effortor reasoning budget is translated to thethinkingobject Claude endpoints take, and the reasoning they return is surfaced on the standard reasoning fields - Remote images -
http(s)image URLs are fetched and inlined, since Claude endpoints on Databricks accept only inline image data - Two auth methods - a personal access token, or OAuth machine-to-machine with a service principal (Databricks’ production recommendation)
- Governance tags - optionally forward Bifrost virtual key / team / customer names to Databricks usage tracking
Supported Operations
Custom MLflow models served at
/serving-endpoints/{name}/invocations take a per-model input schema rather than a canonical chat or embedding request, so they are outside this provider’s scope. Use a custom provider with a request path override for those.Prerequisites
- A Databricks workspace. Its URL looks like
https://dbc-1234abcd-5678.cloud.databricks.com(AWS),https://adb-1234567890.azuredatabricks.net(Azure), orhttps://1234567890.gcp.databricks.com(GCP). - Credentials, either:
- a personal access token, from Settings > Developer > Access tokens; or
- an OAuth M2M service principal client ID and secret, from Settings > Identity and access > Service principals. Databricks recommends this for production.
- Access to at least one model:
- Model Serving: pay-per-token endpoints are preconfigured in most workspaces; availability varies by region.
- Unity AI Gateway: every account user can query
system.aimodels with no setup. Querying a user-created model service needsUSE CATALOG,USE SCHEMAandEXECUTEon it.
Setup & Configuration
- Web UI
- config.json (personal access token)
- config.json (OAuth M2M)
- Go SDK
- Navigate to Models > Model Providers. Look for Databricks under Configured Providers. If it is missing, click Add New Provider and select Databricks.
- Click Add Key or edit an existing key.
- Set a name for your key.
- Enter your Workspace URL directly or as an environment variable (for example,
env.DATABRICKS_WORKSPACE_URL). A scheme and trailing slash are fine. - Choose an Inference Surface. Leave it on Auto unless you want to pin one surface — see Choosing a surface.
- Pick an Authentication Method:
- Personal Access Token - paste the token or use
env.DATABRICKS_TOKEN. - OAuth M2M (Service Principal) - enter the client ID and secret. Leave the token blank.
- Personal Access Token - paste the token or use
- Optionally enable Forward Governance Tags to attribute usage on the Databricks side.
- Set Allowed Models to All Models (default) or a specific allowlist.
- Save the provider configuration.
Key configuration reference
A key needs either a
value (personal access token) or both client_id and client_secret. Setting only one half of the service principal pair is rejected at configuration time.
Choosing a surface
Withapi_format: "auto" (the default) the model name decides:
- A catalog-qualified name —
system.ai.claude-sonnet-4-5, or a<catalog>.<schema>.<service>Unity Catalog model service — routes to the Unity AI Gateway. - A
databricks-*name —databricks-claude-sonnet-4-5— is a pay-per-token endpoint and routes to Model Serving. - A model name that came from a key alias is sent exactly as configured; when it is bare it routes to Model Serving.
- Any other bare name —
gpt-5.5,claude-opus-5— is treated as a short name for a ready-to-usesystem.aimodel and routes to the Unity AI Gateway with thesystem.ai.prefix added (see below).
api_format explicitly to pin one surface. For production, an explicit setting is safer than relying on the naming convention. Provisioned-throughput endpoints with a custom name need api_format: "model_serving" or a key alias, since nothing in their name separates them from a short gateway name.
Model aliases need no extra configuration: alias resolution rewrites the model to its upstream model_id before the provider sees it, so the surface is chosen from the upstream name.
Short names on the Unity AI Gateway
The AI Gateway addresses models by their full Unity Catalog name, but you do not have to spell out thesystem.ai catalog yourself. When a request targets the AI Gateway and the model name is not already catalog-qualified, Bifrost prefixes it with system.ai. before sending it upstream:
A name counts as catalog-qualified when it has at least two dots, so a version dot such as the one in
gpt-5.5 does not stop the prefix from being added. Names that come from a key alias are always sent exactly as the alias’s model_id, so an alias is the way to point a short name at a model service in a different catalog.
Usage
- Model Serving
- Unity AI Gateway
- Embeddings
- Responses API
Reasoning on Claude endpoints
Claude endpoints on Databricks rejectreasoning_effort and take Anthropic’s thinking object instead. Bifrost translates for you: a reasoning budget (reasoning.max_tokens) is sent as budget_tokens, an effort label is scaled to a budget the same way the Anthropic provider does, and models that only offer adaptive thinking get {"type": "adaptive"}. max_completion_tokens is raised above the budget when you have not set one, because the endpoint requires the ceiling to exceed the budget.
{"type": "reasoning", "summary": [...]}). Bifrost lifts it onto reasoning and reasoning_details (with the signature) on the message and on each streaming delta, so clients read it the same way as from any other provider.
To speak the Databricks dialect directly, send thinking as a top-level field; Bifrost passes it through untouched and does not add its own.
Image inputs
Claude endpoints on Databricks accept only inline image data (Http data URLs are not supported by Claude). Bifrost fetches any http(s) image_url and sends it as a base64 data URL. The fetch uses the same SSRF-safe downloader as the Bedrock and Anthropic providers: only http/https, no private or loopback addresses, 25 MiB cap. A URL that cannot be fetched fails the request with a 400 rather than sending the model a request with the image silently removed.
Provider-specific parameters
Databricks accepts fields Bifrost does not model canonically, such asservice_tier for priority pay-per-token inference. Send them as top-level fields and Bifrost forwards them:
Parameter support per model
Bifrost’s parameter set is wider than any single Databricks endpoint accepts, and both surfaces answer an unknown field with a 400 rather than ignoring it. Each optional parameter is therefore checked against the Bifrost datasheet record for the model behind the endpoint before the request goes out:
A model the datasheet does not describe keeps every field except
reasoning_effort and parallel_tool_calls — Databricks endpoint names are workspace-defined, so there is nothing to identify an unknown endpoint by, and a silent drop is worse than an upstream error you can read. parallel_tool_calls is the exception because both surfaces reject it outright; it is only sent when the datasheet row opts in.
reasoning_effort resolves through unsupported_fields → supports_reasoning_effort (or a published effort ladder) → a reasoning_effort entry in the row’s parameters → the model reasons and is not Claude-family. Claude endpoints on Databricks reason through a thinking budget and reject an effort label, so it is translated to thinking there (see Reasoning on Claude endpoints). A requested effort is clamped onto the levels the model publishes.
Anthropic-native fields — context_management, cache_control, speed, inference_geo, task_budget, container, mcp_servers — are always dropped, whatever the model. Both Databricks surfaces are OpenAI-shaped and reject them even on Claude endpoints.
Provisioned throughput
A provisioned-throughput endpoint uses the same request format as a pay-per-token one. Address it by its endpoint name on the Model Serving surface — no separate configuration.Usage attribution
Withforward_gateway_tags enabled, Bifrost sends the resolved governance labels to Databricks on each request:
x-bf-eh-databricks-ai-gateway-request-tags request header) takes precedence.
Bifrost’s own telemetry is unaffected and remains the cross-provider view.
Cost tracking
Databricks pricing depends on the model, the serving configuration (pay-per-token vs provisioned throughput), the service tier, and cached-token behaviour. Bifrost prices Databricks requests from the model catalog rather than a flat provider rate. Databricks returns the standard OpenAI-shapedusage object, including reasoning_tokens and cache_read_input_tokens / cache_creation_input_tokens where the model supports them.
Troubleshooting
Legacy AI Gateway endpoints
Earlier Databricks workspaces expose a Beta-generation AI Gateway on a separate host,https://<workspace-id>.ai-gateway.cloud.databricks.com, with the MLflow surface at /mlflow/v1 rather than /ai-gateway/mlflow/v1.
That host is not addressable by this provider’s base paths. To use it, configure a custom provider with base_provider_type: openai and a Base URL of https://<workspace-id>.ai-gateway.cloud.databricks.com/mlflow, keyless, with an Authorization: Bearer <PAT> extra header. For the native Anthropic Messages surface on that host, use a second custom provider with base_provider_type: anthropic and a Base URL of .../anthropic, with List Models disabled.

