> ## Documentation Index
> Fetch the complete documentation index at: https://bifrost-backport-mcp-oauth2-server.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Databricks

> Databricks Model Serving (Foundation Model APIs) and Unity AI Gateway model services - chat, streaming, embeddings, Responses API, PAT and OAuth M2M authentication

## Overview

Databricks serves foundation models through two surfaces, both reachable under your workspace host and both OpenAI-compatible on the wire. Bifrost's `databricks` provider covers both behind one provider, so you configure a workspace once and address models by name.

| Surface | Base path | Model name looks like | What it covers |
| - | - | - | - |
| **Model Serving** (Foundation Model APIs) | `/serving-endpoints` | `databricks-claude-sonnet-4-5` | Every pay-per-token endpoint and provisioned-throughput endpoint |
| **Unity AI Gateway** (model services) | `/ai-gateway/mlflow/v1` | `system.ai.claude-sonnet-4-5` or `<catalog>.<schema>.<service>` | Ready-to-use `system.ai` models and user-created Unity Catalog model services, governed by Unity Catalog |

Key characteristics:

* **One provider, two surfaces** - the surface is chosen per request, by model name or by an explicit setting
* **OpenAI-compatible** - Bifrost uses its shared OpenAI converters, so tool calling, structured outputs and streaming work unchanged
* **Responses API** - works on every endpoint, so coding agents (Claude Code, Cursor, Codex CLI) work through Bifrost; served natively where the endpoint supports it and emulated through Chat Completions elsewhere
* **Reasoning on Claude** - a `reasoning_effort` or reasoning budget is translated to the `thinking` object Claude endpoints take, and the reasoning they return is surfaced on the standard reasoning fields
* **Remote images** - `http(s)` image URLs are fetched and inlined, since Claude endpoints on Databricks accept only inline image data
* **Two auth methods** - a personal access token, or OAuth machine-to-machine with a service principal (Databricks' production recommendation)
* **Governance tags** - optionally forward Bifrost virtual key / team / customer names to Databricks usage tracking

### Supported Operations

| Operation | Non-Streaming | Streaming | Endpoint |
| - | - | - | - |
| Chat Completions | ✅ | ✅ | `<base>/chat/completions` |
| Responses API | ✅ | ✅ | `/serving-endpoints/responses` where the endpoint supports it; emulated via chat everywhere else |
| Embeddings | ✅ | - | `<base>/embeddings` |
| List Models | ✅ | - | Served from Bifrost's model catalog; no workspace API is called |
| Text Completions | ❌ | ❌ | - |
| Image Generation | ❌ | ❌ | - |
| Speech (TTS) / Transcriptions (STT) | ❌ | ❌ | - |
| Batch / Files | ❌ | ❌ | - |

<Note>
  Custom MLflow models served at `/serving-endpoints/{name}/invocations` take a per-model input schema rather than a canonical chat or embedding request, so they are outside this provider's scope. Use a [custom provider](/providers/custom-providers) with a request path override for those.
</Note>

## Prerequisites

1. A Databricks workspace. Its URL looks like `https://dbc-1234abcd-5678.cloud.databricks.com` (AWS), `https://adb-1234567890.azuredatabricks.net` (Azure), or `https://1234567890.gcp.databricks.com` (GCP).
2. Credentials, either:
   * a **personal access token**, from **Settings > Developer > Access tokens**; or
   * an **OAuth M2M service principal** client ID and secret, from **Settings > Identity and access > Service principals**. Databricks recommends this for production.
3. Access to at least one model:
   * **Model Serving**: pay-per-token endpoints are preconfigured in most workspaces; availability varies by region.
   * **Unity AI Gateway**: every account user can query `system.ai` models with no setup. Querying a user-created model service needs `USE CATALOG`, `USE SCHEMA` and `EXECUTE` on it.

## Setup & Configuration

<Tabs>
  <Tab title="Web UI">
    1. Navigate to **Models** > **Model Providers**. Look for **Databricks** under **Configured Providers**. If it is missing, click **Add New Provider** and select **Databricks**.
    2. Click **Add Key** or edit an existing key.
    3. Set a name for your key.
    4. Enter your **Workspace URL** directly or as an environment variable (for example, `env.DATABRICKS_WORKSPACE_URL`). A scheme and trailing slash are fine.
    5. Choose an **Inference Surface**. Leave it on **Auto** unless you want to pin one surface — see [Choosing a surface](#choosing-a-surface).
    6. Pick an **Authentication Method**:
       * **Personal Access Token** - paste the token or use `env.DATABRICKS_TOKEN`.
       * **OAuth M2M (Service Principal)** - enter the client ID and secret. Leave the token blank.
    7. Optionally enable **Forward Governance Tags** to attribute usage on the Databricks side.
    8. Set **Allowed Models** to **All Models** (default) or a specific allowlist.
    9. Save the provider configuration.
  </Tab>

  <Tab title="config.json (personal access token)">
    ```json theme={null}
    {
      "providers": {
        "databricks": {
          "keys": [
            {
              "name": "databricks-key-1",
              "value": "env.DATABRICKS_TOKEN",
              "models": ["*"],
              "weight": 1.0,
              "databricks_key_config": {
                "workspace_url": "env.DATABRICKS_WORKSPACE_URL",
                "api_format": "auto"
              }
            }
          ]
        }
      }
    }
    ```
  </Tab>

  <Tab title="config.json (OAuth M2M)">
    ```json theme={null}
    {
      "providers": {
        "databricks": {
          "keys": [
            {
              "name": "databricks-service-principal",
              "models": ["*"],
              "weight": 1.0,
              "databricks_key_config": {
                "workspace_url": "env.DATABRICKS_WORKSPACE_URL",
                "api_format": "auto",
                "client_id": "env.DATABRICKS_CLIENT_ID",
                "client_secret": "env.DATABRICKS_CLIENT_SECRET",
                "forward_gateway_tags": true
              }
            }
          ]
        }
      }
    }
    ```

    Leave the key's `value` unset on the OAuth path. Bifrost mints a token from `https://<workspace>/oidc/v1/token` with a `client_credentials` grant, caches it, and refreshes it before expiry — one token per credential set, not one per request.
  </Tab>

  <Tab title="Go SDK">
    ```go theme={null}
    schemas.Key{
        Value:  *schemas.NewSecretVar("env.DATABRICKS_TOKEN"),
        Models: []string{"*"},
        Weight: 1.0,
        DatabricksKeyConfig: &schemas.DatabricksKeyConfig{
            WorkspaceURL: *schemas.NewSecretVar("env.DATABRICKS_WORKSPACE_URL"),
            APIFormat:    schemas.DatabricksAPIFormatAuto,
        },
    }
    ```
  </Tab>
</Tabs>

### Key configuration reference

| Field | Required | Description |
| - | - | - |
| `workspace_url` | ✅ | Databricks workspace URL. A scheme and trailing path are tolerated. |
| `api_format` | | `auto` (default), `model_serving`, or `ai_gateway`. See below. |
| `client_id` | | OAuth M2M service principal client ID. Set with `client_secret`. |
| `client_secret` | | OAuth M2M service principal secret. Set with `client_id`. |
| `forward_gateway_tags` | | Forward Bifrost governance labels as Databricks request tags. Default `false`. |

A key needs either a `value` (personal access token) or both `client_id` and `client_secret`. Setting only one half of the service principal pair is rejected at configuration time.

## Choosing a surface

With `api_format: "auto"` (the default) the model name decides:

* A catalog-qualified name — `system.ai.claude-sonnet-4-5`, or a `<catalog>.<schema>.<service>` Unity Catalog model service — routes to the **Unity AI Gateway**.
* A `databricks-*` name — `databricks-claude-sonnet-4-5` — is a pay-per-token endpoint and routes to **Model Serving**.
* A model name that came from a key alias is sent exactly as configured; when it is bare it routes to **Model Serving**.
* Any other bare name — `gpt-5.5`, `claude-opus-5` — is treated as a short name for a ready-to-use `system.ai` model and routes to the **Unity AI Gateway** with the `system.ai.` prefix added (see below).

Set `api_format` explicitly to pin one surface. For production, an explicit setting is safer than relying on the naming convention. Provisioned-throughput endpoints with a custom name need `api_format: "model_serving"` or a key alias, since nothing in their name separates them from a short gateway name.

Model aliases need no extra configuration: alias resolution rewrites the model to its upstream `model_id` before the provider sees it, so the surface is chosen from the upstream name.

### Short names on the Unity AI Gateway

The AI Gateway addresses models by their full Unity Catalog name, but you do not have to spell out the `system.ai` catalog yourself. When a request targets the AI Gateway and the model name is not already catalog-qualified, Bifrost prefixes it with `system.ai.` before sending it upstream:

| You send | Databricks receives |
| - | - |
| `databricks/gpt-5.5` (under `auto` or `ai_gateway`) | `system.ai.gpt-5.5` |
| `databricks/claude-opus-5` (under `auto` or `ai_gateway`) | `system.ai.claude-opus-5` |
| `databricks/system.ai.gpt-5.5` | `system.ai.gpt-5.5` |
| `databricks/main.default.my-service` | `main.default.my-service` |
| `databricks/databricks-claude-sonnet-4-5` under `auto` | `databricks-claude-sonnet-4-5` (Model Serving, no prefix) |

A name counts as catalog-qualified when it has at least two dots, so a version dot such as the one in `gpt-5.5` does not stop the prefix from being added. Names that come from a key alias are always sent exactly as the alias's `model_id`, so an alias is the way to point a short name at a model service in a different catalog.

## Usage

<Tabs>
  <Tab title="Model Serving">
    ```bash theme={null}
    curl -X POST http://localhost:8080/v1/chat/completions \
      -H "Content-Type: application/json" \
      -d '{
        "model": "databricks/databricks-claude-sonnet-4-5",
        "messages": [{"role": "user", "content": "Hello!"}]
      }'
    ```
  </Tab>

  <Tab title="Unity AI Gateway">
    ```bash theme={null}
    curl -X POST http://localhost:8080/v1/chat/completions \
      -H "Content-Type: application/json" \
      -d '{
        "model": "databricks/system.ai.claude-sonnet-4-5",
        "messages": [{"role": "user", "content": "Hello!"}]
      }'
    ```
  </Tab>

  <Tab title="Embeddings">
    ```bash theme={null}
    curl -X POST http://localhost:8080/v1/embeddings \
      -H "Content-Type: application/json" \
      -d '{
        "model": "databricks/databricks-gte-large-en",
        "input": "What is Databricks?"
      }'
    ```
  </Tab>

  <Tab title="Responses API">
    ```bash theme={null}
    curl -X POST http://localhost:8080/v1/responses \
      -H "Content-Type: application/json" \
      -d '{
        "model": "databricks/databricks-gpt-5",
        "input": "What is a mixture of experts model?"
      }'
    ```

    Bifrost first tries the native `/serving-endpoints/responses` route. Pay-per-token foundation model endpoints commonly decline it (`Responses API passthrough is not supported for model ...`), in which case the request is replayed through Chat Completions and converted back, and later requests for that endpoint skip straight to the emulated path. The Unity AI Gateway is chat-only and is always emulated.
  </Tab>
</Tabs>

### Reasoning on Claude endpoints

Claude endpoints on Databricks reject `reasoning_effort` and take Anthropic's `thinking` object instead. Bifrost translates for you: a reasoning budget (`reasoning.max_tokens`) is sent as `budget_tokens`, an effort label is scaled to a budget the same way the Anthropic provider does, and models that only offer adaptive thinking get `{"type": "adaptive"}`. `max_completion_tokens` is raised above the budget when you have not set one, because the endpoint requires the ceiling to exceed the budget.

```bash theme={null}
curl -X POST http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "databricks/databricks-claude-sonnet-4-5",
    "messages": [{"role": "user", "content": "What is 17 * 23?"}],
    "reasoning": {"effort": "high", "max_tokens": 2048}
  }'
```

Every Databricks endpoint, Claude or not, returns reasoning as a content block (`{"type": "reasoning", "summary": [...]}`). Bifrost lifts it onto `reasoning` and `reasoning_details` (with the signature) on the message and on each streaming delta, so clients read it the same way as from any other provider.

To speak the Databricks dialect directly, send `thinking` as a top-level field; Bifrost passes it through untouched and does not add its own.

### Image inputs

Claude endpoints on Databricks accept only inline image data (`Http data URLs are not supported by Claude`). Bifrost fetches any `http(s)` `image_url` and sends it as a base64 data URL. The fetch uses the same SSRF-safe downloader as the Bedrock and Anthropic providers: only `http`/`https`, no private or loopback addresses, 25 MiB cap. A URL that cannot be fetched fails the request with a 400 rather than sending the model a request with the image silently removed.

### Provider-specific parameters

Databricks accepts fields Bifrost does not model canonically, such as `service_tier` for priority pay-per-token inference. Send them as top-level fields and Bifrost forwards them:

```bash theme={null}
curl -X POST http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "databricks/databricks-claude-sonnet-4-5",
    "messages": [{"role": "user", "content": "Hello!"}],
    "service_tier": "priority"
  }'
```

### Parameter support per model

Bifrost's parameter set is wider than any single Databricks endpoint accepts, and both surfaces answer an unknown field with a 400 rather than ignoring it. Each optional parameter is therefore checked against the Bifrost datasheet record for the model behind the endpoint before the request goes out:

| Parameter | Dropped when the datasheet says |
| - | - |
| `reasoning_effort` | the model does not take an effort label |
| `temperature`, `top_p`, `top_k` | `supports_sampling_params: false` — the adaptive-thinking Claude endpoints reject them |
| `tool_choice` | `supports_tool_choice: false` |
| `parallel_tool_calls` | `supports_parallel_function_calling: false` |
| `response_format` (`text.format` on Responses) | `supports_response_schema: false` |
| `stop`, `presence_penalty`, `frequency_penalty` | the field is listed in `unsupported_fields` |

A model the datasheet does not describe keeps every field except `reasoning_effort` and `parallel_tool_calls` — Databricks endpoint names are workspace-defined, so there is nothing to identify an unknown endpoint by, and a silent drop is worse than an upstream error you can read. `parallel_tool_calls` is the exception because both surfaces reject it outright; it is only sent when the datasheet row opts in.

`reasoning_effort` resolves through `unsupported_fields` → `supports_reasoning_effort` (or a published effort ladder) → a `reasoning_effort` entry in the row's parameters → the model reasons and is not Claude-family. Claude endpoints on Databricks reason through a thinking budget and reject an effort label, so it is translated to `thinking` there (see [Reasoning on Claude endpoints](#reasoning-on-claude-endpoints)). A requested effort is clamped onto the levels the model publishes.

Anthropic-native fields — `context_management`, `cache_control`, `speed`, `inference_geo`, `task_budget`, `container`, `mcp_servers` — are always dropped, whatever the model. Both Databricks surfaces are OpenAI-shaped and reject them even on Claude endpoints.

### Provisioned throughput

A provisioned-throughput endpoint uses the same request format as a pay-per-token one. Address it by its endpoint name on the Model Serving surface — no separate configuration.

## Usage attribution

With `forward_gateway_tags` enabled, Bifrost sends the resolved governance labels to Databricks on each request:

```http theme={null}
Databricks-Ai-Gateway-Request-Tags: {"customer":"acme","team":"platform","virtual_key":"vk-prod"}
```

Databricks records these against the request for its own usage tracking, so Databricks-side cost reporting can be sliced the same way Bifrost's is. Only display names are sent, never user identifiers. A tag header you supply yourself (via an `x-bf-eh-databricks-ai-gateway-request-tags` request header) takes precedence.

Bifrost's own telemetry is unaffected and remains the cross-provider view.

## Cost tracking

Databricks pricing depends on the model, the serving configuration (pay-per-token vs provisioned throughput), the service tier, and cached-token behaviour. Bifrost prices Databricks requests from the model catalog rather than a flat provider rate. Databricks returns the standard OpenAI-shaped `usage` object, including `reasoning_tokens` and `cache_read_input_tokens` / `cache_creation_input_tokens` where the model supports them.

## Troubleshooting

| Symptom | Cause |
| - | - |
| `databricks workspace url is not set` | No `workspace_url` on the key. `config.json` and the UI require it; only the Go SDK also accepts a provider-level `base_url` as the workspace host. |
| `databricks key has no credentials` | Neither a key `value` nor a complete service principal pair. |
| 403 with `PERMISSION_DENIED` | The principal lacks `EXECUTE` on the model service, or query access to the serving endpoint. Bifrost surfaces this as an authorization error rather than a missing model. |
| 404 on a model that exists | The request went to the wrong surface. Set `api_format` explicitly. |
| A model is missing from the model list | Databricks models are listed from Bifrost's model catalog, not from the workspace. Any serving endpoint or model service can still be called by name; `api_format` only affects routing, not listing. |

## Legacy AI Gateway endpoints

Earlier Databricks workspaces expose a Beta-generation AI Gateway on a separate host, `https://<workspace-id>.ai-gateway.cloud.databricks.com`, with the MLflow surface at `/mlflow/v1` rather than `/ai-gateway/mlflow/v1`.

That host is not addressable by this provider's base paths. To use it, configure a [custom provider](/providers/custom-providers) with `base_provider_type: openai` and a Base URL of `https://<workspace-id>.ai-gateway.cloud.databricks.com/mlflow`, keyless, with an `Authorization: Bearer <PAT>` extra header. For the native Anthropic Messages surface on that host, use a second custom provider with `base_provider_type: anthropic` and a Base URL of `.../anthropic`, with List Models disabled.

## Reference Links

* [Query model APIs (model services)](https://docs.databricks.com/aws/en/ai-gateway/query-model-services)
* [Databricks Foundation Model APIs](https://docs.databricks.com/aws/en/machine-learning/foundation-model-apis/)
* [Foundation model REST API reference](https://docs.databricks.com/aws/en/machine-learning/foundation-model-apis/api-reference)
* [Supported foundation models on Model Serving](https://docs.databricks.com/aws/en/machine-learning/model-serving/foundation-model-overview)
* [OAuth machine-to-machine authentication](https://docs.databricks.com/aws/en/dev-tools/auth/oauth-m2m)
* [Track model usage](https://docs.databricks.com/aws/en/ai-gateway/usage-tracking)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.