> ## Documentation Index
> Fetch the complete documentation index at: https://bifrost-backport-mcp-oauth2-server.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Prometheus

> Monitor Bifrost metrics with Prometheus scraping or Push Gateway for multi-node deployments

## Overview

Bifrost exposes Prometheus metrics via two methods:

1. **Pull-based (Scraping)**: Traditional `/metrics` endpoint that Prometheus can scrape
2. **Push-based (Push Gateway)**: Push metrics to a Prometheus Push Gateway for cluster deployments

<Note>
  **For multi-node deployments**: Use the Push Gateway method to ensure accurate metric aggregation. Traditional scraping may miss nodes behind load balancers.
</Note>

***

## Pull-based Scraping

Bifrost automatically exposes a `/metrics` endpoint when the telemetry plugin is enabled (enabled by default). No additional configuration is needed.

<Info>
  When Bifrost's authentication is enabled (`auth_config.is_enabled = true`), the `/metrics` endpoint requires credentials. You can authenticate the scraper with either the admin **Basic auth** credentials (`admin_username` / `admin_password` from your `auth_config`) or, on Enterprise, a **Bifrost API key** with the Metrics permission. Without valid credentials, Prometheus receives `401 Unauthorized` responses and scraping silently fails.
</Info>

### Prometheus Configuration

Add Bifrost to your Prometheus `prometheus.yml`:

```yaml theme={null}
scrape_configs:
  - job_name: 'bifrost'
    static_configs:
      - targets: ['bifrost-host:8080']
    scrape_interval: 15s
```

If Bifrost authentication is enabled, add `basic_auth` to your scrape config:

```yaml theme={null}
scrape_configs:
  - job_name: 'bifrost'
    static_configs:
      - targets: ['bifrost-host:8080']
    scrape_interval: 15s
    basic_auth:
      username: '<admin_username>'
      password: '<admin_password>'
```

<Warning>
  Prometheus scrapes over plain `http` by default, which sends the Basic auth credentials or API key in cleartext. When Bifrost is served over TLS, set `scheme: https` (and any required `tls_config`) in the scrape config so credentials are not exposed in transit.
</Warning>

#### Authenticating with an API Key <sup>Enterprise</sup>

On Enterprise deployments, you can scrape `/metrics` with a Bifrost API key instead of the admin Basic auth credentials. Create an API key with the **Metrics** permission (included in all default roles) from **Settings → API Keys** (see [Creating API Keys](/api/procuring-api-keys)), then pass it as a bearer token in your scrape config:

```yaml theme={null}
scrape_configs:
  - job_name: 'bifrost'
    static_configs:
      - targets: ['bifrost-host:8080']
    scrape_interval: 15s
    authorization:
      type: Bearer
      credentials: '<bfst-your-api-key>'
```

Older Prometheus versions that lack the `authorization` block can use `bearer_token` instead:

```yaml theme={null}
    bearer_token: '<bfst-your-api-key>'
```

<Note>
  API-key auth for `/metrics` is an Enterprise feature. On the open-source build, the `/metrics` endpoint accepts only Basic auth (the `admin_username` / `admin_password` above).
</Note>

### Endpoint

```
GET /metrics
```

Returns metrics in Prometheus exposition format.

***

## Push-based (Push Gateway)

For multi-node cluster deployments, the Prometheus plugin pushes metrics to a [Prometheus Push Gateway](https://github.com/prometheus/pushgateway). This ensures all nodes' metrics are captured regardless of load balancer routing.

### Configuration

| Field | Type | Required | Default | Description |
| - | - | - | - | - |
| `push_gateway_url` | `string \| EnvVar` | ✅ Yes | - | Push Gateway URL — supports `env.VAR_NAME` |
| `job_name` | `string` | ❌ No | `bifrost` | Job label for pushed metrics |
| `instance_id` | `string` | ❌ No | hostname | Instance identifier for metric grouping |
| `push_interval` | `integer` | ❌ No | `15` | Push interval in seconds (1-300) |
| `basic_auth` | `object` | ❌ No | - | Basic auth credentials |

### Basic Auth Configuration

| Field | Type | Required | Description |
| - | - | - | - |
| `username` | `string \| EnvVar` | ✅ Yes | Basic auth username — supports `env.VAR_NAME` |
| `password` | `string \| EnvVar` | ✅ Yes | Basic auth password — supports `env.VAR_NAME` |

***

## Setup

<Tabs group="setup-method">
  <Tab title="UI">
    1. Navigate to **Observability** → **Prometheus** in the Bifrost UI
    2. The `/metrics` endpoint is shown at the top for scraping configuration
    3. To enable Push Gateway:
       * Enter the **Push Gateway URL**
       * Configure **Job Name** and **Push Interval** as needed
       * Optionally set a custom **Instance ID**
       * Enable **Basic Authentication** if required
       * Toggle **Enable Push Gateway** on
       * Click **Save Prometheus Configuration**
  </Tab>

  <Tab title="Config File">
    ```json theme={null}
    {
      "plugins": [
        {
          "name": "telemetry",
          "enabled": true,
          "config": {
            "push_gateway": {
              "enabled": true,
              "push_gateway_url": "http://pushgateway:9091",
              "job_name": "bifrost",
              "push_interval": 15
            }
          }
        }
      ]
    }
    ```

    ### With Basic Auth

    ```json theme={null}
    {
      "plugins": [
        {
          "name": "telemetry",
          "enabled": true,
          "config": {
            "push_gateway": {
              "enabled": true,
              "push_gateway_url": "http://pushgateway:9091",
              "job_name": "bifrost",
              "push_interval": 15,
              "instance_id": "bifrost-node-1",
              "basic_auth": {
                "username": "admin",
                "password": "secret"
              }
            }
          }
        }
      ]
    }
    ```

    ### With Environment Variables

    Use `env.VAR_NAME` to reference environment variables for the Push Gateway URL and credentials:

    ```json theme={null}
    {
      "plugins": [
        {
          "name": "telemetry",
          "enabled": true,
          "config": {
            "push_gateway": {
              "enabled": true,
              "push_gateway_url": "env.PUSHGATEWAY_URL",
              "job_name": "bifrost",
              "push_interval": 15,
              "basic_auth": {
                "username": "env.PUSHGATEWAY_USER",
                "password": "env.PUSHGATEWAY_PASS"
              }
            }
          }
        }
      ]
    }
    ```
  </Tab>
</Tabs>

***

## Available Metrics

The following metrics are available from both the `/metrics` endpoint and Push Gateway:

### HTTP Metrics

| Metric | Type | Description |
| - | - | - |
| `http_requests_total` | Counter | Total HTTP requests by path, method, status |
| `http_request_duration_seconds` | Histogram | HTTP request latency |
| `http_request_size_bytes` | Histogram | Request body size |
| `http_response_size_bytes` | Histogram | Response body size |

<Note>
  The `path` label contains the matched route template (e.g. `/genai/v1beta/models/{model:*}`, `/v1/messages/batches/{batch_id}`), not the raw URL path. This keeps metric cardinality bounded by the number of registered routes instead of growing with every model name or resource ID that appears in a URL. For per-model breakdowns, use the `model` and `provider` labels on the `bifrost_*` metrics.
</Note>

### Bifrost LLM Metrics

| Metric | Type | Description |
| - | - | - |
| `bifrost_upstream_requests_total` | Counter | Total requests to LLM providers |
| `bifrost_upstream_latency_seconds` | Histogram | Provider request latency |
| `bifrost_overhead_latency_microseconds` | Histogram | Total Bifrost overhead per request in microseconds (Bifrost's own work, excluding upstream provider time) |
| `bifrost_overhead_component_microseconds` | Histogram | That same overhead broken down by internal component via the `overhead_component` label. Opt-in, and populated only when tracing is active — see [Overhead Breakdown](#overhead-breakdown) |
| `bifrost_success_requests_total` | Counter | Successful provider requests |
| `bifrost_error_requests_total` | Counter | Failed requests, by raw `status_code` and normalized `error_type` — see [Error Types](#error-types) |
| `bifrost_input_tokens_total` | Counter | Total input tokens processed |
| `bifrost_output_tokens_total` | Counter | Total output tokens generated |
| `bifrost_cost_total` | Counter | Total cost in USD |
| `bifrost_cache_hits_total` | Counter | Cache hits by type |
| `bifrost_stream_first_token_latency_seconds` | Histogram | Time to first token (streaming) |
| `bifrost_stream_inter_token_latency_seconds` | Histogram | Inter-token latency (streaming) |
| `bifrost_active_requests` | Gauge | LLM requests currently in-flight (labeled by `method` only) |
| `bifrost_provider_key_up` | Gauge | Per-key health. `1` after a successful attempt, `0` after a failed attempt. Labels: `provider`, `key_id`, `key_name`. |
| `bifrost_key_rotation_events_total` | Counter | Key rotations triggered by per-key failures — rate-limit (429), auth (401/403), or billing (402) — see below <sup>v1.5.0-prerelease4+</sup> |
| `bifrost_request_retries` | Histogram | Number of retries used per request (observed once per request; buckets `0,1,2,3,5,10`). |
| `bifrost_routing_embedding_requests_total` | Counter | Embedding calls made by semantic complexity routing. Labels: `provider`, `model` (the embedding provider/model, not the request's), `phase` (`request` classification vs `warmup` exemplar embedding). |
| `bifrost_routing_embedding_cost_total` | Counter | Cost in USD of semantic routing embeddings (same labels as above). Recorded regardless of whether embedding usage counts toward budgets. |

### Error Types

`bifrost_error_requests_total` carries two error dimensions. `status_code` is the raw fact. `error_type` is the normalized interpretation: a closed, low-cardinality vocabulary that answers **whose fault was it**, which is what alarms actually need.

Values are prefixed by fault domain, so a success-rate alarm that should ignore caller mistakes is one clause rather than a list of reasons that grows over time:

```promql theme={null}
# Failure rate, excluding faults the caller caused
sum(rate(bifrost_error_requests_total{error_type!~"caller_.*"}[5m]))
  / sum(rate(bifrost_upstream_requests_total[5m]))

# What are callers getting wrong, and which teams
sum by (error_type, team_name, model) (rate(bifrost_error_requests_total{error_type=~"caller_.*"}[5m]))

# Our policy refusals vs the upstream's
sum by (error_type) (rate(bifrost_error_requests_total{error_type=~"policy_.*|provider_.*"}[5m]))
```

#### Vocabulary

The set is closed, so this list is exhaustive. **Declared** values are stated by the code that produced the failure — it knew the reason, so there is no guessing. **Inferred** values are derived from the upstream's status code, because a provider's own error vocabulary cannot be trusted to mean the same thing twice. Three values are both. `caller_cancelled` and `provider_timeout` are named by the code that recognises them and are also implied by status 499 and 504. `bifrost_internal` is named on the paths that recognise themselves and is otherwise the fallback for a Bifrost-origin failure that named nothing.

| `error_type` | Meaning | Precision |
| - | - | - |
| `caller_model_not_available` | No configured key serves the requested model | Declared |
| `caller_model_unknown` | The provider rejected the model name | Inferred |
| `caller_invalid_request` | Malformed or unsatisfiable request (4xx) | Inferred |
| `caller_cancelled` | The caller hung up before the response completed | Declared or Inferred (status 499) |
| `policy_model_blocked` | The grant does not permit this model on this provider | Declared |
| `policy_provider_blocked` | The grant does not permit this provider | Declared |
| `policy_access_denied` | The credential did not resolve, or the grant refuses the caller | Declared |
| `policy_budget_exceeded` | A configured budget is spent | Declared |
| `policy_rate_limited` | A configured governance limit refused the request | Declared |
| `policy_tool_blocked` | The grant does not permit this MCP tool | Declared |
| `provider_auth_failed` | The upstream rejected the credential (401/403) | Inferred |
| `provider_billing` | The upstream account cannot pay (402) | Inferred |
| `provider_rate_limited` | The upstream rate-limited us (429) | Inferred |
| `provider_overloaded` | The upstream is at capacity (503, Anthropic's 529) | Inferred |
| `provider_server_error` | The upstream failed (other 5xx) | Inferred |
| `provider_timeout` | The upstream did not answer in time | Declared or Inferred (status 504) |
| `provider_connection_failed` | The connection never established (DNS, refused, TLS) | Declared |
| `provider_credentials_exhausted` | Every key in the pool returned a permanent per-key error | Declared |
| `bifrost_dropped` | Request shed because the provider queue was full | Declared |
| `bifrost_internal` | Bifrost failed for a reason of its own | Declared or Inferred (fallback) |
| `_OTHER` | No rule matched | — |

<Note>
  **Alarm on `_OTHER` directly.** It means no producer declared a reason and no rule inferred one. The name is the OTel catch-all, matching the `error_type` label on the MCP metrics. It should be flat at zero; a rise means failures are going unclassified — most likely a new refusal reason was added somewhere without declaring its type — not that a new kind of failure is benign.
</Note>

#### What the distinction buys you

The same status code means different things depending on who produced it, and `error_type` is what separates them:

* A **429** is `policy_rate_limited` when a governance limit refused the request and `provider_rate_limited` when the upstream did. The first means your own limits are too tight; the second means you need more upstream capacity.
* A **403** is `policy_model_blocked` when a grant refuses the model and `provider_auth_failed` when the upstream rejects the key.
* A **503** is `bifrost_dropped` when Bifrost shed the request under queue pressure and `provider_overloaded` when the upstream is at capacity.

Classification always prefers what Bifrost knows over what a status code suggests, so a decision Bifrost made itself is never re-attributed to the provider.

The same value is stamped on the provider-attempt span as `bifrost.error.type`, so trace-based exporters (OpenTelemetry, Datadog) classify identically instead of re-deriving from the raw provider `error.type`. Pre-dispatch refusals produce no attempt span, so they reach the metrics but not the span.

#### Precision caveats on the inferred values

`caller_model_unknown` requires HTTP **404** on a request whose only addressable resource is the model: text completion, embedding, speech, transcription, image generation, rerank, count-tokens, and their streaming variants.

Everything else classifies as `caller_invalid_request`, because its 404 may be about something other than the model:

* file, batch, video and container operations address that resource directly;
* OCR, image edit/variation and video generation reference a remote source asset (`document_url`, `image.url`, `input_reference`, `video_uri`);
* chat completions and responses carry file ids, and responses additionally carry `previous_response_id`, which 404s on its own when stale.

Chat and responses are the notable absence. A wrong model name there lands in `caller_invalid_request` rather than being distinguished, because a 404 on those requests is genuinely ambiguous. Distinguishing it needs the provider to say so explicitly rather than Bifrost inferring it from a status code — see the per-provider normalization note above. The exact `caller_model_not_available` is unaffected: it is declared by Bifrost, not inferred, and covers every request type.

The inferred values deliberately do not key off the provider's own `error.type` / `error.code`. Those are passed through verbatim and disagree across providers for the same condition — OpenAI sends `type="invalid_request_error"` with `code="model_not_found"`, Anthropic sends `type="not_found_error"` and has no `code` field at all, and Bedrock and Databricks each use their own vocabulary. OpenAI even sends `invalid_request_error` for a 401. Both fields remain available per-request on the span and in the logs.

Two consequences, both accepted rather than papered over:

* A provider that rejects an unknown model with a **400** lands in `caller_invalid_request`, not `caller_model_unknown`. This under-counts rather than mislabelling generic invalid-request traffic as a bad model name.
* Some providers (Vertex) use 404 for *"not found **or** your project does not have access to it"*, so a permissions problem can land in `caller_model_unknown`. Separating the two needs per-provider error normalization; the **Declared** rows above do not depend on it.

### Overhead Breakdown

`bifrost_overhead_component_microseconds` decomposes the same overhead measured by `bifrost_overhead_latency_microseconds` into per-component histograms. It carries the same base labels as `bifrost_overhead_latency_microseconds` (see [Default Labels](#default-labels)) plus an `overhead_component` label naming the internal component, and shares the same bucket boundaries. Summing every component for a given label set reconstructs the scalar total.

`overhead_component` takes one of a fixed set of ten values. Each rolls the individual pipeline spans up into a category, matching the categories the Bifrost UI's [log-detail overhead breakdown](/features/observability/latency-breakdown) groups into, so the metric and the UI agree:

| Value | UI label | Component |
| - | - | - |
| `serialization` | Serialization | JSON parsing and encoding at the request edges (request/response unmarshal and marshal) |
| `conversion` | Conversion | Translating between Bifrost's unified schema and the provider's native shape, including per-chunk stream conversion |
| `plugins` | Plugins | All plugin hook spans combined, across every configured plugin |
| `middleware` | Middleware | HTTP transport auth and access-control middleware |
| `routing` | Key selection | Selecting which provider API key to use (key-pool lookup and key selection) |
| `processing` | Processing | Internal pipeline glue: request setup, pre/post-hook loops, worker setup and handoff, queue wait, and attribute population |
| `networking` | Networking | Handling between client, gateway, and provider: provider-side processing, request context, response headers, response finalize, request signing, and credential fetch |
| `streaming` | Client delivery | Streaming egress: backpressure and writing chunks back to the client |
| `miscellaneous` | Miscellaneous | Small glue work not worth its own span, plus the residual overhead not attributed to any phase span |
| `other` | Other | Any unmapped bucket (its presence signals a new bucket needs a category) |

The set is bounded, so the list above is exhaustive.

This metric is **off by default**. Enable it with the telemetry plugin's `overhead_breakdown_enabled` config field (a sibling of `metrics_enabled`), or toggle **Enable Overhead Breakdown** on the pull-based tab of the **Observability → Prometheus** page in the UI.

```json theme={null}
{
  "plugins": [
    {
      "name": "telemetry",
      "enabled": true,
      "config": {
        "overhead_breakdown_enabled": true
      }
    }
  ]
}
```

<Note>
  The breakdown is computed from completed trace spans, so it only populates when tracing/observability is active for the request. With tracing off, `bifrost_overhead_component_microseconds` stays empty even when `overhead_breakdown_enabled` is on.
</Note>

### Bifrost MCP Metrics

Emitted for MCP (Model Context Protocol) tool calls executed through Bifrost:

| Metric | Type | Description |
| - | - | - |
| `bifrost_mcp_client_operation_duration_seconds` | Histogram | Duration of an MCP tool call, observed by Bifrost (the MCP client). `_count` is call volume; a non-empty `error_type` marks failures. |

Labels: `mcp_client` (server label), `mcp_tool_name`, `mcp_method` (`tools/call`), `error_type` (`auth_required` / `_OTHER` on failure, empty on success), plus the governance labels `virtual_key_id`/`virtual_key_name`, `team_id`/`team_name`, `customer_id`/`customer_name`, `business_unit_id`/`business_unit_name`, `project_id`/`project_name`, and any custom labels. Only tool executions are recorded (lifecycle `ping`/`list_tools` and codemode tools are skipped); provider/model and `network_transport` are not labels here.

### Default Labels

Most request-level Bifrost LLM metrics include these labels (the `bifrost_key_rotation_events_total` counter is an exception — see [Key Rotation Events](#key-rotation-events) below for its narrower label set):

* `provider` - LLM provider name
* `model` - Model identifier
* `alias` - Alias resolved to this model (empty if none)
* `method` - Request type (chat, completion, embedding, etc.)
* `virtual_key_id` / `virtual_key_name` - Virtual key identifiers
* `routing_engine_used` - Comma-separated list of routing engines that contributed to the decision (e.g. `governance`, `routing-rule`, `loadbalancing`, `model-catalog`, `core`). `core` is emitted when the Bifrost orchestrator itself makes a routing decision — fallback transitions or retry transitions.
* `routing_rule_id` / `routing_rule_name` - Routing rule that matched the request
* `complexity_tier` - Complexity tier used for routing (`SIMPLE` / `MEDIUM` / `COMPLEX`); empty when no routing rule referenced `complexity_tier`
* `complexity_mechanism` - How the effective complexity tier was determined (`semantic`, `llm`, `session`, or `skipped` when no tier was produced). The raw complexity score is deliberately not a label because it has unbounded cardinality; it remains available in request logs and trace attributes
* `selected_key_id` / `selected_key_name` - API key that successfully served the request (`""` when all attempts failed)
* `fallback_index` - Fallback position
* `team_id` / `team_name` - Team identifiers (empty when governance is not used)
* `customer_id` / `customer_name` - Customer identifiers (empty when governance is not used)
* `project_id` / `project_name` - Project the request was scoped to (empty when the request named no project). A request is scoped to at most one project, so these stay singular where team and customer identifiers can fan out

`user_id` / `user_name` are **not** included by default — see [User Labels](#user-labels).

<Note>
  **v1.5.0-prerelease4+**: `selected_key_id` / `selected_key_name` are only populated when the request succeeds. On final errors both are empty — use the `attempt_trail` log field to see which keys were tried.
</Note>

### User Labels

`user_id` and `user_name` identify the end user a request was made on behalf of. They are available on every other observability surface — BigQuery columns, Splunk event fields, Datadog tags, span attributes — but are **off by default** on Prometheus metrics.

Enable them with the telemetry plugin's `user_labels_enabled` config field, or toggle **User labels** on the pull-based tab of the **Observability → Prometheus** page in the UI.

```json theme={null}
{
  "plugins": [
    {
      "name": "telemetry",
      "enabled": true,
      "config": {
        "user_labels_enabled": true
      }
    }
  ]
}
```

<Note>
  Values are populated by the enterprise auth middleware that resolves the calling user. On an OSS build with no user resolution, the labels are present but empty.
</Note>

### Key Rotation Events <sup>v1.5.0-prerelease4+</sup>

`bifrost_key_rotation_events_total` is incremented once per **actual key rotation** — i.e. when a per-key failure causes the next retry to switch to a different key. Rotation-triggering failures are bound to the specific key/account rather than the request:

* `429 Too Many Requests` — this key is rate-limited; another may have capacity.
* `401 Unauthorized` / `403 Forbidden` — bad / revoked key, or key lacks permission.
* `402 Payment Required` — billing issue on this key's account.

It is **not** incremented for:

* terminal failures (no retry happens, including `max_retries = 0` or every key permanently dead),
* same-key retries on transient 5xx / network errors,
* non-retryable request-bound 4xx (400/404/422/...).

Labels are attributed to the key that failed and triggered the rotation:

| Label | Values | Description |
| - | - | - |
| `provider` | e.g. `openai` | LLM provider |
| `requested_model` | e.g. `gpt-4o` | Model as requested (before any alias resolution) |
| `key_id` | UUID | The provider API key that failed and was rotated away from |
| `key_name` | string | Human-readable name of the provider API key |
| `fail_reason` | error type string | Reason the rotation fired: `rate_limit_error` (429), `authentication_error` (401/403), `billing_error` (402), or a provider-supplied error type for non-status-coded rate-limit messages |

To inspect every attempted key on a failed request (including terminal failures that did not rotate), read the `attempt_trail` field on the corresponding log entry instead.

**Example queries:**

```promql theme={null}
# Rate of key rotations per provider
sum by (provider) (
  rate(bifrost_key_rotation_events_total[5m])
)

# Which specific keys are hitting rate limits most often
topk(5, sum by (provider, key_name) (
  rate(bifrost_key_rotation_events_total[1h])
))
```

***

## Push Gateway Setup

If you don't have a Push Gateway running, deploy one:

### Docker

```bash theme={null}
docker run -d -p 9091:9091 prom/pushgateway
```

### Kubernetes (Helm)

```bash theme={null}
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm install pushgateway prometheus-community/prometheus-pushgateway
```

### Configure Prometheus to Scrape Push Gateway

Add to your `prometheus.yml`:

```yaml theme={null}
scrape_configs:
  - job_name: 'pushgateway'
    honor_labels: true
    static_configs:
      - targets: ['pushgateway:9091']
```

<Note>
  The `honor_labels: true` setting is important - it preserves the `job` and `instance` labels pushed by Bifrost instead of overwriting them with the Push Gateway's labels.
</Note>

***

## Pull vs Push: When to Use Each

| Scenario | Recommended Method |
| - | - |
| Single Bifrost instance | Pull (scraping) |
| Multiple instances, direct access | Pull (scraping) |
| Multiple instances behind load balancer | **Push (Push Gateway)** |
| Kubernetes with service mesh | Pull or Push |
| Serverless / ephemeral instances | **Push (Push Gateway)** |

### Why Push for Clusters?

When multiple Bifrost instances run behind a load balancer:

1. **Scraping randomness**: Each scrape may hit different nodes, missing metrics from others
2. **Instance tracking**: Push Gateway properly tracks per-instance metrics via `instance` label
3. **Aggregation**: Downstream tools (Grafana, Datadog) can aggregate across all instances

***

## Troubleshooting

### Push Gateway Connection Failed

```
failed to push metrics to push gateway: connection refused
```

* Verify the Push Gateway URL is correct and reachable from Bifrost
* Check firewall rules between Bifrost and Push Gateway
* Ensure Push Gateway is running: `curl http://pushgateway:9091/metrics`

### Metrics Not Appearing

* Verify the telemetry plugin is enabled (required for metrics collection)
* Check Bifrost logs for push errors
* Verify Prometheus is scraping the Push Gateway with `honor_labels: true`

### Authentication Failed

* Double-check username and password
* Ensure basic auth is configured on the Push Gateway side
* Check for special characters that may need escaping


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.