Skip to main content

Overview

Guardrail redaction lets Bifrost rewrite sensitive text detected by supported guardrail providers instead of only detecting or blocking it. Bifrost applies span-based runtime redaction for: Providers return findings with ranges and entity types such as EMAIL, PHONE_NUMBER, AWS_ACCESS_TOKEN, or a custom regex entity_type. For providers with Bifrost redaction controls, Bifrost applies the configured action, strategy, and mode. Singulr chooses the action in its own policy configuration; Bifrost applies its valid redaction spans with fixed runtime replacement.
Redaction only applies to text that a guardrail provider detects. If a provider does not detect a value, Bifrost cannot redact that value in runtime payloads, logs, or connector exports.

Bifrost-Managed vs Provider-Managed Rewrites

Bifrost-managed redaction is different from provider-managed transformation.
  • Bifrost-managed redaction means the guardrail provider returns findings, and Bifrost applies the replacement using redaction_strategy and redaction_mode.
  • Provider-managed transformation means an external guardrail provider returns already-transformed text for Bifrost to apply.
Only one rewrite owner can apply to a given request or response phase. If the same phase produces both provider-managed transformed text and Bifrost-managed redaction findings, Bifrost fails closed with a guardrail intervention instead of trying to merge the two rewritten outputs. Bifrost also rejects multiple provider-managed transformed outputs for the same phase because the final replacement would be ambiguous. Detection-only and blocking guardrails can still run alongside Bifrost-managed redaction. The restriction applies when more than one guardrail path attempts to rewrite the same input or output content. Check Point’s AI Agent Security uses Bifrost-managed redaction, but its response shape is different from a local detector. Check Point returns optional message-content spans for supported findings, and Bifrost maps those spans back to the original text before applying the configured strategy and mode. Check Point documents maskable payload spans for PII, profanity, and custom regular-expression detectors. A flagged Check Point result without safely mappable spans fails closed instead of forwarding the original content. Singulr AI uses Singulr-controlled, Bifrost-applied redaction. It selects allow, block, or redact in its own policy configuration. For a redact decision, it returns spans and Bifrost validates then applies fixed runtime replacement to the matching request or response text. Singulr profiles do not expose Bifrost action, strategy, or mode settings. A missing or unsafe span fails closed instead of forwarding the original content.

Actions

The action field controls what happens when the provider finds sensitive text. redaction_strategy and redaction_mode only change request, response, log, or trace content when action is redact.

Redaction Strategies

Strategies control the replacement value used by the non-reversible runtime mode. Reversible modes use numbered placeholders such as [EMAIL-1] so a permitted user can reveal the original values in Bifrost logs.

Redaction Modes

Redaction mode decides where Bifrost applies the rewrite.
Regex guardrail configuration showing the redaction mode selector with Runtime, Logs only, and Runtime plus reversible logs options
For streaming output, runtime redaction checks buffered text segments before releasing their redacted content. Logs-only redaction does not delay client delivery. If the same matched rule set can also block, Bifrost holds the complete stream until the final guardrail decision; see Streaming Output Guardrails.

LLM and MCP Payloads

The same modes apply at both guardrail targets: For reversible MCP redaction, the MCP tool log stores phase-scoped input and output mappings just like an LLM log. A caller with Logs:Reveal can use those mappings on the MCP log detail view; callers without that permission receive only the placeholderized content.

Runtime (runtime)

Use runtime when sensitive text should not reach the model provider or the caller. Bifrost rewrites detected text in the live request or response and stores the already-redacted value in Bifrost logs. Example with redaction_strategy: "replace":
becomes:

Logs only (logs_only)

Use logs_only when the model should receive the original text, but Bifrost logs and trace exports should not store raw sensitive values. Runtime content stays unchanged. Bifrost logs and trace-export connectors receive placeholders:
The placeholder mapping is stored with the Bifrost log row for reveal. It is not sent to connectors.

Runtime + reversible logs (runtime_reversible)

Use runtime_reversible when runtime content should be redacted, but authorized users still need a controlled way to view the original values in Bifrost logs. Runtime content, Bifrost logs, and trace-export connectors use the same placeholder style:

Reveal

Reveal is Enterprise-only and applies only to Bifrost logs. Users need the Logs:Reveal permission to reveal original values for a log that has reversible redaction data. When the caller has that permission, the log detail response can include the placeholder mapping for that log, for example:
Important details:
  • Reveal is scoped to Bifrost logs, not external destinations.
  • The mapping is stored with the log row and is deleted when the log row is deleted.
  • When an encryption key is configured, the mapping is encrypted before storage.
  • The reveal response is marked Cache-Control: no-store.
  • Data Access Control still applies when fetching or revealing a log.
If content logging is disabled, Bifrost does not persist LLM request/response or MCP argument/result content, or redaction reveal data for that log. In that setup, there is nothing to reveal later.

Connector Exports

For trace-export connectors, Bifrost applies raw-to-placeholder replacements before the completed trace is exported. This keeps exported span content aligned with Bifrost log redaction for reversible modes, while keeping the reversible mapping inside Bifrost.
This section describes Bifrost’s completed-trace export path. Integrations that do not consume completed Bifrost traces should not be assumed to receive the same connector redaction behavior.

Provider Defaults

For providers with Bifrost action controls, set action: "redact" explicitly. Relying on defaults is usually the wrong move here, especially for Presidio and Azure AI Language PII. Configure the redact decision in Singulr for Singulr AI profiles.

LLM Tool-Call Arguments

Custom Regex, Secrets Detection, Microsoft Presidio, and Azure AI Language PII include tool-call arguments by default when an LLM rule applies. Input rules cover assistant tool calls in the selected conversation history; output rules cover calls generated by the model. This includes Chat function arguments and Responses function arguments or custom-tool input. Tool names, IDs, and definitions are not redaction targets. Arguments use the same redaction strategy and mode as other fields. Runtime redaction can change what a bash, grep, or other command does; Bifrost returns the redacted arguments without repairing or restoring the command. logs_only preserves the call sent to the client and records redaction mappings for logs and traces. For streaming requests that declare tools, runtime redaction holds output until the complete arguments have been evaluated. Bifrost rewrites argument deltas and terminal copies before replay; it does not call the guardrail provider for every argument fragment. This adds holdback latency. Requests without tools retain the existing text-segment behavior. Under active runtime redaction, unexpected tool calls on requests without declared tools or tool-call history are rejected, including calls arriving before any text. MCP-targeted rules continue to evaluate actual MCP execution arguments and results through their existing adapters. Provider-managed transformations and Lakera retain their existing argument-mapping restrictions.

Edge Cases

  • Redaction is text-based. It does not inspect image pixels, audio, or arbitrary binary content.
  • Custom Regex uses Go’s RE2-compatible regexp engine.
  • Overlapping findings are resolved into a non-overlapping set before replacement.
  • Bifrost-managed redaction cannot be combined with provider-managed transformed output for the same request or response phase.
  • Check Point tool-call arguments are screened but are not rewritten from Check Point message-content spans. A flagged argument that cannot be mapped safely fails closed in redact mode.
  • Singulr tool-call arguments and tool-result content can be blocked but are never redacted.
  • Input redaction cannot safely run together with raw-body passthrough transformations; Bifrost fails closed rather than forwarding an inconsistent payload.
  • If a request uses both input and output redaction, Bifrost carries replacements forward so raw log fields and exported trace content are redacted consistently.