Enterprise AI Security: A Reference Architecture for Governing Model Traffic
A reference architecture for enterprise AI security, covering virtual keys, RBAC, SSO, guardrails and audit logs for model traffic.
TL;DR
- Enterprise AI security controls, monitors, and governs every request between employees, applications, and model providers, covering identity, access, content inspection, network isolation, and compliance logging.
- Point-to-point integrations cannot be governed consistently, because each application holds its own provider keys and there is no single place to enforce policy or record access.
- The reference architecture puts an AI gateway in front of every provider, with four control layers: virtual keys for identity and cost, RBAC and SSO for authorization, guardrails for content, and audit logs for evidence.
- For regulated environments the gateway runs in-VPC, on-premises, or air-gapped, keeping all processing inside a boundary the organization controls.
- A gateway governs only the traffic pointed at it, so endpoint AI on employee machines stays ungoverned until something routes it.
Enterprise AI security is the practice of controlling, monitoring, and governing every request that flows between employees, applications, and large language model providers. Most organizations connect dozens of applications to model APIs with keys embedded in code, no central policy layer, and no audit trail of what data left the building. Bifrost, the open-source AI gateway built in Go by Maxim AI, is designed for enterprises that need to route, govern, and secure model traffic from a single control plane. This reference architecture describes how to place an AI gateway at the center of enterprise AI security, extend that governance to every endpoint, and deploy it inside your own network boundary.
What Is Enterprise AI Security
Enterprise AI security is the set of controls that authenticate, authorize, filter, and audit traffic between an organization and large language model providers, so that sensitive data, credentials, and spend stay under policy. It spans identity, access control, content inspection, network isolation, and compliance logging across every model, provider, and application in use.
The attack surface is larger than most teams assume. Employees paste customer records into browser chat tools, applications ship provider keys in plaintext, coding agents connect to external tool servers, and no one can produce a record of what was sent where. IBM's 2025 Cost of a Data Breach report found that breaches involving unsanctioned "shadow AI" added roughly USD 670,000 to the global average breach cost, and that 63% of the organizations studied had no AI governance policy in place to prevent it. The OWASP Top 10 for LLM Applications ranks prompt injection and sensitive information disclosure among the most critical risks facing production AI systems.
The control-by-control view of the same problem is in the controls that apply to LLM traffic, and the ungoverned half of the surface in shadow AI in enterprises.
Why Model Traffic Needs a Central Control Plane
Model traffic needs a central control plane because point-to-point integrations cannot be governed consistently. When every application holds its own provider keys and calls model APIs directly, there is no single place to enforce access rules, apply content filters, cap spend, or record who accessed which model. Security controls end up duplicated, inconsistent, or missing entirely.
An AI gateway resolves this by becoming the single entry point for all model traffic. Every request passes through one layer where policy is defined and enforced. The Bifrost gateway unifies access to 10,000+ models across 25+ providers through one OpenAI-compatible API, which means teams can standardize governance without rewriting each integration. The role that layer plays is set out in what an AI gateway is. Consolidating traffic through one gateway produces:
- A single policy surface: access control, budgets, rate limits, and content filters are configured once and applied to every request.
- Full request visibility: every prompt and response routed through the gateway is observable in one place, with content logging configurable per deployment.
- Provider abstraction: applications call one endpoint, and the gateway routes to the correct provider without exposing raw keys to application code.
- Consistent compliance: audit trails, retention, and data-handling rules apply uniformly across all models and providers.
The Reference Architecture: AI Gateway as the Control Plane
The core of this reference architecture is the Bifrost AI gateway operating as the enterprise AI security control plane. All applications, services, and users authenticate to the gateway rather than to individual providers, and the gateway enforces identity, access, content, and cost policy on every request before forwarding it to a model. The sections below describe each control layer.
| Control layer | What it decides | Primary object |
|---|---|---|
| Identity and cost | Who is calling, which models they may reach, and how much they may spend | Virtual keys with budgets and rate limits |
| Authorization | Which operations an operator may perform, and which rows they may see | RBAC roles plus data access control scopes |
| Content | Whether a prompt or response violates policy | Guardrail rules and profiles |
| Evidence | Who changed which control, and when | Audit logs with signing, retention, and export |
Each layer fails differently. Without the identity layer there is no attribution; without the content layer sensitive data reaches providers unredacted; without the evidence layer none of it can be demonstrated to an auditor. Policy-level framing for the same stack is in policy-based governance at the gateway.
Virtual keys: the primary governance entity
Virtual keys are the primary governance entity in Bifrost. Instead of distributing raw provider credentials, teams issue virtual keys that carry their own access permissions, budgets, and rate limits. A virtual key can be scoped to specific models and providers, attached to a team or a customer, and enabled or disabled instantly. Real provider keys stay inside the gateway and are never handled by application code.
Each virtual key supports independent budgets, token and request rate limits, and model and provider filtering. Budgets can be set with rolling or calendar-aligned reset windows, and hierarchical cost control operates across four levels: customer, team, virtual key, and individual provider configuration, checked cumulatively. Rate limits attach at the virtual key and provider configuration levels rather than to teams or customers. This is the governance foundation that lets a platform team allocate spend per project and revoke access without touching provider consoles.
Issuing and managing those credentials is covered in managing virtual keys and budgets.
RBAC and SSO: identity-driven access
Role-based access control governs what each user can view, create, update, or delete across gateway resources. RBAC in Bifrost ships with three system roles, Admin, Developer, and Viewer, and supports custom roles for teams such as security, QA, or compliance. Permissions follow the principle of least privilege, so contractors, auditors, and project teams receive only the access they need.
Identity is federated through your existing provider. Bifrost supports OIDC user provisioning with Okta, Microsoft Entra, Keycloak, Google Workspace, and other identity providers, so accounts, groups, and lifecycle state stay in sync. Roles can be assigned automatically from IdP groups and claims. Data Access Control adds row-level scoping on top of RBAC: a developer on one team cannot see virtual keys, prompts, or routing rules owned by another team unless their role grants broader scope.
Both layers are covered together in the guide to access profiles, RBAC, and DAC, and the enterprise identity pattern in RBAC, SSO, and virtual keys for AI traffic.
Guardrails and PII redaction: content-layer enforcement
Guardrails inspect prompts and responses in real time and block, redact, or reject content that violates policy. Bifrost guardrails are built from reusable rules and profiles, where rules are defined in Common Expression Language and profiles configure the underlying providers. A guardrail runs before a prompt reaches a model and before a response returns to the user.
Coverage spans native and external providers:
- Secrets Detection: Gitleaks-backed detection catches leaked API keys, tokens, private keys, and credentials before they leave the network.
- Custom Regex with PII Detection: in-process regex rules, including a built-in PII Detection template, redact or reject sensitive patterns.
- External providers: Presidio, Azure AI Language PII, AWS Bedrock Guardrails, Azure Content Safety, Google Model Armor, CrowdStrike AIDR, Gray Swan Cygnal, Patronus AI, Check Point's AI Agent Security, and Repello Argus extend content filtering, prompt-injection defense, and safety evaluation.
Because guardrails are configured once at the gateway, PII redaction and credential detection apply to every request without per-application setup. Rule design is covered in enterprise AI guardrails for PII, injection, and toxicity and guardrails at the gateway layer for enterprise AI security.
Audit logs, rate limits, and budgets: the compliance layer
Audit logs record administrative activity so operators can review who changed what, when, and to which resource. Bifrost audit logs can be signed with an HMAC key so entries are verifiable, retained for a configurable period, filtered in the dashboard, and exported as JSON, JSON Lines, or Syslog. Those are the artifacts a SOC 2, ISO 27001, or HIPAA review asks for when it tests a control, and the same records support GDPR accountability obligations. Tool-level activity is covered in MCP audit logs for enterprise compliance.
Rate limits and budgets provide the cost-and-abuse controls that keep model traffic within policy. Rate limits cap token and request volume per period on a virtual key or a provider configuration, and budgets cap spend at the customer, team, virtual key, and provider-configuration levels. When a budget is exhausted, requests against it are blocked and the affected provider is excluded from routing, which prevents a single compromised key or runaway agent from generating unbounded cost. The mechanics are covered in LLM API rate limiting with virtual keys and budgets.
Deploying Inside Your Network Boundary
For regulated industries and strict data-residency requirements, the gateway can run entirely inside your own infrastructure. In-VPC deployment keeps all traffic within a private cloud boundary across AWS, Google Cloud, Azure, Cloudflare, and Vercel, so data processing stays in an environment you control. This model is built to meet HIPAA, SOC 2, and GDPR requirements with network isolation and least-privilege access.
Bifrost supports the deployment patterns enterprises require:
| Deployment model | What it gives you | When to choose it |
|---|---|---|
| In-VPC and private cloud | No external network dependencies, full data sovereignty across AWS, Google Cloud, Azure, Cloudflare, and Vercel | Data-residency requirements, or a policy that model traffic stays in a controlled environment |
| On-premises and air-gapped | Isolated infrastructure where no traffic leaves the facility | Classified or strictly regulated environments |
| High-availability clustering | Multiple nodes with automatic service discovery, gossip-based state sync, and zero-downtime deployments | Any production deployment where the gateway is on the critical path |
Deployment guidance is in open-source AI gateways for in-VPC teams and Bifrost cluster mode for enterprise AI.
Bifrost Enterprise builds on the open-source gateway and adds the controls a regulated deployment needs: SAML single sign-on and role-based access control, cluster mode with automatic failover, adaptive load balancing, audit logs and log exports, secret management through HashiCorp Vault and cloud secret managers, guardrails, and VPC deployment. Teams evaluating options can review the enterprise deployment model to map controls to their compliance requirements.
Extending Governance to the Endpoint with Bifrost Edge
A gateway only governs the traffic that is configured to flow through it. In practice, employees install Claude Desktop, use ChatGPT in the browser, run coding agents in the terminal, and wire MCP servers into their tools, none of which point at the gateway by default. That ungoverned usage is shadow AI, and it is where sensitive data leaves the organization with no audit trail, no budget control, and no guardrails. The scale of the problem is covered in what is shadow AI, and the detection side in shadow AI detection tools for security teams.
The combination is AI Gateway + Bifrost Edge: the gateway is the control plane and policy engine, and Edge extends that same governance to every machine. Bifrost Edge runs on each computer and routes all AI traffic through your Bifrost, so the virtual keys, budgets, guardrails, and audit logs you already configured apply on the laptop, not just in the data center. Edge is currently in alpha, and teams register to be onboarded.
How the combined architecture closes the shadow AI gap:
- Zero per-app setup: Edge routes traffic transparently at the machine level, with no base URLs to change and no SDKs to swap. Users sign in once through the organization's existing SSO, and governance follows the user.
- Endpoint guardrails: because Edge routes through Bifrost, every guardrail configured at the gateway applies to endpoint AI automatically, so PII redaction and secrets detection catch sensitive content before it leaves the machine.
- Fleet-wide rollout: Edge deploys through existing device management platforms including Jamf, Microsoft Intune, Kandji, Workspace ONE, and JumpCloud, and runs natively on macOS, Windows, and Linux.
Edge gives fleet-wide visibility into which AI apps and MCP servers exist, then lets administrators allow or deny each one, enforced on the device rather than as an advisory. This makes the same governance that protects gateway traffic reach the AI running on every desk. The division of work between the two layers is described in closing the last mile of AI governance, and the rollout mechanics in rolling out AI governance with MDM.
Key Considerations for Implementation
A gateway rollout succeeds when it is staged: route traffic first and observe it, bind access to identity rather than to keys, set budgets and rate limits per project, keep signed audit logs exported, and extend to endpoints only once gateway policy is stable. Switching every control to blocking on day one is what gets a gateway routed around.
Each practice below reduces a specific failure mode:
- Start with visibility, then enforce: route traffic through the gateway and observe usage before switching guardrails and budgets from monitoring to blocking. The staged version of this is in the secure AI deployment checklist.
- Model access on identity, not keys: bind virtual keys to teams and roles through your IdP so access changes follow the same lifecycle as employee accounts.
- Set budgets and rate limits per project: hierarchical budgets and rate limits contain the blast radius of a compromised key or misbehaving agent.
- Keep audit logs signed and exported: HMAC-signed audit trails, configurable retention, and automated log exports support compliance review and incident response.
- Extend to endpoints last: once gateway policy is stable, use Bifrost Edge to bring shadow AI under the same controls.
Enterprises comparing gateway options can consult the LLM Gateway Buyer's Guide for a capability matrix covering governance, security, and deployment, and the governance resources for a deeper look at virtual keys and access control.
Frequently Asked Questions
What is enterprise AI security?
Enterprise AI security is the set of controls that authenticate, authorize, filter, and audit traffic between an organization and model providers, so sensitive data, credentials, and spend stay under policy. It spans identity, access control, content inspection, network isolation, and compliance logging across every model, provider, and application in use.
Why is an AI gateway the right place to enforce it?
A gateway is the one component every model request already passes through, so policy applied there is not reimplemented per codebase and cannot be skipped by a service that missed a middleware. Enforcing inside each application instead produces duplicated, inconsistent, or missing controls, with no single record of who accessed which model.
How does a gateway stop provider API keys from leaking?
Applications receive a virtual key rather than a provider credential. Real provider keys stay inside the gateway and are never handled by application code, so a leaked application credential exposes a scoped, revocable key instead of the account. Secrets detection additionally catches credentials being sent in prompt content.
Can prompts and responses be inspected without changing application code?
Yes. Guardrails run at the gateway, before a prompt reaches a model and before a response returns, so PII redaction, secrets detection, and prompt-injection defense apply to every request regardless of which application produced it. No per-application setup or SDK change is required.
Does this architecture work in an air-gapped environment?
Yes. The gateway deploys in-VPC, on-premises, or fully air-gapped, so no traffic leaves the facility and all processing stays in infrastructure the organization controls. Clustering keeps that deployment highly available, which matters because the gateway sits on the critical path for every model call.
What does this architecture not cover?
A gateway governs only the traffic configured to route through it. Desktop chat apps, browser AI, coding agents, and the MCP servers they connect to stay outside policy until something routes them, which is the gap endpoint AI governance closes. Model behavior itself is a separate concern from access and content policy.
Getting Started with Enterprise AI Security on Bifrost
Enterprise AI security depends on a control plane that can authenticate every request, enforce access and content policy, cap spend, and produce an audit trail across all model traffic. Bifrost, built by Maxim AI, provides that control plane as an open-source AI gateway, with virtual keys, RBAC, SSO, guardrails, PII redaction, and audit logs configured once and applied everywhere, and Bifrost Edge extends the same governance to every endpoint. Deployed in-VPC, on-premises, or air-gapped, it keeps model traffic inside your network boundary, in a deployment model built to meet HIPAA, SOC 2, and GDPR requirements through network isolation and least-privilege access.
To see how Bifrost can secure and govern your organization's model traffic end to end, book a demo with the Bifrost team.