Evidence-required mode
Fail closed before new AI or MCP work when durable evidence is unavailable, with health-aware recovery when the sink returns.
AssuranceOps is the self-hosted governance control plane for AI agents, MCP, and model traffic — enforcing policy inline and producing durable, independently verifiable evidence.
One identity, one governed decision, one independently verifiable evidence bundle.
Any team can call any model provider, enable any tool, or connect any MCP server — with no central record.
When compliance asks who used AI, for what, and at what cost, there is no answer to give them.
Approved-model lists live in a wiki. Nothing in the request path actually stops a violation.
The latest release tightens the full governance loop: identity lifecycle, least-privilege review, MCP supply-chain and runtime threats, policy-safe resilience, and evidence that can fail closed and verify offline. New skill-execution and storage foundations are published with their production gates visible.
Fail closed before new AI or MCP work when durable evidence is unavailable, with health-aware recovery when the sink returns.
Export one consistent snapshot with chain verification, read-model checks, canonical bytes, and optional Ed25519 attestation.
Scan tool arguments before execution and tool results before they re-enter agent context; result findings remain alert-only.
Provision, deactivate, reactivate, or delete users through SCIM 2.0 and suspend every matching gateway principal in one action.
Match declared MCP packages to a reviewed local OSV artifact with conservative handling for unknown versions.
Compare policy-declared access with observed behavior to expose unused privilege, denied attempts, and current-policy drift.
Retry once through a configured backup only when the same policy and sensitive-data approvals authorize it.
Detect machine secrets deterministically and add opt-in model-based NER for unstructured sensitive narratives.
Tenant-owned publisher and approval evidence binds exact readiness material, short-lived grants, and canonical hash-chained run receipts. The boundary remains fail closed until endpoint execution proof is complete.
The endpoint executor accepts only grant-derived capability profiles with closed filesystem, network, process, and environment modes plus bounded resources — never caller-supplied commands or paths.
MySQL now covers 27 of 29 store contracts with live conformance on the skill-execution slices; SQL Server has an explicit adapter gate. Complete runtime parity and provider certification remain open.
JSONL and database verification use fixed-size readers and per-row streams with a one-row lookahead, so peak memory no longer scales with chain length.
Governance does not stop at the gateway. AssuranceOps brings the same fail-closed discipline to pilot endpoint releases: exact signed artifacts, bounded delivery, anti-rollback checks, health-gated rollout, and evidence that follows the update.
A closed release manifest binds the exact version, payload, and updater identity. Endpoints verify it against a locally pinned Ed25519 key.
The agent accepts only a fixed HTTPS origin, with redirects disabled, strict size bounds, streamed SHA-256 checks, and no-clobber local staging.
Monotonic release sequence and semantic-version checks refuse stale or replayed artifacts before an update can advance.
Target rollout rings and scheduled health gates keep a release scoped until observed fleet health earns the next promotion.
The updater is fixed-purpose: sealed product profiles, fixed service adapters, no arbitrary shell execution, and rollback after a failed health confirmation.
Content-minimized, predecessor-linked receipts preserve staged-to-committed or rollback history for the operator evidence plane.
Clients hold virtual keys. Real provider credentials never leave the gateway.
Providers, models, tools, MCP servers, input size, and budgets are evaluated per request.
Allowed requests pass through unmodified. Denied requests fail before a token is spent.
Every request is written to an immutable trail, priced, and attributed by activity.
Three steps stand between a pilot and production — none of them is an application rewrite.
The gateway is a single static binary with a Docker Compose quick start. For production, the shipped Helm chart runs it highly available on Kubernetes — budget and rate-limit counters are shared across replicas through Postgres, and separate liveness and readiness probes keep rolling updates safe.
Clients keep calling the standard OpenAI and Anthropic API paths. Adoption is a base-URL change plus a virtual key — no SDK swap, no application rewrite, and streaming responses pass straight through.
Each client key or directory group maps to a named policy — or starts from a shipped compliance pack. Policy is plain YAML with secrets referenced by environment variable, so it is reviewable, versionable, and safe to commit.
The properties security and platform leaders ask about first — each one is architectural, not a configuration flag someone can forget.
Real provider keys are referenced by environment variable, resolved at startup, and injected only at the moment a request is forwarded. The config file never contains a secret, and clients only ever hold revocable virtual keys.
Users authenticate with JWTs from Entra ID, Okta, Google Workspace, or ADFS. Directory groups map to policies — fail closed — and every audit record carries the actual user, not just a key. Service accounts keep virtual keys.
Anything the gateway cannot authenticate, parse, route, or that policy does not explicitly permit is rejected before a token is spent. An unknown policy is a deny, not an error.
Allowed, denied, over budget, or rejected at the front door — every terminal path writes exactly one audit record. There is no code path that skips the trail.
Allow/deny rules over providers, models, tools, and MCP servers — per team. Deny always wins; default-deny allow-lists.
Regulated deployments can refuse new AI and MCP work when the durable evidence sink is unhealthy, then recover automatically when the store is available again.
Cache-aware provider pricing, reasoning-token visibility, proactive budget warnings, and hard daily or monthly caps enforced inline.
Tool arguments are scanned before execution; untrusted tool results are scanned before they re-enter agent context, with content-free findings written to evidence.
Deterministic credential and identifier detection combines with an optional model tier for names, addresses, dates of birth, medical narratives, and other unstructured data.
Versioned, policy-gated templates with typed variables, static lint scores, and runtime quality metrics.
A pre-authorized backup can handle an unreachable, 5xx, or rate-limited primary only when the same policy and data-protection checks approve it.
One repeatable database snapshot is cross-checked against its read model, canonically serialized, and optionally signed with Ed25519 for offline verification.
Entra ID, Okta, Google Workspace, and ADFS identities map to policy; SCIM 2.0 deactivation suspends every principal key derived for that user.
An approved-app catalog enforced by an inline forward proxy and decision API, with a shadow-IT discovery feed for what is not yet catalogued.
Webhook notifications to Slack, Teams, or a generic endpoint on budget breaches, denial spikes, and sensitive-data blocks — with per-rule cooldowns.
An optional cap per key or user on top of the policy budget, plus token-bucket rate limits — over-limit requests get 429 before any upstream call.
Managed machines enroll once through the aops-agent binary and heartbeat on a scoped, one-time device credential — rotate or retire any endpoint from the operator UI without touching the team's virtual key.
MCP servers and skills are cataloged and reviewed through a transactional approval workflow before they're discoverable; a caller still only sees skills its own policy explicitly allows.
Each declared MCP implementation package can be checked against a reviewed local OSV artifact without live dependency resolution or runtime package installation.
Compare policy-declared models, providers, tools, and MCP servers with observed use to surface unused access, denied attempts, and policy drift.
The audit trail, budgets, and approvals persist to PostgreSQL you control — AWS RDS, Azure, Google Cloud SQL, on-prem, or Kubernetes — with engine-neutral configuration and versioned migrations.
Stage only pinned, signed, hash-verified endpoint artifacts; reject rollback before a fixed-purpose updater can act, and retain content-minimized lifecycle evidence.
Tenant-owned publisher, approval, readiness, and capability evidence can issue exact short-lived grants and canonical run receipts; native capability evaluation and production execution proof still gate activation.
Verify JSONL and database chains through fixed-size readers and row streams so peak memory stays bounded as evidence grows, without changing verification results.
The gateway supplies the enforcement and evidence. The operations layer turns that evidence into posture, investigation, response, and reporting workflows your security and compliance teams can use.
Turn policy, provider, and platform checks into a 0-1000 posture score with prioritized remediation and framework references.
Coalesce governance events into an operational queue, then approve or pre-authorize playbooks that notify, suspend a key, or change policy.
Detect provider traffic that bypasses the gateway, import existing proxy logs, and correlate new host or model indicators against prior activity.
Produce scheduled governance reports, top-offender views, and a single pass/warn/fail sweep for leadership, auditors, or CI.
Scope usage, alerts, reports, identities, and governance views by organization for business units, regulated subsidiaries, or managed services.
Track posture drift, review sampled AI usage for nuanced policy evasion, and flag risky configuration changes before they merge.
Apply policy to durable agent memory operations, preserve attribution, and keep memory-related threat findings in the same evidence and response workflows.
Every request runs the same ordered gates. Deny always wins, and each allow-list is default-deny — anything not explicitly permitted is rejected.
The gateway, policy, and audit trail run inside your infrastructure. Provider credentials never reach clients, and request content is sent only to the provider or MCP server that policy approved.
The built-in dashboard at /admin — shown with sample data.
Three starter packs pre-wire policy controls to named regulatory requirements. Every allow-list ships empty or as a placeholder — your team names the approved providers, models, and prompt templates before it goes live.
Restrictive-by-default baseline for protected health information — provider and model allow-lists scoped to your BAA, PHI detection that blocks on match, and minimum-necessary limits on prompt templates and input size.
Baseline for personal data of EU/EEA subjects — provider and model allow-lists pinned to adequacy- or SCC-covered endpoints, with detection set to flag rather than block so lawful-basis processing isn't rejected outright.
Baseline for nonpublic personal information and cardholder data — Luhn-validated card detection, vendor-scoped provider allow-lists, and an optional full-transcript mode for books-and-records retention.
SOC 2 is a control mapping rather than a policy pack. Virtual keys and policy scoping support logical access; the audit trail and denial reason codes support monitoring; and config-as-code — reviewable YAML with every secret referenced by environment variable, never written in the file — supports change management.
Pattern-based detection on request input, reported by kind and count only — the matched value itself is never logged or stored anywhere.
Every request — allowed or denied — writes one audit record: who made it, which policy applied, the disposition and matched rule, PII flags, and latency.
Input scanning stops sensitive data going to a provider. Response scanning catches it coming back — PHI the model was given elsewhere and echoed, or a generated value that looks real.
Every record is SHA-256 hash-chained to its predecessor, so inserting, deleting, reordering, or editing a committed record forces recomputing every hash after it — and is detectable.
One line in the append-only audit trail — attributed, priced, and policy-tagged. Denied here before a token was spent.
Compliance packs and detection accelerate the work — they are scaffolding, not legal advice or a certification. Pattern-based detection has real gaps: unformatted identifiers, names, and free-text narratives are not caught, so layer a dedicated DLP or NER service where your obligations demand it. Full control mapping in docs/COMPLIANCE.md ↗.
Provider and MCP proxying, streaming, policy, identity, budgets, scanning, transcripts, and tamper-evident audit evidence are implemented and under active pre-1.0 validation.
Posture scoring, alerts, response automation, threat intelligence, executive reports, multi-tenancy, and guardrail observability are implemented.
Versioned templates, quality metrics, LLM-as-judge review, content-hash-bound approvals, and approval-gated rendering are implemented.
OIDC authorization-code sign-in with PKCE, tenant-scoped RBAC, and short-lived Postgres-backed browser sessions are implemented.
Docker Compose, Kubernetes and Helm, shared Postgres counters, health/readiness probes, concurrent migration safety, and a Coolify staging profile are available.
The skill-execution control plane and sandbox substrate are delivered as a fail-closed boundary; native capability evaluation, complete adapter parity, provider certification, end-user apps, external anchoring, and broader production proof remain.
Request access for deployment guidance and roadmap conversations, or run the open-source demo now to see policy enforcement and verifiable evidence end to end.
By joining, you agree to our Privacy Policy and Terms.