Home/Products/Bud SENTRY
Bud SENTRY · Layer 06 · Security & Governance

One plane for security, governance & FinOps — built in, not bolted on.

Secure Evaluation and Runtime Trust for your models. Zero-trust ingestion, CPU-native guardrails at 0.70ms, 160+ policies, budgets enforced at admission — one audit trace across the whole stack.

Overview

One plane for the whole stack.

Security, governance, and FinOps on the same plane: every model verified at ingestion, every call guarded and budget-checked, every step in one immutable audit trace.

0.70ms guardrail p50 latency 300+ ready-to-use probes 160+ policies 15–18× cost-performance vs. GPU guardrails
  • Beat 1 — Where it sits. Bud SENTRY is Layer 06 of the eight-layer Bud stack — the Security and Governance layer. One plane for security, governance, FinOps and audit; a zero-trust model supply chain, verified before it ever runs; Bud Sentinel inside, guardrails at 0.70ms p50 on CPUs; 160+ policies, your IdP on day one, one immutable trace.
  • Beat 2 — The trust problem. A downloaded model isn't just weights: it can carry executable scripts, obfuscated binaries, or embedded malware. And governance bolted on across separate tools leaves seams between every pair.
  • Beat 3 — The zero-trust supply chain. Every model passes six gates before it ever runs: sandboxed fetch, multi-format scan (.pickle, .safetensors, GGUF, ONNX), exploit and trojaned-weight detection, a gated registry that admits verified artifacts only, deployment with a verified-state seal, and continuous monitoring against that baseline. Verified before it ever runs.
  • Beat 4 — Sentinel, the guardrail engine inside SENTRY. Every input and output classified for jailbreaks, prompt injection, PII across 11 regions, and content safety at 0.70ms p50 — CPU-native, no GPU anywhere. 65,536 tokens classified natively. About $0.10 per 1M classifications versus about $24 on GPUs — 678× faster on the same silicon.
  • Beat 5 — One gate for policy, identity and FinOps. Admission is a single gate: RBAC validated against your existing IdP (Okta, Entra ID, Google Workspace), 160+ policies checked, rate limits and budgets applied. An over-budget request is held, not destroyed — cost governance is enforcement on the same plane as security.
  • Beat 6 — One trace, the whole lifecycle. A production request crosses the platform in eleven governed steps — Studio, SENTRY admission, Sentinel input guard, Gateway, Agent, MCP Foundry, Runtime, Sentinel output guard, Scaler, SENTRY audit record, ART. SENTRY owns four of them, and the result is one immutable, tamper-evident trace — so audit is an export, not a forensic exercise across scattered logs.
  • Beat 7 — Benchmarked, not asserted. Every number is measured, not asserted: 0.70ms p50 per classification on CPU; $0.50 per hour versus $2–3 for GPU guardrails; 15–18× cost-performance; $0.10 per million classifications at CPU pricing throughout; ~25ms edge inference, 96× faster on-device; 124 million classifications a day. Methodology: single node, Intel Xeon, sustained load. Intel Corporation: "By optimizing Sentinel for Intel Xeon processors, Bud enables enterprises to scale AI safety with high performance and significantly lower total cost of ownership, without the cost and complexity of GPUs." — Sangeeta Roy, Director, Global Partner Business Leadership, Intel.
  • Payoff — Built in. Not bolted on. Bud SENTRY — the governance plane — runs everywhere: cloud, on-prem, air-gapped, edge. Nothing needs a GPU and nothing calls home.

Zero-trust model ingestion

Every download triggers the security pipeline: sandboxed fetch, multi-format scanning (.pickle, .safetensors, GGUF, ONNX, more), exploit and trojaned-weight detection, gated registry.

verify before it ever runs

Sentinel — runtime guardrails

The guardrail engine inside SENTRY: 23 specialised models across 33 variants, CPU-native via Resource Aware Attention — 0.70ms p50 at 10K concurrent connections, #1 balanced accuracy (84.56%).

0.70ms p50 · CPU-native

Compliance enforcement

SOC 2, GDPR, and EU AI Act-aligned controls with full forensic audit trails — honest posture: certified vs. in-progress is documented on the Trust hub.

audit-ready by default

RBAC & identity

Granular role-based permissions across modules, projects, and APIs. Identity brokering via Okta, Entra ID, Google Workspace (OIDC, OAuth 2.0, SAML 2.0).

your IdP · day one

Continuous monitoring

Cluster, system, and traffic monitoring on deployed models — drift from verified state raises the alarm.

drift → alarm

Immutable audit

Tamper-evident logs for prompts, responses, model versions, and every access event — one trace spans the full agent lifecycle.

one trace · full lifecycle
Value proposition

Five bolted-on tools become one plane — security, governance & FinOps.

Guardrails, compliance, audit, access control, and spend limits usually live in three to five tools with no shared data model — and every boundary is a seam.

Guardrail p50 · CPU0.70ms
Attack success & false refusalboth <20%
Throughput · one Xeon node4,400+ req/s
Native context65,536 tokens
CPU-native serving benchmarks · methodology in the product brief
Key features

The only deployable operating point.

Six reasons SENTRY doesn't trade accuracy for latency or hardware.

01

Accuracy you can ship

The only evaluated model with attack-success and false-refusal rates both under 20%.

<20% ASR & FRR15.97% / 14.92% · only occupant
02

CPU-native speed

Single-digit milliseconds on commodity CPUs — faster than the baselines running on an A100.

678× fastersame silicon · 0.70ms p50
03

Long context, one call

Full-length inputs classified natively — a single call, not a chunk-and-vote pipeline. Transformer baselines cap at 512 tokens.

65,536 tokensnative · baselines cap at 512
04

Zero-trust supply chain

Every model pull passes six gates — sandboxed fetch, multi-format scan, exploit and trojaned-weight detection — before it ever runs.

6 gatesevery pull · verified before it runs
05

One immutable trace

SENTRY closes the loop on every production request with a tamper-evident record — audit is an export.

1 trace11 steps · 4 owned by SENTRY
06

No GPU, no egress

The same enforcement plane runs in the cloud, on-prem, air-gapped, and at the edge — nothing needs a GPU and nothing calls home.

~25ms at the edge96× faster than rivals on edge CPUs
Featured components

What's inside the plane.

The guardrail engine, the serving path in front of it, and the evaluation machinery around them — the parts of SENTRY worth knowing by name.

Component page

Bud Sentinel

The guardrail engine inside Bud SENTRY — 23 specialised models across 33 variants classify every input and output, CPU-native via Resource Aware Attention.

0.70ms p50 · 300+ probes
Serving path

Bud Guardrail Gateway

The path every classification travels — Sentinel's model family served from a single binary, on the CPU fleet the application already runs on.

one binary · 33 variants
Evaluation

Bud Evals

Native evaluation and red-teaming — immediate feedback on accuracy, safety, and compliance during agent development, before anything ships.

140+ benchmarks · 300+ probes
Go deeper

The full story, in depth.

The architecture, threat coverage, benchmarks, and methodology behind the claims, in full.

Get started with Bud

Put your data on it.

The fastest way to see what an integrated AI operating system does for your enterprise is a proof-of-concept on your infrastructure, with your data.

01 Identify a use case where complexity, cost, or governance is a known pain point.
02 Joint discovery — Bud maps your AI pain points to platform capabilities.
03 POC in days, on your hardware, with your data.