Home/ Products/ Bud Data Foundry/ Product Brief
Bud Data Foundry overview
Product Brief · Layer 06 · Data & Context

Bud Data Foundry

AI that knows your whole organization — built on the systems you already run, with the permissions you already set. Documents, databases, APIs, live streams, logs and metrics across 200+ sources become one living model that people and AI agents can ask, analyze and act on. Everyone sees only what they are allowed to see; every answer shows where it came from; nothing about your existing systems has to change.

Product reference v1.0 October 2026 ~12 min read
01At a glance

A living model of your organization — for people and for AI agents.

Data Foundry turns everything your organization knows — across every system you use — into one living model that people and agents can ask, analyze and act on. Every one of its five capabilities checks what a person or agent is allowed to see at every step, not just at the end. The data stays where it lives. The answer is sourced, or Data Foundry says it doesn't know.

Sources connected200+
Permission checks on every answer2
Layers in the digital twin3
Systems migrated or replaced0
the five capabilities are detailed in §03 · every number's basis in §06
What it is
A single end-to-end data system for the enterprise — connectors, preprocessing, indexing, retrieval, an ontology and knowledge graph, and data agents at scale — the layer of the Bud Novaria AI OS that turns enterprise data into context
A digital twin of the organization: a shared vocabulary of what the business is about, a knowledge graph of how people, teams, customers, products, systems and datasets relate, and a private understanding of who is asking
Built for AI agents as the main user — real users with their own identity and limits, or acting for a person and seeing only what that person sees — with oversight that can see oversharing, not just avoid it
What it is not
✕Not a chat window over your documents — text fragments nothing can reason over, with access worked out afterwards, if at all
✕Not another copy of your data to protect — the working copy can be rebuilt from your systems at any time and thrown away without touching them, and the most sensitive sources are asked live rather than copied
✕Not a cloud-only service — it runs on your own infrastructure, in a fully offline environment, or as a managed service, with your knowledge and your permissions never leaving your control
02Where it fits

Layer 06 — the ontology layer, the one role nothing else in the stack has.

Data Foundry grounds everything above it. Bud Agent and Bud Studio ask it; Bud MCP Foundry's tools act through it; Bud SENTRY governs it; Bud AI Foundry serves the models that read it, on fleets GPU Foundry and LayerZero run.

Feeds upward — agents in Bud Agent and Bud Studio reason over the knowledge graph and act through Bud MCP Foundry's tools, under Bud SENTRY's policies, seeing only what the person they work for sees.

Builds on Bud AI Foundry's serving plane for the models that read and answer, and on Bud GPU Foundry and LayerZero for the capacity they run on — any model, any silicon, inside your boundary.

Why most "AI over your company's knowledge" projects stall

Most organizations start the obvious way: connect a few systems, put a chat window on top, and see what happens. It's quick to stand up — and three things reliably sink it before it gets past the pilot. Data Foundry was built to avoid all three.

  1. Gathering everything first and thinking about access later.Copying documents into one place and working out who may see what afterwards — if at all. Once the original access rules are left behind they're hard to put back, and the filtering added later leaks through result counts, previews and suggestions. Point an AI assistant at it and the leak runs itself.
  2. Treating your knowledge as a heap of text.Chop it up, look for bits that sound similar, hope for the best. No sense of who works on what, which system holds which data, or how things relate — and nothing for an agent to reason over. It can't connect the dots, can't tell two people apart, and gives the CFO and the intern the same answer.
  3. Sending your most sensitive information somewhere you don't control.Many tools run only in their own cloud. For banks, governments, defense and anyone with rules about where data may live and who may hold it, that's not a trade-off — it's a deal-breaker before the conversation starts.
03Capabilities, in full

Five things that build on each other — and the oversight around them.

Every one of them checks what a person or agent is allowed to see at every step, not just at the end.

01Connect200+ sources — the tools people use every day and the data underneath them: databases, warehouses, APIs, live streams, logs, metrics, time-series · access rules picked up with the content · kept current in the background · missing one, we add it200+ sources
02ComprehendEvery kind of data read the way a specialist would · documents keep layout, tables and formulas · databases and spreadsheets stay structured and queryable · streams, logs and metrics keep their timing and shape · scans recovered as well as the original allows · what couldn't be read is flaggedtables stay tables
03OrganizeAn ontology — a shared vocabulary and a running catalog down to each table and field · a knowledge graph of people, teams, customers, products, systems, datasets and how they relate · a private understanding of each person and each agent3 layers
04Find & answerFor people: exact-phrase and meaning-based search, plain-English questions with sourced answers · for agents: the same through a standard interface, plus direct questions of the data — totals, trends, comparisons, anomalies — with every number traced back · all inside what the asker is allowed to seesourced only
05Act & learnAgents analyze, decide and act under control · every action previewed, checked against your rules, carried out, then confirmed · anything that can't be undone needs a person · what happens feeds back into the model4 steps per action
06Permissions & oversightThe check before the search, for the specific asker, against live rules — and again on every answer · losing access jumps the queue · three sensitivity tiers · oversharing detection · one complete, exportable record of who saw what, when and why2 checks · 1 record

Permissions built in — not bolted on

Data Foundry starts from one simple idea: who is allowed to see what is part of your knowledge, and deserves the same care. Four things follow.

rule 1Access rules come in with the contentWhen Data Foundry connects to a system, it brings the documents and the rules about who can see them together, from day one. It recognizes that "j.smith" in one system and "Jane Smith" in another are the same person — and when it isn't sure, it asks rather than guesses.
rule 2The check happens before the AI sees anythingIt doesn't gather widely and then hide things. It only looks at what you're allowed to see in the first place — so nothing else can leak through a count, a preview or a suggestion.
rule 3Losing access comes firstAn out-of-date document is a small problem. An out-of-date permission is a breach. So when someone's access is removed, that change jumps the queue.
rule 4Every answer is double-checked as it's givenEach source behind an answer is re-checked against the live rules at that moment — so access removed a minute ago can't slip through.

The guarantee in one line. The permission check happens before the search, for the specific person asking, against the live rules — and again on every answer. It isn't a setting anyone has to remember to turn on.

Data gravity — nothing to migrate, nothing to replace

Data has gravity. It piles up where it's created — in the databases, warehouses, streams and systems that run the business — and the bigger and more sensitive it gets, the harder and riskier it is to move. Data Foundry is built so you never have to.

How a source is readWhat it meansTypical sources
A working copyData Foundry reads from your system and keeps its own working copy — but your system stays in charge. That copy can be rebuilt from it at any time and thrown away without touching it. Nothing moves out, nothing gets replaced, and Data Foundry never becomes another place you have to protect.SharePoint · Slack · email · documents · Excel · tickets · wikis · code
Ask the source liveWhere copying doesn't make sense — a very large warehouse, a stream that changes every second, data too sensitive to duplicate — Data Foundry asks the source directly at the moment of the question. For Snowflake or Databricks it can work from their metadata and semantic views instead of copying the data into Bud.Snowflake · Databricks · warehouses · streams · logs · metrics · time-series
You decide, source by sourceThe choice is yours per source, and both paths sit behind the same permission check and the same citation. Unstructured and structured sources are processed alongside each other, so one answer can cite a document and a table.any mix

This is why Data Foundry gets more useful with every source you connect, not more unwieldy. Your data stays where it is; what grows is the understanding of it — and people and agents alike come to it as the one place to ask. It's also why "delete this everywhere" actually means everywhere.

The digital twin — ontology, knowledge graph, and who is asking

Connecting sources is the easy part. What turns a pile of data into something an AI agent can reason over is structure and meaning. Data Foundry builds that in three layers, and together they form a living model of the organization.

LayerWhat it holdsWhat it makes possible
The ontologyA shared vocabulary of what your business is about — people, teams, customers, products, projects, systems, datasets — and a running catalog of everything you have, down to each table and field. People and agents can both enrich it.Every search, answer and agent works from the same definitions. An agent asked about "churn" or "SLA breaches" knows what those mean here, and which data answers them.
The knowledge graphHow it all connects: who belongs to which team, which system holds which data, which customers and products a record or a metric concerns, how work and data flow between them.Reasoning, not just retrieval. An agent can follow a thread across systems — from a customer to their tickets to the service metrics behind them — spot the pattern, and explain why.
Who is askingA private understanding of every individual — and of every agent with its own identity: role, what they work on, what they know, what they've asked before. People can see theirs and correct it; you set each agent's.Two people ask the same question and each gets the answer that fits them. An agent working for someone inherits the same understanding — and the same limits — so it is personal and safe from day one.

Personal never means permissive. Knowing someone changes what Data Foundry shows them first — never what they're allowed to see.

Answers you could put in front of a regulator

An answer is only as good as the reading behind it. Read carelessly, tables become run-on text, spreadsheets lose their formulas, a metric loses its units, a log loses its timing, and the numbers a business runs on quietly go wrong.

A faithful copy of every source

Documents keep their tables, formulas and layout. Databases and streams keep their structure, types and timing. Anything that couldn't be read is flagged rather than glossed over, and Data Foundry can show the exact passage or record an answer came from.

Numbers that stay numbers

Ask about figures — a total, a trend, a comparison — and Data Foundry works it out from the actual data, then shows which table, which rows, as of when. The same for an agent. Nobody needs technical skills or database access.

Sourced, cited — and willing to say no

Every answer comes only from real material you're allowed to see, with sources you can click through to. When the answer isn't there — or isn't there for you — Data Foundry says so. "Not in your accessible knowledge" is a real answer. Making something up is not.

Built for AI agents — with oversight that can see oversharing

People are one kind of user. AI agents will be the main one — and Data Foundry treats them that way from the ground up. An agent connects through a standard interface and gets everything a person gets: search, sourced answers, questions of the data, the knowledge graph to reason over, and the ability to act. And it is a real user, not a borrowed login — whether it is working on someone's behalf or on its own.

An agent that is…seesand
Working for a personOnly what that person sees. Every request is answered as the person the agent is working for.Two people using the same agent get different answers. Handing work to an agent can only ever narrow what's visible, never widen it.
Working on its ownExactly the sources, areas and actions you grant it — just as you would for a new member of staff — changeable at any time.Within those limits it builds up its own knowledge and its own memory of what it has learned, kept separate from any person's. Nothing is inherited by default.
About to actA preview of the action, checked against your rules.Carried out, then confirmed. Anything that can't be undone needs a person's approval first.

It can spot oversharing, not just avoid it. Because Data Foundry understands your access rules, it can see where something is open to more people than it should be — and tell you. Every access, every action, every permission change is kept as evidence you can export: one complete record of who saw what, when, and why.

04How it works

Permission first. Then the search. Then the check again.

How a question is answered, how an agent acts, and the three tiers of handling you choose per area.

How a question is answered

01Who is askingA person, or an agent — working for a person, or with an identity of its own.
02Scope set firstThe permission check runs before the search, for the specific asker, against the live rules. What they may not see is never looked at.
03Meaning resolvedThe ontology knows what the words mean here and which datasets answer them; the graph says how they connect.
04Sources askedThe working copy for documents; the warehouse through its semantic view; the stream at the moment of the question.
05Re-checked and citedEvery source re-checked against the live rules, then the answer cited to the passage, the table and rows, as of when — or "not in your accessible knowledge".

How an agent acts

01PreviewThe action is shown before it happens — what will change, where.
02CheckAgainst your rules and the agent's own limits. Anything that can't be undone waits for a person.
03ActCarried out through the systems you already run, under the identity of the agent or the person it works for.
04Confirm & learnConfirmed and recorded. What happens feeds back into the model, so it gets sharper for every person and every agent.

Different handling for different sensitivities

You choose per area. The tiers change how material is stored and checked — never who may see it.

tier 1Ordinary materialIndexed in the working copy, checked before the search and re-checked on every answer like everything else.
tier 2Sensitive materialRe-checked on every single result, so that a permission removed a moment ago can't surface through a count, a preview or a suggestion.
tier 3The most protected materialData Foundry never stores the text of it at all. It is asked at the source, at the moment of the question, inside the asker's permissions.
05Deployment & compatibility

Runs wherever your rules require.

On your own infrastructure, in a fully offline environment, or as a managed service — with your knowledge and your permissions never leaving your control. That is what a regulated organization can actually sign off.

Where it runs, what it reads, who asks

Where it runs

Your infrastructureFully offlineSovereign cloudManaged service

Inside your own boundary when that's where it needs to be. Runs on the Bud Novaria AI OS — any model, any silicon, on-premise, hybrid or air-gapped — so the models that read and answer never leave the boundary either.

Sources

Shared drivesWikisHelp desksCRMChatEmailCode
DatabasesWarehousesAPIsStreamsLogsMetricsTime-series

200+ sources, including SharePoint, Slack, documents and Excel alongside Snowflake and Databricks through their metadata and semantic views. If one of yours isn't on the list, we add it.

Interfaces

People — search & askAgents — standard interfaceYour codeBud Agent · Bud Studio

People search and ask in plain English. Agents connect through a standard interface and get everything a person gets — search, sourced answers, questions of the data, the graph to reason over, and the ability to act — as real users with their own identity.

Nothing moves out. The only thing that comes in is content together with the rules about who may see it. What never moves is your data: the working copy is rebuildable and disposable, the most sensitive sources are asked live, and "delete this everywhere" actually means everywhere.

06Proof & methodology

Every headline number, with its basis.

The figures on this page are properties of the design, not benchmark results. Each is paired with where it comes from.

200+
Sources connected
BasisThe connector catalog at release — the everyday tools and the data underneath them. Treat it as a floor: a source that isn't on the list is added for the customer that needs it.
2
Permission checks on every answer
BasisThe permission model: one check before the search, for the specific asker, against the live rules; a second on every source behind the answer at the moment it is given. Sensitive material is additionally re-checked on every single result.
3
Layers in the digital twin
BasisThe ontology (vocabulary and catalog), the knowledge graph (how it connects), and who is asking (the private understanding of each person and agent). Together they are what an agent reasons over rather than searches.
0
Systems migrated or replaced
BasisThe data-gravity design: Data Foundry reads from your systems and keeps a working copy it can rebuild or delete at any time, or asks the source live. Your systems stay in charge; nothing moves out and nothing is replaced.

How these claims are framed. Source counts, check counts and layer counts are properties of the design. Deployment-specific numbers — answer latency on your hardware, coverage of your estate, the share of questions answered from the working copy versus the live source — are established in a proof-of-concept on your infrastructure, with your data and your permissions.

07Positioning · as of Q4 2026

Against the usual approach.

Side by side with the way most organizations first try this — a chat window over their documents, built in-house or bought off the shelf.

Chat over your documentsBud Data Foundry
Who sees whatWorked out afterwards, if at all; the assistant usually holds a master keyRules captured with the content, checked before the search, re-checked on every answer
What it coversDocuments, maybe a database200+ sources — documents, databases, APIs, streams, logs, metrics
UnderstandingText fragments; nothing an agent can reason overAn ontology and knowledge graph — a digital twin of the organization
Your existing systemsLoaded into a new store that becomes one more copy to manageStay in charge; the working copy can be rebuilt or deleted at any time
Reading qualityChopped into fragments; tables and numbers flattenedFaithful copy of every source; tables and streams stay data; originals one click away
AnswersFluent, sometimes invented; sources best-effortSourced only, cited to the exact passage; says "I don't know" rather than guess
AI agentsGiven free run of everything, or bolted on as an afterthoughtReal users with their own identity and limits — or acting for a person and seeing only what that person sees
Knows who's askingNo, or only what's in the promptYes — it gets to know each person, strictly inside their permissions
OversightWhatever was bolted onSensitivity tiers, oversharing detection, a full record of who saw what and why
Where it runsDepends on every componentYour infrastructure, fully offline, or managed — your choice
A fair word about established enterprise-search products

The best of them also connect broadly and respect permissions, and this brief doesn't pretend otherwise. Where they are already deployed and the question is document search for people, they do the job.

Where the conversation gets specific

Data beyond documents; a knowledge graph agents can reason over; running inside your own boundary or fully offline; refusing to answer rather than guessing; seeing oversharing rather than just avoiding it; and agents that can analyze and act — under approval. We're glad to go through any of them side by side.

Permission before the search A digital twin, not a heap of text Data stays where it lives Sourced, or it says no Agents as real users

Comparison current as of Q4 2026 · the left column is the category approach, not a named vendor

08Who it's for Optional

Every organization being asked to put agents to work on what it knows.

Especially the ones with intricate access rules, answers that have to be trustworthy, and data that can't leave their control.

Banking & insurance

Answers a regulator can follow

Every answer cited to the passage, the table and the rows, as of when; permissions checked before the search; one exportable record of who saw what — inside the boundary.

Government & defense

Fully offline, with sensitivity tiers

Runs in a disconnected environment; the most protected material is never stored as text; agents have identities of their own with exactly the sources and actions they are granted.

Enterprises with data gravity

The warehouse stays where it is

Snowflake or Databricks read through metadata and semantic views; SharePoint, Slack, documents and Excel processed alongside. Nothing migrated, nothing replaced.

AI platform teams

The grounding layer for agents

A knowledge graph agents reason over, a standard interface they connect through, and an action loop — preview, check, act, confirm — that makes handing them work safe.

Operations-heavy enterprises

One place to ask

Contracts, tickets, drawings, field reports and the systems behind them, connected — so a person or an agent can follow a thread from a customer to their tickets to the metrics behind them.

Service providers & integrators

Delivered as a managed service

Run Data Foundry for your customers on sovereign or dedicated infrastructure, with each customer's knowledge and permissions never leaving their control.

09Go deeper & next steps

The platform around the model.

This brief is the reference for Bud Data Foundry. For the platform-level argument — why data, serving, agents and governance belong on one plane — read the whitepaper, or return to the product overview.

Get started with Bud

Put your data on it.

The fastest way to see what a living model of your organization does for your people and your agents is a proof-of-concept on your infrastructure, with your data and your permissions.

01 Connect three sources — one with documents, one with tables, one that streams — with their permissions.
02 Ask the same question as two people with different access, and read both answers.
03 Hand one workflow to an agent and watch it preview, check, act and confirm.