Home/ Bud Gaia/ Product Brief
Bud Gaia overview
Product Brief · Bud Gaia · PAIOS

Bud Gaia

The Personal AI Operating System — six layers on one machine, running many models and agents at once on the hardware you own. Gaia is the device-side version of Bud Novaria: the same architecture, scaled from the datacenter to a desk, a shelf, or a vehicle.

Product reference v0.1 · pre-release July 2026 ~12 min read
Status: announced, not shipping. Gaia is on a public waitlist. Figures in this brief are hardware facts, published third-party numbers, or Bud design targets — each is labelled as such in §06. Nothing here is a measured production result yet.
01At a glance

An operating system for the box, not another model runner.

Desk-side machines now ship datacenter-class AI compute with a large unified memory pool. The software on them still runs one model at a time, by hand. Gaia is the missing layer: many models and agents resident together, sub-second switching, automatic model choice, and your own work always first in line.

Layers, one install6
Model wake · target~1s
Idle footprint0GB
Host OSes3
design targets · basis in §06
What it is
A personal AI operating system: it installs on top of Windows, Ubuntu or macOS and takes over scheduling of models, skills and agents on the machine
A multi-model runtime — several base brains resident at once, each carrying megabyte-scale skills, blended per request
A discovery layer that chooses and sizes models for the specific chip in front of it, before anything is downloaded
A hub: one box serving every laptop, phone, till and camera on the local network, with no GPU on the client
The device-side version of Bud Novaria — same architecture, personal scale
What it is not
A model runner or chat app — those load one model at a time and leave orchestration to you
A replacement for your operating system — Windows, Ubuntu and macOS stay in charge of the desktop
A cloud service with a local cache — the default path is on-box, and the frontier is opt-in per request
A tax on your machine — Gaia holds nothing when idle and yields to your workload first
The enterprise platform — that's Bud Novaria, the datacenter tier
02Where it fits

Six layers. One box. One install.

Gaia sits between the host operating system and the AI applications above it — the layer that today's local-AI stack is missing. Each layer is addressable on its own; together they turn a single machine into an AI department.

06ExperienceBud UI · App Store · SDKinstall
05AgentsAgent Runtime · ARTdo
04AccessGateway — one API, local and frontierroute
03DiscoveryModel Finder · Bud Simulatorchoose
02IntelligenceOrchestrator · Serverless · Engine Backendserve
01MetalVirtualization (FCSP) · Layer Zerohardware

Above the stack: whole AI applications — agents plus a real interface — and the people and devices around the box: laptops, phones, front desks, cameras, vehicles.

Below the stack: your host OS and your silicon — GB10-class desk-side systems, RTX workstations, Core Ultra and Ryzen AI laptops, Apple silicon. Layer 01 docks them without application changes.

03Capabilities, in full

Six capabilities that need an OS to exist.

What the overview page states, expanded to the specifics an evaluator needs. Each one requires control of the whole machine — none of them can be a feature of a single application.

01Many models resident, instant switchingSeveral base brains stay in unified memory together — language, reasoning, speech, vision — while cold-start acceleration wakes anything else in about a second · no evictions, no manual load-kill-reload cycle, no restart to try a second model · multi-model applications become ordinary software instead of an ops project~1 s wake · target
02Skills: depth without another downloadA brain is gigabytes; a skill is megabytes · several adapters can be live on one brain at once and are blended per request — email tone, ticket triage, sales conversations, a language, your document layouts · overnight, agentic reinforcement training (ART) distils your own agent runs into new skills, scored by a judge before they attachMB-scale skills
03Automatic model discovery and fitApplications declare an outcome, never a filename · Model Finder ranks open models against real jobs, languages, licences and size · Bud Simulator predicts memory and speed for each candidate on your exact chip, beside everything already running, before a byte is downloaded · the same app resolves to a different model on a GB10, a laptop and a Macfit proven pre-download
04Your work always outranks the AIScale to zero: no request, no model in memory — a box full of installed AI apps idles like a box with none · pre-emption: your game, render, build or call reclaims memory and compute immediately, and Gaia steps back first · opportunism: agents and overnight training use the gaps — lunch, evenings, the small hours0 GB idle
05Private by default, frontier by choiceOne standard API fronts local models and frontier clouds; Gaia routes per request by privacy class, difficulty and cost, and reports what ran where · sensitive work is answered on-box or refused · spend caps make a runaway bill structurally impossible · swapping the model behind a call needs no code changeone API
06One box, a whole buildingEvery client on the local network — laptop, phone, tablet, till, front desk, camera, drone — gets the same apps, agents and models with no accelerator of its own · agents live on the hub, so they keep working through closed laptops and weekends · requests stay on the network unless a frontier call is explicitly allowedno GPU on clients
04How it works

One request, end to end — on one machine.

When a person or an agent asks Gaia for something, the request crosses all six layers on the same box, in one scheduling domain. There is no network boundary to pay for and nothing to serialize between steps.

The request lifecycle — eight steps, one machine

01AskAn app or agent states the job and its privacy class
02PermitGateway checks the app's granted scopes; anything more asks you
03RouteOn-box, or frontier by explicit allowance and budget
04ResolveDiscovery maps the job to a brain plus the skills it needs
05AdmitOrchestrator checks the promise fits current capacity
06ServeEngine backend runs it on the partition FCSP granted
07ReleaseIdle models scale to zero; your workload pre-empts first
08LearnThe run is scored; good ones become tonight's skill

Brains and skills — the memory model

Resident unitScaleCoversExample skills attached
Fast brain3B · ~1.9 GBRoutine language work, classification, triageemail tone (14 MB) · ticket triage (11 MB) · routing (9 MB)
Deep brain8B · ~4.6 GBReasoning, long-form drafting, domain worksales conversations (22 MB) · financial analysis (26 MB) · Malayalam (31 MB)
Sensesspeech + vision · ~1.2 GBTranscription, speech, documents, camerayour accent (7 MB) · your document layouts (12 MB) · meeting voices (9 MB)
Frontier (optional)off-boxThe hardest public tasks, on requestrouted via Gateway, inside a spend cap

Sizes are illustrative of the shape, not a fixed bundle: breadth comes from several brains, depth from many megabyte-scale skills — and skills trained on your own runs stay yours.

Agents are declared, not scripted

agent: morning-brief needs: llm: small + finance skill voice: any natural TTS runs: daily at 07:00 promise: ready by 07:05 may: read mail, read calendar # anything more asks you first
Portable by construction

The file names needs, not models. The same agent runs on a GB10, a laptop and a Mac; discovery resolves the best available fit on each.

Accountable

The promise is a scheduling input. The orchestrator admits work it can keep and surfaces what it cannot, rather than silently missing it.

Scoped like a phone app

Permissions are explicit and revocable; anything outside the grant asks before acting.

Discovery, step by step

CandidateGood at the job?Fits the box?Around a second?Verdict
13B generalistDecent26.8 GB — won't fitRejected pre-download
3B chat modelWeak on the domain6.1 GB~0.3 sKept as fallback
7B domain-tuned modelBest in class14.2 GB~0.6 s✓ Selected

Illustrative Bud Simulator output. The point is the order of operations: predict fit on this chip, then download — never the reverse.

05Hardware & deployment shapes

One box, four shapes.

Gaia is the same install in every shape below; what changes is who the box serves. Nothing here requires a datacenter, a cloud account, or an internet connection.

1 · Personal

A workstation or laptop serving one person: agents in the gaps, private drafting, local research — and the machine still feels like yours.

2 · The office hub

One box on a shelf serving every device in a clinic, café, studio, law firm or workshop. No accelerator on the clients, no per-seat meter.

3 · Edge site

A production line, remote school, ship or field office — offline by default, low-latency, sovereign, with frontier access optional.

4 · Physical AI

Robots, drones and vehicles that need on-site reflexes: local brains, hard promises, fault isolation between models.

ShapeWho it servesNetworkFrontier callsNotes
Personal1 personlocal onlyOpt-in per request✓ Full stack, host OS unchanged
Office hub5–50 devicesLANOpt-in, budgeted✓ Clients need no accelerator
Edge sitesite devicesLAN, intermittent WANUsually off✓ Offline-native
Physical AIone machineon-boardOff✓ Fault isolation per model

Target hardware

The 2025–26 generation of desk-side AI machines: roughly a petaFLOP of AI compute and a large unified memory pool, at a one-time $3–4k.

NVIDIA GB10 · DGX SparkRTX AI workstationIntel Core UltraAMD Ryzen AI MaxApple M-series

Host operating systems

Gaia installs on top of the OS you already run. It schedules AI work; your desktop stays in charge of the desktop.

WindowsUbuntumacOS

APIs & standards

One OpenAI-compatible endpoint fronts local models and frontier providers alike, so existing clients point at Gaia by swapping a base URL. Agents and skills are declarative files; the SDK declares needs, never filenames.

06Claims & methodology

Every number, with its basis and its status.

Gaia is pre-release, so this section is deliberately conservative: what is a hardware fact, what is a published third-party figure, and what is a Bud design target we intend to be held to.

~1petaFLOP
Desk-side AI compute · hardware fact
BasisVendor specifications for the 2025–26 generation of personal AI systems (GB10-class desk-side machines and comparable workstation silicon), paired with ~128 GB of unified memory at a one-time $3–4k. Not a Bud measurement — this is the hardware Gaia targets.
~1s
Model wake from cold · design target
BasisGaia's cold-start acceleration and FCSP partitioning are being engineered to a sub-second wake target on supported hardware — roughly the feel of opening an app on a phone. Range, not a guarantee: actual wake depends on model size, chip and what else is resident. Measured figures will publish with the first release.
0GB
Idle memory footprint · design invariant
BasisScale-to-zero is an architectural property, not a tuning result: with no request in flight, no model is resident, so a box with a full app roster idles like a box with none. The corresponding invariant is pre-emption — your workload reclaims memory and compute first, always.
MBvs GB
Skill vs brain · architectural ratio
BasisSkills are adapters in the tens of megabytes, attached to base brains of gigabytes — the ratio is what makes depth affordable on a personal box, and what lets several specialisms be live at once. Overnight ART distils new skills from your own scored agent runs.
ClaimStatusHow it will be measured
Sub-second model switchingdesign targetTime-to-first-token from cold, per model class, on each reference machine, with a stated resident set
Several brains resident togetherdesign targetConcurrent resident models and blended skills within a fixed unified-memory budget
Zero idle footprintdesign invariantResident bytes and accelerator utilisation with no request in flight
Pre-emption in favour of the userdesign invariantTime for a foreground workload to reclaim memory and compute after it starts
Fit predicted before downloaddesign targetBud Simulator predicted vs observed memory and latency, per candidate and chip
Skills improve week over weekdesign targetJudge-scored task win-rate for ART-distilled skills against the base brain, per workflow
Same architecture as the enterprise tierstructuralShared lineage with Bud Novaria's Layer Zero, virtualization, gateway and agent runtime
07Positioning

Against the obvious alternatives.

Nothing available today does the OS job on a personal machine. Each alternative solves one part — and leaves the scheduling, discovery and sharing problems to the user.

AlternativeWhat it gives youWhat Gaia adds
Local model runnersOne model, running offline, quicklyMany models and skills resident together, automatic choice, sub-second switching, fair sharing of the box
Frontier subscriptionsThe best models, per seat, per tokenEveryday work answered on-box at the price of electricity; frontier as an explicit, budgeted choice
Rented cloud GPUsElastic capacity for burstsHardware you own, data that stays in the building, and no meter running on routine work
DIY orchestration on a workstationFull control, if you have the teamDiscovery, admission control, pre-emption and fault isolation as OS services rather than a private project
Bud NovariaThe enterprise AI operating system, in the datacenterThe same architecture at personal scale — one machine, one person or one building, no control plane to run
Six layers, one install Alt-Tab for models Zero idle footprint Model choice automated Private by default One box, whole building
08Who it's for

Where a personal AI OS lands hardest.

The pattern repeats: a small organisation with real work, private data, and no appetite for a per-seat AI bill — or a machine that has to think where the cloud cannot reach.

Small business

An AI staff for the price of power

A café, clinic or law firm runs sales, scheduling, intake and follow-up on one box — no per-seat subscription, no per-token meter, nothing metered as the team grows.

Regulated & sensitive work

Files that must not leave

Client files, patient records and financials are processed in the building. Frontier calls are a deliberate, budgeted exception rather than the default path.

Developers & prosumers

Ship an app, not an install guide

Declare the models and skills your app needs; Gaia resolves them per machine. One binary works on a GB10, a Core Ultra laptop and a Mac.

Edge sites

Where the cloud does not reach

Factory lines, remote schools and clinics, ships and field offices: offline-native, low-latency, sovereign — with the same apps as head office.

OEM & silicon partners

Software that earns the box

An AI workstation is only as valuable as what runs on it. Gaia is the layer that turns a spec sheet into applications a buyer keeps using on day thirty.

Physical AI

On-site reflexes

Robots, drones and vehicles get local brains with hard promises and fault isolation, so one overloaded model cannot take the machine down.

09Go deeper & next steps

Both tiers of the same architecture.

This brief is the Gaia reference. The overview page carries the story and the stack explorer; Novaria is the enterprise tier of the same architecture.

Get started with Bud

Put your data on it.

The fastest way to see what an integrated AI operating system does for your enterprise is a proof-of-concept on your infrastructure, with your data.

01 Identify a use case where complexity, cost, or governance is a known pain point.
02 Joint discovery — Bud maps your AI pain points to platform capabilities.
03 POC in days, on your hardware, with your data.