Size the cluster before you buy it.
A serving deployment fixes six things at once — chip, count, engine, parallelism split, quantization and batching. Bud Simulator predicts the whole space, discards everything that cannot meet your SLO, and returns the cheapest survivor. It answers the same question for training runs, and models CPUs as carefully as GPUs. Open source.
Sizing is a search problem, not a spreadsheet.
vLLM alone carries around 400 million configuration permutations for a single node — and the chip is a variable too, across 72 hardware profiles spanning GPUs, TPUs, ASICs and CPUs. Nobody benchmarks their way through that. Bud Simulator predicts it instead: a closed-form roofline scores every candidate, learned regressors correct the residual, and an event-driven serving simulator settles p99 on the finalists.
- Beat one — where it sits. Bud Simulator is the sizing engine inside Bud AI Foundry, Layer 04 of the eight-layer Bud stack. It fixes six things that are only correct together: chip, count, engine, parallelism split, quantization and batching. Three predictors are stacked by cost and fidelity — an analytical roofline, learned regressors, an event-driven serving simulator. Feasibility is gated before cost. The output is the deployment: the runtime applies the configuration it returns.
- Beat two — one prediction pipeline. A candidate configuration passes a memory fit check (weights plus KV cache against aggregate device memory), an analytical roofline (prefill bound by compute, decode by memory bandwidth), learned regressors fitted to measured runs, an event-driven serving simulator covering queueing and continuous batching, and a compatibility check across model, hardware and engine. It comes out sized — the cheapest configuration that meets the SLO.
- Beat three — the collapse. Roughly thirty billion candidate configurations narrow under one query of use case plus SLO to three finalists on the Pareto frontier, and then to one plan that is deployed. All 72 hardware profiles resolve to a single sized deployment. Cost per million tokens improves six-fold against a single-tier GPU deployment at equal service level. Feasibility is gated first, so cost is only ever optimised among plans that already meet the objective.
- Beat four — the result. Guesswork becomes a search result.
Four hundred million configurations become one plan.
The reason enterprises staff a GenAI systems engineering team is largely that somebody has to guess well, because nobody can measure exhaustively. A benchmark answers for one configuration on hardware you already own. Bud Simulator answers for the whole space, including the chips you are still being quoted — and it answers before the purchase order, not after.
Nine reasons the answer holds.
A sizing tool is only worth the confidence you can place in it. These nine are what separate a prediction from a guess with better typography — across serving, training, and the silicon underneath both.
Three predictors, one funnel
A closed-form roofline scores every candidate at effectively zero cost, learned regressors correct the residual on hardware with measured history, and an event-driven serving simulator settles p99 on the finalists.
Feasibility before cost
Weights plus KV cache at the target concurrency and context must fit in aggregate device memory. That single inequality yields the minimum chip count and discards most of the space before anything expensive runs.
A frontier, not a weighted score
Multi-objective genetic search over cost, latency and throughput returns the Pareto frontier — every configuration you cannot improve on one axis without giving up another. No weights you do not actually know.
The SLO is an input
Chat, code completion, summarisation and agentic traffic want different batch sizes, different cache blocks and often different chips. A pre-built profile carries that shape into the search instead of one latency number.
Hardware you do not own yet
The most valuable sizing question is asked before purchase. Every accelerator reduces to four numbers — compute, memory, bandwidth, interconnect — so a chip that is quoted but not racked can still be compared honestly.
Training, not just serving
14 training stages from SFT through DPO, PPO, GRPO and KTO; six fine-tuning methods; 30+ optimizers; TP, PP, DP, EP and ZeRO 0–3. Memory is the peak of the forward, backward and optimizer phases — not their sum.
CPUs modelled properly
Not a fallback with a fudge factor — an operator framework. AMX, AVX-512, AVX2, SVE and NEON throughput, L1–L3 cache and KV placement, NUMA topology, threading and frequency scaling under thermal limits.
BudEvolve — the search, inverted
Hold the workload fixed and sweep the hardware instead: FLOPS, bandwidth, memory and interconnect as free variables. Plus Morris sensitivity ranking, what-if curves, and LLM-driven evolution of scheduling and cache-eviction algorithms.
Open source, and auditable
Every efficiency constant is sourced from a datasheet, ISA spec, published benchmark or microbenchmark — no constants fitted until the answer looked right. 900+ tests; validated against MLPerf Training, DeepSpeed ZeRO and Megatron-LM.
The full story, in depth.
The roofline, the search, the calibration — and the code behind all three.
Repository
BudEcosystem/simulator
The engine, in the open
Decoder-only, encoder-only and diffusion model simulation for SLO, memory and infrastructure calculations — for both inference and training.
View on GitHubEngineering
Inside Bud Simulator
Sizing before you buy the hardware
Why prefill and decode sit on opposite sides of the roofline, how the Pareto frontier is searched, what the calibration is worth — and where the answer is still only a prior.
Read the deep diveProduct Page
Bud AI Foundry
Deployment & serving · Layer 04
The simulator sizes; the Foundry serves. One control panel for inference, guardrails, observability, evaluations and FinOps — with the configuration applied automatically.
Explore Bud AI FoundryPut your data on it.
The fastest way to see what an integrated AI operating system does for your enterprise is a proof-of-concept on your infrastructure, with your data.