Single GPU · Host included
RTX 3060 12GB
Used purchase estimate
4.4 GiB of spare memory
The lowest priced match. Enough memory and estimated speed for this workload, with more of your budget left over.
Explore the trade-offs. Adjust your budget and speed target to find your sweet spot.
Every point is a setup. Select one to explore its cost, speed and memory.
↑ Faster · ← Lower cost
Central estimates, not guaranteed speeds. Select a point for its range.
21 systems have missing or unplottable values for these axes.
Use Browse every setup below to explore the full hardware library.
Over-budget setups stay visible so you can explore the next step. GPU prices include the configured host PC. Used prices where available; otherwise new.
YOUR HARDWARE SHORTLIST
Every pick fits your model, meets 30 tokens/s per user, and stays within budget.
Single GPU · Host included
Used purchase estimate
4.4 GiB of spare memory
The lowest priced match. Enough memory and estimated speed for this workload, with more of your budget left over.
Complete computer · Unified memory
Used purchase estimate
41.3 GiB of spare memory
An estimated 81 W during inference. About 15 SEK/month at 4 hours a day. A useful option when ongoing energy use matters.
Single GPU · Host included
Used purchase estimate
16.4 GiB of spare memory
Around 4.1× your speed target at the central estimate. A performance-focused option for more responsive generation.
Watch the same answer arrive at your target speed and on a system that fits your workload.
About 4.2 s for this ~125-token response.
Write a short passage about a neighborhood garden.
Press play to see this response take shape.
Requested pace 30 tok/s
About 1.0 s for this ~125-token response · 32 500 SEK
Write a short passage about a neighborhood garden.
Press play to see this response take shape.
Central estimate 122.2 tok/s
Compare local language-model hardware by memory fit, estimated generation speed and price. GPU VRAM, Apple-style unified memory and CPU offload behave differently, so the same model and settings can produce very different trade-offs. Adjust the workload and budget to explore a practical shortlist.
It estimates whether a selected model, quantization and context can run on each listed system, then compares a range of per-user generation speeds. It also shows hardware prices where available. These are planning estimates based on hardware and software assumptions, not measured benchmarks or a promise that a particular setup will perform the same in your environment.
No. A discrete GPU usually has its own VRAM, while unified-memory systems let the processor and GPU share one memory pool. Ordinary system RAM can also hold model data, but moving it to a GPU over PCIe can limit speed. The simulator treats those memory paths differently; installed RAM plus VRAM should not be read as one equally fast pool.
Generation speed varies with backend, drivers, batch size, model implementation and other workload details. A range communicates that uncertainty more honestly than a single precise-looking number. The estimate describes decoding tokens per second per person; prompt processing and time to first token are separate parts of an end-to-end response. Check the assumptions and available measurements before buying hardware.