Guide
Run local models on your computer
Velociti installs a compatible local runtime, measures hardware, and adds the model to your picker. GGUF weights can be placed in system RAM, including DDR5, as well as GPU VRAM. A machine is not limited to video memory alone.
Measure first
Local-model fit depends on measured headroom, not a marketing VRAM number. Velociti profiles the machine, then tunes runtime settings before the model lands in the picker.
Free includes this install path. You do not assemble a separate stack to try a local model.
RAM and VRAM
GGUF weights can sit in system RAM, including DDR5, as well as GPU VRAM. Use VRAM when it is available and large enough. Use system RAM when it is the honest fit.
Suitable models still depend on what the machine actually has. Alpha cohorts may ask about operating system and GPU so matching stays honest.
Hardware classes
A compact laptop with modest VRAM can still run a smaller local model in system RAM for private drafting. A desktop with a larger GPU can keep a heavier GGUF in VRAM for faster local turns.
When the local pass is not enough, send that turn through a supported provider with a key you control. The project stays attached.
Not only a runner
LM Studio and Ollama load models. Velociti keeps the workspace around them: files, instructions, approved memory, receipts, and one-click switches. Read Velociti vs LM Studio and Velociti vs Ollama.
Questions
- Do I need a datacenter GPU?
No. Placement follows measured headroom. Smaller models can live in system RAM.
- Which models can I run locally?
Compatible GGUF models Velociti can install and tune. Provider models such as ChatGPT and Claude enter through keys. The models list names both.
