Local-first workspace

Your AI.Your data.Your models.

Run compatible local models on your computer. Connect supported providers with keys you already control. Organize work into projects so files, instructions, and approved memory stay on your machine unless you send a turn out.

Local record

Prompts, context, responses, and thinking stay on your machine.

The workspace writes the turn to the project on your computer. Search and reopen it on your machine. Work stays with you, and it is not used as training data.

  • Prompts

    What you asked. Stored with the project on your computer.

  • Context

    Files, instructions, and selected memory attached to the turn.

  • Responses

    What came back. Kept on your machine, not in a provider chat history.

  • Thinking

    The model’s reasoning for the turn, when it produces one.

Your PC · Launch plan

  1. Prompt

    Prepare a launch plan. Keep private files on your computer.

  2. Context

    brief.md · prior-decisions.md · approved memory

  3. Thinking

    Reasoning for this turn stays with the project on your PC.

  4. Response

    Draft organized from local sources. Folders stayed on your PC.

Stored on your machine

Velociti / Launch plan /
prompt.json
context.json
thinking.json
response.json

On your computer

Kept with the project.

What changes: AI work gains a durable home. Projects keep related files, conversations, instructions, decisions, and results together. Approved context can strengthen each new chat or Bot without surrendering your history to one provider.

Prepare a launch plan. Keep private files on your computer.

Thread stays attached to Launch plan

FilesInstructionsApproved memoryTool gates
  1. Found relevant notes and previous decisions. Private folders stay on your computer.
  2. Local model · this step

    Organizing sources and checking the draft on your machine.

  3. A follow-up can continue without starting over.
Same project. Same files.Receipt · Local · no provider charge

Your keys · idle

  • Hugging Face
  • OpenRouter
  • Codex
  • Claude

DeepSeek V4 Flash

Compatible local GGUF · chat · tools · your PC

  • This LLM is downloaded and installed.
  • Native local runtime connected. Exact process and model verified.
  • Full health check passes.

GGUF placement

System RAM · DDR5Weights can live here

GPU VRAMNot the only home

GGUF can store the model in system RAM, including DDR5, as well as VRAM. Placement follows measured headroom. Not a speed claim.

In Chat

LLM Configuration Panel

Save DefaultRestore DefaultsPreview payload
1

Max output tokens

Maximum generated tokens.

32,768
2

Max input tokens

Maximum prompt/context budget. Adaptive VRAM optimizer picks a safe cap.

48,000
3

Screen VRAM reserve [in Mb]

VRAM kept free for display and desktop use.

512
4

Enable thinking

Allow supported reasoning mode.

Thinking off
5

Temperature

Randomness and variety.

020.70
6

GGUF memory refresh

Apply the GGUF llama.cpp memory profile.

On
7

Speculative decoding

Local llama.cpp draft prediction. N-gram / MTP.

SimpleVerified activeN 12M 48

Projects and workspaces

Each body of work has its own files, chats, instructions, context, and results. People and agents continue the same work without rebuilding its background.

The real workbench

Work with an editor, terminal, browser, files, diagnostics, extensions, chats, and agents together. The workspace is the place the work lives, not a separate chat tab.

Compatible local models

Velociti installs a compatible local runtime, measures hardware, and adds the model to your picker. GGUF weights can be placed in system RAM, including DDR5, as well as GPU VRAM. Placement follows measured headroom.

Supported providers

Connect supported providers through keys you already control.

Selective memory

Polar is the local vector database for project context. Workspace retrieval selects relevant file segments and references without injecting the entire workspace. Chat memory and procedural memory stay separate.

Visible receipts

Per-message receipts can show model, provider, input, output, cache, and reasoning tokens while you work.

Approved action

A local Chrome DevTools Protocol channel lets Velociti inspect visible IDE actions and verify work without relying on screen scraping. Agents receive scoped, user-approved tool access.

From chat to continuous work

Chats become work that can continue.

Work directly in chat or create a Hermes Bot in Plus. Velociti carries selected context, model choice, agent state, token use, and accepted results from one step to the next.

Models and Bots gain useful context from accepted work. Responses become more relevant, repeated prompting falls, and tailored intelligence survives model changes.

Carried forward

Context
Selected project files and approved memory
Model
Local or a supported provider for this step
State
Attached to the project
Tokens
Visible on this turn
Result
Kept for the next chat or bot

Step 1 of 6: Start in a project. The payload does not reset.

Spend intelligence where it matters

Local handles routine. Frontier handles exceptions.

A frontier model can supervise or review work prepared by a local model. Hard judgment receives stronger intelligence. Private and repeatable steps stay local when appropriate.

  • Selective memory injects only relevant history.
  • Tool definitions enter only when needed.
  • Irrelevant tools are pruned from each turn.
  • Live token counts show input, output, cache, and reasoning use.
Velociti Models screen showing a local Qwen model installed, validated, and connected to a native local runtime.
A compatible local model, installed and ready in the picker.
LLM Configuration Panel with hardware-aware token, VRAM, RAM, and speculative-decoding settings.
Local-model tuning follows your machine.
How Velociti works · Local models, your keys, the project stays