This turn

One click · stays

Local chat activated. Pick a thread.

Write in your project…Send

Weekly research brief

DeepSeek V4 Flash · files stay attached

Ready

Health

Idle

Runtime

00:00:00

  1. Pulled notes on your PC
  2. Sorted material on your PC
  3. Draft waiting for you

Writes stay gated until you allow them.

Launch plan

Several bots · your machine

Plus
  1. Hermes · briefResearch · DeepSeek V4 FlashReady
  2. SourcesGather · Qwen3 32BQueued
  3. Campaign draftCompose · Gemma 3 27BIdle

Handoff carries accepted work. Writes stay gated.

DeepSeek V4 Flash

DeepSeek-V4-Flash-Q4_K_M.gguf

Q4_K_M
Context
33k
Temperature
0.20
GPU layers
28
System RAM · DDR5
74%
GPU VRAM
36%

Launch plan · your PC

  1. Prompt

    Prepare a launch plan. Keep private files on your machine.

  2. Context

    brief.md · prior-decisions.md · memory

  3. Response

    Draft organized from local sources.

prompt.json · context.json · thinking.json · response.json

Folders and tools

  • Launch plan /
  • brief.md
  • secrets /
  • Write files
  • MCP tools

Private Alpha · Limited cohorts

Your AI.Your data.Your models.

A local-first workspace. Run compatible local models on your computer. Connect a supported provider only when you choose to send a turn out.

Local chatCodex → Kimi K3 · On your PC

DeepSeek V4 Flash

Compatible local GGUF · chat · tools · your PC

  • This LLM is downloaded and installed.
  • Native local runtime connected. Exact process and model verified.
  • Full health check passes.

GGUF placement

System RAM · DDR5Weights can live here

GPU VRAMNot the only home

GGUF can store the model in system RAM, including DDR5, as well as VRAM. Placement follows measured headroom. Not a speed claim.

In Chat

LLM Configuration Panel

Save DefaultRestore DefaultsPreview payload
1

Max output tokens

Maximum generated tokens.

32,768
2

Max input tokens

Maximum prompt/context budget. Adaptive VRAM optimizer picks a safe cap.

48,000
3

Screen VRAM reserve [in Mb]

VRAM kept free for display and desktop use.

512
4

Enable thinking

Allow supported reasoning mode.

Thinking off
5

Temperature

Randomness and variety.

020.70
6

GGUF memory refresh

Apply the GGUF llama.cpp memory profile.

On
7

Speculative decoding

Local llama.cpp draft prediction. N-gram / MTP.

SimpleVerified activeN 12M 48

Local record

Prompts, context, responses, and thinking stay on your machine.

The workspace writes the turn to the project on your computer. Search and reopen it on your machine. Work stays with you, and it is not used as training data.

  • Prompts

    What you asked. Stored with the project on your computer.

  • Context

    Files, instructions, and selected memory attached to the turn.

  • Responses

    What came back. Kept on your machine, not in a provider chat history.

  • Thinking

    The model’s reasoning for the turn, when it produces one.

Your PC · Launch plan

  1. Prompt

    Prepare a launch plan. Keep private files on your computer.

  2. Context

    brief.md · prior-decisions.md · approved memory

  3. Thinking

    Reasoning for this turn stays with the project on your PC.

  4. Response

    Draft organized from local sources. Folders stayed on your PC.

Stored on your machine

Velociti / Launch plan /
prompt.json
context.json
thinking.json
response.json

On your computer

Kept with the project.

One click. Your keys.

Change the model. Keep the work on your computer.

Switch between a compatible local model and a supported provider without rebuilding the task. You bring the key.

  • Local GGUF models stay on your machine. Weights can live in system RAM, including DDR5, not only VRAM.
  • Supported providers connect through keys you already control.
  • The project, files, and instructions do not reload.

Prepare a launch plan. Keep private files on your computer.

Thread stays attached to Launch plan

FilesInstructionsApproved memoryTool gates
  1. Found relevant notes and previous decisions. Private folders stay on your computer.
  2. Local model · this step

    Organizing sources and checking the draft on your machine.

  3. A follow-up can continue without starting over.
Same project. Same files.Receipt · Local · no provider charge

Your keys · idle

  • Hugging Face
  • OpenRouter
  • Codex
  • Claude

Hermes · Plus

Start a Hermes Bot in one click.

Name the job, choose what it may use, and start. Hermes stands the bot up on the project already open, usually a local model, so it can keep going after you leave the chat, with last week’s brief still attached. Mount prompts and rules on a model such as Fable 5 only when you need this turn steered. That is an agent: it follows the documents you attached, then the context ends.

  • Create the bot from the project. No separate stack to assemble.
  • It keeps persistent memory on your computer so recurring work does not start from zero.
  • Writes stay gated. Pause keeps the bot’s state. The project does not reset.

Agent · this turn

Mount the brief. Run the turn.

Fable 5 · documents attached

  • voice.mdMounted
  • rules.mdMounted
  • scope.mdMounted

This turn

Attach voice, rules, and scope. Fable 5 follows them against the files you opened. Close the turn and that context is gone. The agent does not keep a diary.

Bot · one click

Weekly research brief

DeepSeek V4 Flash · local · persistent memory

Still on your project

  • Last week’s sources
  • Accepted outline
  • Your edits
  1. Gathered saved sources
  2. Local model sorted the material
  3. Draft ready for you

Start it once. It can run the brief again tomorrow with last week still attached.

Orchestration · Plus

Several bots. One project.

Coordinate Hermes Bots around a shared task. Each bot can carry its own memory. Handoffs pass accepted work, not an entire history, and not another bot’s private memory.

  • Research, gather, and draft stay attached to the same project.
  • A handoff carries accepted work, not an entire history.
  • Writes stay gated. Pause and recovery stay with the task.

Orchestration · Plus

Launch plan

Several bots · one project · your machine

Shared files stay attached
  1. Hermes · Weekly brief

    Research

    Passed accepted notes
  2. Sources

    Gather

    Allowed this turn
  3. Campaign draft

    Compose

    Review ready

A handoff carries accepted work. A full transcript is not dumped into the next bot. Writes stay gated.

Boundaries

Files have a fence. Tools ask first.

Choose which folders an agent may see. Connection is not permission. A tool still stops for current-turn approval when policy requires it.

  • Allow, ask, or deny folders on your computer.
  • MCP and write tools wait for you.
  • A saved connection cannot bypass a turn gate.

Folders on your PC

  • Launch plan /
  • brief.md
  • secrets /
  • prior-decisions.md

Tools

  • Read files
  • Write files
  • MCP tools

Connection is not permission.

MCP tool gate requiring current-turn approval before a scoped tool can run.

Scope tool gate

This request needs current-turn review. Saved authority cannot bypass it.

The tool asked to evaluate an expression. Do not treat this form as a place to paste secrets.

Waiting for a person.

Workspaces · Plus

Share the project, not the account.

Plus brings people into one workspace with shared files, conversations, agents, and accepted results. Each person still uses keys they control.

  • One project. Several people.
  • Shared files and instructions stay attached.
  • Personal provider keys are not the shared object.

Shared with the project

  • brief.md · attached
  • Launch thread · three people
  • Hermes · Weekly research brief

Keys stay with each person. The project is what you share.

One workspace

More than a chat window.

The product brief names a workbench: editor, terminal, browser, files, diagnostics, extensions, chats, and agents together. History stays searchable. Only useful context enters the next turn.

Your real workbench

Editor, terminal, browser, files, diagnostics, extensions, chats, and agents sit in one place. The project is what they share.

Selective memory

Polar is the local vector database for project context. Workspace retrieval, approved procedural memory, and chat memory stay separate. Relevant file segments are selected. The entire workspace is not injected.

Token use you can see

Receipts can show model, provider, input, output, cache, and reasoning tokens. Tool definitions enter only when needed. Irrelevant tools are pruned from each turn.

CDP, not screen scraping

A local Chrome DevTools Protocol channel inspects visible IDE actions so work can be verified without relying on screen scraping.

Right intelligence per step

Keep private preparation local. Use a supported provider when harder judgment earns the added cost. The task decides the route. The project remains.

Download and self-install

Velociti Free

A better place to use AI.

  • Local AI tuned for your machineInstall in one click. Velociti profiles your hardware, tunes model fit and runtime settings, and adds it to your picker. GGUF weights can sit in system RAM, including DDR5, as well as VRAM.
  • Local recordPrompts, attached context, responses, and thinking stay on your computer. Reopen them without a provider-owned chat history.
  • Your models and your keysUse compatible local models and connect supported providers through keys you already control.
  • One-click model switchingChange intelligence without rebuilding the task.
  • Your real workbenchWork with an editor, terminal, browser, files, diagnostics, extensions, chats, and agents together.

For ongoing and coordinated work

Velociti Plus

AI that keeps working with you and for you.

  • Shared workspaces and projectsBring people into a common workspace with shared files, project context, conversations, agents, and accepted results.
  • One-click Hermes Bot setupGive a bot a task, choose what it may use, and start without assembling a separate automation stack.
  • Hermes monitoring and schedulesRun work while you are away. See progress, model use, tokens, results, and anything needing attention.
  • Project continuityCarry relevant state across chats, people, and agents while the work remains attached to its project.
  • Cross-agent orchestrationCoordinate research, drafting, review, and checking around one shared task.

Free is complete. Plus keeps going.

Compare Free and Plus

Velociti Enterprise

AI capability, in your office.

Enterprise is a separate organizational product. It joins company workspaces, private intelligence, governed agents, institutional memory, and customer-owned compute.

  • Customer-owned compute and appliance options
  • Sentinel authority and policy gates
  • Governed Hermes agents
  • IAM, audit evidence, recovery, and lifecycle support
Talk to Velociti Enterprise
Velociti One appliance: an orange chassis with a stainless front plate and visible copper cooling.
Velociti One.

Private Alpha

Join a limited cohort.

Velociti is admitting small cohorts so we can observe setup, fix activation problems, and expand responsibly.

Required: email. Everything else is optional and helps us match a cohort.

For company deployments, talk to Velociti Enterprise.

Talk to Velociti Enterprise

We’ll match you to a cohort as space opens.

FAQ

Straightforward answers.

What is the private Alpha?

A limited, observed rollout. Velociti is admitting small cohorts so we can watch setup, fix activation problems, and expand responsibly.

When will I receive access?

We’ll match you to a cohort as space opens.

What hardware do I need?

Local-model tuning is hardware-aware. GGUF weights can be placed in system RAM, including DDR5, as well as GPU VRAM, so a machine is not limited to video memory alone. Suitable models and settings still depend on measured headroom. We may ask about operating system and GPU so a cohort can be matched honestly.

What is Velociti?

A local-first AI workspace and operating layer. It brings compatible local models, supported providers, project memory, files, and agents together so changing intelligence does not mean starting over. See how Velociti works.

Which models and providers does Velociti support?

Velociti runs compatible local models and connects supported providers with keys you control. The public list is on models Velociti works with.

Where does my work live?

Prompts, context, responses, and thinking stay on your computer with the project. Search and reopen them without a provider-owned chat account. Polar retrieves relevant file segments without injecting the entire workspace. Work is not used as training data. A supported provider is used only when you send a turn out.

What is the difference between Free and Plus?

Free is a complete personal workspace: local models, supported providers, projects, history, receipts, and approved tools. Plus includes Free and adds shared workspaces, Hermes Bots, schedules, orchestration, and recovery. Read Velociti Free vs Plus and Free and Plus.

Is Enterprise just a higher Plus tier?

Enterprise is a separate organizational product with customer-owned compute, company workspaces, Sentinel authority, governed agents, audit evidence, and expert services. See Velociti Enterprise.

What is the workbench?

Velociti holds an editor, terminal, browser, files, diagnostics, extensions, chats, and agents together. The project is the shared object, not a provider chat account.

What is the CDP listener?

A local Chrome DevTools Protocol channel. It lets Velociti inspect visible IDE actions and verify work without relying on screen scraping.

What is the difference between an agent and a bot?

Start a Hermes Bot when the work should continue: one click from the project, usually on a local model, with memory that lasts beyond the chat. Use an agent when you only need this turn steered. Mount prompts and rules on a model such as Fable 5, and it works from that context alone. The agent does not keep a long memory. The bot does. Continuity lives in Plus.

Can a Hermes Bot run on a schedule?

Plus adds schedules and monitoring so a bot can continue while you are away. You can see progress, model use, tokens, results, and anything that needs attention. Pause, stop, restart, and continue from saved state.

How is memory handled?

Polar is the local vector database for project context. Workspace retrieval selects relevant file segments and references without injecting the entire workspace. Project retrieval, approved procedural memory, and chat memory stay separate. A full transcript is not handed to the next model.

Do you publish prices here?

Free is listed on this site. Plus pricing appears when it is public. Enterprise is qualified with you.

Your AI.Your data.Your models.

Apply for Alpha if you want a local-first workspace. Compatible local models run on your machine. Supported providers enter only when you send a turn out.

Velociti · Local-first AI workspace