Overview
vg serve starts Vibgrate AI Context — a local-first MCP (Model Context Protocol) server that gives your AI assistant access to your code map, offline drift score, and version-correct library docs.
All tools are read-only. Your source is not uploaded. No account required. Fully offline under --local; when local docs for a library are thin, thin catalog fall-through to the hosted library may apply unless you pass --local.
The map stays fresh while you edit: each tool call runs a cheap freshness check and rebuilds incrementally when files really changed. There is no separate filesystem watcher.
Usage
vg serve # MCP over stdio — what assistants spawn
vg serve --http # MCP over streamable HTTP on 127.0.0.1:7437
vg serve --local # never touch the network
vg serve --compress # also compress context on 127.0.0.1:8787
vg serve --compress-only # compression only, no code map needed
| Flag | Default | Description |
|---|---|---|
--http | — | Serve over streamable HTTP instead of stdio |
--port <n> | 7437 | Port for --http |
--host <h> | 127.0.0.1 | Host for --http |
--savings | — | Record local, counts-only usage savings (opt-in; off by default) |
--share-stats | — | Also upload the counts-only usage ledger to Vibgrate (opt-in; implies --savings; disabled under --local; never code, paths, or questions) |
--dedup | — | Collapse a node's heavy relation lists on repeat reads within a session (opt-in; saves tokens) |
--surface <mode> | full | Tool listing: hot lists only the navigation core (orient, search_symbols, query_graph, get_node) to cut per-step schema tokens; every tool stays callable. Env VG_MCP_SURFACE |
--tools <names> | — | Comma-separated tool names to list (listing only — all tools remain callable). Env VG_MCP_TOOLS |
--no-refresh | — | Serve the map as built — skip auto-rebuild when files change |
--no-watch | — | Disable the event-driven file watcher; freshness falls back to the periodic probe |
--memory | — | Expose cross-agent memory tools (memory_search / memory_save) scoped to this project. Env VG_MEMORY=1 |
--local | — | Never touch the network |
Context compression flags
--compress adds a second listener to the same process — an Anthropic- and OpenAI-compatible endpoint that shrinks tool output and older turns on their way to the model, plus the compression MCP tools. One runtime, two listeners; there is no separate server to start or stop. Full description: context compression.
| Flag | Default | Description |
|---|---|---|
--compress | off | Turn on context compression: the listener, plus the compress_content / retrieve_original / compression_stats MCP tools |
--compress-port <n> | 8787 | Port for the compression listener (independent of --port; env VG_PROXY_PORT) |
--compress-only | off | Compression without a code map: none is built, none is required (implies --compress) |
--compress-mode <m> | cache | cache (newest turn only, prompt-cache safe) or token (maximum removal) |
--background | off | Start the compression listener as a background process — or reuse the healthy one already on the port — and return; no code map, no MCP. vg install <agent> --compress runs this for you; vg serve stop ends it |
--profile <p> | coding | coding, balanced, aggressive, or general |
[-- <agent> …] | — | Run one agent session through the listener, then restore its environment — nothing written |
vg serve --http --compress # MCP over HTTP and compression, one process
vg serve --compress --profile aggressive # compress harder, for long autonomous runs
vg serve --compress --background # start (or reuse) the listener in the background, then return
vg serve --compress claude # one Claude Code session, environment only — everything after the agent name is Claude's
If a healthy listener is already on the port, a second vg serve --compress attaches to it instead of failing — several assistants each spawning their own vg serve is the normal case.
Subcommands
| Subcommand | Description |
|---|---|
vg serve status | What the local runtime is serving: compression listener, attached agents, sign-in state (--compress-port <n>, --json) |
vg serve stop | Stop a background compression listener (graceful, then SIGTERM; --compress-port <n>) |
vg serve config | Print settings.json and the effective value and source of every VG_* knob |
vg serve config set <KEY> <VALUE> / unset <KEY> | Persist or remove a VG_* setting; hot knobs apply on the next request |
vg serve compress [file] | Run the compression pipeline over a file or stdin by hand — the debug path |
vg serve retrieve <hash> | Expand a compression marker back to its original, or a slice of it (--grep, --lines, --head, --tail, --json-path); --list, --purge, --stats |
vg serve memory | Cross-agent project memory: list, search, add, delete, stats, export, import, clear |
What it serves
All tools are read-only. Over the code map and dependency data: orient, query_graph, get_node, find_path, impact_of, tests_for, check_drift, list_vulnerabilities, vuln_attribution, upgrade_impact, and the version-correct resolve_library / library_docs. With --compress: compress_content, retrieve_original, compression_stats. With --memory: memory_search, memory_save.
While vg serve runs in a terminal, a live status block shows uptime, which AI clients are connected, calls and timing per tool, and the context tokens served against the estimated grep-and-read baseline. It is in-memory only; recording (--savings) and sharing (--share-stats) stay opt-in.
Wiring to your AI assistant
Use vg install to wire vg serve into your assistant's MCP config, and --compress to route it through the compression listener as well:
vg install claude
vg install --detect
vg install --all
vg install claude --compress
Or wire manually — point your assistant to vg serve as an stdio MCP server.
Related
- vg install — auto-wire to your AI assistant
- Context compression — what
--compressdoes, and its settings - vg savings — what the map and compression saved
- vg build — build the code map first
- vg doctor — check MCP launch, map freshness, and the compression listener
- vg lib — version-correct library docs