Token-cost benchmarks: AI sessions with and without the code map
Every question you ask an AI assistant about your codebase starts with the assistant re-reading your codebase — and that re-reading is most of the token bill. Vibgrate AI Context replaces it with answers from a pre-built code map. This page publishes what that is worth, measured: the same tasks, the same model, with and without the map, at equal verified success — including the repo sizes where the map is not worth it.
~27% fewer tokens, same results
The same coding tasks, the same model, with and without the Vibgrate code map — on a ~300-file codebase. We count only the tasks both runs solved, so it is fewer tokens for the same result, never “cheaper because it gave up”.
Measured with a frontier coding model. On tiny repositories grep already finds the code and the map costs more — we publish that part of the curve too.
How the numbers are measured
Same model, twice
Each coding task runs twice with the same model on two fresh copies of the repo: once with generic file tools (list, search, read, write), once with Vibgrate AI Context (vg serve) for discovery.
Only discovery differs
Both runs can read and write files. The only difference is how the model finds the right code: walking and grepping the tree, or asking the pre-built code map.
Success is verified, not assumed
After each run, a deterministic verifier executes the edited code in the language’s real runtime and checks behaviour. A run that fails the verifier counts for nothing.
Savings only count when results match
Reduction is computed only across tasks both runs solved. The headline is “fewer tokens, same results” — never “cheaper because it gave up”. Tiers where the map saves nothing are published too.
The tasks are symptom-described — “the total is wrong after a promotion”, “refunds omit the tax” — the shape a real bug report takes, not “change function X”. They run against a ~300-file commerce codebase plus small fixtures, so the curve shows where a code map earns its keep and where grep already wins. The mix is fixed in advance (a written pre-registration) so the result can’t be tuned after the fact. Tokens are the model’s own reported usage summed across every step. The run is expensive (a live model over the whole corpus), so it is measured intermittently against a pinned CLI build and reviewed before publication — not on every release. The report below is pinned to the version it was last measured on.
Latest report — Vibgrate CLI v2026.921.1
Model gpt-5.4-mini · measured 2026-09-28 · step cap 30 · 34 of 38 tasks solved by both runs. Tokens 27.1% fewer overall · turns 0% fewer (175 → 175) · elapsed time 13.1% more. This is one A/B run against the shipped build; each task ran three times and the arm number is the median (per-task variance is large, so the median is the honest figure).
The curve — result by repository size
| Repo size | Tasks compared | File-tools tokens | AI Context tokens | Tokens | Turns | Turn result |
|---|---|---|---|---|---|---|
| Small (3–5 files) | 5/6 | 14,355 | 12,573 | 12.4% fewer | 20 → 16 | 20% fewer |
| Mid (~300 files) | 28/31 | 729,920 | 551,777 | 24.4% fewer | 148 → 154 | 4.1% more |
| XL (~1,000 files) | 1/1 | 49,753 | 14,416 | 71% fewer | 7 → 5 | 28.6% fewer |
A code map earns little on tiny repositories — any tool finds the code instantly, so the map’s fixed overhead shows as a cost. The saving appears as the repository grows and discovery becomes the dominant token cost. We publish the whole curve, not a single flattering number. Turns are model steps — each one re-bills the growing context — so fewer turns is the sequential-latency saving.
Every task, every repository
| Task | Repository | Language | Files | File tools | AI Context | Result |
|---|---|---|---|---|---|---|
| Fix order total to apply the discount | orders-api | JavaScript | 5 | 3,111 tok · 4 steps | 3,450 tok · 4 steps | 10.9% more |
| Fix subscription invoices billing 100× too high (mid repo) | meridian-commerce | JavaScript | 323 | 12,446 tok · 4 steps | 23,779 tok · 6 steps | 91.1% more |
| Apply the checkout promotion to invoice totals (XL repo) | meridian-commerce-xl | JavaScript | 963 | 49,753 tok · 7 steps | 14,416 tok · 5 steps | 71% fewer |
| Make slugify lowercase and punctuation-safe | slug-lib | JavaScript | 4 | 2,524 tok · 4 steps | 2,105 tok · 3 steps | 16.6% fewer |
| Apply gift cards as a fixed dollar amount (mid repo) | meridian-commerce | JavaScript | 323 | 11,706 tok · 4 steps | 15,569 tok · 5 steps | 33% more |
| Apply the checkout promotion to invoice totals (Python) | atlas-crm | Python | 257 | 10,796 tok · 5 steps | 15,891 tok · 5 steps | 47.2% more |
| Stop billing shipping on packaging weight (mid repo) | meridian-commerce | JavaScript | 323 | 19,348 tok · 5 steps | 10,561 tok · 4 steps | 45.4% fewer |
| Round totals to the nearest cent, not truncate (mid repo) | meridian-commerce | JavaScript | 323 | 24,313 tok · 5 steps | 17,446 tok · 6 steps | 28.2% fewer |
| Fix tiny-validate 1.2 call site from library docs | libdocs-mini | JavaScript | 4 | 2,501 tok · 4 steps | 2,121 tok · 3 steps | 15.2% fewer |
| Charge the card fee on the original captured amount (mid repo) | meridian-commerce | JavaScript | 323 | 36,351 tok · 5 steps | 13,624 tok · 5 steps | 62.5% fewer |
| Round refund amounts to cents (mid repo) | meridian-commerce | JavaScript | 323 | 82,141 tok · 7 steps | 59,996 tok · 11 steps | 27% fewer |
| Prefix normalizeId once; keep all callers on core | impact-mini | JavaScript | 5 | 2,726 tok · 4 steps | 2,053 tok · 3 steps | 24.7% fewer |
| Free shipping should use the pre-discount subtotal (mid repo) | meridian-commerce | JavaScript | 323 | 19,963 tok · 5 steps | 13,449 tok · 5 steps | 32.6% fewer |
| Fix the tax-exemption boundary (mid repo) | meridian-commerce | JavaScript | 323 | 10,663 tok · 4 steps | 9,602 tok · 4 steps | 10% fewer |
| Implement parseArgs from how it is used | argv-cli | JavaScript | 3 | did not finish | did not finish | — |
| Implement the statement balance from its usage (mid repo) | meridian-commerce | JavaScript | 323 | 18,220 tok · 6 steps | 20,193 tok · 5 steps | 10.8% more |
| Apply the checkout promotion to invoice totals (TypeScript) | harbor-api | TypeScript | 248 | 22,179 tok · 5 steps | 14,704 tok · 5 steps | 33.7% fewer |
| Apply the checkout promotion to invoice totals (mid repo) | meridian-commerce | JavaScript | 323 | 20,794 tok · 6 steps | 14,273 tok · 5 steps | 31.4% fewer |
| Implement the late-fee calculator from its usage (mid repo) | meridian-commerce | JavaScript | 323 | 12,287 tok · 4 steps | 20,359 tok · 5 steps | 65.7% more |
| Apply the checkout promotion to invoice totals (Go) | freight-routing | Go | 226 | did not finish | 57,169 tok · 12 steps | — |
| Implement the refund amount calculator (mid repo) | meridian-commerce | JavaScript | 323 | 37,511 tok · 8 steps | 25,837 tok · 6 steps | 31.1% fewer |
| Fix the monthly-to-annual rate conversion (mid repo) | meridian-commerce | JavaScript | 323 | 59,881 tok · 5 steps | 46,716 tok · 6 steps | 22% fewer |
| Apply the checkout promotion to invoice totals (Ruby) | ledger-app | Ruby | 236 | 10,487 tok · 4 steps | 9,729 tok · 4 steps | 7.2% fewer |
| Fix the cents/dollars mixup in shipping quotes (mid repo) | meridian-commerce | JavaScript | 323 | 16,886 tok · 4 steps | 10,841 tok · 4 steps | 35.8% fewer |
| Fix the sign on account running balances (mid repo) | meridian-commerce | JavaScript | 323 | did not finish | 35,351 tok · 4 steps | — |
| Apply the checkout promotion to invoice totals (Java) | inventory-forge | Java | 228 | 18,873 tok · 4 steps | 15,838 tok · 5 steps | 16.1% fewer |
| Stop charging sales tax on the shipping fee (mid repo) | meridian-commerce | JavaScript | 323 | 21,978 tok · 4 steps | 10,523 tok · 4 steps | 52.1% fewer |
| Raise the sales-tax rate everywhere it is applied (mid repo) | meridian-commerce | JavaScript | 323 | 32,872 tok · 9 steps | 33,785 tok · 9 steps | 2.8% more |
| Apply the checkout promotion to invoice totals (PHP) | storefront-cart | PHP | 226 | 26,676 tok · 6 steps | 19,723 tok · 6 steps | 26.1% fewer |
| Fix wholesale lines billing below the agreed price (mid repo) | meridian-commerce | JavaScript | 323 | 14,496 tok · 4 steps | 10,244 tok · 4 steps | 29.3% fewer |
| Stop reporting amounts from dropping cents (mid repo) | meridian-commerce | JavaScript | 323 | 70,976 tok · 7 steps | 43,621 tok · 8 steps | 38.5% fewer |
| Apply the checkout promotion to invoice totals (Bash) | ops-toolkit | Bash | 235 | 16,387 tok · 5 steps | 17,718 tok · 6 steps | 8.1% more |
| Include the sales tax in return credits (mid repo) | meridian-commerce | JavaScript | 323 | 29,636 tok · 5 steps | 13,672 tok · 5 steps | 53.9% fewer |
| Raise the free-shipping minimum everywhere (mid repo) | meridian-commerce | JavaScript | 323 | 29,734 tok · 8 steps | 16,248 tok · 6 steps | 45.4% fewer |
| Apply the checkout promotion to invoice totals (C) | sensor-hub | C | 230 | 9,110 tok · 4 steps | did not finish | — |
| Stop loyalty credits from reducing taxable amount (mid repo) | meridian-commerce | JavaScript | 323 | 21,891 tok · 6 steps | 18,175 tok · 6 steps | 17% fewer |
| Make the two discount helpers agree (mid repo) | meridian-commerce | JavaScript | 323 | 20,429 tok · 4 steps | 9,661 tok · 4 steps | 52.7% fewer |
| Apply discount with README cents rounding | commerce-docs | JavaScript | 6 | 3,493 tok · 4 steps | 2,844 tok · 3 steps | 18.6% fewer |
“did not finish” means that run failed the behaviour verifier; such tasks are excluded from every aggregate on this page.
Previous reports
- v2026.917.125.1% fewer overall · 36/38 tasks comparable2026-09-21
- v2026.911.123.2% fewer overall · 35/38 tasks comparable2026-09-14
- v2026.903.337% fewer overall · 36/38 tasks comparable2026-09-07
- v2026.829.116.7% fewer overall · 35/38 tasks comparable2026-08-31
- v2026.825.219.7% fewer overall · 36/38 tasks comparable2026-08-25
- v2026.819.321.7% fewer overall · 36/38 tasks comparable2026-08-24
- v2026.814.226.9% fewer overall · 35/38 tasks comparable2026-08-17
- v2026.722.240.7% fewer overall · 31/35 tasks comparable2026-07-22
- v2026.718.222.6% fewer overall · 33/35 tasks comparable2026-07-20
- v2026.711.233.6% fewer overall · 32/35 tasks comparable2026-07-13
- v2026.708.332.5% fewer overall · 32/35 tasks comparable2026-07-09
- v2026.704.122.9% fewer overall · 24/25 tasks comparable2026-07-04
These reports sit alongside the CLI release benchmarks (scan correctness, code-graph extraction, retrieval accuracy, performance) and follow the same rule: measured on a pinned corpus, reviewed by a human before anything appears here. Unlike those, which run with every release, this token-savings run is measured intermittently because it drives a live model.
See the code graph in action
A replay of the actual CLI querying the code map — ask a question in plain English, trace a change’s blast radius, follow a path between symbols. Nothing executes in your browser.