02 Visp Code open source · beta · CLI & MCP · Node 22.16+

Done means checked.

Visp Code is a harness for your coding agent. It works one slice at a time inside an authorized file scope, runs the checks each slice declares, and has an independent reviewer judge the result against your original request.

npm install -g visp-coder@beta
AgentClaim

agentThe retry button is in, and the tests pass. Done.

visp done
EvidenceExecuted
  1. C001node --test test/save.test.mjspassed
  2. C002browser journey · 390×844 · 2 capturespassed
  3. Criticread-only · sees the request, source, results and screenshots, not the claimreviewed
  4. F001A second failed save leaves the retry button disabled.required
  5. nextreopen T001 · repair F001 · rerun C002
An illustrative run. The claim is not evidence; what ran is.

Start from the real request.

Visp Code keeps the request verbatim. Under Claude Code and Codex it reads it from your recorded prompt, so the agent’s paraphrase can’t replace it. With the Codex reviewer configured, an independent tester starts writing acceptance tests from the request alone.

visp feature "Add a retry button to the failed-save banner"

RecordsThe original request, and the critic configuration pinned for this feature.

Authorize one slice.

A slice needs an outcome, a bounded write scope and a runnable check. Without a check, work stops and hands back a patch that adds one. Small changes take the light path: the whole request as one slice, checked by your tests.

visp work --check "npm test"

RecordsThe authorized scope, and the context delivered: outcomes, source excerpts, findings and notes.

Build inside the scope.

The agent edits only the slice’s files. Host hooks refuse out-of-scope writes, a Git pre-commit hook checks every commit, and an optional CI job checks the pull request against the committed brief.

visp guard

RecordsNothing yet. The agent’s own account of its work is never evidence.

Run what was declared.

Visp Code runs the slice’s checks itself: commands without a shell, and browser journeys in an isolated Chrome with screenshots. A blocked sandbox is an environment failure, not a product failure.

visp done

RecordsEach execution, bound to the source, the verifier files and the environment it ran against.

Review it independently.

When every check passes, a separate read-only reviewer judges the work against the verbatim request. It sees source, results and screenshots, never the agent’s verdicts. Required findings reopen the slice and go back as the next repair.

visp next

RecordsThe reviewer’s attributed assessment of each outcome, and at most three findings.

Check the assembled product.

Acceptance reruns every check against the whole product, including the pinned independent tests, and needs a current assessment of every mandatory outcome. Passing commands alone don’t satisfy it.

visp accept

RecordsThe acceptance result, or the outcomes that are still open.

Hand a person the evidence.

The loop always ends: in acceptance, or, once the review budget is spent, in a handoff. Either way the reviewer document collects the request, outcomes, executed results, findings and every intent change.

visp pr

RecordsA Markdown document for the pull request. It publishes nothing.

The reviewer sees

  • Your original request, verbatim
  • The brief’s outcomes
  • The current source
  • Check results and screenshots
ReturnsAn assessment of each outcome, and at most three findings. It runs read-only.
The independent reviewer judges the work against your verbatim request, not the agent’s account of it. Its findings come back as the next repair, three reviews per feature by default.
Checks
visp work won’t authorize a slice until a runnable check exercises it. Commands run without a shell; browser journeys run in an isolated Chrome with screenshots.
Scope
Edits stay inside the slice’s files. Host hooks, a Git pre-commit check and an optional CI job all call visp guard.
Tests
On new projects, an independent tester writes acceptance tests from the request alone. They are pinned only if they fail before the work starts.
Hand-off
visp pr prints the request, outcomes, executed results, findings and every intent change for the person who reviews the pull request.
Hosts
Claude Code, Codex, Cursor, Copilot, OpenCode, or any MCP client. Claude Code and Codex get the fullest hooks. Release 0.5 is a beta.

Get started

From install to a checked change.

Node.js 22.16 or newer, inside a Git repository. Swap claude-code for codex, cursor, copilot, opencode or generic.

Full setup guide
  1. 1

    Install the CLI.

    One package, three executables: visp, visp-migrate and the optional visp-runner.

    npm install -g visp-coder@beta
  2. 2

    Set up your coding host.

    init writes visp.yml and .visp/. install adds the host’s instructions, MCP registration and hooks. Commit the setup before feature work.

    visp init --harness claude-code
    visp install --harness claude-code
    visp doctor
  3. 3

    Start a feature.

    From here your agent runs the loop itself. The installed hooks send it back when it stops early.

    visp feature "Add a retry button to the failed-save banner"
    visp work --check "npm test"

Results

Measured, with limits.

Hidden checks passed on Visp Code’s own benchmark, with Claude Haiku 4.5 as the coding agent. Each dot is one run.

Reservations API44 hidden checks
303744Visp Code · run 1 · 44 of 44Visp Code · run 2 · 44 of 44Visp Code · run 3 · 44 of 44Bare · run 1 · 38 of 44Bare · run 2 · 39 of 44Bare · run 3 · 40 of 44Spec Kit · run 1 · 39 of 44Spec Kit · run 2 · 40 of 44Spec Kit · run 3 · 37 of 44BMAD · run 1 · 38 of 44BMAD · run 2 · 38 of 44BMAD · run 3 · 39 of 44
Spreadsheet engine38 hidden checks
202938Visp Code · run 1 · 36 of 38Visp Code · run 2 · 35 of 38Visp Code · run 3 · 37 of 38Bare · run 1 · 28 of 38Bare · run 2 · 32 of 38Bare · run 3 · 26 of 38Spec Kit · run 1 · 28 of 38Spec Kit · run 2 · 28 of 38Spec Kit · run 3 · 27 of 38BMAD · run 1 · 33 of 38BMAD · run 2 · 32 of 38BMAD · run 3 · 35 of 38
Bundles on an existing service60 hidden checks
505560Visp Code · run 1 · 60 of 60Visp Code · run 2 · 60 of 60Visp Code · run 3 · 60 of 60Bare · run 1 · 60 of 60Bare · run 2 · 60 of 60Bare · run 3 · 60 of 60Spec Kit · run 1 · 60 of 60Spec Kit · run 2 · 59 of 60Spec Kit · run 3 · 59 of 60BMAD · run 1 · 60 of 60BMAD · run 2 · 60 of 60BMAD · run 3 · 60 of 60
Extending an existing engine63 hidden checks
405263Visp Code · run 1 · 57 of 63Visp Code · run 2 · 62 of 63Visp Code · run 3 · 63 of 63Bare · run 1 · 51 of 63Bare · run 2 · 53 of 63Bare · run 3 · 53 of 63Spec Kit · run 1 · 56 of 63Spec Kit · run 2 · 59 of 63Spec Kit · run 3 · 41 of 63BMAD · run 1 · 56 of 63BMAD · run 2 · 51 of 63BMAD · run 3 · 50 of 63

Visp Code Other workflows · dashed line: every check passed

Show as a table, with wall time
Task (hidden checks)Visp CodeBareSpec KitBMAD
Reservations API (44)44, 44, 44 · 8–10 min38, 39, 40 · 2–4 min39, 40, 37 · 10–13 min38, 38, 39 · 3–6 min
Spreadsheet engine (38)36, 35, 37 · 10–16 min28, 32, 26 · 5–8 min28, 28, 27 · 12–15 min33, 32, 35 · 3–22 min
Bundles on an existing service (60)60, 60, 60 · 5–9 min60, 60, 60 · 2–3 min60, 59, 59 · 8–11 min60, 60, 60 · 3–5 min
Extending an existing engine (63)57, 62, 63 · 14–16 min51, 53, 53 · 9–11 min56, 59, 41 · 11–13 min56, 51, 50 · 7–13 min

Visp Code matched or beat every other arm on every task and took about as long as Spec Kit, longer than bare coding. The runs are few, the tasks are Visp Code’s own, and the reviewer was a stronger model than the worker. With a strong coding agent the differences were within noise, including on SWE-bench Verified Mini in October. Every number and its limits.

Keep going