Wallaby for Coding Agents

  29 Sep 2026   16 min read

Coding agents can write a lot of code in very little time. Reviewing every change by hand does not scale, so tests, especially fast unit tests, carry more of the load. They give agents a repeatable way to check their work in seconds, often at a low compute cost.

But a passing test suite can still be a weak one. A final block of terminal output tells an agent what passed or failed, but often not whether the code behaved as intended.

Wallaby gives coding agents the runtime feedback they need to move fast with confidence. It continuously runs your existing JavaScript, TypeScript, and Python tests in the background and keeps results and coverage current as files change. The agent can inspect the test suite as a live runtime model instead of stopping at pass or fail.

A coding agent inspecting a highlighted code branch through Wallaby test results, branch coverage, execution trace, and runtime values Wallaby lets an agent move from a line of code to the tests that cover it, the path they took, and the values produced at runtime.

With Wallaby, an agent can ask for the evidence it needs next: a failing test, the tests that cover one line, branch coverage, an execution trace, or a runtime value from a specific test context.

That changes how an agent can approach the work:

  1. Start with a baseline. Before implementing a feature, fixing a bug, or refactoring code, the agent can inspect the current results, coverage, and execution traces, even in plan mode. It sees pre-existing failures, coverage gaps, and the tests that already exercise the code it plans to change. Test names and assertions show which expectations the change must preserve.
  2. Check each change. While the agent writes code and tests, Wallaby keeps test, coverage, and runtime evidence available from each execution. If the agent captured a baseline, it can compare the new behavior with the previous state. Either way, it can find failures, confirm the intended behavior, and refine the change based on what actually ran instead of guesswork.
  3. Improve the code and tests. During a pull request, code review, or scheduled quality task, the agent can examine coverage gaps, execution paths, runtime values, test timings, complexity, and change risk. It can use those signals to strengthen the implementation and add focused tests for the behavior and edge cases that matter.

Wallaby keeps the same test and runtime context available throughout the job, from the first plan to the final verification.

If you’re new to Wallaby, Wallaby for coding agents is free during the beta. You can try it on a real project today.

Beyond pass/fail

Wallaby works with your existing test framework. Your Jest, Vitest, and other supported tests remain the source of truth. What changes is how the agent sees and investigates the results.

A conventional test runner is built to report the result of a completed run. The agent gets a block of terminal output and then has to search it, run another command, or add temporary logs whenever it needs more detail. The answer may be buried in thousands of lines or missing from the output entirely.

A conventional coding agent repeatedly runs tests, searches terminal output, and adds temporary logs, while a Wallaby-connected agent starts with a live summary and opens focused failures, coverage, execution traces, and runtime values Wallaby keeps the current test state ready and lets the agent open the failure, coverage, trace, or runtime value it needs next.

Wallaby keeps the same test suite live and queryable. In practice, that means:

  • Wallaby keeps results current in the background and runs affected tests as files change. The agent can check each edit without starting a cold full-suite run for every question.
  • The agent starts with a concise summary and opens more detail only when it needs it. It can focus on one failure, file, test, or location instead of loading the entire run into context.
  • Coverage relationships, execution paths, logs, errors, runtime values, test timings, cyclomatic complexity, and change risk remain available when the agent needs them. They show how the code behaved and whether a test exercised the behavior that matters.

This gives the agent a shorter path from question to answer. It starts with a compact signal and opens only the relevant detail, using fewer tokens than conventional test-runner CLI output.

Wherever the coding agent runs

Wallaby is not tied to an editor. Any coding agent that supports skills can use it from a CLI or graphical app, including Claude Code, Codex, GitHub Copilot, Cursor, Windsurf, OpenCode, and Pi.

That covers developer-guided sessions as well as autonomous work. In either case, Wallaby keeps the testing context ready, so the agent does not have to reconstruct it from terminal output. Wallaby can run:

  • in a supported editor through the Wallaby extension
  • alongside any editor in Standalone Mode
  • directly in a local project through the Wallaby skills
  • inside a container, including a DevContainer
  • in the cloud or another remote environment, with agents such as GitHub Copilot Cloud Agent

A continuous Wallaby test signal connecting local, container, and cloud coding agent environments The agent gets the same live test feedback whether it runs locally, in a container, or in the cloud.

Developers can also combine coding agents with the Wallaby UI and editor extensions. Whether the agent starts Wallaby or connects to an instance already running in the editor, the developer can open the UI to inspect current test results, coverage, and execution data.

A local agent can reuse the Wallaby instance already running in the editor. The developer and agent then work from the same live state, including local changes and current test results. The developer can keep steering the agent over SSH or from a mobile agent app connected to the same machine, which is useful when the full test suite takes a while.

One test session for the whole task

A conventional runner often starts a new process whenever the agent needs an answer. Wallaby stays in the background for the whole agent session, watching files and keeping test results, coverage, and execution data current. The agent can ask for the latest state instead of rebuilding that context. If Wallaby is not running yet, the agent starts it once and leaves it available between requests.

The agent can begin in exclusive mode with tests from selected files. As the task grows, it can add related files or switch the same Wallaby instance to project mode for full-project verification. The test engine keeps running while the scope changes.

Each state request starts with a concise report: the active mode, passed, failed, skipped, and todo counts, and current coverage. Wallaby reports fatal and global errors separately. For relevant failures, it includes assertion details, logs, stack traces, timings, and covered files. Full test, coverage, and timing details are available when the agent needs them.

A coding agent expands from selected test files to full-project verification while the same Wallaby process keeps results current The agent starts with selected tests, widens the scope as needed, and verifies the full project without restarting Wallaby.

What the tests reveal about the code

With Wallaby running, the agent can move from the project summary to one source file, test file, source location, or executed test. Wallaby answers from its current results and coverage, so each question can stay focused without restarting the test process or loading the full run into context.

Map the change surface

Before changing a source file, the agent can inspect the whole file or focus on one line or expression. Covering Tests shows which tests reach that code, along with their status, location, and timing. Errors, logs, stack traces, and covered files are available when needed.

That context matters before the first edit. The agent can see what the current tests expect, spot failures that already existed, understand which behavior is covered, and decide where another focused test is needed. For a test file, the view runs in the other direction: it shows the tests in the file and the source files they exercise.

A selected line in discount.ts connected to three covering tests, with their status, timing, assertions, and related source files Before the agent edits discount.ts, Wallaby shows the tests that cover the selected line, an existing failure, and the files those tests also exercise.

Wallaby’s .wcov format puts coverage markers directly alongside the source. Each coverable line is marked as fully, partially, or not covered. On a partially covered line, Wallaby also points to the expression that did not run.

Filter the view to one test and the markers show only what that test covered. The agent can see which path it took and which path it missed before looking at the test setup, assertions, or implementation. This is useful when planning a change and when a test or coverage result is surprising.

A WCOV view of auth.ts filtered to the admin can edit test, with the reached expression highlighted and missed expressions and lines called out Filtered to one test, WCOV keeps the source visible and marks exactly which lines and expressions ran.

Follow one test end to end

When a test behaves unexpectedly, especially in unfamiliar code, the agent needs to see how execution reached the result. The agent selects a test, and Wallaby shows the exact execution path for that test. The status, execution time, errors, logs, and covered files appear alongside the trace.

Wallaby calls this a Test Execution Trace, an agent-readable version of its Test Story. It lists the source lines in the order they ran, even when execution crosses files. The trace separates imports and setup from the start of the selected test, then follows each call into production code and back again. Coverage tells the agent what ran. The trace shows the route the test took.

A failing coupon test traced in execution order from setup through discount and clock code to the final assertion This failing test moves through setup, the test body, discount code, a clock helper, and coupon logic before returning to the failed assertion.

The trace replaces a stack trace and repeated file reads with one ordered path. The agent can find where execution first departs from the expected flow and catch setup or helper code that would be easy to overlook.

Together with test-specific coverage, the trace helps the agent keep the fix small, correct a misleading test, or add the missing case.

What file metrics can suggest

File metrics can give the agent useful clues about what to do next. Wallaby brings together related test counts and timings with line count, coverage, file size, complexity, and change risk. Taken together, these metrics may point to slow or tightly coupled tests, complex files with weak coverage, or changes that deserve smaller steps and wider verification. Wallaby offers guidance, but the agent makes the decision. It can weigh the metrics against the task and the surrounding code before choosing whether to refactor, add a focused test, or do broader quality work.

A Wallaby file analysis for discount.ts uses coverage and change risk to suggest three possible next steps for the agent to consider Wallaby presents three possible next steps for discount.ts. The agent decides which, if any, fit the task.

This can make high project-wide coverage targets, such as 95% or 100%, practical to maintain. Coding agents can generate tests quickly. Wallaby can point them toward meaningful gaps and flag tests that add cost without useful confidence. The agent still decides whether those signals warrant action.

Inspect runtime values without changing the code

Tests and coverage can show where a problem occurs, but not always why. Often, the missing piece is the value of an expression or variable at one source location. Without runtime inspection, the agent has to form a hypothesis, add temporary logging, and run the test again.

Wallaby lets the agent inspect that expression or variable at the source location without editing the code. It evaluates the selection in every test context that reaches the line and keeps each value tied to its test. If the line runs several times, the agent can narrow the result to one test.

Several expressions and variables can be inspected across different source locations in one request. The agent can review the values together and test a whole hypothesis in one pass.

The order variable at line 42 shown with its object data and rate value for three test contexts, including one failing test Wallaby shows the value of order in every test that reaches line 42, then lets the agent focus on the failing test without adding a log.

In our evaluations, the same agent workflows used fewer tokens with Wallaby’s runtime inspection than with temporary logging and conventional test-runner CLI output.

Get started

Install the Wallaby skills in your project:

gh skill install wallabyjs/skills

or:

npx skills add wallabyjs/skills

The wallaby skill provides workflows for writing new tests and improving existing ones. These workflows use wallaby-cli, Wallaby’s API for agents, to run tests and inspect coverage, execution traces, and runtime values. The agent can also select wallaby-cli automatically during a coding task.

Write new tests

/wallaby write tests for the current changes
/wallaby write tests for account suspension and reactivation
Implement the feature described in <GitHub issue URL> and add tests using /wallaby write

wallaby write works through tests one at a time. Before writing each test, the agent predicts its result and the source paths it should exercise. It then checks that test’s result, coverage, and assertions against the prediction before moving on, including boundary and failure cases within the requested scope.

Improve existing tests

/wallaby improve src/accounts/
/wallaby improve the files changed in this branch
/wallaby improve files with coverage below 90%

wallaby improve finds uncovered branches, weak assertions, and missing boundary or failure cases. You can target files, a directory, or a Git change set and specify criteria such as coverage. Without a target or criteria, it reviews source files with coverage at or below 95%. The agent checks whether the tests would catch plausible defects and explains any remaining gaps.

You can also ask the agent to use Wallaby directly as part of a coding task:

Run the tests and fix all failing tests using wallaby.
Before implementing this feature, use Wallaby to capture the current full-project test and coverage state and identify
the tests that cover the code you expect to change. Then implement the feature, verify the affected tests first, and
check the full project state for failures or coverage regressions.
Use Wallaby to trace the failing "coupon / rejects an expired coupon" test through setup and source code. Inspect the
runtime values that explain the failure, fix the cause, and verify the result.

To make Wallaby part of the agent’s usual workflow, add instructions to your project’s AGENTS.md, including a configuration file if your project uses one:

Use the wallaby-cli skill to run, investigate, and verify unit tests for the project.
Use ./wallaby.unit.js as the configuration file.

If Wallaby is already running in your editor, the skill connects to that instance. Otherwise, it starts a background instance. Ask the agent to open the Wallaby UI in a browser to inspect results or follow its progress.

The skill can download Wallaby CLI when needed. To keep it as a development dependency in your project, install it with:

npm install --save-dev @wallabyjs/cli

For sandboxed agents, follow the sandbox configuration instructions. For containers and cloud agents, see the token and environment setup.

Free during the beta

If you are new to Wallaby, you can use Wallaby for coding agents free during the beta. Request access with:

npx --package @wallabyjs/cli wallaby access

Enter your email address when prompted and follow the activation link. Then return to the agent and start using Wallaby in your project.

Once the skill and CLI are installed, that command activates Wallaby for coding agents. It does not start an editor trial. If you want to try Wallaby in your editor too, you can do that separately:

If you already use Wallaby locally and have a licensed editor extension on the same machine, Wallaby for coding agents is included in your existing license. You do not need to request beta access.

Access through WALLABY_CLI_TOKEN, including remote and hosted workflows, is still part of the beta. It may require separate paid access after the beta, even if you already have a Wallaby license.

With Wallaby, the agent can inspect what the code actually did instead of trying to reconstruct it from partial output.

Once you have tried it, tell us what worked and what did not at hello@wallabyjs.com.