Kaloko

All guides › By agent

By agent

QA with Claude Code: install the skill, ask in plain words, read the canvas

How Claude Code runs Kaloko: the skill it installs, what it may do without asking, what it asks first, the prompts that work, and how it reads reviewer feedback and the inbox.

Claude Code already builds the feature. With the kaloko skill it also walks the acceptance plan, captures the screens, evaluates the criteria and shares the canvas, then reads what the reviewer returned.

The problem

Without a skill the agent improvises QA: it opens the page once, says "looks fine" and moves on. Nobody can see what it looked at.

What changes with Kaloko

The skill tells Claude Code when a walkthrough is due, what "done" means (validated scenario, every step captured or explained, evaluation and build run, the link posted), what it may do without asking (local and review environments, test data there, scenario mechanics) and what needs a person first (changing acceptance criteria, any production run, uploading).

How it works

  1. Install

    npm install --save-dev kaloko, then npx kaloko init --agent claude --org <your-org> copies the skill to .claude/skills/kaloko and creates kaloko.config.yml. npx kaloko doctor checks Chrome, token, environments and the evaluator key.

  2. Ask

    "Walk the acceptance plan of TASK-123 on local and share the result." The skill routes to its references: acceptance, execution, evaluation, environments, collaboration.

  3. Drive the browser your way

    Claude Code uses the shared browser through a Playwright step script or a browser tool over MCP (the CDP endpoint is printed by kaloko browser start).

  4. Close the loop

    after sharing, the skill reads kaloko feedback and kaloko inbox before continuing on the task. Optionally connect the MCP server so the agent reads runs and feedback as tools.

Skills and prompts

Walk the acceptance plan of TASK-123 with Kaloko on local and share the result into the PR.
What did the reviewer return on the last run of TASK-123? Fix it and run again.
Compare the latest run of checkout with the baseline and explain the differences.

Expected outcome in each case: the commands the skill prescribes, a summary line, failing or unreviewed criteria with evidence, and links to the canvas; no production run without asking.

What you get

FAQ

Does Claude Code need a browser tool?

No. Scripted steps run through Playwright inside the CLI; for judgement-driven steps a browser MCP or computer use helps, but kaloko capture --url works from the shell alone.

What does the agent send to the AI evaluator?

Only a reduced text outline of each page with e-mails and tokens masked, and only for criteria marked semantic. Screenshots and raw HTML never leave your machine except to your own Kaloko organization.

What it looks like

A real run of Kaloko on its own public pages, refreshed daily. This is the canvas your team gets.

10 of 10 steps passed · 2026-09-29Open the run in Kaloko ↗

More guides

QA with Codex: the skill in skills/qawalk and a pointer in AGENTS.mdKaloko in CI: headless runs, read-only tokens, one comment per scenario

Install once, then work through your agent

Kaloko runs where your code and your agent are. The service stores and versions the results, shows the canvas and collects approvals.

  1. Add the CLI to the project
    npm install --save-dev kaloko

    Needs Node 20 or newer. Update later with npm update kaloko.

  2. Create the config and install the skill
    npx kaloko init --agent claude --org <your-org>

    The skill is copied to .claude/skills/kaloko. npx kaloko doctor checks Chrome, the config and the token.

  3. Create your organization and a token

    Create an organization; you become its admin. The start page offers a tester token in one click, later under Settings → API tokens. Put it into the project .env:

    KALOKO_TOKEN=qwk_…
    TYPESAFE_API_KEY=…   # optional: semantic evaluator

Then just ask your agent

The skill teaches your agent the whole loop: it writes the acceptance plan and the scenario from the task, walks the screens, evaluates, shares the canvas, reads what reviewers said and fixes it. You do not type the commands; you review the canvas.

What the agent runs (or run it yourself, e.g. in CI)

The same loop by hand:

npx kaloko start --scenario docs/tasks/TASK-123/qa/scenario.yml --env local
npx kaloko walk        # playwright steps; agent/manual steps: kaloko capture
npx kaloko evaluate
npx kaloko share --pr
npx kaloko feedback    # what reviewers said, with ids to answer