Kaloko

All guides › By use case

By use case

Regression before a release: compare the run with the accepted baseline

Run the flows against staging, compare with the baseline run the team accepted, and see what moved: step statuses, criteria, evidence, pixels.

Three agents, four branches, one staging, zero memory of what changed. Before a release we want one answer: what looks different from the version we accepted, and is it on purpose.

The problem

Regression suites tell us whether the code still passes. They do not show us the screens, and they do not know which differences a person would care about. So we open staging, scroll, and hope.

What changes with Kaloko

A run the team accepted becomes the scenario's baseline. Every later run can be compared with it on the canvas: pixel differences per step, criteria whose verdict or evidence changed, added or removed steps. The same comparison prints as text for the agent and the release notes.

How it works

  1. Accept a baseline

    on a good run the reviewer presses *Accept run* and *Set as baseline*. One baseline per scenario.

  2. Run the flows against staging

    kaloko start --scenario qa/flows/checkout.yml --env staging, kaloko walk, kaloko evaluate, kaloko share. Test data belongs to the staging prepare step; the scenario creates only what it needs.

  3. Compare

    on the canvas choose *Compare → baseline*: badges with the share of changed pixels per step, Δ on changed criteria, side-by-side and diff views in the lightbox. In the terminal kaloko compare prints the same.

  4. Decide

    intended changes become the new baseline; unintended ones go back as returned steps with notes.

Skills and prompts

Run every flow in qa/flows against staging, compare each with its baseline and list what changed with the evidence. Do not accept anything.

Expected outcome: shared runs per scenario, a text diff per scenario (statuses, verdicts, evidence), canvas links with #cmp=<baseline>.

Explain the differences between run A and its baseline in two sentences per changed step, for the release notes.

Expected outcome: a short list the agent derives from kaloko compare --json; nothing invented, every line points at a step.

What you get

FAQ

Are small rendering differences noise?

Fonts and anti-aliasing move a few tenths of a percent; the badge shows the share so you can ignore it. Masks for dynamic regions are on the roadmap.

Can I compare two arbitrary runs?

Yes: the *Compare* selector lists the baseline, the previous run and other runs of the scenario; #cmp=<run id> in the URL shares the view.

What it looks like

A real run of Kaloko on its own public pages, refreshed daily. This is the canvas your team gets.

10 of 10 steps passed · 2026-09-29Open the run in Kaloko ↗

More guides

Accepting a task with an AI agent, in the pull requestSEO and landing-page checks on production, read-only, every dayE-mail flows: capture the message next to the screen that sent itMobile apps: walk Android and iOS screens like web pages

Install once, then work through your agent

Kaloko runs where your code and your agent are. The service stores and versions the results, shows the canvas and collects approvals.

  1. Add the CLI to the project
    npm install --save-dev kaloko

    Needs Node 20 or newer. Update later with npm update kaloko.

  2. Create the config and install the skill
    npx kaloko init --agent claude --org <your-org>

    The skill is copied to .claude/skills/kaloko. npx kaloko doctor checks Chrome, the config and the token.

  3. Create your organization and a token

    Create an organization; you become its admin. The start page offers a tester token in one click, later under Settings → API tokens. Put it into the project .env:

    KALOKO_TOKEN=qwk_…
    TYPESAFE_API_KEY=…   # optional: semantic evaluator

Then just ask your agent

The skill teaches your agent the whole loop: it writes the acceptance plan and the scenario from the task, walks the screens, evaluates, shares the canvas, reads what reviewers said and fixes it. You do not type the commands; you review the canvas.

What the agent runs (or run it yourself, e.g. in CI)

The same loop by hand:

npx kaloko start --scenario docs/tasks/TASK-123/qa/scenario.yml --env local
npx kaloko walk        # playwright steps; agent/manual steps: kaloko capture
npx kaloko evaluate
npx kaloko share --pr
npx kaloko feedback    # what reviewers said, with ids to answer