Kaloko

Guides

How teams run QA with their agent: what to install, what to ask, what comes out. Each guide ends with a live demo run and the three install steps.

By use case

Accepting a task with an AI agent, in the pull request

The agent walks the acceptance plan on local or review, Kaloko captures every screen and e-mail, checks the criteria and puts one canvas link into the PR. The reviewer accepts or returns it there.

Read the guide →
Regression before a release: compare the run with the accepted baseline

Run the flows against staging, compare with the baseline run the team accepted, and see what moved: step statuses, criteria, evidence, pixels.

Read the guide →
SEO and landing-page checks on production, read-only, every day

A read-only run of your public pages as a visitor: headings, metadata, hreflang, Open Graph, structured data, overflow, console. Shared, compared with yesterday, fixed before anyone else notices.

Read the guide →
E-mail flows: capture the message next to the screen that sent it

Sign-up confirmation, order receipt, password reset. Kaloko captures the e-mail from Mailpit or letter_opener as a step of the flow, checks its subject and links, and shows it on the canvas next to the form.

Read the guide →
Mobile apps: walk Android and iOS screens like web pages

Kaloko drives the Android emulator over adb and the iOS Simulator, captures each screen with its UI tree, checks labels, touch targets and crashes, and lays the app flow out on the same canvas as your web pages.

Read the guide →
Desktop apps: Electron and macOS apps on the same canvas

Kaloko opens Electron apps through Playwright with every web check, captures macOS app windows with their accessibility tree, and imports screens from any other tool, so desktop flows get the same review as web and mobile.

Read the guide →

By agent

QA with Claude Code: install the skill, ask in plain words, read the canvas

How Claude Code runs Kaloko: the skill it installs, what it may do without asking, what it asks first, the prompts that work, and how it reads reviewer feedback and the inbox.

Read the guide →
QA with Codex: the skill in skills/qawalk and a pointer in AGENTS.md

How Codex runs Kaloko: where the skill lives, how AGENTS.md points to it, the prompts that work, and how Codex reads reviewer feedback before it continues.

Read the guide →
Kaloko in CI: headless runs, read-only tokens, one comment per scenario

Run the same walkthrough from GitHub Actions or any CI: bundled Chromium, a tester token to share, a read-only token to read, results in the PR and on the canvas.

Read the guide →

Skills and tools

Skill catalog: the kaloko skill, its references, and the tools we recommend next to it

What the kaloko skill makes an agent do, what each reference covers, and which third-party skills and tools fit alongside, each with its implications: what runs where, which data leaves the machine, when not to use it.

Read the guide →

What it looks like

A real run of Kaloko on its own public pages, refreshed daily. This is the canvas your team gets.

10 of 10 steps passed · 2026-09-29Open the run in Kaloko ↗

Install once, then work through your agent

Kaloko runs where your code and your agent are. The service stores and versions the results, shows the canvas and collects approvals.

  1. Add the CLI to the project
    npm install --save-dev kaloko

    Needs Node 20 or newer. Update later with npm update kaloko.

  2. Create the config and install the skill
    npx kaloko init --agent claude --org <your-org>

    The skill is copied to .claude/skills/kaloko. npx kaloko doctor checks Chrome, the config and the token.

  3. Create your organization and a token

    Create an organization; you become its admin. The start page offers a tester token in one click, later under Settings → API tokens. Put it into the project .env:

    KALOKO_TOKEN=qwk_…
    TYPESAFE_API_KEY=…   # optional: semantic evaluator

Then just ask your agent

The skill teaches your agent the whole loop: it writes the acceptance plan and the scenario from the task, walks the screens, evaluates, shares the canvas, reads what reviewers said and fixes it. You do not type the commands; you review the canvas.

What the agent runs (or run it yourself, e.g. in CI)

The same loop by hand:

npx kaloko start --scenario docs/tasks/TASK-123/qa/scenario.yml --env local
npx kaloko walk        # playwright steps; agent/manual steps: kaloko capture
npx kaloko evaluate
npx kaloko share --pr
npx kaloko feedback    # what reviewers said, with ids to answer