Sinfin Kaloko
Shipped with AI. Now let’s actually look at it.
We all build faster than we review. Screens appear in branches nobody has opened, in languages nobody has clicked through. That is not an AI problem, it is a looking problem. Kaloko walks the acceptance plan, captures every screen and e-mail on the way, checks each criterion and lays the run out as one canvas: the contact sheet of the sprint, with a stamp of who looked at what. For teams that ship with coding agents: developers get evidence, QA and product owners review in one place, clients approve what they saw.
Free plan, free for good · e-mail sign-in · the first run in ten minutes
Here is what we built and what got approved
A live run of Kaloko on its own public pages, refreshed daily. Findings are genuine; when something fails, it is fixed in the next release, not hidden.
Things that happen
- PR merged at 2:14. Screenshots nowhere.
- The Slovak version exists. Probably.
- Who approved this? The bot did.
- Three agents, four branches, one staging, zero memory of what changed.
- “Works on my machine” is now “works in my context window.”
How it works
- Acceptance plan
An
ACCEPTANCE.mdnext to the task says what each screen must achieve. Ascenario.ymlturns it into steps, branches and criteria the CLI can run. - Walk
A Playwright script, an agent with a browser, or a person walks the flow in a shared browser. Every step is captured the same way: full-page screenshots on desktop and mobile, rendered HTML, console errors, e-mails from the test mailbox.
- Check
Countable criteria (elements, texts, prices, metadata, overflow) are checked in the live page with evidence. Criteria that need judgement get a probability from an AI evaluator that reads a reduced text outline of the page; Kaloko turns it into pass, fail or needs review.
- Canvas and stamps
The run becomes a canvas: screens, branches, locales, evidence. Reviewers approve a step, return it with a note, or decide a manual criterion. The agent reads the notes with
kaloko feedbackand fixes them in the next run.
Three stamps. The human is the point.
A person looked and agreed.
A person looked and wrote what to change.
The evaluator was not sure. Someone decides on the canvas.
What the canvas shows
- Desktop and mobile screenshots of every step, e-mails and third-party screens
- Pass, partial or fail per criterion, per locale and viewport, with the evidence
- Page metadata: title, description, canonical, hreflang, Open Graph, JSON-LD
- Branch, commit, author, PR and environment; the scenario’s history and a pixel diff against the baseline
- Replay: step by step with a choice at every decision, for people who do not read canvases
- Links to a specific step and criterion you can send to colleagues
What teams use it for
The agent walks the plan on local or review, shares the canvas, posts the link into the PR. The reviewer accepts or returns it there.
Run the flows against staging, compare with the accepted baseline, see what moved: statuses, criteria, pixels.
Read-only runs of public pages: headings, metadata, hreflang, overflow, console. Kaloko runs this on itself every day.
Claude Code, Codex or CI run the same commands a person would; the canvas is where the two meet.
Fits the tools you already use
- Claude Code and Codexa skill that knows when and how to walk, capture, evaluate and share
- GitHub
kaloko share --prkeeps one comment per scenario on the pull request - CIthe CLI runs headless; a read-only token is enough to read results
- Mobile and desktop appsAndroid over adb, the iOS Simulator, Electron and macOS apps on the same canvas; screens from Maestro, Appium or XCUITest imported with their UI tree (guide)
- MCPruns, feedback, verdicts and the inbox over MCP at
kaloko.app/mcp - E-mail and Slackdigest of what needs you, immediate mail when your run is returned, one Slack channel for the team
For designers: the design development and QA run against
Design the screens of a flow as live HTML with your own agent and skills, together with your team in real time. Accept a version, and every implementation run is checked against it: pixels, structure, tokens and contrast.
Your agent writes each step as live HTML; the team clicks through the prototype, pins comments on elements and sees each other’s cursors.
Publishing makes a version of the whole flow. Unchanged steps are stored and rendered once; compare shows only what moved.
An accepted version becomes the reference. Developers’ agents take the live files and tokens from it, step by step.
The design pack compares every implementation run with the reference and names the token to use where it drifts.
Install once, then work through your agent
Kaloko runs where your code and your agent are. The service stores and versions the results, shows the canvas and collects approvals.
- Add the CLI to the project
npm install --save-dev kalokoNeeds Node 20 or newer. Update later with
npm update kaloko. - Create the config and install the skill
npx kaloko init --agent claude --org <your-org>The skill is copied to
.claude/skills/kaloko.npx kaloko doctorchecks Chrome, the config and the token.npx kaloko init --agent codex --org <your-org>The skill is copied to
skills/kalokoandAGENTS.mdgets a pointer to it.npx kaloko doctorchecks Chrome, the config and the token.npx kaloko init --org <your-org>No skill is installed. Drive the CLI from scripts or CI;
kaloko helplists every command. - Create your organization and a token
Create an organization; you become its admin. The start page offers a tester token in one click, later under Settings → API tokens. Put it into the project
.env:KALOKO_TOKEN=qwk_… TYPESAFE_API_KEY=… # optional: semantic evaluator
Then just ask your agent
The skill teaches your agent the whole loop: it writes the acceptance plan and the scenario from the task, walks the screens, evaluates, shares the canvas, reads what reviewers said and fixes it. You do not type the commands; you review the canvas.
- From a task to a shared walkthrough
Read the task in docs/tasks/TASK-123, write its acceptance plan and a Kaloko scenario, walk it on local and share the result on the pull request.Creates
ACCEPTANCE.mdandqa/scenario.yml, captures every step in each language and screen size, and posts the link. - Process the feedback
Go through the open Kaloko feedback on TASK-123: fix what reviewers rejected, run the affected steps again, share, and answer each comment with what changed.Reads threads with
kaloko feedback, resolves them with the commit that fixed them. - Check a public site
Draft a Kaloko scenario for https://example.com (home, pricing, contact) with the seo, a11y and perf packs, run it on production read-only and summarise what fails.Read-only on production; check packs add SEO, accessibility and performance criteria without writing them by hand.
- Keep the history tidy
Clean up old Kaloko runs of TASK-123: show me what a prune keeping the newest 3 would remove, then do it once I agree.Always a dry run first; accepted, kept and commented runs stay.
What the agent runs (or run it yourself, e.g. in CI)
The same loop by hand:
npx kaloko start --scenario docs/tasks/TASK-123/qa/scenario.yml --env local
npx kaloko walk # playwright steps; agent/manual steps: kaloko capture
npx kaloko evaluate
npx kaloko share --pr
npx kaloko feedback # what reviewers said, with ids to answerData & security
The AI evaluator receives a reduced text outline of the page with e-mail addresses and tokens masked. It never receives screenshots or raw HTML.
People sign in with a one-time link sent to their work e-mail. A viewer reads and reviews, a tester also pushes runs, an admin also manages members, domains and tokens.
Every run has an expiry date (30 days by default) and is deleted afterwards. Uploaded files are served from a separate origin under short-lived signed links; captured HTML runs in a sandbox without scripts.
A production environment is always read-only in the CLI: public pages as a visitor, no sign-in to back offices, no data created.
Free stays free. Paid plans start with a 14-day trial: card required, cancel before it ends and nothing is charged. Terms of service.
Pricing
Viewers are free and unlimited in every plan: the people who look must never be the bottleneck. You pay for what you ship: shared runs per month, tester and admin seats, retention and storage. Team and Business start with a 14-day trial. For businesses; prices exclude VAT.
Free
free for good
no card needed
For one person and their agent
- 30 shared runs / month
- 2 tester + admin seats
- unlimited viewers
- 30 days retention · 2 GB storage
- canvas, replay, compare, inbox
49 € / month
billed monthly
yearly you save 118 €
39 € / month 49 €
billed 470 € yearly
you save 118 €
For a team shipping every week
- 300 shared runs / month
- 10 tester + admin seats
- unlimited viewers
- 90 days retention · 20 GB storage
- MCP, Slack, export, required approvers
199 € / month
billed monthly
yearly you save 478 €
159 € / month 199 €
billed 1,910 € yearly
you save 478 €
For several teams and documentation
- 1500 shared runs / month
- unlimited tester + admin seats
- unlimited viewers
- 365 days retention · 100 GB storage
- accepted and reference runs never expire; documentation layer coming
on request
yearly contract, invoice
SSO, SLA, DPA, self-hosted
Governance and operations
- unlimited shared runs / month
- unlimited tester + admin seats
- unlimited viewers
- unlimited storage
- SSO/SCIM, custom retention, own files domain, self-hosted, SLA, DPA, invoice
Semantic evaluation is bring-your-own-key in every plan; local previews are always unlimited. Compare all plans →
Questions we get
Does Kaloko replace tests?
No. Unit and end-to-end tests tell you whether the code works. Kaloko shows what was built and lets a person say whether it is what was meant.
Do I need an AI evaluator?
No. Countable criteria run without one. For judgement criteria bring your own key (the default is JEV by TypeSafe); without it they stay “not evaluated” until a reviewer decides.
What does it cost?
Free stays free. Team is 49 € and Business 199 € per month, −20 % with yearly billing, both with a 14-day trial. Data stays readable and exportable whatever the plan.
Where is the data?
Cloudflare storage chosen by Cloudflare; the operator is a Czech company under EU law. Runs are yours and expire on the date you set.
What comes next
A weekly “what changed” canvas assembled from runs, check packs for SEO and accessibility, and a GitHub check that turns green when the run is accepted. We tell admins before anything changes that affects them.