QA guides
The QA guides page builds a manual test guide for a PR: what to set up, what to test first, where the change must not appear, and what not to file as a bug. It is for a tester who has never seen the code.
Generate one
Section titled “Generate one”Enter a PR number on the QA page and click Generate. It needs your Claude account connected: the guide runs on your subscription, in its own worktree, and never writes to GitHub. Guides share the server’s one-heavy-job-at-a-time lock with reviews, so it may show waiting for another job to finish first. You can Stop it.
A QA build is a full agent run and now says what it cost. The model, the tokens and the duration are captured from the run and shown beside the guide, exactly as they are for a review — a guide that quietly burned an hour of someone’s subscription used to leave no trace anywhere. The usage line is written before the completeness checks below, so a run that timed out still accounts for what it spent.
The budget is sized from the diff. At most 5 changed files and 200 additions gets 25 minutes; at most 40 files and 2,000 additions gets 40; anything larger gets 60. RS_QA_TIMEOUT overrides it. A guide is heavier than a Deep review — it reads the diff, the review history and the surrounding code, then writes a long document — and the old hardcoded 25 minutes truncated large PRs.
A truncated guide is never published. The agent’s exit code is checked, a timeout fails with its own message, and the output must carry the end-of-guide marker, all three tier sections, at least one test case and the known-non-defects section before it is copied out. It used to be enough for the file to be non-empty, so a run killed at 60% arrived as “Guide ready — hand it to QA”. A failed build also sends the same notification card a failed review does, and its agent log is readable from the page.
What it reads
Section titled “What it reads”The guide is derived from evidence, not the PR description:
- the diff and the surrounding code (flags, settings, every surface that renders the change, adjacent surfaces that must not change);
- the PR conversation and inline review threads, which become the risk map;
- the branch’s commit history, since late fix-after-review commits mark the least-exercised paths;
- how a tester actually triggers the change in the target environment, including scheduled jobs and environment traps.
The agent has no network and no gh, so everything in that list that lives on GitHub is gathered by ReviewStage before the agent starts and written into the worktree as .rs-pr-context.md, which the prompt tells it to read first. This was previously a claim the code did not support: only the branch name, head SHA, title, URL and author were fetched, and the agent was left to improvise the rest.
Structure
Section titled “Structure”A guide is assembled from these parts, in this order. Two of them appear only when the change calls for them:
- Header — the PR number, any ticket IDs, the scope or pilot audience, and a plain-language line saying what the change is.
- Domain primer — only when understanding the feature needs a concept the tester does not have (a cut-off rule, a billing cycle, an inventory state machine), with a worked example. Most PRs do not get one.
- What changes on screen — stated plainly, including the case where nothing visibly changes and the real evidence is a log line, an export or a database row.
- Before you start — flags, settings, test data, accounts, viewports; exact names, in
code. - How to run it — only when the change is not triggered by ordinary clicking: a scheduled job, a queue worker, a webhook, an import.
- P0 · Test these first — failures that defeat the purpose of the change or silently corrupt what the user sees, plus anything the PR’s review history marked fragile.
- P1 · Does the feature work — the advertised behaviours.
- P2 · Check nothing else broke — each gate independently off, un-gated users see no change, shared components unchanged elsewhere.
- Surface matrix — a table of every surface the change could plausibly appear on, including the deliberate must not appear rows.
- Known — please don’t file these — intentional limitations and out-of-scope surfaces, sourced from the PR conversation.
- Footer — the one failure worth escalating immediately, and the branch head SHA the guide was checked against.
Cases are numbered continuously across P0, P1 and P2 (1…N) so a bug report can say “test 7”.
Every case is a - [ ] checkbox item carrying the data it needs, ordered steps, and an explicit
pass and fail.
Headless note: the skill’s own publishing step does not apply here. run-qa.sh runs the agent
with no Artifact tool and tells it to write GitHub-flavoured markdown to qa.md instead, which
is what the QA page renders. The skill itself has been rewritten so the markdown guide is the
deliverable and the Artifact path is an optional appendix for interactive use.
Output
Section titled “Output”GitHub-flavoured markdown, rendered on the page and copyable. Recent guides are listed with their PR titles, which are cached so they survive the PR leaving the queue.
MIT licensed · Built on Claude Code