Skip to content

Architecture

Read this before changing anything. Most of the layout decisions look arbitrary and are not.

PieceRuns asDoes
bin/pr-watch.shpoller, every 3 minFinds PRs awaiting any signed-in user’s review → queue.json + a notification card. Never runs a review.
bin/server.pythe dashboard service, 127.0.0.1:8899Serves the React SPA shell and a JSON API: sessions, queue, PR pages, runs, posting, approving, skills, learnings, insights, QA.
bin/run-review.shspawned per click, detachedWorktree → claude -p with the chosen skill + output contract → review.json. Never writes to GitHub.
bin/run-qa.shspawned per clickSame shape, runs the QA-guide skill → qa.md.
bin/profile-repo.shspawned per click, or by the pollerDeterministic signals → one Sonnet call with skills/repo-profile → validated profiles/<slug>/profile.json + profile.md.
bin/rs_profile.pyimported + CLIThe profile: signals gathering, prompt, validation against the tree, risk-rule merge, the critical-path prompt block, markdown round-trip, versioning.
bin/rs_diff.pyimportedDiff-anchor validation so GitHub cannot 422 a whole review.
bin/rs_learn.pyimportedLearnings loop: scores kept/edited/dropped, renders recent decisions into the next prompt, scores skills.
bin/rs_agree.pyimportedMatches findings across independent runs; independence-weighted agreement.
bin/rs_rollup.pyimportedInsights aggregation.
bin/rs_md.pyimportedDependency-free markdown → HTML.
bin/rs_queue.pyimportedThe queue + dedup logic the poller and the webhook share: queue.json upsert / stale / done, seen keys, the review_requested payload and signed links (byte-identical to lib-common.sh).
bin/rs_webhook.pyimportedPOST /webhooks/github: X-Hub-Signature-256 verification, the event → queue mapping, webhooks.json. Notify-only, like the poller.
bin/rs_paths.pyimportedThe repository dimension: slugs, base_dir / prdir / udir, iter_prdirs, and the one-time legacy migration.
bin/lib-common.shsourcedConfig (REPOS, REPO_ALLOW_ORG), the bash twins of the path helpers, HMAC link signing, Slack posting (webhook or bot token).
bin/lib-limits.shsourcedMIN_FREE_MB and MIN_FREE_DISK_MB, defined once so lib-common.sh (the job scripts) and doctor.sh cannot disagree about the floor.
dashboard-ui/built onceReact 19 + TypeScript SPA: queue, PR page, stack page, QA, skills, learnings, insights, integrations, tour, command palette.
skills/pr-review/ · skills/pr-qa-guide/installedThe built-in review and QA-guide procedures, as Claude Code skills.

The backend is Python’s standard library plus bash, gh, jq, git, openssl and claude. There is no database; state is files.

browser ──HTTPS──▶ reverse proxy ──▶ server.py (127.0.0.1:8899)
├─ GET / SPA shell (React Router handles the rest)
├─ GET /api/* me · queue · pr?repo=&pr= · qa · skills · learnings · rollup?repo= · stack
├─ POST /api/review start_review → spawn run-review.sh (session leader)
├─ POST /api/post decrypt clicker's token → validate anchors → COMMENT review
├─ POST /api/approve LGTM comment → APPROVE, as the clicker
├─ POST /api/stop · /api/explain · /api/markdone · /api/archive · /api/skill/* · /api/claude/*
├─ GET /oauth/start · /oauth/callback · /health
└─ POST /webhooks/github GitHub → HMAC check → 202 → thread: rs_webhook.handle
GitHub ──HTTPS──▶ (same proxy; this path must be reachable by GitHub) ──▶ /webhooks/github

Every POST carries an HMAC token minted at render time (30 minutes) over action:owner/name#pr:expiry; the old action:pr:expiry form still verifies for RS_SIGNATURE_GRACE_DAYS after the upgrade. Pages are gated by the session cookie, not a signature, so they stay bookmarkable. A PR is addressed as ?repo=owner/name&pr=N; ?pr=N alone resolves when one repository is configured or when the number exists under exactly one repository’s state, and otherwise returns the candidate repos for a picker.

start_review records effort, focus, model and the clicker’s login, archives any previous run into history/<ts>/, then spawns run-review.sh as a new session leader (so Stop can kill the whole process group). The script:

  1. Clones the repository’s base clone if it does not exist yet (repos/<owner>__<name>, blobless), fetches the PR head and creates a git worktree per repo per reviewer per PR, so two people never collide.
  2. Picks the skill, in order: a per-repo override (skills/repos/<owner>__<name>/SKILL.md), else the clicker’s personal skill if selected, else the editable team default, else the installed built-in — and logs which one it used. It runs that skill’s logic and always appends an explicit review.json output contract, so any skill yields the shape the dashboard needs.
  3. Renders recent learnings and the focus note into the prompt, plus a depth instruction for the effort level.
  4. Runs claude -p headless with --allowedTools "Bash Read Glob Grep Write", an explicit --disallowedTools list, an optional --model, and a timeout ceiling by effort (a typical run finishes well inside it — a few minutes on Opus). The environment is built by agent_env: every GitHub, Slack, webhook and HMAC credential is stripped and GH_CONFIG_DIR points at an empty directory, so a gh call inside the agent fails unauthenticated. The clicker’s Claude token is the one credential kept, because the CLI reads it only from the environment.
  5. Counts the PR’s reviews, comments and threads with the service token before and after the run. A difference fails the job as a security incident — see Security.
  6. Parses token usage and the model from the stream log, records the head SHA into both head and meta.json (for staleness and the caches), records a path-based risk flag, validates review.json as an object with a comments array, copies it out, and removes the worktree and its branch through an EXIT trap that also fires on failure.

The prompt tells the agent a human will read explainer and analysis in a dashboard to decide whether to trust the findings. That framing is load-bearing: it is what makes the prose readable rather than a wall of bullets.

A run keyed on (login, head, effort, focus, model, skill) is cached; an identical request reuses the earlier result.

One install reviews many repositories (REPOS; REPO is a single-entry alias; REPO_ALLOW_ORG accepts any repo under an org on demand). Internally a PR is (repo, number). On disk a repository is the slug <owner>__<name> — GitHub owners cannot contain _, so the first __ is always the separator and the mapping is reversible:

ROOT/repos/<owner>__<name>/ base clone (blobless); review worktrees branch off it
ROOT/state/<owner>__<name>/<pr>/ per-PR state
ROOT/state/<owner>__<name>/<pr>/users/<login>/ one reviewer's run and markers
ROOT/skills/repos/<owner>__<name>/SKILL.md optional per-repo team default

rs_paths.py (repo_slug, slug_repo, base_dir, prdir, udir, iter_prdirs, repos_with_pr) and the same-named bash functions in lib-common.sh are the only code that builds these paths.

Migration. Installs from before this layout kept the clone at ROOT/repo and state at ROOT/state/<pr>. On start, migrate_legacy() moves both into the new layout once — when exactly one repository is configured — stamps repo onto queue.json rows, seen lines and learnings.jsonl rows, logs each step and writes ROOT/MIGRATED. Nothing is ever deleted; if a destination already exists the contents are merged file by file without overwriting. With several repositories configured and legacy state present the server refuses to start and prints the env to set, because the state cannot be attributed safely.

ROOT/state/<owner>__<name>/<pr>/ holds the shared PR identity (meta.json, with repo), QA files (qa.md, qa.status, qa_meta.json), the agreement indices and users/<login>/ — each reviewer’s run (review.json, status, effort, focus, skill, head, risk, pid, usage.json, logs, history/, cache/) and markers (opened, posted.json, payload.json, approved, archived). queue.json rows carry repo and requested: [logins], the union of everyone awaiting the PR, and the dashboard filters on both so one poller output serves every user.

meta.json exists because queue.json only holds PRs currently awaiting review; the moment you post, the PR leaves it, and without the cache the dashboard would lose the title of the review you just ran.

  • Posting defaults to a plain COMMENT review. REQUEST_CHANGES only when the reviewer ticks it on the post form: the agent produces notes, not a merge block, and the human approves separately.
  • Approve posts a comment and then approves. Pre-filled LGTM plus a checklist of blocker/should-fix findings, editable first.
  • run-review.sh never touches GitHub. This is the property that makes the whole thing safe to run against a real review queue.
  • Slack dedup is per repo + PR + login, not per head SHA. Each reviewer is pinged once per PR (<repo>:<pr>:<login> in seen); pushing new commits must not re-ping anyone. The dashboard reflects the live queue regardless.
  • One search per user per repo per poll, not a page-through. review-requested:<login> resolves server-side, so the poll costs one API call per user per repository (plus one org-wide search per user when REPO_ALLOW_ORG is set), and page loads never call GitHub.
  • Writes use the acting user’s token. The service token is never used to write: the review step has no GitHub write path, and every post and approval is made with the signed-in user’s own credentials. It is not read-only by permission — GitHub’s fine-grained model requires Pull requests: Read and write on it for this path to reach review threads — so the guarantee is structural, in the code, not enforced by the token’s scope.
  • The dashboard reads .env once. Deliberate: the process should not change behaviour under you because a file was edited; a restart is an explicit act.
  • Learnings (rs_learn.py): once a post has reached GitHub, each original finding is scored dropped/edited/kept and appended to learnings.jsonl as a short gist tagged with its repository, keyed by the run so a retry replaces its rows. The detail log is capped at 300 rows; learnings_totals.json beside it is a never-truncated tally and is where every all-time count comes from. render(repo) folds recent dropped (24) and edited (12) rows into the next prompt, same-repository rows first and the team’s general preferences after. Attributed per user.
  • Skills: the team default is a file in a small git repo on the server; every save is a commit with the editor’s login, and the page shows the log. save_skill refuses a blank; restore_global_skill needs a typed confirm. The commit’s exit code is checked and a failure is surfaced in the save banner rather than hidden, signing is disabled inside that repository, and each commit names only the paths it changed. Quick-add tidies a plain-English rule into a managed ## Team rules section, inserted at the end of that section wherever it sits. Each run records its skill id — including repo:<slug> for a per-repository override, which the dashboard can now actually save — and keep rate is per skill, from the tally — which excludes decisions recorded on a DRY_RUN post.
  • Agreement (rs_agree.py): findings from independent runs on the same head are matched; a finding is confirmed when raised by runs that differ in skill, model or effort.
  • Explain simply: a one-turn Haiku call on the clicker’s account, cached per finding content.
  • Stacked PRs: walks the chain of open PRs whose base is the previous head, on demand.
  • Staleness: the reviewed head SHA is recorded; the page flags a mismatch without re-running. A synchronize webhook refreshes the queue row’s head immediately; the poller does the same on its next pass.
  • Posting (posted_runs.json): the gate is per run, not per PR — one entry per post this dashboard made, naming the head and a content hash of the run — so a second review after a push posts, while the same run cannot post twice. The list keeps its last 20 entries and logs an eviction, since the run that drops off could in principle be posted again. The check, the POST and the write are held under one per-(repo, PR, user) lock, and the client echoes the run’s hash back so a stale tab gets a 409 instead of posting one finding’s text against another’s file and line. Approve takes the same lock, compares the reviewed head against the current one, and refuses a second approval at the same head.
  • Webhooks vs polling (rs_webhook.py, rs_queue.py): both feed queue.json and seen through one module and dedup on <repo>:<pr>:<login>, so a request that arrives twice (webhook, then poll) yields one card. Both now notify before writing seen, and both route draft, bot and over-age PRs to suppressed for re-evaluation instead. The poller merges into queue.json under the shared lock rather than replacing it, and each row records how every requested login was learned, so a poll — whose review-requested: search sees only direct requests — can no longer erase the rows a webhook built by expanding a team. A search that failed produces no verdict at all, and a cycle where nothing succeeded publishes nothing. The receiver replies 202, works on a bounded pool (RS_WEBHOOK_WORKERS, shedding with 503 when saturated), and drops any X-GitHub-Delivery it has already processed.
  • Suggestion blocks: a finding’s suggestion is appended to the body as a ```suggestion block on post.
  • Stop: the run’s pid is recorded; POST /api/stop kills the process group and writes stopped.
  • Handoff: a short signed token lets a session move between alias hostnames of the same server without re-login. Only the server’s own hosts are accepted.
  • Repository profile (rs_profile.py, profile-repo.sh): one profile per repository under $ROOT/profiles/<slug>/. run-review.sh merges its risk_paths into the banner rules and, for Standard and Deep runs, appends a “Critical paths for this repository” section listing only the paths the PR touches (cap 12, why cut to 200 chars) with their checks, the repo’s rules and do-not-flag list; findings may carry critical_path, which the card badges and learnings keep so Insights can report the kept rate on critical paths. The profile hash is part of the re-run cache key. GET/PUT /api/profile, POST /api/profile/run|stop (signed profile token, connected Claude required); POST /api/profile/auto is called by pr-watch.sh with a server-secret HMAC when auto_profile is on and the tree changed materially, and runs as the admin on the admin’s account. See Repository profile.
bin/ server, scripts, Python helpers
dashboard-ui/ React SPA (esbuild, node --test, Playwright)
skills/ built-in Claude Code skills (pr-review, pr-qa-guide) and the seeded team default
docs/ SETUP, ARCHITECTURE, SECURITY, OPERATIONS (source for this site)
website/ this site (Astro + Starlight)
docker-compose.yml, .env.example, bin/doctor.sh

MIT licensed · Built on Claude Code