Skip to content

Skills and learnings

skills · team default · keep rate
The Skills page: tabs for Which skill, Suggested rules, Editors, Per repository, Profiles and Depth; a selector between the team default and your own skill; and a “How each skill scores” card showing the team default’s keep rate. The Skills page: tabs for Which skill, Suggested rules, Editors, Per repository, Profiles and Depth; a selector between the team default and your own skill; and a “How each skill scores” card showing the team default’s keep rate.

A skill is the review procedure: a Claude Code skill file that tells the agent how to read the PR, what to check, and how to decide reply-vs-new. Learnings are your team’s recent accept/reject decisions, folded into the next prompt so the agent stops raising what you drop.

The Skills page has one selector:

  • Team default — a shared, editable skill, seeded on first start from skills/global-review.md in the repository and then owned by you. It is not the built-in pr-review skill: global-review.md is a separate, shorter document written to be edited by a team, while pr-review/SKILL.md is the built-in procedure the agent falls back to. Everyone’s reviews use the team default unless they opt out.
  • Your own skill — a skill you pasted in Integrations. Only your reviews use it.

Above both sits a per-repository override. There are three tiers on disk, and both the dashboard and run-review.sh resolve them in the same order:

TierPath under ROOT/skills/When it wins
Repository overriderepos/<owner>__<name>/SKILL.mdAlways, for that repository.
Your own skill<login>.mdWhen there is no repository override and your selector says own.
Team default_global.mdOtherwise. Falls back to the installed pr-review skill if it is missing.

The repository tier works now. Until this release the editor rendered fine and posted a target the server did not understand, so saving a repository override wrote over your personal skill — behind a green Saved your skill banner — while the repository file was never created and run-review.sh never saw an override. Clear override then deleted the personal skill. If you tried this before upgrading, check your personal skill in Integrations before you trust it. Personal and per-repository skills are written 0600.

ReviewStage runs the chosen skill’s logic and always appends its own output contract, so any Claude Code review skill works: the agent must end by writing a review.json with the assessment, explainer, analysis and a comments array of path, line, severity, body, reply_to, and an optional suggestion.

The team default is a file, edited in the browser. Guard rails:

  • It cannot be blanked. Saving an empty skill is refused.
  • Restore built-in writes the shipped skill back, and requires a typed confirmation. It used to delete _global.md instead, so runs silently fell back to a different document, the editor showed nothing, and the next container start re-seeded the file anyway.
  • Every save commits to a small git repository on the server (ROOT/skills), with the editor’s login as the commit author and a summary as the message; the Revision history panel reads that log. The save never fails on git — but the commit’s exit code is checked, and a failure is reported in the save banner and the server log rather than swallowed. It used to be a silent no-op: a global commit.gpgsign=true on the host was enough to leave every edit uncommitted while the UI just hid the panel. Signing is now disabled inside this repository, and each commit names only the paths that edit touched, so two people saving at the same moment are two commits by two authors. ROOT/skills is one of the directories that must be in your backup (Operations).

Type a preference in plain words — “don’t ask for a ticket link in code comments” — and it is tidied into a managed Team rules section. The rule is inserted at the end of that section, not the end of the file, so a skill whose Team rules are not the last heading no longer collects rules under whatever section happens to come last. You do not edit the whole file to add one rule.

Each review records which skill ran it. On post, each finding is scored kept / edited / dropped, tagged with that skill. How each skill scores shows a keep rate per skill.

Decisions made while DRY_RUN=1 are excluded from this rate, as they are from every rate — nothing reached GitHub — while still feeding the prompt block and the rule suggestions below.

This is the same keep rate the Insights page shows, on a per-skill population: (kept + edited) ÷ (kept + edited + dropped). A finding worth rewording was worth raising. The stricter measure — kept unchanged — is a separate number under a separate name, verbatimRate, so the two can no longer be quoted as if they were one. See Insights for the full definition.

Both surfaces read the never-truncated tally in learnings_totals.json, not the capped detail log, so a skill’s history does not evaporate as old rows fall off.

Below 20 scored decisions no rate is shown at all, on either page — one kept finding is not “100%”. It is a signal for improving the team default, not a leaderboard; a skill that produces many findings with a low keep rate is a skill that costs reviewers time.

Alongside the per-repository skill, the Skills page holds a repository profile: the paths where a mistake hurts most, the checks a reviewer must perform when a PR touches one, the repository’s risk paths and its rules. It is built once from deterministic signals plus one Sonnet call, validated against the tree, and editable in the same markdown editor. Standard and Deep reviews that touch a profiled path are told to walk it explicitly. See Repository profile.

On every post, each original finding is recorded as one of:

OutcomeMeaning
Kept as-isSelected and posted unchanged.
RewordedSelected, but the body was edited first.
Dropped as noiseNot selected.

It is recorded once, after the post actually reached GitHub. The outcomes used to be written before the POST, so a click with nothing ticked logged a full set of drops and three retries through an outage logged every finding four times. The call now happens on the path that succeeded, keyed by the run, and a repeat of the same run replaces its rows instead of appending another set. Four identical retries leave one set of decisions; a genuinely new review of a new commit leaves two.

A dry run is recorded, and flagged. With the shipped DRY_RUN=1 the post never reaches GitHub, but unticking a finding is still the reviewer’s real judgement about whether it was worth saying. Those rows are written with a dry flag: they feed the prompt block and the rule suggestions exactly like any other row, and they are kept out of every published rate and total — the keep rate, the verbatim rate, the per-skill scores, the outcome counts and the per-day series. They are counted on their own instead, as dryDecisions, and both pages say so: the Learnings page adds a Made in dry run tile, a note naming DRY_RUN=1, and a dry run pill on each affected row, and Insights carries the same explanation above its tiles so a pilot does not read as an install where nobody decided anything. A pilot run entirely on the default therefore shows no keep rate at all, rather than one computed from reviews that were never posted.

A short gist of each is appended to one learnings log for the whole install — ROOT/learnings.jsonl, not a file per repository. Each row records the repository it came from, and that log is capped at the most recent 300 rows across every repository.

Beside it sits learnings_totals.json, a never-truncated tally: outcomes by repository, by skill, by critical path and by UTC day. Every “all-time” count on Insights and Skills reads that file, so the numbers no longer freeze or go down as old detail rows fall off the cap. What the cap still governs is the per-finding detail — the recent-decisions list, the per-day keep series, and the prompt block below. An install upgrading into this release seeds the tally from whatever survives in the log and marks it incomplete.

When the next review starts, recent dropped and reworded rows are rendered into the prompt: “the team has recently rejected findings like these; do not raise them again unless the code makes them unavoidable.” The two windows are bounded separately — 24 dropped rows and 12 reworded — and the dashboard reads those numbers from the server rather than hardcoding them. Rows from the repository being reviewed come first, and rows from other repositories fill whatever room is left, so the steering is repository-preferring, not repository-scoped, and a new repository still benefits from the team’s general preferences.

This is not machine learning. It is in-context steering with your own recent decisions, shared per repository and attributed per user. The Learnings page shows the counts and the recent decisions so you can see what the agent is being told.

Repetition is slow, and it only ever learns from rejection. A reviewer reading a finding already knows whether it should be raised again, so each finding card carries a Teach the skill button next to Edit comment.

  1. Direction. The panel opens in the card and asks what the next review should do with this complaint: Don’t raise it again, or Always check it. Nothing is assumed from whether you ticked the finding — the clustering engine infers intent from repetition, this does not have to.
  2. Drafting. One Claude call on your own account, one turn, writes the rule in the house style of the rules already in the target skill, with a one-line rationale. With no Claude account connected the panel says so and offers the link to connect one, rather than a button that cannot work.
  3. Reading it. The drafted sentence lands in a box you can edit, above the name of the exact skill it will be written to. Nothing has been written yet. Redraft asks again.
  4. Adding it. Add this rule goes through the same quick-add path as a typed rule and an accepted suggestion: a bullet in ## Team rules, committed to the skills repository with you as the author. The same targeting applies — the repository’s own team default when it has one, otherwise the shared team default, and never a new per-repository override created from a single rule.

The promotion is recorded against the same vocabulary-keyed signature the clustering engine uses, which is the point of reusing it: a complaint you have taught is not offered back to you later as a fresh suggestion, and the card says Already a rule instead of offering the button again. It counts among the promoted rules on Insights like any other.

From a repeated rejection to a proposed rule

Section titled “From a repeated rejection to a proposed rule”

A rolling window forgets. A finding you dropped six times still arrives on review seven, because the last few dozen decisions are a preference, not a standard. So repetition is promoted deliberately.

  1. Clustering. Dropped rows — and separately reworded ones — are grouped into complaints. Two findings are the same complaint when they carry the same severity, sit under the same top-two directory segments, and their gists share vocabulary: stopwords dropped, words stemmed crudely, and at least half of the shorter gist’s significant words present in the other, with a floor of two shared words. It is cheap, dependency-free string work, in the same spirit as the agreement matching — coarse on purpose.
  2. Qualifying. A cluster becomes evidence at RULE_SUGGEST_MIN findings (default 3) from at least two different PRs. Three drops on one pull request is one bad day; three across three is a pattern. A cluster an existing Team rule already covers — checked with the same similarity function — is never offered again.
  3. Drafting. One Claude call, on the account of whoever opened the page, turns the cluster’s gists into a single imperative sentence in the house style of the rules already in the skill, plus a one-line rationale. It is cached against the cluster’s signature, so reopening the page costs nothing. With no Claude account connected, the cluster still appears with its evidence and simply has no drafted sentence.
  4. Deciding. The Suggested rules section at the top of the Skills page shows each proposal: the sentence, the rationale, the count in words (“from 4 findings you dropped across 3 PRs”), and an expandable list of the actual findings, each linking to its PR.
    • Accept appends it through the same quick-add path a hand-typed rule uses: a bullet in ## Team rules, committed to the skills repository with you as the author and the evidence count in the message. Cross-repository evidence goes to the team default; evidence from a single repository goes to that repository’s own team default, but only when one already exists — accepting a rule must never create an override that silently displaces the shared skill.
    • Dismiss records the cluster so it is not offered again. A dismissal is anchored to the row ids of the findings that were in the cluster when you dismissed it, with a same-directory, high-similarity fallback — it used to be anchored to a gist that drifts as the cluster grows, which both suppressed complaints nobody had dismissed and resurrected the one that had been. Dismissed suggestions stay behind a Show dismissed toggle with an Undo.

Drafting no longer happens on page load. Opening the Skills page used to spend up to two Claude calls of your quota with no consent, and a failure was not cached, so a bad token meant a fresh 90-second subprocess on every load with the error reaching nothing but stdout. A draft is an explicit action now, and a failure is cached with its reason and shown to you.

Nothing is ever written to a skill without a click. The model drafts; a person decides.

Once a cluster is promoted, its rows leave the rolling prompt block — the rule carries them now, and the window is spent on newer signal. A promotion records the row ids it was made over, so the answer does not depend on re-clustering the log again later; the outcome is part of a cluster’s identity, so a dropped-row cluster and a reworded-row cluster can no longer be merged into a third that matches neither. The block says so in its preamble. The Learnings page lists each cluster as a rolling preference or promoted to a rule, and Insights counts the rules promoted from evidence, so you can watch the memory harden instead of guessing.

When more than one reviewer runs the same commit, findings are matched across runs. A finding is confirmed when it was raised by runs with a different skill, model or effort; two runs of the same configuration do not confirm each other. The PR page shows ✓ N independent on such findings and an overall convergence rate. It is a signal to build on, not a score.

Both error directions were real and both are fixed. A body with fewer than two significant words used to match anything structurally nearby, so a one-word typo nit confirmed a null-check six lines away; a body too thin to judge is now an honest “not confirmed”, and two findings whose severities merely collapse into the same bucket must sit on the same line rather than within six of each other. In the other direction, two byte-identical file-level findings never confirmed each other, because a file-level finding has no line number; they now match on the same path, on the same token-overlap terms as any anchored pair.

MIT licensed · Built on Claude Code