86,317 review comments · 21,762 pull requests · 2011 to 2026 · 298 learned conventions
Large projects answer the same contributor questions forever. The answers exist, buried in review threads going back a decade, but they are not retrievable, and a maintainer's correction today does nothing for the person who asks the same thing next month.
Precedent reads a repository's entire review history, distils the durable conventions out of it, and says them back where the work happens: on a pull request, unprompted, cited to the specific pull requests the claim came from. A maintainer can correct it in place, and every later answer reflects that.
Target repository: pandas-dev/pandas.
Ask, correct, ask again. The rule the first answer used is retired, the correction replaces it, and the second answer cites the maintainer rather than the pull request it used to. Nothing was re-ingested and nothing was retrained.
See it happen on a real pull request.
Nobody addressed the agent. A pull request opened touching
pandas/core/groupby/groupby.py and doc/source/whatsnew/v3.0.0.rst, and it
read the changed paths, found the conventions anchored to those files, and
posted them:
From pandas-dev/pandas review history
Nobody asked me. I read the files this pull request changes and found three conventions the project has settled before, each linked to where it was settled.
- Always include a whatsnew entry in the appropriate file for any bug fixes or new features. Established in pandas-dev/pandas#64119 and #61985. Raised because it changes
doc/source/whatsnew/v3.0.0.rst.
One of the three it raised there is not a rule it inferred from the corpus at all. It is a convention a maintainer taught it earlier, through a pull request comment, and it now sits alongside conventions distilled from a decade of review threads.
Identity comes from GitHub, not from a login. Webhook deliveries are signed with
HMAC SHA-256, so sender.login is trustworthy without this application ever
handling a password, running an OAuth flow, or holding a session, and
author_association is GitHub's own answer to whether someone may speak for the
project.
What it does not do is check whether your diff complies. It names what the project has already decided and links where. Claiming a diff violates a convention means being right about the diff, and being wrong there costs more than being unhelpful.
The obvious rebuttal is that gpt-4o-mini has read a lot of pandas and might
answer these questions on its own. So the same 29 evaluation questions went to
both, and the answer is uncomfortable in one column and decisive in the other.
| correct | refusal | citations that resolve | |
|---|---|---|---|
| Precedent | 8/24 | recall 4/5, precision 4/7 | 43/43, 95% CI [92, 100] |
gpt-4o-mini alone |
9/24 | recall 0/5 | 0/24, 95% CI [0, 14] |
Memory did not make it more accurate. Both systems answered the same questions, so the comparison is paired and McNemar's exact test applies: they disagree on 3 of 24 and p = 1.00. Allowing partial credit it gets worse rather than better, because there is no question Precedent got right that the baseline got wrong. It also over-refuses, turning down 3 questions it could have answered, which is the difference between its refusal recall and its precision. By tag it is weakest on typing (0/3) and deprecation (0/2), best on process (4/8) and testing (3/6).
What it changed is whether the answer can be trusted. Five questions are deliberately unanswerable from review history. The baseline answered all five, confidently, inventing project policy on release schedules, governance and credentials. And of the 24 pull requests it cited, none resolve to a real discussion in the corpus: it is generating plausible five-digit numbers. Precedent cited 43 and every one resolves, because citations are verified against retrieved evidence before an answer is released and a failure suppresses the whole answer. Those two intervals do not overlap, which makes this the one comparison a sample of 29 can actually settle.
Memory does not make the model smarter. It makes it accountable, which is the difference that matters when a contributor cannot tell a confident right answer from a confident invented one.
Reproduce with python scripts/run_baseline.py, about a cent and a half, or
re-read the recorded run for nothing with python scripts/analyse_eval.py,
which is where every number above comes from. Raw output in
eval/baseline.json.
Four kinds of memory, one database. Not four services pretending to be one system.
| Store | Table | What it holds |
|---|---|---|
| Episodic | review_comments |
Individual review comments, embedded for semantic search |
| Semantic | rules |
Distilled repo conventions, with confidence and supersession history |
| Provenance | rule_evidence |
Which comments each rule was learned from |
| Entity | contributors |
Per-repo contributor state: volume, tenure, and areas touched |
| Working | sessions, session_turns |
The conversation, and which rules each answer used |
| Corrections | corrections |
What a maintainer corrected, and what it changed |
Every table is keyed on repo_id, and embeddings sit in the same tables as the
data they describe, so retrieval and the operational data can never disagree
about what the project said. How corrections retire a rule without deleting
it is the part worth reading next.
| CockroachDB | How it is used |
|---|---|
| Distributed vector indexing | idx_rules_embedding serves the agent's hot path. EXPLAIN confirms vector search: rules@idx_rules_embedding. Embeddings live beside the rows they describe, so there is no second store to keep consistent. |
| ccloud CLI | Provisioned and manages the serverless cluster in ap-south-1, next to the Lambda that queries it. |
| AWS | How it is used |
|---|---|
| Lambda | Serves the whole application, API, page and GitHub webhook, behind a Function URL. |
| S3 | Holds the raw ingested review history, 3,801 gzipped pages, staged before transform. |
| GitHub | How it is used |
|---|---|
| App webhooks | Signed deliveries make sender.login trustworthy with no login of any kind. author_association decides who may teach the memory. |
| Installation tokens | RS256 JWT exchanged for a short-lived token, so the grant is revocable by uninstalling rather than by rotating a key. |
pip install -e ".[db,api]"
python -m precedent.db.migrate --create-db
uvicorn precedent.api.app:app --port 8000Then open http://localhost:8000. COCKROACH_DSN and OPENAI_API_KEY come from
.env; see .env.example. Deploying it, and the settings that exist because the
demo is a public URL in front of a paid model, are in
docs/deployment.md.
| Architecture | Memory model, correction semantics, and why a model rather than a distance threshold decides what contradicts what |
| Failure modes | Every failure that actually happened, and what handles it |
| Deployment | Local, Lambda, and the two things Mangum's lifespan behaviour broke |
| Ingest and database | GraphQL staging, bulk embedding, CockroachDB driver notes |
MIT. See LICENSE.
