Skip to content

AI review: run two models on each unit and post one comment each - #3070

Closed
marcocastignoli wants to merge 5 commits into
ci/ai-review-collectfrom
ci/ai-review-hello-world
Closed

marcocastignoli wants to merge 5 commits into
ci/ai-review-collectfrom
ci/ai-review-hello-world

Conversation

@marcocastignoli

@marcocastignoli marcocastignoli commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Step 4 of #3057, stacked on #3061: the review step and the comment, end to end. Marco's review points on #3061 and #3069 are in here too.

The review job, after the gate and the collect step:

  • Reads the pull request title, description and discussion through the API into a JSON file, and the collector adds them to every unit as pullRequest (bot comments removed, long texts cut). The prompt treats them as claims to check against the source.
  • Sends every unit to two models, one request each, one turn, no tools: Claude Sonnet 5.5 at low effort (Messages API) and GPT-6 Luna at xhigh (Responses API). The prompt and the trimmed spec come from the base branch as the system prompt; the unit is the user message inside a nonce tag. Both run even if one fails. Two models for now, to compare them on real pull requests; one will stay.
  • Uploads the answers, the token usage and the list-price cost as the artifact ai-review-answers.

The prompt moves to docs/ai-review/REVIEW_PROMPT.md. It gains the pullRequest input, a fourteenth check other with the note that the list is not complete, and a Markdown answer format with fixed sections: summary, Critical, Warning, Info, and "What could not be reviewed" only when something limited the review; each finding a ### block with Check, Where, Why, Evidence, Fix. The JSON schema is gone.

The post job holds pull-requests: write and no key. It renders one comment per model, headed as AI-generated and advisory, with the cost line and the review collapsed unless it has a critical finding, and updates the comment on later runs by its marker. Every answer is untrusted: prose lines are escaped, links, mentions and issue references broken, headings demoted, code fences closed; fenced code is left as written since GitHub renders it literally.

Secrets the Environment ai-review needs: ANTHROPIC_API_KEY and OPENAI_API_KEY.

Tested locally: dry runs on benchmark inputs, both providers' authentication errors, the renderer on a hostile fake answer (mention, links, image tag, script, a </details> inside a fence, an unclosed fence). Not yet run on the fork: it needs the two keys in the fork's Environment first. The benchmark of #3069 was measured with the JSON prompt; a re-run with the Markdown prompt is pending.

🤖 Generated with Claude Code

@github-actions github-actions Bot added the ci Changes to continuous integration label Sep 30, 2026
@marcocastignoli
marcocastignoli force-pushed the ci/ai-review-hello-world branch 2 times, most recently from 8d77861 to fafe93e Compare October 1, 2026 07:32
@marcocastignoli
marcocastignoli force-pushed the ci/ai-review-hello-world branch 2 times, most recently from 5aa4f7e to 7a43a9f Compare October 1, 2026 07:44
marcocastignoli and others added 5 commits October 1, 2026 09:44
The workflow reads the title, the description, the comments, the reviews and
the review comments of the pull request through the API into a JSON file,
never on a shell line, and the collector adds them to every review unit as
`pullRequest`, bot comments removed, long texts cut, the latest sixty
comments kept. They tell the model what the author meant and what reviewers
asked; the prompt treats them as claims to check, not as evidence.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The review job sends every unit to two models, one request each, one turn,
no tools: Claude Sonnet 5.5 at low effort through the Messages API and GPT-6
Luna at xhigh through the Responses API, with the prompt and the trimmed
specification from the base branch as the system prompt and the unit as the
user message inside a nonce tag. Two models run for now to compare them on
real pull requests; one will stay. The answers, the token usage and the
list-price cost are the artifact ai-review-answers.

The model answers in Markdown with fixed sections, critical findings first;
the prompt gains a fourteenth check, `other`, for what the list does not
name, and moves to docs/ai-review/REVIEW_PROMPT.md.

A second job, post, holds pull-requests: write and no key. It renders one
comment per model, headed as AI-generated and advisory with its cost, and
updates it on later runs. Every answer is untrusted: prose is escaped,
links, mentions and issue references are broken, headings demoted, code
fences closed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The answer's last section, "What was reviewed", listed the functions and
files behind every review. It becomes "What could not be reviewed", present
only when a dropped or omitted file, a missing source or an unresolved proxy
limited the review, as Marco asked on the docs page. The runner requires the
three severity sections only.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The first run on the fork failed with "This API key is not scoped to a
workspace, so this request must include the anthropic-workspace-id header".
An optional ANTHROPIC_WORKSPACE_ID secret sets that header; a key created
inside a workspace needs nothing.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@marcocastignoli
marcocastignoli force-pushed the ci/ai-review-hello-world branch from 7a43a9f to f82464d Compare October 1, 2026 07:44
@marcocastignoli
marcocastignoli deleted the ci/ai-review-hello-world branch October 1, 2026 07:45
@marcocastignoli marcocastignoli changed the title AI review: hello world end to end, two models, one comment each AI review: run two models on each unit and post one comment each Oct 1, 2026
@marcocastignoli

Copy link
Copy Markdown
Member Author

Closed by GitHub when the branch was renamed to ci/ai-review-run. The same commits continue in #3081.

Posted with Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci Changes to continuous integration

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

1 participant