Skip to content

review-pr: harden Pre-Verdict Audit against comment-quality halo effect - #56

Draft
warp-agent-staging[bot] wants to merge 1 commit into
eval-base/019ff1e1from
factory/review-pr-audit-halo-effect
Draft

review-pr: harden Pre-Verdict Audit against comment-quality halo effect#56
warp-agent-staging[bot] wants to merge 1 commit into
eval-base/019ff1e1from
factory/review-pr-audit-halo-effect

Conversation

@warp-agent-staging

Copy link
Copy Markdown
Contributor

What

Hardens the Comments bullet of ## Pre-Verdict Audit in .agents/skills/review-pr/SKILL.md.

Why

A reviewer using this skill let a comment's writing quality and technical accuracy ("well-written, explains a genuinely hard bug") stand in for actually checking it against the target repository's commenting guidelines, and missed confirmed rule violations as a result. The old wording ("check every comment ... against the repository's own commenting guidelines") was satisfiable by a holistic impression, which is exactly how a good-looking comment gets skimmed past.

Change

The bullet now:

  • requires enumerating every added or changed comment individually, with its file:line, so none is absorbed into a general impression of the diff;
  • gives a fallback standard for repositories with no written commenting guidelines — the commenting distribution of the existing code (density, tone, what existing comments explain vs. omit);
  • states explicitly that a comment's writing quality, technical accuracy, or the importance of the issue it describes never excuses a violation of an applicable guideline or a clear mismatch with the codebase's norms.

Scope

Documentation only, one line in one section. No specific rule names or categories are introduced, so the skill stays generic across repositories with and without explicit commenting guidelines. The schema, severity labels, safety rules, evidence rules, suggestion-block constraints, and diff-line-annotation contract are untouched.

Conversation: https://staging.warp.dev/conversation/fe82fcbf-bc05-48e9-8d3e-6416a0d87098
Run: https://oz.staging.warp.dev/runs/019ff4a3-1c98-74c9-8eaf-1a02ca29cac2

This PR was generated with Oz.

The Pre-Verdict Audit's comment check allowed a holistic impression to
stand in for a per-comment check, so a well-written comment could be
skimmed past without being measured against the repository's commenting
guidelines. Require enumerating each added or changed comment with its
file:line, give a fallback standard for repos without written guidelines,
and state explicitly that writing quality, technical accuracy, or the
importance of the issue a comment describes never excuses a violation.

Co-Authored-By: Warp <agent@warp.dev>
@warp-agent-staging warp-agent-staging Bot added the factory-ab-eval Factory A/B evaluation replay label Aug 12, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

factory-ab-eval Factory A/B evaluation replay

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant