Skip to content

ci(pr-agent): fall back to Zhipu GLM-4.7-Flash instead of DeepSeek - #471

Merged
buke merged 4 commits into
mainfrom
ci/pr-agent-glm-4.7-flash
Oct 3, 2026
Merged

buke merged 4 commits into
mainfrom
ci/pr-agent-glm-4.7-flash

Conversation

@buke

@buke buke commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Replace the DeepSeek last-resort fallback with Zhipu glm-4.7-flash (zhipuai/glm-4.7-flash) so Gemini outages do not depend on a paid DeepSeek balance.
  • Serialize PR-Agent jobs (concurrency group pr-agent) and turn off /improve parallel chunk calls (max_number_of_calls=3) to match the free-tier one-inflight limit.
  • Set LiteLLM NUM_RETRIES=6 for 429 exponential backoff.

Test plan

  • Add repo secret ZHIPUAI_API_KEY (https://open.bigmodel.cn) before merging; remove unused DEEPSEEK_API_KEY when ready
  • Open or push this PR and confirm Gemini still runs first; force fallback (or wait for Gemini 429) and check logs for zhipuai/glm-4.7-flash without DeepSeek
  • Two overlapping PR-Agent triggers queue instead of overlapping (no cancel-in-progress)
  • /improve posts suggestions without empty output from chunking

Summary by Sourcery

Configure PR-Agent to use Zhipu GLM-4.7-Flash as its primary fallback while managing request concurrency and rate-limit retries.

Enhancements:

  • Use Zhipu GLM-4.7-Flash as the primary fallback before Gemini and DeepSeek alternatives.
  • Serialize PR-Agent workflow runs and limit improvement suggestion calls to accommodate the fallback model’s free-tier request limits.
  • Increase LiteLLM retries to improve resilience to rate limiting.

CI:

  • Queue concurrent PR-Agent workflow triggers instead of cancelling or overlapping them.

Summary by CodeRabbit

  • Chores
    • Automated pull request review runs are now serialized, and in-progress runs are no longer cancelled when another run starts.
    • Review processing can fall back across three AI models, with up to six retries.
    • Code suggestions are generated sequentially, with a limit of three calls per review.

Serialize workflow runs and sequential /improve chunks so the free-tier one-inflight limit does not 429; LiteLLM retries with backoff.

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry @buke, you've used your own review budget of 250,000 diff characters for the last 7 days.

You can request another review in 1 day and 19 hours by commenting @sourcery-ai review. Upgrade to get a review now.

@sourcery-ai

sourcery-ai Bot commented Oct 3, 2026

Copy link
Copy Markdown

Reviewer's Guide

The PR switches PR-Agent’s last-resort model from DeepSeek to Zhipu GLM-4.7-Flash, wires in the new secret and fallback order, and limits workflow and chunk-level concurrency while enabling retries to respect the provider’s free-tier one-request limit.

Sequence diagram for PR-Agent model fallback and retries

sequenceDiagram
    participant Trigger as PR-Agent Trigger
    participant Workflow as GitHub Actions Workflow
    participant Gemini as Gemini
    participant GLM as Zhipu GLM-4.7-Flash

    Trigger->>Workflow: Start PR-Agent job
    Workflow->>Gemini: Request review or suggestions
    alt Gemini succeeds
        Gemini-->>Workflow: Response
    else Gemini unavailable or rate-limited
        Workflow->>GLM: Request using fallback_models
        GLM-->>Workflow: Response
        Workflow->>GLM: Retry with exponential backoff up to NUM_RETRIES=6
    end
    Workflow-->>Trigger: Publish PR-Agent result
Loading

File-Level Changes

Change Details Files
Replace the paid DeepSeek fallback with Zhipu GLM-4.7-Flash and configure its credentials and fallback ordering.
  • Expose the repository Zhipu API key to the action.
  • Use Gemini models first, followed by zhipuai/glm-4.7-flash as the last fallback.
  • Update model-specific comments and remove DeepSeek configuration references.
.github/workflows/pr-agent.yml
Serialize PR-Agent execution and reduce concurrent model requests to accommodate the fallback provider’s free-tier limits.
  • Queue workflow runs in a shared pr-agent concurrency group without cancelling active runs.
  • Disable parallel /improve chunk calls and cap chunk requests at three.
  • Apply the same sequential-call and request cap settings in both workflow environment overrides and project configuration.
.github/workflows/pr-agent.yml
.pr_agent.toml
Add retry behavior for transient rate limiting from the fallback provider.
  • Set LiteLLM NUM_RETRIES to 6 for exponential backoff on HTTP 429 responses.
.github/workflows/pr-agent.yml

Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@coderabbitai

coderabbitai Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Note

Currently processing new changes in this PR. This may take a few minutes, please wait...

⚙️ Run configuration
  • Configuration used: defaults
  • Review profile: CHILL
  • Plan: Essentials
  • Run ID: 982f1c6a-09c6-4c71-a921-9c1070de423f
📥 Commits

Reviewing files that changed from the base of the PR and between 9c86585 and bee8bb0.

📒 Files selected for processing (2)
  • .github/workflows/pr-agent.yml
  • .pr_agent.toml
 _________________________________________________________________
< You sanitized... the wrong thing. That's impressively specific. >
 -----------------------------------------------------------------
  \
   \   (\__/)
       (•ㅅ•)
       /   づ
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown

PR Reviewer Guide 🔍

Here are some key observations to aid the review process:

⏱️ Estimated effort to review: 1 🔵⚪⚪⚪⚪
🧪 No relevant tests
🔒 No security concerns identified
✅ No TODO sections
⚡ No major issues detected

@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown

Failed to generate code suggestions for PR

@codecov

codecov Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ All tests successful. No failed tests found.

📢 Thoughts on this report? Let us know!

PR-Agent v0.46 LiteLLM rejects the zhipuai/ provider prefix; route glm-4.7-flash through open.bigmodel.cn as openai/*.
@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown

Failed to generate code suggestions for PR

Gemini then GLM-4.7-Flash; DeepSeek remains the paid last resort when both free models are unavailable.
@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown

PR Code Suggestions ✨

Explore these optional code suggestions:

CategorySuggestion                                                                                                                                    Impact
Performance
Reorder fallback models to reach GLM sooner

Both Gemini ids share the same GEMINI_API_KEY, so an account-level quota or provider
outage makes the second Gemini hop fail too, burning up to 6 backoff retries before
GLM is ever attempted. Place openai/glm-4.7-flash directly after the primary so a
Gemini outage reaches the GLM (and then DeepSeek) fallback with far less wasted
latency.

.github/workflows/pr-agent.yml [45]

-          config.fallback_models: '["gemini/gemini-3.7-flash", "openai/glm-4.7-flash", "deepseek/deepseek-v4-flash"]'
+          config.fallback_models: '["openai/glm-4.7-flash", "gemini/gemini-3.7-flash", "deepseek/deepseek-v4-flash"]'
Suggestion importance[1-10]: 5

__

Why: Reasonable point that both gemini/* ids share the same GEMINI_API_KEY, so an account-level Gemini outage wastes retries on the second Gemini hop before reaching GLM. The reorder is coherent and the improved_code correctly reflects it, though it is an opinionated latency/robustness trade-off rather than a bug fix.

Low
Possible issue
Scope concurrency per pull request

A single global concurrency group serializes reviews for every PR, but GitHub keeps
only the newest pending run per group and silently cancels older queued ones, so
pushes/comments on other PRs can go unreviewed (a canceled pending run can also
leave a required check unfilled). Scope the group per PR so only rapid re-triggers
of the same PR coalesce, and let NUM_RETRIES absorb the occasional 429.

.github/workflows/pr-agent.yml [11-13]

 concurrency:
-  group: pr-agent
+  group: pr-agent-${{ github.event.pull_request.number || github.event.issue.number }}
   cancel-in-progress: false
Suggestion importance[1-10]: 4

__

Why: The observation that a repo-wide concurrency group queues all PR runs and that older pending runs get cancelled is factually correct, but the fix contradicts the PR's explicit intent (the added comment states runs should be queued repo-wide to avoid GLM 429s / 1-concurrent limits). Scoping per PR re-enables cross-PR overlap, so this is a debatable trade-off rather than a clear improvement.

Low
Align retry count with AI timeout

NUM_RETRIES is applied per LiteLLM call, so with config.ai_timeout: "180" and up to
3 sequential chunk calls, 6 exponential-backoff retries can push a single chunk past
its timeout window and yield empty completions instead of failing fast. A smaller
value (e.g. 3) keeps retries inside the timeout budget, or raise config.ai_timeout
to match the retry window.

.github/workflows/pr-agent.yml [54]

-          NUM_RETRIES: "6"
+          NUM_RETRIES: "3"
Suggestion importance[1-10]: 3

__

Why: The claim that NUM_RETRIES: "6" will overrun config.ai_timeout: "180" is speculative and depends on unstated backoff/PR-Agent internals; the suggested NUM_RETRIES: "3" is an arbitrary tuning change with no demonstrated failure in the diff.

Low

Both Gemini ids share one key; a quota miss on 3.8 would otherwise retry 3.7 before leaving Google. Keep DeepSeek last.
@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown

PR Code Suggestions ✨

Explore these optional code suggestions:

CategorySuggestion                                                                                                                                    Impact
Performance
Reduce global retry count

NUM_RETRIES is global to LiteLLM, so Gemini and DeepSeek calls also retry six times.
With config.ai_timeout: "180", exponential backoff can consume the timeout budget
and delay fallback to the next provider; consider NUM_RETRIES: "3" or raise
ai_timeout to cover the worst case.

.github/workflows/pr-agent.yml [56]

-          NUM_RETRIES: "6"
+          NUM_RETRIES: "3"
Suggestion importance[1-10]: 4

__

Why: A reasonable performance consideration since NUM_RETRIES is global to LiteLLM, but the value 6 was intentionally chosen for GLM backoff, and the suggestion is only a trade-off rather than a clear bug fix.

Low
Possible issue
Harden global concurrency queue behavior

GitHub retains only one pending run per concurrency group and cancels older pending
runs, so the static pr-agent group can silently drop queued reviews for other PRs,
and a single stuck run blocks the whole queue. If that trade-off is not intended,
use cancel-in-progress: true so the latest event supersedes stale runs, or bound the
job with timeout-minutes.

.github/workflows/pr-agent.yml [11-13]

 concurrency:
   group: pr-agent
-  cancel-in-progress: false
+  cancel-in-progress: true
Suggestion importance[1-10]: 3

__

Why: The PR comment explicitly documents the intent to queue runs rather than cancel them (cancel-in-progress: false), so the improved_code directly contradicts the deliberate design. It raises a valid informational caveat about GitHub keeping only one pending run, but the suggested fix is counter to the PR's stated goal.

Low
Reconsider fallback provider ordering

This order sends every Gemini 3.8 failure through GLM first, including
model-specific 3.8 outages where gemini-3.7-flash would be a faster, same-provider
fallback. If account-level Gemini quota exhaustion is not the dominant failure mode,
put gemini/gemini-3.7-flash immediately after the primary and keep GLM next;
otherwise the current order is a deliberate quota-miss trade-off.

.github/workflows/pr-agent.yml [47]

-          config.fallback_models: '["openai/glm-4.7-flash", "gemini/gemini-3.7-flash", "deepseek/deepseek-v4-flash"]'
+          config.fallback_models: '["gemini/gemini-3.7-flash", "openai/glm-4.7-flash", "deepseek/deepseek-v4-flash"]'
Suggestion importance[1-10]: 3

__

Why: The PR comment documents exactly why GLM precedes gemini-3.7-flash (a Gemini account-level quota miss), so the reordering contradicts the deliberate rationale; the suggestion even concedes the current order may be intentional.

Low
Preserve large-PR suggestion coverage

With parallel_calls already false, PR-Agent issues chunk calls sequentially, so
lowering max_number_of_calls from 5 to 3 does not prevent concurrent GLM requests
and mainly reduces suggestion coverage for large PRs. If RPM is the constraint,
prefer the retry/backoff settings or a larger job timeout, and sync the value in
.pr_agent.toml.

.github/workflows/pr-agent.yml [59]

-          pr_code_suggestions.max_number_of_calls: "3"
+          pr_code_suggestions.max_number_of_calls: "5"
Suggestion importance[1-10]: 2

__

Why: The suggestion to revert to 5 contradicts the PR's explicit intent of limiting calls to avoid GLM 429s, and its claim that .pr_agent.toml needs syncing is inaccurate since both files were already set to 3.

Low

@buke
buke merged commit c4381a7 into main Oct 3, 2026
45 of 46 checks passed
@buke
buke deleted the ci/pr-agent-glm-4.7-flash branch October 3, 2026 14:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant