ci(pr-agent): fall back to Zhipu GLM-4.7-Flash instead of DeepSeek - #471
Conversation
Serialize workflow runs and sequential /improve chunks so the free-tier one-inflight limit does not 429; LiteLLM retries with backoff.
Reviewer's GuideThe PR switches PR-Agent’s last-resort model from DeepSeek to Zhipu GLM-4.7-Flash, wires in the new secret and fallback order, and limits workflow and chunk-level concurrency while enabling retries to respect the provider’s free-tier one-request limit. Sequence diagram for PR-Agent model fallback and retriessequenceDiagram
participant Trigger as PR-Agent Trigger
participant Workflow as GitHub Actions Workflow
participant Gemini as Gemini
participant GLM as Zhipu GLM-4.7-Flash
Trigger->>Workflow: Start PR-Agent job
Workflow->>Gemini: Request review or suggestions
alt Gemini succeeds
Gemini-->>Workflow: Response
else Gemini unavailable or rate-limited
Workflow->>GLM: Request using fallback_models
GLM-->>Workflow: Response
Workflow->>GLM: Retry with exponential backoff up to NUM_RETRIES=6
end
Workflow-->>Trigger: Publish PR-Agent result
File-Level Changes
Tips and commandsInteracting with Sourcery
Customizing Your ExperienceAccess your dashboard to:
Getting Help
|
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. Note Currently processing new changes in this PR. This may take a few minutes, please wait... ⚙️ Run configuration
📒 Files selected for processing (2)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
PR Reviewer Guide 🔍Here are some key observations to aid the review process:
|
|
Failed to generate code suggestions for PR |
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
PR-Agent v0.46 LiteLLM rejects the zhipuai/ provider prefix; route glm-4.7-flash through open.bigmodel.cn as openai/*.
|
Failed to generate code suggestions for PR |
Gemini then GLM-4.7-Flash; DeepSeek remains the paid last resort when both free models are unavailable.
PR Code Suggestions ✨Explore these optional code suggestions:
|
Both Gemini ids share one key; a quota miss on 3.8 would otherwise retry 3.7 before leaving Google. Keep DeepSeek last.
PR Code Suggestions ✨Explore these optional code suggestions:
|
Summary
glm-4.7-flash(zhipuai/glm-4.7-flash) so Gemini outages do not depend on a paid DeepSeek balance.concurrencygrouppr-agent) and turn off/improveparallel chunk calls (max_number_of_calls=3) to match the free-tier one-inflight limit.NUM_RETRIES=6for 429 exponential backoff.Test plan
ZHIPUAI_API_KEY(https://open.bigmodel.cn) before merging; remove unusedDEEPSEEK_API_KEYwhen readyzhipuai/glm-4.7-flashwithout DeepSeek/improveposts suggestions without empty output from chunkingSummary by Sourcery
Configure PR-Agent to use Zhipu GLM-4.7-Flash as its primary fallback while managing request concurrency and rate-limit retries.
Enhancements:
CI:
Summary by CodeRabbit