Repository navigation
fix(engines): repair DeepSeek + Perplexity scans, clarify quota failures - #115
Merged
Merged
Conversation
DeepSeek and Perplexity failed on every run. Two distinct causes: - DeepSeek retired the `deepseek-chat` alias; GET /models now serves only deepseek-v4-flash and deepseek-v4-pro, and the old name 400s. Default to flash (matches this engine's "quick, lightweight" billing), overridable via DEEPSEEK_MODEL. The 8K output cap was a deepseek-chat limit — V4 accepts 64K and is a reasoning model whose reasoning_content draws from the same budget, so 8K would truncate the JSON report. - Perplexity's key and model are fine; the shared 90s client timeout was too tight. Sonar does its web retrieval before emitting response headers, and a real audit measures ~221s, so every run died as "Request timed out." Give it 5min x 2 attempts and widen its stuck window to match. Timeouts are now per-engine rather than hardcoded in oa-compat. The SDK clears its timer as soon as fetch() resolves, so on a streaming call it bounds time-to-first-byte only — a provider that opens the stream then stalls was unbounded. Added an idle watchdog that aborts on a gap between chunks, which is what makes raising Perplexity's header timeout safe. OpenAI GPT-5 Mini and Sakana Fugu are not code bugs — both accounts are out of prepaid credit (verified: OpenAI insufficient_quota; Sakana "Prepaid credit balance is exhausted"). Both surfaced as generic rate limits, pointing at a per-minute cap that was never the problem, and Fugu's message hardcoded "after 5 attempts" while maxRetries was 2. They now distinguish out-of-credit from throttling and quote the provider. Verified live against crawlproof.com: DeepSeek 84/100 in 57s, Perplexity 93/100 in 3m41s. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
vu1nz Security Review0 finding(s) in PR #? No security issues found. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This work was committed on this branch earlier but no PR was ever opened, so it never reached
master— meaning prod is still sending the retireddeepseek-chatalias and every DeepSeek scan 400s with:What it fixes
DeepSeek — the
deepseek-chatalias was retired;GET /modelsnow serves onlydeepseek-v4-flashanddeepseek-v4-pro. Defaults to flash (matches this engine's "quick, lightweight" billing), overridable viaDEEPSEEK_MODEL. The old 8K output cap was adeepseek-chatlimit — V4 accepts 64K and is a reasoning model whosereasoning_contentdraws from the same budget, so 8K would truncate the JSON report.Perplexity — key and model were fine; the shared 90s client timeout was too tight. Sonar does its web retrieval before emitting response headers and a real audit measures ~221s, so every run died as "Request timed out." Now 5min × 2 attempts with a matching stuck window.
Timeouts are now per-engine rather than hardcoded in
oa-compat. The SDK clears its timer as soon asfetch()resolves, so on a streaming call it bounded time-to-first-byte only — a provider that opens the stream then stalls was unbounded. An idle watchdog now aborts on a gap between chunks, which is what makes raising Perplexity's header timeout safe.OpenAI GPT-5 Mini and Sakana Fugu are not code bugs — both accounts are out of prepaid credit (verified: OpenAI
insufficient_quota; Sakana "Prepaid credit balance is exhausted"). Both surfaced as generic rate limits, pointing at a per-minute cap that was never the problem, and Fugu's message hardcoded "after 5 attempts" whilemaxRetrieswas 2. They now distinguish out-of-credit from throttling and quote the provider.Verification
Verified live against crawlproof.com when originally written: DeepSeek 84/100 in 57s, Perplexity 93/100 in 3m41s.
Re-verified now on top of the Slop Score merge (
6c1a886): merges cleanly with no conflicts,tsc --noEmitclean, 473 tests pass. Confirmed prod Railway env has noDEEPSEEK_MODELset, so the newdeepseek-v4-flashdefault takes effect on deploy.🤖 Generated with Claude Code