Repository navigation
Apply settings sent to /api/update_crawler_settings - #102
Open
amedipiran wants to merge 1 commit into
Open
amedipiran wants to merge 1 commit into
amedipiran wants to merge 1 commit into
Conversation
The endpoint never read request.get_json(). It re-pushed whatever the
settings manager already held and answered {'success': true}, so a client
posting {'maxUrls': 2000, 'crawlDelay': 1} received a success response and
then crawled with the defaults.
A silent no-op that reports success is worse than an error: a caller that
believes it capped a crawl at 2000 URLs has no reason to check, and the
crawl runs to the 5000000 default against someone else's server.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
/api/update_crawler_settingsnever readsrequest.get_json(). It re-pushes whatever the settings manager already holds and answers{"success": true}.So a client that posts
{"maxUrls": 2000, "crawlDelay": 1, "concurrency": 5}gets a success response, and then crawls with the defaults.I found this the slow way. I was driving LibreCrawl from a script, capped a crawl at 2000 URLs, got
{"success": true, "message": "Crawler settings updated"}, and watched it sail past 2000 and keep going against someone else's server.maxUrlswas still5000000.A silent no-op that reports success is worse than an error here. A caller who believes they capped a crawl has no reason to check.
This applies anything sent with the request via
save_settings, then does what it did before. An empty body behaves exactly as today, so existing callers are unaffected. The response lists which keys were applied, which makes the failure mode visible if a key is ever rejected.🤖 Generated with Claude Code