Skip to content

Apply settings sent to /api/update_crawler_settings - #102

Open
amedipiran wants to merge 1 commit into
PhialsBasement:mainfrom
amedipiran:fix/update-crawler-settings
Open

amedipiran wants to merge 1 commit into
PhialsBasement:mainfrom
amedipiran:fix/update-crawler-settings

Conversation

@amedipiran

Copy link
Copy Markdown

/api/update_crawler_settings never reads request.get_json(). It re-pushes whatever the settings manager already holds and answers {"success": true}.

So a client that posts {"maxUrls": 2000, "crawlDelay": 1, "concurrency": 5} gets a success response, and then crawls with the defaults.

I found this the slow way. I was driving LibreCrawl from a script, capped a crawl at 2000 URLs, got {"success": true, "message": "Crawler settings updated"}, and watched it sail past 2000 and keep going against someone else's server. maxUrls was still 5000000.

A silent no-op that reports success is worse than an error here. A caller who believes they capped a crawl has no reason to check.

This applies anything sent with the request via save_settings, then does what it did before. An empty body behaves exactly as today, so existing callers are unaffected. The response lists which keys were applied, which makes the failure mode visible if a key is ever rejected.

🤖 Generated with Claude Code

The endpoint never read request.get_json(). It re-pushed whatever the
settings manager already held and answered {'success': true}, so a client
posting {'maxUrls': 2000, 'crawlDelay': 1} received a success response and
then crawled with the defaults.

A silent no-op that reports success is worse than an error: a caller that
believes it capped a crawl at 2000 URLs has no reason to check, and the
crawl runs to the 5000000 default against someone else's server.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant