diff --git a/README.md b/README.md
index a683958..bf13af4 100644
--- a/README.md
+++ b/README.md
@@ -94,7 +94,7 @@ The notification bell reports a drop of at least five positions from a previous
Open **Listing history → Track a listing**. Collection is off by default. Select an app, store, country, and frequency; Google Play also supports a language choice. The credit estimate appears before you enable tracking. Each scheduled check uses one SerpApi product request, with an initial baseline check when tracking starts. Manual checks and retries can consume additional credits.
-The timeline highlights changes to the fields returned by the store, including titles, descriptions, versions, pricing, and images. Text comparisons highlight additions and removals. Screenshots can be compared in order, with added and moved images labeled. Supported images are archived locally in the database and included in backups. If an image cannot be archived, the comparison shows an unavailable-image placeholder. AppTrail never substitutes the live image for a missing historical image.
+The timeline highlights changes to the fields returned by the store, including titles, descriptions, versions, pricing, and images. Text comparisons highlight additions and removals. Screenshots can be compared in order, with added and moved images labeled. Supported images are archived locally in the database. Downloaded backups omit archived images and raw SerpApi responses to save space, while keeping saved results, matched evidence, listing text, and image-change records. After restoring, AppTrail explains why omitted images and raw responses are unavailable. It never substitutes a live image for a missing historical image.
Use **Manage** to change frequency, pause or resume collection, or request a manual check. Pausing preserves history. Unchanged checks are hidden until you select **Show unchanged checks**. Collection requires the AppTrail process to remain running, and begins when you enable it; earlier listing versions cannot be reconstructed.
diff --git a/docs/AUTHENTICATION.md b/docs/AUTHENTICATION.md
index 6a95d51..f2043ef 100644
--- a/docs/AUTHENTICATION.md
+++ b/docs/AUTHENTICATION.md
@@ -12,6 +12,8 @@ The server generates a random one-time setup code, stores it in a permission-pro
Choose a username of 3–64 letters, numbers, dots, underscores, or hyphens, starting with a letter or number. Usernames are case-insensitive. Passwords require 8–128 characters, including at least one number (0–9) and one special character, such as `!` or `@`. Spaces are preserved but do not count as special characters.
+If you already have a SQLite backup, select **Have a backup? Restore your data** below **Create account**. Enter this server's setup code and upload the backup. After restoration, sign in with the username and password saved in the backup. This option is available only before an owner account exists.
+
## Sessions and requests
A successful login issues a random cookie with `HttpOnly`, `SameSite=Strict`, and a host-only scope. Requests that AppTrail sees as HTTPS also use `Secure` and the `__Host-` cookie prefix. A proxy can forward the original scheme through the optional [forwarded-header settings](SELF_HOSTING.md#public-https-hosting). Sessions are bound to the scheme, host, and port seen by AppTrail when signing in. SQLite stores a hash of the session token, not the token itself. Tokens are not stored in browser local storage.
@@ -20,10 +22,12 @@ Sessions last 24 hours from sign-in, including time spent away from the app. The
Every workspace route requires a valid session, including reads, searches, discovery, key changes, exports, backups, and the dashboard HTML. Only the login/setup page, its static assets, authentication status and entry endpoints, and the minimal health probe are public. Data responses use `Cache-Control: no-store`.
-Writes require JSON, a custom request header, and a CSRF token tied to the authenticated session. Login and setup require JSON and the custom header before a session exists. Requests marked `cross-site` by the browser's `Sec-Fetch-Site` header are rejected, and no cross-origin access is enabled through CORS. AppTrail does not compare the browser's `Origin` header with the internal server address, so an HTTPS proxy can forward HTTP without blocking account setup or login. Login, setup, and password changes share a persisted limit of ten attempts per client address per five minutes and a global limit of 100 per minute. See the proxy instructions before forwarding client addresses.
+Writes require JSON, a custom request header, and a CSRF token tied to the authenticated session. SQLite restoration accepts a binary upload with `Content-Type: application/vnd.sqlite3` and `X-AppTrail-Confirm-Restore: overwrite`, with the same session, custom header, and CSRF checks. Login and setup require JSON and the custom header before a session exists. Requests marked `cross-site` by the browser's `Sec-Fetch-Site` header are rejected, and no cross-origin access is enabled through CORS. AppTrail does not compare the browser's `Origin` header with the internal server address, so an HTTPS proxy can forward HTTP without blocking account setup or login. Login, setup, and password changes share a persisted limit of ten attempts per client address per five minutes and a global limit of 100 per minute. See the proxy instructions before forwarding client addresses.
These controls follow the [OWASP password storage](https://cheatsheetseries.owasp.org/cheatsheets/Password_Storage_Cheat_Sheet.html), [session management](https://cheatsheetseries.owasp.org/cheatsheets/Session_Management_Cheat_Sheet.html), and [CSRF prevention](https://cheatsheetseries.owasp.org/cheatsheets/Cross-Site_Request_Forgery_Prevention_Cheat_Sheet.html) guidance. They are covered by automated tests and browser checks, not an independent security assessment. There is no MFA or SSO in this version.
+Before account creation, `POST /api/auth/restore` accepts the same SQLite upload and overwrite confirmation with `X-AppTrail-Setup-Token` containing the server's setup code, in place of a session and CSRF token. It requires the custom request header and rejects cross-site requests. Attempts share the setup and login rate limits. AppTrail checks that no owner exists both before accepting the upload and immediately before restoring, then removes the setup code after success.
+
## API clients
Use an HTTP client with a cookie jar. Submit JSON to `POST /api/auth/login` with `username` and `password`, plus `X-AppTrail-Request: 1` and `Content-Type: application/json`. Retain the returned cookie and `csrf_token`. Send the cookie on subsequent requests; writes also need `X-CSRF-Token` with that value and the same JSON/custom headers. `GET /api/auth/status` returns the current session's CSRF token. Never place passwords or session tokens in URLs.
@@ -32,6 +36,8 @@ HTTP 401 means authentication is required. HTTP 403 indicates a failed setup cod
## Recovery and background work
+Restoring a SQLite backup in Settings replaces the owner account and password with those saved in the upload. It revokes all current and uploaded sessions. Sign in again with the restored credentials. Invalid uploads leave the current account and sessions intact. See [backups and restoration](SELF_HOSTING.md#backups-and-restoration) for compatibility limits and recovery copies.
+
Change a known password in Settings. Recover a forgotten password with `apptrail --reset-password` on the server, using the same data directory and operating-system user. Stop the AppTrail service first. Server filesystem access is required; there is no unauthenticated web password-reset endpoint.
The worker waits for an owner account before processing searches. After setup, scheduled monitoring continues while you are signed out. Logging out stops browser access, not scheduled work. Pause queries in the dashboard to stop monitoring.
diff --git a/docs/DEVELOPMENT.md b/docs/DEVELOPMENT.md
index 1656d05..f3a4c12 100644
--- a/docs/DEVELOPMENT.md
+++ b/docs/DEVELOPMENT.md
@@ -28,6 +28,8 @@ Charts use the bundled Chart.js 4.4.9 distribution in `static/vendor/chart.umd.m
| `src/apptrail/cli.py` | Console command, free-port binding, browser launch |
| `src/apptrail/api.py` | FastAPI routes and input validation |
| `src/apptrail/auth.py` | Owner setup, password hashing, session validation, CSRF, and login throttling |
+| `src/apptrail/backups.py` | Restore validation, schema compatibility, and exclusion of concurrent database users during restore |
+| `src/apptrail/storage.py` | Storage measurements, age-based response and image cleanup, and database compaction |
| `src/apptrail/config.py` | Data directory and credentials |
| `src/apptrail/db.py` | SQLAlchemy models and numbered SQLite schema migrations |
| `src/apptrail/engines.py` | Official SerpApi SDK requests and source normalization |
@@ -75,8 +77,8 @@ Never commit keys, credentials, `.env` files, databases, or unredacted provider
```bash
uv build
-uvx --from ./dist/apptrail-0.4.1-py3-none-any.whl apptrail --version
-uvx --from ./dist/apptrail-0.4.1-py3-none-any.whl apptrail --no-browser
+uvx --from ./dist/apptrail-1.0.0-py3-none-any.whl apptrail --version
+uvx --from ./dist/apptrail-1.0.0-py3-none-any.whl apptrail --no-browser
```
The wheel includes the static UI. Verify it from outside the checkout. The command's data directory is independent of the installed package and uv tool cache.
@@ -93,6 +95,10 @@ The separate live workflow runs on the same events and can be started manually.
SQLite uses WAL mode, foreign keys, and a busy timeout. Partial unique indexes prevent duplicate queued/running jobs for the same query or listing. A process lock allows one active AppTrail worker per data directory. Do not run multiple Uvicorn workers against the same workspace.
+Downloaded backups call `Database.backup(compact=True, include_sessions=False)` to remove raw response JSON, archived image blobs, and sessions from the copy, then run `VACUUM` on that copy. Normalized results and snapshot image hashes remain. The `backup_omissions` setting records the last affected run and snapshot IDs so restored views can explain missing content without labelling new checks as incomplete. Internal recovery and migration snapshots use full backups. Restore supports the current schema and schema 6, whose image comparison upgrade does not change the table layout.
+
+Storage cleanup uses the same request gate as restore and stops the worker before deleting content and running `VACUUM` plus a WAL checkpoint. It only clears responses for completed runs before the selected cutoff. Image age comes from the latest snapshot reference across every media field, so deduplicated assets used by newer snapshots survive. Unreferenced assets are also removed. The `storage_cleanup` setting records each category's cutoff for missing-content notices. Measurements read blob and JSON byte lengths without loading their contents into Python; snapshot media references are indexed in a temporary table for shared-image checks.
+
Schema 4 stores normalized results and provider responses in `run_payloads`, separate from the small `runs` records used for polling and scheduling. Load payloads only for evidence, matching, or reanalysis. State and dashboard queries must not fetch them. The upgrade preserves a snapshot in `backups/before-schema-3.sqlite3` before moving existing payloads.
Schema 5 adds notifications, per-query alert state, regional listing watches, listing snapshots, and archived image blobs. Existing workspaces receive no enabled listing watches. The migration backs up a schema-4 database to `backups/before-schema-4.sqlite3`. Listing-history jobs use `watch_id` and a separate partial unique index so different countries can be collected independently. Images share content hashes to avoid storing identical bytes repeatedly, and authenticated asset routes serve them from the database.
@@ -134,7 +140,7 @@ The workflow name is the filename, without `.github/workflows/`. If the project
1. Update `pyproject.toml` and `src/apptrail/__init__.py` to the same version, then run `uv lock` to update `uv.lock`.
2. Merge those changes and the workflows into `main`, and wait for the CI and live test workflows to pass.
-3. Create and publish a GitHub Release with a tag of `v` targeting `main`, for example `v0.4.1` for package version `0.4.1`. Use a version that has not already been published to PyPI.
+3. Create and publish a GitHub Release with a tag of `v` targeting `main`, for example `v1.0.0` for package version `1.0.0`. Use a version that has not already been published to PyPI.
The [Publish to PyPI workflow](https://github.com/serpapi/apptrail/actions/workflows/publish.yml) starts when the release is published, including published prereleases. Saving a draft or pushing a tag alone does not publish a package. Tags without a `v` prefix are ignored; mismatched versions and commits outside `main` fail validation. The workflow reruns Python, JavaScript, and Chromium tests, builds with `uv build`, and smoke-tests both the wheel and source distribution before uploading those artifacts to PyPI. Live tests run separately and are not a publishing-job dependency.
diff --git a/docs/SELF_HOSTING.md b/docs/SELF_HOSTING.md
index 2ad0b05..b14547a 100644
--- a/docs/SELF_HOSTING.md
+++ b/docs/SELF_HOSTING.md
@@ -24,9 +24,48 @@ The image supports Intel/AMD and ARM Linux and uses port `80`. If that port is a
### CapRover and Coolify
-Deploy `serpapi/apptrail:latest`. Mount persistent storage at `/data` and run one instance. Use the setup code in the application logs to create your owner account. If configuring a platform health check, use HTTP `/healthz`.
+Use the Docker Hub image `serpapi/apptrail:latest` and configure persistent storage at `/data` before the first deployment. AppTrail stores your account, history, and saved credentials there. Reusing this storage keeps your data when an update replaces the container. Run one instance.
-Account setup and login work behind the platform's HTTPS proxy without an origin setting or proxy IP configuration. Read [Public HTTPS hosting](#public-https-hosting) for optional forwarded-header settings.
+#### CapRover
+
+1. In **Apps**, create a new app named `apptrail` with **Has Persistent Data** checked.
+2. Open the app's **App Configs** and add a directory under **Persistent Directories**:
+
+ | Setting | Value |
+ |---|---|
+ | Path in App | `/data` |
+ | Label / Volume Name | `apptrail-data` |
+ | Set specific host path | Leave unchecked |
+
+3. Keep **Instance Count** at `1` and click **Save & Update** to save the storage configuration.
+4. Open **Deployment**, find **Deploy via ImageName**, enter `serpapi/apptrail:latest`, and click **Deploy**.
+
+CapRover manages the volume's location on the server. See [CapRover's persistent apps guide](https://caprover.com/docs/persistent-apps) for storage details.
+
+#### Coolify
+
+1. Open your project and environment, select **+ New**, then choose **Docker Image**.
+2. Enter `serpapi/apptrail` as **Image Name** and `latest` as **Tag**, then save to create the application.
+3. Before clicking **Deploy**, open **Configuration > Persistent Storage**, select **Add > Volume Mount**, and enter:
+
+ | Setting | Value |
+ |---|---|
+ | Name | `apptrail-data` |
+ | Source Path | Leave empty |
+ | Destination Path | `/data` |
+
+4. Click **Add** to save the mount. Leaving **Source Path** empty lets Docker manage a named volume.
+5. Configure your domain, keep the application on one server with one instance, then click **Deploy**.
+
+See [Coolify's persistent storage guide](https://coolify.io/docs/core/persistent-storage/storage-mounts/overview) and [volume mount instructions](https://coolify.io/docs/core/persistent-storage/storage-mounts/volume-mounts) for details.
+
+#### First setup and updates
+
+After deploying on either platform, use the one-time setup code in the application logs to create your owner account, then connect SerpApi. If configuring a platform health check, use HTTP `/healthz`. Account setup and login work behind the platform's HTTPS proxy without an origin setting or proxy IP configuration. Read [Public HTTPS hosting](#public-https-hosting) for optional forwarded-header settings.
+
+Before updating, open **Settings → Download SQLite backup** in AppTrail and save the SQLite database backup to your computer. Keep it so you can [restore your data](#backups-and-restoration) if the update causes issues.
+
+Then deploy `serpapi/apptrail:latest` again in the existing CapRover app or click **Redeploy** in the existing Coolify application. Keep the same volume mounted at `/data`; do not delete or recreate it. Your saved data persists across container replacement, though the app may briefly be unavailable while restarting.
## Build from source with Docker Compose
@@ -140,11 +179,25 @@ The worker retries transient failures up to three attempts with a delay. Final e
The browser can close while tracking continues. If the process stops or the computer sleeps, checks pause. On restart, AppTrail queues current checks for overdue queries. It cannot recover past results from the time it was offline. Graceful shutdown waits for the current provider operation to finish; individual requests have a 90-second timeout. Compose allows six minutes for shutdown. Use `--stop-timeout 360` with `docker run` as well, since a Google Play check can fetch three pages. The CLI stops waiting for stalled HTTP responses after ten seconds before shutting down the worker.
+## Storage usage and cleanup
+
+In **Settings → Storage usage**, check the total database size and the space used by saved SerpApi responses and listing-history images. Each category has a one-time cleanup option for content older than **7 days** or **30 days**, with an estimate of the content eligible for deletion. Review the confirmation before deleting. Rankings, saved answers, matched evidence, listing text and image-change records are kept. Images shared with snapshots inside the selected period are also kept.
+
+Cleanup waits for active checks, deletes the selected older content, then compacts the database to return space to disk. Other requests may briefly show that cleanup is in progress. The history views explain when original responses or archived images were deleted. New checks continue saving both; cleanup does not set an automatic retention policy. The total includes SQLite's temporary journal but excludes separate backup files, which cleanup leaves untouched. Compaction needs temporary free disk space; if it cannot finish, AppTrail reports that the deleted pages remain available for SQLite to reuse.
+
## Backups and restoration
-Use **Settings → Download backup** to create a consistent SQLite snapshot while the app is running. The backup includes the owner account and password hash, apps, queries, jobs, account usage snapshots, and saved search history. Keep it private. It excludes login sessions, `credentials.json`, and environment secrets. Sign in again after restoring a downloaded backup.
+Use **Settings → Download SQLite backup** to create a consistent snapshot while the app is running. The backup includes the owner account and password hash, apps, competitors, queries, settings, jobs, account usage snapshots, saved rankings and answers, matched evidence, and listing text history. Keep it private. It excludes raw SerpApi responses, archived listing images, login sessions, `credentials.json`, and environment secrets. AppTrail compacts the copy after removing those records so the downloaded file uses less space. The running database is unchanged. **Export history CSV** downloads search results for spreadsheet analysis; CSV files cannot restore a workspace.
-To restore:
+Restored search details explain that original response data was omitted to save space. Listing comparisons show placeholders for omitted images and keep their hashes and change records, so future checks can still detect changes. New checks save responses and images normally. AppTrail does not fetch live images as substitutes for missing historical images. Automatic recovery and schema-upgrade backups preserve all available response data and images.
+
+Use **Settings → Restore from SQLite** to upload a compatible AppTrail backup of up to 2 GB. Schema versions 6 and 7 are supported; schema 6 is validated and upgraded in a temporary copy before restoration. Confirm the overwrite, then select **Overwrite and restore**. Invalid uploads leave your workspace and login intact.
+
+On a fresh installation without an owner account, select **Have a backup? Restore your data** below **Create account**. Use the setup code from the new server's terminal or container logs, then upload your backup. You can restore before creating another account and sign in with the credentials saved in the backup.
+
+Restoring overwrites all current server data, including the owner account, password, apps, settings, and history. AppTrail waits for active requests and background checks to finish, saves the current database in `backups/before-restore-*.sqlite3`, then replaces it. All sessions are revoked, including sessions in manually copied backups. Sign in with the username and password saved in the uploaded backup. The current server's SerpApi key is kept because it is stored outside SQLite. Restored schedules resume with that key.
+
+To restore a larger backup, upgrade an older backup through the startup migrations, or restore without signing in:
1. Stop every AppTrail process using the destination directory.
2. Move the existing data directory aside as a rollback copy.
@@ -156,7 +209,9 @@ For Docker, perform the same operation inside the named volume with the service
## Updates
-Back up first. For an installation from Docker Hub, pull the new image and replace the container, reusing its data volume:
+Before any update, use **Settings → Download SQLite backup** in AppTrail to save a SQLite database backup to your computer. Keep it so you can [restore your data](#backups-and-restoration) if the update causes issues.
+
+For an installation from Docker Hub, pull the new image and replace the container, reusing its data volume:
```bash
docker pull serpapi/apptrail:latest
diff --git a/pyproject.toml b/pyproject.toml
index 7a04d07..0c07154 100644
--- a/pyproject.toml
+++ b/pyproject.toml
@@ -4,7 +4,7 @@ build-backend = "uv_build"
[project]
name = "apptrail"
-version = "0.4.1"
+version = "1.0.0"
description = "Open-source app visibility tracking across the App Store, Google Play, and AI search."
readme = "README.md"
requires-python = ">=3.11"
diff --git a/src/apptrail/__init__.py b/src/apptrail/__init__.py
index 6958d75..e2dedbd 100644
--- a/src/apptrail/__init__.py
+++ b/src/apptrail/__init__.py
@@ -1,3 +1,3 @@
"""App visibility tracking, on your own machine."""
-__version__ = "0.4.1"
+__version__ = "1.0.0"
diff --git a/src/apptrail/api.py b/src/apptrail/api.py
index d9c9c82..6c8bcd7 100644
--- a/src/apptrail/api.py
+++ b/src/apptrail/api.py
@@ -16,9 +16,11 @@
from sqlalchemy import select, text
from sqlalchemy.exc import IntegrityError
from starlette.background import BackgroundTask
+from starlette.concurrency import run_in_threadpool
-from . import __version__
+from . import __version__, storage
from .auth import Auth
+from .backups import MAX_RESTORE_BYTES, RestoreGate, RestoreMiddleware, prepare_restore
from .config import Config
from .db import (
App,
@@ -32,6 +34,7 @@
Run,
Setting,
Target,
+ backup_omits,
now,
)
from .engines import Gateway, ProviderError, apple_language, http_url
@@ -54,6 +57,12 @@ class KeyInput(Payload):
api_key: SecretStr
+class StorageCleanupInput(Payload):
+ kind: Literal["responses", "images"]
+ days: Literal[7, 30]
+ confirm: Literal[True]
+
+
class ReplaceQueryInput(Payload):
query: str = Field(min_length=1, max_length=500)
@@ -192,6 +201,7 @@ def create_app(directory=None, *, start_worker=True, gateway_factory=Gateway):
raise
service = Service(db, config, gateway_factory)
worker = Worker(service)
+ restore_gate = RestoreGate()
@asynccontextmanager
async def lifespan(app):
@@ -229,6 +239,8 @@ async def security(request: Request, call_next):
response.headers["Cache-Control"] = "no-store"
return response
+ app.add_middleware(RestoreMiddleware, gate=restore_gate)
+
@app.exception_handler(RequestValidationError)
async def invalid(request, exc):
# Pydantic's default errors include submitted input, including API keys.
@@ -560,6 +572,13 @@ def run_details(run_id: int):
raise HTTPException(404, "Run not found.")
return {
**record(item),
+ "responses_omitted": backup_omits(session, "responses_through_run", item.id)
+ and not item.responses,
+ "responses_cleaned": not item.responses
+ and item.status not in {"queued", "running"}
+ and storage.removed_by_cleanup(
+ session, "responses", item.finished_at or item.started_at or item.created_at
+ ),
**(
{
"listing_snapshot_id": session.scalar(
@@ -648,11 +667,31 @@ def export(app_id: int | None = None, country: str | None = None, source: str |
background=BackgroundTask(path.unlink, missing_ok=True),
)
+ @app.get("/api/storage")
+ def storage_usage():
+ return storage.usage(db)
+
+ @app.post("/api/storage/cleanup")
+ def storage_cleanup(payload: StorageCleanupInput):
+ with restore_gate.exclusive(operation="Storage cleanup"):
+ running = worker.thread is not None and worker.thread.is_alive()
+ if running:
+ worker.stop()
+ try:
+ return storage.cleanup(db, payload.kind, payload.days)
+ finally:
+ if running:
+ worker.start()
+
@app.get("/api/backup")
def backup():
with tempfile.NamedTemporaryFile(suffix=".sqlite3", delete=False) as handle:
path = Path(handle.name)
- db.backup(path, include_sessions=False)
+ try:
+ db.backup(path, include_sessions=False, compact=True)
+ except BaseException:
+ path.unlink(missing_ok=True)
+ raise
return FileResponse(
path,
filename="apptrail-backup.sqlite3",
@@ -660,6 +699,67 @@ def backup():
background=BackgroundTask(path.unlink, missing_ok=True),
)
+ def replace_database(request, path):
+ with restore_gate.exclusive():
+ if request.url.path == "/api/auth/restore":
+ auth.authorize_setup_restore(request, throttle=False)
+ elif not auth.identify(request):
+ raise HTTPException(401, "Sign in to AppTrail again before restoring.")
+ running = worker.thread is not None and worker.thread.is_alive()
+ if running:
+ worker.stop()
+ try:
+ backups = config.directory / "backups"
+ backups.mkdir(exist_ok=True, mode=0o700)
+ with tempfile.NamedTemporaryFile(
+ dir=backups, prefix="before-restore-", suffix=".sqlite3", delete=False
+ ) as handle:
+ recovery = Path(handle.name)
+ try:
+ db.backup(recovery, include_sessions=False)
+ except BaseException:
+ recovery.unlink(missing_ok=True)
+ raise
+ db.restore(path)
+ auth.setup_path.unlink(missing_ok=True)
+ finally:
+ if running:
+ worker.start()
+
+ @app.post("/api/restore")
+ @app.post("/api/auth/restore")
+ async def restore(request: Request):
+ if request.url.path == "/api/auth/restore":
+ await run_in_threadpool(auth.authorize_setup_restore, request)
+ if request.headers.get("x-apptrail-confirm-restore") != "overwrite":
+ raise HTTPException(
+ 422, "Confirm that the backup will overwrite all current server data."
+ )
+ if int(request.headers.get("content-length", "0")) > MAX_RESTORE_BYTES:
+ raise HTTPException(413, "SQLite backups must be 2 GB or smaller.")
+ with tempfile.TemporaryDirectory(prefix=".restore-", dir=config.directory) as directory:
+ staging = Path(directory)
+ upload = staging / "upload.sqlite3"
+ size = 0
+ with upload.open("wb") as handle:
+ upload.chmod(0o600)
+ async for chunk in request.stream():
+ size += len(chunk)
+ if size > MAX_RESTORE_BYTES:
+ raise HTTPException(413, "SQLite backups must be 2 GB or smaller.")
+ handle.write(chunk)
+ path = await run_in_threadpool(prepare_restore, upload, staging)
+ await run_in_threadpool(replace_database, request, path)
+ response = JSONResponse({"ok": True})
+ response.delete_cookie(
+ auth.cookie_name(request),
+ path="/",
+ secure=request.url.scheme == "https",
+ httponly=True,
+ samesite="strict",
+ )
+ return response
+
static = Path(__file__).parent / "static"
from .insights_api import register_insights
diff --git a/src/apptrail/auth.py b/src/apptrail/auth.py
index c1cc05e..7ff44d3 100644
--- a/src/apptrail/auth.py
+++ b/src/apptrail/auth.py
@@ -21,7 +21,7 @@
SESSION_SECONDS = 24 * 3600
SAFE_METHODS = {"GET", "HEAD", "OPTIONS"}
-PUBLIC_API = {"/api/auth/status", "/api/auth/setup", "/api/auth/login"}
+PUBLIC_API = {"/api/auth/status", "/api/auth/setup", "/api/auth/login", "/api/auth/restore"}
PUBLIC_FILES = {
"/login",
"/healthz",
@@ -104,6 +104,22 @@ def has_owner(self):
def setup_token(self):
return self.setup_path.read_text().strip() if self.setup_path.exists() else ""
+ def require_setup_token(self, supplied):
+ expected = self.setup_token()
+ if not expected or not secrets.compare_digest(expected.encode(), supplied.encode()):
+ raise HTTPException(
+ 403, "The setup code is incorrect. Use the code from the server terminal."
+ )
+
+ def authorize_setup_restore(self, request, *, throttle=True):
+ if throttle:
+ self.limit(request)
+ if self.has_owner():
+ raise HTTPException(
+ 409, "An account already exists. Sign in and restore from Settings."
+ )
+ self.require_setup_token(request.headers.get("x-apptrail-setup-token", ""))
+
def limit(self, request):
ip = request.client.host if request.client else "unknown"
try:
@@ -222,15 +238,17 @@ async def guard(self, request, call_next):
return RedirectResponse("/login", status_code=303)
return JSONResponse({"detail": "Sign in to AppTrail."}, status_code=401)
if request.method not in SAFE_METHODS:
- if (
- request.headers.get("x-apptrail-request") != "1"
- or request.headers.get("content-type", "").split(";")[0] != "application/json"
- ):
+ content_type = request.headers.get("content-type", "").split(";")[0]
+ allowed_type = content_type == "application/json" or (
+ path in {"/api/restore", "/api/auth/restore"}
+ and content_type == "application/vnd.sqlite3"
+ )
+ if request.headers.get("x-apptrail-request") != "1" or not allowed_type:
return JSONResponse(
- {"detail": "Use the AppTrail interface or authenticated JSON requests."},
+ {"detail": "Use the AppTrail interface or authenticated API requests."},
status_code=403,
)
- if path not in {"/api/auth/setup", "/api/auth/login"}:
+ if path not in {"/api/auth/setup", "/api/auth/login", "/api/auth/restore"}:
supplied = request.headers.get("x-csrf-token", "")
expected = (request.state.auth or {}).get("csrf_token", "")
if not expected or not secrets.compare_digest(supplied.encode(), expected.encode()):
@@ -272,13 +290,7 @@ def setup(payload: SetupInput, request: Request):
session.execute(text("BEGIN IMMEDIATE"))
if session.get(Owner, 1):
raise HTTPException(409, "An account already exists. Sign in instead.")
- expected = self.setup_token()
- if not expected or not secrets.compare_digest(
- expected.encode(), payload.setup_token.get_secret_value().encode()
- ):
- raise HTTPException(
- 403, "The setup code is incorrect. Use the code from the server terminal."
- )
+ self.require_setup_token(payload.setup_token.get_secret_value())
owner = Owner(
id=1,
username=payload.username.lower(),
diff --git a/src/apptrail/backups.py b/src/apptrail/backups.py
new file mode 100644
index 0000000..39e1a58
--- /dev/null
+++ b/src/apptrail/backups.py
@@ -0,0 +1,165 @@
+from __future__ import annotations
+
+import sqlite3
+import threading
+from contextlib import closing, contextmanager
+from pathlib import Path
+
+from argon2 import extract_parameters
+from argon2.exceptions import InvalidHashError
+from fastapi import HTTPException
+from starlette.responses import JSONResponse
+
+from .db import SCHEMA_VERSION, Database, now
+
+MAX_RESTORE_BYTES = 2 * 1024 * 1024 * 1024
+
+
+class RestoreGate:
+ def __init__(self):
+ self.condition = threading.Condition()
+ self.active = 0
+ self.restoring = False
+ self.operation = "A restore"
+
+ @contextmanager
+ def exclusive(self, operation="A restore"):
+ with self.condition:
+ if self.restoring:
+ raise HTTPException(
+ 503, f"{self.operation} is already in progress. Try again shortly."
+ )
+ self.restoring = True
+ self.operation = operation
+ if not self.condition.wait_for(lambda: self.active == 1, timeout=30):
+ self.restoring = False
+ raise HTTPException(503, "The server is busy. Try again shortly.")
+ try:
+ yield
+ finally:
+ with self.condition:
+ self.restoring = False
+
+
+class RestoreMiddleware:
+ def __init__(self, app, gate):
+ self.app, self.gate = app, gate
+
+ async def __call__(self, scope, receive, send):
+ if scope["type"] != "http" or scope["path"] == "/healthz":
+ return await self.app(scope, receive, send)
+ with self.gate.condition:
+ blocked = self.gate.restoring
+ if not blocked:
+ self.gate.active += 1
+ if blocked:
+ response = JSONResponse(
+ {"detail": f"{self.gate.operation} is in progress. Try again shortly."},
+ status_code=503,
+ headers={"Retry-After": "5", "Cache-Control": "no-store"},
+ )
+ return await response(scope, receive, send)
+ try:
+ await self.app(scope, receive, send)
+ finally:
+ with self.gate.condition:
+ self.gate.active -= 1
+ self.gate.condition.notify_all()
+
+
+def schema(connection):
+ # Ignore SQL formatting differences between SQLite and SQLAlchemy versions.
+ return {
+ (kind, name): " ".join(sql.split()).lower() if sql else None
+ for kind, name, sql in connection.execute(
+ "SELECT type, name, sql FROM sqlite_schema WHERE name NOT LIKE 'sqlite_%'"
+ )
+ }
+
+
+def prepare_restore(upload: Path, directory: Path) -> Path:
+ with upload.open("rb") as handle:
+ if handle.read(16) != b"SQLite format 3\0":
+ raise ValueError("Upload an AppTrail SQLite backup (.sqlite3).")
+
+ # Build a trusted schema so uploaded triggers or constraints never reach the server.
+ staged = Database(directory)
+ staged.close()
+ try:
+ with (
+ closing(sqlite3.connect(upload.as_uri() + "?mode=ro", uri=True)) as source,
+ closing(sqlite3.connect(staged.path)) as target,
+ ):
+ source.execute("PRAGMA trusted_schema=OFF")
+ version = source.execute("PRAGMA user_version").fetchone()[0]
+ # Schema 7 updates image comparisons without changing the table layout.
+ if version not in {6, SCHEMA_VERSION}:
+ raise ValueError(
+ "This backup uses an incompatible database format. "
+ "Use a backup from a compatible AppTrail version."
+ )
+ if schema(source) != schema(target):
+ raise ValueError("This file is not a compatible AppTrail backup.")
+ if source.execute("PRAGMA quick_check(1)").fetchone() != ("ok",):
+ raise ValueError("The backup is damaged. Upload a different SQLite backup.")
+
+ target.execute("BEGIN")
+ for (table,) in target.execute(
+ "SELECT name FROM sqlite_schema WHERE type='table' AND name NOT LIKE 'sqlite_%'"
+ ).fetchall():
+ columns = source.execute(f'PRAGMA table_info("{table}")').fetchall()
+ for _, column, kind, *_ in columns:
+ if kind != "JSON":
+ continue
+ if source.execute(
+ f'SELECT 1 FROM "{table}" WHERE NOT json_valid("{column}") LIMIT 1'
+ ).fetchone():
+ raise ValueError("The backup contains invalid app data.")
+ if table != "settings":
+ expected = (
+ "array" if column in {"aliases", "responses", "changes"} else "object"
+ )
+ if source.execute(
+ f'SELECT 1 FROM "{table}" WHERE json_type("{column}") != ? LIMIT 1',
+ (expected,),
+ ).fetchone():
+ raise ValueError("The backup contains incompatible app data.")
+ # Sessions must never be imported, even from a manually copied database.
+ if table in {"login_sessions", "auth_limits"}:
+ continue
+ rows = source.execute(f'SELECT * FROM "{table}"')
+ placeholders = ",".join("?" for _ in columns)
+ while batch := rows.fetchmany(1 if table == "listing_assets" else 100):
+ target.executemany(f'INSERT INTO "{table}" VALUES ({placeholders})', batch)
+
+ if target.execute("PRAGMA foreign_key_check").fetchone():
+ raise ValueError("The backup contains broken app or history references.")
+ owners = target.execute("SELECT id, username, password_hash FROM owner").fetchall()
+ if len(owners) != 1 or owners[0][0] != 1 or not owners[0][1]:
+ raise ValueError("The backup must contain an AppTrail owner account.")
+ try:
+ parameters = extract_parameters(owners[0][2])
+ if not (
+ 1 <= parameters.time_cost <= 10
+ and 8 <= parameters.memory_cost <= 262144
+ and 1 <= parameters.parallelism <= 8
+ ):
+ raise InvalidHashError
+ except (InvalidHashError, TypeError):
+ raise ValueError("The backup contains an invalid account password.") from None
+ target.execute(
+ "UPDATE runs SET status='queued', available_at=?, "
+ "error='Interrupted; resuming after restore.' WHERE status='running'",
+ (now(),),
+ )
+ target.commit()
+ if version == 6:
+ from .listing_history import upgrade_image_comparisons
+
+ with staged.engine.begin() as connection:
+ upgrade_image_comparisons(connection)
+ except sqlite3.DatabaseError:
+ raise ValueError("The backup is damaged or contains incompatible app data.") from None
+ finally:
+ staged.close()
+ return staged.path
diff --git a/src/apptrail/db.py b/src/apptrail/db.py
index 22fe572..00381da 100644
--- a/src/apptrail/db.py
+++ b/src/apptrail/db.py
@@ -1,5 +1,6 @@
from __future__ import annotations
+import json
import os
import sqlite3
import tempfile
@@ -24,6 +25,8 @@
)
from sqlalchemy.orm import DeclarativeBase, Mapped, mapped_column, relationship, sessionmaker
+SCHEMA_VERSION = 7
+
def now() -> float:
return datetime.now(UTC).timestamp()
@@ -197,6 +200,13 @@ class Setting(Base):
value: Mapped[dict] = mapped_column(JSON)
+def backup_omits(session, key, item_id):
+ setting = session.get(Setting, "backup_omissions")
+ metadata = setting.value if setting and isinstance(setting.value, dict) else {}
+ cutoff = metadata.get(key)
+ return isinstance(cutoff, int) and item_id <= cutoff
+
+
class Notification(Base):
__tablename__ = "notifications"
id: Mapped[int] = mapped_column(primary_key=True)
@@ -297,11 +307,11 @@ def configure(connection, _):
with self.engine.begin() as connection:
connection.exec_driver_sql("BEGIN IMMEDIATE")
version = connection.execute(text("PRAGMA user_version")).scalar()
- if version > 7:
+ if version > SCHEMA_VERSION:
raise RuntimeError(
"This database needs a newer AppTrail version. Upgrade AppTrail."
)
- if version and version < 7:
+ if version and version < SCHEMA_VERSION:
backups = directory / "backups"
backups.mkdir(exist_ok=True, mode=0o700)
with tempfile.NamedTemporaryFile(
@@ -408,7 +418,7 @@ def configure(connection, _):
self.path.chmod(0o600)
self.session = sessionmaker(self.engine, expire_on_commit=False)
- def backup(self, destination: Path, *, include_sessions=True):
+ def backup(self, destination: Path, *, include_sessions=True, compact=False):
private_sqlite_file(destination)
with (
closing(sqlite3.connect(self.path)) as source,
@@ -417,8 +427,37 @@ def backup(self, destination: Path, *, include_sessions=True):
source.backup(target)
if not include_sessions:
target.execute("DELETE FROM login_sessions")
- target.commit()
+ if compact:
+ target.execute("DELETE FROM listing_assets")
+ target.execute("UPDATE run_payloads SET responses='[]'")
+ target.execute("UPDATE runs SET responses='[]'")
+ omissions = {
+ "responses_through_run": target.execute(
+ "SELECT COALESCE(MAX(id), 0) FROM runs"
+ ).fetchone()[0],
+ "images_through_snapshot": target.execute(
+ "SELECT COALESCE(MAX(id), 0) FROM listing_snapshots"
+ ).fetchone()[0],
+ }
+ target.execute(
+ "INSERT INTO settings (key, value) VALUES ('backup_omissions', ?) "
+ "ON CONFLICT(key) DO UPDATE SET value=excluded.value",
+ (json.dumps(omissions),),
+ )
+ target.commit()
+ if compact:
+ # Deleting rows alone leaves their pages in the downloaded file.
+ target.execute("VACUUM")
destination.chmod(0o600)
def close(self):
self.engine.dispose()
+
+ def restore(self, source_path: Path):
+ # Call only after HTTP requests and the worker have finished using the database.
+ self.engine.dispose()
+ with (
+ closing(sqlite3.connect(source_path)) as source,
+ closing(sqlite3.connect(self.path)) as target,
+ ):
+ source.backup(target)
diff --git a/src/apptrail/listing_history.py b/src/apptrail/listing_history.py
index e71cde6..adce4ac 100644
--- a/src/apptrail/listing_history.py
+++ b/src/apptrail/listing_history.py
@@ -10,7 +10,8 @@
from PIL import Image, ImageOps
from sqlalchemy import select
-from .db import App, Listing, ListingAsset, ListingSnapshot, ListingWatch, Run, now
+from .db import App, Listing, ListingAsset, ListingSnapshot, ListingWatch, Run, backup_omits, now
+from .storage import removed_by_cleanup
MEDIA_FIELDS = {
"icon",
@@ -362,7 +363,39 @@ def comparison(self, snapshot_id):
.order_by(ListingSnapshot.id.desc())
.limit(1)
)
+
+ def with_image_availability(item):
+ if item is None:
+ return None
+ asset_ids = set()
+ for field in MEDIA_FIELDS:
+ value = item.data.get(field)
+ refs = [value] if field == "icon" else value or []
+ asset_ids.update(
+ ref["asset_id"]
+ for ref in refs
+ if isinstance(ref, dict) and ref.get("asset_id")
+ )
+ available = (
+ set(
+ session.scalars(
+ select(ListingAsset.id).where(ListingAsset.id.in_(asset_ids))
+ )
+ )
+ if asset_ids
+ else set()
+ )
+ missing = sorted(asset_ids - available)
+ return {
+ **record(item),
+ "missing_assets": missing,
+ "images_omitted": bool(missing)
+ and backup_omits(session, "images_through_snapshot", item.id),
+ "images_cleaned": bool(missing)
+ and removed_by_cleanup(session, "images", item.checked_at),
+ }
+
return {
- "snapshot": record(snapshot),
- "previous": record(previous) if previous else None,
+ "snapshot": with_image_availability(snapshot),
+ "previous": with_image_availability(previous),
}
diff --git a/src/apptrail/static/app.js b/src/apptrail/static/app.js
index b30e4dc..e2ce73f 100644
--- a/src/apptrail/static/app.js
+++ b/src/apptrail/static/app.js
@@ -70,7 +70,11 @@ const sourceColors = {
};
let csrfToken = "",
accountName = "",
+ restoring = false,
+ cleaningStorage = false,
leaving = false;
+let storageUsage = null, storageError = "", storagePending = false;
+const storageDays = { responses: 7, images: 7 };
let state,
dashboard,
charts = [],
@@ -216,7 +220,9 @@ async function api(path, options = {}) {
"X-CSRF-Token": csrfToken,
...options.headers,
},
- body: options.body === undefined ? undefined : JSON.stringify(options.body),
+ body: options.body instanceof Blob
+ ? options.body
+ : options.body === undefined ? undefined : JSON.stringify(options.body),
});
if (!response.ok) {
if (response.status === 401 && path !== "auth/password") {
@@ -855,6 +861,132 @@ function activityPage() {
}`
);
}
+function storageBytes(value) {
+ if (value < 1024) return `${value} B`;
+ const unit = Math.min(Math.floor(Math.log(value) / Math.log(1024)), 3);
+ return `${(value / 1024 ** unit).toFixed(1)} ${["B", "KiB", "MiB", "GiB"][unit]}`;
+}
+function storagePanel() {
+ const rows = storageUsage ? [
+ ["responses", "Saved SerpApi responses", "Original response data used for search evidence. Saved results and matched evidence are kept."],
+ ["images", "Listing-history images", "Archived icons and screenshots. Images still used by newer snapshots are kept."],
+ ].map(([kind, title, description]) => {
+ const data = storageUsage[kind], days = storageDays[kind], eligible = data.older_than[days];
+ return `
Includes saved data, indexes, free pages and ${storageBytes(storageUsage.journal_bytes)} of temporary database journal data. Separate backup files are excluded.
` : ""}${rows}
Cleanup runs once when you confirm. It permanently deletes the selected older content and compacts the database to free disk space. New checks continue saving responses and images.
`;
+}
+function renderStorage() {
+ const panel = $("#storage-panel");
+ if (panel) panel.outerHTML = storagePanel();
+}
+async function loadStorage() {
+ if (storagePending || cleaningStorage || restoring) return;
+ storagePending = true;
+ storageError = "";
+ renderStorage();
+ try {
+ storageUsage = await api("storage");
+ } catch (error) {
+ storageUsage = null;
+ storageError = error.message;
+ } finally {
+ storagePending = false;
+ renderStorage();
+ }
+}
+function cleanupStorageDialog(kind) {
+ const days = storageDays[kind], label = kind === "images" ? "listing-history images" : "saved SerpApi responses";
+ showModal("Delete older " + (kind === "images" ? "images" : "responses"), "Free space in this workspace.",
+ ``);
+}
+async function cleanupStorage(form) {
+ cleaningStorage = true;
+ syncVersion++;
+ const controls = $$("button", modal).filter(el => !el.disabled);
+ controls.forEach(el => { el.disabled = true; });
+ $("#cleanup-status").textContent = "Waiting for active checks, deleting older content and compacting the database…";
+ try {
+ const result = await api("storage/cleanup", { method: "POST", body: {
+ kind: form.dataset.kind, days: Number(form.dataset.days), confirm: true,
+ } });
+ storageUsage = result.usage;
+ modal.close();
+ renderStorage();
+ toast(result.compacted ? `Deleted ${storageBytes(result.removed_bytes)} of older content. Database compacted.` : "Older content was deleted, but the database could not be compacted. Its freed pages can still be reused by new data.", !result.compacted);
+ } finally {
+ cleaningStorage = false;
+ controls.forEach(el => { el.disabled = false; });
+ if ($("#cleanup-status")) $("#cleanup-status").textContent = "";
+ }
+}
+document.addEventListener("change", event => {
+ const kind = event.target.dataset.storageDays;
+ if (!kind) return;
+ storageDays[kind] = Number(event.target.value);
+ renderStorage();
+ $(`[data-storage-days="${kind}"]`)?.focus();
+});
+function backupsPanel() {
+ return `
Data & backups
+
History CSV
Export saved search results for spreadsheets and analysis. CSV files cannot restore your workspace and do not include account details or settings.
Save your apps, competitors, queries, settings, rankings, saved answers and listing text history. Includes your account and saved password. Keep this file private.
Raw SerpApi responses and listing-history images are omitted to keep backups small. Saved results, matched evidence and image-change records are kept. After restoring, omitted content is marked as unavailable.
Login sessions and your SerpApi key are not included. Your key is stored separately on the server.
+
Restore from SQLite
Upload a compatible AppTrail backup to replace the current server data. Your account, password and app data will be reset to the uploaded backup. Everyone will be signed out.
${button("Restore from SQLite", "restore-backup", "danger")}
AppTrail validates the file before restoring and saves a recovery backup on the server. The server’s current SerpApi key is kept.
+ `;
+}
+function restoreBackupDialog() {
+ showModal(
+ "Restore from SQLite",
+ "Replace this server’s workspace with an AppTrail backup.",
+ ``,
+ );
+}
+async function restoreBackup(form) {
+ const file = form.elements.backup.files[0];
+ if (!file || !file.size) throw new Error("Choose a SQLite backup file.");
+ if (file.size > 2 * 1024 * 1024 * 1024)
+ throw new Error("SQLite backups must be 2 GB or smaller.");
+ if (!form.elements.confirm.checked)
+ throw new Error("Confirm that the backup will overwrite all current server data.");
+ restoring = true;
+ syncVersion++;
+ const controls = $$("input, button", modal).filter((el) => !el.disabled);
+ controls.forEach((el) => { el.disabled = true; });
+ $("#restore-status").textContent = "Uploading, validating and restoring your backup. Waiting for active checks may take a few minutes. Keep this page open.";
+ try {
+ await api("restore", {
+ method: "POST",
+ headers: {
+ "Content-Type": "application/vnd.sqlite3",
+ "X-AppTrail-Confirm-Restore": "overwrite",
+ },
+ body: file,
+ });
+ leaving = true;
+ localStorage.removeItem("apptrail-app");
+ document.body.replaceChildren();
+ location.replace("/login?restored=1");
+ } finally {
+ restoring = false;
+ controls.forEach((el) => { el.disabled = false; });
+ if ($("#restore-status")) $("#restore-status").textContent = "";
+ }
+}
+function missingKeyBanner() {
+ if (
+ state.configured ||
+ (!state.apps.length && state.onboarding?.visible !== false)
+ )
+ return "";
+ return `
${state.sync_paused ? "Estimate for when workspace sync resumes. " : ""}Based on active queries and enabled listing histories. Discovery, retries, and manual checks are additional.
${state.sync_paused ? "Estimate for when workspace sync resumes. " : ""}Based on active queries and enabled listing histories. Discovery, retries, and manual checks are additional.
No matching evidence in this check.${result.kind === "ai" ? " If the answer uses another name, add it as an alias in My apps → Manage, then reanalyze history." : ""}
`}`).join("")}`;
+ if (r.responses_omitted)
+ body = '
Original SerpApi responses were omitted from this backup to save space. Saved results and matched evidence are still available.
' + body;
+ else if (r.responses_cleaned)
+ body = '
Original SerpApi responses were deleted during storage cleanup. Saved results and matched evidence are still available.