A surveillance camera is just "eyes" — it sees but doesn't understand. We give it a brain, but that brain must meet three harsh conditions at once: fast (alarm ≤ 1s), cheap (nearly zero long-term cost), and obedient (every rule is defined by the admin; the system never decides on its own).
Core idea: events are first-class; uncertain models are wrapped in deterministic engineering. The fast system (frame-difference gating + NanoDet detection) discovers and records events with or without any admin rule; admin rules only decide what earns extra attention and alarms. The slow system (local VLM) appends clearly-labeled supplementary descriptions to in-progress events — it can never mutate recorded facts, suppress notifications, or lift alarms.
- Rule-free event discovery — every appearance becomes a traceable event fact (start / end / reason / immutable initial observation), even with zero zones and zero rules configured
- Factual notifications first — every event creates one factual notification; the workbench separates "client displayed" (auto receipt) from "user confirmed" (explicit click), both persisted and restart-safe
- In-progress semantic updates, safely labeled — a local VLM appends versioned descriptions (v2, v3, …) with observed/inference split; every model version is presented as "model supplement, may be wrong, verify against evidence" and can never cancel events or alarms
- Fast/slow layering — T0 gating 0.4ms per frame → T1 detection on motion → the alarm path has zero model latency; the slow system works within a bounded budget
- Grid zone selection — the admin paints cells on the frame to define managed areas
- Custom alarm templates — structural conditions evaluated in real time
- Dual model channels — local ollama (data never leaves the machine) or a cloud OpenAI-compatible API, each call explicitly confirmed
- First-use environment baseline (Win11) — a guided, versioned scene baseline with hash-verified raw frames; recognition never auto-runs and never becomes a rule
- Events archived as they happen — SQLite stays fully auditable; the workbench is loopback-only
| Edition | Dedicated entry point | Product focus | Release gate |
|---|---|---|---|
| Linux NVR | python -m scam.linux_nvr |
Always-on service, multi-camera, recording, remote operations, 24h soak | Linux CI + systemd + on-device soak |
| Windows 11 Workstation | python -m scam.win11 |
Local interactive use, first-use baseline, event/confirmation chain, delivery packaging | Windows CI + native Windows 11 acceptance (file-source closed loop passed; real RTSP pending) |
Both editions share the scam domain core, while entry points, defaults, deployment chains and acceptance reports are maintained separately. python -m scam.nvr is kept for legacy deployments only.
Operations (install / upgrade / rollback / backup and restore) are covered by the operations runbook.
git clone https://github.com/CommitStrip/semantic-camera.git
cd semantic-camera
python -m venv .venv && . .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -e ".[detect,vision]" # the core needs only numpy; these add detection and video decoding
# optional: pip install -e ".[discover]" LAN camera discovery (WS-Discovery)
# optional: pip install -e ".[mqtt]" notification outletDetection and embedding weights are not distributed in the package; fetch them with the script:
python scripts/fetch_models.py --list # manifest and on-disk status
python scripts/fetch_models.py --model nanodet # download + verify once the hash is frozen
python scripts/fetch_models.py --verify-local <path> --expect <sha256> # verify your own exportpython -m scam.discover # discover cameras on your LAN (writes a cameras.json draft)
# edit cameras.json: fill in RTSP address and credentials
python -m scam.linux_nvr # start monitoring + workbench
# open http://127.0.0.1:8600
# the workbench binds localhost only and has no auth — it is never exposed to the LAN;
# remote access: ssh -L 8600:127.0.0.1:8600 <NVR> and open the same addressbash deploy/install.sh
sudo systemctl enable --now scam-nvrUpgrades, rollbacks and backup/restore are covered by the operations runbook, where every destructive step requires explicit human confirmation.
powershell -ExecutionPolicy Bypass -File deploy\install-win11.ps1
deploy\start-win11.batEquivalent entry point: python -m scam.win11.
RTSP camera
│ FrameSource (VUS first, automatic fallback to cv2)
▼
┌─────────────── Fast system (per-frame, zero model) ────────────────┐
│ T0 frame gating (0.4ms) → T1 NanoDet (on motion / on patrol) │
│ → track confirm → grid decision → four-state verdict → alarm ≤ 1s │
└────────────────────────────────────────────────────────────────────┘
▼ motion trigger
┌─────────────── Slow system (budget-controlled, novelty only) ──────┐
│ single execution core: claim CAS · whole-snapshot CAS · one backlog │
│ T2a V-JEPA embedding match → known pattern names itself (zero VLM) │
│ T2b VLM naming → new-event archive │
└────────────────────────────────────────────────────────────────────┘
▼
SQLite, fully auditable
| Metric | Value | Basis and conditions |
|---|---|---|
| Tests | 1379 passed / 16 skipped | Full suite on this Windows 11 machine (2026-09-29); CI additionally runs Linux 3.10 / 3.12 and Windows jobs |
| Alarm latency (structural trigger) | ≤ 1s | End-to-end assertion in tests/test_monitor.py: an alarm must fire within one second of the target entering a managed cell (synthetic frames, detection verdict path) |
| T0 frame gating | 0.39ms per frame (P95 0.43ms) | 1080p → 96×54 grayscale → gate decision, pure Python; measured over 199 frames |
| T1 detection | NanoDet-Plus ONNX, 416×416 input | Weights are not distributed; an earlier local measurement gave 23–24ms per inference, not re-measured since |
| Win11 file-source closed loop | 79 events / 79 notifications / 79 displayed receipts / 1 user confirmation; events, versions, receipts and confirmations all survived a real process restart | Native Windows 11 run on a recorded clip with real NanoDet ONNX + real local VLM; file source only — not real-camera evidence |
| Client display latency (exploratory) | 5 samples: 2.4–9.5s | Only measured while a browser tab was actively polling (3s cadence); sample size too small for percentiles — exploratory, not a steady-state claim |
| Semantic description quality | Not verified — risk samples recorded | Human frame-level review of 5 model outputs: 1 faithful, 4 with errors (scene misidentification, fabricated timestamps, overconfident labels). The UI therefore labels all model output as an unverified supplement |
CI runs three jobs (workflow):
- test (Linux, Python 3.10 and 3.12) — deploy script syntax (
bash -n), single-source systemd unit rendering +systemd-analyze verify, full test suite (including Linux service-level smoke), Linux NVR edition contract, release packaging gate (sensitive-surface double gate + reproducible packaging self-check) - test-windows — full test suite (service-level smoke skips automatically), release packaging dry-run, platform-layer smoke (UTF-8 output and data directory), Win11 entry smoke, deploy scripts parsed by both PowerShell 5.1 and 7
Most recent CI run on main: #61 (2026-09-23), all three jobs green; this
release commit re-runs the same workflows.
scam/ domain core (pure Python)
config.py venue profile, fail-closed validation
gate.py T0 frame-difference gating
detect.py T1 NanoDet-Plus ONNX detection
track.py track confirmation
zones.py grid zone selection
verdict.py four-state verdict
monitor.py fast-system monitoring loop
source.py camera source (VUS first, cv2 fallback)
db.py SQLite persistence (events / facts / descriptions /
notifications / escalations / baselines / zones)
event facts layer rule-free event discovery, immutable v1 observation,
factual notifications, crash-window backfill
semantic_updater.py in-progress semantic updates (versioned, safely
labeled, cannot mutate facts or alarms)
environment.py first-use environment baseline (Win11)
evidence.py evidence index and path fencing
server.py workbench HTTP (loopback only)
sinks.py notify.py alarm sinks and outlets (webhook / MQTT, optional)
recording.py recorder.py recording and clip export
slow_core.py single slow-system execution core
health.py nvr.py health water level; always-on entry
win11*.py Windows 11 entry, launcher, setup, baseline and
on-device acceptance
linux_*.py operations line: host probe, backup, upgrade, replay,
soak, unit rendering
models.py dual VLM channels (local / cloud, explicit opt-in)
packaging/ Windows delivery: single-instance launcher, build
script, forbidden-content gate, third-party notices
scripts/ acceptance, model fetch and verification, offline
measurement report, resource snapshot
tests/ pytest (1400+ test cases)
docs/ operations runbook
scripts/measure_report.py additionally turns any event database into a
read-only latency/fact measurement report (clock-source aware, integrity
checked, rebuildable).
| Document | Contents |
|---|---|
| Operations runbook | Install / upgrade / rollback / backup and restore, with human confirmation points |
- Real RTSP cameras are not yet verified. All Win11 closed-loop evidence (the 79-event chain above) was produced on recorded clips replayed as a camera stream. Real-camera acceptance needs an authorized RTSP source and is the next milestone.
- Semantic description quality is not verified. In a 5-sample frame-level human review, 1 output was faithful and 4 contained errors (scene misidentification, fabricated timestamps, overconfident labels). The UI therefore presents every model version as an unverified supplement; model text can never close events, suppress notifications or lift alarms.
- The Linux native operations chain, a 24-hour soak, and privacy network audits are not verified; dual-camera and long-run evidence is pending.
- Recording is off by default; enable it per camera if you want clip evidence.
- Detection and embedding weights are not distributed (
scripts/fetch_models.py); automatic download is refused fail-closed until the hash is frozen. - The workbench is a single-admin loopback form: no authentication, no multi-user; remote access goes through an SSH tunnel.
- Windows CI proves automated compatibility; the delivery zip must be built and accepted on a real Windows 11 machine.
- VUS — the single upstream: perception, budget mechanism and stream service are all reused
- NanoDet-Plus — person detection model (Apache-2.0)
- ollama — local VLM inference
- Meta AI — V-JEPA 2 video representation model (CC-BY-NC 4.0)
MIT, see LICENSE.