Skip to content

Repository files navigation

Semantic Camera

Semantic Camera v2

Budget-aware on-device semantic video event runtime — Linux NVR and Windows 11 editions

CI License: MIT Tests

English · 简体中文


A surveillance camera is just "eyes" — it sees but doesn't understand. We give it a brain, but that brain must meet three harsh conditions at once: fast (alarm ≤ 1s), cheap (nearly zero long-term cost), and obedient (every rule is defined by the admin; the system never decides on its own).

Core idea: events are first-class; uncertain models are wrapped in deterministic engineering. The fast system (frame-difference gating + NanoDet detection) discovers and records events with or without any admin rule; admin rules only decide what earns extra attention and alarms. The slow system (local VLM) appends clearly-labeled supplementary descriptions to in-progress events — it can never mutate recorded facts, suppress notifications, or lift alarms.

✨ Core features

  • Rule-free event discovery — every appearance becomes a traceable event fact (start / end / reason / immutable initial observation), even with zero zones and zero rules configured
  • Factual notifications first — every event creates one factual notification; the workbench separates "client displayed" (auto receipt) from "user confirmed" (explicit click), both persisted and restart-safe
  • In-progress semantic updates, safely labeled — a local VLM appends versioned descriptions (v2, v3, …) with observed/inference split; every model version is presented as "model supplement, may be wrong, verify against evidence" and can never cancel events or alarms
  • Fast/slow layering — T0 gating 0.4ms per frame → T1 detection on motion → the alarm path has zero model latency; the slow system works within a bounded budget
  • Grid zone selection — the admin paints cells on the frame to define managed areas
  • Custom alarm templates — structural conditions evaluated in real time
  • Dual model channels — local ollama (data never leaves the machine) or a cloud OpenAI-compatible API, each call explicitly confirmed
  • First-use environment baseline (Win11) — a guided, versioned scene baseline with hash-verified raw frames; recognition never auto-runs and never becomes a rule
  • Events archived as they happen — SQLite stays fully auditable; the workbench is loopback-only

🧭 Product editions

Edition Dedicated entry point Product focus Release gate
Linux NVR python -m scam.linux_nvr Always-on service, multi-camera, recording, remote operations, 24h soak Linux CI + systemd + on-device soak
Windows 11 Workstation python -m scam.win11 Local interactive use, first-use baseline, event/confirmation chain, delivery packaging Windows CI + native Windows 11 acceptance (file-source closed loop passed; real RTSP pending)

Both editions share the scam domain core, while entry points, defaults, deployment chains and acceptance reports are maintained separately. python -m scam.nvr is kept for legacy deployments only.

Operations (install / upgrade / rollback / backup and restore) are covered by the operations runbook.

🚀 Quick start

Install

git clone https://github.com/CommitStrip/semantic-camera.git
cd semantic-camera
python -m venv .venv && . .venv/bin/activate        # Windows: .venv\Scripts\activate
pip install -e ".[detect,vision]"                   # the core needs only numpy; these add detection and video decoding
# optional: pip install -e ".[discover]"   LAN camera discovery (WS-Discovery)
# optional: pip install -e ".[mqtt]"       notification outlet

Detection and embedding weights are not distributed in the package; fetch them with the script:

python scripts/fetch_models.py --list                              # manifest and on-disk status
python scripts/fetch_models.py --model nanodet                     # download + verify once the hash is frozen
python scripts/fetch_models.py --verify-local <path> --expect <sha256>   # verify your own export

Linux NVR

python -m scam.discover        # discover cameras on your LAN (writes a cameras.json draft)
# edit cameras.json: fill in RTSP address and credentials
python -m scam.linux_nvr       # start monitoring + workbench
# open http://127.0.0.1:8600
# the workbench binds localhost only and has no auth — it is never exposed to the LAN;
# remote access: ssh -L 8600:127.0.0.1:8600 <NVR> and open the same address

NVR deployment (systemd)

bash deploy/install.sh
sudo systemctl enable --now scam-nvr

Upgrades, rollbacks and backup/restore are covered by the operations runbook, where every destructive step requires explicit human confirmation.

Windows 11 Workstation

powershell -ExecutionPolicy Bypass -File deploy\install-win11.ps1
deploy\start-win11.bat

Equivalent entry point: python -m scam.win11.

🏗️ Architecture

RTSP camera
   │ FrameSource (VUS first, automatic fallback to cv2)
   ▼
┌─────────────── Fast system (per-frame, zero model) ────────────────┐
│ T0 frame gating (0.4ms) → T1 NanoDet (on motion / on patrol)       │
│ → track confirm → grid decision → four-state verdict → alarm ≤ 1s  │
└────────────────────────────────────────────────────────────────────┘
   ▼ motion trigger
┌─────────────── Slow system (budget-controlled, novelty only) ──────┐
│ single execution core: claim CAS · whole-snapshot CAS · one backlog │
│ T2a V-JEPA embedding match → known pattern names itself (zero VLM)  │
│ T2b VLM naming → new-event archive                                  │
└────────────────────────────────────────────────────────────────────┘
   ▼
SQLite, fully auditable

📊 Measured performance

Metric Value Basis and conditions
Tests 1379 passed / 16 skipped Full suite on this Windows 11 machine (2026-09-29); CI additionally runs Linux 3.10 / 3.12 and Windows jobs
Alarm latency (structural trigger) ≤ 1s End-to-end assertion in tests/test_monitor.py: an alarm must fire within one second of the target entering a managed cell (synthetic frames, detection verdict path)
T0 frame gating 0.39ms per frame (P95 0.43ms) 1080p → 96×54 grayscale → gate decision, pure Python; measured over 199 frames
T1 detection NanoDet-Plus ONNX, 416×416 input Weights are not distributed; an earlier local measurement gave 23–24ms per inference, not re-measured since
Win11 file-source closed loop 79 events / 79 notifications / 79 displayed receipts / 1 user confirmation; events, versions, receipts and confirmations all survived a real process restart Native Windows 11 run on a recorded clip with real NanoDet ONNX + real local VLM; file source only — not real-camera evidence
Client display latency (exploratory) 5 samples: 2.4–9.5s Only measured while a browser tab was actively polling (3s cadence); sample size too small for percentiles — exploratory, not a steady-state claim
Semantic description quality Not verified — risk samples recorded Human frame-level review of 5 model outputs: 1 faithful, 4 with errors (scene misidentification, fabricated timestamps, overconfident labels). The UI therefore labels all model output as an unverified supplement

✅ Quality gates

CI runs three jobs (workflow):

  • test (Linux, Python 3.10 and 3.12) — deploy script syntax (bash -n), single-source systemd unit rendering + systemd-analyze verify, full test suite (including Linux service-level smoke), Linux NVR edition contract, release packaging gate (sensitive-surface double gate + reproducible packaging self-check)
  • test-windows — full test suite (service-level smoke skips automatically), release packaging dry-run, platform-layer smoke (UTF-8 output and data directory), Win11 entry smoke, deploy scripts parsed by both PowerShell 5.1 and 7

Most recent CI run on main: #61 (2026-09-23), all three jobs green; this release commit re-runs the same workflows.

📁 Repository layout

scam/                    domain core (pure Python)
  config.py              venue profile, fail-closed validation
  gate.py                T0 frame-difference gating
  detect.py              T1 NanoDet-Plus ONNX detection
  track.py               track confirmation
  zones.py               grid zone selection
  verdict.py             four-state verdict
  monitor.py             fast-system monitoring loop
  source.py              camera source (VUS first, cv2 fallback)
  db.py                  SQLite persistence (events / facts / descriptions /
                         notifications / escalations / baselines / zones)
  event facts layer      rule-free event discovery, immutable v1 observation,
                         factual notifications, crash-window backfill
  semantic_updater.py    in-progress semantic updates (versioned, safely
                         labeled, cannot mutate facts or alarms)
  environment.py         first-use environment baseline (Win11)
  evidence.py            evidence index and path fencing
  server.py              workbench HTTP (loopback only)
  sinks.py notify.py     alarm sinks and outlets (webhook / MQTT, optional)
  recording.py recorder.py   recording and clip export
  slow_core.py           single slow-system execution core
  health.py nvr.py       health water level; always-on entry
  win11*.py              Windows 11 entry, launcher, setup, baseline and
                         on-device acceptance
  linux_*.py             operations line: host probe, backup, upgrade, replay,
                         soak, unit rendering
  models.py              dual VLM channels (local / cloud, explicit opt-in)
packaging/               Windows delivery: single-instance launcher, build
                         script, forbidden-content gate, third-party notices
scripts/                 acceptance, model fetch and verification, offline
                         measurement report, resource snapshot
tests/                   pytest (1400+ test cases)
docs/                    operations runbook

scripts/measure_report.py additionally turns any event database into a read-only latency/fact measurement report (clock-source aware, integrity checked, rebuildable).

📖 Documentation

Document Contents
Operations runbook Install / upgrade / rollback / backup and restore, with human confirmation points

⚠️ Known limitations (honest list — v0.1.0-beta)

  • Real RTSP cameras are not yet verified. All Win11 closed-loop evidence (the 79-event chain above) was produced on recorded clips replayed as a camera stream. Real-camera acceptance needs an authorized RTSP source and is the next milestone.
  • Semantic description quality is not verified. In a 5-sample frame-level human review, 1 output was faithful and 4 contained errors (scene misidentification, fabricated timestamps, overconfident labels). The UI therefore presents every model version as an unverified supplement; model text can never close events, suppress notifications or lift alarms.
  • The Linux native operations chain, a 24-hour soak, and privacy network audits are not verified; dual-camera and long-run evidence is pending.
  • Recording is off by default; enable it per camera if you want clip evidence.
  • Detection and embedding weights are not distributed (scripts/fetch_models.py); automatic download is refused fail-closed until the hash is frozen.
  • The workbench is a single-admin loopback form: no authentication, no multi-user; remote access goes through an SSH tunnel.
  • Windows CI proves automated compatibility; the delivery zip must be built and accepted on a real Windows 11 machine.

🙏 Acknowledgements

  • VUS — the single upstream: perception, budget mechanism and stream service are all reused
  • NanoDet-Plus — person detection model (Apache-2.0)
  • ollama — local VLM inference
  • Meta AI — V-JEPA 2 video representation model (CC-BY-NC 4.0)

License

MIT, see LICENSE.

About

Semantic camera: venue-mode realtime video understanding on device (语义摄像头:场所模式端侧实时视频理解)

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages