Skip to content

Commit 2f93e9b

Browse files
Merge pull request #3 from Computer-Hardware-Club/feat/peter-foreground
Bring Peter's club-member redesign to a staged release
2 parents 6cb3693 + 7abce15 commit 2f93e9b

110 files changed

Lines changed: 15419 additions & 737 deletions

File tree

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎Dockerfile‎

Lines changed: 3 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,4 +1,4 @@
1-
FROM python:3.12-slim AS base
1+
FROM python:3.12-slim@sha256:2f17fc044b579bab302c2e8054d3a686e2cb9a83de48e70534b94cd8ebbe06a9 AS base
22

33
ENV PYTHONDONTWRITEBYTECODE=1 \
44
PYTHONUNBUFFERED=1 \
@@ -20,6 +20,7 @@ RUN python3 -m pip install --no-cache-dir -r requirements.txt
2020
COPY bot.py README.md config.json .env.example club-knowledge.md ./
2121
COPY docker ./docker
2222
COPY peterbot ./peterbot
23+
COPY deploy/housekeeping.py deploy/state_backup.py ./deploy/
2324

2425
RUN chmod +x docker/entrypoint.sh \
2526
&& mkdir -p /app/peterbot-data /app/logs \
@@ -31,6 +32,7 @@ ENTRYPOINT ["/usr/bin/tini", "--", "./docker/entrypoint.sh"]
3132

3233
FROM base AS bot
3334
ARG PETERBOT_REVISION=unknown
35+
ENV PETERBOT_REVISION=$PETERBOT_REVISION
3436
LABEL org.opencontainers.image.revision=$PETERBOT_REVISION
3537

3638
FROM ghcr.io/ggml-org/llama.cpp:server AS llama_cpp_server

‎README.md‎

Lines changed: 34 additions & 243 deletions
Large diffs are not rendered by default.

‎compose.hermes.yml‎

Lines changed: 3 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -49,6 +49,9 @@ services:
4949
PETERBOT_RUNNER_SCOPE: peterbot
5050
PETERBOT_RUNNER_CONCURRENCY: 1
5151
PETERBOT_RUNNER_TIMEOUT: 1230
52+
# Bounded build profile (4 CPU / 4g / 2g workspace); rustc linking peaks above the
53+
# old 2g ceiling. Operator can downgrade to standard/small per host capacity.
54+
PETERBOT_WORKER_PROFILE: ${PETERBOT_WORKER_PROFILE:-build}
5255
# Trusted supervisor only. Worker containers never receive this socket.
5356
volumes: ['/var/run/docker.sock:/var/run/docker.sock']
5457
read_only: true

‎config.json‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -71,10 +71,10 @@
7171
"max_response_chars": 6000,
7272
"max_prompt_chars": 4000,
7373
"max_concurrent": 2,
74-
"user_requests_per_minute": 3,
75-
"guild_requests_per_minute": 15,
74+
"user_requests_per_minute": 20,
75+
"guild_requests_per_minute": 120,
7676
"allowed_guild_ids": [],
7777
"allow_dms": false,
7878
"vision_enabled": true
7979
}
80-
}
80+
}

‎deploy/HERMES.md‎

Lines changed: 50 additions & 45 deletions
Original file line numberDiff line numberDiff line change
@@ -1,5 +1,12 @@
11
# Operating Hermes-backed Peter
22

3+
Current P910 deployment (September 23, 2026) uses the [dedicated worker VM](../docs/worker-vm.md),
4+
member work access, and private officer controls in `#officers` and `#testing`.
5+
See the [live release record](../docs/release-evidence.md) and
6+
[cutover runbook](../docs/p910-cutover.md) for current limits, image revisions,
7+
backups, and rollback. The pilot defaults below document the earlier Docker
8+
stage and must not be used as the live P910 configuration.
9+
310
Hermes is an immutable upstream source dependency, not a fork or submodule. The worker Dockerfile installs the pinned revision in `requirements-hermes.txt` using the upstream-required editable installation. Runtime root files remain read-only. Hermes streaming is explicitly disabled because Peter's capability proxy returns non-streamed completions; reasoning remains enabled. Do not update Hermes without the adapter tests and a real-model smoke test.
411

512
## Services and authority
@@ -12,48 +19,43 @@ Hermes is an immutable upstream source dependency, not a fork or submodule. The
1219

1320
The worker only receives a short-lived capability for its job. The gateway verifies the requester's current Discord roles/channel access on every model/tool request. Personal memory is isolated by guild and user; club memory is public and officer-writable. Record versions and immutable revisions support audit. Authority never comes from memory. No private officer knowledge store is enabled in this pilot.
1421

15-
## Discord pilot
16-
17-
Ordinary mentions are conversational: Peter answers in the original channel without creating a thread or showing task IDs/status messages. A short model turn decides whether tools are needed. If so, work runs quietly in the existing sandbox and the useful answer/files are returned as a reply to the original message. Shared-channel workers receive public club memory and same-requester context only; personal memory is unavailable, including ID-based updates/deletes. Social context from other speakers stays in the conversational turn and is not sent to sandbox tools.
18-
19-
`officer_only: true` preserves the current tool pilot for configured officer role IDs. Member mentions and `/ask` keep the existing bounded conversational path. Explicit `/task` remains an optional private workspace, never the default for pings.
20-
21-
- `/task prompt [attachment]`: start work in a new private, non-invitable task thread.
22-
- `/ask` stays a private conversational answer. Officer mentions use the conversational/tool-routing path above.
23-
- Private work is explicitly continued through `/continue_task`; ordinary thread messages are not automatically converted into tasks.
24-
- `/tasks`: list your task IDs and statuses.
25-
- `/cancel_task task_id`: revoke the task and request immediate container termination.
26-
- `/continue_task task_id prompt`: continue a finished/interrupted task in its original private thread.
27-
- `/memory scope query`: privately inspect personal/public club memory.
28-
- `/forget memory_id version`: remove an authorized memory from recall. Restricted audit revisions remain.
29-
30-
For explicitly requested private tasks, Discord server administrators and members with Manage Threads may be able to access private threads; they are not confidential from server administration. Task ownership still prevents another user from taking over a task. The pilot accepts at most three UTF-8 text/code attachments totaling 128 KiB (one attachment in `/task`, multiple through mentions/follow-ups). Generated artifacts total at most 8 MiB. Images, Office/PDF uploads, arbitrary internet/package access, outbound messaging tools, server administration, native global memory/session search, cron and subagents are not exposed yet.
31-
32-
One sandbox agent task runs at a time. At most two tasks per user and 20 globally may be pending. Default limits are 20 minutes per task, 30 Hermes iterations, 8192 tokens per response, and a total allocated output budget of 131072 tokens. Thinking is enabled by the trusted model proxy regardless of caller flags. These are independent of legacy member-chat budgets. `deploy/prepare_hermes_config.py` generates a bot config with the stable persona, thinking enabled for member chat too, a 4096-token response allowance and a 240-second legacy request limit. Preserve a backup before replacing production JSON.
33-
34-
## Fast conversational turn
35-
36-
The first model turn on a mention decides whether to answer or hand the request to the sandbox. That turn runs with thinking enabled, because the deployed reasoning model does not emit tool calls reliably without it, and its completion budget (4096, and never below that) leaves room for thinking as well as the answer. Thinking is billed against the same budget, so a 2k allowance truncates mid-thought and returns an empty answer.
37-
38-
Reliability rules for that turn, all enforced in `conversation.py`:
39-
40-
- A blank answer is retried once with thinking disabled, which is the reliably non-empty path, plus an instruction to answer plainly.
41-
- Two blank answers return a short human line. Members never see an internal error string from a model wobble.
42-
- A blank answer never starts sandbox work by itself: only an explicit handoff does.
43-
- A tool name or argument shape the fast model invented is treated as a handoff, not an error. The sandbox re-checks authority and honours only its own allowlist, so failing toward doing the work is the safe direction.
44-
- Only wall-clock that is actually left is spent: the turn honours `inference.timeout_seconds`, and the retry shares the remaining budget instead of getting a fresh one.
45-
- The deadline is sized for a reasoning model (420 seconds deployed). A hard question can spend minutes thinking before it emits a byte, and a shorter deadline turns that into a member-facing failure.
46-
- The thinking attempt is capped at `TOTAL_ATTEMPT_SECONDS` and holds back `RETRY_RESERVE_SECONDS` for the cheap retry, keeping at least half of what is left if the deadline is short. Attempts log their budget, so a slow turn is distinguishable from a dead one.
47-
- Requests stream, and a stream that goes quiet for `STREAM_IDLE_SECONDS` is treated as dead. Non-streamed, vLLM sends nothing until the whole completion is finished, so a total deadline cannot tell "still thinking" from "server gone" and always loses to a long turn.
48-
- The fast turn is told to decide promptly and hand off rather than attempt real work itself.
49-
50-
Club facts come from `club-knowledge.md`, baked into the gateway image and loaded through `paths.knowledge_file`. The file must exist: a missing one fails config load rather than silently letting Peter answer club questions from guesses. Both the conversational turn and the sandbox persona receive the same excerpt.
51-
52-
Sandbox model calls get their own deadline (up to 600 seconds, bounded by `job_timeout`). The session-wide client deadline is far too short for a reasoning model writing thousands of tokens.
53-
54-
The worker sets an explicit `HERMES_API_CALL_STALE_TIMEOUT` (600 seconds) before it builds the agent. The trusted proxy answers non-streamed, and Hermes abandons a non-streamed call it has heard nothing from: the upstream floor for this model family is 180 seconds, while a 4k-token reasoning turn needs up to about 250 seconds before its first byte. Without the override, long tasks die as `model_failed` after the stale retry collides with the proxy's one-call-per-task lock. Setting it explicitly also prevents the run-budget calculation from halving it mid-job.
55-
56-
A heavy task can still exceed the 20-minute job budget, because the deployed model decodes at roughly 15-20 tokens per second and a coding task spends minutes reasoning. That ends as an honest `timeout` with the artifacts collected, not as a model error. The same limit decides whether a member's long request should hand off early rather than be attempted in the conversational turn.
22+
## Discord member workflow
23+
24+
The current P910 configuration has member work enabled in its configured listen
25+
channels. Peter responds when named, mentioned, replied to, or addressed through
26+
a recent scoped follow-up. Bare greetings are answered locally. Ordinary
27+
questions get a conversational model turn; requests that need tools or promised
28+
files go to a disposable worker and return to the original message. The
29+
foreground scheduler queues competing turns and gives an honest wait notice.
30+
31+
`officer_only` in `deploy/hermes.example.json` is an earlier pilot default, not
32+
the current protected production value. Authorization still comes from current
33+
Discord identity, channel, and roles. Shared-channel work receives public club
34+
context and the requester's scoped context; private personal memory and task
35+
history do not flow into a public work request.
36+
37+
- `/task prompt [attachment]` starts explicit work in a private task thread.
38+
- `/tasks`, `/continue_task`, and `/cancel_task` inspect, resume, or stop owned
39+
work. A cancellation preserves valid files collected before teardown.
40+
- `/memory` and `/forget` inspect and remove authorized recall entries; audit
41+
revisions remain.
42+
- `/ask`, `/recap`, `/suggest`, and `/remindme` remain available.
43+
44+
Private task threads can still be visible to Discord server administrators and
45+
members with Manage Threads; task ownership prevents another member from
46+
continuing or cancelling one. The task accepts at most three UTF-8 text/code
47+
attachments totaling 128 KiB. Generated artifacts total at most 8 MiB. The
48+
Hermes worker does not receive general browser/network, Discord administration,
49+
cron, or subagent authority. Exact pinned PyPI wheels and crates.io crates are
50+
available only through the authenticated [dependency broker](../docs/dependency-access.md).
51+
52+
One worker task runs at a time. At most two tasks per user and 20 globally may
53+
be pending. Defaults include a 20-minute task deadline, 30 Hermes iterations,
54+
8192 output tokens per model response, and a 131072-token task output budget.
55+
For the current conversation routing, tier budgets, retries, and live Qwen
56+
measurements, see [model latency](../docs/model-latency.md). Club facts come
57+
from `club-knowledge.md` and versioned officer updates; missing configured
58+
knowledge fails startup rather than making Peter guess.
5759

5860
## Presence: one message that becomes the answer
5961

@@ -95,6 +97,9 @@ Three smoke scripts, in increasing distance from the sandbox:
9597

9698
Note that a handoff for "who are the current club officers?" is correct: that answer needs the live roster tool, not the static knowledge file.
9799

98-
For rollback, stop/remove only the new `peterbot` container, restore the saved Compose file, `.env`, and `config.production.json`, then recreate `peterbot` from the previous gateway image (currently `peterbot-hermes-gateway:088c670-flashnext-v2`). Stop the new runner after active workers are gone. Preserve new SQLite state for diagnosis or later reuse. The dedicated worker firewall may safely remain installed.
99-
100-
This Docker pilot shares p910's kernel. Move execution to a dedicated VM before widening to general member access, arbitrary network/package downloads, or more privileged capabilities. Command allowlists and model instructions are not substitutes for OS/network isolation.
100+
For a later image switch or rollback, follow the
101+
[cutover runbook](../docs/p910-cutover.md). It requires a fresh verified state
102+
snapshot, protected copies of deployment config, one Discord gateway at a
103+
time, and reconciliation of uncertain delivery before any replay. The old
104+
same-host Docker pilot is a historical topology; current member work runs in
105+
the dedicated VM boundary.

‎deploy/check_hermes_isolation.py‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -118,6 +118,8 @@ def record(name, passed, **details):
118118
exposed = [path for path in FORBIDDEN_PATHS if accessible(path)]
119119
record("no_host_credentials_or_docker_socket", not exposed, accessible_paths=exposed)
120120
record("root_write_denied", not probe_write("/"))
121+
# The pinned rustc/cargo/node toolchain is trusted image content: read-only like /app.
122+
record("toolchain_write_denied", not probe_write("/usr/local/bin"))
121123
record("workspace_write_allowed", probe_write("/workspace"))
122124

123125
try:

‎deploy/hermes-firewall.sh‎

Lines changed: 12 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -9,3 +9,15 @@ $IPT -w -A PETERBOT-WORKER-HOST -s 192.168.240.2/32 -j RETURN
99
$IPT -w -A PETERBOT-WORKER-HOST -j DROP
1010
$IPT -w -C INPUT -s 192.168.240.0/24 -j PETERBOT-WORKER-HOST 2>/dev/null || \
1111
$IPT -w -I INPUT 1 -s 192.168.240.0/24 -j PETERBOT-WORKER-HOST
12+
13+
# VM workers reach the gateway through Docker's host-only published port.
14+
# DOCKER-USER sees the packet AFTER Docker DNAT, so match its original host
15+
# destination with conntrack. The source and ingress interface are pinned to
16+
# the dedicated guest; no tailnet or other Docker service is opened.
17+
while $IPT -w -C DOCKER-USER -i virbr-ctl -s 192.168.241.2/32 -p tcp \
18+
-m conntrack --ctorigdst 192.168.241.1 --ctorigdstport 8770 -j ACCEPT 2>/dev/null; do
19+
$IPT -w -D DOCKER-USER -i virbr-ctl -s 192.168.241.2/32 -p tcp \
20+
-m conntrack --ctorigdst 192.168.241.1 --ctorigdstport 8770 -j ACCEPT
21+
done
22+
$IPT -w -I DOCKER-USER 1 -i virbr-ctl -s 192.168.241.2/32 -p tcp \
23+
-m conntrack --ctorigdst 192.168.241.1 --ctorigdstport 8770 -j ACCEPT

‎deploy/hermes.example.json‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -4,10 +4,10 @@
44
"owner_user_ids": [],
55
"listen_channel_ids": [],
66
"control_channel_ids": [],
7-
"conversation_lease_seconds": 120,
7+
"conversation_lease_seconds": 300,
88
"officer_only": true,
99
"runner_url": "http://runner:8780",
10-
"tool_service_url": "http://gateway:8770",
10+
"tool_service_url": "http://192.168.240.2:8770",
1111
"state_dir": "/app/peterbot-data/hermes",
1212
"max_iterations": 30,
1313
"max_tokens": 8192,

‎deploy/housekeeping.py‎

Lines changed: 98 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,98 @@
1+
"""Operator CLI for private diagnostics, retention, and snapshot checks (PETER-16).
2+
3+
Subcommands:
4+
5+
diagnose STATE_DIR read-only aggregate health report (JSON)
6+
retention STATE_DIR [--apply ...] retention plan (dry-run by default)
7+
check-snapshot SNAPSHOT STAGING verify snapshot, restore to staging, diagnose
8+
9+
Nothing here touches the network, the model, or Discord. `diagnose` and the
10+
dry-run plan only read; `--apply` first writes a verified snapshot (including
11+
the project-blob cross-check) and aborts before any deletion if that check
12+
fails. Output carries aggregate counts and fixed reason tokens only: never
13+
prompts, answers, memory text, Discord IDs, job IDs, tokens, or paths.
14+
15+
Requires the repository root on sys.path (the container runs from /app; the
16+
bootstrap below also allows direct `python deploy/housekeeping.py`).
17+
"""
18+
from __future__ import annotations
19+
20+
import argparse
21+
import json
22+
import os
23+
from pathlib import Path
24+
import sys
25+
26+
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
27+
28+
from deploy.state_backup import backup, restore, verify # noqa: E402
29+
from peterbot.operator_ops import ( # noqa: E402
30+
RetentionConfig, diagnose, retention_apply, retention_plan,
31+
)
32+
33+
34+
def _emit(report: object) -> None:
35+
print(json.dumps(report, indent=2, sort_keys=True))
36+
37+
38+
def _config(args: argparse.Namespace) -> RetentionConfig:
39+
return RetentionConfig(
40+
conversations_days=args.conversations_days,
41+
metrics_days=args.metrics_days,
42+
terminal_jobs_days=args.terminal_jobs_days,
43+
settled_receipts_days=args.settled_receipts_days,
44+
include_projects=args.include_projects,
45+
)
46+
47+
48+
def main(argv: list[str] | None = None) -> int:
49+
parser = argparse.ArgumentParser(description=__doc__.splitlines()[0])
50+
sub = parser.add_subparsers(dest="command", required=True)
51+
52+
diag = sub.add_parser("diagnose", help="read-only aggregate health report")
53+
diag.add_argument("state_dir", type=Path,
54+
default=Path(os.environ.get("PETERBOT_STATE_DIR", "peterbot-data")),
55+
nargs="?")
56+
57+
ret = sub.add_parser("retention", help="retention plan (dry-run) or explicit apply")
58+
ret.add_argument("state_dir", type=Path,
59+
default=Path(os.environ.get("PETERBOT_STATE_DIR", "peterbot-data")),
60+
nargs="?")
61+
ret.add_argument("--apply", action="store_true",
62+
help="delete eligible rows; requires --backup-destination")
63+
ret.add_argument("--backup-destination", type=Path,
64+
help="snapshot written and verified before any deletion")
65+
ret.add_argument("--conversations-days", type=int, default=90)
66+
ret.add_argument("--metrics-days", type=int, default=30)
67+
ret.add_argument("--terminal-jobs-days", type=int, default=90)
68+
ret.add_argument("--settled-receipts-days", type=int, default=180)
69+
ret.add_argument("--include-projects", action="store_true",
70+
help="also run ProjectStore.retention_sweep under its own policy")
71+
72+
snap = sub.add_parser("check-snapshot",
73+
help="verify a snapshot, restore to staging, diagnose the copy")
74+
snap.add_argument("snapshot", type=Path)
75+
snap.add_argument("staging", type=Path)
76+
77+
args = parser.parse_args(argv)
78+
if args.command == "diagnose":
79+
_emit(diagnose(args.state_dir))
80+
elif args.command == "retention":
81+
config = _config(args)
82+
if not args.apply:
83+
_emit(retention_plan(args.state_dir, config))
84+
return 0
85+
if args.backup_destination is None:
86+
parser.error("--apply requires --backup-destination")
87+
backup(args.state_dir, args.backup_destination)
88+
verify(args.backup_destination) # redundant with restore; explicit gate
89+
_emit(retention_apply(args.state_dir, config))
90+
else:
91+
verify(args.snapshot)
92+
restore(args.snapshot, args.staging)
93+
_emit(diagnose(args.staging))
94+
return 0
95+
96+
97+
if __name__ == "__main__":
98+
raise SystemExit(main())

0 commit comments

Comments
 (0)