Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
29 commits
Select commit Hold shift + click to select a range
2cda1d8
Preserve scoped conversations across gateway restarts
Sep 23, 2026
1f26be5
Make release latency and failure evidence private and measurable
Sep 23, 2026
8443b72
Keep authorized announcement sends recoverable after uncertainty
Sep 23, 2026
f83b761
Let verified officers update club facts and Peter’s voice safely
Sep 23, 2026
5b6cac4
Answer with the latest authorized club context and voice
Sep 23, 2026
79f618f
Measure when a model turn first becomes actionable
Sep 23, 2026
90bc6ad
Send authorized announcements with an enforced Discord nonce
Sep 23, 2026
9a317f5
Keep member project files scoped and recoverable across tasks
Sep 23, 2026
7d3b95e
Keep Peter's state recoverable during the staged rollout
Sep 23, 2026
e439b5e
Give Peter one durable turn at a time
Sep 23, 2026
8124757
Let Peter build real projects inside a contained worker
Sep 23, 2026
3636b9e
Make Peter's conversations and club actions coherent across restarts
Sep 23, 2026
053d059
Keep revoked work and package fetches inside their task limits
Sep 23, 2026
80c1002
Reserve the broker address so workers can actually reach it
Sep 23, 2026
5ca4193
Let the VM reach Peter's real broker through Docker's host firewall
Sep 23, 2026
8f014d7
Keep quick conversation free of false queue notices
Sep 23, 2026
e9e8a5f
Require real work before Peter promises requested files
Sep 23, 2026
66ffd54
Keep private project follow-ups honest through delivery
Sep 23, 2026
274c86a
Meet explicit work handoff latency without weakening delivery
Sep 23, 2026
cb95984
Make the deployed Peter rollout auditable
Sep 23, 2026
62e2b9f
Let Peter answer casual turns with less ceremony
Sep 23, 2026
64ce84f
Keep the voice rollout verifiable under CI load
Sep 23, 2026
87aaf6d
Keep club conversations alive through natural followups
Sep 23, 2026
eaaeb9d
Keep the conversation lease contract aligned with rollout
Sep 23, 2026
f254581
Prove officer controls in the private testing channel
Sep 23, 2026
08c8968
Keep Peter truthful and useful when officers correct his memory
Sep 23, 2026
2800d97
Recognize the exact model correction Oliver used
Sep 23, 2026
c900f0e
Record the live Qwen and memory followup rollout
Sep 23, 2026
7abce15
Make the published release understandable to club operators
Sep 23, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 3 additions & 1 deletion Dockerfile
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
FROM python:3.12-slim AS base
FROM python:3.12-slim@sha256:2f17fc044b579bab302c2e8054d3a686e2cb9a83de48e70534b94cd8ebbe06a9 AS base

ENV PYTHONDONTWRITEBYTECODE=1 \
PYTHONUNBUFFERED=1 \
Expand All @@ -20,6 +20,7 @@ RUN python3 -m pip install --no-cache-dir -r requirements.txt
COPY bot.py README.md config.json .env.example club-knowledge.md ./
COPY docker ./docker
COPY peterbot ./peterbot
COPY deploy/housekeeping.py deploy/state_backup.py ./deploy/

RUN chmod +x docker/entrypoint.sh \
&& mkdir -p /app/peterbot-data /app/logs \
Expand All @@ -31,6 +32,7 @@ ENTRYPOINT ["/usr/bin/tini", "--", "./docker/entrypoint.sh"]

FROM base AS bot
ARG PETERBOT_REVISION=unknown
ENV PETERBOT_REVISION=$PETERBOT_REVISION
LABEL org.opencontainers.image.revision=$PETERBOT_REVISION

FROM ghcr.io/ggml-org/llama.cpp:server AS llama_cpp_server
Expand Down
277 changes: 34 additions & 243 deletions README.md

Large diffs are not rendered by default.

3 changes: 3 additions & 0 deletions compose.hermes.yml
Original file line number Diff line number Diff line change
Expand Up @@ -49,6 +49,9 @@ services:
PETERBOT_RUNNER_SCOPE: peterbot
PETERBOT_RUNNER_CONCURRENCY: 1
PETERBOT_RUNNER_TIMEOUT: 1230
# Bounded build profile (4 CPU / 4g / 2g workspace); rustc linking peaks above the
# old 2g ceiling. Operator can downgrade to standard/small per host capacity.
PETERBOT_WORKER_PROFILE: ${PETERBOT_WORKER_PROFILE:-build}
# Trusted supervisor only. Worker containers never receive this socket.
volumes: ['/var/run/docker.sock:/var/run/docker.sock']
read_only: true
Expand Down
6 changes: 3 additions & 3 deletions config.json
Original file line number Diff line number Diff line change
Expand Up @@ -71,10 +71,10 @@
"max_response_chars": 6000,
"max_prompt_chars": 4000,
"max_concurrent": 2,
"user_requests_per_minute": 3,
"guild_requests_per_minute": 15,
"user_requests_per_minute": 20,
"guild_requests_per_minute": 120,
"allowed_guild_ids": [],
"allow_dms": false,
"vision_enabled": true
}
}
}
95 changes: 50 additions & 45 deletions deploy/HERMES.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,12 @@
# Operating Hermes-backed Peter

Current P910 deployment (September 23, 2026) uses the [dedicated worker VM](../docs/worker-vm.md),
member work access, and private officer controls in `#officers` and `#testing`.
See the [live release record](../docs/release-evidence.md) and
[cutover runbook](../docs/p910-cutover.md) for current limits, image revisions,
backups, and rollback. The pilot defaults below document the earlier Docker
stage and must not be used as the live P910 configuration.

Hermes is an immutable upstream source dependency, not a fork or submodule. The worker Dockerfile installs the pinned revision in `requirements-hermes.txt` using the upstream-required editable installation. Runtime root files remain read-only. Hermes streaming is explicitly disabled because Peter's capability proxy returns non-streamed completions; reasoning remains enabled. Do not update Hermes without the adapter tests and a real-model smoke test.

## Services and authority
Expand All @@ -12,48 +19,43 @@ Hermes is an immutable upstream source dependency, not a fork or submodule. The

The worker only receives a short-lived capability for its job. The gateway verifies the requester's current Discord roles/channel access on every model/tool request. Personal memory is isolated by guild and user; club memory is public and officer-writable. Record versions and immutable revisions support audit. Authority never comes from memory. No private officer knowledge store is enabled in this pilot.

## Discord pilot

Ordinary mentions are conversational: Peter answers in the original channel without creating a thread or showing task IDs/status messages. A short model turn decides whether tools are needed. If so, work runs quietly in the existing sandbox and the useful answer/files are returned as a reply to the original message. Shared-channel workers receive public club memory and same-requester context only; personal memory is unavailable, including ID-based updates/deletes. Social context from other speakers stays in the conversational turn and is not sent to sandbox tools.

`officer_only: true` preserves the current tool pilot for configured officer role IDs. Member mentions and `/ask` keep the existing bounded conversational path. Explicit `/task` remains an optional private workspace, never the default for pings.

- `/task prompt [attachment]`: start work in a new private, non-invitable task thread.
- `/ask` stays a private conversational answer. Officer mentions use the conversational/tool-routing path above.
- Private work is explicitly continued through `/continue_task`; ordinary thread messages are not automatically converted into tasks.
- `/tasks`: list your task IDs and statuses.
- `/cancel_task task_id`: revoke the task and request immediate container termination.
- `/continue_task task_id prompt`: continue a finished/interrupted task in its original private thread.
- `/memory scope query`: privately inspect personal/public club memory.
- `/forget memory_id version`: remove an authorized memory from recall. Restricted audit revisions remain.

For explicitly requested private tasks, Discord server administrators and members with Manage Threads may be able to access private threads; they are not confidential from server administration. Task ownership still prevents another user from taking over a task. The pilot accepts at most three UTF-8 text/code attachments totaling 128 KiB (one attachment in `/task`, multiple through mentions/follow-ups). Generated artifacts total at most 8 MiB. Images, Office/PDF uploads, arbitrary internet/package access, outbound messaging tools, server administration, native global memory/session search, cron and subagents are not exposed yet.

One sandbox agent task runs at a time. At most two tasks per user and 20 globally may be pending. Default limits are 20 minutes per task, 30 Hermes iterations, 8192 tokens per response, and a total allocated output budget of 131072 tokens. Thinking is enabled by the trusted model proxy regardless of caller flags. These are independent of legacy member-chat budgets. `deploy/prepare_hermes_config.py` generates a bot config with the stable persona, thinking enabled for member chat too, a 4096-token response allowance and a 240-second legacy request limit. Preserve a backup before replacing production JSON.

## Fast conversational turn

The first model turn on a mention decides whether to answer or hand the request to the sandbox. That turn runs with thinking enabled, because the deployed reasoning model does not emit tool calls reliably without it, and its completion budget (4096, and never below that) leaves room for thinking as well as the answer. Thinking is billed against the same budget, so a 2k allowance truncates mid-thought and returns an empty answer.

Reliability rules for that turn, all enforced in `conversation.py`:

- A blank answer is retried once with thinking disabled, which is the reliably non-empty path, plus an instruction to answer plainly.
- Two blank answers return a short human line. Members never see an internal error string from a model wobble.
- A blank answer never starts sandbox work by itself: only an explicit handoff does.
- A tool name or argument shape the fast model invented is treated as a handoff, not an error. The sandbox re-checks authority and honours only its own allowlist, so failing toward doing the work is the safe direction.
- Only wall-clock that is actually left is spent: the turn honours `inference.timeout_seconds`, and the retry shares the remaining budget instead of getting a fresh one.
- The deadline is sized for a reasoning model (420 seconds deployed). A hard question can spend minutes thinking before it emits a byte, and a shorter deadline turns that into a member-facing failure.
- The thinking attempt is capped at `TOTAL_ATTEMPT_SECONDS` and holds back `RETRY_RESERVE_SECONDS` for the cheap retry, keeping at least half of what is left if the deadline is short. Attempts log their budget, so a slow turn is distinguishable from a dead one.
- Requests stream, and a stream that goes quiet for `STREAM_IDLE_SECONDS` is treated as dead. Non-streamed, vLLM sends nothing until the whole completion is finished, so a total deadline cannot tell "still thinking" from "server gone" and always loses to a long turn.
- The fast turn is told to decide promptly and hand off rather than attempt real work itself.

Club facts come from `club-knowledge.md`, baked into the gateway image and loaded through `paths.knowledge_file`. The file must exist: a missing one fails config load rather than silently letting Peter answer club questions from guesses. Both the conversational turn and the sandbox persona receive the same excerpt.

Sandbox model calls get their own deadline (up to 600 seconds, bounded by `job_timeout`). The session-wide client deadline is far too short for a reasoning model writing thousands of tokens.

The worker sets an explicit `HERMES_API_CALL_STALE_TIMEOUT` (600 seconds) before it builds the agent. The trusted proxy answers non-streamed, and Hermes abandons a non-streamed call it has heard nothing from: the upstream floor for this model family is 180 seconds, while a 4k-token reasoning turn needs up to about 250 seconds before its first byte. Without the override, long tasks die as `model_failed` after the stale retry collides with the proxy's one-call-per-task lock. Setting it explicitly also prevents the run-budget calculation from halving it mid-job.

A heavy task can still exceed the 20-minute job budget, because the deployed model decodes at roughly 15-20 tokens per second and a coding task spends minutes reasoning. That ends as an honest `timeout` with the artifacts collected, not as a model error. The same limit decides whether a member's long request should hand off early rather than be attempted in the conversational turn.
## Discord member workflow

The current P910 configuration has member work enabled in its configured listen
channels. Peter responds when named, mentioned, replied to, or addressed through
a recent scoped follow-up. Bare greetings are answered locally. Ordinary
questions get a conversational model turn; requests that need tools or promised
files go to a disposable worker and return to the original message. The
foreground scheduler queues competing turns and gives an honest wait notice.

`officer_only` in `deploy/hermes.example.json` is an earlier pilot default, not
the current protected production value. Authorization still comes from current
Discord identity, channel, and roles. Shared-channel work receives public club
context and the requester's scoped context; private personal memory and task
history do not flow into a public work request.

- `/task prompt [attachment]` starts explicit work in a private task thread.
- `/tasks`, `/continue_task`, and `/cancel_task` inspect, resume, or stop owned
work. A cancellation preserves valid files collected before teardown.
- `/memory` and `/forget` inspect and remove authorized recall entries; audit
revisions remain.
- `/ask`, `/recap`, `/suggest`, and `/remindme` remain available.

Private task threads can still be visible to Discord server administrators and
members with Manage Threads; task ownership prevents another member from
continuing or cancelling one. The task accepts at most three UTF-8 text/code
attachments totaling 128 KiB. Generated artifacts total at most 8 MiB. The
Hermes worker does not receive general browser/network, Discord administration,
cron, or subagent authority. Exact pinned PyPI wheels and crates.io crates are
available only through the authenticated [dependency broker](../docs/dependency-access.md).

One worker task runs at a time. At most two tasks per user and 20 globally may
be pending. Defaults include a 20-minute task deadline, 30 Hermes iterations,
8192 output tokens per model response, and a 131072-token task output budget.
For the current conversation routing, tier budgets, retries, and live Qwen
measurements, see [model latency](../docs/model-latency.md). Club facts come
from `club-knowledge.md` and versioned officer updates; missing configured
knowledge fails startup rather than making Peter guess.

## Presence: one message that becomes the answer

Expand Down Expand Up @@ -95,6 +97,9 @@ Three smoke scripts, in increasing distance from the sandbox:

Note that a handoff for "who are the current club officers?" is correct: that answer needs the live roster tool, not the static knowledge file.

For rollback, stop/remove only the new `peterbot` container, restore the saved Compose file, `.env`, and `config.production.json`, then recreate `peterbot` from the previous gateway image (currently `peterbot-hermes-gateway:088c670-flashnext-v2`). Stop the new runner after active workers are gone. Preserve new SQLite state for diagnosis or later reuse. The dedicated worker firewall may safely remain installed.

This Docker pilot shares p910's kernel. Move execution to a dedicated VM before widening to general member access, arbitrary network/package downloads, or more privileged capabilities. Command allowlists and model instructions are not substitutes for OS/network isolation.
For a later image switch or rollback, follow the
[cutover runbook](../docs/p910-cutover.md). It requires a fresh verified state
snapshot, protected copies of deployment config, one Discord gateway at a
time, and reconciliation of uncertain delivery before any replay. The old
same-host Docker pilot is a historical topology; current member work runs in
the dedicated VM boundary.
2 changes: 2 additions & 0 deletions deploy/check_hermes_isolation.py
Original file line number Diff line number Diff line change
Expand Up @@ -118,6 +118,8 @@ def record(name, passed, **details):
exposed = [path for path in FORBIDDEN_PATHS if accessible(path)]
record("no_host_credentials_or_docker_socket", not exposed, accessible_paths=exposed)
record("root_write_denied", not probe_write("/"))
# The pinned rustc/cargo/node toolchain is trusted image content: read-only like /app.
record("toolchain_write_denied", not probe_write("/usr/local/bin"))
record("workspace_write_allowed", probe_write("/workspace"))

try:
Expand Down
12 changes: 12 additions & 0 deletions deploy/hermes-firewall.sh
Original file line number Diff line number Diff line change
Expand Up @@ -9,3 +9,15 @@ $IPT -w -A PETERBOT-WORKER-HOST -s 192.168.240.2/32 -j RETURN
$IPT -w -A PETERBOT-WORKER-HOST -j DROP
$IPT -w -C INPUT -s 192.168.240.0/24 -j PETERBOT-WORKER-HOST 2>/dev/null || \
$IPT -w -I INPUT 1 -s 192.168.240.0/24 -j PETERBOT-WORKER-HOST

# VM workers reach the gateway through Docker's host-only published port.
# DOCKER-USER sees the packet AFTER Docker DNAT, so match its original host
# destination with conntrack. The source and ingress interface are pinned to
# the dedicated guest; no tailnet or other Docker service is opened.
while $IPT -w -C DOCKER-USER -i virbr-ctl -s 192.168.241.2/32 -p tcp \
-m conntrack --ctorigdst 192.168.241.1 --ctorigdstport 8770 -j ACCEPT 2>/dev/null; do
$IPT -w -D DOCKER-USER -i virbr-ctl -s 192.168.241.2/32 -p tcp \
-m conntrack --ctorigdst 192.168.241.1 --ctorigdstport 8770 -j ACCEPT
done
$IPT -w -I DOCKER-USER 1 -i virbr-ctl -s 192.168.241.2/32 -p tcp \
-m conntrack --ctorigdst 192.168.241.1 --ctorigdstport 8770 -j ACCEPT
4 changes: 2 additions & 2 deletions deploy/hermes.example.json
Original file line number Diff line number Diff line change
Expand Up @@ -4,10 +4,10 @@
"owner_user_ids": [],
"listen_channel_ids": [],
"control_channel_ids": [],
"conversation_lease_seconds": 120,
"conversation_lease_seconds": 300,
"officer_only": true,
"runner_url": "http://runner:8780",
"tool_service_url": "http://gateway:8770",
"tool_service_url": "http://192.168.240.2:8770",
"state_dir": "/app/peterbot-data/hermes",
"max_iterations": 30,
"max_tokens": 8192,
Expand Down
98 changes: 98 additions & 0 deletions deploy/housekeeping.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,98 @@
"""Operator CLI for private diagnostics, retention, and snapshot checks (PETER-16).

Subcommands:

diagnose STATE_DIR read-only aggregate health report (JSON)
retention STATE_DIR [--apply ...] retention plan (dry-run by default)
check-snapshot SNAPSHOT STAGING verify snapshot, restore to staging, diagnose

Nothing here touches the network, the model, or Discord. `diagnose` and the
dry-run plan only read; `--apply` first writes a verified snapshot (including
the project-blob cross-check) and aborts before any deletion if that check
fails. Output carries aggregate counts and fixed reason tokens only: never
prompts, answers, memory text, Discord IDs, job IDs, tokens, or paths.

Requires the repository root on sys.path (the container runs from /app; the
bootstrap below also allows direct `python deploy/housekeeping.py`).
"""
from __future__ import annotations

import argparse
import json
import os
from pathlib import Path
import sys

sys.path.insert(0, str(Path(__file__).resolve().parents[1]))

from deploy.state_backup import backup, restore, verify # noqa: E402
from peterbot.operator_ops import ( # noqa: E402
RetentionConfig, diagnose, retention_apply, retention_plan,
)


def _emit(report: object) -> None:
print(json.dumps(report, indent=2, sort_keys=True))


def _config(args: argparse.Namespace) -> RetentionConfig:
return RetentionConfig(
conversations_days=args.conversations_days,
metrics_days=args.metrics_days,
terminal_jobs_days=args.terminal_jobs_days,
settled_receipts_days=args.settled_receipts_days,
include_projects=args.include_projects,
)


def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(description=__doc__.splitlines()[0])
sub = parser.add_subparsers(dest="command", required=True)

diag = sub.add_parser("diagnose", help="read-only aggregate health report")
diag.add_argument("state_dir", type=Path,
default=Path(os.environ.get("PETERBOT_STATE_DIR", "peterbot-data")),
nargs="?")

ret = sub.add_parser("retention", help="retention plan (dry-run) or explicit apply")
ret.add_argument("state_dir", type=Path,
default=Path(os.environ.get("PETERBOT_STATE_DIR", "peterbot-data")),
nargs="?")
ret.add_argument("--apply", action="store_true",
help="delete eligible rows; requires --backup-destination")
ret.add_argument("--backup-destination", type=Path,
help="snapshot written and verified before any deletion")
ret.add_argument("--conversations-days", type=int, default=90)
ret.add_argument("--metrics-days", type=int, default=30)
ret.add_argument("--terminal-jobs-days", type=int, default=90)
ret.add_argument("--settled-receipts-days", type=int, default=180)
ret.add_argument("--include-projects", action="store_true",
help="also run ProjectStore.retention_sweep under its own policy")

snap = sub.add_parser("check-snapshot",
help="verify a snapshot, restore to staging, diagnose the copy")
snap.add_argument("snapshot", type=Path)
snap.add_argument("staging", type=Path)

args = parser.parse_args(argv)
if args.command == "diagnose":
_emit(diagnose(args.state_dir))
elif args.command == "retention":
config = _config(args)
if not args.apply:
_emit(retention_plan(args.state_dir, config))
return 0
if args.backup_destination is None:
parser.error("--apply requires --backup-destination")
backup(args.state_dir, args.backup_destination)
verify(args.backup_destination) # redundant with restore; explicit gate
_emit(retention_apply(args.state_dir, config))
else:
verify(args.snapshot)
restore(args.snapshot, args.staging)
_emit(diagnose(args.staging))
return 0


if __name__ == "__main__":
raise SystemExit(main())
Loading