feat(sandbox): write agent output to the container log - #4005
Conversation
|
🌿 Preview your docs: https://nvidia-preview-pr-4005.docs.buildwithfern.com/openshell |
73f0ef6 to
a7aa90d
Compare
|
I like this change, especially making the agent output available through the normal container logs again. I’m still pretty new to this kind of logging setup, but I was curious why you chose to use a separate forwarder thread and queue for each stream instead of writing directly to the container log? |
|
The channel moves the blocking |
a7aa90d to
68d8a58
Compare
|
🤖
The highlighted risk is that there are things in the sandbox log that shouldn't be copied out to the gateway as part of exit diagnostics; I don't know enough to say if that's an interesting risk or not. |
Since #2726 the canonical main process's stdout and stderr are captured in pipes that feed only the in-memory replay buffer used by sandbox connect. Agent output therefore never reaches the container's own stdout and stderr, so it is missing from kubectl logs, docker logs, and podman logs and from anything that collects container logs. Before #2726 the entrypoint inherited the container's descriptors and its output appeared there. Copy the main process's output to the launcher's stdout and stderr in addition to the replay buffer, restoring the earlier behavior: - Output is copied byte for byte to the matching stream from a forwarder thread per stream, after it is published to the replay buffer. When the container runtime falls behind on a stream, that stream's reader waits instead of dropping output, so backpressure reaches the agent as it did with inherited descriptors, while the other stream and attachments keep receiving output. - Before the main process's exit is published, the output readers finish and queued output is drained to the container log, so an agent's final lines are not lost at shutdown. A 30 second deadline covers both; when it expires, readers waiting on the container log are released and drain the pipes into the replay buffer only, so a stalled container log cannot block exit reporting. - PTY-mode processes are not copied. The terminal stream carries escape sequences and echoed input, and terminal commands never reached the container log before #2726. - Exec, SSH, and SFTP sessions are not copied. Launcher log lines keep their existing format and remain in the container's stderr. They are written as whole lines, and a newline is inserted first when the agent left stderr mid-line, so launcher and agent lines do not merge. The Docker and VM drivers appended the tail of the workload's output to failure messages: Docker the workload container's log, and the VM driver the guest console, which carries the launcher's stdout and stderr. Those messages land in the sandbox's Ready condition and in platform events that the gateway republishes to the sandbox event stream. With agent output in that log, those messages would carry arbitrary agent output, including anything sensitive the agent prints, into gateway status and events. The supervisor starts its health endpoint only after the agent starts, so every Docker failure path could include agent output, and the VM driver reports one whenever the VM or host supervisor exits. Forward only the supervisor's log tail, matching the Podman driver, which reads the workload log solely to match fixed launcher markers and never forwards raw workload output. The workload's output remains available through docker logs and the VM's rootfs-console.log. Document where main process output appears in the logging docs and the cluster debugging skill. Closes #3928 Signed-off-by: Kris Hicks <khicks@nvidia.com>
68d8a58 to
049ca16
Compare
Summary
Since #2726, the sandbox's main process writes its stdout and stderr into pipes that only feed the
sandbox connectreplay buffer, so agent output is missing fromkubectl logs,docker logs, andpodman logs. This PR copies that output to the container's stdout and stderr as well, restoring the pre-#2726 behavior (v0.0.110 and earlier). It also stops the Docker and VM drivers from forwarding the workload's now agent-bearing log into gateway status and events.Related Issue
Closes #3928
Changes
MainSessionpublishes each chunk it reads from the main process to the replay buffer, then copies it to the matching container stream (stdout or stderr) through a separate forwarder thread and bounded queue per stream (container_log.rs).MainSession::finishpublishes the exit, the output readers finish and queued output is drained to the container log, so an agent's final lines are not lost. A single 30-second deadline covers both. When it expires, readers waiting on the container log are released and drain the pipes into the replay buffer only, so a stalled container log cannot block exit reporting.exec, SSH, and SFTP sessions are not copied either.Readycondition and platform events republished to the sandbox event stream) no longer include the workload's output, only the supervisor's log tail. Otherwise arbitrary agent output, including anything sensitive, would land in gateway status and events: up to 80 lines of the Docker workload container's log, or the last 8 KiB of the VM guest console, which carries the launcher's stdout and stderr.warn!that fix(runtime): recover SSH relays and bound startup diagnostics #4011 added to the readiness failure path no longer logs the workload tail, and fix(runtime): recover SSH relays and bound startup diagnostics #4011's 1 KiB cap on status-message tails is kept.ProcessExitedis reported whenever the VM or host supervisor exits, including after the main process ends. It keeps the host supervisor's stderr tail.docker logsand the VM'srootfs-console.log.docs/observability/accessing-logs.mdxdescribes where main process output appears.debug-openshell-clusternotes what the sandbox and supervisor container logs contain.Testing
mise run pre-commitpassescontainer_log: output written unmodified, the newline guard for launcher lines, line assembly from fragmented writes, ordering, waiting on a stalled log without dropping output, stderr unaffected by a stalled stdout, and drain.main_session: pipe output copied to the matching stream, PTY output not copied,finishwaiting for queued output,finishreleasing readers blocked on a saturated and permanently stalled log, stderr and attachments flowing while the stdout log is stalled, and attachments receiving a chunk whose container-log send is waiting.openshell-driver-vm:process_exit_status_omits_guest_console_outputdrivesmonitor_sandboxthrough the VM and host-supervisor exit paths and checks that the condition and platform event omit the guest console while keeping the supervisor's stderr tail.openshell-driver-docker:startup_error_log_tails_fit_grpc_header_budget(from fix(runtime): recover SSH relays and bound startup diagnostics #4011) now builds the status from the supervisor tail only.main_sessiontests. Linux-target clippy also passes foropenshell-sandbox(with and withoutperf-harness) andopenshell-driver-vm.kubectl logs <pod> -c agentshowed the launcher's startup warning followed by the agent's stdout and stderr lines, and the node's container log file tagged each line with the correct stream.openshell sandbox connectstill replayed from the buffer. That revision predates the waiting, PTY, and drain changes and the Docker driver change; those are covered by the unit tests above.Checklist