Commit Graph

1 Commits

Author SHA1 Message Date
jasonwitty 1114482046 Fix actix worker wedge on session teardown; grant CAP_KILL; bump to 0.3.12
Build and Deploy to K3s / test (push) Successful in 2m6s
Build and Deploy to K3s / lint (push) Successful in 1m2s
Build and Deploy to K3s / build-and-push (push) Successful in 9m23s
Build and Deploy to K3s / deploy (push) Successful in 2m13s
Terminal::stopping did child.kill() (error ignored) and then a blocking
child.wait() on the actix worker thread. Since the 0.3.11 privilege drop,
sessions run as `demo` and the server had no CAP_KILL, so kill() failed
with EPERM and wait() blocked forever. With two workers, exactly every
other request to :8082 then timed out — 5 days of readiness/liveness
flapping on socktop.io and orphaned restricted-shell/socktop processes
piling up in the pods.

- Signal the session's whole process group (portable-pty setsid()s the
  child), not just the shell: a non-interactive bash defers signals while
  a foreground command runs, so the old SIGHUP/SIGKILL to the shell alone
  orphaned socktop anyway.
- Reap on a dedicated thread: SIGHUP, 2 s grace, SIGKILL, 10 s deadline,
  then a blocking wait on that thread only. Signal failures are logged.
- Drop the pty handles before handing off the child.
- kubernetes/03-deployment.yaml: add CAP_KILL (sessions still lose every
  cap via setpriv). CI never applies the manifest — it was patched live.
- tests/reaper_tests.rs pins the non-blocking return, process-group kill,
  and SIGHUP→SIGKILL escalation.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-29 17:53:53 -07:00