Fixes from Jason's hands-on verification of the branch:
1. Old-agent messaging regression (this branch): a detail-request timeout
went through the loud poison/reconnect flow, burying the ProcessDetails
modal's 'Agent Update Required' message under a connection-error modal.
Old agents IGNORE unknown messages (no late reply, no desync), so the
optional per-PID endpoints now use quiet_reconnect(): swap the stream
silently (still safe against merely-slow agents) and let the modal show
its message. Only a failed reconnect surfaces loudly. Verified against
a real v1.40.0 agent: message shows, session stays healthy.
2. Draw starvation (this branch): an agent that never answers get_metrics
put the loop in fetch->timeout->poison->restart cycles that never
reached the draw call — permanently blank TUI. The iteration now paints
before fetching, and a second consecutive metrics timeout trips a
circuit breaker: persistent 'Agent is not responding' error, recovery
left to the manual/30s retry paths. Verified against a 0.9-era agent.
3. Command & Details pane blank (pre-existing on master): the minimal-
refresh optimization dropped cmd/exe/cwd from the detail endpoint's
ProcessRefreshKind, so process.cmd() had nothing to return. Restored
with UpdateKind::OnlyIfNotSet — immutable values, read once per PID.
Regression test added; journal E2E re-verified (100 entries render).
4. Scatter-plot axis misalignment: Y labels used a fixed 4-char field from
the era when CPU times were 1000x too small; honest millisecond values
(e.g. 136114) blew through it. Labels now right-align to the widest
value per frame and X labels/titles share the dynamic padding.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Lightweight:
- GPU collection moves to a dedicated worker thread that owns the gfxinfo
handle for the process lifetime. gfxinfo's active_gpu() runs a full NVML
init/teardown (~20ms, blocking) and we were paying it on the async
runtime for every collect — measured at ~80% of the agent's entire
active CPU on a GPU machine. The handle holds Rc<Nvml> (not Send), so a
thread + mpsc/oneshot channel pair confines it; a zero-total-VRAM reply
is treated as a dead session (driver reload) and re-probed.
- journalctl now runs via tokio::process instead of blocking one of the
two runtime workers for the duration of the subprocess.
- TtlCell (state.rs) replaces the four hand-rolled static TTL caches; a
cached negative result now counts as fresh, so hosts with no matching
temp sensor or GPU stop rescanning every request. Single lock+clone on
the GPU cache hit path (was two).
Correctness:
- Process/child CPU times are now microseconds as documented; they were
milliseconds, rendering 1000x too small next to (correct) thread times.
- Non-Linux per-process CPU%% clamps AFTER dividing by core count; a
4-cores-busy process on an 8-core box reported 12.5% instead of 50%.
- Journal timestamps are real RFC 3339 UTC plus an additive timestamp_us
field (sorting is now numeric); the old strings were Debug-formatted
SystemTime mangled by string replace.
- Partition detection uses /sys/block on Linux: whole-disk filesystems on
names like nvme0n1 or zram1 are no longer misclassified as partitions.
One shared parent_disk_name() replaces two inline copies.
- New sampled_at_ms on the metrics payload (additive) records when the
snapshot was actually collected, so clients can compute exact rates
across the agent's TTL cache.
Security/robustness:
- key.pem is created 0600 (was umask default 0644, world-readable) and
pre-1.51 keys are tightened on startup.
- Per-PID detail/journal caches now evict (60s max age, 64 entries max);
they previously grew without bound under PID-walking clients.
- The two per-PID ws handlers collapse into one generic helper.
- /proc/<pid>/stat parsing unified in one comm-safe module.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This check in offers alpha support for per process metrics, you can view threads, process CPU usage over time, IO, memory, CPU time, parent process, command, uptime and journal entries. This is unfinished but all major functionality is available and I wanted to make it available to feedback and testing.