* chore: dead-code sweep
- Delete socktop_connector/src/connector.rs: orphaned since 08f248c removed
'pub mod connector;' during the modularization refactor. Never compiled
(verified under default, wasm, and workspace feature combos) but shipped
in the crates.io tarball and contained an outdated copy of the TLS
verifier — a trap for anyone patching the pinning bug in the dead copy.
- Delete empty socktop/src/ws.rs, tracked editor backup ui/.modal.rs.backup,
and stray test_thiserror.rs at the repo root.
- Delete the two LEGACY #[allow(dead_code)] process input handlers; the
header-click render test now exercises the live _with_selection handler
instead (better coverage of the real path).
- Drop unused sysinfo dependency from the socktop client.
- Replace stale 'temporarily increased for testing' comment on
COMPRESSION_THRESHOLD (it already held the production value).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(connector): make certificate pinning real; disable Nagle
Security: with --verify-hostname off (the default), the old NoVerify
verifier accepted ANY server certificate — the CA loaded from --tls-ca
was never consulted, so the documented pinning was a no-op and the
connection was trivially MITM-able. Replace it with PinnedCertVerifier:
the presented end-entity cert must be byte-identical to a cert in the
--tls-ca file (any cert in a multi-cert PEM matches, supporting
rotation). Signature validation now uses the ring provider's full
algorithm set instead of a hardcoded 3-scheme list. Empty PEM files
fail fast instead of failing closed per-handshake.
The --verify-hostname path is unchanged (WebPki root-store validation).
Also: the third argument of connect_async_tls_with_config is
tungstenite's disable_nagle flag, not a verification toggle — we were
passing verify_hostname there, leaving Nagle ON for default users. Pass
true unconditionally, and disable Nagle on the plain ws:// path too;
socktop exchanges small request/response frames where Nagle only adds
latency.
Client now consumes the connector via a dual path+version dep so these
fixes are in local builds and CI before the crates.io publish (cargo
strips the path on publish). Connector version -> 1.51.0.
Verified E2E: agent A's cert connects to agent A; agent B's cert
against agent A fails the handshake (the rpi-worker-1 wrong-PEM
scenario); --verify-hostname against a 127.0.0.1 SAN still connects.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(agent): GPU worker thread, async journalctl, correctness + cache fixes
Lightweight:
- GPU collection moves to a dedicated worker thread that owns the gfxinfo
handle for the process lifetime. gfxinfo's active_gpu() runs a full NVML
init/teardown (~20ms, blocking) and we were paying it on the async
runtime for every collect — measured at ~80% of the agent's entire
active CPU on a GPU machine. The handle holds Rc<Nvml> (not Send), so a
thread + mpsc/oneshot channel pair confines it; a zero-total-VRAM reply
is treated as a dead session (driver reload) and re-probed.
- journalctl now runs via tokio::process instead of blocking one of the
two runtime workers for the duration of the subprocess.
- TtlCell (state.rs) replaces the four hand-rolled static TTL caches; a
cached negative result now counts as fresh, so hosts with no matching
temp sensor or GPU stop rescanning every request. Single lock+clone on
the GPU cache hit path (was two).
Correctness:
- Process/child CPU times are now microseconds as documented; they were
milliseconds, rendering 1000x too small next to (correct) thread times.
- Non-Linux per-process CPU%% clamps AFTER dividing by core count; a
4-cores-busy process on an 8-core box reported 12.5% instead of 50%.
- Journal timestamps are real RFC 3339 UTC plus an additive timestamp_us
field (sorting is now numeric); the old strings were Debug-formatted
SystemTime mangled by string replace.
- Partition detection uses /sys/block on Linux: whole-disk filesystems on
names like nvme0n1 or zram1 are no longer misclassified as partitions.
One shared parent_disk_name() replaces two inline copies.
- New sampled_at_ms on the metrics payload (additive) records when the
snapshot was actually collected, so clients can compute exact rates
across the agent's TTL cache.
Security/robustness:
- key.pem is created 0600 (was umask default 0644, world-readable) and
pre-1.51 keys are tightened on startup.
- Per-PID detail/journal caches now evict (60s max age, 64 entries max);
they previously grew without bound under PID-walking clients.
- The two per-PID ws handlers collapse into one generic helper.
- /proc/<pid>/stat parsing unified in one comm-safe module.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(tui): responsive input, request timeouts, poisoned-stream reconnect
R1 — input latency: the event loop drained input once per iteration, then
slept the whole metrics interval; keys and wheel events queued for up to
500ms (or the full interval at slower rates) and applied in bursts. The
input block is extracted to drain_input() and the tail sleep replaced by
a deadline wait in <=33ms poll slices that handles and repaints input the
moment it arrives. Verified: help modal opens <150ms into a 2000ms tick.
R2 — freeze-proofing: metrics/processes/disks requests had no timeout; a
half-dead connection left ws.next() pending forever and froze the TUI
with no way to quit (raw mode eats Ctrl+C as an unread key event). All
requests now carry a 5s budget.
C3 — desync: replies are matched to requests by order alone, so a timed-
out request's late reply would shift every subsequent reply off by one.
Any timeout now treats the stream as poisoned and goes through the
reconnect flow — a fresh stream is aligned by construction. The modal
endpoints additionally mark process details unsupported (flag resets on
modal close/selection change) so a detail-less agent doesn't cause a
reconnect loop. While disconnected the fetch path idles: recovery belongs
to the manual/auto retry paths instead of 5s-timeout hammering.
C7 — fit::truncate_middle_cols replaces util::truncate_middle: display-
width aware and char-boundary safe; the byte-slicing version panicked the
draw loop on non-ASCII device names.
Verified live: agent kill -9 mid-session -> error modal in <3s, q exits
while disconnected, r reconnects and resumes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore: version 1.51.0, path-dep the wasm examples, README notes
- socktop, socktop_agent, socktop_connector -> 1.51.0.
- socktop_wasm_test and zellij_socktop_plugin consume the in-repo
connector via path deps so wasm-feature API drift is caught at PR time
instead of after publish. Immediately proved out: the wasm requests
module needed the new sampled_at_ms field, invisible to native builds.
- zellij plugin gains the standalone [workspace] marker (it could not be
cargo-checked in-tree at all before). NOTE: its lib.rs has pre-existing
compile errors unrelated to the connector (static mut STATE conflicts
with register_plugin!, missing BTreeMap import) — needs its own rework,
out of scope here.
- README: sampled_at_ms in the example payload.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(review): restore Agent Update Required flow, command field, axis alignment
Fixes from Jason's hands-on verification of the branch:
1. Old-agent messaging regression (this branch): a detail-request timeout
went through the loud poison/reconnect flow, burying the ProcessDetails
modal's 'Agent Update Required' message under a connection-error modal.
Old agents IGNORE unknown messages (no late reply, no desync), so the
optional per-PID endpoints now use quiet_reconnect(): swap the stream
silently (still safe against merely-slow agents) and let the modal show
its message. Only a failed reconnect surfaces loudly. Verified against
a real v1.40.0 agent: message shows, session stays healthy.
2. Draw starvation (this branch): an agent that never answers get_metrics
put the loop in fetch->timeout->poison->restart cycles that never
reached the draw call — permanently blank TUI. The iteration now paints
before fetching, and a second consecutive metrics timeout trips a
circuit breaker: persistent 'Agent is not responding' error, recovery
left to the manual/30s retry paths. Verified against a 0.9-era agent.
3. Command & Details pane blank (pre-existing on master): the minimal-
refresh optimization dropped cmd/exe/cwd from the detail endpoint's
ProcessRefreshKind, so process.cmd() had nothing to return. Restored
with UpdateKind::OnlyIfNotSet — immutable values, read once per PID.
Regression test added; journal E2E re-verified (100 entries render).
4. Scatter-plot axis misalignment: Y labels used a fixed 4-char field from
the era when CPU times were 1000x too small; honest millisecond values
(e.g. 136114) blew through it. Labels now right-align to the widest
value per frame and X labels/titles share the dynamic padding.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: journal access notice, 1.60.0, install script, changelog, riscv protoc fallback
- Journal pane now distinguishes 'no entries' from 'no journal access':
journalctl exits 0 with empty output when the agent's user simply can't
see the target's entries (demo mode / user-run agents), explaining
itself only on stderr. The agent forwards that hint as an additive
JournalResponse.notice and the client renders it with practical advice.
Verified E2E via a stub journalctl emulating the unprivileged case.
- Version 1.60.0 across all crates (1.51 would read fine, but the repo's
scheme is 1.40/1.50/…, and a literal 1.6.0 would sort BELOW 1.50.0 in
semver). All user-facing version strings already come from
CARGO_PKG_VERSION — a stale binary was the only way to see an old one.
- scripts/install.sh: build-from-source install/upgrade for the test
fleet (Linux + macOS). Detects in-repo checkouts, installs rustup when
missing, replaces a systemd socktop-agent service binary in place and
restarts it, requires system protoc on riscv64.
- build.rs (agent + connector): fall back to $PROTOC / PATH when
protoc-bin-vendored has no binary for the host (riscv64) — native SBC
builds previously panicked in the build script.
- CHANGELOG.md covering v1.50.0 -> 1.60.0.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: add notice field to cache test initializer
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: untrack zellij plugin build dir; installer updates all PATH copies
- Remove zellij_socktop_plugin/target from git (3,577 files committed by
accident in bf6ac87): the root .gitignore anchors /target to the repo
root, so the standalone plugin's own build dir wasn't covered. Ignore
target/ at any depth (also fixes the pre-existing
'/socktop-wasm-test/target' entry, which pointed at a hyphenated path
that doesn't exist).
- install.sh now updates EVERY copy of socktop/socktop_agent on PATH,
not just $PREFIX: a stale 'cargo install' in ~/.cargo/bin shadows
/usr/local/bin on most PATHs, so an install could 'succeed' while
'socktop --version' kept reporting the old release. Extra copies that
can't be written are warned about, not fatal, and the script now
prints which binary is actually active on PATH at the end.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(installer): manage the socktop-agent systemd service
Upgrade path (unit already present): NEVER touch the unit file — it is
the operator's config (SSL, tokens, ports live there as Environment=
lines). Only the binary at the unit's own ExecStart path is replaced,
then the service restarts. Flags/args preserved by construction.
Fresh path (no unit): full first-time setup mirroring the deb postinst
and the agent-service docs — create the socktop system user/group and
/var/lib/socktop, install docs/socktop-agent.service (ExecStart rewritten
to wherever this run installed the agent; embedded fallback for old
refs), daemon-reload, enable --now, and print how to turn on TLS/token.
Also: system-level operations get their own sudo decision (SYS_SUDO) —
previously they inherited the PREFIX sudo flag, so a writable --prefix
made the service section run groupadd/systemctl unprivileged and die.
No sudo at all now skips service management with a warning instead of
failing the install.
Both branches dry-run verified with stubbed systemctl/sudo.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(installer): don't bind fresh agent services onto occupied ports
The Orange Pi install put the new service straight into a crash-restart
loop: the unit's default --port 3000 collided with a Docker service
already publishing 3000 (Umami; Gitea and friends default there too).
Fresh installs now scan 3000/3001/3010/3231/3232 via ss and configure
the unit on the first free port, warning loudly when 3000 was taken and
printing the resulting ws:// URL. Upgrades still never touch the unit.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(agent): detect NVIDIA GPUs on distros without the unversioned NVML soname
On Debian and derivatives the NVIDIA driver ships only libnvidia-ml.so.1
(the unversioned symlink belongs to the dev package), and nvml-wrapper's
default init dlopens the unversioned name — so gfxinfo reported 'No GPU
found' on a fully functional RTX A2000 host while nvidia-smi worked
fine. Arch-family distros ship the symlink, which is why the desktop
never showed this.
The GPU worker now falls back to initializing NVML directly with the
versioned soname when gfxinfo's probe fails, collecting name/util/vram
through the same handle-caching path. nvml-wrapper was already in the
tree via gfxinfo — same version, no new build cost.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(agent): box the NVML handle variant (clippy large_enum_variant)
CI clippy runs with -D warnings; Nvml is a large struct next to the
16-byte Box<dyn Gpu> variant.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(installer): survive self-modification mid-run; sturdier unit detection
Root cause of the mixed-up second install on the A2000 host: when run
from the clone it manages, the script's own git checkout/merge REPLACES
scripts/install.sh while bash is still executing it. Bash reads scripts
lazily by byte offset, so it resumed parsing the NEW file at the OLD
offset and executed an arbitrary tail of it — observed as the fresh-
service path running on a host whose unit already existed: the port scan
saw the still-running old service on 3000 and silently wrote a new unit
on 3001, while enable --now on the already-active service changed
nothing until a manual daemon-reload.
Fix: the whole script now runs inside main(), invoked as
'main "$@"; exit $?' so bash parses everything up front and never
reads the file again after main returns (the exit lives in the same
parse unit — demonstrated necessary: with a bare 'main "$@"' ending,
bash still executed the swapped file's trailing content after main
returned).
Also: unit existence is now checked with 'systemctl cat' instead of
grepping the full list-unit-files output.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(installer): use a durable ref in the usage example
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
In a short window the fixed root layout runs out of rows and the CPU graph and
per-core bars are what collapse first: the header, gauges and process table hold
fixed heights, so at ~18 rows the top row is left with no drawable interior and
both panes disappear entirely.
Add a second layout, entered automatically once the Disks pane can no longer show
even one complete disk card:
- Disks is dropped — it is the pane that degrades worst when partly drawn.
- Memory and Swap move side by side into the row Disks vacated.
- GPU collapses to a single full-width line (utilisation and VRAM, no device
name), and is omitted entirely on a host that reports no GPU.
- The reclaimed rows go to the CPU panes, with the surplus above their floor
shared with the bottom half so the process table still grows with the window.
`--compact` pins the layout at any size.
The root layout was duplicated in three places (the draw path and both input
hit-testing paths), which would have drifted the moment a second layout existed.
Move it into ui::layout as the single source of truth and have all three callers
go through it.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* Add Debian packaging support with cargo-deb
- Add cargo-deb metadata to socktop and socktop_agent Cargo.toml
- Create systemd service file for socktop_agent
- Add postinst/postrm maintainer scripts for user/group management
- Create GitHub Actions workflow to build .deb packages for AMD64 and ARM64
- Add comprehensive documentation in docs/DEBIAN_PACKAGING.md
- Packages will be available as artifacts on every push
- Automatic GitHub releases for version tags
* Add summary documentation for debian packaging
* fix unit test, move to macro cargo_bin!
* hotfix for issue with socktop agent not creating ssl certificate on first launch after upgrade of axum server version.
* Add helpful post-install message to guide users on enabling socktop-agent service
* Fix CI build by installing libdrm development dependencies
* Fix package rename script - cargo-deb already includes architecture in filename
* Make GPU support optional to enable RISC-V builds without libdrm
- Add 'gpu' feature flag (enabled by default)
- Make gfxinfo dependency optional
- Provide no-op GPU metrics when gpu feature disabled
- Disable GPU support for RISC-V builds in CI (libdrm unavailable)
- All other architectures (amd64, arm64, armhf) still get GPU support
* feature gate GPU stats for arm v7
* specify correct package names.
* install aarch64-linux-gnu-gcc build dep
* specify correct package name
* add RISC-V GCC compiler
* add .cargo to gitignore to elimicate issue with riscv64-linux-gnu-gcc linker in config.toml
* add gcc-arm-linux-gnueabihf linker fore armv7
* set correct x-compile lib gcc-aarch64-linux-gnu for arm64 builds.
* add ports.ubuntu.com to sources
* Add ARM64 as a foreign architecture
* fixe for ARM64 build.
* security.ubuntu.com` aNNOYING
* apt repo github page
* copy output to apt repo
* fix secrets path
* fix secrets path
* change build dep
* Fix postinst message box alignment
* ci(deb): restrict APT publish to v* release tags
Previously the workflow built and published on every push to master
and feature/debian-packaging in addition to v* tags. That meant the
gh-pages APT repo got overwritten on every commit with same-version
.debs, causing apt clients to see a phantom "update available" each
time and burning ~5-10 min of cross-compile CI per merge.
After this change:
- PRs into master still cross-build .debs as a sanity check.
- v* tags build, publish to gh-pages, and create a GitHub release.
- workflow_dispatch remains as the manual escape hatch.
- master pushes no longer trigger this workflow (ci.yml still runs).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Here are concise release notes you can paste into your GitHub release.
Release notes — 2025-08-12
Highlights
Agent back to near-zero CPU when idle (request-driven, no background samplers).
Accurate per-process CPU% via /proc deltas; only top-level processes (threads hidden).
TUI: processes pane gets scrollbar, click-to-sort (CPU% or Mem) with indicator, stable total count.
Network panes made taller; disks slightly reduced.
README revamped: rustup prereqs, crates.io install, update/systemd instructions.
Clippy cleanups across agent and client.
Agent
Reverted precompressed caches and background samplers; WebSocket path is request-driven again.
Ensured on-demand gzip for larger replies; no per-request overhead when small.
Processes: switched to refresh_processes_specifics with ProcessRefreshKind::everything().without_tasks() to exclude threads.
Per-process CPU% now computed from /proc jiffies deltas using a small ProcCpuTracker (fixes “always 0%”/scaling issues).
Optional metrics and light caching:
CPU temp and GPU metrics gated by env (SOCKTOP_AGENT_TEMP=0, SOCKTOP_AGENT_GPU=0).
Tiny TTL caches via once_cell to avoid rescanning sensors every tick.
Dependencies: added once_cell = "1.19".
No API changes to WS endpoints.
Client (TUI)
Processes pane:
Scrollbar (mouse wheel, drag; keyboard arrows/PageUp/PageDown/Home/End).
Click header to sort by CPU% or Mem; dot indicator on active column.
Preserves process_count across fast metrics updates to avoid flicker.
UI/theme:
Shared scrollbar colors moved to ui/theme.rs; both CPU and Processes reuse them.
Cached pane rect to fix input handling; removed unused vars.
Layout: network download/upload get more vertical space; disks shrink slightly.
Clippy fixes: derive Default for ProcSortBy; style/import cleanups.
Docs
README: added rustup install steps (with proper shell reload), install via cargo install socktop and cargo install socktop_agent, and a clear Updating section (systemd service steps included).
Features list updated; roadmap marks independent cadences as done.
Upgrade notes
Agent: cargo install socktop_agent --force, then restart your systemd service; if unit changed, systemctl daemon-reload.
TUI: cargo install socktop --force.
Optional envs to trim overhead: SOCKTOP_AGENT_GPU=0, SOCKTOP_AGENT_TEMP=0.
No config or API breaking changes.
Release highlights
Introduced split client/agent architecture with a ratatui-based TUI and a lightweight WebSocket agent.
Added adaptive (idle-aware) sampler: agent samples fast only when clients are connected; sleeps when idle.
Implemented metrics JSON caching for instant ws replies; cold-start does one-off collection.
Port configuration: --port/-p, positional PORT, or SOCKTOP_PORT env (default 3000).
Optional token auth: SOCKTOP_TOKEN on agent, ws://HOST:PORT/ws?token=VALUE in client.
Logging via tracing with RUST_LOG control.
CI workflow (fmt, clippy, build) for Linux and Windows.
Systemd unit example for always-on agent.
TUI features
CPU: overall sparkline + per-core history with trend arrows and color thresholds.
Memory/Swap gauges with humanized labels.
Disks panel with per-device usage and icons.
Network download/upload sparklines (KB/s) with peak tracking.
Top processes table (PID, name, CPU%, mem, mem%).
Header with hostname and CPU temperature indicator.
Agent changes
sysinfo 0.36.1 targeted refresh: refresh_cpu_all, refresh_memory, refresh_processes_specifics(ProcessesToUpdate::All, ProcessRefreshKind::new().with_cpu().with_memory(), true).
WebSocket handler: client counting with wake notifications, cold-start handling, proper Response returns.
Sampler uses MissedTickBehavior::Skip to avoid catch-up bursts.
Docs
README updates: running instructions, port configuration, optional token auth, platform notes, example JSON.
Added socktop-agent.service systemd unit.
Platform notes
Linux (AMD/Intel) supported; tested on AMD, targeting Intel next.
Raspberry Pi supported (availability of temps varies by model).
Windows builds/run; CPU temperature may be unavailable (shows N/A).
Known/next
Roadmap includes configurable refresh interval, TUI filtering/sorting, TLS/WSS, and export to file.
Add Context...
README.md