- Journal pane now distinguishes 'no entries' from 'no journal access':
journalctl exits 0 with empty output when the agent's user simply can't
see the target's entries (demo mode / user-run agents), explaining
itself only on stderr. The agent forwards that hint as an additive
JournalResponse.notice and the client renders it with practical advice.
Verified E2E via a stub journalctl emulating the unprivileged case.
- Version 1.60.0 across all crates (1.51 would read fine, but the repo's
scheme is 1.40/1.50/…, and a literal 1.6.0 would sort BELOW 1.50.0 in
semver). All user-facing version strings already come from
CARGO_PKG_VERSION — a stale binary was the only way to see an old one.
- scripts/install.sh: build-from-source install/upgrade for the test
fleet (Linux + macOS). Detects in-repo checkouts, installs rustup when
missing, replaces a systemd socktop-agent service binary in place and
restarts it, requires system protoc on riscv64.
- build.rs (agent + connector): fall back to $PROTOC / PATH when
protoc-bin-vendored has no binary for the host (riscv64) — native SBC
builds previously panicked in the build script.
- CHANGELOG.md covering v1.50.0 -> 1.60.0.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Fixes from Jason's hands-on verification of the branch:
1. Old-agent messaging regression (this branch): a detail-request timeout
went through the loud poison/reconnect flow, burying the ProcessDetails
modal's 'Agent Update Required' message under a connection-error modal.
Old agents IGNORE unknown messages (no late reply, no desync), so the
optional per-PID endpoints now use quiet_reconnect(): swap the stream
silently (still safe against merely-slow agents) and let the modal show
its message. Only a failed reconnect surfaces loudly. Verified against
a real v1.40.0 agent: message shows, session stays healthy.
2. Draw starvation (this branch): an agent that never answers get_metrics
put the loop in fetch->timeout->poison->restart cycles that never
reached the draw call — permanently blank TUI. The iteration now paints
before fetching, and a second consecutive metrics timeout trips a
circuit breaker: persistent 'Agent is not responding' error, recovery
left to the manual/30s retry paths. Verified against a 0.9-era agent.
3. Command & Details pane blank (pre-existing on master): the minimal-
refresh optimization dropped cmd/exe/cwd from the detail endpoint's
ProcessRefreshKind, so process.cmd() had nothing to return. Restored
with UpdateKind::OnlyIfNotSet — immutable values, read once per PID.
Regression test added; journal E2E re-verified (100 entries render).
4. Scatter-plot axis misalignment: Y labels used a fixed 4-char field from
the era when CPU times were 1000x too small; honest millisecond values
(e.g. 136114) blew through it. Labels now right-align to the widest
value per frame and X labels/titles share the dynamic padding.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- socktop, socktop_agent, socktop_connector -> 1.51.0.
- socktop_wasm_test and zellij_socktop_plugin consume the in-repo
connector via path deps so wasm-feature API drift is caught at PR time
instead of after publish. Immediately proved out: the wasm requests
module needed the new sampled_at_ms field, invisible to native builds.
- zellij plugin gains the standalone [workspace] marker (it could not be
cargo-checked in-tree at all before). NOTE: its lib.rs has pre-existing
compile errors unrelated to the connector (static mut STATE conflicts
with register_plugin!, missing BTreeMap import) — needs its own rework,
out of scope here.
- README: sampled_at_ms in the example payload.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
R1 — input latency: the event loop drained input once per iteration, then
slept the whole metrics interval; keys and wheel events queued for up to
500ms (or the full interval at slower rates) and applied in bursts. The
input block is extracted to drain_input() and the tail sleep replaced by
a deadline wait in <=33ms poll slices that handles and repaints input the
moment it arrives. Verified: help modal opens <150ms into a 2000ms tick.
R2 — freeze-proofing: metrics/processes/disks requests had no timeout; a
half-dead connection left ws.next() pending forever and froze the TUI
with no way to quit (raw mode eats Ctrl+C as an unread key event). All
requests now carry a 5s budget.
C3 — desync: replies are matched to requests by order alone, so a timed-
out request's late reply would shift every subsequent reply off by one.
Any timeout now treats the stream as poisoned and goes through the
reconnect flow — a fresh stream is aligned by construction. The modal
endpoints additionally mark process details unsupported (flag resets on
modal close/selection change) so a detail-less agent doesn't cause a
reconnect loop. While disconnected the fetch path idles: recovery belongs
to the manual/auto retry paths instead of 5s-timeout hammering.
C7 — fit::truncate_middle_cols replaces util::truncate_middle: display-
width aware and char-boundary safe; the byte-slicing version panicked the
draw loop on non-ASCII device names.
Verified live: agent kill -9 mid-session -> error modal in <3s, q exits
while disconnected, r reconnects and resumes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Lightweight:
- GPU collection moves to a dedicated worker thread that owns the gfxinfo
handle for the process lifetime. gfxinfo's active_gpu() runs a full NVML
init/teardown (~20ms, blocking) and we were paying it on the async
runtime for every collect — measured at ~80% of the agent's entire
active CPU on a GPU machine. The handle holds Rc<Nvml> (not Send), so a
thread + mpsc/oneshot channel pair confines it; a zero-total-VRAM reply
is treated as a dead session (driver reload) and re-probed.
- journalctl now runs via tokio::process instead of blocking one of the
two runtime workers for the duration of the subprocess.
- TtlCell (state.rs) replaces the four hand-rolled static TTL caches; a
cached negative result now counts as fresh, so hosts with no matching
temp sensor or GPU stop rescanning every request. Single lock+clone on
the GPU cache hit path (was two).
Correctness:
- Process/child CPU times are now microseconds as documented; they were
milliseconds, rendering 1000x too small next to (correct) thread times.
- Non-Linux per-process CPU%% clamps AFTER dividing by core count; a
4-cores-busy process on an 8-core box reported 12.5% instead of 50%.
- Journal timestamps are real RFC 3339 UTC plus an additive timestamp_us
field (sorting is now numeric); the old strings were Debug-formatted
SystemTime mangled by string replace.
- Partition detection uses /sys/block on Linux: whole-disk filesystems on
names like nvme0n1 or zram1 are no longer misclassified as partitions.
One shared parent_disk_name() replaces two inline copies.
- New sampled_at_ms on the metrics payload (additive) records when the
snapshot was actually collected, so clients can compute exact rates
across the agent's TTL cache.
Security/robustness:
- key.pem is created 0600 (was umask default 0644, world-readable) and
pre-1.51 keys are tightened on startup.
- Per-PID detail/journal caches now evict (60s max age, 64 entries max);
they previously grew without bound under PID-walking clients.
- The two per-PID ws handlers collapse into one generic helper.
- /proc/<pid>/stat parsing unified in one comm-safe module.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Security: with --verify-hostname off (the default), the old NoVerify
verifier accepted ANY server certificate — the CA loaded from --tls-ca
was never consulted, so the documented pinning was a no-op and the
connection was trivially MITM-able. Replace it with PinnedCertVerifier:
the presented end-entity cert must be byte-identical to a cert in the
--tls-ca file (any cert in a multi-cert PEM matches, supporting
rotation). Signature validation now uses the ring provider's full
algorithm set instead of a hardcoded 3-scheme list. Empty PEM files
fail fast instead of failing closed per-handshake.
The --verify-hostname path is unchanged (WebPki root-store validation).
Also: the third argument of connect_async_tls_with_config is
tungstenite's disable_nagle flag, not a verification toggle — we were
passing verify_hostname there, leaving Nagle ON for default users. Pass
true unconditionally, and disable Nagle on the plain ws:// path too;
socktop exchanges small request/response frames where Nagle only adds
latency.
Client now consumes the connector via a dual path+version dep so these
fixes are in local builds and CI before the crates.io publish (cargo
strips the path on publish). Connector version -> 1.51.0.
Verified E2E: agent A's cert connects to agent A; agent B's cert
against agent A fails the handshake (the rpi-worker-1 wrong-PEM
scenario); --verify-hostname against a 127.0.0.1 SAN still connects.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Delete socktop_connector/src/connector.rs: orphaned since 08f248c removed
'pub mod connector;' during the modularization refactor. Never compiled
(verified under default, wasm, and workspace feature combos) but shipped
in the crates.io tarball and contained an outdated copy of the TLS
verifier — a trap for anyone patching the pinning bug in the dead copy.
- Delete empty socktop/src/ws.rs, tracked editor backup ui/.modal.rs.backup,
and stray test_thiserror.rs at the repo root.
- Delete the two LEGACY #[allow(dead_code)] process input handlers; the
header-click render test now exercises the live _with_selection handler
instead (better coverage of the real path).
- Drop unused sysinfo dependency from the socktop client.
- Replace stale 'temporarily increased for testing' comment on
COMPRESSION_THRESHOLD (it already held the production value).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Three panes painted two independent pieces of text onto one row with nothing
reserving space between them, so below roughly 105 columns the right-hand piece
was simply drawn over the tail of the left one:
socktop — host: cachyos-gaming | 🔒✗ TLS | (a: about⏱ 500ms metrics | 2000ms
The width arithmetic used str::len(), a byte count, so the emoji in these strings
overstated their width and left orphaned glyphs at the right edge as well. The
process table had the same problem in a different form: it handed the layout
solver a fixed, over-constrained column set, so a narrow pane crushed the
percentage-sized Name column to nothing while fixed-width PID and Mem % kept
their full width — losing the one field that identifies a process.
Add ui::fit (measure in terminal columns, truncate on character boundaries, pick
the richest wording that fits) and give each pane a priority ladder:
- Header: drop the key hints, then the TLS/token badges, then the
"socktop — host:" prefix, then the metrics/procs words, and only then shorten
the hostname. Hostname and intervals are what survive longest.
- CPU pane: drop the "CPU Temp:" label, then the now:/avg: labels, then the
average, then the temperature's decimal, and the temperature itself last — a
thermal warning outranks a second decimal place.
- Process table: Name is unconditional; CPU %, then Mem, then PID, then Mem %
are added as the pane widens, so Mem % is the first to go and Name the last.
Also fixes sort-header clicks, which resolved against a Layout that omitted the
column spacing the Table renders with, so a click landed off by up to four
columns. Covered by a test that clicks each label where it is actually drawn.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
In a short window the fixed root layout runs out of rows and the CPU graph and
per-core bars are what collapse first: the header, gauges and process table hold
fixed heights, so at ~18 rows the top row is left with no drawable interior and
both panes disappear entirely.
Add a second layout, entered automatically once the Disks pane can no longer show
even one complete disk card:
- Disks is dropped — it is the pane that degrades worst when partly drawn.
- Memory and Swap move side by side into the row Disks vacated.
- GPU collapses to a single full-width line (utilisation and VRAM, no device
name), and is omitted entirely on a host that reports no GPU.
- The reclaimed rows go to the CPU panes, with the surplus above their floor
shared with the bottom half so the process table still grows with the window.
`--compact` pins the layout at any size.
The root layout was duplicated in three places (the draw path and both input
hit-testing paths), which would have drifted the moment a second layout existed.
Move it into ui::layout as the single source of truth and have all three callers
go through it.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Demo mode spawns a socktop_agent child process, but the agent is a separate
crate that `cargo install socktop` does not pull in. When it was missing, the
raw spawn error propagated to main and printed as
`Error: Os { code: 2, kind: NotFound, message: "No such file or directory" }`,
which gives the user nothing to act on.
Introduce DemoAgentError so a NotFound spawn failure is distinguishable from
other io errors, and print the path we looked for plus the `cargo install
socktop_agent` fix. Other spawn errors still propagate as before.
Co-authored-by: Jason Witty <jason@localhost-live.localdomain>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Resolves all five open Dependabot alerts (GHSA-9f94-5g5w-gf6r,
GHSA-394x-vwmw-crm3, GHSA-hfpc-8r3f-gw53, GHSA-65p9-r9h6-22vj,
GHSA-vw5v-4f2q-w9xf) by moving aws-lc-sys from 0.33.0 to 0.42.0.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
The Windows CI matrix surfaced an `unused_variables` warning at
metrics.rs:846 — `let now = std::time::Instant::now();` was bound
unconditionally but only consumed inside a `#[cfg(feature = "logging")]`
tracing::debug! call.
This block lives in the non-Linux `collect_processes_all`, so the Linux
CI never compiles it and never sees the warning. Same shape as the
`processes_ttl_ms` Windows fix from earlier: a binding whose only
consumer is cfg-gated needs to be cfg-gated too.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* Add Debian packaging support with cargo-deb
- Add cargo-deb metadata to socktop and socktop_agent Cargo.toml
- Create systemd service file for socktop_agent
- Add postinst/postrm maintainer scripts for user/group management
- Create GitHub Actions workflow to build .deb packages for AMD64 and ARM64
- Add comprehensive documentation in docs/DEBIAN_PACKAGING.md
- Packages will be available as artifacts on every push
- Automatic GitHub releases for version tags
* Add summary documentation for debian packaging
* fix unit test, move to macro cargo_bin!
* hotfix for issue with socktop agent not creating ssl certificate on first launch after upgrade of axum server version.
* Add helpful post-install message to guide users on enabling socktop-agent service
* Fix CI build by installing libdrm development dependencies
* Fix package rename script - cargo-deb already includes architecture in filename
* Make GPU support optional to enable RISC-V builds without libdrm
- Add 'gpu' feature flag (enabled by default)
- Make gfxinfo dependency optional
- Provide no-op GPU metrics when gpu feature disabled
- Disable GPU support for RISC-V builds in CI (libdrm unavailable)
- All other architectures (amd64, arm64, armhf) still get GPU support
* feature gate GPU stats for arm v7
* specify correct package names.
* install aarch64-linux-gnu-gcc build dep
* specify correct package name
* add RISC-V GCC compiler
* add .cargo to gitignore to elimicate issue with riscv64-linux-gnu-gcc linker in config.toml
* add gcc-arm-linux-gnueabihf linker fore armv7
* set correct x-compile lib gcc-aarch64-linux-gnu for arm64 builds.
* add ports.ubuntu.com to sources
* Add ARM64 as a foreign architecture
* fixe for ARM64 build.
* security.ubuntu.com` aNNOYING
* apt repo github page
* copy output to apt repo
* fix secrets path
* fix secrets path
* change build dep
* Fix postinst message box alignment
* ci(deb): restrict APT publish to v* release tags
Previously the workflow built and published on every push to master
and feature/debian-packaging in addition to v* tags. That meant the
gh-pages APT repo got overwritten on every commit with same-version
.debs, causing apt clients to see a phantom "update available" each
time and burning ~5-10 min of cross-compile CI per merge.
After this change:
- PRs into master still cross-build .debs as a sanity check.
- v* tags build, publish to gh-pages, and create a GitHub release.
- workflow_dispatch remains as the manual escape hatch.
- master pushes no longer trigger this workflow (ci.yml still runs).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore: update ratatui from 0.28 to 0.30
* style: cargo fmt
* fix: replace manual zero-guarded divisions with checked_div
* fix: collapse nested if into match guard
* style: cargo fmt
* bump crossterm and optimize various types, remove stale code.
* fix windows build
* only show parent level processes on main tui
- Add Up/Down arrow key handling in help modal
- Display scrollbar when content exceeds viewport
- Update title to indicate scrollability
- Fixes content cutoff on small terminal windows
This commit implements several major improvements to the TUI experience:
1. CPU Average Display in Main Window
- Show average CPU usage over monitoring period alongside current value
- Format: "CPU avg (now: 45.2% | avg: 52.3%)"
- Helps identify sustained vs momentary CPU spikes
2. Max Memory Tracking in Process Details Modal
- Track and display peak memory usage since monitoring started
- Shown as "Max Memory: 67.8 MB" in yellow for emphasis
- Helps identify memory leaks and usage patterns
- Resets when switching to different process
3. Fuzzy Process Search
- Press / to activate search mode with bordered search box
- Type to fuzzy-match process names (case-insensitive)
- Press Enter to auto-select first result
- Navigate results with arrow keys while typing
- Press c to clear filter
- Press / again to edit existing search
Search box features:
- Yellow bordered box for high visibility
- Active mode: "Search: query_"
- Filter mode: "Filter: query (press / to edit, c to clear)"
Technical implementation:
- Centralized filtering with get_filtered_sorted_indices()
- Consistent filtering across display, navigation, mouse, and auto-scroll
- Proper content area offset calculation for search box
- Real-time filtering as user types
4. Code Quality Improvements
- Created ProcessDisplayParams and ProcessKeyParams structs
- Created MemoryIoParams struct for process modal rendering
- Reduced function arguments to stay under clippy limits
- Exported get_filtered_sorted_indices for reuse
Files Modified:
- socktop/src/app.rs: Search state, auto-scroll with filtering, max memory tracking
- socktop/src/ui/cpu.rs: CPU average calculation and display
- socktop/src/ui/processes.rs: Fuzzy search, filtering, parameter structs
- socktop/src/ui/modal.rs: Updated help modal with new shortcuts
- socktop/src/ui/modal_process.rs: Max memory display, MemoryIoParams struct
- socktop/src/ui/modal_types.rs: Added max_mem_bytes field
Testing:
- All tests pass
- No clippy warnings
- Cargo fmt applied
- Tested search, navigation, mouse clicks, and auto-scroll
- Verified on both filtered and unfiltered process lists
Breaking Changes:
- None (all changes are additive features)
Closes: (performance monitoring improvements)