* chore: dead-code sweep - Delete socktop_connector/src/connector.rs: orphaned since08f248cremoved 'pub mod connector;' during the modularization refactor. Never compiled (verified under default, wasm, and workspace feature combos) but shipped in the crates.io tarball and contained an outdated copy of the TLS verifier — a trap for anyone patching the pinning bug in the dead copy. - Delete empty socktop/src/ws.rs, tracked editor backup ui/.modal.rs.backup, and stray test_thiserror.rs at the repo root. - Delete the two LEGACY #[allow(dead_code)] process input handlers; the header-click render test now exercises the live _with_selection handler instead (better coverage of the real path). - Drop unused sysinfo dependency from the socktop client. - Replace stale 'temporarily increased for testing' comment on COMPRESSION_THRESHOLD (it already held the production value). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(connector): make certificate pinning real; disable Nagle Security: with --verify-hostname off (the default), the old NoVerify verifier accepted ANY server certificate — the CA loaded from --tls-ca was never consulted, so the documented pinning was a no-op and the connection was trivially MITM-able. Replace it with PinnedCertVerifier: the presented end-entity cert must be byte-identical to a cert in the --tls-ca file (any cert in a multi-cert PEM matches, supporting rotation). Signature validation now uses the ring provider's full algorithm set instead of a hardcoded 3-scheme list. Empty PEM files fail fast instead of failing closed per-handshake. The --verify-hostname path is unchanged (WebPki root-store validation). Also: the third argument of connect_async_tls_with_config is tungstenite's disable_nagle flag, not a verification toggle — we were passing verify_hostname there, leaving Nagle ON for default users. Pass true unconditionally, and disable Nagle on the plain ws:// path too; socktop exchanges small request/response frames where Nagle only adds latency. Client now consumes the connector via a dual path+version dep so these fixes are in local builds and CI before the crates.io publish (cargo strips the path on publish). Connector version -> 1.51.0. Verified E2E: agent A's cert connects to agent A; agent B's cert against agent A fails the handshake (the rpi-worker-1 wrong-PEM scenario); --verify-hostname against a 127.0.0.1 SAN still connects. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(agent): GPU worker thread, async journalctl, correctness + cache fixes Lightweight: - GPU collection moves to a dedicated worker thread that owns the gfxinfo handle for the process lifetime. gfxinfo's active_gpu() runs a full NVML init/teardown (~20ms, blocking) and we were paying it on the async runtime for every collect — measured at ~80% of the agent's entire active CPU on a GPU machine. The handle holds Rc<Nvml> (not Send), so a thread + mpsc/oneshot channel pair confines it; a zero-total-VRAM reply is treated as a dead session (driver reload) and re-probed. - journalctl now runs via tokio::process instead of blocking one of the two runtime workers for the duration of the subprocess. - TtlCell (state.rs) replaces the four hand-rolled static TTL caches; a cached negative result now counts as fresh, so hosts with no matching temp sensor or GPU stop rescanning every request. Single lock+clone on the GPU cache hit path (was two). Correctness: - Process/child CPU times are now microseconds as documented; they were milliseconds, rendering 1000x too small next to (correct) thread times. - Non-Linux per-process CPU%% clamps AFTER dividing by core count; a 4-cores-busy process on an 8-core box reported 12.5% instead of 50%. - Journal timestamps are real RFC 3339 UTC plus an additive timestamp_us field (sorting is now numeric); the old strings were Debug-formatted SystemTime mangled by string replace. - Partition detection uses /sys/block on Linux: whole-disk filesystems on names like nvme0n1 or zram1 are no longer misclassified as partitions. One shared parent_disk_name() replaces two inline copies. - New sampled_at_ms on the metrics payload (additive) records when the snapshot was actually collected, so clients can compute exact rates across the agent's TTL cache. Security/robustness: - key.pem is created 0600 (was umask default 0644, world-readable) and pre-1.51 keys are tightened on startup. - Per-PID detail/journal caches now evict (60s max age, 64 entries max); they previously grew without bound under PID-walking clients. - The two per-PID ws handlers collapse into one generic helper. - /proc/<pid>/stat parsing unified in one comm-safe module. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(tui): responsive input, request timeouts, poisoned-stream reconnect R1 — input latency: the event loop drained input once per iteration, then slept the whole metrics interval; keys and wheel events queued for up to 500ms (or the full interval at slower rates) and applied in bursts. The input block is extracted to drain_input() and the tail sleep replaced by a deadline wait in <=33ms poll slices that handles and repaints input the moment it arrives. Verified: help modal opens <150ms into a 2000ms tick. R2 — freeze-proofing: metrics/processes/disks requests had no timeout; a half-dead connection left ws.next() pending forever and froze the TUI with no way to quit (raw mode eats Ctrl+C as an unread key event). All requests now carry a 5s budget. C3 — desync: replies are matched to requests by order alone, so a timed- out request's late reply would shift every subsequent reply off by one. Any timeout now treats the stream as poisoned and goes through the reconnect flow — a fresh stream is aligned by construction. The modal endpoints additionally mark process details unsupported (flag resets on modal close/selection change) so a detail-less agent doesn't cause a reconnect loop. While disconnected the fetch path idles: recovery belongs to the manual/auto retry paths instead of 5s-timeout hammering. C7 — fit::truncate_middle_cols replaces util::truncate_middle: display- width aware and char-boundary safe; the byte-slicing version panicked the draw loop on non-ASCII device names. Verified live: agent kill -9 mid-session -> error modal in <3s, q exits while disconnected, r reconnects and resumes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: version 1.51.0, path-dep the wasm examples, README notes - socktop, socktop_agent, socktop_connector -> 1.51.0. - socktop_wasm_test and zellij_socktop_plugin consume the in-repo connector via path deps so wasm-feature API drift is caught at PR time instead of after publish. Immediately proved out: the wasm requests module needed the new sampled_at_ms field, invisible to native builds. - zellij plugin gains the standalone [workspace] marker (it could not be cargo-checked in-tree at all before). NOTE: its lib.rs has pre-existing compile errors unrelated to the connector (static mut STATE conflicts with register_plugin!, missing BTreeMap import) — needs its own rework, out of scope here. - README: sampled_at_ms in the example payload. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(review): restore Agent Update Required flow, command field, axis alignment Fixes from Jason's hands-on verification of the branch: 1. Old-agent messaging regression (this branch): a detail-request timeout went through the loud poison/reconnect flow, burying the ProcessDetails modal's 'Agent Update Required' message under a connection-error modal. Old agents IGNORE unknown messages (no late reply, no desync), so the optional per-PID endpoints now use quiet_reconnect(): swap the stream silently (still safe against merely-slow agents) and let the modal show its message. Only a failed reconnect surfaces loudly. Verified against a real v1.40.0 agent: message shows, session stays healthy. 2. Draw starvation (this branch): an agent that never answers get_metrics put the loop in fetch->timeout->poison->restart cycles that never reached the draw call — permanently blank TUI. The iteration now paints before fetching, and a second consecutive metrics timeout trips a circuit breaker: persistent 'Agent is not responding' error, recovery left to the manual/30s retry paths. Verified against a 0.9-era agent. 3. Command & Details pane blank (pre-existing on master): the minimal- refresh optimization dropped cmd/exe/cwd from the detail endpoint's ProcessRefreshKind, so process.cmd() had nothing to return. Restored with UpdateKind::OnlyIfNotSet — immutable values, read once per PID. Regression test added; journal E2E re-verified (100 entries render). 4. Scatter-plot axis misalignment: Y labels used a fixed 4-char field from the era when CPU times were 1000x too small; honest millisecond values (e.g. 136114) blew through it. Labels now right-align to the widest value per frame and X labels/titles share the dynamic padding. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: journal access notice, 1.60.0, install script, changelog, riscv protoc fallback - Journal pane now distinguishes 'no entries' from 'no journal access': journalctl exits 0 with empty output when the agent's user simply can't see the target's entries (demo mode / user-run agents), explaining itself only on stderr. The agent forwards that hint as an additive JournalResponse.notice and the client renders it with practical advice. Verified E2E via a stub journalctl emulating the unprivileged case. - Version 1.60.0 across all crates (1.51 would read fine, but the repo's scheme is 1.40/1.50/…, and a literal 1.6.0 would sort BELOW 1.50.0 in semver). All user-facing version strings already come from CARGO_PKG_VERSION — a stale binary was the only way to see an old one. - scripts/install.sh: build-from-source install/upgrade for the test fleet (Linux + macOS). Detects in-repo checkouts, installs rustup when missing, replaces a systemd socktop-agent service binary in place and restarts it, requires system protoc on riscv64. - build.rs (agent + connector): fall back to $PROTOC / PATH when protoc-bin-vendored has no binary for the host (riscv64) — native SBC builds previously panicked in the build script. - CHANGELOG.md covering v1.50.0 -> 1.60.0. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: add notice field to cache test initializer Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: untrack zellij plugin build dir; installer updates all PATH copies - Remove zellij_socktop_plugin/target from git (3,577 files committed by accident inbf6ac87): the root .gitignore anchors /target to the repo root, so the standalone plugin's own build dir wasn't covered. Ignore target/ at any depth (also fixes the pre-existing '/socktop-wasm-test/target' entry, which pointed at a hyphenated path that doesn't exist). - install.sh now updates EVERY copy of socktop/socktop_agent on PATH, not just $PREFIX: a stale 'cargo install' in ~/.cargo/bin shadows /usr/local/bin on most PATHs, so an install could 'succeed' while 'socktop --version' kept reporting the old release. Extra copies that can't be written are warned about, not fatal, and the script now prints which binary is actually active on PATH at the end. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(installer): manage the socktop-agent systemd service Upgrade path (unit already present): NEVER touch the unit file — it is the operator's config (SSL, tokens, ports live there as Environment= lines). Only the binary at the unit's own ExecStart path is replaced, then the service restarts. Flags/args preserved by construction. Fresh path (no unit): full first-time setup mirroring the deb postinst and the agent-service docs — create the socktop system user/group and /var/lib/socktop, install docs/socktop-agent.service (ExecStart rewritten to wherever this run installed the agent; embedded fallback for old refs), daemon-reload, enable --now, and print how to turn on TLS/token. Also: system-level operations get their own sudo decision (SYS_SUDO) — previously they inherited the PREFIX sudo flag, so a writable --prefix made the service section run groupadd/systemctl unprivileged and die. No sudo at all now skips service management with a warning instead of failing the install. Both branches dry-run verified with stubbed systemctl/sudo. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(installer): don't bind fresh agent services onto occupied ports The Orange Pi install put the new service straight into a crash-restart loop: the unit's default --port 3000 collided with a Docker service already publishing 3000 (Umami; Gitea and friends default there too). Fresh installs now scan 3000/3001/3010/3231/3232 via ss and configure the unit on the first free port, warning loudly when 3000 was taken and printing the resulting ws:// URL. Upgrades still never touch the unit. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(agent): detect NVIDIA GPUs on distros without the unversioned NVML soname On Debian and derivatives the NVIDIA driver ships only libnvidia-ml.so.1 (the unversioned symlink belongs to the dev package), and nvml-wrapper's default init dlopens the unversioned name — so gfxinfo reported 'No GPU found' on a fully functional RTX A2000 host while nvidia-smi worked fine. Arch-family distros ship the symlink, which is why the desktop never showed this. The GPU worker now falls back to initializing NVML directly with the versioned soname when gfxinfo's probe fails, collecting name/util/vram through the same handle-caching path. nvml-wrapper was already in the tree via gfxinfo — same version, no new build cost. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(agent): box the NVML handle variant (clippy large_enum_variant) CI clippy runs with -D warnings; Nvml is a large struct next to the 16-byte Box<dyn Gpu> variant. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(installer): survive self-modification mid-run; sturdier unit detection Root cause of the mixed-up second install on the A2000 host: when run from the clone it manages, the script's own git checkout/merge REPLACES scripts/install.sh while bash is still executing it. Bash reads scripts lazily by byte offset, so it resumed parsing the NEW file at the OLD offset and executed an arbitrary tail of it — observed as the fresh- service path running on a host whose unit already existed: the port scan saw the still-running old service on 3000 and silently wrote a new unit on 3001, while enable --now on the already-active service changed nothing until a manual daemon-reload. Fix: the whole script now runs inside main(), invoked as 'main "$@"; exit $?' so bash parses everything up front and never reads the file again after main returns (the exit lives in the same parse unit — demonstrated necessary: with a bare 'main "$@"' ending, bash still executed the swapped file's trailing content after main returned). Also: unit existence is now checked with 'systemctl cat' instead of grepping the full list-unit-files output. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(installer): use a durable ref in the usage example Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
socktop_agent (server)
Lightweight on‑demand metrics WebSocket server for the socktop TUI.
Highlights:
- Collects system metrics only when requested (keeps idle CPU <1%)
- Optional TLS (self‑signed cert auto‑generated & pinned by client)
- JSON for fast metrics / disks; protobuf (optionally gzipped) for processes
- Accurate per‑process CPU% on Linux via /proc jiffies delta
- Optional GPU & temperature metrics (disable via env vars)
- Simple token auth (?token=...) support
Run (no TLS):
cargo install socktop_agent
socktop_agent --port 3000
Enable TLS:
SOCKTOP_ENABLE_SSL=1 socktop_agent --port 8443
# cert/key stored under $XDG_DATA_HOME/socktop_agent/tls
Environment toggles:
- SOCKTOP_AGENT_GPU=0 (disable GPU collection)
- SOCKTOP_AGENT_TEMP=0 (disable temperature)
- SOCKTOP_TOKEN=secret (require token param from client)
- SOCKTOP_AGENT_METRICS_TTL_MS=250 (cache fast metrics window)
- SOCKTOP_AGENT_PROCESSES_TTL_MS=1000
- SOCKTOP_AGENT_DISKS_TTL_MS=1000
NOTE ON ENV vars
Generally these have been added for debugging purposes. you do not need to configure them, default values are tuned and GPU will deisable itself after the first poll if not available.
Systemd unit example & full docs: https://github.com/jasonwitty/socktop
WebSocket API Integration Guide
The socktop_agent exposes a WebSocket API that can be directly integrated with your own applications. This allows you to build custom monitoring dashboards or analysis tools using the agent's metrics.
WebSocket Endpoint
ws://HOST:PORT/ws # Without TLS
wss://HOST:PORT/ws # With TLS
With authentication token (if configured):
ws://HOST:PORT/ws?token=YOUR_TOKEN
wss://HOST:PORT/ws?token=YOUR_TOKEN
Communication Protocol
All communication uses JSON format for requests and responses, except for the process list which uses Protocol Buffers (protobuf) format with optional gzip compression.
Request Types
Send a JSON message with a type field to request specific metrics:
{"type": "metrics"} // Request fast-changing metrics (CPU, memory, network)
{"type": "disks"} // Request disk information
{"type": "processes"} // Request process list (returns protobuf)
Response Formats
- Fast Metrics (JSON):
{
"cpu_total": 12.4,
"cpu_per_core": [11.2, 15.7],
"mem_total": 33554432,
"mem_used": 18321408,
"swap_total": 0,
"swap_used": 0,
"hostname": "myserver",
"cpu_temp_c": 42.5,
"networks": [{"name":"eth0","received":12345678,"transmitted":87654321}],
"gpus": [{"name":"nvidia-0","usage":56.7,"memory_total":8589934592,"memory_used":1073741824,"temp_c":65.0}]
}
- Disks (JSON):
[
{"name":"nvme0n1p2","total":512000000000,"available":320000000000},
{"name":"sda1","total":1000000000000,"available":750000000000}
]
- Processes (Protocol Buffers):
Processes are returned in Protocol Buffers format, optionally gzip-compressed for large process lists. The protobuf schema is:
syntax = "proto3";
message Process {
uint32 pid = 1;
string name = 2;
float cpu_usage = 3;
uint64 mem_bytes = 4;
}
message ProcessList {
uint32 process_count = 1;
repeated Process processes = 2;
}
Example Integration (JavaScript/Node.js)
const WebSocket = require('ws');
// Connect to the agent
const ws = new WebSocket('ws://localhost:3000/ws');
ws.on('open', function open() {
console.log('Connected to socktop_agent');
// Request metrics immediately on connection
ws.send(JSON.stringify({type: 'metrics'}));
// Set up regular polling
setInterval(() => {
ws.send(JSON.stringify({type: 'metrics'}));
}, 1000);
// Request processes every 3 seconds
setInterval(() => {
ws.send(JSON.stringify({type: 'processes'}));
}, 3000);
});
ws.on('message', function incoming(data) {
// Check if the response is JSON or binary (protobuf)
try {
const jsonData = JSON.parse(data);
console.log('Received JSON data:', jsonData);
} catch (e) {
console.log('Received binary data (protobuf), length:', data.length);
// Process binary protobuf data with a library like protobufjs
}
});
ws.on('close', function close() {
console.log('Disconnected from socktop_agent');
});
Example Integration (Python)
import json
import asyncio
import websockets
async def monitor_system():
uri = "ws://localhost:3000/ws"
async with websockets.connect(uri) as websocket:
print("Connected to socktop_agent")
# Request initial metrics
await websocket.send(json.dumps({"type": "metrics"}))
# Set up regular polling
while True:
# Request metrics
await websocket.send(json.dumps({"type": "metrics"}))
# Receive and process response
response = await websocket.recv()
# Check if response is JSON or binary (protobuf)
try:
data = json.loads(response)
print(f"CPU: {data['cpu_total']}%, Memory: {data['mem_used']/data['mem_total']*100:.1f}%")
except json.JSONDecodeError:
print(f"Received binary data, length: {len(response)}")
# Process binary protobuf data with a library like protobuf
# Wait before next poll
await asyncio.sleep(1)
asyncio.run(monitor_system())
Notes for Integration
-
Error Handling: The WebSocket connection may close unexpectedly; implement reconnection logic in your client.
-
Rate Limiting: Avoid excessive polling that could impact the system being monitored. Recommended intervals:
- Metrics: 500ms or slower
- Processes: 2000ms or slower
- Disks: 5000ms or slower
-
Authentication: If the agent is configured with a token, always include it in the WebSocket URL.
-
Protocol Buffers Handling: For processing the binary process list data, use a Protocol Buffers library for your language and the schema provided in the
proto/processes.protofile. -
Compression: Process lists may be gzip-compressed. Check if the response starts with the gzip magic bytes (
0x1f, 0x8b) and decompress if necessary.
LLM Integration Guide
If you're using an LLM to generate code for integrating with socktop_agent, this section provides structured information to help the model understand the API better.
API Schema
# WebSocket API Schema for socktop_agent
endpoint: ws://HOST:PORT/ws or wss://HOST:PORT/ws (with TLS)
authentication:
type: query parameter
parameter: token
example: ws://HOST:PORT/ws?token=YOUR_TOKEN
requests:
- type: metrics
format: JSON
example: {"type": "metrics"}
description: Fast-changing metrics (CPU, memory, network)
- type: disks
format: JSON
example: {"type": "disks"}
description: Disk information
- type: processes
format: JSON
example: {"type": "processes"}
description: Process list (returns protobuf)
responses:
- request_type: metrics
format: JSON
schema:
cpu_total: float # percentage of total CPU usage
cpu_per_core: [float] # array of per-core CPU usage percentages
mem_total: uint64 # total memory in bytes
mem_used: uint64 # used memory in bytes
swap_total: uint64 # total swap in bytes
swap_used: uint64 # used swap in bytes
hostname: string # system hostname
cpu_temp_c: float? # CPU temperature in Celsius (optional)
networks: [
{
name: string # network interface name
received: uint64 # total bytes received
transmitted: uint64 # total bytes transmitted
}
]
gpus: [
{
name: string # GPU device name
usage: float # GPU usage percentage
memory_total: uint64 # total GPU memory in bytes
memory_used: uint64 # used GPU memory in bytes
temp_c: float # GPU temperature in Celsius
}
]?
- request_type: disks
format: JSON
schema:
[
{
name: string # disk name
total: uint64 # total space in bytes
available: uint64 # available space in bytes
}
]
- request_type: processes
format: Protocol Buffers (optionally gzip-compressed)
schema: See protobuf definition below
Protobuf Schema (processes.proto)
syntax = "proto3";
message Process {
uint32 pid = 1;
string name = 2;
float cpu_usage = 3;
uint64 mem_bytes = 4;
}
message ProcessList {
uint32 process_count = 1;
repeated Process processes = 2;
}
Step-by-Step Integration Pseudocode
1. Establish WebSocket connection to ws://HOST:PORT/ws
- Add token if required: ws://HOST:PORT/ws?token=YOUR_TOKEN
2. For regular metrics updates:
- Send: {"type": "metrics"}
- Parse JSON response
- Extract CPU, memory, network info
3. For disk information:
- Send: {"type": "disks"}
- Parse JSON response
- Extract disk usage data
4. For process list:
- Send: {"type": "processes"}
- Check if response is binary
- If starts with 0x1f, 0x8b bytes:
- Decompress using gzip
- Parse binary data using protobuf schema
- Extract process information
5. Implement reconnection logic:
- On connection close/error
- Use exponential backoff
6. Respect rate limits:
- metrics: ≥ 500ms interval
- disks: ≥ 5000ms interval
- processes: ≥ 2000ms interval
Common Implementation Patterns
Pattern 1: Periodic Polling
// Set up separate timers for different metric types
const metricsInterval = setInterval(() => ws.send(JSON.stringify({type: 'metrics'})), 500);
const disksInterval = setInterval(() => ws.send(JSON.stringify({type: 'disks'})), 5000);
const processesInterval = setInterval(() => ws.send(JSON.stringify({type: 'processes'})), 2000);
// Clean up on disconnect
ws.on('close', () => {
clearInterval(metricsInterval);
clearInterval(disksInterval);
clearInterval(processesInterval);
});
Pattern 2: Processing Binary Protobuf Data
// Using protobufjs
const root = protobuf.loadSync('processes.proto');
const ProcessList = root.lookupType('ProcessList');
ws.on('message', function(data) {
if (typeof data !== 'string') {
// Check for gzip compression
if (data[0] === 0x1f && data[1] === 0x8b) {
data = gunzipSync(data); // Use appropriate decompression library
}
// Decode protobuf
const processes = ProcessList.decode(new Uint8Array(data));
console.log(`Total processes: ${processes.process_count}`);
processes.processes.forEach(p => {
console.log(`PID: ${p.pid}, Name: ${p.name}, CPU: ${p.cpu_usage}%`);
});
}
});
Pattern 3: Reconnection Logic
function connect() {
const ws = new WebSocket('ws://localhost:3000/ws');
ws.on('open', () => {
console.log('Connected');
// Start polling
});
ws.on('close', () => {
console.log('Connection lost, reconnecting...');
setTimeout(connect, 1000); // Reconnect after 1 second
});
// Handle other events...
}
connect();