Server monitoring · open source
The server agent, in the open
Before you run anything as root on your server you should be able to read it. This page lists exactly what the agent reads and sends, shows its source code, and explains how to build it yourself and check that you get the same binary we ship.
- MIT licence
- Go standard library only, no dependencies
- Static binary, about 8 MB
- Reproducible builds
What it reads, and when it is sent
The agent samples locally every 60 seconds. Very little of that goes over the network: see the schedule below.
- CPU usage, load average, memory and swap use, swap-in/out rateRead from:/proc/stat, /proc/loadavg, /proc/meminfo, /proc/vmstatSent:Heartbeat (CPU, memory) every 5 min; per-minute average and peak in the 15-minute rollup
- Disk and inode usage per real filesystem (ext4, xfs, btrfs, zfs…; not tmpfs, overlays, snaps or network mounts), fill rate over the last 6 hours and “full in N hours”Read from:/proc/self/mountinfo, statfs()Sent:Fullest disk and soonest ETA in the heartbeat; all filesystems in the rollup
- Where the disk space goes, every 6 hours (postponed while the server is busy): the biggest directories a few levels deep, the ones that grew, the largest and fastest-growing files (16 MB or more) with a kind (log, database, backup, cache, Docker). Paths and sizes only, never contents; symbolic links are not followed, other mounts are not entered; at most 20 seconds of work per scanRead from:a walk of each real filesystem (directory entries and their allocated size)Sent:Rollup, after each scan
- Which programs write to disk: bytes written per program (name, systemd service, container), the top 10Read from:/proc/PID/io (write_bytes), every 15 minutesSent:Rollup
- Network receive / transmit rate (physical interfaces; not lo, docker, veth or bridges)Read from:/proc/net/devSent:Rollup
- Hostname, OS name, kernel, CPU count, architecture, boot time, uptime, virtualization (KVM, VMware, Docker, LXC…)Read from:/proc/sys/kernel, /etc/os-release, /proc/uptime, /sys/class/dmi/id, /run/systemd/containerSent:Rollup
- Top 10 processes by CPU and by memory: pid, name, user, executable path, systemd service, start time, command line cut to 200 characters with secrets masked (password=…, --token …, Authorization headers, user:pass@, mysql -p… and well-known token formats)Read from:/proc/PID/stat every minute; cmdline, exe and owner when a process appears and every 15 minutes; cgroupSent:Rollup
- Listening TCP ports and UDP services (protocol, address, port) and WHY each one is open: the program behind it, its user, path, systemd service and Docker container (for a port Docker publishes: the container it forwards to)Read from:/proc/net/tcp, tcp6, udp, udp6; /proc/PID/fd, exe, cgroup only when the ports change; Docker's config.v2.json (name and address)Sent:With every change and at least daily; added or removed ports at once
- Logins: who logged in, from which address, on which terminal, when, and when they logged out; who is logged in now; each account's last login; boots and shutdowns with the kernel versionRead from:/var/log/wtmp, /run/utmp, /var/log/lastlog (read incrementally)Sent:Rollup, when something is new
- SSH logins with their method (key or password), failed logins per source address with the time of the last try and the names of accounts that EXIST on the server it tried (names an attacker guesses are never sent), sudo use (who, allowed or refused; never the command)Read from:/var/log/auth.log or /var/log/secure, read incrementallySent:Rollup
- Active crontab lines, with secrets masked like command linesRead from:/etc/crontab, /etc/cron.d, /var/spool/cronSent:Only changes
- SSH keys of root and of accounts with a login shell: type, SHA-256 fingerprint and comment (never the key)Read from:~/.ssh/authorized_keys, authorized_keys2Sent:Only changes
- setuid / setgid files (path, mode, owner), every 6 hoursRead from:/bin, /sbin, /usr/bin, /usr/sbin, /usr/lib, /usr/libexec, /usr/local, /opt, /tmp, /var/tmp, /dev/shmSent:Only changes
- Only if you configure web roots: names of new, changed and removed .php, .phtml, .phar, .inc, .js, .mjs, .html, .htm, .shtml, .htaccess and .user.ini files, every 15 minutes. Compared by modification time and size; the content is never readRead from:the folders you list in web_rootsSent:Only changes (at most 50 names per list)
- Processes killed by the out-of-memory killer, disk I/O per device (busy %, wait per request), TCP connections by state, open files and the conntrack table against their limitsRead from:/proc/vmstat, /proc/diskstats, /proc/net/tcp, /proc/sys/fs/file-nr, nf_conntrack_countSent:Rollup; an event at once when a rule trips
- Failed systemd services and why: description, result, exit code or signal, restarts, since when, and once per failure its last 10 log lines (secrets masked)Read from:systemd over D-Bus (ListUnits and the unit's properties, read-only); /var/log/syslog or /var/log/messages (no journalctl: it runs no commands)Sent:Rollup; an event at once
- Security settings, once an hour: effective sshd settings (root login, passwords, keys, MaxAuthTries, AllowUsers, Match blocks), key counts per account, accounts that can log in or use sudo, extra UID 0 accounts, sudoers NOPASSWD rules, fail2ban and its jails, the host firewall with the ports its rules open, automatic updates and when packages were last updated, databases on a public address, a pending reboot, waiting updates. Setting values only, never file contentsRead from:/etc/ssh/sshd_config (+ Include), /etc/passwd, /etc/group (never /etc/shadow), /etc/sudoers, /etc/fail2ban, /etc/ufw, firewalld zones, /etc/nftables.conf, /etc/iptables, /etc/apt, /var/lib/apt/periodic, /run/reboot-required, update-notifierSent:Rollup, when it changes and at least once a day
- The time sync daemon, and steps of the system clockRead from:the process list, /run/systemd/timesyncSent:Rollup
- The agent's own CPU and memory, so you can check what it costsRead from:/proc/self/statSent:Rollup
- Once a day: a hash of each list above, so we can notice that our copy driftedRead from:–Sent:Rollup, once a day
What goes over the network
Heartbeat · every 5 minutes · about 150 bytes
Time, agent version, uptime, CPU %, memory %, the fullest disk and its projected-full ETA, the rules that are tripped now. Fifteen minutes without one is the “server or agent silent” alert.
Rollup · every 15 minutes · about 2 KB
Per-minute averages and peaks since the previous rollup (CPU, memory, swap, load, network, fullest disk, connections, busiest disk), filesystems, disk I/O, the top processes, the agent's own CPU and memory, and health: failed services with their cause, out-of-memory kills, SSH logins and failures, time sync. When they change (at least once a day): the security settings and the listening ports with the programs behind them. When new: logins, boots, sudo use. After a disk scan: where the space goes. This is what the charts and the server page show.
Event · at once, at most one a minute
When a local rule trips (disk at 90% or full within 48 hours, inodes at 90%, memory at 95% for 10 minutes, heavy swapping for 10 minutes, load above twice the cores for 15 minutes, a process killed for lack of memory, the open-file or conntrack table at 90%, a disk saturated for 15 minutes, a failed service, a security signal) or a list changes (ports, crontab, SSH keys, setuid, web roots), with the evidence.
Inventory · rarely
The full lists of ports, crontab lines, SSH key fingerprints and setuid files: once after install or an upgrade, and again only if our daily hash check says our copy differs.
Measured: about 240 KB of data a day; with HTTPS overhead (a TLS handshake for most messages) roughly 2 MB. Everything is gzip JSON over HTTPS, POST https://approvalens.com/api/agent/v1/report, with the server's token. When it cannot reach us the agent keeps events, inventories and chart data for up to 24 hours (8 MB at most) and sends them later, retrying less and less often.
Checking your public ports from outside
For the ports the agent reports as listening on a public address, Approvalens tries a plain TCP connection from the internet to the address your server reports from: only that address, only those ports (at most 30), at most once every 6 hours, 3-second timeout, nothing is sent over the connection. It tells you whether a firewall really blocks a port. By installing the agent you agree to this check; to switch it off for a server, add reachability to disable: in /etc/approvalens-agent.yaml.
What it never does
No remote commands
It never executes commands or shells out (not even journalctl: a failed service's log lines are read from the syslog file), and never runs code it receives. The only thing it reads from an answer is the HTTP status code: 2xx means delivered, 205 means “send your full lists again”.
No inbound connections
It opens no listening socket. The only traffic is an outgoing HTTPS POST to the configured URL (plain HTTP is refused except to 127.0.0.1). Locally it only asks systemd, over the D-Bus system socket, for its list of units and the properties of failed ones (what systemctl list-units and systemctl status show).
No file contents, no secrets
It never reads the content of your files: setting values only, paths and sizes for disk usage, names for web roots, fingerprints for SSH keys. It never reads /etc/shadow, private keys, other processes' memory or environment. In the detailed mode the unit hides /etc/shadow and the SSH host keys from it and its system call filter blocks ptrace and process_vm_readv/writev.
No changes
It never kills, quarantines or deletes anything. It writes only its own state folder, /var/lib/approvalens-agent (unsent messages for at most 24 hours, what it last sent, where it stopped reading the logs, the web root file list).
No noticeable load
Gentle mode: the one-minute sample reads one small file per process; everything heavier runs on one background worker, one task at a time, at the lowest CPU and I/O priority (nice 19, idle I/O, CPUWeight=1), in 2-millisecond slices with pauses, and waits while your server is busy. CPUQuota=5% and 64 MB of memory are hard caps.
Security signals are heuristics
A match is a reason to look, not proof, and no match proves nothing: this is not an antivirus. The signals are: a process name or argument from the crypto-miner list (signatures.txt below); an executable under /tmp, /var/tmp or /dev/shm; an executable deleted after it started (outside the system folders, which upgrades do legitimately) or running from memory (memfd); one process using a whole core for 15 minutes (databases, compressors, compilers and backup tools excepted); and anything new in the lists above after the first 24 hours (the baseline). You can acknowledge a signal and it joins the baseline.
Permissions
Since 0.3.0 the default is the detailed mode: the agent runs as its own user, approvalens-agent, with two read-only capabilities set by the systemd unit. CAP_DAC_READ_SEARCH lets it read files only root may read (sshd_config, sudoers, the auth log, authorized_keys, other users' crontabs, firewall rules, the syslog file); it grants no write access, and its filesystem is read-only anyway. CAP_SYS_PTRACE is what the kernel checks before it shows another user's /proc/PID/fd, exe and io: the agent uses it only to see which program listens on a port, its path and how much it writes. Because that capability would also allow attaching to processes, the unit's system call filter blocks ptrace, process_vm_readv/writev, kcmp and pidfd_getfd, and /etc/shadow and the SSH host keys are hidden from it. Install with --unprivileged for the minimal mode without capabilities (your account then lists what it cannot see), or with --root (root, limited to the same two capabilities). Upgrading an older unprivileged install switches it to the detailed mode, and the installer says what changes.
Resource limits and systemd hardening
Goal: well under 0.1% of one core on average, no burst above 2% of a core, and under 20 MB of memory, on a server with hundreds of processes. Gentle mode does the rest: heavy work in small slices on one low-priority worker, postponed while your server is busy. The unit enforces hard caps and locks the process down:
[Unit]
Description=Approvalens server agent (reads system statistics, reports over HTTPS)
Documentation=https://approvalens.com/agent
After=network-online.target
Wants=network-online.target
[Service]
# Mode: detailed (the default: its own user, two read-only capabilities)
Type=simple
User=approvalens-agent
ExecStart=/usr/local/bin/approvalens-agent run --config /etc/approvalens-agent.yaml
Restart=always
RestartSec=15
# 78: a configuration error (the log says which line); restarting would not fix it.
RestartPreventExitStatus=78
TimeoutStopSec=20
StateDirectory=approvalens-agent
StateDirectoryMode=0700
# Hardening: everything read-only except its own state directory, no privilege gain.
NoNewPrivileges=yes
ProtectSystem=strict
ProtectHome=read-only
ReadOnlyPaths=/
ReadWritePaths=-/var/lib/approvalens-agent
# Two read-only capabilities: CAP_DAC_READ_SEARCH reads files only root may read (sshd_config,
# sudoers, the auth log, authorized_keys); CAP_SYS_PTRACE is what the kernel checks before it shows
# another user's /proc/<pid>/fd, exe and io links (PTRACE_MODE_READ). Nothing else.
CapabilityBoundingSet=CAP_DAC_READ_SEARCH CAP_SYS_PTRACE
AmbientCapabilities=CAP_DAC_READ_SEARCH CAP_SYS_PTRACE
# Secrets it never needs stay out of reach even with CAP_DAC_READ_SEARCH: password hashes and the
# SSH host keys are hidden from the agent ("-": fine if a file does not exist).
InaccessiblePaths=-/etc/shadow -/etc/shadow- -/etc/gshadow -/etc/gshadow- -/etc/security/opasswd
InaccessiblePaths=-/etc/ssh/ssh_host_rsa_key -/etc/ssh/ssh_host_ecdsa_key -/etc/ssh/ssh_host_ed25519_key -/etc/ssh/ssh_host_dsa_key
ProtectKernelTunables=yes
ProtectKernelModules=yes
ProtectKernelLogs=yes
ProtectControlGroups=yes
ProtectClock=yes
ProtectHostname=yes
# Not PrivateTmp, PrivateDevices or ProtectProc: they would hide /tmp, /dev/shm and the other
# processes, which is what the agent watches. AF_UNIX is for systemd's unit list (D-Bus).
RestrictAddressFamilies=AF_INET AF_INET6 AF_UNIX
RestrictNamespaces=yes
RestrictRealtime=yes
RestrictSUIDSGID=yes
LockPersonality=yes
MemoryDenyWriteExecute=yes
RemoveIPC=yes
KeyringMode=private
DevicePolicy=closed
SystemCallArchitectures=native
SystemCallFilter=@system-service
# ...minus every way to attach to, read or write another process (what CAP_SYS_PTRACE would allow).
SystemCallFilter=~@debug ptrace process_vm_readv process_vm_writev kcmp pidfd_getfd
SystemCallErrorNumber=EPERM
UMask=0077
# Footprint: the agent gives way to your workload, never the other way round. Lowest CPU and I/O
# priority and weight; CPUQuota is the hard cap.
CPUQuota=5%
CPUWeight=1
IOWeight=1
MemoryHigh=48M
MemoryMax=64M
TasksMax=64
LimitNOFILE=1024
OOMScoreAdjust=500
Nice=19
IOSchedulingClass=idle
[Install]
WantedBy=multi-user.targetMeasured on a test machine with 614 processes: 0.05% of one core on average, never more than 2% in any second, 17 MB of memory.
Install
Create a server in your account to get its token, then on the server (for many servers, use an enrollment key instead: see below): /account
curl -fsSL https://approvalens.com/agent/install.sh | sudo bash -s -- --token <TOKEN>
See every step first, without changing anything:
curl -fsSL https://approvalens.com/agent/install.sh | sudo bash -s -- --token <TOKEN> --dry-run
Install a binary you built yourself:
curl -fsSLO https://approvalens.com/agent/install.sh sudo bash install.sh --token <TOKEN> --binary ./approvalens-agent-linux-amd64
Print exactly what it would send, without sending. This is the line the installer prints: it runs the agent as its own user with the same two read-only capabilities the service has, so it sees what the service sees:
sudo setpriv --reuid=approvalens-agent --regid=approvalens-agent --init-groups \ --inh-caps=+dac_read_search,+sys_ptrace --ambient-caps=+dac_read_search,+sys_ptrace \ /usr/local/bin/approvalens-agent check --config /etc/approvalens-agent.yaml
Installed with --unprivileged (no capabilities)? Then run it the plain way, as the service does in that mode; with --root, use sudo alone:
sudo -u approvalens-agent /usr/local/bin/approvalens-agent check --config /etc/approvalens-agent.yaml
Many servers: one enrollment key
Create an enrollment key in your account and use it in your provisioning. On each server the installer exchanges the key for that server's own token (the server appears in your account under its host name, or --name), and only that token is written to the configuration: the agent never sees the key, and revoking the key later does not disconnect the servers already enrolled. Running the line again on an installed server changes nothing. Your plan's server limit applies, and a key enrolls at most 30 servers an hour.
curl -fsSL https://approvalens.com/agent/install.sh | sudo bash -s -- --enroll alek_YOUR_KEY --name web-1Keep the key in your secrets store (Ansible Vault, your cloud's secret manager), not in a public repository. Anyone with it can add servers to your account until you revoke it.
Uninstall
Stops the service and removes the binary, the config, the state folder, the unit and the user:
curl -fsSL https://approvalens.com/agent/install.sh | sudo bash -s -- --uninstall
Build it yourself and compare
The builds are reproducible: the same Go version (the toolchain line in go.mod; Go downloads it by itself) produces byte-identical binaries on any machine. Build from the source tarball and compare the SHA-256 with ours:
curl -fsSLO https://approvalens.com/agent/0.3.0/approvalens-agent-0.3.0-src.tar.gz
curl -fsSLO https://approvalens.com/agent/0.3.0/SHA256SUMS
sha256sum --ignore-missing -c SHA256SUMS
tar xzf approvalens-agent-0.3.0-src.tar.gz && cd approvalens-agent-0.3.0
go test ./...
for arch in amd64 arm64; do
CGO_ENABLED=0 GOOS=linux GOARCH=$arch go build -trimpath -buildvcs=false \
-ldflags "-s -w -buildid= -X main.version=0.3.0" -o approvalens-agent-linux-$arch .
done
sha256sum approvalens-agent-linux-amd64 approvalens-agent-linux-arm64
grep linux ../SHA256SUMSThe sums must equal the lines in SHA256SUMS. If they do not, do not install it, and tell us.
Downloads · version 0.3.0
- Source code (tar.gz) approvalens-agent-0.3.0-src.tar.gz
- SHA256SUMS SHA256SUMS
- Linux x86-64 binary approvalens-agent-linux-amd64
- Linux ARM64 binary approvalens-agent-linux-arm64
- install.sh
16ff543e9ffbb6a36e6dec15cc90b5b8c012feb1d77a863a26fe8678353e396d approvalens-agent-linux-amd64 d8b62bed2010f1b7460e8dc651a7f34782ec4812f52cf27093c1921f535d9dc8 approvalens-agent-linux-arm64 a9b8252d0a35d292d03690da2749ae7c8c6cdfa74a5a99cba4c6ca9842d06d5e approvalens-agent-0.3.0-src.tar.gz
Source code
The main files of version 0.3.0, exactly as built. The tarball above also has the tests and their fixtures.
main.go435 lines
// Command approvalens-agent: the optional Approvalens server agent.
//
// It reads /proc, /sys, statfs and a few files (crontabs, authorized_keys, sshd and sudoers
// settings, the setuid bits of binaries, the size and time of files in configured web roots, the
// auth log's SSH lines), asks systemd for its unit list over D-Bus, aggregates a sample taken
// every 60 seconds and posts compact reports to approvalens.com. It never executes commands,
// never runs code it receives, opens no listening socket and changes nothing on the machine apart
// from its own state directory.
package main
import (
"context"
"encoding/json"
"errors"
"flag"
"fmt"
"log/slog"
"os"
"os/signal"
"runtime"
"runtime/debug"
"strings"
"sync/atomic"
"syscall"
"time"
)
var version = "dev" // set by build.sh (-ldflags "-X main.version=...")
// Exit codes: 78 (EX_CONFIG) for a configuration error, which the systemd unit does not restart on
// (RestartPreventExitStatus=78); 70 when the watchdog finds the sampling loop stuck.
const (
exitConfig = 78
exitStuck = 70
)
func usage() {
fmt.Fprintf(os.Stderr, `approvalens-agent %s
Usage:
approvalens-agent run [--config FILE] [--dry-run]
sample every 60 s and report (the systemd service); with --dry-run the messages are
printed instead of sent and no state is written
approvalens-agent check [--config FILE]
take one sample and print every message it would send; sends nothing
approvalens-agent --check [--config FILE] (also: check-config)
validate the configuration, show the effective settings and what is skipped here
approvalens-agent --version (also: version)
Config: %s. Documentation: https://approvalens.com/agent
`, version, DefaultConfigPath)
}
func main() {
setupLogging()
if len(os.Args) < 2 {
usage()
os.Exit(2)
}
cmd, args := os.Args[1], os.Args[2:]
switch cmd {
case "--check", "-check":
cmd = "check-config"
case "--dry-run", "-dry-run":
cmd, args = "run", append([]string{"--dry-run"}, args...)
case "--version", "-version", "-v":
cmd = "version"
case "--help", "-help", "-h", "help":
usage()
return
}
fs := flag.NewFlagSet(cmd, flag.ContinueOnError)
cfgPath := fs.String("config", envOr("APPROVALENS_AGENT_CONFIG", DefaultConfigPath), "config file")
dry := fs.Bool("dry-run", false, "print the messages instead of sending them")
switch cmd {
case "version":
fmt.Printf("approvalens-agent %s (%s, %s/%s)\n", version, runtime.Version(), runtime.GOOS, runtime.GOARCH)
return
case "run", "check", "check-config":
if err := fs.Parse(args); err != nil {
os.Exit(2)
}
if fs.NArg() > 0 {
fmt.Fprintf(os.Stderr, "unexpected argument %q\n", fs.Arg(0))
os.Exit(2)
}
default:
usage()
os.Exit(2)
}
switch cmd {
case "run":
os.Exit(run(*cfgPath, *dry))
case "check":
os.Exit(check(*cfgPath))
case "check-config":
os.Exit(checkConfig(*cfgPath))
}
}
func envOr(k, d string) string {
if v := os.Getenv(k); v != "" {
return v
}
return d
}
// Bounded footprint: two OS threads run Go code at most, a soft memory limit well under the
// unit's MemoryMax, and an eager GC.
func tuneRuntime() {
runtime.GOMAXPROCS(2)
debug.SetGCPercent(50)
debug.SetMemoryLimit(24 << 20)
}
var panicCount atomic.Int64
// safe runs fn and turns a panic into a logged error: one bad reading never stops the agent.
func safe(name string, fn func()) (ok bool) {
defer func() {
if r := recover(); r != nil {
panicCount.Add(1)
logLimited("panic:"+name, 15*time.Minute, slog.LevelError, "internal error recovered", "in", name, "panic", fmt.Sprint(r),
"stack", string(debug.Stack()))
ok = false
}
}()
fn()
return true
}
func safeBool(name string, fn func() bool) (res bool) {
if !safe(name, func() { res = fn() }) {
return false
}
return res
}
// watchdog exits the process when the sampling loop has not made progress for `limit` (a read
// of /proc or statfs() hanging on a sick device). systemd then restarts the agent.
func watchdog(beat *atomic.Int64, base time.Time, limit time.Duration) {
t := time.NewTicker(limit / 4)
defer t.Stop()
for range t.C {
if stuck := time.Since(base) - time.Duration(beat.Load()); stuck > limit {
buf := make([]byte, 32<<10)
n := runtime.Stack(buf, true)
slog.Error("the sampling loop is stuck; exiting so that systemd restarts the agent", "stuck_for", stuck.Round(time.Second).String(),
"goroutines", string(buf[:n]))
os.Exit(exitStuck)
}
}
}
func run(path string, dry bool) int {
tuneRuntime()
cfg, err := LoadConfig(path)
if err != nil {
slog.Error("cannot start: configuration error (fix it, then: systemctl restart approvalens-agent)", "err", err)
return exitConfig
}
for _, w := range cfg.Warnings {
slog.Warn(w)
}
c := NewCollector(cfg, version)
o := NewOutbox(c, !dry)
s := NewSender(cfg, version)
if dry {
s.Dry = os.Stdout
}
s.OnResync = o.Resync
s.LoadSpool()
if !dry {
c.LoadWebState()
}
c.LoadLoginState(!dry)
ctx, stop := signal.NotifyContext(context.Background(), syscall.SIGINT, syscall.SIGTERM)
defer stop()
signal.Ignore(syscall.SIGHUP, syscall.SIGPIPE)
sendDone := make(chan struct{})
go func() {
defer close(sendDone)
for ctx.Err() == nil {
safe("sender", func() { s.Run(ctx) })
}
}()
slog.Info("agent started", "version", version, "url", cfg.URL, "uid", os.Geteuid(), "sample", cfg.SampleInterval.String(),
"heartbeat", cfg.HeartbeatInterval.String(), "rollup", cfg.RollupInterval.String(), "spooled", s.Pending(), "dry_run", dry)
if d := c.Disabled(); len(d) > 0 {
slog.Info("checks disabled in the config", "checks", strings.Join(d, ","))
}
// Everything heavier than the sample runs on ONE background worker (gentle mode, gentle.go),
// one task at a time, so a sample is never late and the agent never makes a burst. A task that
// is already queued is not queued twice.
w := newWorker(c, !dry)
go w.loop(ctx)
queue := func(name string) {
j := job{name: name, notBefore: time.Now().Add(jitter(20*time.Second, 0.75))} // never on the sample's second
switch name {
case "posture":
j.in = c.PostureSnapshot()
case "owners":
j.inodes, j.procs = c.InodesSnapshot(), c.procRefs()
j.notBefore = time.Now().Add(2 * time.Second)
case "writers":
j.procs, j.boot = c.procRefs(), c.hostInfo.BootTime
case "disk":
j.mounts = c.MountsSnapshot()
}
w.add(j)
}
base := time.Now()
var beat atomic.Int64
go watchdog(&beat, base, max(4*cfg.SampleInterval, 10*time.Minute))
tick := time.NewTicker(cfg.SampleInterval)
defer tick.Stop()
start := time.Now()
safe("first sample", func() {
c.Sample(start)
o.RefreshInventory()
})
queue("units")
queue("auth")
queue("posture")
if c.ownersWant {
c.ownersWant = false
queue("owners")
}
safe("first messages", func() {
s.Enqueue(o.Heartbeat(start.UTC(), s.Pending())) // the account shows "connected" right away
if o.NeedFull() {
s.Enqueue(o.Full(start.UTC()))
}
})
nextHB := start.Add(cfg.HeartbeatInterval)
nextRollup := start.Add(time.Minute) // the first chart points soon after install, then every 15 min
nextInv := start.Add(cfg.HeartbeatInterval)
nextSetuid := start.Add(jitter(5*time.Minute, 0.5))
nextWeb := start
nextChecks := start.Add(cfg.HeartbeatInterval)
nextPosture := start.Add(jitter(time.Hour, 0.1))
nextDisk := start.Add(jitter(15*time.Minute, 0.5))
lastStats := start
step := func(now time.Time) {
c.Sample(now)
if c.ownersWant {
c.ownersWant = false
queue("owners")
}
if !now.Before(nextSetuid) {
nextSetuid = now.Add(jitter(cfg.SetuidInterval, 0.1))
queue("setuid")
}
if !now.Before(nextDisk) {
nextDisk = now.Add(jitter(6*time.Hour, 0.1))
queue("disk")
}
if len(cfg.WebRoots) > 0 && !now.Before(nextWeb) {
nextWeb = now.Add(cfg.WebRootInterval)
queue("web")
}
if !now.Before(nextChecks) {
nextChecks = now.Add(cfg.HeartbeatInterval)
queue("units")
queue("auth")
}
if !now.Before(nextPosture) {
nextPosture = now.Add(jitter(time.Hour, 0.1))
queue("posture")
}
if !now.Before(nextInv) { // crontabs and SSH keys: read every heartbeat interval
nextInv = now.Add(cfg.HeartbeatInterval)
o.RefreshInventory()
}
utc := now.UTC()
if o.NeedFull() {
o.RefreshInventory()
s.Enqueue(o.Full(utc))
}
if m := o.Event(utc, cfg.EventGap); m != nil {
s.Enqueue(*m)
}
if !now.Before(nextRollup) {
nextRollup = now.Add(cfg.RollupInterval)
s.Enqueue(o.Rollup(utc))
queue("writers") // read the write counters now: their difference goes into the next rollup
}
if !now.Before(nextHB) {
nextHB = now.Add(cfg.HeartbeatInterval)
s.Enqueue(o.Heartbeat(utc, s.Pending()))
}
if now.Sub(lastStats) >= 24*time.Hour {
msgs, bytes := s.Sent()
slog.Info("daily statistics", "messages", msgs, "bytes", bytes, "spooled", s.Pending(), "dropped", s.dropped.Load(),
"recovered_errors", panicCount.Load(), "worker_busy_ms", w.busy.Load(), "postponed", w.postponed.Load())
lastStats = now
}
}
for {
select {
case <-ctx.Done():
tick.Stop()
stop()
<-sendDone
s.FinalFlush(5 * time.Second)
msgs, bytes := s.Sent()
slog.Info("stopping", "messages_sent", msgs, "bytes_sent", bytes, "left_in_spool", s.Pending())
return 0
case now := <-tick.C:
beat.Store(int64(time.Since(base)))
safe("sample", func() { step(now) })
}
}
}
// check prints the messages the agent would send now. Nothing is sent and no state is written.
func check(path string) int {
cfg, err := LoadConfig(path)
if err != nil {
// Transparency first: show what would be collected even without a valid config.
fmt.Fprintf(os.Stderr, "config: %v (continuing with defaults; nothing is sent by `check` anyway)\n", err)
cfg = DefaultConfig()
cfg.path = path
}
lowerThreadPriority()
c := NewCollector(cfg, version)
c.LoadWebState()
c.LoadLoginState(false) // reads history (backfill), saves nothing
o := NewOutbox(c, false)
c.Sample(time.Now())
c.SampleWriters(c.procRefs(), c.hostInfo.BootTime, nil, time.Now())
c.FindPortOwners(c.InodesSnapshot(), c.procRefs(), nil)
time.Sleep(2 * time.Second) // a second sample turns counters into rates
c.Sample(time.Now())
c.SampleWriters(c.procRefs(), c.hostInfo.BootTime, nil, time.Now())
c.ScanSetuidNow(nil)
c.ScanWebNow(false, nil)
c.CheckUnits(time.Now())
c.ReadAuthLog() // positions at the end (the history goes into the login block); failures are counted from now on
c.ReadLogins()
c.CheckPosture(time.Now(), c.PostureSnapshot())
c.x.disk.MaxWork, c.x.disk.MaxWall = 5*time.Second, time.Minute // `check` shows a shortened disk scan
c.ScanDisks(c.MountsSnapshot(), nil, time.Now())
o.RefreshInventory()
now := time.Now().UTC()
msgs := []any{o.Heartbeat(now, 0), o.Full(now), o.Rollup(now)}
if m := o.Event(now, 0); m != nil {
msgs = append(msgs, *m)
}
total := 0
for _, m := range msgs {
b, _ := json.MarshalIndent(m, "", " ")
gz, _ := Encode(m)
total += len(gz)
fmt.Println(string(b))
fmt.Fprintf(os.Stderr, "-- %s: %d bytes gzip (%d bytes JSON)\n", msgKind(m), len(gz), len(b))
}
fmt.Fprintf(os.Stderr, "\nWould POST these to %s with token %s. Schedule: heartbeat every %s, rollup every %s, events at once (at most one per %s), the full inventory only after install, an upgrade or when the server asks for a resync.\n",
cfg.Endpoint(), maskToken(cfg.Token), cfg.HeartbeatInterval, cfg.RollupInterval, cfg.EventGap)
fmt.Fprintf(os.Stderr, "Running as uid %d, mode %s (capabilities: %s). Checks skipped here: %s. Disabled in the config: %s\n", os.Geteuid(),
c.ident.Mode, joinOr(c.ident.Caps, "none"), joinOr(c.Skipped(), "none"), joinOr(c.Disabled(), "none"))
fmt.Fprintln(os.Stderr, "Nothing was sent. The agent never executes commands and has no listening socket.")
return 0
}
// checkConfig validates the configuration and shows what the agent will do with it.
func checkConfig(path string) int {
cfg, err := LoadConfig(path)
if err != nil {
fmt.Fprintf(os.Stderr, "config error: %v\n", err)
return exitConfig
}
fmt.Printf("%s: OK\n", path)
fmt.Printf(" reports to %s (token %s)\n", cfg.Endpoint(), maskToken(cfg.Token))
fmt.Printf(" schedule sample %s, heartbeat %s, rollup %s, events at most one per %s\n", cfg.SampleInterval, cfg.HeartbeatInterval,
cfg.RollupInterval, cfg.EventGap)
fmt.Printf(" web roots %s\n", joinOr(cfg.WebRoots, "none"))
c := NewCollector(cfg, version)
fmt.Printf(" disabled %s\n", joinOr(c.Disabled(), "none"))
fmt.Printf(" auth log %s\n", orStr(c.auth.Path, "none found"))
fmt.Printf(" state dir %s (%s)\n", cfg.StateDir, dirState(cfg.StateDir))
fmt.Printf(" running as uid %d, mode %s (capabilities: %s)\n", os.Geteuid(), c.ident.Mode, joinOr(c.ident.Caps, "none"))
for _, w := range cfg.Warnings {
fmt.Printf("warning: %s\n", w)
}
return 0
}
func dirState(p string) string {
fi, err := os.Stat(p)
switch {
case errors.Is(err, os.ErrNotExist):
return "created by systemd when the service starts"
case err != nil:
return err.Error()
case !fi.IsDir():
return "not a directory"
}
f, err := os.CreateTemp(p, ".check-*")
if err != nil {
return "not writable by this user"
}
f.Close()
os.Remove(f.Name())
return "writable"
}
func maskToken(t string) string {
if len(t) > 9 {
return t[:9] + "…"
}
return t
}
func orStr(s, d string) string {
if s == "" {
return d
}
return s
}
func joinOr(xs []string, d string) string {
if len(xs) == 0 {
return d
}
return strings.Join(xs, ", ")
}
config.go367 lines
package main
import (
"bufio"
"bytes"
"errors"
"fmt"
"net"
"net/url"
"os"
"regexp"
"sort"
"strconv"
"strings"
"time"
)
const DefaultConfigPath = "/etc/approvalens-agent.yaml"
// Checks that can be switched off with `disable: [...]`.
var checkNames = map[string]string{
"cmdline": "command lines of the top processes (names, users and paths are still sent)",
"auth_log": "SSH login failures counted from /var/log/auth.log or /var/log/secure",
"systemd": "failed systemd units and the state of firewall / update / time services (read over D-Bus)",
"posture": "the daily security settings check (sshd, firewall, automatic updates, sudoers, UID 0, exposed databases)",
"crontabs": "crontab lines",
"ssh_keys": "SSH key fingerprints",
"setuid": "the setuid / setgid file scan",
// 0.3.0
"logins": "login history (wtmp, utmp, lastlog) and the logins and sudo use counted from the auth log",
"disk_scan": "the disk usage scan (directory and file sizes, a few times a day) and the per-process disk writes",
"unit_logs": "the last log lines of a unit that failed (from /var/log/syslog or /var/log/messages)",
"reachability": "nothing on the server: asks Approvalens not to test from outside whether the public ports answer",
}
// Config is /etc/approvalens-agent.yaml. Only a small YAML subset is read: `key: value`
// lines and lists written as `- item` lines under a key (or `[a, b]`).
type Config struct {
Token string
URL string
WebRoots []string
WebRootExclude []string
WebRootMaxFiles int
SampleInterval time.Duration
HeartbeatInterval time.Duration
RollupInterval time.Duration
EventGap time.Duration
SetuidInterval time.Duration
WebRootInterval time.Duration
StateDir string
AuthLog string // "" = auto (/var/log/auth.log, then /var/log/secure)
Disabled map[string]bool // check name -> off
Warnings []string // not fatal: shown by --check and logged at start
path string
}
func DefaultConfig() Config {
return Config{
URL: "https://approvalens.com",
WebRootMaxFiles: 20000,
SampleInterval: 60 * time.Second,
HeartbeatInterval: 5 * time.Minute,
RollupInterval: 15 * time.Minute,
EventGap: 60 * time.Second,
SetuidInterval: 6 * time.Hour,
WebRootInterval: 15 * time.Minute,
StateDir: "/var/lib/approvalens-agent",
Disabled: map[string]bool{},
}
}
// On reports whether a check is enabled.
func (c Config) On(check string) bool { return !c.Disabled[check] }
// LoadConfig reads, parses and validates the config file. Every error names the file and,
// for syntax errors, the line.
func LoadConfig(path string) (Config, error) {
c := DefaultConfig()
c.path = path
b, err := os.ReadFile(path)
if err != nil {
if errors.Is(err, os.ErrNotExist) {
return c, fmt.Errorf("%s does not exist (install the agent with the line from your Approvalens account)", path)
}
if errors.Is(err, os.ErrPermission) {
return c, fmt.Errorf("%s is not readable by uid %d (it must be owned by the agent's user, mode 0600)", path, os.Geteuid())
}
return c, err
}
if err := c.parse(b); err != nil {
return c, fmt.Errorf("%s: %w", path, err)
}
if env := os.Getenv("APPROVALENS_AGENT_TOKEN"); env != "" {
c.Token = env
}
if fi, err := os.Stat(path); err == nil && fi.Mode().Perm()&0o004 != 0 {
c.Warnings = append(c.Warnings, fmt.Sprintf("%s is readable by every user and holds the token: chmod 600 %s", path, path))
}
if err := c.Validate(); err != nil {
return c, fmt.Errorf("%s: %w", path, err)
}
return c, nil
}
var scalarKeys = map[string]bool{"token": true, "url": true, "web_root_max_files": true, "sample_interval": true, "heartbeat_interval": true,
"rollup_interval": true, "event_min_gap": true, "setuid_interval": true, "web_root_interval": true, "state_dir": true, "auth_log": true}
var listKeys = map[string]bool{"web_roots": true, "web_root_exclude": true, "disable": true}
func (c *Config) parse(b []byte) error {
sc := bufio.NewScanner(bytes.NewReader(b))
sc.Buffer(make([]byte, 64*1024), 1024*1024)
listKey := ""
seen := map[string]int{}
n := 0
for sc.Scan() {
n++
raw := sc.Text()
if n == 1 {
raw = strings.TrimPrefix(raw, "\xef\xbb\xbf") // a BOM from a Windows editor
}
line := strings.TrimSpace(stripComment(raw))
if line == "" {
continue
}
if strings.HasPrefix(line, "- ") || line == "-" {
if listKey == "" {
return fmt.Errorf("line %d: list item without a key above it", n)
}
if err := c.set(listKey, unquote(strings.TrimSpace(strings.TrimPrefix(line, "-")))); err != nil {
return fmt.Errorf("line %d: %w", n, err)
}
continue
}
k, v, ok := strings.Cut(line, ":")
if !ok {
return fmt.Errorf("line %d: expected `key: value`, got %q", n, truncate(line, 60))
}
k, v = strings.TrimSpace(k), strings.TrimSpace(v)
if err := knownKey(k); err != nil {
return fmt.Errorf("line %d: %w", n, err)
}
if prev, dup := seen[k]; dup && scalarKeys[k] { // 0.1.0 took the later value silently; keep that, but say so
c.Warnings = append(c.Warnings, fmt.Sprintf("line %d: %s is already set on line %d; the later value is used", n, k, prev))
}
seen[k] = n
listKey = ""
if v == "" {
if !listKeys[k] {
return fmt.Errorf("line %d: %s needs a value", n, k)
}
listKey = k
continue
}
if strings.HasPrefix(v, "[") && strings.HasSuffix(v, "]") {
if !listKeys[k] {
return fmt.Errorf("line %d: %s takes one value, not a list", n, k)
}
for _, item := range strings.Split(strings.Trim(v, "[]"), ",") {
if item = unquote(strings.TrimSpace(item)); item != "" {
if err := c.set(k, item); err != nil {
return fmt.Errorf("line %d: %w", n, err)
}
}
}
continue
}
if err := c.set(k, unquote(v)); err != nil {
return fmt.Errorf("line %d: %w", n, err)
}
}
return sc.Err()
}
func knownKey(k string) error {
if scalarKeys[k] || listKeys[k] {
return nil
}
var all []string
for x := range scalarKeys {
all = append(all, x)
}
for x := range listKeys {
all = append(all, x)
}
sort.Strings(all)
best, bestD := "", 4
for _, x := range all {
if d := editDistance(k, x); d < bestD {
best, bestD = x, d
}
}
if best != "" {
return fmt.Errorf("unknown key %q (did you mean %q?)", k, best)
}
return fmt.Errorf("unknown key %q (known keys: %s)", k, strings.Join(all, ", "))
}
func editDistance(a, b string) int {
prev := make([]int, len(b)+1)
cur := make([]int, len(b)+1)
for j := range prev {
prev[j] = j
}
for i := 1; i <= len(a); i++ {
cur[0] = i
for j := 1; j <= len(b); j++ {
cost := 1
if a[i-1] == b[j-1] {
cost = 0
}
cur[j] = min(prev[j]+1, cur[j-1]+1, prev[j-1]+cost)
}
prev, cur = cur, prev
}
return prev[len(b)]
}
func stripComment(s string) string {
inQ := byte(0)
for i := 0; i < len(s); i++ {
ch := s[i]
switch {
case inQ != 0 && ch == inQ:
inQ = 0
case inQ == 0 && (ch == '"' || ch == '\''):
inQ = ch
case inQ == 0 && ch == '#' && (i == 0 || s[i-1] == ' ' || s[i-1] == '\t'):
return s[:i]
}
}
return s
}
func unquote(s string) string {
if len(s) >= 2 && (s[0] == '"' && s[len(s)-1] == '"' || s[0] == '\'' && s[len(s)-1] == '\'') {
return s[1 : len(s)-1]
}
return s
}
func (c *Config) set(k, v string) error {
switch k {
case "token":
c.Token = v
case "url":
c.URL = strings.TrimRight(v, "/")
case "web_roots":
c.WebRoots = append(c.WebRoots, v)
case "web_root_exclude":
c.WebRootExclude = append(c.WebRootExclude, v)
case "disable":
v = strings.ToLower(v)
if _, ok := checkNames[v]; !ok {
names := make([]string, 0, len(checkNames))
for n := range checkNames {
names = append(names, n)
}
sort.Strings(names)
return fmt.Errorf("disable: unknown check %q (known: %s)", v, strings.Join(names, ", "))
}
c.Disabled[v] = true
case "auth_log":
if v != "auto" && !strings.HasPrefix(v, "/") {
return fmt.Errorf("auth_log: %q is not an absolute path (or `auto`)", v)
}
if v == "auto" {
v = ""
}
c.AuthLog = v
case "web_root_max_files":
n, err := strconv.Atoi(v)
if err != nil || n < 1 || n > 1_000_000 {
return fmt.Errorf("web_root_max_files: %q is not a number between 1 and 1000000", v)
}
c.WebRootMaxFiles = n
case "sample_interval", "heartbeat_interval", "rollup_interval", "event_min_gap", "setuid_interval", "web_root_interval":
d, err := parseDuration(v)
if err != nil {
return fmt.Errorf("%s: %q is not a duration (write 60s, 5m or 6h)", k, v)
}
switch k {
case "sample_interval":
c.SampleInterval = d
case "heartbeat_interval":
c.HeartbeatInterval = d
case "rollup_interval":
c.RollupInterval = d
case "event_min_gap":
c.EventGap = d
case "setuid_interval":
c.SetuidInterval = d
case "web_root_interval":
c.WebRootInterval = d
}
case "state_dir":
c.StateDir = v
default:
return knownKey(k)
}
return nil
}
// parseDuration accepts Go durations ("90s", "5m", "6h") or plain seconds.
func parseDuration(v string) (time.Duration, error) {
if n, err := strconv.Atoi(v); err == nil {
return time.Duration(n) * time.Second, nil
}
return time.ParseDuration(v)
}
// The token format the web app accepts (web/app/api/agent/v1/report/route.ts).
var tokenRe = regexp.MustCompile(`^alsa_[A-Za-z0-9_-]{20,90}$`)
func (c Config) Validate() error {
if c.Token == "" {
return errors.New("token is missing (copy the install line from your Approvalens account)")
}
if !tokenRe.MatchString(c.Token) {
return errors.New("token does not look like an Approvalens server token (alsa_ followed by 20-90 letters, digits, - or _)")
}
u, err := url.Parse(c.URL)
if err != nil || u.Host == "" {
return fmt.Errorf("url %q is not valid", c.URL)
}
if u.User != nil || u.RawQuery != "" || u.Fragment != "" {
return fmt.Errorf("url %q must not carry a user, a query or a fragment", c.URL)
}
// Plain HTTP only to this machine (local testing); anything else must be HTTPS.
if u.Scheme != "https" {
host := u.Hostname()
ip := net.ParseIP(host)
if u.Scheme != "http" || !(host == "localhost" || (ip != nil && ip.IsLoopback())) {
return fmt.Errorf("url must be https:// (got %q)", c.URL)
}
}
if c.SampleInterval < 5*time.Second || c.SampleInterval > 10*time.Minute {
return fmt.Errorf("sample_interval must be between 5s and 10m (got %s)", c.SampleInterval)
}
if c.HeartbeatInterval < 30*time.Second || c.HeartbeatInterval > 10*time.Minute || c.HeartbeatInterval < c.SampleInterval {
return fmt.Errorf("heartbeat_interval must be between 30s and 10m, and not shorter than sample_interval (got %s)", c.HeartbeatInterval)
}
if c.RollupInterval < c.HeartbeatInterval || c.RollupInterval > time.Hour {
return fmt.Errorf("rollup_interval must be between heartbeat_interval and 1h (got %s)", c.RollupInterval)
}
if c.EventGap < 10*time.Second || c.EventGap > 10*time.Minute {
return fmt.Errorf("event_min_gap must be between 10s and 10m (got %s)", c.EventGap)
}
if c.SetuidInterval < time.Minute {
return fmt.Errorf("setuid_interval must be at least 1m (got %s)", c.SetuidInterval)
}
if c.WebRootInterval < time.Minute {
return fmt.Errorf("web_root_interval must be at least 1m (got %s)", c.WebRootInterval)
}
for _, r := range c.WebRoots {
if !strings.HasPrefix(r, "/") {
return fmt.Errorf("web_roots: %q is not an absolute path", r)
}
}
if !strings.HasPrefix(c.StateDir, "/") {
return fmt.Errorf("state_dir: %q is not an absolute path", c.StateDir)
}
return nil
}
// Endpoint the reports go to.
func (c Config) Endpoint() string { return c.URL + "/api/agent/v1/report" }
worker.go140 lines
package main
// The background worker: one goroutine, pinned to one OS thread at nice 19 with idle I/O priority,
// running one task at a time (gentle.go). Tasks that can wait are postponed while the machine is
// busy (load above 0.8 per core or heavy I/O wait), for at most two hours; each task starts a few
// seconds after it was queued (jittered) so it never shares a second with the sample.
import (
"context"
"sort"
"sync"
"sync/atomic"
"time"
)
type job struct {
name string
notBefore time.Time
queuedAt time.Time
in PostureInput
inodes map[uint64]bool
procs []procRef
mounts []Mount
boot int64
}
// priority: lower runs first. Cheap and time-sensitive tasks before the long scans.
var jobPriority = map[string]int{"owners": 0, "units": 1, "auth": 2, "web": 3, "writers": 4, "posture": 5, "setuid": 6, "disk": 7}
// deferrable: may wait while the machine is busy.
var deferrable = map[string]bool{"posture": true, "setuid": true, "disk": true}
type worker struct {
c *Collector
persist bool
mu sync.Mutex
pending map[string]job
wake chan struct{}
busy atomic.Int64 // milliseconds of work since start
postponed atomic.Int64
sleep func(time.Duration)
}
func newWorker(c *Collector, persist bool) *worker {
return &worker{c: c, persist: persist, pending: map[string]job{}, wake: make(chan struct{}, 1), sleep: time.Sleep}
}
// add queues a task; one already waiting with the same name is replaced (newer snapshot).
func (w *worker) add(j job) {
j.queuedAt = time.Now()
w.mu.Lock()
if old, ok := w.pending[j.name]; ok {
j.queuedAt = old.queuedAt
}
w.pending[j.name] = j
w.mu.Unlock()
select {
case w.wake <- struct{}{}:
default:
}
}
// next picks the runnable task with the best priority, or says how long to wait.
func (w *worker) next(now time.Time) (job, bool, time.Duration) {
w.mu.Lock()
defer w.mu.Unlock()
busy := w.c.load.Busy()
names := make([]string, 0, len(w.pending))
for n := range w.pending {
names = append(names, n)
}
sort.Slice(names, func(i, j int) bool { return jobPriority[names[i]] < jobPriority[names[j]] })
wait := time.Minute
for _, n := range names {
j := w.pending[n]
if now.Before(j.notBefore) {
wait = min(wait, j.notBefore.Sub(now))
continue
}
if busy && deferrable[n] && now.Sub(j.queuedAt) < 2*time.Hour {
w.postponed.Add(1)
j.notBefore = now.Add(10 * time.Minute)
w.pending[n] = j
wait = min(wait, 10*time.Minute)
continue
}
delete(w.pending, n)
return j, true, 0
}
return job{}, false, wait
}
func (w *worker) loop(ctx context.Context) {
lowerThreadPriority()
g := NewGentle(w.c.load.Busy)
g.Sleep = w.sleep
t := time.NewTimer(time.Second)
defer t.Stop()
for ctx.Err() == nil {
j, ok, wait := w.next(time.Now())
if !ok {
t.Reset(max(wait, 100*time.Millisecond))
select {
case <-ctx.Done():
return
case <-w.wake:
case <-t.C:
}
continue
}
g.Begin()
safe("job "+j.name, func() { w.run(j, g) })
w.busy.Add(g.Work().Milliseconds())
}
}
func (w *worker) run(j job, g *Gentle) {
c := w.c
now := time.Now()
switch j.name {
case "setuid":
c.ScanSetuidNow(g)
case "web":
c.ScanWebNow(w.persist, g)
case "units":
c.CheckUnits(now)
case "auth":
c.ReadAuthLog()
c.ReadLogins()
case "posture":
c.CheckPosture(now, j.in)
case "owners":
c.FindPortOwners(j.inodes, j.procs, g)
case "writers":
c.SampleWriters(j.procs, j.boot, g, now)
case "disk":
c.ScanDisks(j.mounts, g, now)
}
}
gentle.go101 lines
package main
// Gentle mode: the agent must never be noticed. Everything heavier than the one-minute sample (the
// port owner search, the login history, the hourly settings check, the disk usage scan, the
// per-process write counters) runs on ONE background worker, one task at a time, on a thread with
// the lowest CPU priority (nice 19) and idle I/O priority. Inside a task the work is cut into
// slices of about 2 ms; after each slice the worker sleeps so that it uses about 1% of one core
// while it works, and four times less while the machine is busy (load above 0.8 per core or a
// lot of I/O wait). Tasks that can wait (the disk scan, the setuid scan, the settings check) are
// postponed while the machine is busy, up to two hours.
import (
"math/rand"
"sync/atomic"
"time"
)
type Gentle struct {
Slice time.Duration // work before a pause
Pause time.Duration // pause after a slice (4x while the host is busy)
Busy func() bool
Sleep func(time.Duration)
start time.Time
work time.Duration
n int
}
// NewGentle: 2 ms slices, 200 ms pauses: about 1% of one core while a task runs.
func NewGentle(busy func() bool) *Gentle {
return &Gentle{Slice: 2 * time.Millisecond, Pause: 200 * time.Millisecond, Busy: busy, Sleep: time.Sleep}
}
// Begin starts timing a task.
func (g *Gentle) Begin() {
if g == nil {
return
}
g.start, g.work, g.n = time.Now(), 0, 0
}
// Tick is called in the loops of a task; every few calls it checks the time spent since the last
// pause and pauses when the slice is used up.
func (g *Gentle) Tick() {
if g == nil {
return
}
g.n++
if g.n&3 != 0 { // the clock is looked at every 4 steps: a slice overshoots little
return
}
el := time.Since(g.start)
if el < g.Slice {
return
}
g.work += el
p := g.Pause
if g.Busy != nil && g.Busy() {
p *= 4
}
if g.Sleep != nil {
g.Sleep(p)
}
g.start = time.Now()
}
// Work is the time spent working (pauses excluded) since Begin.
func (g *Gentle) Work() time.Duration {
if g == nil {
return 0
}
return g.work + time.Since(g.start)
}
// HostLoad is what the sampling loop knows about how busy the machine is, for the worker.
type HostLoad struct {
perCore atomic.Uint64 // load1 / cores, x1000
iowait atomic.Uint64 // % of CPU time in I/O wait over the last sample, x10
}
func (h *HostLoad) Set(load1 float64, cores int, iowaitPct float64) {
if cores < 1 {
cores = 1
}
h.perCore.Store(uint64(max(0, load1/float64(cores)*1000)))
h.iowait.Store(uint64(max(0, iowaitPct*10)))
}
// Busy: load above 0.8 per core, or more than 20% of CPU time waiting for I/O.
func (h *HostLoad) Busy() bool {
return h.perCore.Load() > 800 || h.iowait.Load() > 200
}
// jitter returns d plus or minus up to frac of it (timers never line up on the same second).
func jitter(d time.Duration, frac float64) time.Duration {
if d <= 0 || frac <= 0 {
return d
}
j := time.Duration((rand.Float64()*2 - 1) * frac * float64(d))
return d + j
}
collector.go1120 lines
package main
import (
"bytes"
"crypto/sha256"
"encoding/hex"
"encoding/json"
"log/slog"
"os"
"path/filepath"
"runtime"
"sort"
"strconv"
"strings"
"sync"
"time"
)
const (
hogPercent = 90.0 // one process at this share of a core ...
hogDuration = 15 * time.Minute // ... for this long, without a break
fillWindow = 6 * 3600 // disk fill projection looks at the last 6 hours
clockStep = 30 * time.Second // a wall clock change this much larger than the elapsed time is a jump
)
// ---------------------------------------------------------------------------
// Data shared by the messages (messages.go). Field names are short: they go over the network.
// ---------------------------------------------------------------------------
type Minute struct {
T int64 `json:"t"` // start of the minute, unix seconds (UTC)
CPU float64 `json:"cpu"` // % of all cores, average
CPUMax float64 `json:"cpu_max"` // highest sample in the minute
Load1 float64 `json:"load1"`
Mem float64 `json:"mem"` // % used (MemAvailable based)
MemMax float64 `json:"mem_max"`
Swap float64 `json:"swap"` // % of swap used
SwapIO float64 `json:"swap_io"` // pages swapped in + out per second
RX float64 `json:"rx"` // bytes per second
TX float64 `json:"tx"`
Disk float64 `json:"disk"` // fullest mount, % used
TCPEst float64 `json:"tcp_est"` // established TCP connections, average
IO float64 `json:"io"` // busiest block device, % of time busy (highest sample)
OOM uint64 `json:"oom"` // OOM kills in the minute
N int `json:"n"` // samples in the minute
}
type DiskOut struct {
Mount string `json:"mount"`
FS string `json:"fs"`
Total uint64 `json:"total"`
Used uint64 `json:"used"`
Pct float64 `json:"pct"`
InodesPct float64 `json:"inodes_pct"`
RateBPH *float64 `json:"rate_bph,omitempty"` // bytes per hour over the last 6 h (least squares)
FullInH *float64 `json:"full_in_h,omitempty"` // hours until full at that rate
}
type PortOut struct {
Proto string `json:"proto"`
Addr string `json:"addr"`
Port int `json:"port"`
Local bool `json:"local"` // bound to loopback only
PID int `json:"pid,omitempty"`
Name string `json:"name,omitempty"`
User string `json:"user,omitempty"`
Exe string `json:"exe,omitempty"` // 0.3.0: the owner's executable
Unit string `json:"unit,omitempty"` // 0.3.0: its systemd unit
Container string `json:"container,omitempty"` // 0.3.0: Docker container name
CID string `json:"cid,omitempty"` // 0.3.0: container id, 12 hex digits
}
func (p PortOut) Key() string {
k := p.Proto + "/" + strconv.Itoa(p.Port)
if p.Local {
k += "/local"
}
return k
}
type MemOut struct {
Total uint64 `json:"total"`
Available uint64 `json:"available"`
UsedPct float64 `json:"used_pct"`
SwapTotal uint64 `json:"swap_total"`
SwapUsed uint64 `json:"swap_used"`
SwapPct float64 `json:"swap_pct"`
}
// TableOut is a kernel table and how full it is (open files, conntrack entries).
type TableOut struct {
Used uint64 `json:"used"`
Max uint64 `json:"max"`
Pct float64 `json:"pct"`
}
func newTable(used, max uint64) *TableOut {
t := &TableOut{Used: used, Max: max}
if max > 0 {
t.Pct = round2(clamp(100*float64(used)/float64(max), 0, 100))
}
return t
}
type Current struct {
At int64 `json:"at"`
CPU float64 `json:"cpu"`
Load [3]float64 `json:"load"`
Mem MemOut `json:"mem"`
Disks []DiskOut `json:"disks"`
RX float64 `json:"rx"`
TX float64 `json:"tx"`
Procs int `json:"procs"`
TopCPU []Proc `json:"top_cpu"`
TopMem []Proc `json:"top_mem"`
Ports []PortOut `json:"ports"`
TCP *TCPStates `json:"tcp,omitempty"`
FD *TableOut `json:"fd,omitempty"`
Conntrack *TableOut `json:"conntrack,omitempty"` // only when the conntrack module is loaded
IO []IOOut `json:"io,omitempty"`
}
type Signal struct {
Kind string `json:"kind"` // miner | tmp_exe | deleted_exe | cpu_hog
Key string `json:"key"` // stable across samples and restarts of the process
At int64 `json:"at"` // first seen in this report window
Detail map[string]any `json:"detail"`
}
type WebRootOut struct {
Roots []string `json:"roots"`
Files int `json:"files"`
LastScan int64 `json:"last_scan,omitempty"`
Baseline bool `json:"baseline,omitempty"` // first scan: nothing to compare with yet
Changes *WebChanges `json:"changes,omitempty"`
}
type AgentInfo struct {
Version string `json:"version"`
Root bool `json:"root"`
UID int `json:"uid"`
Skipped []string `json:"skipped"`
Disabled []string `json:"disabled"` // checks switched off in the config
SampleS int `json:"sample_s"`
ReportS int `json:"report_s"`
StartedAt int64 `json:"started_at"`
Panics int `json:"panics,omitempty"` // internal errors recovered since start
// 0.3.0
Mode string `json:"mode,omitempty"` // unprivileged | detailed | root
Caps []string `json:"caps,omitempty"` // effective capabilities that matter (dac_read_search, sys_ptrace)
CPUPct float64 `json:"cpu_pct"` // the agent's own CPU, % of one core, average over the window
CPUPeak float64 `json:"cpu_peak"` // highest one-sample-interval average in the window
RSS int64 `json:"rss"` // bytes, average of the samples in the window
RSSMax int64 `json:"rss_max"`
WindowS int64 `json:"window_s"`
}
type HostInfo struct {
Hostname string `json:"hostname"`
OS string `json:"os"`
Kernel string `json:"kernel"`
Arch string `json:"arch"`
Cores int `json:"cores"`
BootTime int64 `json:"boot_time"`
UptimeS int64 `json:"uptime_s"`
Virt string `json:"virt,omitempty"` // 0.3.0: kvm, vmware, docker, lxc, none… ("" unknown)
Vendor string `json:"vendor,omitempty"` // 0.3.0: DMI system vendor
Product string `json:"product,omitempty"` // 0.3.0: DMI product name
}
// HealthOut goes in every rollup: what is not a chart but tells whether the machine is well.
type HealthOut struct {
Units *UnitsOut `json:"units,omitempty"` // nil: systemd's unit list could not be read
OOMKills uint64 `json:"oom_kills"` // since the previous rollup
OOMTotal uint64 `json:"oom_total"` // since boot
SSH *SSHAuthOut `json:"ssh,omitempty"` // nil: no readable auth log
Time TimeOut `json:"time"`
}
type UnitsOut struct {
Failed []string `json:"failed"`
At int64 `json:"at"`
Detail []UnitDetail `json:"detail,omitempty"` // 0.3.0: what systemd says about each failed unit
}
// ---------------------------------------------------------------------------
// Collector
// ---------------------------------------------------------------------------
type minuteAgg struct {
n int
cpu, cpuMax, load, mem, memMax, swap, swapIO, rx, tx, disk float64
tcpEst, io float64
oom uint64
}
type unitsState struct {
failed []string
byName map[string]UnitInfo
at int64
}
type Collector struct {
cfg Config
version string
proc string // /proc
sys string // /sys
host string // / (root of the other files read)
users *UserCache
procs *ProcSampler
started time.Time
statfs func(string) (DiskUsage, error)
busPath string
units func(socket string, timeout time.Duration) ([]UnitInfo, error)
prevT time.Time
prevCPU CPUTimes
prevNet NetCounters
prevVM VMStat
havePrev bool
fill map[string]*FillTracker
minutes map[int64]*minuteAgg
sampleSig map[string]Signal // process signals seen in the latest sample
hog map[pkey]time.Time // pid:start -> first sample over hogPercent
rules *Rules
pending []Event // tripped since the last event message
io *IOTracker
prevTopMem map[pkey]Proc // pid:start -> the largest processes of the previous sample
procNames map[string]bool // process names of the latest sample (time and firewall daemons)
oomWindow uint64
oomTotal uint64
clockJumps int
portInodes map[uint64]bool
portOwners map[uint64]Owner // filled by the worker (FindOwners), read under mu
ownersWant bool // the socket set changed: the worker should look for owners
docker *DockerIndex
load HostLoad
// cached between samples: what changes rarely
mountRaw []byte
mounts []Mount
sysBlk map[string]bool
sysBlkAt time.Time
portLo int
portHi int
portRangeAt time.Time
hostAt time.Time
hostStatic HostInfo
netBuf []byte
lastProcs []Proc
namesGen int
namesAt int
ident Identity
// the agent's own footprint, per rollup window
self selfStats
x extra
current Current
hostInfo HostInfo
samples int
errors int
mu sync.Mutex // guards what the slow-checks goroutine fills in, below
setuid []string // nil until the first scan
setuidAt int64
web *WebRootOut
webPrev map[string]string
webPend *WebChanges
unitSt *unitsState // nil: never read
unitsErr string
auth AuthLog
authWin AuthCounts
authFrom time.Time
authErr string // why the auth log is not read ("" when it is)
posture *PostureOut
sshdNoRd bool
sudoNoRd bool
cronUnreadable []string
sshUnreadable []string
}
func NewCollector(cfg Config, version string) *Collector {
users := NewUserCache("/etc/passwd")
sendCmdline = cfg.On("cmdline")
c := &Collector{
cfg: cfg, version: version, proc: "/proc", sys: "/sys", host: "/", users: users,
procs: NewProcSampler("/proc", users), started: time.Now(), statfs: statDisk, busPath: SystemBusPath(), units: ListUnits,
fill: map[string]*FillTracker{}, minutes: map[int64]*minuteAgg{}, sampleSig: map[string]Signal{},
hog: map[pkey]time.Time{}, portInodes: map[uint64]bool{}, portOwners: map[uint64]Owner{}, rules: NewRules(),
io: NewIOTracker(), prevTopMem: map[pkey]Proc{}, procNames: map[string]bool{},
}
c.auth.Path = FindAuthLog(c.host, cfg.AuthLog)
c.authFrom = c.started
c.docker = NewDockerIndex(c.host)
c.ident = DetectIdentity(c.proc)
c.initExtra()
c.procs.Chunk, c.procs.Pause, c.procs.Sleep = 100, 250*time.Millisecond, time.Sleep
return c
}
// withRoot points the collector at another tree (tests use testdata/root).
func (c *Collector) withRoot(root string) *Collector {
c.proc, c.sys, c.host = filepath.Join(root, "proc"), filepath.Join(root, "sys"), root
c.users = NewUserCache(filepath.Join(root, "etc/passwd"))
c.procs = NewProcSampler(c.proc, c.users)
c.auth.Path = FindAuthLog(root, c.cfg.AuthLog)
c.docker = NewDockerIndex(root)
c.initExtra()
return c
}
func (c *Collector) read(name string) ([]byte, error) {
return readSmall(filepath.Join(c.proc, name), 1<<20)
}
// Sample takes one reading of everything that is sampled every interval.
func (c *Collector) Sample(now time.Time) {
c.samples++
ok := true
var cur Current
cur.At = now.Unix()
b, err := c.read("stat")
ps, perr := ParseProcStat(b)
if err != nil || perr != nil {
ok = false
}
b, _ = c.read("meminfo")
mem, merr := ParseMemInfo(b)
if merr != nil {
ok = false
}
b, _ = c.read("vmstat")
vm := ParseVMStat(b)
b, _ = c.read("net/dev")
netc := ParseNetDev(b)
b, _ = c.read("loadavg")
cur.Load, _ = ParseLoadAvg(b)
b, _ = c.read("uptime")
up, _ := ParseUptime(b)
// Intervals use the monotonic clock; a step of the wall clock (NTP correction, manual change,
// a VM resumed) is only counted and logged.
elapsed := now.Sub(c.prevT).Seconds()
if c.havePrev {
if d := now.Round(0).Sub(c.prevT.Round(0)) - now.Sub(c.prevT); d > clockStep || d < -clockStep {
c.clockJumps++
logLimited("clock-jump", time.Hour, slog.LevelWarn, "the wall clock jumped", "by", d.Round(time.Second).String())
}
}
var swapIO float64
var oomDelta uint64
if c.havePrev && elapsed > 0 {
cur.CPU = round1(CPUPercent(c.prevCPU, ps.CPU))
if dt := float64(ps.CPU.Total()) - float64(c.prevCPU.Total()); dt > 0 && ps.CPU.IOWait >= c.prevCPU.IOWait {
c.load.Set(cur.Load[0], ps.Cores, 100*float64(ps.CPU.IOWait-c.prevCPU.IOWait)/dt)
}
if netc.RX >= c.prevNet.RX {
cur.RX = float64(netc.RX-c.prevNet.RX) / elapsed
}
if netc.TX >= c.prevNet.TX {
cur.TX = float64(netc.TX-c.prevNet.TX) / elapsed
}
if vm.PswpIn >= c.prevVM.PswpIn && vm.PswpOut >= c.prevVM.PswpOut {
swapIO = float64(vm.PswpIn-c.prevVM.PswpIn+vm.PswpOut-c.prevVM.PswpOut) / elapsed
}
if vm.HasOOM && c.prevVM.HasOOM && vm.OOMKill >= c.prevVM.OOMKill {
oomDelta = vm.OOMKill - c.prevVM.OOMKill
}
}
cur.RX, cur.TX = float64(int64(cur.RX)), float64(int64(cur.TX))
if vm.HasOOM || !c.havePrev {
c.prevVM = vm
} else {
c.prevVM = VMStat{PswpIn: vm.PswpIn, PswpOut: vm.PswpOut}
}
c.prevCPU, c.prevNet, c.prevT, c.havePrev = ps.CPU, netc, now, true
c.oomTotal = vm.OOMKill
cur.Mem = MemOut{Total: mem.Total, Available: mem.Available, UsedPct: round1(mem.UsedPct()), SwapTotal: mem.SwapTotal,
SwapUsed: mem.SwapUsed(), SwapPct: round1(mem.SwapPct())}
// Kernel tables
if b, err := c.read("sys/fs/file-nr"); err == nil {
if used, max, err := ParseFileNR(b); err == nil {
cur.FD = newTable(used, max)
}
}
if b, err := c.read("sys/net/netfilter/nf_conntrack_count"); err == nil {
if n, err := ParseUintFile(b); err == nil {
if b, err := c.read("sys/net/netfilter/nf_conntrack_max"); err == nil {
if max, err := ParseUintFile(b); err == nil && max > 0 {
cur.Conntrack = newTable(n, max)
}
}
}
}
// Disks
b, _ = c.read("self/mountinfo")
if !bytes.Equal(b, c.mountRaw) { // parsed again only when the mounts changed
c.mountRaw, c.mounts = b, ParseMountInfo(b)
}
maxDisk := 0.0
seenMounts := make(map[string]bool, len(c.mounts))
mono := now.Sub(c.started).Seconds() // the fill trend is not fooled by a clock step
for _, m := range c.mounts {
d, err := c.statfs(m.Point)
if err != nil || d.Total == 0 {
continue
}
seenMounts[m.Point] = true
ft := c.fill[m.Point]
if ft == nil {
ft = NewFillTracker(fillWindow)
c.fill[m.Point] = ft
}
ft.Add(int64(mono), d.Used)
o := DiskOut{Mount: m.Point, FS: m.FSType, Total: d.Total, Used: d.Used, Pct: round1(d.UsedPct), InodesPct: round1(d.InodesPct)}
if r, ok := ft.Rate(); ok {
r = float64(int64(r))
o.RateBPH = &r
if h, ok := ft.FullInHours(d.Avail); ok {
h = round1(h)
o.FullInH = &h
}
}
if o.Pct > maxDisk {
maxDisk = o.Pct
}
cur.Disks = append(cur.Disks, o)
if len(cur.Disks) >= 40 {
break
}
}
for k := range c.fill {
if !seenMounts[k] {
delete(c.fill, k)
}
}
b, _ = c.read("diskstats")
if c.sysBlk == nil || now.Sub(c.sysBlkAt) > 10*time.Minute || now.Before(c.sysBlkAt) {
c.sysBlk, c.sysBlkAt = c.sysBlock(), now
}
cur.IO = c.io.Update(now, ParseDiskStats(b, wholeDisk(c.sysBlk)))
// Processes
c.procs.bootTime = ps.BootTime
procs := c.procs.Sample(now)
cur.Procs = len(procs)
cur.TopCPU = c.withUnits(TopBy(procs, 10, func(p Proc) float64 { return p.CPU }))
cur.TopMem = c.withUnits(TopBy(procs, 10, func(p Proc) float64 { return float64(p.RSS) }))
var victims []Proc
if oomDelta > 0 {
victims = c.vanished(procs)
}
c.rememberTopMem(procs)
c.processSignals(now, procs)
c.lastProcs = procs
c.namesGen++
c.self.sample(c.proc, now)
// Sockets: listening ports and TCP connections by state
var tcp TCPStates
cur.Ports = c.ports(procs, &tcp)
cur.TCP = &tcp
c.current = cur
c.hostInfo = c.readHost(now, ps, up)
// Local rules: what trips now goes out in the next event message (and only then).
c.mu.Lock()
var failed []string
unitsKnown := c.unitSt != nil && c.unitsErr == ""
if unitsKnown {
failed = c.unitSt.failed
}
c.mu.Unlock()
evs, _ := c.rules.Update(now, Reading{Mem: cur.Mem.UsedPct, SwapIO: swapIO, SwapPct: cur.Mem.SwapPct, Load1: cur.Load[0],
Cores: ps.Cores, Disks: cur.Disks, TopCPU: cur.TopCPU, TopMem: cur.TopMem, Signals: c.sampleSig,
OOMKills: oomDelta, OOMTotal: vm.OOMKill, OOMVictims: victims, FD: cur.FD, Conntrack: cur.Conntrack, IO: cur.IO,
FailedUnits: failed, UnitsKnown: unitsKnown})
c.pending = append(c.pending, evs...)
if len(c.pending) > 100 {
c.pending = c.pending[len(c.pending)-100:]
}
c.oomWindow += oomDelta
if !ok {
c.errors++
return
}
t := now.Unix() - now.Unix()%60
a := c.minutes[t]
if a == nil {
if len(c.minutes) >= 24*60 { // rollups stopped somehow: keep the memory bounded
c.minutes = map[int64]*minuteAgg{}
}
a = &minuteAgg{}
c.minutes[t] = a
}
worstIO := 0.0
for _, d := range cur.IO {
worstIO = maxF(worstIO, d.Util)
}
a.n++
a.cpu += cur.CPU
a.cpuMax = maxF(a.cpuMax, cur.CPU)
a.load += cur.Load[0]
a.mem += cur.Mem.UsedPct
a.memMax = maxF(a.memMax, cur.Mem.UsedPct)
a.swap += cur.Mem.SwapPct
a.swapIO += swapIO
a.rx += cur.RX
a.tx += cur.TX
a.disk = maxF(a.disk, maxDisk)
a.tcpEst += float64(tcp.Established)
a.io = maxF(a.io, worstIO)
a.oom += oomDelta
}
// sysBlock lists /sys/block (whole devices); empty when it cannot be read.
func (c *Collector) sysBlock() map[string]bool {
ents, err := os.ReadDir(filepath.Join(c.sys, "block"))
if err != nil {
return nil
}
m := make(map[string]bool, len(ents))
for _, e := range ents {
m[e.Name()] = true
}
return m
}
// pkey identifies a process across samples (a pid is reused; pid + start time is not).
type pkey struct {
pid int
start uint64
}
func procID(p Proc) pkey { return pkey{p.PID, p.startTime} }
func (c *Collector) rememberTopMem(procs []Proc) {
top := topRaw(procs, 10, func(p Proc) float64 { return float64(p.RSS) })
m := make(map[pkey]Proc, len(top))
for _, p := range top {
m[procID(p)] = p.Reported()
}
c.prevTopMem = m
}
// vanished: the largest processes of the previous sample that are gone now (after an OOM kill,
// probably its victims).
func (c *Collector) vanished(procs []Proc) []Proc {
live := make(map[pkey]bool, len(procs))
for _, p := range procs {
live[procID(p)] = true
}
var out []Proc
for id, p := range c.prevTopMem {
if !live[id] {
out = append(out, p)
}
}
sort.Slice(out, func(i, j int) bool { return out[i].RSS > out[j].RSS })
if len(out) > 3 {
out = out[:3]
}
return out
}
func (c *Collector) readHost(now time.Time, ps ProcStat, up float64) HostInfo {
if c.hostStatic.Kernel == "" || now.Sub(c.hostAt) > 10*time.Minute || now.Before(c.hostAt) {
c.hostStatic, c.hostAt = c.readHostStatic(), now
}
h := c.hostStatic
h.Arch, h.Cores, h.BootTime, h.UptimeS = runtime.GOARCH, ps.Cores, ps.BootTime, int64(up)
return h
}
func (c *Collector) readHostStatic() HostInfo {
var h HostInfo
v := DetectVirt(c.host, c.proc, c.sys)
h.Virt, h.Vendor, h.Product = v.Virt, v.Vendor, v.Product
if b, err := c.read("sys/kernel/hostname"); err == nil {
h.Hostname = truncate(strings.TrimSpace(string(b)), 100)
}
if b, err := c.read("sys/kernel/osrelease"); err == nil {
h.Kernel = truncate(strings.TrimSpace(string(b)), 100)
}
if b, err := readSmall(filepath.Join(c.host, "etc/os-release"), 64*1024); err == nil {
h.OS = ParseOSRelease(b)
} else if b, err := readSmall(filepath.Join(c.host, "usr/lib/os-release"), 64*1024); err == nil {
h.OS = ParseOSRelease(b)
}
return h
}
func (c *Collector) processSignals(now time.Time, procs []Proc) {
c.sampleSig = map[string]Signal{}
live := make(map[pkey]bool, len(procs))
for _, p := range procs {
if p.kernel || p.PID == os.Getpid() {
continue
}
// Crypto miner by name or arguments
why := p.minerWhy
if !p.sigDone {
why = sigs.MinerMatch(p.Name, p.argv0, p.cmdRaw)
}
if why != "" {
c.addSignal(now, "miner", "miner:"+strings.ToLower(p.Name)+":"+exeOrArgv0(p), p, map[string]any{"match": why})
}
// Executable in a temporary directory, or deleted after start
exe := strings.TrimSuffix(p.Exe, " (deleted)")
switch {
case p.exeKnown && InTmp(exe):
c.addSignal(now, "tmp_exe", "tmp_exe:"+exe, p, map[string]any{"source": "exe"})
case !p.exeKnown && strings.HasPrefix(p.argv0, "/") && InTmp(p.argv0):
c.addSignal(now, "tmp_exe", "tmp_exe:"+p.argv0, p, map[string]any{"source": "argv0"})
}
if p.exeKnown && SuspiciousDeleted(p.Exe) {
c.addSignal(now, "deleted_exe", "deleted_exe:"+exe, p, map[string]any{"source": "exe"})
}
// Sustained CPU by one process
id := procID(p)
live[id] = true
if p.CPU >= hogPercent {
first, seen := c.hog[id]
if !seen {
c.hog[id] = now
first = now
}
if now.Sub(first) >= hogDuration && !sigs.Busy[strings.ToLower(p.Name)] {
c.addSignal(now, "cpu_hog", "cpu_hog:"+strings.ToLower(p.Name), p,
map[string]any{"minutes": int(now.Sub(first).Minutes()), "since": first.Unix()})
}
} else {
delete(c.hog, id)
}
}
for id := range c.hog {
if !live[id] {
delete(c.hog, id)
}
}
}
func exeOrArgv0(p Proc) string {
if p.exeKnown {
return strings.TrimSuffix(p.Exe, " (deleted)")
}
return p.argv0
}
func (c *Collector) addSignal(now time.Time, kind, key string, p Proc, detail map[string]any) {
if len(c.sampleSig) >= 50 {
return
}
key = truncate(key, 300)
detail["process"] = p.Reported()
c.sampleSig[key] = Signal{Kind: kind, Key: key, At: now.Unix(), Detail: detail}
}
// localPortRange: UDP sockets bound to a port in this range are clients (resolvers,
// time sync), not services.
func (c *Collector) localPortRange() (int, int) {
if c.portLo > 0 && time.Since(c.portRangeAt) < 10*time.Minute {
return c.portLo, c.portHi
}
c.portLo, c.portHi = c.readPortRange()
c.portRangeAt = time.Now()
return c.portLo, c.portHi
}
func (c *Collector) readPortRange() (int, int) {
b, err := c.read("sys/net/ipv4/ip_local_port_range")
if err == nil {
f := strings.Fields(string(b))
if len(f) == 2 {
lo, e1 := strconv.Atoi(f[0])
hi, e2 := strconv.Atoi(f[1])
if e1 == nil && e2 == nil {
return lo, hi
}
}
}
return 32768, 60999
}
func (c *Collector) ports(procs []Proc, tcp *TCPStates) []PortOut {
var socks []Socket
if c.netBuf == nil {
c.netBuf = make([]byte, 64*1024)
}
for _, f := range [][2]string{{"net/tcp", "tcp"}, {"net/tcp6", "tcp"}, {"net/udp", "udp"}, {"net/udp6", "udp"}} {
fh, err := os.Open(filepath.Join(c.proc, f[0]))
if err != nil {
continue
}
var st *TCPStates
if f[1] == "tcp" {
st = tcp
}
socks = scanNetFast(fh, f[1], maxNetLines, st, c.netBuf, socks)
fh.Close()
}
lo, hi := c.localPortRange()
inodes := map[uint64]bool{}
var kept []Socket
for _, s := range socks {
if s.Proto == "udp" && s.Port >= lo && s.Port <= hi {
continue
}
kept = append(kept, s)
if s.Inode != 0 {
inodes[s.Inode] = true
}
}
if !sameSet(inodes, c.portInodes) {
c.portInodes = inodes
c.ownersWant = true // the worker looks for the owners (FindOwners)
}
c.mu.Lock()
owners := c.portOwners
c.mu.Unlock()
seen := map[string]bool{}
var out []PortOut
for _, s := range kept {
po := PortOut{Proto: s.Proto, Addr: s.Addr, Port: s.Port, Local: s.Local()}
if o, ok := owners[s.Inode]; ok && o.PID > 0 {
po.PID, po.Name, po.User, po.Exe, po.Unit, po.Container, po.CID = o.PID, o.Name, o.User, o.Exe, o.Unit, o.Container, o.CID
if po.User == "" {
po.User = c.users.Name(s.UID)
}
} else {
po.User = c.users.Name(s.UID)
}
k := po.Key() + "|" + po.Addr
if seen[k] {
continue
}
seen[k] = true
out = append(out, po)
}
sort.Slice(out, func(i, j int) bool {
if out[i].Proto != out[j].Proto {
return out[i].Proto < out[j].Proto
}
if out[i].Port != out[j].Port {
return out[i].Port < out[j].Port
}
return out[i].Addr < out[j].Addr
})
if len(out) > 200 {
out = out[:200]
}
return out
}
func sameSet(a, b map[uint64]bool) bool {
if len(a) != len(b) {
return false
}
for k := range a {
if !b[k] {
return false
}
}
return true
}
// ---------------------------------------------------------------------------
// Slow work, on its own goroutine: setuid scan (6 h), web roots (15 min), systemd units and the
// auth log (every heartbeat interval), the posture check (hourly).
// ---------------------------------------------------------------------------
func (c *Collector) ScanSetuidNow(g *Gentle) {
if !c.cfg.On("setuid") {
return
}
items, _ := ScanSetuid(setuidRoots, 200000, g)
if len(items) > 500 {
items = items[:500]
}
if items == nil {
items = []string{}
}
c.mu.Lock()
c.setuid, c.setuidAt = items, time.Now().Unix()
c.mu.Unlock()
}
// CheckUnits reads systemd's unit list over D-Bus.
func (c *Collector) CheckUnits(now time.Time) {
if !c.cfg.On("systemd") {
return
}
list, err := c.units(c.busPath, 3*time.Second)
c.mu.Lock()
defer c.mu.Unlock()
if err != nil {
if c.unitsErr == "" {
slog.Info("systemd unit list not available; failed units are not checked", "err", err)
}
c.unitsErr = truncate(err.Error(), 200)
return
}
st := &unitsState{byName: make(map[string]UnitInfo, len(list)), at: now.Unix(), failed: []string{}}
for _, u := range list {
st.byName[u.Name] = u
if u.ActiveState == "failed" && len(st.failed) < maxUnitRules {
st.failed = append(st.failed, truncate(u.Name, 120))
}
}
sort.Strings(st.failed)
c.unitSt, c.unitsErr = st, ""
failed := st.failed
c.mu.Unlock()
c.unitDetails(now, list, failed)
c.mu.Lock()
}
// ReadAuthLog adds what was appended to the auth log since the previous look.
func (c *Collector) ReadAuthLog() {
if !c.cfg.On("auth_log") {
return
}
c.mu.Lock()
defer c.mu.Unlock()
if c.auth.Path == "" {
c.auth.Path = FindAuthLog(c.host, c.cfg.AuthLog)
if c.auth.Path == "" {
c.authErr = "missing"
return
}
}
got, err := c.auth.Read()
switch {
case err == nil:
c.authErr = ""
c.authWin.add(got)
c.saveLoginState()
case os.IsPermission(err):
c.authErr = "unreadable"
case os.IsNotExist(err):
c.authErr = "missing"
c.auth = AuthLog{} // look for it again next time
default:
c.authErr = "error"
logLimited("auth-log", time.Hour, slog.LevelWarn, "cannot read the auth log", "path", c.auth.Path, "err", err)
}
}
// PostureInput is what the posture check needs from the sampling loop, copied there so that the
// slow goroutine never reads the sampler's state.
type PostureInput struct {
Ports []PortOut
ProcNames map[string]bool
Kernel string
BootTime int64
Procs []procRef
}
// PostureSnapshot is taken on the sampling goroutine (each sample replaces these, never mutates them).
func (c *Collector) PostureSnapshot() PostureInput {
return PostureInput{Ports: c.current.Ports, ProcNames: c.procNameSet(), Kernel: c.hostInfo.Kernel, BootTime: c.hostInfo.BootTime,
Procs: c.procRefs()}
}
// CheckPosture runs the security settings check.
func (c *Collector) CheckPosture(now time.Time, in PostureInput) {
if !c.cfg.On("posture") {
return
}
c.mu.Lock()
var units map[string]UnitInfo
if c.unitSt != nil && c.unitsErr == "" {
units = c.unitSt.byName
}
c.mu.Unlock()
p := &PostureOut{At: now.Unix()}
p.SSH = ReadSSHD(c.host)
if p.SSH != nil {
p.SSH.Running = in.ProcNames["sshd"]
}
p.Firewall = DetectFirewall(c.host, units, in.ProcNames)
p.AutoUpd = DetectAutoUpdates(c.host, units)
p.UID0 = ExtraUID0(c.users.All())
p.Sudo = ReadSudoers(c.host)
p.PublicDB = PublicDBPorts(in.Ports)
p.Reboot = CheckReboot(c.host, in.Kernel, in.BootTime)
p.Updates = ReadUpdates(c.host)
// 0.3.0
users := c.users.All()
var keys map[string]KeyCount
keysOK := false
if c.cfg.On("ssh_keys") {
keys, keysOK = CountSSHKeys(c.host, users)
if p.SSH != nil {
p.SSH.Keys = keyList(keys)
}
}
p.Fail2ban = ReadFail2ban(c.host, in.ProcNames)
c.mu.Lock()
last := c.x.last
c.mu.Unlock()
p.Users = BuildUsers(c.host, users, p.Sudo, last, keys, keysOK)
p.FW = ReadFirewallDetail(c.host, units, in.ProcNames, p.Firewall)
p.Pkg = ReadPkg(c.host)
c.mu.Lock()
c.posture = p
c.sshdNoRd = p.SSH != nil && !p.SSH.Readable
c.sudoNoRd = !p.Sudo.Readable
c.mu.Unlock()
}
// Posture returns the latest security settings check (nil before the first one).
func (c *Collector) Posture() *PostureOut {
c.mu.Lock()
defer c.mu.Unlock()
return c.posture
}
// Health closes the health window: units, OOM kills and SSH failures since the previous rollup.
func (c *Collector) Health(now time.Time, reset bool) HealthOut {
c.mu.Lock()
defer c.mu.Unlock()
h := HealthOut{OOMKills: c.oomWindow, OOMTotal: c.oomTotal}
if c.unitSt != nil && c.unitsErr == "" {
h.Units = &UnitsOut{Failed: c.unitSt.failed, At: c.unitSt.at, Detail: c.unitDetailList(c.unitSt.failed, reset)}
}
if c.cfg.On("auth_log") && c.authErr == "" && c.auth.started {
h.SSH = c.authWin.Out(filepath.Base(c.auth.Path), int64(now.Sub(c.authFrom).Seconds()))
}
var units map[string]UnitInfo
if c.unitSt != nil && c.unitsErr == "" {
units = c.unitSt.byName
}
h.Time = TimeSync(c.host, units, c.procNameSet())
h.Time.Jumps = c.clockJumps
if reset {
c.oomWindow = 0
c.authWin = AuthCounts{}
c.authFrom = now
}
return h
}
func (c *Collector) webStatePath() string { return filepath.Join(c.cfg.StateDir, "webroot.json") }
// LoadWebState reads the previous web root scan (kept across restarts).
func (c *Collector) LoadWebState() {
b, err := readSmall(c.webStatePath(), 32<<20)
if err != nil {
return
}
var m map[string]string
if json.Unmarshal(b, &m) == nil {
c.mu.Lock()
c.webPrev = m
c.mu.Unlock()
}
}
// ScanWebNow compares the web roots with the previous scan. persist=false (the `check`
// command, --dry-run) leaves the saved state alone.
func (c *Collector) ScanWebNow(persist bool, g *Gentle) {
if len(c.cfg.WebRoots) == 0 {
return
}
files, truncated := ScanWebRoots(c.cfg.WebRoots, c.cfg.WebRootExclude, c.cfg.WebRootMaxFiles, g)
c.mu.Lock()
defer c.mu.Unlock()
out := &WebRootOut{Roots: c.cfg.WebRoots, Files: len(files), LastScan: time.Now().Unix()}
if c.webPrev == nil {
out.Baseline = true
} else {
ch := DiffWebRoots(c.webPrev, files, 50)
ch.Truncated = truncated
if !ch.Empty() {
c.webPend = mergeWeb(c.webPend, &ch)
}
}
c.web = out
c.webPrev = files
if persist {
if b, err := json.Marshal(files); err == nil {
if err := writeFileAtomic(c.webStatePath(), b, 0o600); err != nil {
logLimited("web-state", time.Hour, slog.LevelWarn, "cannot save the web root state", "err", err)
}
}
}
}
func mergeWeb(a, b *WebChanges) *WebChanges {
if a == nil {
return b
}
m := *a
m.Added = capList(dedupSorted(sortedCopy(append(m.Added, b.Added...))), 50)
m.Changed = capList(dedupSorted(sortedCopy(append(m.Changed, b.Changed...))), 50)
m.Removed = capList(dedupSorted(sortedCopy(append(m.Removed, b.Removed...))), 50)
m.NAdded += b.NAdded
m.NChanged += b.NChanged
m.NRemoved += b.NRemoved
m.Files = b.Files
m.Truncated = m.Truncated || b.Truncated
return &m
}
func capList(xs []string, n int) []string {
if len(xs) > n {
return xs[:n]
}
return xs
}
// ---------------------------------------------------------------------------
// Building the report
// ---------------------------------------------------------------------------
// Skipped names the checks this agent cannot do here, with its permissions or on this system.
func (c *Collector) Skipped() []string {
var s []string
if !c.ident.SeesOthers() {
s = append(s, "exe_paths_other_users", "port_owners_other_users")
}
if len(c.cronUnreadable) > 0 {
s = append(s, "crontabs_unreadable")
}
if len(c.sshUnreadable) > 0 {
s = append(s, "ssh_keys_unreadable")
}
if len(c.cfg.WebRoots) == 0 {
s = append(s, "web_roots_not_configured")
}
c.mu.Lock()
defer c.mu.Unlock()
if c.cfg.On("systemd") && c.unitsErr != "" {
s = append(s, "systemd_units_unavailable")
}
if c.cfg.On("auth_log") {
switch c.authErr {
case "unreadable":
s = append(s, "auth_log_unreadable")
case "missing":
s = append(s, "auth_log_missing")
}
}
if c.cfg.On("posture") && c.sshdNoRd {
s = append(s, "sshd_config_unreadable")
}
if c.cfg.On("posture") && c.sudoNoRd {
s = append(s, "sudoers_unreadable")
}
if c.cfg.On("logins") {
switch c.x.wtmpErr {
case "unreadable":
s = append(s, "wtmp_unreadable")
case "missing":
s = append(s, "wtmp_missing")
}
}
if c.cfg.On("unit_logs") && c.x.syslogErr == "unreadable" {
s = append(s, "syslog_unreadable")
}
return s
}
// hashJSON: a short hash of a value's JSON (to send something only when it changed).
func hashJSON(v any) string {
b, _ := json.Marshal(v)
sum := sha256.Sum256(b)
return hex.EncodeToString(sum[:8])
}
func (c *Collector) Disabled() []string {
out := []string{}
for k := range c.cfg.Disabled {
out = append(out, k)
}
sort.Strings(out)
return out
}
func maxF(a, b float64) float64 {
if a > b {
return a
}
return b
}
func round2(v float64) float64 { return float64(int64(v*100+0.5)) / 100 }
func writeFileAtomic(path string, b []byte, mode os.FileMode) error {
if err := os.MkdirAll(filepath.Dir(path), 0o700); err != nil {
return err
}
tmp := path + ".tmp"
f, err := os.OpenFile(tmp, os.O_WRONLY|os.O_CREATE|os.O_TRUNC, mode)
if err != nil {
return err
}
if _, err := f.Write(b); err != nil {
f.Close()
os.Remove(tmp)
return err
}
if err := f.Close(); err != nil {
os.Remove(tmp)
return err
}
return os.Rename(tmp, path)
}
collector_extra.go465 lines
package main
// The 0.3.0 parts of the collector: login history, failed-unit details, port owners, disk usage
// scans and disk writers. All of it runs on the background worker (main.go), never on the sampling
// loop; what it finds waits under c.mu until the next rollup takes it.
import (
"encoding/json"
"log/slog"
"os"
"path/filepath"
"sort"
"time"
)
type loginState struct {
Version string `json:"version"`
Wtmp wtmpState `json:"wtmp"`
Auth *AuthPos `json:"auth,omitempty"`
}
// LoginsOut is the rollup's login block (see the wire format in README.md).
type LoginsOut struct {
Backfill bool `json:"backfill,omitempty"`
Wtmp string `json:"wtmp,omitempty"` // ok | unreadable | missing
Auth string `json:"auth,omitempty"` // ok | unreadable | missing
Recent []LoginRec `json:"recent,omitempty"`
Boots []BootRec `json:"boots,omitempty"`
Sessions []SessionRec `json:"sessions"`
Last map[string]int64 `json:"last,omitempty"`
Accepted []AcceptedRec `json:"accepted,omitempty"`
Failed []FailedIP `json:"failed,omitempty"`
Sudo []SudoRec `json:"sudo,omitempty"`
}
// FailedIP: one source address's failed SSH attempts in the rollup window.
type FailedIP struct {
IP string `json:"ip"`
N int `json:"n"`
Invalid int `json:"invalid"`
LastAt int64 `json:"last_at,omitempty"`
Users []string `json:"users,omitempty"` // existing accounts only
}
type extra struct {
persist bool
// logins (guarded by c.mu)
lg loginState
lgLoaded bool
wtmpErr string
recent []LoginRec
boots []BootRec
sessions []SessionRec
sessHash string
sessSent string
last map[string]int64
lastSent string
backfill bool
loginsSent bool // since this start
lastlogMt time.Time
lastlog map[string]int64
authDirty bool
// failed units (c.mu)
unitProps func(string, time.Duration, []UnitInfo) (map[string]UnitDetail, error)
unitDet map[string]UnitDetail
unitDetAt time.Time
unitKey string
logged map[string]int64 // unit -> the failure (since) whose log lines were already sent
syslogErr string
// disk (c.mu)
disk *DiskScanner
diskOut *DiskScanOut
writers *WriteTracker
writersOut []WriterOut
// rollup ports (sampling goroutine)
portsHash string
}
func (c *Collector) initExtra() {
c.x.unitProps = UnitDetails
c.x.unitDet = map[string]UnitDetail{}
c.x.logged = map[string]int64{}
c.x.disk = NewDiskScanner()
c.x.writers = NewWriteTracker()
c.x.last = map[string]int64{}
}
func (c *Collector) loginStatePath() string { return filepath.Join(c.cfg.StateDir, "logins.json") }
// LoadLoginState restores where the wtmp and auth log readers stopped (persist: the service; the
// `check` command and --dry-run neither read nor write it and backfill every time).
func (c *Collector) LoadLoginState(persist bool) {
c.mu.Lock()
defer c.mu.Unlock()
c.x.persist = persist
c.x.lgLoaded = true
if persist {
if b, err := readSmall(c.loginStatePath(), 4<<20); err == nil {
var st loginState
if json.Unmarshal(b, &st) == nil {
c.x.lg = st
}
}
}
c.auth.Resume(c.x.lg.Auth)
c.auth.Backfill = !c.auth.started
c.x.backfill = c.x.lg.Wtmp.Ino == 0
}
func (c *Collector) saveLoginState() {
if !c.x.persist {
return
}
c.x.lg.Version = c.version
c.x.lg.Auth = c.auth.Pos()
if b, err := json.Marshal(c.x.lg); err == nil {
if err := writeFileAtomic(c.loginStatePath(), b, 0o600); err != nil {
logLimited("login-state", time.Hour, slog.LevelWarn, "cannot save the login reader state", "err", err)
}
}
}
// ReadLogins reads wtmp (new records), utmp (who is logged in) and lastlog (when it changed).
func (c *Collector) ReadLogins() {
if !c.cfg.On("logins") {
return
}
c.mu.Lock()
if !c.x.lgLoaded {
c.mu.Unlock()
c.LoadLoginState(false)
c.mu.Lock()
}
st := c.x.lg.Wtmp // a copy: the read happens without the lock
if st.Open != nil {
cp := make(map[string]LoginRec, len(st.Open))
for k, v := range st.Open {
cp[k] = v
}
st.Open = cp
}
if st.Last != nil {
cp := make(map[string]int64, len(st.Last))
for k, v := range st.Last {
cp[k] = v
}
st.Last = cp
}
backfill := c.x.backfill
c.mu.Unlock()
ev, err := ReadWtmp(filepath.Join(c.host, "var/log/wtmp"), &st, backfill)
wtmpErr := ""
switch {
case err == nil:
case os.IsNotExist(err):
wtmpErr = "missing"
case os.IsPermission(err):
wtmpErr = "unreadable"
default:
wtmpErr = "error"
logLimited("wtmp", time.Hour, slog.LevelWarn, "cannot read wtmp", "err", err)
}
sessions, serr := ReadUtmpSessions(c.host, c.proc)
if serr != nil {
sessions = nil
}
var lastlog map[string]int64
llPath := filepath.Join(c.host, "var/log/lastlog")
fi, llErr := os.Stat(llPath)
c.mu.Lock()
defer c.mu.Unlock()
if llErr == nil && !fi.ModTime().Equal(c.x.lastlogMt) {
rs := 292
if st.Size == 400 {
rs = 296
}
lastlog = ReadLastlog(llPath, c.users.All(), rs)
c.x.lastlogMt, c.x.lastlog = fi.ModTime(), lastlog
}
c.x.wtmpErr = wtmpErr
if err == nil {
c.x.lg.Wtmp = st
c.x.backfill = false
c.x.recent = mergeLogins(c.x.recent, ev.Logins, maxLogins)
c.x.boots = mergeBoots(c.x.boots, ev.Boots, maxBoots)
}
if serr == nil {
c.x.sessions = sessions
}
last := map[string]int64{}
for u, t := range c.x.lastlog {
last[u] = t
}
for u, t := range c.x.lg.Wtmp.Last {
if t > last[u] {
last[u] = t
}
}
c.x.last = last
c.saveLoginState()
}
func mergeLogins(have, add []LoginRec, n int) []LoginRec {
if len(add) == 0 {
return have
}
key := func(l LoginRec) string { return itoa64(l.At) + "|" + l.User + "|" + l.TTY }
idx := map[string]int{}
for i, l := range have {
idx[key(l)] = i
}
for _, l := range add {
if i, ok := idx[key(l)]; ok {
have[i] = l
continue
}
idx[key(l)] = len(have)
have = append(have, l)
}
sort.SliceStable(have, func(i, j int) bool { return have[i].At < have[j].At })
return lastN(have, n)
}
func mergeBoots(have, add []BootRec, n int) []BootRec {
for _, b := range add {
replaced := false
for i := range have {
if have[i].At == b.At {
have[i] = b
replaced = true
}
}
if !replaced {
have = append(have, b)
}
}
sort.SliceStable(have, func(i, j int) bool { return have[i].At < have[j].At })
return lastN(have, n)
}
// takeLogins builds the rollup's login block (nil when there is nothing new). Called with c.mu
// held, before Health resets the auth window.
func (c *Collector) takeLogins() *LoginsOut {
if !c.cfg.On("logins") {
return nil
}
o := &LoginsOut{Backfill: !c.x.loginsSent, Recent: c.x.recent, Boots: c.x.boots, Sessions: c.x.sessions}
o.Wtmp = orStr(c.x.wtmpErr, "ok")
if c.cfg.On("auth_log") {
o.Auth = orStr(c.authErr, "ok")
if !c.auth.started && c.authErr == "" {
o.Auth = ""
}
}
if o.Sessions == nil {
o.Sessions = []SessionRec{}
}
sh := hashJSON(o.Sessions)
lh := hashJSON(c.x.last)
if lh != c.x.lastSent {
o.Last = c.x.last
}
w := c.authWin
o.Accepted = lastN(w.Logins, maxLogins)
o.Sudo = lastN(w.Sudo, maxLogins)
valid := map[string]bool{}
for _, u := range c.users.All() {
valid[u.Name] = true
}
for ip, n := range w.ByIP {
f := FailedIP{IP: ip, N: n}
if m := w.Fail[ip]; m != nil {
f.Invalid, f.LastAt = m.Invalid, m.Last
for u := range m.Users {
if valid[u] && len(f.Users) < 10 {
f.Users = append(f.Users, u)
}
}
sort.Strings(f.Users)
}
o.Failed = append(o.Failed, f)
}
sort.Slice(o.Failed, func(i, j int) bool {
if o.Failed[i].N != o.Failed[j].N {
return o.Failed[i].N > o.Failed[j].N
}
return o.Failed[i].IP < o.Failed[j].IP
})
if len(o.Failed) > 20 {
o.Failed = o.Failed[:20]
}
if c.x.loginsSent && len(o.Recent) == 0 && len(o.Boots) == 0 && o.Last == nil && len(o.Accepted) == 0 && len(o.Sudo) == 0 &&
len(o.Failed) == 0 && sh == c.x.sessSent {
return nil
}
c.x.loginsSent = true
c.x.sessSent, c.x.lastSent = sh, lh
c.x.recent, c.x.boots = nil, nil
return o
}
// unitDetails refreshes what systemd says about the failed units (when the failed set changes, and
// hourly) and, for a unit that newly failed, its last log lines. Called on the worker.
func (c *Collector) unitDetails(now time.Time, list []UnitInfo, failed []string) {
key := ""
byName := map[string]UnitInfo{}
for _, u := range list {
byName[u.Name] = u
}
var want []UnitInfo
for _, f := range failed {
key += f + ","
if u, ok := byName[f]; ok {
want = append(want, u)
}
}
c.mu.Lock()
fresh := key != c.x.unitKey || now.Sub(c.x.unitDetAt) > time.Hour
c.mu.Unlock()
if !fresh {
return
}
det := map[string]UnitDetail{}
if len(want) > 0 && c.x.unitProps != nil {
if d, err := c.x.unitProps(c.busPath, 3*time.Second, want); err == nil {
det = d
}
}
for _, u := range want { // a unit systemd would not describe still gets its name
if _, ok := det[u.Name]; !ok {
det[u.Name] = UnitDetail{Unit: truncate(u.Name, 120), Description: truncate(u.Description, 200)}
}
}
// log lines, once per failure
c.mu.Lock()
var need []string
for name, d := range det {
if since, ok := c.x.logged[name]; !ok || since != d.Since {
need = append(need, name)
}
}
c.mu.Unlock()
syslogErr := ""
if len(need) > 0 && c.cfg.On("unit_logs") {
sort.Strings(need)
path := SyslogPath(c.host)
tails, err := UnitLogTails(path, need)
switch {
case path == "":
syslogErr = "missing"
case err != nil && os.IsPermission(err):
syslogErr = "unreadable"
}
for _, name := range need {
d := det[name]
d.Log = tails[name]
det[name] = d
}
}
c.mu.Lock()
for name := range c.x.logged {
if _, ok := det[name]; !ok {
delete(c.x.logged, name) // recovered: a new failure gets its lines again
}
}
for _, name := range need {
c.x.logged[name] = det[name].Since
}
// keep log lines that were not sent yet
for name, old := range c.x.unitDet {
if d, ok := det[name]; ok && len(d.Log) == 0 && len(old.Log) > 0 && old.Since == d.Since {
d.Log = old.Log
det[name] = d
}
}
c.x.unitDet, c.x.unitDetAt, c.x.unitKey, c.x.syslogErr = det, now, key, syslogErr
c.mu.Unlock()
}
// unitDetailList: the failed units' details in the failed list's order; log lines go out once.
// Called with c.mu held.
func (c *Collector) unitDetailList(failed []string, reset bool) []UnitDetail {
var out []UnitDetail
for _, f := range failed {
if d, ok := c.x.unitDet[f]; ok {
out = append(out, d)
if reset && len(d.Log) > 0 {
d.Log = nil
c.x.unitDet[f] = d
}
}
}
return out
}
// FindPortOwners runs on the worker when the set of listening sockets changed.
func (c *Collector) FindPortOwners(inodes map[uint64]bool, procs []procRef, g *Gentle) {
c.mu.Lock()
known := c.portOwners
c.mu.Unlock()
owners := FindOwners(c.proc, inodes, known, procs, c.users, c.docker, g, 200_000)
c.mu.Lock()
c.portOwners = owners
c.mu.Unlock()
}
// ScanDisks runs the disk usage scan (worker).
func (c *Collector) ScanDisks(mounts []Mount, g *Gentle, now time.Time) {
if !c.cfg.On("disk_scan") {
return
}
out := c.x.disk.Scan(mounts, c.statfs, g, now)
c.mu.Lock()
c.x.diskOut = out
c.mu.Unlock()
slog.Info("disk usage scan finished", "entries", out.Entries, "took_s", out.TookS, "work_ms", out.WorkMs, "partial", out.Partial)
}
// SampleWriters reads every process's write counter (worker, once per rollup).
func (c *Collector) SampleWriters(procs []procRef, bootTime int64, g *Gentle, now time.Time) {
if !c.cfg.On("disk_scan") {
return
}
out := c.x.writers.Sample(c.proc, procs, bootTime, c.users, c.docker, g, now)
if out == nil {
return
}
c.mu.Lock()
c.x.writersOut = out
c.mu.Unlock()
}
// takeDisk returns (and clears) the pending disk scan and writers. Called with c.mu held.
func (c *Collector) takeDisk() (*DiskScanOut, []WriterOut) {
d, w := c.x.diskOut, c.x.writersOut
c.x.diskOut, c.x.writersOut = nil, nil
return d, w
}
// MountsSnapshot: the real filesystems of the latest sample (for the worker).
func (c *Collector) MountsSnapshot() []Mount {
out := make([]Mount, 0, len(c.mounts))
for _, m := range c.mounts {
if !skipMountPoint(m.Point) {
out = append(out, m)
}
}
return out
}
// InodesSnapshot copies the listening sockets' inodes (for the worker).
func (c *Collector) InodesSnapshot() map[uint64]bool {
m := make(map[uint64]bool, len(c.portInodes))
for k := range c.portInodes {
m[k] = true
}
return m
}
messages.go413 lines
package main
// What goes over the network. Sampling is local (every 60 s); the network sees four kinds
// of message, all gzip JSON posted to <url>/api/agent/v1/report:
//
// heartbeat every 5 min a few numbers (CPU, memory, worst disk and its fill ETA, active rules)
// rollup every 15 min per-minute averages / maxima since the last rollup, top processes,
// disks, network, disk I/O, TCP connections, kernel tables, health
// (failed units, OOM kills, SSH failures, time sync); the security
// settings check when it changed (at least daily); once a day also a
// hash of each inventory list
// event at once a local rule tripped (disk, inodes, memory, swap, load, OOM kill,
// open files / conntrack table, disk I/O, a failed unit, a security
// signal) or an inventory list changed; at most one per minute
// inventory rarely the full lists (listening ports, crontab lines, SSH key
// fingerprints, setuid files): after install or an upgrade, and when
// the server answers a daily hash with 205 (its copy drifted)
//
// Protocol version 1 since 0.1.0: 0.2.0 and 0.3.0 only add fields, which older servers ignore.
// 0.3.0 adds to the rollup: the agent's mode and own CPU / memory, the machine's virtualization,
// `ports` (the listening sockets with their owner, unit and container; when they change and at
// least daily), `logins` (wtmp / utmp / lastlog history and the auth log's logins, failures by
// address and sudo use), more posture (sshd details, accounts, key counts, fail2ban, firewall
// rules, package-manager activity), failed-unit details, `disk_scan` and `writers`.
import (
"crypto/sha256"
"encoding/hex"
"encoding/json"
"log/slog"
"os"
"path/filepath"
"sort"
"strings"
"sync/atomic"
"time"
)
type Header struct {
V int `json:"v"`
Type string `json:"type"`
T int64 `json:"t"` // unix seconds, UTC
Seq int64 `json:"seq"`
Version string `json:"version"`
}
type HeartbeatMsg struct {
Header
UptimeS int64 `json:"uptime_s"`
CPU float64 `json:"cpu"`
Mem float64 `json:"mem"`
Disk float64 `json:"disk"` // fullest filesystem, % used
DiskMount string `json:"disk_mount"`
FullInH *float64 `json:"full_in_h,omitempty"` // soonest projected full, hours
Active []string `json:"active"` // rule keys tripped now
Spooled int `json:"spooled"`
}
type RollupMsg struct {
Header
Agent AgentInfo `json:"agent"`
Host HostInfo `json:"host"`
Minutes []Minute `json:"minutes"`
Current Current `json:"current"`
Active []string `json:"active"`
InvHash map[string]string `json:"inv_hash,omitempty"` // once a day
WebRoot *WebRootOut `json:"web_root,omitempty"`
Health *HealthOut `json:"health,omitempty"`
Posture *PostureOut `json:"posture,omitempty"` // when it changed, and at least once a day
// 0.3.0
Ports []PortOut `json:"ports,omitempty"` // when they changed, and at least once a day
Logins *LoginsOut `json:"logins,omitempty"` // when something is new
DiskScan *DiskScanOut `json:"disk_scan,omitempty"` // after a scan finished
Writers []WriterOut `json:"writers,omitempty"` // bytes written in the window, by program
}
type SetDiff struct {
Added []string `json:"added,omitempty"`
Removed []string `json:"removed,omitempty"`
}
type EventMsg struct {
Header
Events []Event `json:"events,omitempty"`
InvDiff map[string]SetDiff `json:"inv_diff,omitempty"`
Ports []PortOut `json:"ports,omitempty"` // details of newly listening ports
Web *WebChanges `json:"web,omitempty"`
}
type InventoryMsg struct {
Header
Reason string `json:"reason"` // install | upgrade | resync
Inventory Inv `json:"inventory"`
Ports []PortOut `json:"ports"`
}
// Inv is the inventory as last sent. Setuid is nil until the first scan.
type Inv struct {
Ports []string `json:"ports"`
Cron []string `json:"cron"`
SSHKeys []string `json:"ssh_keys"`
Setuid []string `json:"setuid"`
}
func (i *Inv) fields() map[string]*[]string {
return map[string]*[]string{"ports": &i.Ports, "cron": &i.Cron, "ssh_keys": &i.SSHKeys, "setuid": &i.Setuid}
}
// HashItems is what the server recomputes over its copy of a list: sha256 of the sorted
// items joined by newlines, first 16 bytes in hex.
func HashItems(items []string) string {
s := sortedCopy(items)
sum := sha256.Sum256([]byte(strings.Join(s, "\n")))
return hex.EncodeToString(sum[:16])
}
// DiffInv compares the inventory as sent with the current one. A list that is nil in
// cur (setuid before its first scan) is left out.
func DiffInv(sent, cur Inv) map[string]SetDiff {
out := map[string]SetDiff{}
sf, cf := sent.fields(), cur.fields()
for name, cp := range cf {
if *cp == nil {
continue
}
a, r := DiffSets(*sf[name], *cp)
if len(a)+len(r) > 0 {
out[name] = SetDiff{Added: a, Removed: r}
}
}
return out
}
// invState is kept in the state directory: what the server has, as far as the agent knows.
type invState struct {
Sent Inv `json:"sent"`
InstalledAt int64 `json:"installed_at"`
HashDay int64 `json:"hash_day"`
AgentVersion string `json:"agent_version,omitempty"` // the version that sent Sent (0.1.0 did not write it)
PostureHash string `json:"posture_hash,omitempty"`
PostureDay int64 `json:"posture_day,omitempty"`
PortsHash string `json:"ports_hash,omitempty"`
PortsDay int64 `json:"ports_day,omitempty"`
}
type Outbox struct {
c *Collector
persist bool
state invState
haveFull bool // the server has a full inventory from us
resync atomic.Bool
seq int64
lastEvt time.Time
cur Inv
postureSent bool // since this start
portsSent bool // since this start
}
func (o *Outbox) path() string { return filepath.Join(o.c.cfg.StateDir, "inventory.json") }
func NewOutbox(c *Collector, persist bool) *Outbox {
o := &Outbox{c: c, persist: persist}
if persist {
if b, err := readSmall(o.path(), 8<<20); err == nil {
if err := json.Unmarshal(b, &o.state); err != nil {
slog.Warn("inventory state unreadable, sending the full lists again", "path", o.path(), "err", err)
o.state = invState{}
} else if o.state.InstalledAt > 0 {
o.haveFull = true
}
}
}
return o
}
func (o *Outbox) save() {
if !o.persist {
return
}
if b, err := json.Marshal(o.state); err == nil {
if err := writeFileAtomic(o.path(), b, 0o600); err != nil {
logLimited("inv-state", time.Hour, slog.LevelWarn, "cannot save the inventory state", "err", err)
}
}
}
func (o *Outbox) header(typ string, now time.Time) Header {
o.seq++
return Header{V: 1, Type: typ, T: now.Unix(), Seq: o.seq, Version: o.c.version}
}
// Resync is called when the server answered 205: the next message is the full inventory.
func (o *Outbox) Resync() { o.resync.Store(true) }
// RefreshInventory reads the lists that are cheap to read (ports from the last sample,
// crontabs and authorized_keys now) and takes the latest setuid scan.
func (o *Outbox) RefreshInventory() {
c := o.c
pk := map[string]bool{}
for _, p := range c.current.Ports {
pk[p.Key()] = true
}
ports := make([]string, 0, len(pk))
for k := range pk {
ports = append(ports, k)
}
sort.Strings(ports)
var cron, keys []string
if c.cfg.On("crontabs") {
var cu []string
cron, cu = ReadCrontabs(c.host)
cron, c.cronUnreadable = nonNil(cron), cu
}
if c.cfg.On("ssh_keys") {
var ku []string
keys, ku = ReadSSHKeys(c.host, c.users.All())
keys, c.sshUnreadable = nonNil(keys), ku
}
c.mu.Lock()
setuid := c.setuid
c.mu.Unlock()
// A list switched off with `disable:` stays nil: it is neither sent nor compared.
o.cur = Inv{Ports: ports, Cron: cron, SSHKeys: keys, Setuid: setuid}
}
func nonNil(xs []string) []string {
if xs == nil {
return []string{}
}
return xs
}
func (o *Outbox) portDetails(keys []string) []PortOut {
want := map[string]bool{}
for _, k := range keys {
want[k] = true
}
var out []PortOut
for _, p := range o.c.current.Ports {
if want[p.Key()] {
out = append(out, p)
}
}
return out
}
// NeedFull: after install (no saved state), after an upgrade (lists may be written differently:
// 0.2.0 redacts secrets in crontab lines) or a 205 from the server.
func (o *Outbox) NeedFull() bool { return !o.haveFull || o.resync.Load() || o.upgraded() }
func (o *Outbox) upgraded() bool {
return o.persist && o.haveFull && o.state.AgentVersion != o.c.version
}
func (o *Outbox) Full(now time.Time) InventoryMsg {
reason := "install"
switch {
case o.upgraded() && !o.resync.Load():
reason = "upgrade"
case o.haveFull:
reason = "resync"
}
m := InventoryMsg{Header: o.header("inventory", now), Reason: reason, Inventory: o.cur, Ports: o.portDetails(o.cur.Ports)}
if m.Inventory.Setuid == nil && o.state.Sent.Setuid != nil {
m.Inventory.Setuid = o.state.Sent.Setuid
}
o.state.Sent = m.Inventory
o.state.AgentVersion = o.c.version
if o.state.InstalledAt == 0 {
o.state.InstalledAt = now.Unix()
}
o.haveFull = true
o.resync.Store(false)
o.save()
return m
}
func (o *Outbox) Heartbeat(now time.Time, spooled int) HeartbeatMsg {
cur := o.c.current
m := HeartbeatMsg{Header: o.header("heartbeat", now), UptimeS: o.c.hostInfo.UptimeS, CPU: cur.CPU, Mem: cur.Mem.UsedPct,
Active: o.c.rules.Active(), Spooled: spooled}
for _, d := range cur.Disks {
if d.Pct >= m.Disk {
m.Disk, m.DiskMount = d.Pct, d.Mount
}
if d.FullInH != nil && (m.FullInH == nil || *d.FullInH < *m.FullInH) {
v := *d.FullInH
m.FullInH = &v
}
}
return m
}
// Rollup closes the chart window: per-minute aggregates since the last rollup.
func (o *Outbox) Rollup(now time.Time) RollupMsg {
c := o.c
m := RollupMsg{Header: o.header("rollup", now), Host: c.hostInfo, Active: c.rules.Active()}
m.Agent = AgentInfo{Version: c.version, Root: os.Geteuid() == 0, UID: os.Geteuid(), SampleS: int(c.cfg.SampleInterval / time.Second),
ReportS: int(c.cfg.RollupInterval / time.Second), StartedAt: c.started.Unix(), Skipped: nonNil(c.Skipped()),
Disabled: c.Disabled(), Panics: int(panicCount.Load()), Mode: c.ident.Mode, Caps: c.ident.Caps}
m.Agent.CPUPct, m.Agent.CPUPeak, m.Agent.RSS, m.Agent.RSSMax, m.Agent.WindowS = c.self.window(now)
keys := make([]int64, 0, len(c.minutes))
for t := range c.minutes {
keys = append(keys, t)
}
sort.Slice(keys, func(i, j int) bool { return keys[i] < keys[j] })
for _, t := range keys {
a := c.minutes[t]
n := float64(a.n)
m.Minutes = append(m.Minutes, Minute{T: t, CPU: round1(a.cpu / n), CPUMax: round1(a.cpuMax), Load1: round2(a.load / n),
Mem: round1(a.mem / n), MemMax: round1(a.memMax), Swap: round1(a.swap / n), SwapIO: round1(a.swapIO / n),
RX: float64(int64(a.rx / n)), TX: float64(int64(a.tx / n)), Disk: round1(a.disk), TCPEst: round1(a.tcpEst / n),
IO: round1(a.io), OOM: a.oom, N: a.n})
}
c.minutes = map[int64]*minuteAgg{}
m.Current = c.current
m.Current.Ports = nil // ports travel as inventory diffs
c.mu.Lock()
if c.web != nil {
w := *c.web
w.Changes = nil // changes travel in event messages
m.WebRoot = &w
}
c.mu.Unlock()
c.mu.Lock()
m.Logins = c.takeLogins()
m.DiskScan, m.Writers = c.takeDisk()
c.mu.Unlock()
h := c.Health(now, true)
m.Health = &h
day := now.Unix() / 86400
if ports := c.current.Ports; ports != nil || !o.portsSent {
ph := hashJSON(ports)
if !o.portsSent || ph != o.state.PortsHash || day != o.state.PortsDay {
m.Ports = nonNilPorts(ports)
o.portsSent = true
o.state.PortsHash, o.state.PortsDay = ph, day
o.save()
}
}
if p := c.Posture(); p != nil {
cp := *p
cp.At = 0
b, _ := json.Marshal(cp)
sum := sha256.Sum256(b)
ph := hex.EncodeToString(sum[:8])
if !o.postureSent || ph != o.state.PostureHash || day != o.state.PostureDay {
m.Posture = p
o.postureSent = true
o.state.PostureHash, o.state.PostureDay = ph, day
o.save()
}
}
if o.haveFull && day != o.state.HashDay {
m.InvHash = map[string]string{}
for name, p := range o.state.Sent.fields() {
if *p != nil {
m.InvHash[name] = HashItems(*p)
}
}
o.state.HashDay = day
o.save()
}
return m
}
// Event returns the pending event message (tripped rules, inventory changes, web root
// changes), or nil when there is nothing or the last one went out less than gap ago.
func (o *Outbox) Event(now time.Time, gap time.Duration) *EventMsg {
if now.Sub(o.lastEvt) < gap {
return nil
}
c := o.c
var diff map[string]SetDiff
if o.haveFull {
diff = DiffInv(o.state.Sent, o.cur)
}
c.mu.Lock()
web := c.webPend
c.mu.Unlock()
if len(c.pending) == 0 && len(diff) == 0 && web == nil {
return nil
}
m := &EventMsg{Header: o.header("event", now), Events: c.pending, Web: web}
if len(diff) > 0 {
m.InvDiff = diff
if d, ok := diff["ports"]; ok {
m.Ports = o.portDetails(d.Added)
}
for name, p := range o.cur.fields() {
if *p != nil {
*o.state.Sent.fields()[name] = *p
}
}
o.save()
}
c.pending = nil
c.mu.Lock()
c.webPend = nil
c.mu.Unlock()
o.lastEvt = now
return m
}
func (h Header) kind() string { return h.Type }
func nonNilPorts(p []PortOut) []PortOut {
if p == nil {
return []PortOut{}
}
return p
}
rules.go285 lines
package main
import (
"sort"
"time"
)
// Local rules decide on the server when something is worth reporting at once (an "event"
// message) instead of waiting for the 15-minute rollup. Each rule has a trip condition, a
// clear condition with hysteresis and a minimum duration, so a single odd sample never trips it.
const (
diskBad, diskGood = 90.0, 87.0
fullSoonH, fullOkH = 48.0, 72.0
memBad, memGood = 95.0, 90.0
swapBad, swapGood = 500.0, 100.0 // pages swapped in + out per second
loadFactor, loadGood = 2.0, 1.5 // x CPU cores
memFor, swapFor = 10 * time.Minute, 10 * time.Minute
loadFor = 15 * time.Minute
diskFor = 2 * time.Minute // two samples in a row at the default interval
clearFor = 5 * time.Minute
signalGoneFor = 5 * time.Minute
fdBad, fdGood = 90.0, 80.0 // % of the kernel's open-file (or conntrack) table
fdFor = 5 * time.Minute
ioUtilBad, ioWaitBad = 90.0, 50.0 // % busy and ms per request, both at once ...
ioFor = 15 * time.Minute
ioUtilOk, ioWaitOk = 70.0, 20.0 // ... clears when either is back under these
oomQuiet = 30 * time.Minute // an OOM incident ends after this long without a kill
maxUnitRules = 50
)
// Event is one tripped rule, with its evidence.
type Event struct {
Kind string `json:"kind"` // disk | inode | memory | swap | load | oom | fd | conntrack | io | unit | miner | tmp_exe | deleted_exe | cpu_hog
Key string `json:"key"`
At int64 `json:"at"`
Detail map[string]any `json:"detail"`
}
type ruleState struct {
badSince time.Time // zero: not bad now
goodSince time.Time
active bool
}
// Rules keeps the state of every rule key ("disk:/", "memory", "miner:xmrig:/tmp/xmrig", ...).
type Rules struct {
st map[string]*ruleState
lastSeen map[string]time.Time // process signals: last sample that showed them
lastOOM time.Time
}
func NewRules() *Rules { return &Rules{st: map[string]*ruleState{}, lastSeen: map[string]time.Time{}} }
// step advances one key. bad/good are this sample's readings; it returns true when the key
// trips now (it was not active) and false otherwise; cleared reports the reverse.
func (r *Rules) step(key string, now time.Time, bad, good bool, need time.Duration) (tripped, cleared bool) {
s := r.st[key]
if s == nil {
s = &ruleState{}
r.st[key] = s
}
if bad {
if s.badSince.IsZero() {
s.badSince = now
}
} else {
s.badSince = time.Time{}
}
if good {
if s.goodSince.IsZero() {
s.goodSince = now
}
} else {
s.goodSince = time.Time{}
}
if !s.active && bad && now.Sub(s.badSince) >= need-time.Second {
s.active = true
return true, false
}
if s.active && good && now.Sub(s.goodSince) >= clearFor-time.Second {
s.active = false
return false, true
}
return false, false
}
// Reading is what the rules look at in one sample.
type Reading struct {
Mem, SwapIO, SwapPct, Load1 float64
Cores int
Disks []DiskOut
TopCPU, TopMem []Proc
Signals map[string]Signal // process signals seen in this sample
OOMKills uint64 // processes killed by the OOM killer since the previous sample
OOMTotal uint64 // since boot
OOMVictims []Proc // large processes that disappeared at the same time (probably the victims)
FD, Conntrack *TableOut // nil: not readable here
IO []IOOut // busiest device first
FailedUnits []string
UnitsKnown bool // the systemd check works (FailedUnits is meaningful)
}
// Update runs every rule on one sample. Returns the events that tripped now and the keys
// that cleared now.
func (r *Rules) Update(now time.Time, in Reading) (events []Event, cleared []string) {
add := func(kind, key string, detail map[string]any) {
events = append(events, Event{Kind: kind, Key: key, At: now.Unix(), Detail: detail})
}
mounts := map[string]bool{}
for _, d := range in.Disks {
mounts[d.Mount] = true
soon := d.FullInH != nil && *d.FullInH < fullSoonH
good := d.Pct < diskGood && (d.FullInH == nil || *d.FullInH >= fullOkH)
if t, c := r.step("disk:"+d.Mount, now, d.Pct >= diskBad || soon, good, diskFor); t {
add("disk", "disk:"+d.Mount, map[string]any{"mount": d.Mount, "pct": d.Pct, "full_in_h": d.FullInH, "rate_bph": d.RateBPH,
"total": d.Total, "used": d.Used})
} else if c {
cleared = append(cleared, "disk:"+d.Mount)
}
if t, c := r.step("inode:"+d.Mount, now, d.InodesPct >= diskBad, d.InodesPct < diskGood, diskFor); t {
add("inode", "inode:"+d.Mount, map[string]any{"mount": d.Mount, "pct": d.InodesPct})
} else if c {
cleared = append(cleared, "inode:"+d.Mount)
}
}
for key, s := range r.st { // a filesystem that went away ends its rules
for _, p := range []string{"disk:", "inode:"} {
if len(key) > len(p) && key[:len(p)] == p && !mounts[key[len(p):]] {
if s.active {
cleared = append(cleared, key)
}
delete(r.st, key)
}
}
}
top := func(ps []Proc) any {
if len(ps) == 0 {
return nil
}
return ps[0]
}
minutes := func(key string) int { return int(now.Sub(r.st[key].badSince).Minutes() + 0.5) }
if t, c := r.step("memory", now, in.Mem >= memBad, in.Mem < memGood, memFor); t {
add("memory", "memory", map[string]any{"pct": in.Mem, "minutes": minutes("memory"), "top": top(in.TopMem)})
} else if c {
cleared = append(cleared, "memory")
}
if t, c := r.step("swap", now, in.SwapIO >= swapBad, in.SwapIO < swapGood, swapFor); t {
add("swap", "swap", map[string]any{"rate": in.SwapIO, "pct": in.SwapPct, "minutes": minutes("swap")})
} else if c {
cleared = append(cleared, "swap")
}
cores := float64(maxInt(in.Cores, 1))
if t, c := r.step("load", now, in.Load1 > loadFactor*cores, in.Load1 < loadGood*cores, loadFor); t {
add("load", "load", map[string]any{"load": in.Load1, "cores": in.Cores, "limit": loadFactor * cores, "minutes": minutes("load"),
"top": top(in.TopCPU)})
} else if c {
cleared = append(cleared, "load")
}
// Out-of-memory kills: at once, and the incident lasts until 30 minutes pass without a kill.
if in.OOMKills > 0 {
r.lastOOM = now
if s := r.st["oom"]; s == nil || !s.active {
r.st["oom"] = &ruleState{active: true, badSince: now}
victims := in.OOMVictims
if victims == nil {
victims = []Proc{}
}
add("oom", "oom", map[string]any{"kills": in.OOMKills, "total": in.OOMTotal, "victims": victims})
}
} else if s := r.st["oom"]; s != nil && s.active && now.Sub(r.lastOOM) >= oomQuiet {
delete(r.st, "oom")
cleared = append(cleared, "oom")
}
// Kernel tables close to full: open files (EMFILE / ENFILE for everyone) and conntrack (new
// connections dropped).
for _, tb := range []struct {
key string
t *TableOut
}{{"fd", in.FD}, {"conntrack", in.Conntrack}} {
if tb.t == nil {
continue
}
if t, c := r.step(tb.key, now, tb.t.Pct >= fdBad, tb.t.Pct < fdGood, fdFor); t {
add(tb.key, tb.key, map[string]any{"pct": tb.t.Pct, "used": tb.t.Used, "max": tb.t.Max, "minutes": minutes(tb.key)})
} else if c {
cleared = append(cleared, tb.key)
}
}
// Disk I/O saturated: busy nearly all the time AND slow to answer (an NVMe drive can be "100%
// busy" and still fast), for 15 minutes.
var worst IOOut
for _, d := range in.IO {
if d.Util > worst.Util || (d.Util == worst.Util && d.AwaitMs > worst.AwaitMs) {
worst = d
}
}
ioBad := worst.Util >= ioUtilBad && worst.AwaitMs >= ioWaitBad
ioGood := worst.Util < ioUtilOk || worst.AwaitMs < ioWaitOk
if t, c := r.step("io", now, ioBad, ioGood, ioFor); t {
add("io", "io", map[string]any{"dev": worst.Dev, "util": worst.Util, "await_ms": worst.AwaitMs, "rbps": worst.RBps, "wbps": worst.WBps,
"minutes": minutes("io")})
} else if c {
cleared = append(cleared, "io")
}
// Failed systemd units: at once; a unit that is no longer failed clears 5 minutes later.
if in.UnitsKnown {
failed := map[string]bool{}
for _, u := range in.FailedUnits {
if len(failed) < maxUnitRules {
failed["unit:"+u] = true
}
}
keys := make([]string, 0, len(failed))
for k := range failed {
keys = append(keys, k)
}
sort.Strings(keys)
for _, k := range keys {
if t, _ := r.step(k, now, true, false, 0); t {
add("unit", k, map[string]any{"unit": k[len("unit:"):]})
}
}
for k, s := range r.st {
if len(k) <= 5 || k[:5] != "unit:" || failed[k] {
continue
}
if _, c := r.step(k, now, false, true, 0); c {
cleared = append(cleared, k)
}
if !s.active {
delete(r.st, k)
}
}
}
// Process signals: active while seen, cleared after 5 minutes without them.
keys := make([]string, 0, len(in.Signals))
for k := range in.Signals {
keys = append(keys, k)
}
sort.Strings(keys)
for _, k := range keys {
sig := in.Signals[k]
r.lastSeen[k] = now
s := r.st[k]
if s == nil || !s.active {
r.st[k] = &ruleState{active: true, badSince: now}
add(sig.Kind, k, sig.Detail)
}
}
for k, seen := range r.lastSeen {
if _, now2 := in.Signals[k]; now2 {
continue
}
if now.Sub(seen) >= signalGoneFor {
delete(r.lastSeen, k)
if s := r.st[k]; s != nil && s.active {
cleared = append(cleared, k)
}
delete(r.st, k)
}
}
sort.Strings(cleared)
return events, cleared
}
// Active lists the keys that are tripped now.
func (r *Rules) Active() []string {
out := []string{}
for k, s := range r.st {
if s.active {
out = append(out, k)
}
}
sort.Strings(out)
return out
}
func maxInt(a, b int) int {
if a > b {
return a
}
return b
}
security.go523 lines
package main
// Security signals. All of them are heuristics: a match is a reason to look, never proof,
// and the absence of a match proves nothing either. The agent only reads: it never kills,
// quarantines or changes anything.
import (
"bufio"
"bytes"
"crypto/sha256"
_ "embed"
"encoding/base64"
"errors"
"io/fs"
"os"
"path/filepath"
"sort"
"strings"
)
//go:embed signatures.txt
var signaturesTxt []byte
// Signatures from signatures.txt.
type Signatures struct {
Names map[string]bool
Args []string
Busy map[string]bool
}
func ParseSignatures(b []byte) Signatures {
s := Signatures{Names: map[string]bool{}, Busy: map[string]bool{}}
sc := bufio.NewScanner(bytes.NewReader(b))
for sc.Scan() {
line := strings.TrimSpace(sc.Text())
if line == "" || strings.HasPrefix(line, "#") {
continue
}
kind, val, ok := strings.Cut(line, " ")
val = strings.ToLower(strings.TrimSpace(val))
if !ok || val == "" {
continue
}
switch kind {
case "name":
s.Names[val] = true
case "arg":
s.Args = append(s.Args, val)
case "busy":
s.Busy[val] = true
}
}
return s
}
var sigs = ParseSignatures(signaturesTxt)
// Programs that show other programs' names in their arguments (someone looking for a
// miner with `grep xmrig` is not running one).
var lookers = map[string]bool{
"grep": true, "egrep": true, "fgrep": true, "rg": true, "ag": true, "pgrep": true, "pkill": true, "ps": true,
"less": true, "more": true, "tail": true, "head": true, "cat": true, "vi": true, "vim": true, "nano": true,
"emacs": true, "journalctl": true, "man": true, "find": true, "locate": true, "htop": true, "top": true,
"approvalens-age": true, "approvalens-agent": true,
}
// MinerMatch: why a process looks like a crypto miner ("" = it does not).
func (s Signatures) MinerMatch(comm, argv0, cmdline string) string {
c := strings.ToLower(comm)
a0 := strings.ToLower(filepath.Base(argv0))
if s.Names[c] {
return "name:" + c
}
if argv0 != "" && s.Names[a0] {
return "name:" + a0
}
if lookers[c] || lookers[a0] {
return ""
}
lc := strings.ToLower(cmdline)
for _, a := range s.Args {
if strings.Contains(lc, a) {
return "arg:" + a
}
}
return ""
}
var tmpDirs = []string{"/tmp/", "/var/tmp/", "/dev/shm/", "/run/shm/"}
// InTmp: an executable that runs from a world-writable temporary directory.
func InTmp(path string) bool {
for _, d := range tmpDirs {
if strings.HasPrefix(path, d) {
return true
}
}
return false
}
// SuspiciousDeleted: the process runs a binary that was deleted after it started. A
// package upgrade does that to every running daemon, so binaries under the system
// directories are left out; a memfd (a program that never touched the disk) is not.
func SuspiciousDeleted(exe string) bool {
if !strings.HasSuffix(exe, " (deleted)") {
return false
}
p := strings.TrimSuffix(exe, " (deleted)")
if strings.HasPrefix(p, "/memfd:") {
return true
}
for _, d := range []string{"/usr/", "/bin/", "/sbin/", "/lib/", "/lib64/", "/snap/", "/opt/", "/nix/", "/var/lib/docker/"} {
if strings.HasPrefix(p, d) {
return false
}
}
return true
}
// ---------------------------------------------------------------------------
// crontabs
// ---------------------------------------------------------------------------
// ParseCrontab returns the active lines of one crontab ("<source>: <line>"), whitespace
// collapsed, comments and blank lines dropped, secrets redacted (see Redact: a job like
// `mysqldump -pSECRET` or `curl -H "Authorization: Bearer ..."` is sent with *** instead).
// Variable lines (PATH=..., MAILTO=...) are kept: a changed PATH is a known way to hijack a job.
func ParseCrontab(content []byte, source string) []string {
var out []string
sc := bufio.NewScanner(bytes.NewReader(content))
sc.Buffer(make([]byte, 64*1024), 1024*1024)
for sc.Scan() {
line := strings.Join(strings.Fields(sc.Text()), " ")
if line == "" || strings.HasPrefix(line, "#") {
continue
}
out = append(out, truncate(source+": "+Redact(line), 300))
}
return out
}
// DiffSets returns the items of cur that are not in prev, and the items of prev that are
// not in cur (both sorted).
func DiffSets(prev, cur []string) (added, removed []string) {
p := make(map[string]bool, len(prev))
for _, x := range prev {
p[x] = true
}
c := make(map[string]bool, len(cur))
for _, x := range cur {
c[x] = true
if !p[x] {
added = append(added, x)
}
}
for _, x := range prev {
if !c[x] {
removed = append(removed, x)
}
}
sort.Strings(added)
sort.Strings(removed)
return dedupSorted(added), dedupSorted(removed)
}
func dedupSorted(xs []string) []string {
if len(xs) < 2 {
return xs
}
out := xs[:1]
for _, x := range xs[1:] {
if x != out[len(out)-1] {
out = append(out, x)
}
}
return out
}
// ReadCrontabs collects every system and user crontab it can read. unreadable lists the
// places it could not (other users' crontabs need root).
func ReadCrontabs(root string) (entries []string, unreadable []string) {
files := []string{filepath.Join(root, "etc/crontab")}
for _, dir := range []string{"etc/cron.d", "var/spool/cron/crontabs", "var/spool/cron"} {
d := filepath.Join(root, dir)
ents, err := os.ReadDir(d)
if err != nil {
if errors.Is(err, fs.ErrPermission) {
unreadable = append(unreadable, "/"+dir)
}
continue
}
for _, e := range ents {
if e.Type().IsRegular() && !strings.HasPrefix(e.Name(), ".") {
files = append(files, filepath.Join(d, e.Name()))
}
}
}
for _, f := range files {
b, err := readSmall(f, 256*1024)
if err != nil {
if errors.Is(err, fs.ErrPermission) {
unreadable = append(unreadable, strings.TrimPrefix(f, strings.TrimSuffix(root, "/")))
}
continue
}
entries = append(entries, ParseCrontab(b, strings.TrimPrefix(f, strings.TrimSuffix(root, "/")))...)
}
sort.Strings(entries)
entries = dedupSorted(entries)
if len(entries) > 300 {
entries = entries[:300]
}
return entries, unreadable
}
// ---------------------------------------------------------------------------
// SSH authorized keys
// ---------------------------------------------------------------------------
// ParseAuthorizedKeys returns "<user> <type> SHA256:<fingerprint> <comment>" per key.
func ParseAuthorizedKeys(b []byte, user string) []string {
var out []string
sc := bufio.NewScanner(bytes.NewReader(b))
sc.Buffer(make([]byte, 64*1024), 1024*1024)
for sc.Scan() {
line := strings.TrimSpace(sc.Text())
if line == "" || strings.HasPrefix(line, "#") {
continue
}
f := strings.Fields(line)
idx := -1
for i, x := range f {
if keyType(x) {
idx = i
break
}
}
if idx < 0 || idx+1 >= len(f) {
continue
}
blob, err := base64.StdEncoding.DecodeString(f[idx+1])
if err != nil {
continue
}
sum := sha256.Sum256(blob)
item := user + " " + f[idx] + " SHA256:" + base64.RawStdEncoding.EncodeToString(sum[:])
if idx > 0 {
item += " [options]"
}
if idx+2 < len(f) {
item += " " + truncate(strings.Join(f[idx+2:], " "), 60)
}
out = append(out, item)
}
return out
}
func keyType(s string) bool {
return strings.HasPrefix(s, "ssh-") || strings.HasPrefix(s, "ecdsa-sha2-") || strings.HasPrefix(s, "sk-ssh-") ||
strings.HasPrefix(s, "sk-ecdsa-")
}
// ReadSSHKeys reads authorized_keys of root and of every account with a login shell.
func ReadSSHKeys(root string, users []PasswdEntry) (keys []string, unreadable []string) {
seen := map[string]bool{}
for _, u := range users {
if u.UID != 0 && !LoginShell(u.Shell) {
continue
}
if u.Home == "" || u.Home == "/" || seen[u.Home] {
continue
}
seen[u.Home] = true
for _, name := range []string{"authorized_keys", "authorized_keys2"} {
p := filepath.Join(root, u.Home, ".ssh", name)
b, err := readSmall(p, 512*1024)
if err != nil {
if errors.Is(err, fs.ErrPermission) {
unreadable = append(unreadable, filepath.Join(u.Home, ".ssh", name))
}
continue
}
keys = append(keys, ParseAuthorizedKeys(b, u.Name)...)
}
}
sort.Strings(keys)
keys = dedupSorted(keys)
if len(keys) > 200 {
keys = keys[:200]
}
return keys, dedupSorted(sortedCopy(unreadable))
}
// ---------------------------------------------------------------------------
// setuid / setgid binaries
// ---------------------------------------------------------------------------
var setuidRoots = []string{"/bin", "/sbin", "/usr/bin", "/usr/sbin", "/usr/local/bin", "/usr/local/sbin", "/usr/lib",
"/usr/libexec", "/usr/local/lib", "/opt", "/tmp", "/var/tmp", "/dev/shm"}
// ScanSetuid walks the usual binary directories (and the temporary ones) for setuid or
// setgid files: "<path> <mode> uid=<n>". It stays light: no symlinks are followed and
// it stops after maxFiles entries.
func ScanSetuid(roots []string, maxFiles int, g *Gentle) (items []string, truncated bool) {
seen := map[string]bool{}
count := 0
for _, r := range roots {
real, err := filepath.EvalSymlinks(r)
if err != nil || seen[real] {
continue
}
seen[real] = true
_ = filepath.WalkDir(real, func(p string, d fs.DirEntry, err error) error {
if err != nil {
if d != nil && d.IsDir() {
return fs.SkipDir
}
return nil
}
count++
g.Tick()
if count > maxFiles {
truncated = true
return fs.SkipAll
}
if d.IsDir() {
if strings.Count(strings.TrimPrefix(p, real), "/") > 8 {
return fs.SkipDir
}
return nil
}
if !d.Type().IsRegular() {
return nil
}
info, err := d.Info()
if err != nil {
return nil
}
m := info.Mode()
if m&(fs.ModeSetuid|fs.ModeSetgid) == 0 {
return nil
}
items = append(items, p+" "+modeString(m)+" uid="+itoa(fileUID(info)))
return nil
})
if truncated {
break
}
}
sort.Strings(items)
return items, truncated
}
func modeString(m fs.FileMode) string {
v := uint32(m.Perm())
if m&fs.ModeSetuid != 0 {
v |= 0o4000
}
if m&fs.ModeSetgid != 0 {
v |= 0o2000
}
if m&fs.ModeSticky != 0 {
v |= 0o1000
}
s := []byte("0000")
for i := 3; i >= 0; i-- {
s[i] = byte('0' + v&7)
v >>= 3
}
return string(s)
}
// ---------------------------------------------------------------------------
// web roots
// ---------------------------------------------------------------------------
var webExts = map[string]bool{".php": true, ".phtml": true, ".php3": true, ".php4": true, ".php5": true, ".php7": true,
".phar": true, ".inc": true, ".js": true, ".mjs": true, ".html": true, ".htm": true, ".shtml": true, ".htaccess": true,
".user.ini": true}
var webSkipDirs = map[string]bool{"node_modules": true, ".git": true, ".svn": true, ".hg": true, "cache": true, ".cache": true}
// WebFile key: modification time and size, never the content.
func webWatched(name string) bool {
if name == ".htaccess" || name == ".user.ini" {
return true
}
return webExts[strings.ToLower(filepath.Ext(name))]
}
// ScanWebRoots maps each PHP / JS / HTML file (and .htaccess) under the roots to
// "<mtime>:<size>". Stops at maxFiles (truncated=true).
func ScanWebRoots(roots []string, exclude []string, maxFiles int, g *Gentle) (files map[string]string, truncated bool) {
files = map[string]string{}
skip := map[string]bool{}
for k := range webSkipDirs {
skip[k] = true
}
for _, x := range exclude {
skip[x] = true
}
for _, r := range roots {
_ = filepath.WalkDir(r, func(p string, d fs.DirEntry, err error) error {
if err != nil {
if d != nil && d.IsDir() {
return fs.SkipDir
}
return nil
}
g.Tick()
if d.IsDir() {
if p != r && (skip[d.Name()] || skip[p]) {
return fs.SkipDir
}
return nil
}
if !d.Type().IsRegular() || !webWatched(d.Name()) {
return nil
}
if len(files) >= maxFiles {
truncated = true
return fs.SkipAll
}
info, err := d.Info()
if err != nil {
return nil
}
files[p] = itoa64(info.ModTime().Unix()) + ":" + itoa64(info.Size())
return nil
})
if truncated {
break
}
}
return files, truncated
}
// WebChanges between two scans: new, changed and removed file names (capped).
type WebChanges struct {
Added []string `json:"added"`
Changed []string `json:"changed"`
Removed []string `json:"removed"`
NAdded int `json:"n_added"`
NChanged int `json:"n_changed"`
NRemoved int `json:"n_removed"`
Files int `json:"files"`
Truncated bool `json:"truncated,omitempty"`
}
func (w WebChanges) Empty() bool { return w.NAdded+w.NChanged+w.NRemoved == 0 }
func DiffWebRoots(prev, cur map[string]string, limit int) WebChanges {
var w WebChanges
w.Files = len(cur)
for p, v := range cur {
old, ok := prev[p]
switch {
case !ok:
w.NAdded++
w.Added = append(w.Added, p)
case old != v:
w.NChanged++
w.Changed = append(w.Changed, p)
}
}
for p := range prev {
if _, ok := cur[p]; !ok {
w.NRemoved++
w.Removed = append(w.Removed, p)
}
}
for _, xs := range []*[]string{&w.Added, &w.Changed, &w.Removed} {
sort.Strings(*xs)
if len(*xs) > limit {
*xs = (*xs)[:limit]
}
}
return w
}
// ---------------------------------------------------------------------------
// helpers
// ---------------------------------------------------------------------------
func readSmall(path string, max int64) ([]byte, error) {
f, err := os.Open(path)
if err != nil {
return nil, err
}
defer f.Close()
buf := make([]byte, 0, 4096)
tmp := make([]byte, 32*1024)
for int64(len(buf)) < max {
n, err := f.Read(tmp)
buf = append(buf, tmp[:n]...)
if err != nil {
break
}
}
if int64(len(buf)) > max {
buf = buf[:max]
}
return buf, nil
}
func truncate(s string, n int) string {
if len(s) <= n {
return s
}
// keep valid UTF-8: back off to a rune start
cut := n
for cut > 0 && (s[cut]&0xC0) == 0x80 {
cut--
}
return s[:cut] + "…"
}
func sortedCopy(xs []string) []string {
out := append([]string(nil), xs...)
sort.Strings(out)
return out
}
posture.go672 lines
package main
// Security posture: a handful of settings that make a server easy to break into or leave it
// unpatched, read once an hour from configuration files and from what the agent already knows
// (listening ports, accounts, systemd units, processes). Only verdict-sized facts leave the
// machine: the effective value of a few sshd options, which host firewall was found, whether
// automatic updates are on, the names of extra UID 0 accounts, the number of sudoers NOPASSWD
// rules and the users or groups they name, the database ports reachable from the network,
// whether a reboot is pending and how many updates wait. Never the content of a file.
import (
"bufio"
"bytes"
"errors"
"io/fs"
"net"
"os"
"path/filepath"
"regexp"
"sort"
"strconv"
"strings"
)
type PostureOut struct {
At int64 `json:"at"`
SSH *SSHPosture `json:"ssh,omitempty"` // nil: no sshd configuration on this machine
Firewall FirewallOut `json:"firewall"`
AutoUpd AutoUpdOut `json:"auto_updates"`
UID0 []string `json:"uid0"` // accounts other than root with UID 0
Sudo SudoOut `json:"sudo_nopasswd"`
PublicDB []PublicPort `json:"public_db"`
Reboot RebootOut `json:"reboot"`
Updates *UpdatesOut `json:"updates,omitempty"` // nil: this system keeps no update count the agent can read
// 0.3.0
Fail2ban *Fail2banOut `json:"fail2ban,omitempty"` // nil: fail2ban is not installed
Users []UserOut `json:"users,omitempty"`
FW *FWOut `json:"fw,omitempty"`
Pkg *PkgOut `json:"pkg,omitempty"`
}
// ---------------------------------------------------------------------------
// sshd
// ---------------------------------------------------------------------------
type SSHPosture struct {
Readable bool `json:"readable"`
PermitRootLogin string `json:"permit_root_login,omitempty"` // effective values (defaults applied)
PasswordAuth string `json:"password_auth,omitempty"`
KbdInteractive string `json:"kbd_interactive,omitempty"`
UsePAM string `json:"use_pam,omitempty"`
PermitEmpty string `json:"permit_empty,omitempty"`
AuthMethods string `json:"auth_methods,omitempty"`
PasswordLogin bool `json:"password_login"` // a password can log in, by the values above
Ports []int `json:"ports,omitempty"`
Running bool `json:"running"` // an sshd process runs
// 0.3.0
PubkeyAuth string `json:"pubkey_auth,omitempty"`
MaxAuthTries int `json:"max_auth_tries,omitempty"`
AllowUsers []string `json:"allow_users,omitempty"`
AllowGroups []string `json:"allow_groups,omitempty"`
DenyUsers []string `json:"deny_users,omitempty"`
DenyGroups []string `json:"deny_groups,omitempty"`
MatchBlocks int `json:"match_blocks,omitempty"`
MatchPassword bool `json:"match_password,omitempty"` // a Match block turns password logins on
MatchRoot bool `json:"match_root,omitempty"` // a Match block sets PermitRootLogin yes
Keys []KeyCount `json:"keys,omitempty"` // authorized_keys per account: count and key types
}
// KeyCount: how many keys an account's authorized_keys holds, and of which types (no fingerprints).
type KeyCount struct {
User string `json:"user"`
Count int `json:"count"`
Types []string `json:"types"`
}
type sshdParse struct {
root string
vals map[string]string
lists map[string][]string // allowusers & co. accumulate over lines (sshd appends them)
ports []int
files int
unreadable bool
matches int
matchPw bool
matchRoot bool
}
// ReadSSHD reads /etc/ssh/sshd_config and the files it includes the way sshd does: keywords are
// case-insensitive, the first value of a keyword wins, Include is followed (globs, in order) and
// Match blocks (settings for some users or addresses only) are left out.
func ReadSSHD(root string) *SSHPosture {
main := filepath.Join(root, "etc/ssh/sshd_config")
if _, err := os.Lstat(main); err != nil {
alt := filepath.Join(root, "usr/etc/ssh/sshd_config") // openSUSE keeps the vendor file here
if _, err2 := os.Lstat(alt); err2 != nil {
return nil
}
main = alt
}
p := &sshdParse{root: root, vals: map[string]string{}, lists: map[string][]string{}}
p.file(main, 0)
out := &SSHPosture{Readable: p.files > 0 && !p.unreadable}
if !out.Readable {
return out
}
get := func(k, def string) string {
if v, ok := p.vals[k]; ok {
return v
}
return def
}
out.PermitRootLogin = get("permitrootlogin", "prohibit-password")
if out.PermitRootLogin == "without-password" {
out.PermitRootLogin = "prohibit-password"
}
out.PasswordAuth = get("passwordauthentication", "yes")
out.KbdInteractive = get("kbdinteractiveauthentication", get("challengeresponseauthentication", "yes"))
out.UsePAM = get("usepam", "no")
out.PermitEmpty = get("permitemptypasswords", "no")
out.AuthMethods = truncate(get("authenticationmethods", "any"), 120)
methodsAllow := out.AuthMethods == "any" || strings.Contains(out.AuthMethods, "password") || strings.Contains(out.AuthMethods, "keyboard-interactive")
out.PasswordLogin = methodsAllow && (out.PasswordAuth == "yes" || (out.KbdInteractive == "yes" && out.UsePAM == "yes"))
out.Ports = p.ports
if len(out.Ports) == 0 {
out.Ports = []int{22}
}
out.PubkeyAuth = get("pubkeyauthentication", "yes")
out.MaxAuthTries = 6
if n, err := strconv.Atoi(get("maxauthtries", "6")); err == nil && n > 0 && n < 10000 {
out.MaxAuthTries = n
}
out.AllowUsers, out.AllowGroups = p.lists["allowusers"], p.lists["allowgroups"]
out.DenyUsers, out.DenyGroups = p.lists["denyusers"], p.lists["denygroups"]
out.MatchBlocks, out.MatchPassword, out.MatchRoot = p.matches, p.matchPw, p.matchRoot
return out
}
func (p *sshdParse) file(path string, depth int) {
if depth > 8 || p.files >= 64 {
return
}
b, err := readSmall(path, 256*1024)
if err != nil {
if errors.Is(err, fs.ErrPermission) {
p.unreadable = true
}
return
}
p.files++
inMatch := false
sc := bufio.NewScanner(bytes.NewReader(b))
sc.Buffer(make([]byte, 16*1024), 64*1024)
for sc.Scan() {
line := strings.TrimSpace(sc.Text())
if line == "" || line[0] == '#' {
continue
}
key, rest := splitSSHDLine(line)
if key == "match" {
inMatch = strings.ToLower(rest) != "all"
if inMatch {
p.matches++
}
continue
}
if inMatch {
v := strings.ToLower(unquote(firstField(rest)))
switch {
case (key == "passwordauthentication" || key == "kbdinteractiveauthentication") && v == "yes":
p.matchPw = true
case key == "permitrootlogin" && v == "yes":
p.matchRoot = true
}
continue
}
switch key {
case "include":
for _, pat := range strings.Fields(rest) {
pat = unquote(pat)
if !strings.HasPrefix(pat, "/") {
pat = "/etc/ssh/" + pat
}
matches, _ := filepath.Glob(filepath.Join(p.root, pat))
sort.Strings(matches)
for _, m := range matches {
p.file(m, depth+1)
}
}
case "port":
if n, err := strconv.Atoi(firstField(rest)); err == nil && n > 0 && n < 65536 && len(p.ports) < 16 {
p.ports = append(p.ports, n)
}
case "allowusers", "allowgroups", "denyusers", "denygroups":
for _, w := range strings.Fields(rest) {
if len(p.lists[key]) < 30 {
p.lists[key] = append(p.lists[key], truncate(unquote(w), 64))
}
}
default:
if _, set := p.vals[key]; !set && rest != "" {
if key == "authenticationmethods" {
p.vals[key] = strings.ToLower(rest)
} else {
p.vals[key] = strings.ToLower(unquote(firstField(rest)))
}
}
}
}
}
// splitSSHDLine: "Keyword value", "Keyword=value" or "Keyword = value".
func splitSSHDLine(line string) (string, string) {
i := strings.IndexAny(line, " \t=")
if i < 0 {
return strings.ToLower(line), ""
}
key := strings.ToLower(line[:i])
rest := strings.TrimLeft(line[i:], " \t")
rest = strings.TrimSpace(strings.TrimPrefix(rest, "="))
return key, rest
}
func firstField(s string) string {
if f := strings.Fields(s); len(f) > 0 {
return f[0]
}
return ""
}
// ---------------------------------------------------------------------------
// host firewall
// ---------------------------------------------------------------------------
type FirewallOut struct {
Found []string `json:"found"` // ufw | firewalld | nftables | iptables | csf | shorewall
Known bool `json:"known"` // false: the agent could not tell (no unit list and nothing found)
}
func unitActive(units map[string]UnitInfo, name string) bool {
u, ok := units[name]
return ok && u.ActiveState == "active"
}
var dropRe = regexp.MustCompile(`(?i)\b(drop|reject)\b`)
// rulesFileFilters: a ruleset file that drops or rejects something (Debian ships an
// /etc/nftables.conf that accepts everything). Unreadable or with includes: assumed to filter.
func rulesFileFilters(path string) bool {
b, err := readSmall(path, 256*1024)
if err != nil {
return !errors.Is(err, fs.ErrNotExist)
}
var kept []string
for _, l := range strings.Split(string(b), "\n") {
l = strings.TrimSpace(l)
if l == "" || l[0] == '#' {
continue
}
kept = append(kept, l)
}
s := strings.Join(kept, "\n")
return dropRe.MatchString(s) || strings.Contains(s, "include")
}
// DetectFirewall looks for an active host firewall. Cloud firewalls (security groups) are outside
// the machine and cannot be seen from it.
func DetectFirewall(root string, units map[string]UnitInfo, procNames map[string]bool) FirewallOut {
var found []string
if b, err := readSmall(filepath.Join(root, "etc/ufw/ufw.conf"), 64*1024); err == nil {
for _, l := range strings.Split(string(b), "\n") {
if strings.EqualFold(strings.ReplaceAll(strings.TrimSpace(l), " ", ""), "ENABLED=yes") {
found = append(found, "ufw")
break
}
}
}
if unitActive(units, "firewalld.service") || procNames["firewalld"] {
found = append(found, "firewalld")
}
if unitActive(units, "nftables.service") &&
(rulesFileFilters(filepath.Join(root, "etc/nftables.conf")) || rulesFileFilters(filepath.Join(root, "etc/sysconfig/nftables.conf"))) {
found = append(found, "nftables")
}
if unitActive(units, "netfilter-persistent.service") || unitActive(units, "iptables.service") || unitActive(units, "ip6tables.service") {
found = append(found, "iptables")
}
if unitActive(units, "csf.service") || procNames["lfd"] {
found = append(found, "csf")
}
if unitActive(units, "shorewall.service") {
found = append(found, "shorewall")
}
if found == nil {
found = []string{}
}
return FirewallOut{Found: found, Known: units != nil || len(found) > 0}
}
// ---------------------------------------------------------------------------
// automatic security updates
// ---------------------------------------------------------------------------
type AutoUpdOut struct {
Kind string `json:"kind"` // unattended-upgrades | dnf-automatic | "" (a system the agent does not know)
Enabled bool `json:"enabled"`
Detail string `json:"detail,omitempty"` // not_installed | disabled | timer_inactive
}
var aptUnattendedRe = regexp.MustCompile(`(?i)^\s*APT::Periodic::Unattended-Upgrade\s+"([^"]*)"`)
func DetectAutoUpdates(root string, units map[string]UnitInfo) AutoUpdOut {
aptDir := filepath.Join(root, "etc/apt/apt.conf.d")
if ents, err := os.ReadDir(aptDir); err == nil {
out := AutoUpdOut{Kind: "unattended-upgrades"}
installed := exists(filepath.Join(root, "usr/bin/unattended-upgrade")) || exists(filepath.Join(root, "usr/bin/unattended-upgrades"))
value := ""
names := make([]string, 0, len(ents))
for _, e := range ents {
names = append(names, e.Name())
}
sort.Strings(names) // apt reads them in this order; a later file overrides an earlier one
for _, n := range names {
b, err := readSmall(filepath.Join(aptDir, n), 64*1024)
if err != nil {
continue
}
for _, l := range strings.Split(string(b), "\n") {
if m := aptUnattendedRe.FindStringSubmatch(l); m != nil {
value = m[1]
}
}
}
switch {
case !installed:
out.Detail = "not_installed"
case value == "" || value == "0":
out.Detail = "disabled"
case units != nil && !unitActive(units, "apt-daily-upgrade.timer"):
out.Detail = "timer_inactive"
default:
out.Enabled = true
}
return out
}
if exists(filepath.Join(root, "etc/dnf")) || exists(filepath.Join(root, "etc/yum.conf")) {
out := AutoUpdOut{Kind: "dnf-automatic"}
apply := false
if b, err := readSmall(filepath.Join(root, "etc/dnf/automatic.conf"), 64*1024); err == nil {
for _, l := range strings.Split(string(b), "\n") {
k, v, ok := strings.Cut(strings.TrimSpace(l), "=")
if ok && strings.TrimSpace(k) == "apply_updates" {
v = strings.ToLower(strings.TrimSpace(v))
apply = v == "yes" || v == "true" || v == "1"
}
}
}
timer := unitActive(units, "dnf-automatic.timer") || unitActive(units, "dnf5-automatic.timer")
switch {
case unitActive(units, "dnf-automatic-install.timer"):
out.Enabled = true
case units == nil && apply:
out.Enabled = true // no unit list to check the timer against
case !apply:
out.Detail = "disabled"
case !timer:
out.Detail = "timer_inactive"
default:
out.Enabled = true
}
return out
}
return AutoUpdOut{}
}
func exists(p string) bool {
_, err := os.Stat(p)
return err == nil
}
// ---------------------------------------------------------------------------
// accounts: extra UID 0, sudoers NOPASSWD
// ---------------------------------------------------------------------------
func ExtraUID0(users []PasswdEntry) []string {
out := []string{}
for _, u := range users {
if u.UID == 0 && u.Name != "root" && len(out) < 20 {
out = append(out, truncate(u.Name, 64))
}
}
sort.Strings(out)
return dedupSorted(out)
}
type SudoOut struct {
Readable bool `json:"readable"`
Count int `json:"count"` // rules with NOPASSWD
Names []string `json:"names"` // the users and %groups they name
}
// ReadSudoers counts the NOPASSWD rules of /etc/sudoers and /etc/sudoers.d. Only the first word of
// a rule (who it applies to) is kept; the commands are not.
func ReadSudoers(root string) SudoOut {
out := SudoOut{Names: []string{}}
main := filepath.Join(root, "etc/sudoers")
files := []string{main}
if ents, err := os.ReadDir(filepath.Join(root, "etc/sudoers.d")); err == nil {
for _, e := range ents {
n := e.Name()
if strings.Contains(n, ".") || strings.HasSuffix(n, "~") || !e.Type().IsRegular() { // sudo skips these too
continue
}
files = append(files, filepath.Join(root, "etc/sudoers.d", n))
}
} else if errors.Is(err, fs.ErrPermission) {
return out
}
readAny := false
for _, f := range files {
b, err := readSmall(f, 256*1024)
if err != nil {
if errors.Is(err, fs.ErrPermission) {
return SudoOut{Names: []string{}}
}
continue
}
readAny = true
text := strings.ReplaceAll(string(b), "\\\n", " ")
for _, l := range strings.Split(text, "\n") {
l = strings.TrimSpace(l)
if l == "" || l[0] == '#' || l[0] == '@' || !strings.Contains(l, "NOPASSWD") {
continue
}
who := firstField(l)
switch who {
case "Defaults", "Cmnd_Alias", "User_Alias", "Runas_Alias", "Host_Alias":
continue
}
if strings.HasPrefix(who, "Defaults") {
continue
}
out.Count++
if len(out.Names) < 20 {
out.Names = append(out.Names, truncate(who, 64))
}
}
}
out.Readable = readAny || !exists(main)
sort.Strings(out.Names)
out.Names = dedupSorted(out.Names)
return out
}
// ---------------------------------------------------------------------------
// databases reachable from the network
// ---------------------------------------------------------------------------
type PublicPort struct {
Proto string `json:"proto"`
Port int `json:"port"`
Addr string `json:"addr"`
Service string `json:"service"`
Process string `json:"process,omitempty"`
}
var riskyPorts = map[int]string{
3306: "MySQL/MariaDB", 33060: "MySQL X", 5432: "PostgreSQL", 6379: "Redis", 27017: "MongoDB", 27018: "MongoDB",
9200: "Elasticsearch/OpenSearch", 9300: "Elasticsearch transport", 11211: "Memcached", 5984: "CouchDB", 9042: "Cassandra",
8086: "InfluxDB", 2379: "etcd", 1433: "SQL Server", 1521: "Oracle", 28015: "RethinkDB", 2375: "Docker API (no TLS)",
}
var cgnat = &net.IPNet{IP: net.IPv4(100, 64, 0, 0), Mask: net.CIDRMask(10, 32)}
// exposedAddr: an address other machines on the internet can reach: all interfaces, or a
// public address (not loopback, private, link-local or carrier-grade NAT).
func exposedAddr(addr string) bool {
ip := net.ParseIP(addr)
if ip == nil || ip.IsLoopback() {
return false
}
if ip.IsUnspecified() {
return true
}
return !(ip.IsPrivate() || ip.IsLinkLocalUnicast() || cgnat.Contains(ip))
}
// PublicDBPorts: database and similar ports listening on an address reachable from the network.
// A firewall may still block them; the agent cannot tell.
func PublicDBPorts(ports []PortOut) []PublicPort {
out := []PublicPort{}
seen := map[string]bool{}
for _, p := range ports {
svc, risky := riskyPorts[p.Port]
if !risky || p.Local || !exposedAddr(p.Addr) || (p.Proto == "udp" && p.Port != 11211) {
continue
}
k := p.Proto + "/" + strconv.Itoa(p.Port)
if seen[k] {
continue
}
seen[k] = true
out = append(out, PublicPort{Proto: p.Proto, Port: p.Port, Addr: p.Addr, Service: svc, Process: p.Name})
}
sort.Slice(out, func(i, j int) bool {
return out[i].Proto+strconv.Itoa(out[i].Port) < out[j].Proto+strconv.Itoa(out[j].Port)
})
return out
}
// ---------------------------------------------------------------------------
// pending reboot, waiting updates
// ---------------------------------------------------------------------------
type RebootOut struct {
Required bool `json:"required"`
Reason string `json:"reason,omitempty"` // reboot-required (Debian, Ubuntu) | new-kernel
Since int64 `json:"since,omitempty"`
Pkgs []string `json:"pkgs,omitempty"` // from /run/reboot-required.pkgs
Kernel string `json:"kernel,omitempty"` // the newer kernel installed after boot
}
// CheckReboot: Debian and Ubuntu flag a needed reboot in /run/reboot-required. Elsewhere (and as a
// second opinion) a kernel newer than the running one, installed after the last boot, means a reboot.
func CheckReboot(root, running string, bootTime int64) RebootOut {
for _, p := range []string{"run/reboot-required", "var/run/reboot-required"} {
fi, err := os.Stat(filepath.Join(root, p))
if err != nil {
continue
}
out := RebootOut{Required: true, Reason: "reboot-required", Since: fi.ModTime().Unix()}
if b, err := readSmall(filepath.Join(root, p+".pkgs"), 64*1024); err == nil {
for _, l := range strings.Split(string(b), "\n") {
if l = strings.TrimSpace(l); l != "" && len(out.Pkgs) < 20 {
out.Pkgs = append(out.Pkgs, truncate(l, 80))
}
}
sort.Strings(out.Pkgs)
out.Pkgs = dedupSorted(out.Pkgs)
}
return out
}
if running == "" {
return RebootOut{}
}
for _, dir := range []string{"lib/modules", "usr/lib/modules"} {
ents, err := os.ReadDir(filepath.Join(root, dir))
if err != nil || len(ents) == 0 {
continue
}
newest := ""
var newestAt int64
for _, e := range ents {
if !e.IsDir() {
continue
}
if newest == "" || kernelCmp(e.Name(), newest) > 0 {
if fi, err := e.Info(); err == nil {
newest, newestAt = e.Name(), fi.ModTime().Unix()
}
}
}
if newest != "" && newest != running && kernelCmp(newest, running) > 0 && (bootTime == 0 || newestAt > bootTime+60) {
return RebootOut{Required: true, Reason: "new-kernel", Since: newestAt, Kernel: truncate(newest, 80)}
}
return RebootOut{}
}
return RebootOut{}
}
var digitsRe = regexp.MustCompile(`\d+`)
// kernelCmp compares two kernel release strings by their numbers in order (6.8.0-47 > 6.8.0-45).
func kernelCmp(a, b string) int {
x, y := digitsRe.FindAllString(a, 12), digitsRe.FindAllString(b, 12)
for i := 0; i < len(x) && i < len(y); i++ {
p, _ := strconv.ParseUint(x[i], 10, 64)
q, _ := strconv.ParseUint(y[i], 10, 64)
if p != q {
if p > q {
return 1
}
return -1
}
}
switch {
case len(x) > len(y):
return 1
case len(x) < len(y):
return -1
}
return strings.Compare(a, b)
}
type UpdatesOut struct {
Source string `json:"source"` // update-notifier (Ubuntu)
Total int `json:"total"`
Security int `json:"security"`
At int64 `json:"at"` // when the counts were computed (the file's time)
}
var (
updTotalRe = regexp.MustCompile(`(?m)^\s*(\d+)\s+(?:updates?|packages?)\s+can\s+be\s+(?:applied\s+immediately|updated)`)
updSecRe = regexp.MustCompile(`(?m)^\s*(\d+)\s+(?:of\s+these\s+)?updates?\s+(?:is\s+a\s+|are\s+|is\s+)?(?:standard\s+)?security\s+updates?`)
)
// ParseUpdatesAvailable reads Ubuntu's /var/lib/update-notifier/updates-available (the text of the
// login banner, refreshed by apt). ok=false when the text is in a form it does not know.
func ParseUpdatesAvailable(b []byte) (total, security int, ok bool) {
s := string(b)
if strings.TrimSpace(s) == "" {
return 0, 0, true // older releases leave it empty when nothing is pending
}
m := updTotalRe.FindStringSubmatch(s)
if m == nil {
return 0, 0, false
}
total, _ = strconv.Atoi(m[1])
if m := updSecRe.FindStringSubmatch(s); m != nil {
security, _ = strconv.Atoi(m[1])
}
return total, security, true
}
func ReadUpdates(root string) *UpdatesOut {
p := filepath.Join(root, "var/lib/update-notifier/updates-available")
fi, err := os.Stat(p)
if err != nil {
return nil
}
b, err := readSmall(p, 64*1024)
if err != nil {
return nil
}
total, sec, ok := ParseUpdatesAvailable(b)
if !ok {
return nil
}
return &UpdatesOut{Source: "update-notifier", Total: total, Security: sec, At: fi.ModTime().Unix()}
}
// ---------------------------------------------------------------------------
// time synchronisation
// ---------------------------------------------------------------------------
type TimeOut struct {
Daemon string `json:"daemon"` // chronyd | ntpd | systemd-timesyncd | ptp4l | "" (none found)
Synced *bool `json:"synced"` // true when the daemon says so in a file the agent can read; null: cannot tell
Jumps int `json:"jumps"` // wall clock steps of more than 30 s seen since the agent started
}
// TimeSync names the time daemon that runs (from the process list, or the unit list) and, for
// systemd-timesyncd, whether it has synchronised (/run/systemd/timesync/synchronized).
func TimeSync(root string, units map[string]UnitInfo, procNames map[string]bool) TimeOut {
var out TimeOut
switch {
case procNames["chronyd"] || unitActive(units, "chrony.service") || unitActive(units, "chronyd.service"):
out.Daemon = "chronyd"
case procNames["ntpd"] || unitActive(units, "ntp.service") || unitActive(units, "ntpd.service") || unitActive(units, "ntpsec.service"):
out.Daemon = "ntpd"
case procNames["systemd-timesyn"] || unitActive(units, "systemd-timesyncd.service"):
out.Daemon = "systemd-timesyncd"
if exists(filepath.Join(root, "run/systemd/timesync/synchronized")) {
t := true
out.Synced = &t
}
case procNames["ptp4l"] || procNames["phc2sys"]:
out.Daemon = "ptp4l"
}
return out
}
accounts.go337 lines
package main
// Accounts, SSH key counts, fail2ban and package-manager activity, for the hourly settings check.
// Read from /etc/passwd, /etc/group (never /etc/shadow), ~/.ssh/authorized_keys (counted, the
// keys themselves stay on the server), fail2ban's configuration files, and the time stamps apt and
// dnf leave behind.
import (
"bufio"
"bytes"
"os"
"path/filepath"
"sort"
"strconv"
"strings"
)
// UserOut is one account that can log in, has UID 0, or may use sudo.
type UserOut struct {
Name string `json:"name"`
UID int `json:"uid"`
Shell string `json:"shell"`
Login bool `json:"login"`
Sudo bool `json:"sudo"` // member of sudo, wheel or admin (or root)
Groups []string `json:"groups,omitempty"` // its groups (at most 12)
NoPasswd bool `json:"nopasswd,omitempty"` // a sudoers NOPASSWD rule names it (or one of its groups)
LastLogin int64 `json:"last_login,omitempty"`
Keys int `json:"keys"` // authorized_keys entries (-1: unreadable)
}
// GroupEntry from /etc/group.
type GroupEntry struct {
Name string
GID int
Members []string
}
func ParseGroup(b []byte) []GroupEntry {
var out []GroupEntry
sc := bufio.NewScanner(bytes.NewReader(b))
sc.Buffer(make([]byte, 64*1024), 1024*1024)
for sc.Scan() {
f := strings.Split(sc.Text(), ":")
if len(f) < 4 || strings.HasPrefix(f[0], "#") {
continue
}
gid, err := strconv.Atoi(f[2])
if err != nil {
continue
}
g := GroupEntry{Name: f[0], GID: gid}
for _, m := range strings.Split(f[3], ",") {
if m = strings.TrimSpace(m); m != "" {
g.Members = append(g.Members, m)
}
}
out = append(out, g)
}
return out
}
// passwdGIDs: /etc/passwd's primary group of each user (ParsePasswd drops it).
func passwdGIDs(b []byte) map[string]int {
out := map[string]int{}
sc := bufio.NewScanner(bytes.NewReader(b))
for sc.Scan() {
f := strings.Split(sc.Text(), ":")
if len(f) < 7 {
continue
}
if gid, err := strconv.Atoi(f[3]); err == nil {
out[f[0]] = gid
}
}
return out
}
var sudoGroups = map[string]bool{"sudo": true, "wheel": true, "admin": true}
// BuildUsers lists the accounts worth showing: login shell, UID 0, or a sudo group.
func BuildUsers(root string, users []PasswdEntry, sudo SudoOut, last map[string]int64, keys map[string]KeyCount, keysRead bool) []UserOut {
var groups []GroupEntry
if b, err := readSmall(filepath.Join(root, "etc/group"), 4<<20); err == nil {
groups = ParseGroup(b)
}
var gids map[string]int
if b, err := readSmall(filepath.Join(root, "etc/passwd"), 4<<20); err == nil {
gids = passwdGIDs(b)
}
byGID := map[int]string{}
member := map[string][]string{}
for _, g := range groups {
byGID[g.GID] = g.Name
for _, m := range g.Members {
member[m] = append(member[m], g.Name)
}
}
nopw := map[string]bool{}
for _, n := range sudo.Names {
nopw[n] = true
}
var out []UserOut
for _, u := range users {
gs := append([]string(nil), member[u.Name]...)
if g, ok := byGID[gids[u.Name]]; ok {
gs = append(gs, g)
}
sort.Strings(gs)
gs = dedupSorted(gs)
inSudo := u.UID == 0
np := nopw[u.Name]
for _, g := range gs {
if sudoGroups[g] {
inSudo = true
}
if nopw["%"+g] {
np = true
}
}
login := LoginShell(u.Shell)
if !login && u.UID != 0 && !inSudo {
continue
}
o := UserOut{Name: truncate(u.Name, 32), UID: u.UID, Shell: truncate(u.Shell, 64), Login: login, Sudo: inSudo, NoPasswd: np,
LastLogin: last[u.Name], Keys: -1}
if len(gs) > 12 {
gs = gs[:12]
}
o.Groups = gs
if keysRead {
o.Keys = keys[u.Name].Count
}
out = append(out, o)
if len(out) >= 50 {
break
}
}
return out
}
// CountSSHKeys counts the keys in authorized_keys(2) of root and of every account with a login
// shell, by type. ok is false when a file could not be read (permissions).
func CountSSHKeys(root string, users []PasswdEntry) (map[string]KeyCount, bool) {
out := map[string]KeyCount{}
ok := true
seen := map[string]bool{}
for _, u := range users {
if u.UID != 0 && !LoginShell(u.Shell) {
continue
}
if u.Home == "" || u.Home == "/" || seen[u.Home] {
continue
}
seen[u.Home] = true
kc := KeyCount{User: truncate(u.Name, 32)}
for _, name := range []string{"authorized_keys", "authorized_keys2"} {
b, err := readSmall(filepath.Join(root, u.Home, ".ssh", name), 512*1024)
if err != nil {
if os.IsPermission(err) {
ok = false
}
continue
}
sc := bufio.NewScanner(bytes.NewReader(b))
sc.Buffer(make([]byte, 64*1024), 1024*1024)
for sc.Scan() {
line := strings.TrimSpace(sc.Text())
if line == "" || line[0] == '#' {
continue
}
for _, f := range strings.Fields(line) {
if keyType(f) {
kc.Count++
if !hasStr(kc.Types, f) && len(kc.Types) < 8 {
kc.Types = append(kc.Types, f)
}
break
}
}
}
}
if kc.Count > 0 {
sort.Strings(kc.Types)
out[u.Name] = kc
}
}
return out, ok
}
func hasStr(xs []string, x string) bool {
for _, y := range xs {
if y == x {
return true
}
}
return false
}
func keyList(m map[string]KeyCount) []KeyCount {
out := make([]KeyCount, 0, len(m))
for _, k := range m {
out = append(out, k)
}
sort.Slice(out, func(i, j int) bool { return out[i].User < out[j].User })
if len(out) > 50 {
out = out[:50]
}
return out
}
// ---------------------------------------------------------------------------
// fail2ban
// ---------------------------------------------------------------------------
type Fail2banOut struct {
Installed bool `json:"installed"`
Running bool `json:"running"`
Jails []string `json:"jails"` // enabled jails, from the configuration files
Banned *int `json:"banned"` // not read (it lives in fail2ban's database): null
}
// ReadFail2ban: installed (its configuration or server exists), running (its process or socket),
// and the jails its configuration enables, read the way fail2ban does: jail.conf, jail.d/*.conf,
// jail.local, jail.d/*.local, later files overriding earlier ones, [DEFAULT] enabled applying to
// jails that do not set it.
func ReadFail2ban(root string, procNames map[string]bool) *Fail2banOut {
dir := filepath.Join(root, "etc/fail2ban")
installed := exists(filepath.Join(dir, "jail.conf")) || exists(filepath.Join(root, "usr/bin/fail2ban-server"))
if !installed {
return nil
}
out := &Fail2banOut{Installed: true, Jails: []string{}}
out.Running = procNames["fail2ban-server"] || procNames["fail2ban-serve"] ||
exists(filepath.Join(root, "run/fail2ban/fail2ban.sock")) || exists(filepath.Join(root, "var/run/fail2ban/fail2ban.sock"))
files := []string{filepath.Join(dir, "jail.conf")}
globD := func(ext string) []string {
m, _ := filepath.Glob(filepath.Join(dir, "jail.d", "*"+ext))
sort.Strings(m)
return m
}
files = append(files, globD(".conf")...)
files = append(files, filepath.Join(dir, "jail.local"))
files = append(files, globD(".local")...)
enabled := map[string]string{} // jail -> "true"/"false"/"" (unset)
defEnabled := "false"
for _, f := range files {
b, err := readSmall(f, 1<<20)
if err != nil {
continue
}
sec := ""
for _, l := range strings.Split(string(b), "\n") {
l = strings.TrimSpace(l)
if l == "" || l[0] == '#' || l[0] == ';' {
continue
}
if l[0] == '[' && strings.HasSuffix(l, "]") {
sec = strings.TrimSpace(l[1 : len(l)-1])
if sec != "DEFAULT" && sec != "INCLUDES" && sec != "Definition" {
if _, ok := enabled[sec]; !ok && len(enabled) < 500 {
enabled[sec] = ""
}
}
continue
}
k, v, ok := strings.Cut(l, "=")
if !ok || strings.TrimSpace(strings.ToLower(k)) != "enabled" {
continue
}
v = strings.ToLower(strings.TrimSpace(v))
on := "false"
if v == "true" || v == "yes" || v == "1" || v == "on" {
on = "true"
}
if sec == "DEFAULT" {
defEnabled = on
} else if sec != "" && sec != "INCLUDES" {
enabled[sec] = on
}
}
}
for j, e := range enabled {
if e == "true" || (e == "" && defEnabled == "true") {
out.Jails = append(out.Jails, truncate(j, 64))
}
}
sort.Strings(out.Jails)
if len(out.Jails) > 30 {
out.Jails = out.Jails[:30]
}
return out
}
// ---------------------------------------------------------------------------
// package manager activity
// ---------------------------------------------------------------------------
type PkgOut struct {
Manager string `json:"manager"` // apt | dnf | yum | zypper | apk
LastUpdate int64 `json:"last_update,omitempty"` // the package lists were last refreshed
LastUpgrade int64 `json:"last_upgrade,omitempty"` // an automatic upgrade last ran
LastChange int64 `json:"last_change,omitempty"` // packages were last installed / upgraded / removed
}
func mtimeOf(paths ...string) int64 {
var best int64
for _, p := range paths {
if fi, err := os.Stat(p); err == nil && fi.ModTime().Unix() > best {
best = fi.ModTime().Unix()
}
}
return best
}
// ReadPkg: apt's periodic stamps (/var/lib/apt/periodic), unattended-upgrades' log and apt's
// history; dnf's makecache stamp and history database; only their modification times.
func ReadPkg(root string) *PkgOut {
j := func(p string) string { return filepath.Join(root, p) }
switch {
case exists(j("var/lib/dpkg/status")):
return &PkgOut{Manager: "apt",
LastUpdate: mtimeOf(j("var/lib/apt/periodic/update-success-stamp"), j("var/lib/apt/periodic/update-stamp")),
LastUpgrade: mtimeOf(j("var/lib/apt/periodic/upgrade-stamp"), j("var/log/unattended-upgrades/unattended-upgrades.log")),
LastChange: mtimeOf(j("var/log/apt/history.log"), j("var/log/dpkg.log"))}
case exists(j("etc/dnf")) || exists(j("usr/bin/dnf")):
return &PkgOut{Manager: "dnf", LastUpdate: mtimeOf(j("var/cache/dnf/last_makecache"), j("var/cache/dnf/expired_repos.json")),
LastChange: mtimeOf(j("var/lib/dnf/history.sqlite"), j("var/log/dnf.rpm.log"))}
case exists(j("etc/yum.conf")):
return &PkgOut{Manager: "yum", LastChange: mtimeOf(j("var/log/yum.log"))}
case exists(j("etc/zypp")):
return &PkgOut{Manager: "zypper", LastChange: mtimeOf(j("var/log/zypp/history"))}
case exists(j("etc/apk")):
return &PkgOut{Manager: "apk", LastChange: mtimeOf(j("lib/apk/db/installed"))}
}
return nil
}
firewall.go364 lines
package main
// The host firewall in more detail than "found / not found": which backend is in charge, its
// default inbound policy and the inbound ports its saved rules open, as far as its files tell.
// Best effort and read-only: ufw's user.rules (the "### tuple ###" lines ufw itself writes),
// firewalld's zone files, /etc/nftables.conf and iptables-save files. Rules added at run time (by
// Docker, for example) are not in these files; Docker-published ports bypass ufw.
import (
"bufio"
"bytes"
"path/filepath"
"regexp"
"sort"
"strconv"
"strings"
)
type FWOut struct {
Backend string `json:"backend"` // ufw | firewalld | nftables | iptables | none | unknown
Active bool `json:"active"`
DefaultIn string `json:"default_in"` // drop | reject | accept | unknown
Rules int `json:"rules"`
Allowed []int `json:"allowed"` // inbound ports open to any address
Limited []int `json:"limited"` // ufw "limit": open, rate limited
Restricted []int `json:"restricted"` // allowed from some addresses only
Denied []int `json:"denied"`
AllPorts bool `json:"all_ports,omitempty"` // a rule accepts every port
}
type portSet map[int]bool
func (s portSet) list() []int {
out := make([]int, 0, len(s))
for p := range s {
out = append(out, p)
}
sort.Ints(out)
if len(out) > 100 {
out = out[:100]
}
return out
}
// addPorts parses "22", "80,443", "8000:8010" / "8000-8010" (ranges of at most 64 ports listed in
// full, larger ones by their first port).
func (s portSet) addPorts(spec string) {
for _, part := range strings.Split(spec, ",") {
part = strings.TrimSpace(strings.Trim(part, "{} "))
if part == "" {
continue
}
sep := strings.IndexAny(part, ":-")
if sep > 0 {
a, e1 := strconv.Atoi(part[:sep])
b, e2 := strconv.Atoi(part[sep+1:])
if e1 == nil && e2 == nil && a > 0 && b >= a && b < 65536 {
if b-a > 64 {
b = a
}
for p := a; p <= b && len(s) < 1000; p++ {
s[p] = true
}
}
continue
}
if p, err := strconv.Atoi(part); err == nil && p > 0 && p < 65536 && len(s) < 1000 {
s[p] = true
}
}
}
type fwAcc struct {
allowed, limited, restricted, denied portSet
rules int
all bool
}
func newFwAcc() *fwAcc {
return &fwAcc{allowed: portSet{}, limited: portSet{}, restricted: portSet{}, denied: portSet{}}
}
func (a *fwAcc) out(backend string, active bool, def string) *FWOut {
return &FWOut{Backend: backend, Active: active, DefaultIn: def, Rules: a.rules, Allowed: a.allowed.list(), Limited: a.limited.list(),
Restricted: a.restricted.list(), Denied: a.denied.list(), AllPorts: a.all}
}
func anyAddr(src string) bool {
return src == "" || src == "any" || src == "0.0.0.0/0" || src == "::/0" || src == "anywhere"
}
// ParseUFWRules reads ufw's user.rules / user6.rules: "### tuple ### allow tcp 22 0.0.0.0/0 any 0.0.0.0/0 in"
// (action proto dport dst sport src [dapp sapp] direction).
func ParseUFWRules(b []byte, a *fwAcc) {
sc := bufio.NewScanner(bytes.NewReader(b))
sc.Buffer(make([]byte, 64*1024), 1024*1024)
for sc.Scan() {
l := strings.TrimSpace(sc.Text())
if !strings.HasPrefix(l, "### tuple ### ") {
continue
}
f := strings.Fields(strings.TrimPrefix(l, "### tuple ### "))
if len(f) < 7 || f[len(f)-1] != "in" || strings.HasPrefix(f[0], "route:") {
continue
}
a.rules++
action := strings.TrimSuffix(f[0], "_log-all")
action = strings.TrimSuffix(action, "_log")
dport, src := f[2], f[5]
target := a.allowed
switch action {
case "allow":
if !anyAddr(src) {
target = a.restricted
}
case "limit":
target = a.limited
if !anyAddr(src) {
target = a.restricted
}
case "deny", "reject":
target = a.denied
default:
continue
}
if dport == "any" {
if action == "allow" && anyAddr(src) {
a.all = true
}
continue
}
target.addPorts(dport)
}
}
var ufwDefaultRe = regexp.MustCompile(`(?m)^\s*DEFAULT_INPUT_POLICY\s*=\s*"?([A-Za-z-]+)"?`)
func policyWord(s string) string {
switch strings.ToLower(strings.TrimSpace(s)) {
case "drop", "deny":
return "drop"
case "reject":
return "reject"
case "accept", "allow":
return "accept"
}
return "unknown"
}
// firewalld services: the ports of the predefined services most often enabled.
var firewalldServices = map[string]string{
"ssh": "22", "http": "80", "https": "443", "http3": "443", "dns": "53", "smtp": "25", "smtps": "465", "smtp-submission": "587",
"imap": "143", "imaps": "993", "pop3": "110", "pop3s": "995", "ftp": "21", "mysql": "3306", "postgresql": "5432", "redis": "6379",
"mongodb": "27017", "elasticsearch": "9200", "cockpit": "9090", "prometheus": "9090", "grafana": "3000", "vnc-server": "5900-5903",
"openvpn": "1194", "wireguard": "51820", "ldap": "389", "ldaps": "636", "nfs": "2049", "rpc-bind": "111", "samba": "139,445",
"docker-registry": "5000", "jenkins": "8080", "zabbix-agent": "10050", "zabbix-server": "10051", "kubelet": "10250",
"kube-apiserver": "6443", "ntp": "123", "dhcpv6-client": "546", "mdns": "5353", "samba-client": "137,138",
}
var (
xmlServiceRe = regexp.MustCompile(`<service\s+name="([^"]+)"`)
xmlPortRe = regexp.MustCompile(`<port\s+[^>]*port="([^"]+)"`)
xmlTargetRe = regexp.MustCompile(`<zone[^>]*target="([^"]+)"`)
xmlSourceRe = regexp.MustCompile(`<source\s+address=`)
fwdZoneRe = regexp.MustCompile(`(?m)^\s*DefaultZone\s*=\s*(\S+)`)
)
// ParseFirewalldZone reads one zone file: its services and ports are open to the zone's sources
// (every address for the default zone without <source>).
func ParseFirewalldZone(b []byte, a *fwAcc) string {
s := string(b)
target := a.allowed
if xmlSourceRe.MatchString(s) {
target = a.restricted
}
for _, m := range xmlServiceRe.FindAllStringSubmatch(s, 100) {
a.rules++
if p, ok := firewalldServices[m[1]]; ok {
target.addPorts(p)
}
}
for _, m := range xmlPortRe.FindAllStringSubmatch(s, 100) {
a.rules++
target.addPorts(m[1])
}
if m := xmlTargetRe.FindStringSubmatch(s); m != nil {
switch strings.ToUpper(m[1]) {
case "ACCEPT":
return "accept"
case "DROP":
return "drop"
case "%%REJECT%%", "REJECT":
return "reject"
}
}
return "reject" // firewalld's default target rejects what the zone does not allow
}
var (
nftChainRe = regexp.MustCompile(`(?s)chain\s+(\S+)\s*\{(.*?)\n\s*\}`)
nftPolicyRe = regexp.MustCompile(`policy\s+(accept|drop|reject)`)
nftDportRe = regexp.MustCompile(`(tcp|udp|th)\s+dport\s+(\{[^}]*\}|\S+)`)
)
// ParseNftables reads the input chain(s) of an nftables ruleset file: the policy and the ports
// of rules that accept ("ip saddr …" makes a rule restricted).
func ParseNftables(b []byte, a *fwAcc) string {
policy := "unknown"
found := false
for _, m := range nftChainRe.FindAllStringSubmatch(string(b), 50) {
body := m[2]
if !strings.Contains(body, "hook input") {
continue
}
found = true
if p := nftPolicyRe.FindStringSubmatch(body); p != nil {
policy = p[1]
} else if policy == "unknown" {
policy = "accept"
}
for _, l := range strings.Split(body, "\n") {
l = strings.TrimSpace(l)
if l == "" || l[0] == '#' {
continue
}
if strings.Contains(l, "dport") {
a.rules++
}
dm := nftDportRe.FindStringSubmatch(l)
if dm == nil {
continue
}
ports := strings.Trim(dm[2], "{}")
switch {
case strings.HasSuffix(l, "accept") || strings.Contains(l, " accept"):
if strings.Contains(l, "saddr") {
a.restricted.addPorts(ports)
} else {
a.allowed.addPorts(ports)
}
case strings.Contains(l, " drop") || strings.Contains(l, " reject"):
a.denied.addPorts(ports)
}
}
}
if !found {
return "unknown"
}
return policyWord(policy)
}
var iptPolicyRe = regexp.MustCompile(`(?m)^:INPUT\s+(ACCEPT|DROP|REJECT)`)
// ParseIptablesSave reads an iptables-save file (/etc/iptables/rules.v4, /etc/sysconfig/iptables).
func ParseIptablesSave(b []byte, a *fwAcc) string {
policy := "unknown"
if m := iptPolicyRe.FindSubmatch(b); m != nil {
policy = policyWord(string(m[1]))
}
for _, l := range strings.Split(string(b), "\n") {
if !strings.HasPrefix(l, "-A INPUT") {
continue
}
a.rules++
f := strings.Fields(l)
var ports, src, jump string
for i := 0; i+1 < len(f); i++ {
switch f[i] {
case "--dport", "--dports", "--destination-port":
ports = f[i+1]
case "-s", "--source":
src = f[i+1]
case "-j", "--jump":
jump = f[i+1]
}
}
if ports == "" {
continue
}
switch jump {
case "ACCEPT":
if anyAddr(src) {
a.allowed.addPorts(ports)
} else {
a.restricted.addPorts(ports)
}
case "DROP", "REJECT":
a.denied.addPorts(ports)
}
}
return policy
}
// ReadFirewallDetail: the backend in charge (the first found of ufw, firewalld, nftables,
// iptables-persistent), with what its files say.
func ReadFirewallDetail(root string, units map[string]UnitInfo, procNames map[string]bool, found FirewallOut) *FWOut {
j := func(p string) string { return filepath.Join(root, p) }
has := func(n string) bool {
for _, f := range found.Found {
if f == n {
return true
}
}
return false
}
switch {
case has("ufw"):
a := newFwAcc()
for _, f := range []string{"etc/ufw/user.rules", "etc/ufw/user6.rules", "lib/ufw/user.rules", "lib/ufw/user6.rules"} {
if b, err := readSmall(j(f), 1<<20); err == nil {
ParseUFWRules(b, a)
}
}
def := "unknown"
if b, err := readSmall(j("etc/default/ufw"), 64*1024); err == nil {
if m := ufwDefaultRe.FindSubmatch(b); m != nil {
def = policyWord(string(m[1]))
}
}
return a.out("ufw", true, def)
case has("firewalld"):
a := newFwAcc()
zone := "public"
if b, err := readSmall(j("etc/firewalld/firewalld.conf"), 64*1024); err == nil {
if m := fwdZoneRe.FindSubmatch(b); m != nil {
zone = string(m[1])
}
}
def := "reject"
for _, f := range []string{"etc/firewalld/zones/" + zone + ".xml", "usr/lib/firewalld/zones/" + zone + ".xml"} {
if b, err := readSmall(j(f), 1<<20); err == nil {
def = ParseFirewalldZone(b, a)
break
}
}
return a.out("firewalld", true, def)
case has("nftables"):
a := newFwAcc()
def := "unknown"
for _, f := range []string{"etc/nftables.conf", "etc/sysconfig/nftables.conf"} {
if b, err := readSmall(j(f), 1<<20); err == nil {
def = ParseNftables(b, a)
break
}
}
return a.out("nftables", true, def)
case has("iptables"):
a := newFwAcc()
def := "unknown"
for _, f := range []string{"etc/iptables/rules.v4", "etc/sysconfig/iptables"} {
if b, err := readSmall(j(f), 1<<20); err == nil {
def = ParseIptablesSave(b, a)
break
}
}
return a.out("iptables", true, def)
case len(found.Found) > 0:
return &FWOut{Backend: found.Found[0], Active: true, DefaultIn: "unknown", Allowed: []int{}, Limited: []int{}, Restricted: []int{}, Denied: []int{}}
case found.Known:
return &FWOut{Backend: "none", DefaultIn: "accept", Allowed: []int{}, Limited: []int{}, Restricted: []int{}, Denied: []int{}}
}
return &FWOut{Backend: "unknown", DefaultIn: "unknown", Allowed: []int{}, Limited: []int{}, Restricted: []int{}, Denied: []int{}}
}
authlog.go496 lines
package main
// SSH logins and sudo use from the system's auth log (/var/log/auth.log on Debian and Ubuntu,
// /var/log/secure on RHEL and its relatives). The agent reads only what was appended since its
// previous look (it remembers where it stopped across restarts), at most authLogMaxRead bytes at a
// time, and keeps:
// - counters: failed and invalid-user attempts, accepted logins, the source addresses that tried
// most (with the time of their last try and the names of EXISTING accounts they tried: the
// names an attacker guesses are never sent, they may be someone's mistyped password);
// - successful SSH logins: time, account, source address and port, method (publickey, password…);
// - sudo use: time, account, allowed or refused (never the command).
// The first look after install also reads the last part of the file (authLogMaxRead) for the
// logins and sudo history (not the counters). On systems that log only to journald the file does
// not exist and the check is skipped. The file is readable by root and the adm group only: the
// detailed install (CAP_DAC_READ_SEARCH), --root or --auth-log (adm group) read it.
import (
"bytes"
"errors"
"io"
"io/fs"
"net"
"os"
"path/filepath"
"sort"
"strconv"
"strings"
"syscall"
"time"
)
const authLogMaxRead = 2 << 20
// SSHAuthOut is what the rollup carries: counts since the previous rollup.
type SSHAuthOut struct {
Source string `json:"source"`
Failed int `json:"failed"` // failed password / key / keyboard-interactive attempts
Invalid int `json:"invalid"` // connections naming a user that does not exist
Accepted int `json:"accepted"` // successful logins
IPs int `json:"ips"` // distinct source addresses of failures
Top []IPCount `json:"top"` // the 5 addresses with the most failures
WindowS int64 `json:"window_s"`
Truncated bool `json:"truncated,omitempty"` // more was written than the agent reads at once
}
type IPCount struct {
IP string `json:"ip"`
N int `json:"n"`
}
// AuthCounts accumulates between two rollups.
type AuthCounts struct {
Failed, Invalid, Accepted int
ByIP map[string]int
Truncated bool
Fail map[string]*IPFail // 0.3.0: per source address
Logins []AcceptedRec
Sudo []SudoRec
}
// IPFail: the failures of one source address.
type IPFail struct {
Invalid int
Last int64
Users map[string]int // account names tried (filtered to existing accounts before sending)
}
// AcceptedRec: "Accepted <method> for <user> from <ip> port <port>".
type AcceptedRec struct {
At int64 `json:"at"`
User string `json:"user"`
IP string `json:"ip"`
Method string `json:"method"`
Port int `json:"port,omitempty"`
}
// SudoRec: someone ran sudo (allowed or refused). The command is never kept.
type SudoRec struct {
At int64 `json:"at"`
User string `json:"user"`
OK bool `json:"ok"`
}
const maxAuthEvents = 200
func (a *AuthCounts) add(b AuthCounts) {
a.Failed += b.Failed
a.Invalid += b.Invalid
a.Accepted += b.Accepted
a.Truncated = a.Truncated || b.Truncated
a.Logins = lastN(append(a.Logins, b.Logins...), maxAuthEvents)
a.Sudo = lastN(append(a.Sudo, b.Sudo...), maxAuthEvents)
for ip, f := range b.Fail {
if a.Fail == nil {
a.Fail = map[string]*IPFail{}
}
g := a.Fail[ip]
if g == nil {
if len(a.Fail) >= 10000 {
continue
}
g = &IPFail{}
a.Fail[ip] = g
}
g.Invalid += f.Invalid
g.Last = max(g.Last, f.Last)
for u, n := range f.Users {
if g.Users == nil {
g.Users = map[string]int{}
}
if _, ok := g.Users[u]; ok || len(g.Users) < 20 {
g.Users[u] += n
}
}
}
for ip, n := range b.ByIP {
if a.ByIP == nil {
a.ByIP = map[string]int{}
}
if _, ok := a.ByIP[ip]; ok || len(a.ByIP) < 10000 {
a.ByIP[ip] += n
}
}
}
func (a AuthCounts) Out(source string, window int64) *SSHAuthOut {
o := &SSHAuthOut{Source: source, Failed: a.Failed, Invalid: a.Invalid, Accepted: a.Accepted, IPs: len(a.ByIP), WindowS: window,
Truncated: a.Truncated, Top: []IPCount{}}
for ip, n := range a.ByIP {
o.Top = append(o.Top, IPCount{ip, n})
}
sort.Slice(o.Top, func(i, j int) bool {
if o.Top[i].N != o.Top[j].N {
return o.Top[i].N > o.Top[j].N
}
return o.Top[i].IP < o.Top[j].IP
})
if len(o.Top) > 5 {
o.Top = o.Top[:5]
}
return o
}
// ParseAuthLines counts the sshd lines of a chunk of the auth log and collects the successful
// logins and the sudo use, with their times (now: the reference for syslog dates without a year).
func ParseAuthLines(b []byte) AuthCounts { return parseAuth(b, time.Now(), false) }
func parseAuth(b []byte, now time.Time, eventsOnly bool) AuthCounts {
c := AuthCounts{ByIP: map[string]int{}, Fail: map[string]*IPFail{}}
for len(b) > 0 {
var line []byte
if i := bytes.IndexByte(b, '\n'); i >= 0 {
line, b = b[:i], b[i+1:]
} else {
line, b = b, nil
}
if bytes.Contains(line, []byte(" sudo")) {
if r, ok := parseSudo(line, now); ok && len(c.Sudo) < 10*maxAuthEvents {
c.Sudo = append(c.Sudo, r)
}
continue
}
if !bytes.Contains(line, []byte("sshd")) { // sshd[123] and, since OpenSSH 9.8, sshd-session[123]
continue
}
s := string(line)
switch {
case strings.Contains(s, "]: Failed password for ") || strings.Contains(s, "]: Failed publickey for ") ||
strings.Contains(s, "]: Failed keyboard-interactive/pam for ") || strings.Contains(s, "]: Failed none for "):
if eventsOnly {
continue
}
c.Failed++
// "for invalid user x": its "Invalid user x from IP" line came first and counted the address.
if !strings.Contains(s, " for invalid user ") {
ip := c.countIP(s)
if f := c.fail(ip); f != nil {
f.Last = max(f.Last, syslogTime(line, now))
if u := wordAfter(s, " for "); u != "" && len(f.Users) < 20 {
f.Users[truncate(u, 32)]++
}
}
}
case strings.Contains(s, "]: Invalid user "):
if eventsOnly {
continue
}
c.Invalid++
ip := c.countIP(s)
if f := c.fail(ip); f != nil {
f.Invalid++
f.Last = max(f.Last, syslogTime(line, now))
}
case strings.Contains(s, "]: Accepted "):
if !eventsOnly {
c.Accepted++
}
if r, ok := parseAccepted(s, line, now); ok && len(c.Logins) < 10*maxAuthEvents {
c.Logins = append(c.Logins, r)
}
}
}
c.Logins = lastN(c.Logins, maxAuthEvents)
c.Sudo = lastN(c.Sudo, maxAuthEvents)
return c
}
func (c *AuthCounts) fail(ip string) *IPFail {
if ip == "" {
return nil
}
f := c.Fail[ip]
if f == nil {
if len(c.Fail) >= 10000 {
return nil
}
f = &IPFail{Users: map[string]int{}}
c.Fail[ip] = f
}
return f
}
func (c *AuthCounts) countIP(s string) string {
i := strings.LastIndex(s, " from ")
if i < 0 {
return ""
}
f := strings.Fields(s[i+6:])
if len(f) == 0 || net.ParseIP(f[0]) == nil {
return ""
}
if _, ok := c.ByIP[f[0]]; ok || len(c.ByIP) < 10000 {
c.ByIP[f[0]]++
}
return f[0]
}
// wordAfter returns the word that follows the first occurrence of marker.
func wordAfter(s, marker string) string {
i := strings.Index(s, marker)
if i < 0 {
return ""
}
return firstField(s[i+len(marker):])
}
// parseAccepted: "sshd[812]: Accepted publickey for kaan from 203.0.113.5 port 52144 ssh2: ED25519 SHA256:…"
func parseAccepted(s string, line []byte, now time.Time) (AcceptedRec, bool) {
i := strings.Index(s, "]: Accepted ")
if i < 0 {
return AcceptedRec{}, false
}
rest := s[i+len("]: Accepted "):]
f := strings.Fields(rest)
// method for user from ip port n
if len(f) < 5 || f[1] != "for" {
return AcceptedRec{}, false
}
r := AcceptedRec{At: syslogTime(line, now), Method: authMethod(f[0]), User: truncate(f[2], 32)}
for k := 3; k+1 < len(f); k++ {
switch f[k] {
case "from":
if net.ParseIP(f[k+1]) != nil {
r.IP = f[k+1]
}
case "port":
r.Port, _ = strconv.Atoi(f[k+1])
}
}
return r, r.User != "" && r.At > 0
}
func authMethod(m string) string {
switch m {
case "publickey", "password", "keyboard-interactive/pam", "keyboard-interactive", "hostbased", "gssapi-with-mic", "gssapi-keyex", "none":
return strings.TrimSuffix(m, "/pam")
}
return "other"
}
// parseSudo: "sudo: kaan : TTY=pts/0 ; PWD=/home/kaan ; USER=root ; COMMAND=/usr/bin/apt update" (allowed),
// "sudo: kaan : 3 incorrect password attempts ; TTY=…", "sudo: bob : user NOT in sudoers ; …" (refused).
// The command is never kept.
func parseSudo(line []byte, now time.Time) (SudoRec, bool) {
s := string(line)
i := strings.Index(s, " sudo: ")
if i < 0 {
i = strings.Index(s, " sudo[")
if i < 0 {
return SudoRec{}, false
}
j := strings.Index(s[i:], "]: ")
if j < 0 {
return SudoRec{}, false
}
i += j + 2
} else {
i += len(" sudo:")
}
rest := strings.TrimLeft(s[i:], " ")
if strings.HasPrefix(rest, "pam_unix") || strings.HasPrefix(rest, "pam_") {
return SudoRec{}, false
}
k := strings.Index(rest, " : ")
if k <= 0 {
return SudoRec{}, false
}
user := strings.TrimSpace(rest[:k])
what := rest[k+3:]
if user == "" || strings.ContainsAny(user, " ;=") {
return SudoRec{}, false
}
ok := strings.Contains(what, "COMMAND=")
if strings.Contains(what, "incorrect password") || strings.Contains(what, "NOT in sudoers") || strings.Contains(what, "command not allowed") ||
strings.Contains(what, "authentication failure") || strings.Contains(what, "a password is required") {
ok = false
} else if !ok {
return SudoRec{}, false
}
at := syslogTime(line, now)
if at == 0 {
return SudoRec{}, false
}
return SudoRec{At: at, User: truncate(user, 32), OK: ok}, true
}
// syslogTime reads the time at the start of a syslog line: RFC 3339 ("2026-10-07T10:12:01.123456+03:00",
// Ubuntu 24.04 and rsyslog's high-precision format) or the traditional "Oct 7 10:12:01" (local time,
// no year: the year that puts it in the past). 0 when neither.
func syslogTime(line []byte, now time.Time) int64 {
if len(line) >= 19 && line[4] == '-' && line[10] == 'T' {
end := bytes.IndexByte(line, ' ')
if end < 0 {
end = len(line)
}
if t, err := time.Parse(time.RFC3339Nano, string(line[:end])); err == nil {
return t.Unix()
}
return 0
}
if len(line) < 15 {
return 0
}
t, err := time.ParseInLocation("Jan _2 15:04:05", string(line[:15]), time.Local)
if err != nil {
return 0
}
t = time.Date(now.Year(), t.Month(), t.Day(), t.Hour(), t.Minute(), t.Second(), 0, time.Local)
if t.After(now.Add(48 * time.Hour)) {
t = t.AddDate(-1, 0, 0)
}
return t.Unix()
}
func lastN[T any](xs []T, n int) []T {
if len(xs) > n {
return append([]T(nil), xs[len(xs)-n:]...)
}
return xs
}
// AuthLog follows one log file across rotations.
type AuthLog struct {
Path string
dev, ino uint64
off int64
tail []byte // the bytes just before off: if they changed, the file was truncated and rewritten (copytruncate)
started bool
Backfill bool // the first Read also returns the logins and sudo use of the file's last authLogMaxRead bytes
}
// AuthPos is the reader's position, kept in the state directory across restarts.
type AuthPos struct {
Path string `json:"path"`
Dev uint64 `json:"dev"`
Ino uint64 `json:"ino"`
Off int64 `json:"off"`
Tail []byte `json:"tail"`
}
func (a *AuthLog) Pos() *AuthPos {
if !a.started {
return nil
}
return &AuthPos{Path: a.Path, Dev: a.dev, Ino: a.ino, Off: a.off, Tail: a.tail}
}
// Resume continues from a saved position (same path): what was written while the agent was
// stopped is counted too.
func (a *AuthLog) Resume(p *AuthPos) {
if p == nil || p.Path != a.Path || a.Path == "" {
return
}
a.dev, a.ino, a.off, a.tail, a.started = p.Dev, p.Ino, p.Off, p.Tail, true
}
var errNoAuthLog = errors.New("no auth log")
// FindAuthLog returns the configured path, or the first of /var/log/auth.log and /var/log/secure that exists.
func FindAuthLog(root, configured string) string {
if configured != "" {
return configured
}
for _, p := range []string{"var/log/auth.log", "var/log/secure"} {
if _, err := os.Stat(filepath.Join(root, p)); err == nil {
return filepath.Join(root, p)
}
}
return ""
}
// Read returns the counts in what was appended since the previous call. The first call only
// positions at the end of the file (history before the agent started is not counted).
func (a *AuthLog) Read() (AuthCounts, error) {
if a.Path == "" {
return AuthCounts{}, errNoAuthLog
}
f, err := os.Open(a.Path)
if err != nil {
return AuthCounts{}, err
}
defer f.Close()
fi, err := f.Stat()
if err != nil {
return AuthCounts{}, err
}
if !fi.Mode().IsRegular() {
return AuthCounts{}, &fs.PathError{Op: "read", Path: a.Path, Err: errors.New("not a regular file")}
}
var dev, ino uint64
if st, ok := fi.Sys().(*syscall.Stat_t); ok {
dev, ino = uint64(st.Dev), uint64(st.Ino)
}
size := fi.Size()
if !a.started {
a.started, a.dev, a.ino, a.off = true, dev, ino, size
a.tail = readTail(f, size)
if !a.Backfill || size == 0 {
return AuthCounts{}, nil
}
// History for the login and sudo lists (counters start from now).
from := max(0, size-authLogMaxRead)
buf := make([]byte, size-from)
n, err := f.ReadAt(buf, from)
if err != nil && !errors.Is(err, io.EOF) {
return AuthCounts{}, nil
}
buf = buf[:n]
if from > 0 {
if i := bytes.IndexByte(buf, '\n'); i >= 0 {
buf = buf[i+1:]
}
}
return parseAuth(buf, time.Now(), true), nil
}
if dev != a.dev || ino != a.ino || size < a.off || !bytes.Equal(readTail(f, a.off), a.tail) {
a.dev, a.ino, a.off = dev, ino, 0 // rotated, or truncated in place: the file from its start
}
var out AuthCounts
if size-a.off > authLogMaxRead { // a flood: count the latest part only
a.off = size - authLogMaxRead
out.Truncated = true
}
if size == a.off {
return out, nil
}
buf := make([]byte, size-a.off)
n, err := f.ReadAt(buf, a.off)
if err != nil && !errors.Is(err, io.EOF) {
return out, err
}
buf = buf[:n]
last := bytes.LastIndexByte(buf, '\n')
if last < 0 { // no complete line yet
return out, nil
}
a.off += int64(last + 1)
a.tail = readTail(f, a.off)
c := parseAuth(buf[:last+1], time.Now(), false)
c.Truncated = c.Truncated || out.Truncated
return c, nil
}
// readTail returns up to 64 bytes ending at off.
func readTail(f *os.File, off int64) []byte {
n := min(off, 64)
if n <= 0 {
return nil
}
b := make([]byte, n)
if _, err := f.ReadAt(b, off-n); err != nil {
return nil
}
return b
}
logins.go415 lines
package main
// Login history from the system's own records, read-only:
//
// /var/log/wtmp every login, logout, boot and shutdown (what `last` shows), binary utmp records
// /run/utmp who is logged in now (what `who` shows)
// /var/log/lastlog the last login of each account (what `lastlog` shows), indexed by uid
//
// wtmp is read incrementally (only what was appended since the previous look, following its
// monthly rotation to wtmp.1), at most wtmpMaxRead bytes at a time; the first look after install
// (or when the agent's state is lost) reads the last wtmpBackfill bytes for history. What leaves
// the machine: user name, source address (or host name), terminal, login and logout times, boot
// and shutdown times with the kernel version. Never anything typed in a session.
import (
"bytes"
"encoding/binary"
"errors"
"io"
"net"
"os"
"path/filepath"
"runtime"
"sort"
"strings"
"syscall"
"time"
)
const (
utEmpty = 0
utRunLvl = 1
utBootTime = 2
utNewTime = 3
utOldTime = 4
utInitProcess = 5
utLoginProcess = 6
utUserProcess = 7
utDeadProcess = 8
utAccounting = 9
wtmpMaxRead = 1 << 20
wtmpBackfill = 4 << 20
maxLogins = 50
maxBoots = 20
)
// UtmpRec is one decoded utmp / wtmp record.
type UtmpRec struct {
Type int16
PID int32
Line string
User string
Host string
Sec int64
Addr [16]byte
}
// utmpLayout: glibc's struct utmp is 384 bytes where the time fields are 32-bit
// (__WORDSIZE_TIME64_COMPAT32: x86_64) and 400 bytes where they are 64-bit.
type utmpLayout struct {
size int
sec int // offset of tv_sec
secSize int
addr int
lastlogRS int // the matching lastlog record size
}
var (
layout384 = utmpLayout{size: 384, sec: 340, secSize: 4, addr: 348, lastlogRS: 292}
layout400 = utmpLayout{size: 400, sec: 344, secSize: 8, addr: 360, lastlogRS: 296}
)
func cstr(b []byte) string {
if i := bytes.IndexByte(b, 0); i >= 0 {
b = b[:i]
}
return strings.ToValidUTF8(string(b), "?")
}
func (l utmpLayout) decode(b []byte) UtmpRec {
r := UtmpRec{Type: int16(binary.LittleEndian.Uint16(b[0:])), PID: int32(binary.LittleEndian.Uint32(b[4:])),
Line: cstr(b[8:40]), User: cstr(b[44:76]), Host: cstr(b[76:332])}
if l.secSize == 4 {
r.Sec = int64(int32(binary.LittleEndian.Uint32(b[l.sec:])))
} else {
r.Sec = int64(binary.LittleEndian.Uint64(b[l.sec:]))
}
copy(r.Addr[:], b[l.addr:l.addr+16])
return r
}
// plausible: a record that looks like real utmp data in this layout.
func (l utmpLayout) plausible(b []byte) bool {
t := int16(binary.LittleEndian.Uint16(b[0:]))
if t < utEmpty || t > utAccounting {
return false
}
if t == utEmpty {
return true
}
r := l.decode(b)
return r.Sec > 315532800 && r.Sec < time.Now().Unix()+366*86400 // 1980 .. a year ahead
}
// DetectUtmpLayout picks the record size whose first records decode sensibly (the native one first).
func DetectUtmpLayout(b []byte, size int64) utmpLayout {
cands := []utmpLayout{layout384, layout400}
if runtime.GOARCH != "amd64" && runtime.GOARCH != "386" {
cands = []utmpLayout{layout400, layout384}
}
for _, l := range cands {
if size > 0 && size%int64(l.size) != 0 {
continue
}
ok, checked := true, 0
for off := 0; off+l.size <= len(b) && checked < 8; off += l.size {
if !l.plausible(b[off : off+l.size]) {
ok = false
break
}
checked++
}
if ok && (checked > 0 || len(b) == 0) {
return l
}
}
return cands[0]
}
// ParseUtmp decodes whole records.
func ParseUtmp(b []byte, l utmpLayout) []UtmpRec {
out := make([]UtmpRec, 0, len(b)/l.size)
for off := 0; off+l.size <= len(b); off += l.size {
out = append(out, l.decode(b[off:off+l.size]))
}
return out
}
// remoteOf: the source address of a login (ut_addr_v6, else ut_host when it is an address) and the
// host name when ut_host is a name.
func remoteOf(r UtmpRec) (ip, host string) {
h := strings.TrimSpace(r.Host)
if i := strings.LastIndexByte(h, ':'); i > 0 && strings.Count(h, ":") == 1 { // "10.0.0.5:0" (X display)
h = h[:i]
}
if net.ParseIP(h) != nil {
return h, ""
}
var zero [16]byte
if r.Addr != zero {
if bytes.Equal(r.Addr[4:], zero[4:]) {
return net.IP(r.Addr[:4]).String(), truncate(h, 64)
}
return net.IP(r.Addr[:]).String(), truncate(h, 64)
}
return "", truncate(h, 64)
}
// LoginRec is one successful login.
type LoginRec struct {
At int64 `json:"at"`
User string `json:"user"`
IP string `json:"ip,omitempty"`
Host string `json:"host,omitempty"`
TTY string `json:"tty,omitempty"`
End *int64 `json:"end"` // logout; null while open or unknown
Src string `json:"src"` // wtmp
}
// BootRec is one boot, with the kernel it started and the shutdown that ended it (when recorded).
type BootRec struct {
At int64 `json:"at"`
Kernel string `json:"kernel,omitempty"`
End *int64 `json:"end"`
}
type SessionRec struct {
User string `json:"user"`
IP string `json:"ip,omitempty"`
Host string `json:"host,omitempty"`
TTY string `json:"tty,omitempty"`
Since int64 `json:"since"`
}
// wtmpState is what the reader keeps between looks (persisted in the state directory).
type wtmpState struct {
Dev, Ino uint64
Off int64
Size int // record size
Open map[string]LoginRec // tty -> open login
LastBoot *BootRec
Last map[string]int64 // user -> last login seen in wtmp
}
// WtmpEvents: what one look found.
type WtmpEvents struct {
Logins []LoginRec // new logins, and earlier ones that now have their logout time
Boots []BootRec
}
// ApplyWtmp turns records into logins and boots, using and updating the open sessions.
func ApplyWtmp(st *wtmpState, recs []UtmpRec) WtmpEvents {
var ev WtmpEvents
if st.Open == nil {
st.Open = map[string]LoginRec{}
}
if st.Last == nil {
st.Last = map[string]int64{}
}
closeAll := func(at int64) {
for tty, l := range st.Open {
e := at
l.End = &e
ev.Logins = append(ev.Logins, l)
delete(st.Open, tty)
}
}
for _, r := range recs {
switch {
case r.Type == utUserProcess && r.User != "" && r.Line != "":
if old, ok := st.Open[r.Line]; ok { // the previous one on this terminal never logged out
e := r.Sec
old.End = &e
ev.Logins = append(ev.Logins, old)
}
ip, host := remoteOf(r)
l := LoginRec{At: r.Sec, User: truncate(r.User, 32), IP: ip, Host: host, TTY: truncate(r.Line, 32), Src: "wtmp"}
st.Open[r.Line] = l
ev.Logins = append(ev.Logins, l)
if r.Sec > st.Last[l.User] {
st.Last[l.User] = r.Sec
}
if len(st.Open) > 512 { // never-closed records (a crashed logger): keep the map bounded
for k := range st.Open {
delete(st.Open, k)
break
}
}
case (r.Type == utDeadProcess || (r.Type == utUserProcess && r.User == "")) && r.Line != "":
if l, ok := st.Open[r.Line]; ok {
e := r.Sec
l.End = &e
ev.Logins = append(ev.Logins, l)
delete(st.Open, r.Line)
}
case r.Type == utBootTime || (r.Type == utRunLvl && r.User == "reboot"):
closeAll(r.Sec)
b := BootRec{At: r.Sec, Kernel: truncate(r.Host, 80)}
st.LastBoot = &b
ev.Boots = append(ev.Boots, b)
case r.Type == utRunLvl && r.User == "shutdown":
closeAll(r.Sec)
if st.LastBoot != nil && st.LastBoot.End == nil {
e := r.Sec
st.LastBoot.End = &e
ev.Boots = append(ev.Boots, *st.LastBoot)
}
}
}
return ev
}
// ReadWtmp reads what was appended to wtmp since the previous look (or, the first time, its last
// wtmpBackfill bytes). Rotation: when the file was replaced, the rest of the old one is read from
// wtmp.1 (same inode) before the new file.
func ReadWtmp(path string, st *wtmpState, backfill bool) (WtmpEvents, error) {
f, err := os.Open(path)
if err != nil {
return WtmpEvents{}, err
}
defer f.Close()
fi, err := f.Stat()
if err != nil {
return WtmpEvents{}, err
}
if !fi.Mode().IsRegular() {
return WtmpEvents{}, errors.New("wtmp: not a regular file")
}
dev, ino := devIno(fi)
size := fi.Size()
var all WtmpEvents
if st.Size == 0 {
head := make([]byte, min(int64(8*400), size))
_, _ = f.ReadAt(head, 0)
st.Size = DetectUtmpLayout(head, size).size
}
l := layout384
if st.Size == 400 {
l = layout400
}
switch {
case backfill || st.Ino == 0:
start := max(0, size-wtmpBackfill)
start -= start % int64(l.size)
st.Off = start
case dev != st.Dev || ino != st.Ino:
// rotated: finish the old file (now wtmp.1) if it is still there, then the new one from its start
if of, err := os.Open(path + ".1"); err == nil {
if ofi, err := of.Stat(); err == nil {
if d, i := devIno(ofi); d == st.Dev && i == st.Ino && ofi.Size() > st.Off {
b, _ := readRange(of, st.Off, min(ofi.Size()-st.Off, wtmpMaxRead), l.size)
ev := ApplyWtmp(st, ParseUtmp(b, l))
all.Logins = append(all.Logins, ev.Logins...)
all.Boots = append(all.Boots, ev.Boots...)
}
}
of.Close()
}
st.Off = 0
case size < st.Off: // truncated in place
st.Off = 0
}
st.Dev, st.Ino = dev, ino
if size-st.Off > wtmpMaxRead && !backfill { // a flood: the latest part only
st.Off = size - wtmpMaxRead
st.Off -= st.Off % int64(l.size)
}
if size > st.Off {
b, err := readRange(f, st.Off, size-st.Off, l.size)
if err != nil {
return all, err
}
st.Off += int64(len(b))
ev := ApplyWtmp(st, ParseUtmp(b, l))
all.Logins = append(all.Logins, ev.Logins...)
all.Boots = append(all.Boots, ev.Boots...)
}
return all, nil
}
// readRange reads n bytes at off, cut to whole records.
func readRange(f *os.File, off, n int64, rec int) ([]byte, error) {
if n <= 0 {
return nil, nil
}
b := make([]byte, n)
k, err := f.ReadAt(b, off)
if err != nil && !errors.Is(err, io.EOF) {
return nil, err
}
k -= k % rec
return b[:k], nil
}
func devIno(fi os.FileInfo) (uint64, uint64) {
if st, ok := fi.Sys().(*syscall.Stat_t); ok {
return uint64(st.Dev), uint64(st.Ino)
}
return 0, 0
}
// ReadUtmpSessions lists who is logged in now (live USER_PROCESS records of /run/utmp).
func ReadUtmpSessions(root, procRoot string) ([]SessionRec, error) {
var b []byte
var err error
for _, p := range []string{"run/utmp", "var/run/utmp"} {
if b, err = readSmall(filepath.Join(root, p), 1<<20); err == nil {
break
}
}
if err != nil {
return nil, err
}
l := DetectUtmpLayout(b, int64(len(b)))
out := []SessionRec{}
for _, r := range ParseUtmp(b, l) {
if r.Type != utUserProcess || r.User == "" {
continue
}
if r.PID > 0 && procRoot != "" {
if _, err := os.Stat(filepath.Join(procRoot, itoa(int(r.PID)))); err != nil {
continue // a stale record: its process is gone
}
}
ip, host := remoteOf(r)
out = append(out, SessionRec{User: truncate(r.User, 32), IP: ip, Host: host, TTY: truncate(r.Line, 32), Since: r.Sec})
if len(out) >= maxLogins {
break
}
}
sort.Slice(out, func(i, j int) bool { return out[i].Since > out[j].Since })
return out, nil
}
// ReadLastlog returns the last login time of each account in /var/log/lastlog (a sparse file of
// fixed-size records indexed by uid). Accounts never logged in are left out.
func ReadLastlog(path string, users []PasswdEntry, rs int) map[string]int64 {
f, err := os.Open(path)
if err != nil {
return nil
}
defer f.Close()
out := map[string]int64{}
buf := make([]byte, rs)
for _, u := range users {
if u.UID < 0 || u.UID > 1<<24 || (u.UID != 0 && !LoginShell(u.Shell)) {
continue
}
if _, err := f.ReadAt(buf, int64(u.UID)*int64(rs)); err != nil {
continue
}
var t int64
if rs == 292 {
t = int64(int32(binary.LittleEndian.Uint32(buf)))
} else {
t = int64(binary.LittleEndian.Uint64(buf))
}
if t > 315532800 {
out[u.Name] = t
}
}
return out
}
owners.go297 lines
package main
// Who listens on a port: the process holding the socket (found through /proc/<pid>/fd links,
// which needs CAP_SYS_PTRACE for other users' processes: the kernel checks PTRACE_MODE_READ before
// it shows another process's fd or exe link), its systemd unit (from /proc/<pid>/cgroup) and, when
// it runs in a Docker container, the container's name (from Docker's own metadata file
// /var/lib/docker/containers/<id>/config.v2.json, read-only). A port published by Docker belongs
// to docker-proxy: its -container-ip argument names the container it forwards to.
//
// The search runs on the background worker only when the set of listening sockets changed, a few
// processes at a time (newest first: a new port usually belongs to a new process), and stops as
// soon as every socket has an owner.
import (
"bytes"
"encoding/json"
"os"
"path/filepath"
"sort"
"strconv"
"strings"
"sync"
"syscall"
"time"
)
// Owner is the process behind a listening socket.
type Owner struct {
PID int
Name string
User string
Exe string
Unit string
Container string
CID string
}
// ParseCgroup returns the systemd unit and the container id (first 12 hex digits) of a process
// from its /proc/<pid>/cgroup: the cgroup v2 line "0::/…", else the v1 name=systemd line.
func ParseCgroup(b []byte) (unit, cid string) {
var path []byte
for len(b) > 0 {
var line []byte
if i := bytes.IndexByte(b, '\n'); i >= 0 {
line, b = b[:i], b[i+1:]
} else {
line, b = b, nil
}
if bytes.HasPrefix(line, []byte("0::")) {
path = line[3:]
break
}
if i := bytes.Index(line, []byte(":name=systemd:")); i >= 0 && path == nil {
path = line[i+len(":name=systemd:"):]
}
}
if path == nil {
return "", ""
}
parts := strings.Split(string(path), "/")
for i, p := range parts {
if cid == "" {
if id := containerID(p); id != "" {
cid = id
} else if i > 0 && (parts[i-1] == "docker" || parts[i-1] == "libpod" || parts[i-1] == "containerd") && isHex(p, 64) {
cid = p[:12]
}
}
if unit == "" && (strings.HasSuffix(p, ".service") || strings.HasSuffix(p, ".scope")) && containerID(p) == "" {
unit = truncate(p, 120)
}
}
return unit, cid
}
// containerID: "docker-<64 hex>.scope", "libpod-<id>.scope", "cri-containerd-<id>.scope".
func containerID(p string) string {
for _, pre := range []string{"docker-", "libpod-", "cri-containerd-", "crio-"} {
if strings.HasPrefix(p, pre) {
id := strings.TrimSuffix(strings.TrimPrefix(p, pre), ".scope")
if isHex(id, 64) {
return id[:12]
}
}
}
return ""
}
func isHex(s string, n int) bool {
if len(s) != n {
return false
}
for i := 0; i < len(s); i++ {
c := s[i]
if !(c >= '0' && c <= '9' || c >= 'a' && c <= 'f') {
return false
}
}
return true
}
// DockerIndex maps container ids and container IP addresses to container names, from
// /var/lib/docker/containers/*/config.v2.json (read again only when that directory changes).
type DockerIndex struct {
root string
mu sync.Mutex
mtime time.Time
byID map[string]string // 12-hex id -> name
byIP map[string]string
loaded bool
}
func NewDockerIndex(root string) *DockerIndex { return &DockerIndex{root: root} }
func (d *DockerIndex) refresh() {
dir := filepath.Join(d.root, "var/lib/docker/containers")
fi, err := os.Stat(dir)
if err != nil {
d.byID, d.byIP, d.loaded = nil, nil, true
return
}
if d.loaded && fi.ModTime().Equal(d.mtime) {
return
}
d.mtime, d.loaded = fi.ModTime(), true
d.byID, d.byIP = map[string]string{}, map[string]string{}
ents, err := os.ReadDir(dir)
if err != nil {
return
}
for i, e := range ents {
if i >= 500 {
break
}
if !e.IsDir() || !isHex(e.Name(), 64) {
continue
}
b, err := readSmall(filepath.Join(dir, e.Name(), "config.v2.json"), 2<<20)
if err != nil {
continue
}
var cfg struct {
Name string
NetworkSettings struct {
Networks map[string]struct{ IPAddress, GlobalIPv6Address string }
}
}
if json.Unmarshal(b, &cfg) != nil {
continue
}
name := truncate(strings.TrimPrefix(cfg.Name, "/"), 100)
if name == "" {
continue
}
d.byID[e.Name()[:12]] = name
for _, n := range cfg.NetworkSettings.Networks {
if n.IPAddress != "" {
d.byIP[n.IPAddress] = name
}
if n.GlobalIPv6Address != "" {
d.byIP[n.GlobalIPv6Address] = name
}
}
}
}
// Name of a container by its 12-hex id ("" when unknown).
func (d *DockerIndex) Name(cid string) string {
d.mu.Lock()
defer d.mu.Unlock()
d.refresh()
return d.byID[cid]
}
// ByIP: the container that has this address on a Docker network.
func (d *DockerIndex) ByIP(ip string) string {
d.mu.Lock()
defer d.mu.Unlock()
d.refresh()
return d.byIP[ip]
}
// procRef is the little the owner search needs about a process, copied from the sample.
type procRef struct {
PID int
Start uint64
Kernel bool
}
// FindOwners maps socket inodes to the process holding them, keeping the owners already known for
// sockets that are still there. budget bounds the fd entries looked at (a busy proxy can hold a
// million); g paces the work.
func FindOwners(procRoot string, inodes map[uint64]bool, known map[uint64]Owner, procs []procRef, users *UserCache,
docker *DockerIndex, g *Gentle, budget int) map[uint64]Owner {
out := make(map[uint64]Owner, len(inodes))
missing := 0
for ino := range inodes {
if o, ok := known[ino]; ok {
out[ino] = o
} else {
missing++
}
}
if missing == 0 {
return out
}
order := append([]procRef(nil), procs...)
sort.Slice(order, func(i, j int) bool { return order[i].Start > order[j].Start }) // newest first
lbuf := make([]byte, 128)
found := map[int]bool{}
var pids []int
byPid := map[int][]uint64{}
for _, p := range order {
if budget <= 0 || missing == 0 {
break
}
if p.Kernel || p.PID <= 0 {
continue
}
g.Tick()
fdDir := procRoot + "/" + strconv.Itoa(p.PID) + "/fd"
d, err := os.Open(fdDir)
if err != nil {
continue
}
names, _ := d.Readdirnames(budget)
d.Close()
budget -= len(names)
for _, n := range names {
k, err := readlinkInto(fdDir+"/"+n, lbuf)
if err != nil || k < 9 || !bytes.HasPrefix(lbuf[:k], []byte("socket:[")) || lbuf[k-1] != ']' {
continue
}
ino, ok := parseUintBytes(lbuf[8 : k-1])
if !ok || !inodes[ino] {
continue
}
if _, have := out[ino]; have {
continue
}
out[ino] = Owner{PID: p.PID}
missing--
if !found[p.PID] {
found[p.PID] = true
pids = append(pids, p.PID)
}
byPid[p.PID] = append(byPid[p.PID], ino)
}
}
for _, pid := range pids {
g.Tick()
o := describeOwner(procRoot, pid, users, docker)
for _, ino := range byPid[pid] {
out[ino] = o
}
}
return out
}
// describeOwner reads name, user, executable, unit and container of one process.
func describeOwner(procRoot string, pid int, users *UserCache, docker *DockerIndex) Owner {
o := Owner{PID: pid}
dir := procRoot + "/" + strconv.Itoa(pid)
buf := make([]byte, 4096)
if n, err := readInto(dir+"/stat", buf); err == nil {
var st PidStat
var f [24][]byte
if comm, ok := parsePidStatInto(buf[:n], &st, f[:]); ok {
o.Name = truncate(string(comm), 64)
}
}
var st syscall.Stat_t
if syscall.Stat(dir, &st) == nil {
o.User = users.Name(int(st.Uid))
}
if n, err := readlinkInto(dir+"/exe", buf); err == nil {
o.Exe = truncate(string(buf[:n]), 200)
}
if n, err := readInto(dir+"/cgroup", buf); err == nil {
o.Unit, o.CID = ParseCgroup(buf[:n])
}
if o.CID != "" && docker != nil {
o.Container = docker.Name(o.CID)
}
if docker != nil && o.Container == "" && (o.Name == "docker-proxy" || strings.HasSuffix(o.Exe, "/docker-proxy")) {
if n, err := readInto(dir+"/cmdline", buf); err == nil {
args := strings.Split(strings.TrimRight(string(buf[:n]), "\x00"), "\x00")
for i := 0; i+1 < len(args); i++ {
if args[i] == "-container-ip" {
o.Container = docker.ByIP(args[i+1])
}
}
}
}
return o
}
dbus.go609 lines
package main
// A minimal D-Bus client: just enough of the wire protocol
// (https://dbus.freedesktop.org/doc/dbus-specification.html) to ask systemd for its units over the
// system bus, as `systemctl list-units` does, without running systemctl. It authenticates as its
// own uid (EXTERNAL), says Hello and calls read-only methods any user may call:
// org.freedesktop.systemd1.Manager.ListUnits, and, for a failed unit only,
// org.freedesktop.DBus.Properties.GetAll on that unit (its description, result, exit code or
// signal, restart count and when it failed: what `systemctl status` shows). Every read is bounded
// and the whole exchange has a deadline.
import (
"bufio"
"encoding/binary"
"encoding/hex"
"errors"
"fmt"
"io"
"net"
"os"
"strconv"
"strings"
"time"
)
const (
dbusMethodCall = 1
dbusMethodReturn = 2
dbusError = 3
dbusSignal = 4
dbusMaxMessage = 16 << 20
dbusMaxDepth = 32
)
// UnitInfo is one loaded systemd unit.
type UnitInfo struct {
Name, Description, LoadState, ActiveState, SubState string
Path string // D-Bus object path
}
// UnitDetail is what systemd says about a failed unit.
type UnitDetail struct {
Unit string `json:"unit"`
Description string `json:"description,omitempty"`
Result string `json:"result,omitempty"` // exit-code | signal | core-dump | timeout | start-limit-hit | resources | oom-kill …
ExitCode *int `json:"exit_code"` // the main process's exit status (203 = EXEC: binary missing …)
Signal *string `json:"signal"` // SIGKILL … when it was killed
Restarts *int `json:"restarts"`
Since int64 `json:"since,omitempty"` // when it entered the failed state
Log []string `json:"log,omitempty"` // the last lines it logged (redacted), only once after it newly failed
}
var signalNames = map[int]string{1: "SIGHUP", 2: "SIGINT", 3: "SIGQUIT", 4: "SIGILL", 5: "SIGTRAP", 6: "SIGABRT", 7: "SIGBUS",
8: "SIGFPE", 9: "SIGKILL", 10: "SIGUSR1", 11: "SIGSEGV", 12: "SIGUSR2", 13: "SIGPIPE", 14: "SIGALRM", 15: "SIGTERM", 24: "SIGXCPU", 25: "SIGXFSZ"}
// unitIface: the interface that carries Result for a unit type.
func unitIface(name string) string {
i := strings.LastIndexByte(name, '.')
if i < 0 {
return ""
}
switch name[i+1:] {
case "service":
return "org.freedesktop.systemd1.Service"
case "socket":
return "org.freedesktop.systemd1.Socket"
case "mount":
return "org.freedesktop.systemd1.Mount"
case "timer":
return "org.freedesktop.systemd1.Timer"
case "swap":
return "org.freedesktop.systemd1.Swap"
case "path":
return "org.freedesktop.systemd1.Path"
case "scope":
return "org.freedesktop.systemd1.Scope"
case "automount":
return "org.freedesktop.systemd1.Automount"
}
return ""
}
// UnitDetails asks systemd for the properties of a few units (at most 20) over one connection.
func UnitDetails(socket string, timeout time.Duration, units []UnitInfo) (map[string]UnitDetail, error) {
out := map[string]UnitDetail{}
if len(units) == 0 {
return out, nil
}
conn, err := net.DialTimeout("unix", socket, timeout)
if err != nil {
return nil, err
}
defer conn.Close()
_ = conn.SetDeadline(time.Now().Add(timeout))
d := &dbusConn{w: conn, r: bufio.NewReaderSize(conn, 32*1024)}
if err := d.auth(); err != nil {
return nil, err
}
if _, _, err := d.call("org.freedesktop.DBus", "/org/freedesktop/DBus", "org.freedesktop.DBus", "Hello"); err != nil {
return nil, fmt.Errorf("hello: %w", err)
}
for i, u := range units {
if i >= 20 || !strings.HasPrefix(u.Path, "/org/freedesktop/systemd1/unit/") {
continue
}
props := map[string]any{}
for _, iface := range []string{"org.freedesktop.systemd1.Unit", unitIface(u.Name)} {
if iface == "" {
continue
}
h, body, err := d.callArgs("org.freedesktop.systemd1", u.Path, "org.freedesktop.DBus.Properties", "GetAll", iface)
if err != nil || h.signature != "a{sv}" {
continue
}
for k, v := range parseProps(body, h.order) {
props[k] = v
}
}
out[u.Name] = unitDetailFrom(u, props)
}
return out, nil
}
func unitDetailFrom(u UnitInfo, p map[string]any) UnitDetail {
ud := UnitDetail{Unit: truncate(u.Name, 120), Description: truncate(u.Description, 200)}
if s, ok := p["Description"].(string); ok && s != "" {
ud.Description = truncate(s, 200)
}
if s, ok := p["Result"].(string); ok {
ud.Result = truncate(s, 40)
}
code, hasCode := p["ExecMainCode"].(int32)
status, hasStatus := p["ExecMainStatus"].(int32)
if hasCode && hasStatus {
switch code {
case 1: // CLD_EXITED
v := int(status)
ud.ExitCode = &v
case 2, 3: // CLD_KILLED, CLD_DUMPED
name := signalNames[int(status)]
if name == "" {
name = "SIG" + strconv.Itoa(int(status))
}
ud.Signal = &name
}
}
if n, ok := p["NRestarts"].(uint32); ok {
v := int(n)
ud.Restarts = &v
}
for _, k := range []string{"StateChangeTimestamp", "InactiveEnterTimestamp"} {
if t, ok := p[k].(uint64); ok && t > 0 {
ud.Since = int64(t / 1_000_000)
break
}
}
return ud
}
// parseProps decodes the a{sv} reply of Properties.GetAll, keeping strings, integers and booleans.
func parseProps(body []byte, bo binary.ByteOrder) map[string]any {
out := map[string]any{}
d := &dbusDec{b: body, bo: bo}
n := int(d.u32())
d.align(8)
end := d.off + n
if d.err != nil || end > len(body) {
return out
}
for d.err == nil && d.off < end && len(out) < 1000 {
d.align(8)
key := d.str()
sig := d.sig()
if d.err != nil {
break
}
if k, err := typeEnd(sig, 0, 0); err != nil || k != len(sig) {
break
}
switch sig {
case "s", "o":
out[key] = d.str()
case "i":
out[key] = int32(d.u32())
case "u":
out[key] = d.u32()
case "t", "x":
d.align(8)
if d.need(8) {
out[key] = d.bo.Uint64(d.b[d.off:])
d.off += 8
}
case "b":
out[key] = d.u32() != 0
default:
d.skip(sig, 1)
}
}
return out
}
// SystemBusPath: $DBUS_SYSTEM_BUS_ADDRESS (unix:path=...) or the standard socket.
func SystemBusPath() string {
if a := os.Getenv("DBUS_SYSTEM_BUS_ADDRESS"); strings.HasPrefix(a, "unix:path=") {
p := strings.TrimPrefix(a, "unix:path=")
if i := strings.IndexByte(p, ','); i >= 0 {
p = p[:i]
}
return p
}
return "/run/dbus/system_bus_socket"
}
// ListUnits returns the units systemd has loaded (active, failed, and inactive ones still referenced).
func ListUnits(socket string, timeout time.Duration) ([]UnitInfo, error) {
conn, err := net.DialTimeout("unix", socket, timeout)
if err != nil {
return nil, err
}
defer conn.Close()
_ = conn.SetDeadline(time.Now().Add(timeout))
d := &dbusConn{w: conn, r: bufio.NewReaderSize(conn, 32*1024)}
if err := d.auth(); err != nil {
return nil, err
}
if _, _, err := d.call("org.freedesktop.DBus", "/org/freedesktop/DBus", "org.freedesktop.DBus", "Hello"); err != nil {
return nil, fmt.Errorf("hello: %w", err)
}
h, body, err := d.call("org.freedesktop.systemd1", "/org/freedesktop/systemd1", "org.freedesktop.systemd1.Manager", "ListUnits")
if err != nil {
return nil, fmt.Errorf("ListUnits: %w", err)
}
if h.signature != "a(ssssssouso)" {
return nil, fmt.Errorf("ListUnits: unexpected reply signature %q", h.signature)
}
return parseUnitList(body, h.order)
}
type dbusConn struct {
w io.Writer
r *bufio.Reader
serial uint32
}
func (d *dbusConn) auth() error {
uid := strconv.Itoa(os.Geteuid())
if _, err := d.w.Write([]byte("\x00AUTH EXTERNAL " + hex.EncodeToString([]byte(uid)) + "\r\n")); err != nil {
return err
}
line, err := d.r.ReadSlice('\n')
if err != nil {
return fmt.Errorf("auth: %w", err)
}
if !strings.HasPrefix(string(line), "OK ") {
return fmt.Errorf("auth refused: %q", truncate(strings.TrimSpace(string(line)), 80))
}
_, err = d.w.Write([]byte("BEGIN\r\n"))
return err
}
// call sends a method call without arguments and waits for its reply, skipping signals.
func (d *dbusConn) call(dest, path, iface, member string) (dbusHeader, []byte, error) {
return d.callArgs(dest, path, iface, member)
}
// callArgs sends a method call with string arguments (the interface name for Properties.GetAll)
// and waits for its reply, skipping signals.
func (d *dbusConn) callArgs(dest, path, iface, member string, args ...string) (dbusHeader, []byte, error) {
d.serial++
if _, err := d.w.Write(encodeCall(d.serial, dest, path, iface, member, args...)); err != nil {
return dbusHeader{}, nil, err
}
for i := 0; i < 64; i++ {
h, body, err := d.read()
if err != nil {
return h, nil, err
}
if h.replySerial != d.serial || (h.typ != dbusMethodReturn && h.typ != dbusError) {
continue // NameAcquired and other signals
}
if h.typ == dbusError {
return h, nil, errors.New(truncate(h.errName, 120))
}
return h, body, nil
}
return dbusHeader{}, nil, errors.New("no reply")
}
// encodeCall builds a METHOD_CALL message (little endian) whose body is the string arguments.
func encodeCall(serial uint32, dest, path, iface, member string, args ...string) []byte {
body := &dbusEnc{}
for _, a := range args {
body.str(a)
}
e := &dbusEnc{}
e.b = append(e.b, 'l', dbusMethodCall, 0, 1)
e.u32(0) // body length
e.u32(serial)
at := len(e.b)
e.u32(0) // header fields array length, filled below
start := len(e.b)
field := func(code byte, typ, val string) {
e.align(8)
e.b = append(e.b, code)
e.sig(typ)
e.str(val)
}
field(1, "o", path)
field(2, "s", iface)
field(3, "s", member)
field(6, "s", dest)
if len(args) > 0 {
e.align(8)
e.b = append(e.b, 8)
e.sig("g")
e.sig(strings.Repeat("s", len(args)))
}
binary.LittleEndian.PutUint32(e.b[at:], uint32(len(e.b)-start))
e.align(8)
binary.LittleEndian.PutUint32(e.b[4:], uint32(len(body.b)))
return append(e.b, body.b...)
}
type dbusEnc struct{ b []byte }
func (e *dbusEnc) align(n int) {
for len(e.b)%n != 0 {
e.b = append(e.b, 0)
}
}
func (e *dbusEnc) u32(v uint32) { e.align(4); e.b = binary.LittleEndian.AppendUint32(e.b, v) }
func (e *dbusEnc) str(s string) { e.u32(uint32(len(s))); e.b = append(append(e.b, s...), 0) }
func (e *dbusEnc) sig(s string) { e.b = append(append(append(e.b, byte(len(s))), s...), 0) }
type dbusHeader struct {
typ byte
order binary.ByteOrder
replySerial uint32
signature string
errName string
}
// read reads one message: its header fields of interest and its body.
func (d *dbusConn) read() (dbusHeader, []byte, error) {
var h dbusHeader
var fixed [16]byte
if _, err := io.ReadFull(d.r, fixed[:]); err != nil {
return h, nil, err
}
switch fixed[0] {
case 'l':
h.order = binary.LittleEndian
case 'B':
h.order = binary.BigEndian
default:
return h, nil, fmt.Errorf("bad endianness byte %q", fixed[0])
}
h.typ = fixed[1]
bodyLen := h.order.Uint32(fixed[4:])
fieldsLen := h.order.Uint32(fixed[12:])
if bodyLen > dbusMaxMessage || fieldsLen > dbusMaxMessage {
return h, nil, errors.New("message too large")
}
hdrLen := 16 + int(fieldsLen)
pad := (8 - hdrLen%8) % 8
total := hdrLen + pad + int(bodyLen)
if total > dbusMaxMessage {
return h, nil, errors.New("message too large")
}
buf := make([]byte, total)
copy(buf, fixed[:])
if _, err := io.ReadFull(d.r, buf[16:]); err != nil {
return h, nil, err
}
hd := &dbusDec{b: buf[:hdrLen], off: 12, bo: h.order}
n := int(hd.u32())
end := hd.off + n
for hd.err == nil && hd.off < end {
hd.align(8)
code := hd.u8()
t := hd.sig()
if hd.err != nil {
break
}
if k, err := typeEnd(t, 0, 0); err != nil || k != len(t) {
return h, nil, errors.New("bad header field type")
}
switch {
case code == 5 && t == "u":
h.replySerial = hd.u32()
case code == 8 && t == "g":
h.signature = hd.sig()
case code == 4 && t == "s":
h.errName = hd.str()
default:
hd.skip(t, 0)
}
}
if hd.err != nil {
return h, nil, fmt.Errorf("header: %w", hd.err)
}
return h, buf[hdrLen+pad:], nil
}
// parseUnitList decodes the a(ssssssouso) reply of ListUnits.
func parseUnitList(body []byte, bo binary.ByteOrder) ([]UnitInfo, error) {
d := &dbusDec{b: body, bo: bo}
n := int(d.u32())
d.align(8)
end := d.off + n
if d.err == nil && end > len(body) {
return nil, errors.New("unit list: truncated")
}
var out []UnitInfo
for d.err == nil && d.off < end && len(out) < 50000 {
d.align(8)
u := UnitInfo{Name: d.str(), Description: d.str(), LoadState: d.str(), ActiveState: d.str(), SubState: d.str()}
d.str() // following
u.Path = d.str() // unit object path
d.u32() // job id
d.str() // job type
d.str() // job object path
if d.err == nil {
out = append(out, u)
}
}
if d.err != nil {
return nil, fmt.Errorf("unit list: %w", d.err)
}
return out, nil
}
var errShort = errors.New("truncated message")
type dbusDec struct {
b []byte
off int
bo binary.ByteOrder
err error
}
func (d *dbusDec) need(n int) bool {
if d.err != nil {
return false
}
if n < 0 || d.off+n > len(d.b) {
d.err = errShort
return false
}
return true
}
func (d *dbusDec) align(n int) {
p := (d.off + n - 1) / n * n
if d.need(p - d.off) {
d.off = p
}
}
func (d *dbusDec) u8() byte {
if !d.need(1) {
return 0
}
d.off++
return d.b[d.off-1]
}
func (d *dbusDec) u32() uint32 {
d.align(4)
if !d.need(4) {
return 0
}
d.off += 4
return d.bo.Uint32(d.b[d.off-4:])
}
func (d *dbusDec) str() string {
n := int(d.u32())
if !d.need(n + 1) {
return ""
}
s := string(d.b[d.off : d.off+n])
d.off += n + 1
return s
}
func (d *dbusDec) sig() string {
n := int(d.u8())
if !d.need(n + 1) {
return ""
}
s := string(d.b[d.off : d.off+n])
d.off += n + 1
return s
}
func dbusAlign(c byte) int {
switch c {
case 'n', 'q':
return 2
case 'b', 'i', 'u', 'h', 's', 'o', 'a':
return 4
case 'x', 't', 'd', '(', '{':
return 8
}
return 1 // y, g, v
}
// typeEnd returns the index just past the single complete type starting at sig[i].
func typeEnd(sig string, i, depth int) (int, error) {
if depth > dbusMaxDepth || i >= len(sig) {
return 0, errors.New("bad signature")
}
switch c := sig[i]; c {
case 'y', 'b', 'n', 'q', 'i', 'u', 'x', 't', 'd', 'h', 's', 'o', 'g', 'v':
return i + 1, nil
case 'a':
return typeEnd(sig, i+1, depth+1)
case '(', '{':
closer := byte(')')
if c == '{' {
closer = '}'
}
j := i + 1
for j < len(sig) && sig[j] != closer {
k, err := typeEnd(sig, j, depth+1)
if err != nil {
return 0, err
}
j = k
}
if j >= len(sig) || j == i+1 {
return 0, errors.New("bad signature")
}
return j + 1, nil
}
return 0, errors.New("bad signature")
}
// skip moves past one value of the single complete type t.
func (d *dbusDec) skip(t string, depth int) {
if d.err != nil {
return
}
if depth > dbusMaxDepth || t == "" {
d.err = errors.New("bad signature")
return
}
switch t[0] {
case 'y':
d.u8()
case 'n', 'q':
d.align(2)
if d.need(2) {
d.off += 2
}
case 'b', 'i', 'u', 'h':
d.u32()
case 'x', 't', 'd':
d.align(8)
if d.need(8) {
d.off += 8
}
case 's', 'o':
d.str()
case 'g':
d.sig()
case 'v':
inner := d.sig()
if k, err := typeEnd(inner, 0, depth+1); err != nil || k != len(inner) {
d.err = errors.New("bad variant")
return
}
d.skip(inner, depth+1)
case 'a':
n := int(d.u32())
elem := t[1:]
d.align(dbusAlign(elem[0]))
if !d.need(n) {
return
}
end := d.off + n
for d.err == nil && d.off < end {
before := d.off
d.skip(elem, depth+1)
if d.off <= before {
d.err = errors.New("bad array")
}
}
if d.err == nil && d.off != end {
d.err = errors.New("bad array length")
}
case '(', '{':
d.align(8)
j := 1
for d.err == nil && j < len(t)-1 {
k, err := typeEnd(t, j, depth+1)
if err != nil {
d.err = err
return
}
d.skip(t[j:k], depth+1)
j = k
}
default:
d.err = errors.New("bad signature")
}
}
unitlog.go103 lines
package main
// The last lines a failed unit logged, for its explanation on the account page. Read from the
// tail of the classic syslog file (/var/log/syslog on Debian and Ubuntu, /var/log/messages on
// RHEL), which rsyslog fills from the journal; the agent never runs journalctl (it executes no
// commands at all). Systems that log only to journald send none. Only done once per unit when it
// newly fails, at most unitLogRead bytes; each line is redacted like a command line and cut to 300
// characters.
import (
"bytes"
"os"
"path/filepath"
"strings"
)
const unitLogRead = 1 << 20
// SyslogPath: the first syslog file that exists.
func SyslogPath(root string) string {
for _, p := range []string{"var/log/syslog", "var/log/messages"} {
if fi, err := os.Stat(filepath.Join(root, p)); err == nil && fi.Mode().IsRegular() {
return filepath.Join(root, p)
}
}
return ""
}
// UnitLogTails returns the last (at most 10) lines about each unit in the syslog file's tail:
// lines naming the unit ("cloud-init.service: Main process exited…") or written by its program
// ("cloud-init[612]: …").
func UnitLogTails(path string, units []string) (map[string][]string, error) {
out := map[string][]string{}
if path == "" || len(units) == 0 {
return out, nil
}
f, err := os.Open(path)
if err != nil {
return nil, err
}
defer f.Close()
fi, err := f.Stat()
if err != nil {
return nil, err
}
from := max(0, fi.Size()-unitLogRead)
buf := make([]byte, fi.Size()-from)
n, _ := f.ReadAt(buf, from)
buf = buf[:n]
if from > 0 {
if i := bytes.IndexByte(buf, '\n'); i >= 0 {
buf = buf[i+1:]
}
}
type pat struct{ unit, ident []byte }
pats := make([]pat, 0, len(units))
for _, u := range units {
base := u
if i := strings.LastIndexByte(u, '.'); i > 0 {
base = u[:i]
}
ident := []byte(" " + base + "[")
if strings.Contains(base, "@") { // user@1000: its own lines come from "systemd[pid]"
ident = nil
}
pats = append(pats, pat{unit: []byte(u), ident: ident})
}
lines := bytes.Split(buf, []byte("\n"))
for i := len(lines) - 1; i >= 0; i-- {
l := lines[i]
for k, p := range pats {
u := units[k]
if len(out[u]) >= 10 {
continue
}
if bytes.Contains(l, p.unit) || (p.ident != nil && bytes.Contains(l, p.ident)) {
out[u] = append(out[u], truncate(Redact(stripSyslogPrefix(string(l))), 300))
}
}
}
for u, ls := range out { // oldest first
for i, j := 0, len(ls)-1; i < j; i, j = i+1, j-1 {
ls[i], ls[j] = ls[j], ls[i]
}
out[u] = ls
}
return out, nil
}
// stripSyslogPrefix drops the time and host name: "2026-10-07T10:00:00+00:00 host prog[1]: msg" or
// "Oct 7 10:00:00 host prog[1]: msg" become "prog[1]: msg".
func stripSyslogPrefix(l string) string {
f := strings.Fields(l)
skip := 2 // RFC 3339 time, host
if len(f) > 0 && len(f[0]) == 3 && !strings.ContainsAny(f[0], "0123456789") {
skip = 4 // month day time host
}
if len(f) <= skip {
return strings.TrimSpace(l)
}
return strings.Join(f[skip:], " ")
}
identity.go303 lines
package main
// What the agent is allowed to see here (its mode), what the machine is (virtualization), and the
// agent's own footprint (CPU and memory, reported in every rollup so you can check its cost).
import (
"bytes"
"os"
"path/filepath"
"strconv"
"strings"
"time"
)
// Linux capability numbers (include/uapi/linux/capability.h).
const (
capDacOverride = 1
capDacReadSearch = 2
capSysPtrace = 19
capSysAdmin = 21
)
// Identity: how the agent runs.
type Identity struct {
UID int
Root bool
DacRead bool // CAP_DAC_READ_SEARCH: may read any file (never writes: the unit's filesystem is read-only)
Ptrace bool // CAP_SYS_PTRACE: may look at other users' /proc/<pid>/fd, exe and io (PTRACE_MODE_READ checks)
Caps []string
Mode string // unprivileged | detailed | root
}
// DetectIdentity reads the effective capabilities from /proc/self/status (CapEff).
func DetectIdentity(procRoot string) Identity {
id := Identity{UID: os.Geteuid()}
id.Root = id.UID == 0
var eff uint64
if b, err := readSmall(filepath.Join(procRoot, "self/status"), 64*1024); err == nil {
eff = ParseCapEff(b)
}
has := func(n uint) bool { return eff&(1<<n) != 0 }
id.DacRead = has(capDacReadSearch) || has(capDacOverride)
id.Ptrace = has(capSysPtrace)
if has(capDacReadSearch) {
id.Caps = append(id.Caps, "dac_read_search")
}
if has(capDacOverride) {
id.Caps = append(id.Caps, "dac_override")
}
if has(capSysPtrace) {
id.Caps = append(id.Caps, "sys_ptrace")
}
if has(capSysAdmin) {
id.Caps = append(id.Caps, "sys_admin")
}
switch {
case id.Root:
id.Mode = "root"
case id.DacRead || id.Ptrace:
id.Mode = "detailed"
default:
id.Mode = "unprivileged"
}
return id
}
// ParseCapEff returns the CapEff mask of a /proc/<pid>/status file.
func ParseCapEff(b []byte) uint64 {
for _, l := range bytes.Split(b, []byte("\n")) {
if bytes.HasPrefix(l, []byte("CapEff:")) {
v, err := strconv.ParseUint(strings.TrimSpace(string(l[7:])), 16, 64)
if err == nil {
return v
}
}
}
return 0
}
// SeesOthers: the kernel shows this agent other users' fd / exe / io links.
func (id Identity) SeesOthers() bool { return id.Root || id.Ptrace }
// ---------------------------------------------------------------------------
// virtualization
// ---------------------------------------------------------------------------
type VirtInfo struct {
Virt string // kvm | qemu | vmware | microsoft | xen | oracle | amazon | google | openstack | parallels | docker | lxc | podman | wsl | container | vm | none | ""
Vendor string
Product string
}
// DetectVirt tells what the machine is without running systemd-detect-virt: container markers
// first (/run/systemd/container, /.dockerenv, /run/.containerenv, container= in PID 1's
// environment), then WSL, then the DMI strings in /sys/class/dmi/id, /sys/hypervisor and the
// "hypervisor" CPU flag.
func DetectVirt(root, proc, sys string) VirtInfo {
var v VirtInfo
rd := func(p string) string {
b, err := readSmall(p, 4096)
if err != nil {
return ""
}
return strings.TrimSpace(string(b))
}
v.Vendor = truncate(rd(filepath.Join(sys, "class/dmi/id/sys_vendor")), 80)
v.Product = truncate(rd(filepath.Join(sys, "class/dmi/id/product_name")), 80)
if c := rd(filepath.Join(root, "run/systemd/container")); c != "" {
v.Virt = containerKind(c)
return v
}
if exists(filepath.Join(root, ".dockerenv")) {
v.Virt = "docker"
return v
}
if exists(filepath.Join(root, "run/.containerenv")) {
v.Virt = "podman"
return v
}
if b, err := readSmall(filepath.Join(proc, "1/environ"), 64*1024); err == nil {
for _, kv := range bytes.Split(b, []byte{0}) {
if bytes.HasPrefix(kv, []byte("container=")) {
v.Virt = containerKind(string(kv[len("container="):]))
return v
}
}
}
if strings.Contains(strings.ToLower(rd(filepath.Join(proc, "sys/kernel/osrelease"))), "microsoft") {
v.Virt = "wsl"
return v
}
if k := VirtFromDMI(v.Vendor, v.Product, rd(filepath.Join(sys, "class/dmi/id/bios_vendor"))); k != "" {
v.Virt = k
return v
}
if strings.Contains(strings.ToLower(rd(filepath.Join(sys, "hypervisor/type"))), "xen") {
v.Virt = "xen"
return v
}
if b, err := readSmall(filepath.Join(proc, "cpuinfo"), 256*1024); err == nil {
for _, l := range strings.Split(string(b), "\n") {
if strings.HasPrefix(l, "flags") || strings.HasPrefix(l, "Features") {
if strings.Contains(" "+l+" ", " hypervisor ") {
v.Virt = "vm"
} else if v.Vendor != "" {
v.Virt = "none"
}
break
}
}
}
return v
}
func containerKind(s string) string {
s = strings.ToLower(strings.TrimSpace(s))
switch {
case strings.HasPrefix(s, "docker"):
return "docker"
case strings.HasPrefix(s, "lxc"):
return "lxc"
case strings.HasPrefix(s, "podman"):
return "podman"
case s == "":
return ""
}
return "container"
}
// VirtFromDMI maps the DMI vendor / product / BIOS vendor strings to a hypervisor.
func VirtFromDMI(vendor, product, bios string) string {
all := strings.ToLower(vendor + " | " + product + " | " + bios)
switch {
case strings.Contains(all, "amazon ec2") || strings.Contains(all, "amazon"):
return "amazon"
case strings.Contains(all, "google"):
return "google"
case strings.Contains(all, "vmware"):
return "vmware"
case strings.Contains(all, "microsoft corporation") && strings.Contains(all, "virtual"):
return "microsoft"
case strings.Contains(all, "innotek") || strings.Contains(all, "virtualbox"):
return "oracle"
case strings.Contains(all, "xen"):
return "xen"
case strings.Contains(all, "parallels"):
return "parallels"
case strings.Contains(all, "openstack"):
return "openstack"
case strings.Contains(all, "kvm") || strings.Contains(all, "digitalocean") || strings.Contains(all, "hetzner") ||
strings.Contains(all, "alibaba") || strings.Contains(all, "rhev") || strings.Contains(all, "linode") || strings.Contains(all, "vultr") ||
strings.Contains(all, "scaleway") || strings.Contains(all, "ovh") || strings.Contains(all, "contabo"):
return "kvm"
case strings.Contains(all, "qemu") || strings.Contains(all, "bochs") || strings.Contains(all, "seabios"):
return "qemu"
}
return ""
}
// ---------------------------------------------------------------------------
// the agent's own footprint
// ---------------------------------------------------------------------------
// selfStats: the agent's CPU time and resident memory, from /proc/self/stat at each sample.
type selfStats struct {
buf []byte
f [24][]byte
prevTicks uint64
prevT time.Time
have bool
winTicks uint64
winT time.Time
peak float64
rssSum int64
rssMax int64
rssN int64
}
func (s *selfStats) sample(procRoot string, now time.Time) {
if s.buf == nil {
s.buf = make([]byte, 1024)
}
n, err := readInto(procRoot+"/self/stat", s.buf)
if err != nil {
return
}
var st PidStat
if _, ok := parsePidStatInto(s.buf[:n], &st, s.f[:]); !ok {
return
}
ticks := st.UTime + st.STime
if s.have {
if el := now.Sub(s.prevT).Seconds(); el > 0 && ticks >= s.prevTicks {
s.peak = maxF(s.peak, 100*float64(ticks-s.prevTicks)/clockTicks/el)
}
} else {
s.winTicks, s.winT = ticks, now
}
s.prevTicks, s.prevT, s.have = ticks, now, true
rss := st.RSSPages * pageSize
s.rssSum += rss
s.rssN++
if rss > s.rssMax {
s.rssMax = rss
}
}
// window closes the rollup window: average CPU (% of one core) and memory since the previous call.
func (s *selfStats) window(now time.Time) (cpu, peak float64, rss, rssMax, windowS int64) {
if !s.have {
return 0, 0, 0, 0, 0
}
if el := s.prevT.Sub(s.winT).Seconds(); el > 0 && s.prevTicks >= s.winTicks {
cpu = 100 * float64(s.prevTicks-s.winTicks) / clockTicks / el
}
windowS = int64(now.Sub(s.winT).Seconds())
if s.rssN > 0 {
rss = s.rssSum / s.rssN
}
peak, rssMax = s.peak, s.rssMax
s.winTicks, s.winT = s.prevTicks, s.prevT
s.peak, s.rssSum, s.rssN, s.rssMax = 0, 0, 0, 0
return round3(cpu), round3(peak), rss, rssMax, windowS
}
func round3(v float64) float64 { return float64(int64(v*1000+0.5)) / 1000 }
// ---------------------------------------------------------------------------
// helpers of the collector
// ---------------------------------------------------------------------------
// procNameSet: the process names of the latest sample (time, firewall and fail2ban daemons).
func (c *Collector) procNameSet() map[string]bool {
if c.namesAt == c.namesGen && c.procNames != nil {
return c.procNames
}
names := make(map[string]bool, 128)
for _, p := range c.lastProcs {
if !p.kernel && len(names) < 4096 {
names[p.Name] = true
}
}
c.procNames, c.namesAt = names, c.namesGen
return names
}
// procRefs copies what the worker needs about the processes of the latest sample.
func (c *Collector) procRefs() []procRef {
out := make([]procRef, 0, len(c.lastProcs))
for _, p := range c.lastProcs {
out = append(out, procRef{PID: p.PID, Start: p.startTime, Kernel: p.kernel})
}
return out
}
// withUnits adds the systemd unit to the top processes (their cgroup is read once per process).
func (c *Collector) withUnits(ps []Proc) []Proc {
for i := range ps {
ps[i].Unit = c.procs.UnitOf(ps[i].PID)
}
return ps
}
procs.go394 lines
package main
import (
"bytes"
"os"
"regexp"
"strconv"
"strings"
"sync"
"syscall"
"time"
)
const clockTicks = 100 // USER_HZ: 100 on every Linux build for amd64 and arm64
var pageSize = int64(os.Getpagesize())
// Proc is one process as reported (top lists and signal evidence).
type Proc struct {
PID int `json:"pid"`
Name string `json:"name"`
User string `json:"user"`
Exe string `json:"exe,omitempty"`
Cmd string `json:"cmd,omitempty"`
CPU float64 `json:"cpu"` // percent of one core over the last sample interval
RSS int64 `json:"rss"` // bytes
Unit string `json:"unit,omitempty"` // systemd unit (0.3.0), from /proc/<pid>/cgroup
Start int64 `json:"started_at,omitempty"` // unix seconds (0.3.0)
// internal
uid int
cmdRaw string // the command line as read, never sent (Reported() redacts it into Cmd)
argv0 string
exeKnown bool
startTime uint64
kernel bool
minerWhy string // MinerMatch, computed when the command line was read
sigDone bool // minerWhy is set (Procs built by hand in tests leave it false)
cmdRed string // Redact(cmdRaw), computed when the command line was read
redDone bool
}
// procEntry is what the sampler remembers about one process (pid + start time) between samples:
// the CPU ticks of the previous sample, and what is read only when the process is first seen and
// then again every refreshEvery samples (owner, executable, command line, the miner verdict).
type procEntry struct {
start uint64
ticks uint64 // utime+stime at the previous sample
at time.Time // when they were read (the sample is spread over a few seconds)
gen uint32
uid int
user string
name string
exe string
exeKnown bool
argv0 string
cmdRaw string
cmdRed string // redacted once per read of the command line, not at every report
minerWhy string
unit string
unitRead bool
}
// ProcSampler reads /proc/<pid>/stat of every process once per sample (one small read, parsed
// without allocating) and keeps the CPU ticks of the previous sample to turn them into a
// percentage. The owner, the executable path and the command line of a process are read when it
// is first seen and then once every refreshEvery samples, spread over the pids (so a sample never
// re-reads all of them at once): a command line rewritten later, or a binary deleted after start,
// is still noticed within that period.
type ProcSampler struct {
root string
ents map[int]*procEntry
prevT time.Time
users *UserCache
gen uint32
refreshEvery uint32
bootTime int64
buf []byte // /proc/<pid>/stat
lbuf []byte // readlink
cbuf []byte // cmdline, cgroup
out []Proc
fields [24][]byte
pbuf []byte
// Gentle mode: the processes are read in chunks of Chunk with a Pause between chunks (the
// kernel spends most of a sample generating the stat files), so a sample of a few hundred
// processes is spread over a second or two instead of one burst. Sleep nil: no pauses (tests).
Chunk int
Pause time.Duration
Sleep func(time.Duration)
}
func NewProcSampler(root string, users *UserCache) *ProcSampler {
return &ProcSampler{root: root, ents: map[int]*procEntry{}, users: users, refreshEvery: 15,
buf: make([]byte, 1024), lbuf: make([]byte, 4096), cbuf: make([]byte, 4096)}
}
// maxProcs bounds the work per sample on a machine with an extreme number of processes.
const maxProcs = 32768
// Sample returns every process. The slice is reused by the next call: callers copy what they keep.
func (s *ProcSampler) Sample(now time.Time) []Proc {
d, err := os.Open(s.root)
if err != nil {
return nil
}
names, _ := d.Readdirnames(maxProcs + 512)
d.Close()
s.gen++
out := s.out[:0]
readAt := time.Now()
done := 0
for _, name := range names {
if len(name) == 0 || name[0] < '0' || name[0] > '9' {
continue
}
pid, err := strconv.Atoi(name)
if err != nil || pid <= 0 {
continue
}
if done++; s.Sleep != nil && s.Chunk > 0 && done%s.Chunk == 0 {
s.Sleep(s.Pause)
readAt = time.Now()
}
s.pbuf = append(append(append(append(s.pbuf[:0], s.root...), '/'), name...), "/stat\x00"...)
n, err := readIntoZ(s.pbuf, s.buf)
if err != nil || n == 0 {
continue
}
var st PidStat
comm, ok := parsePidStatInto(s.buf[:n], &st, s.fields[:])
if !ok {
continue
}
kernel := pid == 2 || st.PPid == 2
e := s.ents[pid]
isNew := e == nil || e.start != st.StartTime
if isNew {
e = &procEntry{start: st.StartTime}
s.ents[pid] = e
}
e.gen = s.gen
reread := isNew || (uint32(pid)+s.gen)%s.refreshEvery == 0
if e.name != string(comm) {
e.name = string(comm)
reread = true // renamed (exec): read it again
}
if reread {
s.fill(e, s.root+"/"+name, kernel)
}
ticks := st.UTime + st.STime
var cpu float64
at := readAt
if s.Sleep == nil { // no pauses: the sample's own time (tests drive it)
at = now
}
if elapsed := at.Sub(e.at).Seconds(); !isNew && elapsed > 0 && ticks >= e.ticks {
cpu = round1(100 * float64(ticks-e.ticks) / clockTicks / elapsed)
}
e.ticks, e.at = ticks, at
p := Proc{PID: pid, Name: e.name, User: e.user, Exe: e.exe, CPU: cpu, RSS: st.RSSPages * pageSize, uid: e.uid,
cmdRaw: e.cmdRaw, argv0: e.argv0, exeKnown: e.exeKnown, startTime: st.StartTime, kernel: kernel,
minerWhy: e.minerWhy, sigDone: true, Unit: e.unit, cmdRed: e.cmdRed, redDone: true}
if s.bootTime > 0 {
p.Start = s.bootTime + int64(st.StartTime/clockTicks)
}
out = append(out, p)
if len(out) >= maxProcs {
break
}
}
for pid, e := range s.ents {
if e.gen != s.gen {
delete(s.ents, pid)
}
}
s.out = out
s.prevT = now
return out
}
// fill reads the owner, executable and command line of one process.
func (s *ProcSampler) fill(e *procEntry, dir string, kernel bool) {
var st syscall.Stat_t
if err := syscall.Stat(dir, &st); err == nil {
e.uid = int(st.Uid)
} else {
e.uid = -1
}
e.user = s.users.Name(e.uid)
if kernel {
return
}
if n, err := readlinkInto(dir+"/exe", s.lbuf); err == nil {
if e.exe != string(s.lbuf[:n]) {
e.exe = string(s.lbuf[:n])
}
e.exeKnown = true
} else {
e.exeKnown = false
}
e.argv0, e.cmdRaw = "", ""
if n, err := readInto(dir+"/cmdline", s.cbuf); err == nil && n > 0 {
cb := bytes.TrimRight(s.cbuf[:n], "\x00")
if i := bytes.IndexByte(cb, 0); i >= 0 {
e.argv0 = string(cb[:i])
} else {
e.argv0 = string(cb)
}
e.cmdRaw = strings.ReplaceAll(string(cb), "\x00", " ")
}
e.minerWhy = sigs.MinerMatch(e.name, e.argv0, e.cmdRaw)
e.cmdRed = truncate(Redact(e.cmdRaw), 200)
}
// UnitOf returns the systemd unit of a process the last sample saw (read from its cgroup once).
func (s *ProcSampler) UnitOf(pid int) string {
e := s.ents[pid]
if e == nil {
return ""
}
if !e.unitRead {
e.unitRead = true
if n, err := readInto(s.root+"/"+strconv.Itoa(pid)+"/cgroup", s.cbuf); err == nil {
e.unit, _ = ParseCgroup(s.cbuf[:n])
}
}
return e.unit
}
// sendCmdline: the `cmdline` check (on unless `disable: [cmdline]`).
var sendCmdline = true
// Reported is the process as it goes over the network: the command line redacted (see Redact)
// and cut to 200 characters, or left out with `disable: [cmdline]`.
func (p Proc) Reported() Proc {
q := p
q.Cmd = ""
if sendCmdline {
if p.redDone {
q.Cmd = p.cmdRed
} else {
q.Cmd = truncate(Redact(p.cmdRaw), 200)
}
}
q.Exe = truncate(p.Exe, 200)
q.Name = truncate(p.Name, 64)
q.cmdRaw, q.argv0, q.cmdRed = "", "", ""
return q
}
var (
// key=value and key: value where the key names a secret (DB_PASSWORD=, api_key:, ?token=)
redactKV = regexp.MustCompile(`(?i)((?:pass(?:word|wd|phrase)?|pwd|secret|token|api[_-]?key|apikey|access[_-]?key|auth|credentials?|private[_-]?key|session[_-]?key)[\w.-]*[=:])[^\s&;,'"]+`)
// --password VALUE, --token=VALUE ...
redactFlag = regexp.MustCompile(`(?i)(--?(?:pass(?:word|wd|phrase)?|secret|token|api-?key|auth(?:-token)?|access-key|secret-key|private-key|client-secret)[ =])[^\s'"]+`)
// HTTP headers on a command line (curl -H "Authorization: Bearer ...")
redactHeader = regexp.MustCompile(`(?i)((?:authorization|proxy-authorization|x-api-key|x-auth-token|cookie)\s*[:=]\s*(?:(?:bearer|basic|token|digest)\s+)?)[^\s'"]+`)
redactBearer = regexp.MustCompile(`(?i)(\bbearer\s+)[A-Za-z0-9._~+/=-]{8,}`)
// user:password@ in URLs, curl -u user:password, sshpass -p
redactURL = regexp.MustCompile(`(://[^/\s:@]+:)[^@\s/]+@`)
redactUser = regexp.MustCompile(`((?:^|\s)(?:-u|--user)\s*[^\s:'"]+:)[^\s'"]+`)
redactPass = regexp.MustCompile(`(\bsshpass\s+-p\s*)\S+`)
redactMy = regexp.MustCompile(`(\s-p)[^\s]+`)
// Well-known token formats, wherever they appear: JWTs, AWS access keys, GitHub / GitLab /
// Slack / Stripe tokens, PEM private keys pasted into an argument.
redactToken = regexp.MustCompile(`\beyJ[A-Za-z0-9_-]{8,}\.[A-Za-z0-9_-]{8,}\.[A-Za-z0-9_-]{8,}|\b(?:AKIA|ASIA)[0-9A-Z]{16}\b|` +
`\bgh[pousr]_[A-Za-z0-9]{30,}|\bgithub_pat_[A-Za-z0-9_]{30,}|\bglpat-[A-Za-z0-9_-]{20,}|\bxox[abposr]-[A-Za-z0-9-]{10,}|` +
`\b[rs]k_(?:live|test)_[A-Za-z0-9]{16,}|-----BEGIN [A-Z ]*PRIVATE KEY-----[^-]*`)
)
// Redact hides values that look like secrets in a command line or a crontab line: key=value
// pairs named like passwords or tokens, --password values, Authorization headers, bearer tokens,
// user:password@ in URLs, curl -u user:password, sshpass -p, mysql-style -pSECRET and tokens in
// well-known formats. It errs on the side of hiding.
func Redact(cmd string) string {
cmd = redactToken.ReplaceAllString(cmd, "***")
cmd = redactHeader.ReplaceAllString(cmd, "${1}***")
cmd = redactBearer.ReplaceAllString(cmd, "${1}***")
cmd = redactKV.ReplaceAllString(cmd, "${1}***")
cmd = redactFlag.ReplaceAllString(cmd, "${1}***")
cmd = redactURL.ReplaceAllString(cmd, "${1}***@")
cmd = redactUser.ReplaceAllString(cmd, "${1}***")
cmd = redactPass.ReplaceAllString(cmd, "${1}***")
if strings.Contains(cmd, "mysql") || strings.Contains(cmd, "mariadb") {
cmd = redactMy.ReplaceAllString(cmd, "${1}***")
}
return cmd
}
// TopBy returns the n processes with the highest key (ties: lower pid first), as reported.
func TopBy(ps []Proc, n int, key func(Proc) float64) []Proc {
cp := topRaw(ps, n, key)
out := make([]Proc, len(cp))
for i, p := range cp {
out[i] = p.Reported()
}
return out
}
// topRaw keeps the n best in a small sorted array: one pass, no copy or sort of the whole table.
func topRaw(ps []Proc, n int, key func(Proc) float64) []Proc {
if n <= 0 {
return nil
}
top := make([]Proc, 0, n)
keys := make([]float64, 0, n)
better := func(k float64, p Proc, i int) bool {
return k > keys[i] || (k == keys[i] && p.PID < top[i].PID)
}
for _, p := range ps {
k := key(p)
if len(top) == n && !better(k, p, n-1) {
continue
}
i := len(top)
if i < n {
top = append(top, p)
keys = append(keys, k)
} else {
i = n - 1
}
for i > 0 && better(k, p, i-1) {
top[i], keys[i] = top[i-1], keys[i-1]
i--
}
top[i], keys[i] = p, k
}
return top
}
// UserCache maps uids to names from /etc/passwd, reloading when the file changes.
type UserCache struct {
path string
mu sync.Mutex
mtime time.Time
checked time.Time
names map[int]string
all []PasswdEntry
}
func NewUserCache(path string) *UserCache { return &UserCache{path: path, names: map[int]string{}} }
func (u *UserCache) refresh() {
if now := time.Now(); now.Sub(u.checked) < 30*time.Second && u.names != nil && len(u.all) > 0 {
return
} else {
u.checked = now
}
fi, err := os.Stat(u.path)
if err != nil || fi.ModTime().Equal(u.mtime) {
return
}
b, err := os.ReadFile(u.path)
if err != nil {
return
}
u.mtime = fi.ModTime()
u.all = ParsePasswd(b)
u.names = make(map[int]string, len(u.all))
for _, e := range u.all {
if _, ok := u.names[e.UID]; !ok {
u.names[e.UID] = e.Name
}
}
}
func (u *UserCache) Name(uid int) string {
u.mu.Lock()
defer u.mu.Unlock()
u.refresh()
if n, ok := u.names[uid]; ok {
return n
}
return strconv.Itoa(uid)
}
func (u *UserCache) All() []PasswdEntry {
u.mu.Lock()
defer u.mu.Unlock()
u.refresh()
return append([]PasswdEntry(nil), u.all...)
}
func fileUID(fi os.FileInfo) int {
if st, ok := fi.Sys().(*syscall.Stat_t); ok {
return int(st.Uid)
}
return -1
}
func round1(v float64) float64 { return float64(int64(v*10+0.5)) / 10 }
func itoa(n int) string { return strconv.Itoa(n) }
func itoa64(n int64) string { return strconv.FormatInt(n, 10) }
procfs.go638 lines
package main
// Parsers for the plain-text files under /proc. Each parser takes the file's
// bytes so the tests can feed fixtures (testdata/); the readers that open the
// real files live next to the code that samples them.
import (
"bufio"
"bytes"
"encoding/hex"
"errors"
"fmt"
"io"
"net"
"strconv"
"strings"
)
// CPUTimes is the aggregate "cpu" line of /proc/stat, in clock ticks.
type CPUTimes struct {
User, Nice, System, Idle, IOWait, IRQ, SoftIRQ, Steal uint64
}
func (c CPUTimes) Total() uint64 {
return c.User + c.Nice + c.System + c.Idle + c.IOWait + c.IRQ + c.SoftIRQ + c.Steal
}
// Busy is everything but idle and iowait.
func (c CPUTimes) Busy() uint64 { return c.Total() - c.Idle - c.IOWait }
// ProcStat is what the sampler needs from /proc/stat.
type ProcStat struct {
CPU CPUTimes
Cores int
BootTime int64
}
func ParseProcStat(b []byte) (ProcStat, error) {
var ps ProcStat
found := false
sc := bufio.NewScanner(bytes.NewReader(b))
sc.Buffer(make([]byte, 64*1024), 1024*1024)
for sc.Scan() {
f := strings.Fields(sc.Text())
if len(f) == 0 {
continue
}
switch {
case f[0] == "cpu":
if len(f) < 5 {
return ps, errors.New("short cpu line")
}
v := make([]uint64, 8)
for i := 0; i < 8 && i+1 < len(f); i++ {
n, err := strconv.ParseUint(f[i+1], 10, 64)
if err != nil {
return ps, fmt.Errorf("cpu field %d: %w", i, err)
}
v[i] = n
}
ps.CPU = CPUTimes{v[0], v[1], v[2], v[3], v[4], v[5], v[6], v[7]}
found = true
case strings.HasPrefix(f[0], "cpu"):
ps.Cores++
case f[0] == "btime" && len(f) > 1:
ps.BootTime, _ = strconv.ParseInt(f[1], 10, 64)
}
}
if !found {
return ps, errors.New("no cpu line")
}
if ps.Cores == 0 {
ps.Cores = 1
}
return ps, nil
}
// CPUPercent between two readings of the aggregate line (0..100).
func CPUPercent(prev, cur CPUTimes) float64 {
dt := float64(cur.Total()) - float64(prev.Total())
if dt <= 0 {
return 0
}
db := float64(cur.Busy()) - float64(prev.Busy())
if db < 0 {
db = 0
}
return clamp(100*db/dt, 0, 100)
}
// MemInfo in bytes.
type MemInfo struct {
Total, Available, Free, Buffers, Cached, SwapTotal, SwapFree uint64
}
func (m MemInfo) UsedPct() float64 {
if m.Total == 0 {
return 0
}
return clamp(100*float64(m.Total-minU(m.Available, m.Total))/float64(m.Total), 0, 100)
}
func (m MemInfo) SwapUsed() uint64 {
if m.SwapFree > m.SwapTotal {
return 0
}
return m.SwapTotal - m.SwapFree
}
func (m MemInfo) SwapPct() float64 {
if m.SwapTotal == 0 {
return 0
}
return clamp(100*float64(m.SwapUsed())/float64(m.SwapTotal), 0, 100)
}
func ParseMemInfo(b []byte) (MemInfo, error) {
var m MemInfo
hasAvail := false
sc := bufio.NewScanner(bytes.NewReader(b))
for sc.Scan() {
line := sc.Text()
i := strings.IndexByte(line, ':')
if i < 0 {
continue
}
f := strings.Fields(line[i+1:])
if len(f) == 0 {
continue
}
n, err := strconv.ParseUint(f[0], 10, 64)
if err != nil {
continue
}
if len(f) > 1 && f[1] == "kB" {
n *= 1024
}
switch line[:i] {
case "MemTotal":
m.Total = n
case "MemAvailable":
m.Available, hasAvail = n, true
case "MemFree":
m.Free = n
case "Buffers":
m.Buffers = n
case "Cached":
m.Cached = n
case "SwapTotal":
m.SwapTotal = n
case "SwapFree":
m.SwapFree = n
}
}
if m.Total == 0 {
return m, errors.New("no MemTotal")
}
if !hasAvail { // kernels before 3.14
m.Available = m.Free + m.Buffers + m.Cached
}
return m, nil
}
// VMStat holds the /proc/vmstat counters the agent uses.
type VMStat struct {
PswpIn, PswpOut uint64
OOMKill uint64 // processes killed by the OOM killer since boot
HasOOM bool // the counter exists (Linux 4.13+)
}
func ParseVMStat(b []byte) VMStat {
var v VMStat
sc := bufio.NewScanner(bytes.NewReader(b))
for sc.Scan() {
f := strings.Fields(sc.Text())
if len(f) != 2 {
continue
}
n, err := strconv.ParseUint(f[1], 10, 64)
if err != nil {
continue
}
switch f[0] {
case "pswpin":
v.PswpIn = n
case "pswpout":
v.PswpOut = n
case "oom_kill":
v.OOMKill, v.HasOOM = n, true
}
}
return v
}
// ParseFileNR reads /proc/sys/fs/file-nr: "allocated unused max". used = allocated - unused.
func ParseFileNR(b []byte) (used, max uint64, err error) {
f := strings.Fields(string(b))
if len(f) != 3 {
return 0, 0, errors.New("file-nr: expected three numbers")
}
var v [3]uint64
for i := range f {
if v[i], err = strconv.ParseUint(f[i], 10, 64); err != nil {
return 0, 0, fmt.Errorf("file-nr: %w", err)
}
}
if v[1] > v[0] || v[2] == 0 {
return 0, 0, errors.New("file-nr: inconsistent numbers")
}
return v[0] - v[1], v[2], nil
}
// ParseUintFile reads a file holding one number (nf_conntrack_count, nf_conntrack_max).
func ParseUintFile(b []byte) (uint64, error) {
return strconv.ParseUint(strings.TrimSpace(string(b)), 10, 64)
}
// DiskStat is one line of /proc/diskstats (counters since boot).
type DiskStat struct {
Name string
Reads, Writes uint64 // completed requests
SectorsRead, SectorsWritten uint64 // 512-byte sectors
ReadMs, WriteMs uint64 // time the requests spent (queued and served)
IOMs uint64 // time the device had at least one request in flight
}
// ParseDiskStats returns the devices for which keep(name) is true.
func ParseDiskStats(b []byte, keep func(string) bool) []DiskStat {
var out []DiskStat
sc := bufio.NewScanner(bytes.NewReader(b))
for sc.Scan() {
f := strings.Fields(sc.Text())
if len(f) < 14 || !keep(f[2]) {
continue
}
var v [11]uint64
ok := true
for i := 0; i < 11; i++ {
n, err := strconv.ParseUint(f[3+i], 10, 64)
if err != nil {
ok = false
break
}
v[i] = n
}
if !ok {
continue
}
// after the name: reads merged sectors ms | writes merged sectors ms | in-flight io_ms weighted_ms
out = append(out, DiskStat{Name: f[2], Reads: v[0], SectorsRead: v[2], ReadMs: v[3], Writes: v[4], SectorsWritten: v[6],
WriteMs: v[7], IOMs: v[9]})
if len(out) >= 256 {
break
}
}
return out
}
func ParseLoadAvg(b []byte) ([3]float64, error) {
var l [3]float64
f := strings.Fields(string(b))
if len(f) < 3 {
return l, errors.New("short loadavg")
}
for i := 0; i < 3; i++ {
v, err := strconv.ParseFloat(f[i], 64)
if err != nil {
return l, err
}
l[i] = v
}
return l, nil
}
func ParseUptime(b []byte) (float64, error) {
f := strings.Fields(string(b))
if len(f) < 1 {
return 0, errors.New("empty uptime")
}
return strconv.ParseFloat(f[0], 64)
}
// NetCounters are the summed byte counters of the interfaces that carry real traffic.
type NetCounters struct{ RX, TX uint64 }
// virtualIface: loopback and the host side of containers / bridges (their traffic is
// counted again on the physical interface).
func virtualIface(name string) bool {
if name == "lo" {
return true
}
for _, p := range []string{"veth", "docker", "br-", "virbr", "cni", "flannel", "cali", "vxlan", "tun", "tap", "kube", "lxc", "vnet"} {
if strings.HasPrefix(name, p) {
return true
}
}
return false
}
func ParseNetDev(b []byte) NetCounters {
var n NetCounters
sc := bufio.NewScanner(bytes.NewReader(b))
for sc.Scan() {
line := sc.Text()
i := strings.IndexByte(line, ':')
if i < 0 {
continue
}
name := strings.TrimSpace(line[:i])
if virtualIface(name) {
continue
}
f := strings.Fields(line[i+1:])
if len(f) < 9 {
continue
}
rx, _ := strconv.ParseUint(f[0], 10, 64)
tx, _ := strconv.ParseUint(f[8], 10, 64)
n.RX += rx
n.TX += tx
}
return n
}
// Socket is one listening socket from /proc/net/{tcp,tcp6,udp,udp6}.
type Socket struct {
Proto string // tcp | udp (the address family is in Addr)
Addr string
Port int
UID int
Inode uint64
}
// Local: bound to a loopback address, so not reachable from outside.
func (s Socket) Local() bool {
ip := net.ParseIP(s.Addr)
return ip != nil && ip.IsLoopback()
}
// TCPStates counts the TCP sockets of /proc/net/tcp and tcp6 by state.
type TCPStates struct {
Listen int `json:"listen"`
Established int `json:"established"`
TimeWait int `json:"time_wait"`
CloseWait int `json:"close_wait"`
SynRecv int `json:"syn_recv"`
Other int `json:"other"`
Truncated bool `json:"truncated,omitempty"` // a table was longer than the agent reads
}
// maxNetLines bounds the work on a server with a very large connection table.
const maxNetLines = 250_000
// ParseNetSockets reads one /proc/net/{tcp,udp}[6] table and returns the listening sockets.
func ParseNetSockets(b []byte, proto string) ([]Socket, error) {
return ScanNetTable(bytes.NewReader(b), proto, 0, nil)
}
// ScanNetTable reads one /proc/net/{tcp,udp}[6] table line by line (on a busy server it can
// hold hundreds of thousands of connections, so it is never read whole) and returns the
// listening sockets: TCP in state LISTEN (0A), UDP unconnected (07) with no remote peer. For TCP
// it also counts every socket by state into st. maxLines 0 means no limit.
func ScanNetTable(r io.Reader, proto string, maxLines int, st *TCPStates) ([]Socket, error) {
var out []Socket
sc := bufio.NewScanner(r)
sc.Buffer(make([]byte, 4096), 64*1024)
first := true
n := 0
for sc.Scan() {
if first { // header
first = false
continue
}
n++
if maxLines > 0 && n > maxLines {
if st != nil {
st.Truncated = true
}
break
}
f := strings.Fields(sc.Text())
if len(f) < 10 {
continue
}
state := f[3]
if proto == "tcp" && st != nil {
switch state {
case "01":
st.Established++
case "03":
st.SynRecv++
case "06":
st.TimeWait++
case "08":
st.CloseWait++
case "0A":
st.Listen++
default:
st.Other++
}
}
if proto == "tcp" && state != "0A" {
continue
}
if proto == "udp" {
if state != "07" {
continue
}
if _, rport, err := splitHexAddr(f[2]); err != nil || rport != 0 {
continue
}
}
addr, port, err := splitHexAddr(f[1])
if err != nil {
return nil, err
}
uid, _ := strconv.Atoi(f[7])
inode, _ := strconv.ParseUint(f[9], 10, 64)
out = append(out, Socket{Proto: proto, Addr: addr, Port: port, UID: uid, Inode: inode})
}
return out, sc.Err()
}
// scanNetFast is ScanNetTable for the sampling loop: the same result, appended to out, with the
// caller's buffer and no allocation for the (many) lines that are not listening sockets.
func scanNetFast(r io.Reader, proto string, maxLines int, st *TCPStates, buf []byte, out []Socket) []Socket {
sc := bufio.NewScanner(r)
sc.Buffer(buf[:0:cap(buf)], 64*1024)
var f [12][]byte
first := true
n := 0
for sc.Scan() {
if first {
first = false
continue
}
n++
if maxLines > 0 && n > maxLines {
if st != nil {
st.Truncated = true
}
break
}
if fieldsBytes(sc.Bytes(), f[:10]) < 10 {
continue
}
state := f[3]
if len(state) != 2 {
continue
}
if proto == "tcp" {
if st != nil {
switch string(state) {
case "01":
st.Established++
case "03":
st.SynRecv++
case "06":
st.TimeWait++
case "08":
st.CloseWait++
case "0A":
st.Listen++
default:
st.Other++
}
}
if string(state) != "0A" {
continue
}
} else {
if string(state) != "07" {
continue
}
if _, rport, err := splitHexAddr(string(f[2])); err != nil || rport != 0 {
continue
}
}
addr, port, err := splitHexAddr(string(f[1]))
if err != nil {
continue
}
uid, _ := parseIntBytes(f[7])
inode, _ := parseUintBytes(f[9])
out = append(out, Socket{Proto: proto, Addr: addr, Port: port, UID: int(uid), Inode: inode})
}
return out
}
// splitHexAddr decodes "0100007F:1F90" (IPv4) or the 32-hex-digit IPv6 form. The kernel
// prints each 32-bit word in host byte order (little endian on amd64 and arm64).
func splitHexAddr(s string) (string, int, error) {
i := strings.IndexByte(s, ':')
if i < 0 {
return "", 0, fmt.Errorf("bad address %q", s)
}
raw, err := hex.DecodeString(s[:i])
if err != nil || (len(raw) != 4 && len(raw) != 16) {
return "", 0, fmt.Errorf("bad address %q", s)
}
port, err := strconv.ParseUint(s[i+1:], 16, 16)
if err != nil {
return "", 0, fmt.Errorf("bad port %q", s)
}
ip := make(net.IP, len(raw))
for w := 0; w < len(raw); w += 4 {
ip[w], ip[w+1], ip[w+2], ip[w+3] = raw[w+3], raw[w+2], raw[w+1], raw[w]
}
return ip.String(), int(port), nil
}
// PidStat is what we use from /proc/<pid>/stat.
type PidStat struct {
Comm string
State byte
PPid int
UTime uint64
STime uint64
StartTime uint64 // clock ticks after boot
RSSPages int64
}
// ParsePidStat handles a comm that contains spaces or parentheses: it runs from the
// first '(' to the last ')'.
func ParsePidStat(b []byte) (PidStat, error) {
var p PidStat
s := string(b)
l, r := strings.IndexByte(s, '('), strings.LastIndexByte(s, ')')
if l < 0 || r < l {
return p, errors.New("bad stat")
}
p.Comm = s[l+1 : r]
f := strings.Fields(s[r+1:])
// f[0]=state (field 3) ... utime is field 14 -> f[11], stime f[12], starttime field 22 -> f[19], rss field 24 -> f[21]
if len(f) < 22 {
return p, errors.New("short stat")
}
p.State = f[0][0]
p.PPid, _ = strconv.Atoi(f[1])
p.UTime, _ = strconv.ParseUint(f[11], 10, 64)
p.STime, _ = strconv.ParseUint(f[12], 10, 64)
p.StartTime, _ = strconv.ParseUint(f[19], 10, 64)
p.RSSPages, _ = strconv.ParseInt(f[21], 10, 64)
return p, nil
}
// parsePidStatInto is ParsePidStat for the hot path: no allocation; comm is a subslice of b.
func parsePidStatInto(b []byte, p *PidStat, f [][]byte) (comm []byte, ok bool) {
l, r := bytes.IndexByte(b, '('), bytes.LastIndexByte(b, ')')
if l < 0 || r < l || len(f) < 22 {
return nil, false
}
comm = b[l+1 : r]
if fieldsBytes(b[r+1:], f[:22]) < 22 {
return nil, false
}
p.State = f[0][0]
pp, _ := parseIntBytes(f[1])
p.PPid = int(pp)
p.UTime, _ = parseUintBytes(f[11])
p.STime, _ = parseUintBytes(f[12])
p.StartTime, _ = parseUintBytes(f[19])
p.RSSPages, _ = parseIntBytes(f[21])
return comm, true
}
// ParseOSRelease returns PRETTY_NAME (or NAME VERSION_ID).
func ParseOSRelease(b []byte) string {
vals := map[string]string{}
sc := bufio.NewScanner(bytes.NewReader(b))
for sc.Scan() {
k, v, ok := strings.Cut(strings.TrimSpace(sc.Text()), "=")
if !ok {
continue
}
vals[k] = strings.Trim(v, `"'`)
}
if v := vals["PRETTY_NAME"]; v != "" {
return v
}
return strings.TrimSpace(vals["NAME"] + " " + vals["VERSION_ID"])
}
// Users from /etc/passwd: uid -> name, plus the accounts that can log in.
type PasswdEntry struct {
Name string
UID int
Home string
Shell string
}
func ParsePasswd(b []byte) []PasswdEntry {
var out []PasswdEntry
sc := bufio.NewScanner(bytes.NewReader(b))
for sc.Scan() {
f := strings.Split(sc.Text(), ":")
if len(f) < 7 || strings.HasPrefix(f[0], "#") {
continue
}
uid, err := strconv.Atoi(f[2])
if err != nil {
continue
}
out = append(out, PasswdEntry{Name: f[0], UID: uid, Home: f[5], Shell: f[6]})
}
return out
}
// LoginShell: an account whose shell lets someone log in.
func LoginShell(shell string) bool {
if shell == "" {
return false
}
base := shell[strings.LastIndexByte(shell, '/')+1:]
switch base {
case "nologin", "false", "sync", "halt", "shutdown", "true":
return false
}
return true
}
func clamp(v, lo, hi float64) float64 {
if v < lo {
return lo
}
if v > hi {
return hi
}
return v
}
func minU(a, b uint64) uint64 {
if a < b {
return a
}
return b
}
fastio.go147 lines
package main
// Cheap reads for the hot path. A sample reads a few small files per process every minute; with
// os.ReadFile each one costs an os.File, an fstat and a fresh buffer. These helpers read into a
// buffer the caller keeps (open, read, close: three system calls, no allocation besides the path)
// and parse numbers straight from the bytes.
import (
"os"
"syscall"
)
// readInto reads the start of a small file (a /proc or /sys entry) into buf and returns the
// number of bytes read. A file longer than buf is cut: callers size buf for what they parse.
func readInto(path string, buf []byte) (int, error) {
fd, err := syscall.Open(path, syscall.O_RDONLY|syscall.O_CLOEXEC, 0)
if err != nil {
return 0, &os.PathError{Op: "open", Path: path, Err: err}
}
n := 0
for n < len(buf) {
m, err := syscall.Read(fd, buf[n:])
if err == syscall.EINTR {
continue
}
if err != nil {
syscall.Close(fd)
return n, &os.PathError{Op: "read", Path: path, Err: err}
}
if m <= 0 {
break
}
n += m
}
syscall.Close(fd)
return n, nil
}
// readlinkInto returns the target of a symbolic link read into buf (no allocation for the result
// beyond the string conversion the caller decides on).
func readlinkInto(path string, buf []byte) (int, error) {
for {
n, err := syscall.Readlink(path, buf)
if err == syscall.EINTR {
continue
}
if err != nil {
return 0, &os.PathError{Op: "readlink", Path: path, Err: err}
}
return n, nil
}
}
// parseUintBytes reads a decimal number; ok is false for an empty field or a non-digit.
func parseUintBytes(b []byte) (uint64, bool) {
if len(b) == 0 {
return 0, false
}
var n uint64
for _, c := range b {
if c < '0' || c > '9' {
return 0, false
}
n = n*10 + uint64(c-'0')
}
return n, true
}
func parseIntBytes(b []byte) (int64, bool) {
neg := len(b) > 0 && b[0] == '-'
if neg {
b = b[1:]
}
u, ok := parseUintBytes(b)
if !ok {
return 0, false
}
if neg {
return -int64(u), true
}
return int64(u), true
}
// fieldsBytes splits b on spaces and tabs into at most len(dst) fields, without allocating: the
// fields are subslices of b. It returns how many fields it found.
func fieldsBytes(b []byte, dst [][]byte) int {
n := 0
i := 0
for i < len(b) && n < len(dst) {
for i < len(b) && (b[i] == ' ' || b[i] == '\t') {
i++
}
if i >= len(b) {
break
}
j := i
for j < len(b) && b[j] != ' ' && b[j] != '\t' {
j++
}
dst[n] = b[i:j]
n++
i = j
}
return n
}
// uitoa writes a non-negative number into a small buffer (for /proc/<pid>/... paths).
func appendPid(b []byte, pid int) []byte {
var tmp [20]byte
i := len(tmp)
if pid == 0 {
return append(b, '0')
}
for pid > 0 {
i--
tmp[i] = byte('0' + pid%10)
pid /= 10
}
return append(b, tmp[i:]...)
}
// readIntoZ is readInto for a NUL-terminated path held in a byte slice (no allocation at all on
// Linux, where the kernel gets the bytes directly).
func readIntoZ(pathZ []byte, buf []byte) (int, error) {
fd, err := openZ(pathZ)
if err != nil {
return 0, err
}
n := 0
for n < len(buf) {
m, err := syscall.Read(fd, buf[n:])
if err == syscall.EINTR {
continue
}
if err != nil {
syscall.Close(fd)
return n, err
}
if m <= 0 {
break
}
n += m
}
syscall.Close(fd)
return n, nil
}
disk.go182 lines
package main
import (
"bufio"
"bytes"
"sort"
"strconv"
"strings"
"syscall"
)
// Mount is one real filesystem from /proc/self/mountinfo.
type Mount struct {
Point string
FSType string
Source string
}
// Filesystems that hold data. Pseudo filesystems, tmpfs, overlays, snap images
// (squashfs, always 100% full) and network filesystems (statfs can hang) are left out.
var realFS = map[string]bool{
"ext2": true, "ext3": true, "ext4": true, "xfs": true, "btrfs": true, "zfs": true, "f2fs": true,
"jfs": true, "reiserfs": true, "vfat": true, "exfat": true, "ntfs": true, "ntfs3": true,
"bcachefs": true, "hfsplus": true, "ufs": true,
}
// unescapeMount decodes the octal escapes mountinfo uses (\040 = space).
func unescapeMount(s string) string {
if !strings.Contains(s, `\`) {
return s
}
var b strings.Builder
for i := 0; i < len(s); i++ {
if s[i] == '\\' && i+3 < len(s) {
if n, err := strconv.ParseUint(s[i+1:i+4], 8, 8); err == nil {
b.WriteByte(byte(n))
i += 3
continue
}
}
b.WriteByte(s[i])
}
return b.String()
}
// ParseMountInfo returns one mount per real filesystem: a device mounted at several
// places (bind mounts, btrfs subvolumes, the read-only view systemd gives the agent)
// is listed once, at its shortest mount point.
func ParseMountInfo(b []byte) []Mount {
best := map[string]Mount{}
sc := bufio.NewScanner(bytes.NewReader(b))
sc.Buffer(make([]byte, 64*1024), 1024*1024)
for sc.Scan() {
line := sc.Text()
pre, post, ok := strings.Cut(line, " - ")
if !ok {
continue
}
pf, qf := strings.Fields(pre), strings.Fields(post)
if len(pf) < 5 || len(qf) < 2 {
continue
}
fstype, source := qf[0], qf[1]
if !realFS[fstype] {
continue
}
point := unescapeMount(pf[4])
if skipMountPoint(point) {
continue
}
key := source
if !strings.HasPrefix(source, "/dev/") { // zfs datasets, odd sources: the device numbers
key = pf[2] + "|" + source
}
cur, seen := best[key]
if !seen || len(point) < len(cur.Point) {
best[key] = Mount{Point: point, FSType: fstype, Source: source}
}
}
out := make([]Mount, 0, len(best))
for _, m := range best {
out = append(out, m)
}
sort.Slice(out, func(i, j int) bool { return out[i].Point < out[j].Point })
return out
}
func skipMountPoint(p string) bool {
for _, pre := range []string{"/proc", "/sys", "/dev", "/run", "/snap", "/var/lib/docker", "/var/lib/containers", "/var/snap"} {
if p == pre || strings.HasPrefix(p, pre+"/") {
return true
}
}
return false
}
// DiskUsage of one mount, in bytes and inodes.
type DiskUsage struct {
Total, Used, Avail uint64
Inodes, InodesFree uint64
UsedPct, InodesPct float64
}
func statDisk(path string) (DiskUsage, error) {
var st syscall.Statfs_t
if err := syscall.Statfs(path, &st); err != nil {
return DiskUsage{}, err
}
bs := uint64(st.Bsize)
d := DiskUsage{Total: st.Blocks * bs, Avail: st.Bavail * bs, Inodes: st.Files, InodesFree: st.Ffree}
free := st.Bfree * bs
if d.Total >= free {
d.Used = d.Total - free
}
// As df: used / (used + available to unprivileged users), so root's reserve counts as full.
if den := d.Used + d.Avail; den > 0 {
d.UsedPct = 100 * float64(d.Used) / float64(den)
}
if d.Inodes > 0 && d.Inodes >= d.InodesFree {
d.InodesPct = 100 * float64(d.Inodes-d.InodesFree) / float64(d.Inodes)
}
return d, nil
}
// FillTracker keeps the used bytes of one mount over the last window (6 h) and projects
// when the disk is full if the trend goes on.
type FillTracker struct {
Window int64 // seconds
ts []int64
used []float64
}
func NewFillTracker(window int64) *FillTracker { return &FillTracker{Window: window} }
func (f *FillTracker) Add(t int64, used uint64) {
f.ts = append(f.ts, t)
f.used = append(f.used, float64(used))
cut := 0
for cut < len(f.ts) && f.ts[cut] < t-f.Window {
cut++
}
if cut > 0 {
f.ts = append(f.ts[:0], f.ts[cut:]...)
f.used = append(f.used[:0], f.used[cut:]...)
}
}
// Rate is the least-squares slope in bytes per hour (0 without enough data: at least
// 30 minutes and 10 samples).
func (f *FillTracker) Rate() (float64, bool) {
n := len(f.ts)
if n < 10 || f.ts[n-1]-f.ts[0] < 1800 {
return 0, false
}
t0 := float64(f.ts[0])
var sx, sy, sxx, sxy float64
for i := range f.ts {
x := (float64(f.ts[i]) - t0) / 3600
y := f.used[i]
sx += x
sy += y
sxx += x * x
sxy += x * y
}
fn := float64(n)
den := fn*sxx - sx*sx
if den == 0 {
return 0, false
}
return (fn*sxy - sx*sy) / den, true
}
// FullInHours projects when `avail` bytes are used up at the current rate. ok=false
// when the disk is not filling (or not fast enough to matter: under 1 MB per hour).
func (f *FillTracker) FullInHours(avail uint64) (float64, bool) {
rate, ok := f.Rate()
if !ok || rate < 1<<20 {
return 0, false
}
return float64(avail) / rate, true
}
diskio.go94 lines
package main
import (
"regexp"
"sort"
"strings"
"time"
)
// IOOut is one block device over the last sample interval.
type IOOut struct {
Dev string `json:"dev"`
Util float64 `json:"util"` // % of the time with at least one request in flight
AwaitMs float64 `json:"await_ms"` // average time a request took, queueing included
RBps float64 `json:"rbps"` // bytes read per second
WBps float64 `json:"wbps"`
}
// IOTracker turns the /proc/diskstats counters into per-interval figures.
type IOTracker struct {
prev map[string]DiskStat
prevT time.Time
}
func NewIOTracker() *IOTracker { return &IOTracker{prev: map[string]DiskStat{}} }
// Update returns the devices that did any I/O since the previous call, busiest first (at most 8).
func (t *IOTracker) Update(now time.Time, stats []DiskStat) []IOOut {
elapsed := now.Sub(t.prevT)
next := make(map[string]DiskStat, len(stats))
var out []IOOut
for _, s := range stats {
next[s.Name] = s
p, ok := t.prev[s.Name]
if !ok || t.prevT.IsZero() || elapsed <= 0 || s.IOMs < p.IOMs || s.Reads < p.Reads || s.Writes < p.Writes {
continue // first reading, or the counters were reset (device re-added)
}
ios := (s.Reads - p.Reads) + (s.Writes - p.Writes)
ms := float64(elapsed) / float64(time.Millisecond)
o := IOOut{Dev: s.Name, Util: round1(clamp(100*float64(s.IOMs-p.IOMs)/ms, 0, 100))}
if ios > 0 && s.ReadMs >= p.ReadMs && s.WriteMs >= p.WriteMs {
o.AwaitMs = round1(float64((s.ReadMs-p.ReadMs)+(s.WriteMs-p.WriteMs)) / float64(ios))
}
sec := elapsed.Seconds()
if s.SectorsRead >= p.SectorsRead {
o.RBps = float64(int64(float64(s.SectorsRead-p.SectorsRead) * 512 / sec))
}
if s.SectorsWritten >= p.SectorsWritten {
o.WBps = float64(int64(float64(s.SectorsWritten-p.SectorsWritten) * 512 / sec))
}
if ios == 0 && o.Util == 0 {
continue
}
out = append(out, o)
}
t.prev, t.prevT = next, now
sort.Slice(out, func(i, j int) bool {
if out[i].Util != out[j].Util {
return out[i].Util > out[j].Util
}
return out[i].Dev < out[j].Dev
})
if len(out) > 8 {
out = out[:8]
}
return out
}
// Devices that are never interesting: loop images, RAM disks, CD drives, floppies.
func pseudoBlock(name string) bool {
for _, p := range []string{"loop", "ram", "zram", "sr", "fd", "nbd"} {
if strings.HasPrefix(name, p) {
return true
}
}
return false
}
var partitionRe = regexp.MustCompile(`^((s|h|v|xv)d[a-z]+\d+|nvme\d+n\d+p\d+|mmcblk\d+p\d+)$`)
// wholeDisk keeps whole devices, not their partitions (whose I/O is counted again on the disk).
// sysBlock is the list of /sys/block; when it could not be read a name pattern decides.
func wholeDisk(sysBlock map[string]bool) func(string) bool {
return func(name string) bool {
if pseudoBlock(name) {
return false
}
if len(sysBlock) > 0 {
return sysBlock[strings.ReplaceAll(name, "/", "!")]
}
return !partitionRe.MatchString(name)
}
}
diskscan.go389 lines
package main
// Where the disk space goes. A few times a day (every 6 hours, jittered, postponed while the
// machine is busy) the background worker walks each real filesystem once (bind mounts of the same
// device are walked once), adding up the space files actually take (allocated blocks, so sparse
// files count for what they use), without following symbolic links and without crossing into
// other mounts. It keeps the size of directories up to 6 levels deep and of files of 16 MiB or
// more, and reports:
// - the biggest directories, drilling down into the big branches (/var → /var/lib →
// /var/lib/docker → …), and the ones that grew most since the previous scan;
// - the largest files and the fastest-growing ones, with a kind (log, db, backup, cache, docker).
// Only paths and sizes leave the machine, never file contents. The walk is paced by the worker's
// gentle mode (slices of ~2 ms, then a pause) and has hard budgets: 20 s of work, 3 million
// entries, 45 minutes; when one is hit the result is marked partial (sizes are then lower bounds).
// `disable: [disk_scan]` switches it off.
import (
"os"
"path/filepath"
"sort"
"strings"
"time"
)
const (
scanRecordDepth = 6 // directories recorded up to this depth below the mount point
scanMaxDepth = 64 // deeper directories are not entered
scanDirMin = 1 << 20 // directories smaller than this are not recorded
scanFileMin = 16 << 20
scanMaxDirs = 20000
scanMaxFiles = 5000
)
type ScanMount struct {
Mount string `json:"mount"`
FS string `json:"fs"`
Used uint64 `json:"used"` // statfs
Scanned int64 `json:"scanned"` // what the walk added up
}
type ScanDir struct {
Path string `json:"path"`
Bytes int64 `json:"bytes"`
Files int64 `json:"files"`
Delta *int64 `json:"delta"` // since the previous scan; null when not measured then
}
type ScanFile struct {
Path string `json:"path"`
Bytes int64 `json:"bytes"`
Delta *int64 `json:"delta"`
Kind string `json:"kind"`
}
type DiskScanOut struct {
At int64 `json:"at"`
PrevAt *int64 `json:"prev_at"`
TookS int64 `json:"took_s"`
WorkMs int64 `json:"work_ms"`
Entries int64 `json:"entries"`
Partial bool `json:"partial"`
Mounts []ScanMount `json:"mounts"`
Dirs []ScanDir `json:"dirs"`
Files []ScanFile `json:"files"`
}
type dirStat struct {
bytes, files int64
depth int
}
// scanAcc is the state of one walk, shared by the platform walkers.
type scanAcc struct {
g *Gentle
dirs map[string]dirStat
files map[string]int64
fileMin int64
entries int64
maxEnt int64
maxWork time.Duration
deadline time.Time
stopped bool
bufs [][]byte
}
// tick counts an entry and paces the walk; it returns false once a budget is used up.
func (a *scanAcc) tick() bool {
if a.stopped {
return false
}
a.entries++
if a.entries >= a.maxEnt {
a.stopped = true
return false
}
if a.entries&63 == 0 {
a.g.Tick()
if a.g.Work() > a.maxWork || time.Now().After(a.deadline) {
a.stopped = true
return false
}
}
return true
}
func (a *scanAcc) dir(path string, depth int, bytes, files int64) {
if depth > scanRecordDepth || bytes < scanDirMin {
return
}
if len(a.dirs) >= scanMaxDirs && depth > 2 {
return
}
a.dirs[path] = dirStat{bytes: bytes, files: files, depth: depth}
}
func (a *scanAcc) file(path string, size int64) {
if size < a.fileMin {
return
}
a.files[path] = size
if len(a.files) >= 2*scanMaxFiles { // keep the larger half, raise the bar
sizes := make([]int64, 0, len(a.files))
for _, s := range a.files {
sizes = append(sizes, s)
}
sort.Slice(sizes, func(i, j int) bool { return sizes[i] > sizes[j] })
a.fileMin = sizes[scanMaxFiles-1]
for p, s := range a.files {
if s < a.fileMin {
delete(a.files, p)
}
}
}
}
// buf returns the directory-listing buffer of one depth (each level keeps its own while the walk
// is below it).
func (a *scanAcc) buf(depth int) []byte {
for len(a.bufs) <= depth {
a.bufs = append(a.bufs, nil)
}
if a.bufs[depth] == nil {
a.bufs[depth] = make([]byte, 8192)
}
return a.bufs[depth]
}
func joinPath(dir, name string) string {
if strings.HasSuffix(dir, "/") {
return dir + name
}
return dir + "/" + name
}
// DiskScanner keeps the previous scan to report what grew.
type DiskScanner struct {
prevDirs map[string]int64
prevFiles map[string]int64
prevAt int64
MaxEntries int64
MaxWork time.Duration
MaxWall time.Duration
}
func NewDiskScanner() *DiskScanner {
return &DiskScanner{MaxEntries: 3_000_000, MaxWork: 20 * time.Second, MaxWall: 45 * time.Minute}
}
// Scan walks the mounts (real filesystems, as the disk list has them) and builds the report.
func (d *DiskScanner) Scan(mounts []Mount, statfs func(string) (DiskUsage, error), g *Gentle, now time.Time) *DiskScanOut {
start := time.Now()
if g == nil {
g = &Gentle{}
}
g.Begin()
acc := &scanAcc{g: g, dirs: map[string]dirStat{}, files: map[string]int64{}, fileMin: scanFileMin, maxEnt: d.MaxEntries,
maxWork: d.MaxWork, deadline: start.Add(d.MaxWall)}
out := &DiskScanOut{At: now.Unix(), Mounts: []ScanMount{}, Dirs: []ScanDir{}, Files: []ScanFile{}}
// One walk per device: bind mounts of an already walked device are skipped (shortest path first).
ms := append([]Mount(nil), mounts...)
sort.Slice(ms, func(i, j int) bool { return len(ms[i].Point) < len(ms[j].Point) })
seenDev := map[uint64]bool{}
var scanned []ScanMount
for _, m := range ms {
if acc.stopped || len(scanned) >= 16 {
break
}
dev, ok := devOf(m.Point)
if !ok || seenDev[dev] {
continue
}
if fi, err := os.Stat(m.Point); err != nil || !fi.IsDir() { // a bind-mounted file (/etc/hosts in a container)
continue
}
seenDev[dev] = true
bytes, files := walkTree(m.Point, dev, acc)
acc.dirs[m.Point] = dirStat{bytes: bytes, files: files, depth: 0}
sm := ScanMount{Mount: m.Point, FS: m.FSType, Scanned: bytes}
if du, err := statfs(m.Point); err == nil {
sm.Used = du.Used
}
scanned = append(scanned, sm)
}
out.Mounts = scanned
out.Entries = acc.entries
out.Partial = acc.stopped
out.WorkMs = g.Work().Milliseconds()
out.TookS = int64(time.Since(start).Seconds())
if d.prevAt > 0 && !out.Partial {
p := d.prevAt
out.PrevAt = &p
}
if out.Partial { // sizes are lower bounds: no growth is computed from them, and they are not kept as the reference
saved, savedF := d.prevDirs, d.prevFiles
d.prevDirs, d.prevFiles = nil, nil
out.Dirs = d.pickDirs(acc.dirs, scanned)
out.Files = d.pickFiles(acc.files)
d.prevDirs, d.prevFiles = saved, savedF
return out
}
out.Dirs = d.pickDirs(acc.dirs, scanned)
out.Files = d.pickFiles(acc.files)
nd := make(map[string]int64, len(acc.dirs))
for p, s := range acc.dirs {
nd[p] = s.bytes
}
d.prevDirs, d.prevFiles, d.prevAt = nd, acc.files, now.Unix()
return out
}
func (d *DiskScanner) delta(prev map[string]int64, path string, now int64) *int64 {
if prev == nil {
return nil
}
p, ok := prev[path]
if !ok {
return nil
}
v := now - p
return &v
}
// pickDirs: the mount points, then their big children drilled down (at most 8 per parent, each at
// least 64 MiB and either 2% of the mount or a quarter of its parent), then the growers.
func (d *DiskScanner) pickDirs(dirs map[string]dirStat, mounts []ScanMount) []ScanDir {
kids := map[string][]string{}
for p := range dirs {
par := filepath.Dir(p)
if par != p {
kids[par] = append(kids[par], p)
}
}
sel := map[string]bool{}
var order []string
add := func(p string) {
if !sel[p] && len(order) < 60 {
sel[p] = true
order = append(order, p)
}
}
type item struct {
path string
total int64
}
var queue []item
for _, m := range mounts {
add(m.Mount)
queue = append(queue, item{m.Mount, m.Scanned})
}
for len(queue) > 0 && len(order) < 45 {
it := queue[0]
queue = queue[1:]
ch := kids[it.path]
sort.Slice(ch, func(i, j int) bool { return dirs[ch[i]].bytes > dirs[ch[j]].bytes })
parent := dirs[it.path].bytes
for i, c := range ch {
b := dirs[c].bytes
if i >= 8 || b < 64<<20 || (b*50 < it.total && b*4 < parent) {
break
}
add(c)
queue = append(queue, item{c, it.total})
}
}
// growers since the previous scan
type grow struct {
path string
delta int64
}
var gs []grow
for p, s := range dirs {
if prev, ok := d.prevDirs[p]; ok && s.bytes-prev >= 32<<20 {
gs = append(gs, grow{p, s.bytes - prev})
}
}
sort.Slice(gs, func(i, j int) bool { return gs[i].delta > gs[j].delta })
for i, g := range gs {
if i >= 15 {
break
}
add(g.path)
}
out := make([]ScanDir, 0, len(order))
for _, p := range order {
s := dirs[p]
out = append(out, ScanDir{Path: truncate(p, 300), Bytes: s.bytes, Files: s.files, Delta: d.delta(d.prevDirs, p, s.bytes)})
}
return out
}
// pickFiles: the 15 largest and the 15 fastest-growing (by at least 16 MiB) files.
func (d *DiskScanner) pickFiles(files map[string]int64) []ScanFile {
type f struct {
path string
size int64
delta int64
}
all := make([]f, 0, len(files))
for p, s := range files {
x := f{path: p, size: s}
if prev, ok := d.prevFiles[p]; ok {
x.delta = s - prev
} else if d.prevFiles != nil {
x.delta = s // new since the previous scan (or just crossed 16 MiB)
}
all = append(all, x)
}
sel := map[string]bool{}
var out []ScanFile
add := func(x f) {
if sel[x.path] {
return
}
sel[x.path] = true
out = append(out, ScanFile{Path: truncate(x.path, 300), Bytes: x.size, Delta: d.delta(d.prevFiles, x.path, x.size), Kind: FileKind(x.path)})
}
sort.Slice(all, func(i, j int) bool { return all[i].size > all[j].size })
for i := 0; i < len(all) && i < 15; i++ {
add(all[i])
}
sort.Slice(all, func(i, j int) bool { return all[i].delta > all[j].delta })
for i := 0; i < len(all) && i < 15 && all[i].delta >= 16<<20; i++ {
add(all[i])
}
if out == nil {
out = []ScanFile{}
}
return out
}
// FileKind classifies a big file by its path.
func FileKind(p string) string {
lp := strings.ToLower(p)
base := filepath.Base(lp)
has := func(xs ...string) bool {
for _, x := range xs {
if strings.Contains(lp, x) {
return true
}
}
return false
}
suffix := func(xs ...string) bool {
for _, x := range xs {
if strings.HasSuffix(base, x) {
return true
}
}
return false
}
switch {
case strings.HasSuffix(base, "-json.log") || strings.HasPrefix(lp, "/var/log/") || suffix(".log", ".journal", ".journal~") ||
strings.Contains(base, ".log.") || has("/log/", "/logs/"):
return "log"
case has("/var/lib/docker/", "/var/lib/containerd/", "/var/lib/containers/"):
return "docker"
case has("/var/lib/mysql/", "/var/lib/postgresql/", "/var/lib/pgsql/", "/var/lib/mongodb/", "/var/lib/redis/", "/var/lib/clickhouse/",
"/var/lib/elasticsearch/", "/var/lib/influxdb/") || suffix(".sqlite", ".sqlite3", ".db", ".ibd", ".rdb", ".aof", ".mdb", ".wal") ||
strings.HasPrefix(base, "ibdata") || strings.HasPrefix(base, "binlog.") || strings.HasPrefix(base, "mysql-bin."):
return "db"
case suffix(".bak", ".backup", ".tar", ".tar.gz", ".tgz", ".tar.zst", ".tar.xz", ".zip", ".7z", ".rar", ".sql", ".sql.gz", ".dump", ".gz", ".xz", ".zst", ".img", ".iso") ||
has("backup", "/bak/", "snapshot"):
return "backup"
case has("/.cache/", "/var/cache/", "/tmp/", "/var/tmp/", "/cache/"):
return "cache"
}
return "other"
}
walk_linux.go121 lines
//go:build linux
package main
// The Linux directory walker: openat / getdents64 / fstatat on file descriptors, with the entry
// names passed to the kernel straight from the getdents buffer (they are NUL-terminated there), so
// a file costs one fstatat and no allocation. Directories are opened with O_NOFOLLOW (a symbolic
// link is never followed) and a directory on another device (a mount point) is not entered.
import (
"syscall"
"unsafe"
)
const (
atSymlinkNofollow = 0x100
dtUnknown = 0
dtDir = 4
dtReg = 8
openDirFlags = syscall.O_RDONLY | syscall.O_DIRECTORY | syscall.O_NOFOLLOW | syscall.O_CLOEXEC | syscall.O_NONBLOCK
)
func devOf(path string) (uint64, bool) {
var st syscall.Stat_t
if err := syscall.Stat(path, &st); err != nil {
return 0, false
}
return uint64(st.Dev), true
}
func fstatatRaw(dirfd int, name *byte, st *syscall.Stat_t) syscall.Errno {
_, _, e := syscall.Syscall6(sysFstatat, uintptr(dirfd), uintptr(unsafe.Pointer(name)), uintptr(unsafe.Pointer(st)), atSymlinkNofollow, 0, 0)
return e
}
func openatRaw(dirfd int, name *byte) (int, syscall.Errno) {
fd, _, e := syscall.Syscall6(syscall.SYS_OPENAT, uintptr(dirfd), uintptr(unsafe.Pointer(name)), openDirFlags, 0, 0, 0)
return int(fd), e
}
// walkTree returns the allocated bytes and the number of files below root (its own device only).
func walkTree(root string, dev uint64, acc *scanAcc) (int64, int64) {
fd, err := syscall.Open(root, openDirFlags, 0)
if err != nil {
return 0, 0
}
return walkFd(fd, root, 0, dev, acc)
}
func walkFd(fd int, path string, depth int, dev uint64, acc *scanAcc) (bytes, files int64) {
defer syscall.Close(fd)
buf := acc.buf(depth)
var st syscall.Stat_t
for !acc.stopped {
n, err := syscall.ReadDirent(fd, buf)
if err == syscall.EINTR {
continue
}
if err != nil || n <= 0 {
break
}
for off := 0; off < n; {
// struct linux_dirent64 { u64 ino; s64 off; u16 reclen; u8 type; char name[]; }
if off+19 > n {
break
}
reclen := int(*(*uint16)(unsafe.Pointer(&buf[off+16])))
typ := buf[off+18]
if reclen <= 0 || off+reclen > n {
break
}
nameStart, recEnd := off+19, off+reclen
nameEnd := nameStart
for nameEnd < recEnd && buf[nameEnd] != 0 {
nameEnd++
}
off = recEnd
if nameEnd == nameStart || nameEnd >= recEnd { // empty, or not NUL-terminated inside its record
continue
}
name := buf[nameStart:nameEnd]
namePtr := &buf[nameStart]
if (len(name) == 1 && name[0] == '.') || (len(name) == 2 && name[0] == '.' && name[1] == '.') {
continue
}
if !acc.tick() {
break
}
if typ != dtDir && typ != dtReg && typ != dtUnknown {
continue // symbolic links, sockets, devices: no space worth counting
}
if fstatatRaw(fd, namePtr, &st) != 0 {
continue
}
switch st.Mode & syscall.S_IFMT {
case syscall.S_IFREG:
sz := st.Blocks * 512
bytes += sz
files++
if sz >= acc.fileMin {
acc.file(joinPath(path, string(name)), sz)
}
case syscall.S_IFDIR:
if uint64(st.Dev) != dev || depth+1 > scanMaxDepth {
continue // another filesystem (a mount point) or too deep
}
bytes += st.Blocks * 512
cfd, e := openatRaw(fd, namePtr)
if e != 0 {
continue
}
cb, cf := walkFd(cfd, joinPath(path, string(name)), depth+1, dev, acc)
bytes += cb
files += cf
}
}
}
acc.dir(path, depth, bytes, files)
return bytes, files
}
writers.go126 lines
package main
// Who writes to the disk: the bytes each process caused to be written to storage (write_bytes in
// /proc/<pid>/io), sampled once per rollup by the background worker. The difference since the
// previous look is summed per program (name, systemd unit, container) and the ten biggest writers
// go into the rollup. Other users' /proc/<pid>/io needs the detailed mode (PTRACE_MODE_READ).
import (
"bytes"
"sort"
"strconv"
"time"
)
type WriterOut struct {
Name string `json:"name"`
Unit string `json:"unit,omitempty"`
Container string `json:"container,omitempty"`
Bytes int64 `json:"bytes"`
}
type wPrev struct {
start uint64
wb uint64
}
type WriteTracker struct {
prev map[int]wPrev
prevAt time.Time
buf []byte
}
func NewWriteTracker() *WriteTracker {
return &WriteTracker{prev: map[int]wPrev{}, buf: make([]byte, 512)}
}
// parseWriteBytes reads write_bytes from /proc/<pid>/io.
func parseWriteBytes(b []byte) (uint64, bool) {
i := bytes.Index(b, []byte("\nwrite_bytes: "))
if i < 0 {
if !bytes.HasPrefix(b, []byte("write_bytes: ")) {
return 0, false
}
i = -1
}
rest := b[i+1+len("write_bytes: "):]
if j := bytes.IndexByte(rest, '\n'); j >= 0 {
rest = rest[:j]
}
return parseUintBytes(rest)
}
// Sample returns the top writers since the previous call (nil on the first one). bootTime turns
// the processes' start ticks into time: a process started since the previous look counts whole.
func (w *WriteTracker) Sample(procRoot string, procs []procRef, bootTime int64, users *UserCache, docker *DockerIndex, g *Gentle, now time.Time) []WriterOut {
type d struct {
pid int
delta uint64
}
var deltas []d
next := make(map[int]wPrev, len(procs))
first := w.prevAt.IsZero()
prevTicks := uint64(0)
if !first && bootTime > 0 {
prevTicks = uint64(max(0, w.prevAt.Unix()-bootTime)) * clockTicks
}
for _, p := range procs {
if p.Kernel || p.PID <= 0 {
continue
}
g.Tick()
n, err := readInto(procRoot+"/"+strconv.Itoa(p.PID)+"/io", w.buf)
if err != nil {
continue
}
wb, ok := parseWriteBytes(w.buf[:n])
if !ok {
continue
}
next[p.PID] = wPrev{start: p.Start, wb: wb}
if first {
continue
}
if pp, ok := w.prev[p.PID]; ok && pp.start == p.Start {
if wb > pp.wb {
deltas = append(deltas, d{p.PID, wb - pp.wb})
}
} else if p.Start >= prevTicks && wb > 0 { // started since the previous look
deltas = append(deltas, d{p.PID, wb})
}
}
w.prev, w.prevAt = next, now
if first {
return nil
}
sort.Slice(deltas, func(i, j int) bool { return deltas[i].delta > deltas[j].delta })
if len(deltas) > 25 {
deltas = deltas[:25]
}
agg := map[string]*WriterOut{}
var order []string
for _, x := range deltas {
o := describeOwner(procRoot, x.pid, users, docker)
if o.Name == "" {
continue
}
k := o.Name + "|" + o.Unit + "|" + o.Container
a := agg[k]
if a == nil {
a = &WriterOut{Name: o.Name, Unit: o.Unit, Container: o.Container}
agg[k] = a
order = append(order, k)
}
a.Bytes += int64(x.delta)
}
out := make([]WriterOut, 0, len(order))
for _, k := range order {
out = append(out, *agg[k])
}
sort.Slice(out, func(i, j int) bool { return out[i].Bytes > out[j].Bytes })
if len(out) > 10 {
out = out[:10]
}
return out
}
prio_linux.go25 lines
//go:build linux
package main
import (
"runtime"
"syscall"
)
const (
ioprioWhoProcess = 1
ioprioClassIdle = 3
ioprioClassShift = 13
)
// lowerThreadPriority pins the calling goroutine to its OS thread and gives that thread the lowest
// CPU priority (nice 19) and the idle I/O class: the kernel runs it only when nothing else wants
// the CPU or the disk. Errors are ignored (the systemd unit sets the same for the whole process).
func lowerThreadPriority() {
runtime.LockOSThread()
tid := syscall.Gettid()
_ = syscall.Setpriority(syscall.PRIO_PROCESS, tid, 19)
_, _, _ = syscall.Syscall(sysIoprioSet, ioprioWhoProcess, uintptr(tid), ioprioClassIdle<<ioprioClassShift)
}
sender.go432 lines
package main
import (
"bytes"
"compress/gzip"
"context"
"crypto/tls"
"encoding/json"
"errors"
"fmt"
"io"
"log/slog"
"math/rand"
"net"
"net/http"
"os"
"path/filepath"
"sort"
"strconv"
"strings"
"sync"
"sync/atomic"
"time"
)
// Offline buffering. Undelivered messages wait in the spool (memory, plus one file each in
// <state_dir>/spool so they survive a restart) and go out oldest first once the server is
// reachable again. The spool is bounded in age, count and bytes:
// - events, inventories and rollups are kept for 24 hours (an incident or a security signal
// raised during an outage still arrives, and the charts get the outage's minutes);
// - only the newest heartbeat is kept (an old one says nothing the newer one does not);
// - at most spoolMaxItems messages and spoolMaxBytes bytes: past that the oldest heartbeats and
// rollups go first, events last.
var (
spoolMaxAge = 24 * time.Hour
spoolMaxAgeLegacy = time.Hour // 0.1.0 spool files (no message type in the name)
spoolMaxItems = 1500
spoolMaxBytes = 8 << 20
backoffMin = 15 * time.Second
backoffMax = 5 * time.Minute
)
// Sender posts reports (gzip JSON). It only ever sends: the answer's body is discarded and the
// status code is the only thing read from it.
type Sender struct {
cfg Config
version string
client *http.Client
// Dry, when set, receives each message as indented JSON instead of the network (run --dry-run).
Dry io.Writer
// OnResync is called when the server answers 205 Reset Content: its copy of the inventory
// differs from ours, so the agent sends the full lists once.
OnResync func()
mu sync.Mutex
queue []spooled // oldest first
bytes int
nextID uint64
backoff time.Duration
nextTry time.Time
rnd *rand.Rand
wake chan struct{}
flushing sync.Mutex
sentBytes atomic.Int64
sentMsgs atomic.Int64
dropped atomic.Int64
}
type spooled struct {
id uint64
at time.Time
typ string // heartbeat | rollup | event | inventory | "" (a 0.1.0 spool file)
body []byte // gzip JSON
file string
}
func NewSender(cfg Config, version string) *Sender {
tr := &http.Transport{
Proxy: http.ProxyFromEnvironment, // HTTPS_PROXY in the unit's environment, if any
DialContext: (&net.Dialer{Timeout: 10 * time.Second, KeepAlive: 30 * time.Second}).DialContext,
TLSClientConfig: &tls.Config{MinVersion: tls.VersionTLS12},
TLSHandshakeTimeout: 10 * time.Second,
ResponseHeaderTimeout: 20 * time.Second,
MaxIdleConns: 2,
IdleConnTimeout: 90 * time.Second,
ForceAttemptHTTP2: true,
}
return &Sender{cfg: cfg, version: version, wake: make(chan struct{}, 1), rnd: rand.New(rand.NewSource(time.Now().UnixNano())),
client: &http.Client{
Transport: tr,
Timeout: 30 * time.Second,
// Redirects are not followed: the report goes to the configured URL or nowhere.
CheckRedirect: func(*http.Request, []*http.Request) error { return http.ErrUseLastResponse },
}}
}
func (s *Sender) spoolDir() string { return filepath.Join(s.cfg.StateDir, "spool") }
func maxAge(typ string) time.Duration {
if typ == "" || typ == "heartbeat" {
return spoolMaxAgeLegacy
}
return spoolMaxAge
}
// spoolName: "<unix nanoseconds>-<type>.json.gz". 0.1.0 wrote "<unix nanoseconds>.json.gz".
func spoolName(at time.Time, typ string) string {
return strconv.FormatInt(at.UnixNano(), 10) + "-" + typ + ".json.gz"
}
func parseSpoolName(name string) (time.Time, string, bool) {
base, ok := strings.CutSuffix(name, ".json.gz")
if !ok {
return time.Time{}, "", false
}
ts, typ, _ := strings.Cut(base, "-")
n, err := strconv.ParseInt(ts, 10, 64)
if err != nil || n <= 0 {
return time.Time{}, "", false
}
switch typ {
case "", "heartbeat", "rollup", "event", "inventory":
return time.Unix(0, n), typ, true
}
return time.Time{}, "", false
}
// LoadSpool picks up reports a previous run could not send.
func (s *Sender) LoadSpool() {
if s.Dry != nil {
return
}
ents, err := os.ReadDir(s.spoolDir())
if err != nil {
return
}
s.mu.Lock()
defer s.mu.Unlock()
now := time.Now()
for i, e := range ents {
p := filepath.Join(s.spoolDir(), e.Name())
at, typ, ok := parseSpoolName(e.Name())
if !ok || !e.Type().IsRegular() || now.Sub(at) > maxAge(typ) || i >= 4*spoolMaxItems {
_ = os.Remove(p) // half-written (.tmp), unknown or expired
continue
}
b, err := readSmall(p, int64(spoolMaxBytes))
if err != nil || len(b) < 2 || b[0] != 0x1f || b[1] != 0x8b {
_ = os.Remove(p)
continue
}
s.nextID++
s.queue = append(s.queue, spooled{id: s.nextID, at: at, typ: typ, body: b, file: p})
s.bytes += len(b)
}
sort.SliceStable(s.queue, func(i, j int) bool { return s.queue[i].at.Before(s.queue[j].at) })
s.trimLocked(now)
}
func (s *Sender) Pending() int {
s.mu.Lock()
defer s.mu.Unlock()
return len(s.queue)
}
func Encode(r any) ([]byte, error) {
raw, err := json.Marshal(r)
if err != nil {
return nil, err
}
var buf bytes.Buffer
buf.Grow(len(raw)/3 + 64)
// The compressor's state is about 1 MB: it is reused, not allocated for every message.
zw, _ := gzipPool.Get().(*gzip.Writer)
if zw == nil {
zw, _ = gzip.NewWriterLevel(&buf, gzip.BestCompression)
} else {
zw.Reset(&buf)
}
defer gzipPool.Put(zw)
if _, err := zw.Write(raw); err != nil {
return nil, err
}
if err := zw.Close(); err != nil {
return nil, err
}
return buf.Bytes(), nil
}
var gzipPool sync.Pool
func msgKind(r any) string {
if k, ok := r.(interface{ kind() string }); ok {
return k.kind()
}
return "event"
}
// Enqueue adds a message to the spool and wakes the delivery loop. It never blocks on the network.
func (s *Sender) Enqueue(r any) {
typ := msgKind(r)
body, err := Encode(r)
if err != nil {
slog.Error("could not encode a report", "type", typ, "err", err)
return
}
if s.Dry != nil {
pretty, _ := json.MarshalIndent(r, "", " ")
fmt.Fprintf(s.Dry, "# %s message, %d bytes gzip (%d bytes JSON), not sent (--dry-run)\n%s\n", typ, len(body), len(pretty), pretty)
s.sentMsgs.Add(1)
s.sentBytes.Add(int64(len(body)))
return
}
now := time.Now()
s.mu.Lock()
s.nextID++
item := spooled{id: s.nextID, at: now, typ: typ, body: body}
s.mu.Unlock()
if err := os.MkdirAll(s.spoolDir(), 0o700); err == nil {
p := filepath.Join(s.spoolDir(), spoolName(now, typ))
if err := writeFileAtomic(p, body, 0o600); err == nil {
item.file = p
} else {
logLimited("spool-write", time.Hour, slog.LevelWarn, "cannot write the spool; undelivered messages are kept in memory only", "err", err)
}
}
s.mu.Lock()
if typ == "heartbeat" { // superseded
kept := s.queue[:0]
for _, q := range s.queue {
if q.typ == "heartbeat" {
s.dropLocked(q, false)
continue
}
kept = append(kept, q)
}
s.queue = kept
}
s.queue = append(s.queue, item)
s.bytes += len(body)
s.trimLocked(now)
s.mu.Unlock()
select {
case s.wake <- struct{}{}:
default:
}
}
// trimLocked drops expired messages, then the least useful ones while the spool is over its caps.
func (s *Sender) trimLocked(now time.Time) {
kept := s.queue[:0]
for _, q := range s.queue {
if now.Sub(q.at) > maxAge(q.typ) {
s.dropLocked(q, true)
continue
}
kept = append(kept, q)
}
s.queue = kept
for len(s.queue) > 1 && (len(s.queue) > spoolMaxItems || s.bytes > spoolMaxBytes) {
victim := 0 // the oldest message, unless an older-than-events heartbeat or rollup can go first
for i, q := range s.queue {
if q.typ == "heartbeat" || q.typ == "rollup" || q.typ == "" {
victim = i
break
}
}
s.dropLocked(s.queue[victim], true)
s.queue = append(s.queue[:victim], s.queue[victim+1:]...)
}
}
func (s *Sender) dropLocked(q spooled, count bool) {
if q.file != "" {
_ = os.Remove(q.file)
}
s.bytes -= len(q.body)
if count {
s.dropped.Add(1)
logLimited("spool-drop", 15*time.Minute, slog.LevelWarn, "dropped an undelivered message (spool limit)", "type", q.typ,
"age", time.Since(q.at).Round(time.Second).String(), "queued", len(s.queue))
}
}
func (s *Sender) remove(id uint64) {
s.mu.Lock()
defer s.mu.Unlock()
for i, q := range s.queue {
if q.id == id {
s.dropLocked(q, false)
s.queue = append(s.queue[:i], s.queue[i+1:]...)
return
}
}
}
// Run delivers the spool until ctx ends: at once when a message is queued, and after a failure
// with exponential backoff (15 s doubling to 5 minutes, with jitter so that a fleet of agents does
// not retry in step when approvalens.com comes back).
func (s *Sender) Run(ctx context.Context) {
for {
s.mu.Lock()
wait := time.Until(s.nextTry)
s.mu.Unlock()
if wait > 0 {
t := time.NewTimer(wait)
select {
case <-ctx.Done():
t.Stop()
return
case <-t.C:
}
}
if !safeBool("send", func() bool { return s.Flush(ctx) }) {
continue
}
select {
case <-ctx.Done():
return
case <-s.wake:
}
}
}
// FinalFlush tries once more on shutdown (unless the server was unreachable a moment ago), for at
// most `timeout`. What is not delivered stays in the spool for the next start.
func (s *Sender) FinalFlush(timeout time.Duration) {
s.mu.Lock()
backingOff := time.Now().Before(s.nextTry)
s.mu.Unlock()
if backingOff || s.Pending() == 0 {
return
}
ctx, cancel := context.WithTimeout(context.Background(), timeout)
defer cancel()
s.Flush(ctx)
}
var (
errRetry = errors.New("retry later")
errAuth = errors.New("token rejected")
errRefused = errors.New("report refused")
)
// Flush sends the queue in order. It returns false when it stopped on a temporary failure (network
// error, 429, 5xx); the next attempt then waits for the backoff.
func (s *Sender) Flush(ctx context.Context) bool {
s.flushing.Lock()
defer s.flushing.Unlock()
for {
s.mu.Lock()
s.trimLocked(time.Now())
if len(s.queue) == 0 {
s.mu.Unlock()
return true
}
item := s.queue[0]
s.mu.Unlock()
err := s.post(ctx, item.body)
switch {
case err == nil:
s.remove(item.id)
s.mu.Lock()
s.backoff, s.nextTry = 0, time.Time{}
s.mu.Unlock()
case errors.Is(err, errRetry):
if ctx.Err() != nil { // shutting down: not the server's fault, no backoff
return false
}
d := s.fail()
logLimited("send-retry", 15*time.Minute, slog.LevelWarn, "report not delivered, kept in the spool", "queued", s.Pending(),
"retry_in", d.Round(time.Second).String(), "err", err)
return false
case errors.Is(err, errAuth):
s.remove(item.id)
s.fail()
logLimited("send-auth", 15*time.Minute, slog.LevelError, "the server rejected the token: was the server deleted from the account? "+
"Check the token in "+s.cfg.path, "err", err)
return false
default: // refused (400 / 413): this message will never be accepted
s.remove(item.id)
logLimited("send-refused", 15*time.Minute, slog.LevelWarn, "report refused and dropped", "type", item.typ, "err", err)
}
}
}
// fail doubles the backoff and sets the next attempt (half the backoff plus a random part).
func (s *Sender) fail() time.Duration {
s.mu.Lock()
defer s.mu.Unlock()
s.backoff = min(max(2*s.backoff, backoffMin), backoffMax)
d := s.backoff/2 + time.Duration(s.rnd.Int63n(int64(s.backoff/2)+1))
s.nextTry = time.Now().Add(d)
return d
}
func (s *Sender) post(ctx context.Context, body []byte) error {
req, err := http.NewRequestWithContext(ctx, http.MethodPost, s.cfg.Endpoint(), bytes.NewReader(body))
if err != nil {
return err
}
req.Header.Set("Authorization", "Bearer "+s.cfg.Token)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Content-Encoding", "gzip")
req.Header.Set("User-Agent", "approvalens-agent/"+s.version)
res, err := s.client.Do(req)
if err != nil {
return fmt.Errorf("%w: %v", errRetry, err)
}
_, _ = io.Copy(io.Discard, io.LimitReader(res.Body, 4096))
res.Body.Close()
switch {
case res.StatusCode >= 200 && res.StatusCode < 300:
s.sentBytes.Add(int64(len(body)))
s.sentMsgs.Add(1)
if res.StatusCode == http.StatusResetContent && s.OnResync != nil {
s.OnResync()
}
return nil
case res.StatusCode == 401 || res.StatusCode == 403:
return fmt.Errorf("%w (HTTP %d)", errAuth, res.StatusCode)
case res.StatusCode == 413 || res.StatusCode == 400:
return fmt.Errorf("%w (HTTP %d)", errRefused, res.StatusCode)
default:
return fmt.Errorf("%w: HTTP %d", errRetry, res.StatusCode)
}
}
// Sent returns the messages and bytes (compressed, as posted) delivered since start.
func (s *Sender) Sent() (msgs, bytes int64) { return s.sentMsgs.Load(), s.sentBytes.Load() }
log.go153 lines
package main
// Logging. The agent writes one line per log record to stderr, which systemd hands to journald.
// Under systemd (JOURNAL_STREAM is set) each line starts with a syslog priority, "<3>" for errors,
// "<4>" for warnings, "<6>" for information, so `journalctl -u approvalens-agent -p warning`
// works, and carries no timestamp (journald adds its own). The rest is logfmt:
//
// level=warn msg="report not delivered, kept in the spool" queued=3 err="HTTP 503"
import (
"bytes"
"context"
"io"
"log/slog"
"os"
"strconv"
"sync"
"time"
"unicode"
)
type logHandler struct {
mu *sync.Mutex
w io.Writer
journal bool
min slog.Level
attrs []slog.Attr
}
func newLogger(w io.Writer, journal bool, min slog.Level) *slog.Logger {
return slog.New(&logHandler{mu: &sync.Mutex{}, w: w, journal: journal, min: min})
}
// setupLogging installs the default logger: journald format under systemd, plain timestamps otherwise.
func setupLogging() {
_, underSystemd := os.LookupEnv("JOURNAL_STREAM")
min := slog.LevelInfo
if os.Getenv("APPROVALENS_AGENT_DEBUG") == "1" {
min = slog.LevelDebug
}
slog.SetDefault(newLogger(os.Stderr, underSystemd, min))
}
func (h *logHandler) Enabled(_ context.Context, l slog.Level) bool { return l >= h.min }
func (h *logHandler) WithAttrs(as []slog.Attr) slog.Handler {
c := *h
c.attrs = append(append([]slog.Attr(nil), h.attrs...), as...)
return &c
}
func (h *logHandler) WithGroup(string) slog.Handler { return h } // groups are not used
func (h *logHandler) Handle(_ context.Context, r slog.Record) error {
var b bytes.Buffer
if h.journal {
b.WriteString(syslogPrefix(r.Level))
} else {
b.WriteString(r.Time.UTC().Format(time.RFC3339))
b.WriteByte(' ')
}
b.WriteString("level=")
b.WriteString(levelName(r.Level))
b.WriteString(" msg=")
b.WriteString(logValue(r.Message))
for _, a := range h.attrs {
writeAttr(&b, a)
}
r.Attrs(func(a slog.Attr) bool {
writeAttr(&b, a)
return true
})
b.WriteByte('\n')
h.mu.Lock()
defer h.mu.Unlock()
_, err := h.w.Write(b.Bytes())
return err
}
func writeAttr(b *bytes.Buffer, a slog.Attr) {
if a.Key == "" {
return
}
b.WriteByte(' ')
b.WriteString(a.Key)
b.WriteByte('=')
b.WriteString(logValue(a.Value.Resolve().String()))
}
func syslogPrefix(l slog.Level) string {
switch {
case l >= slog.LevelError:
return "<3>"
case l >= slog.LevelWarn:
return "<4>"
case l >= slog.LevelInfo:
return "<6>"
default:
return "<7>"
}
}
func levelName(l slog.Level) string {
switch {
case l >= slog.LevelError:
return "error"
case l >= slog.LevelWarn:
return "warn"
case l >= slog.LevelInfo:
return "info"
default:
return "debug"
}
}
// logValue quotes a value when it has spaces, quotes, '=' or control characters (one record is
// always one line), and cuts it to 2000 bytes.
func logValue(s string) string {
s = truncate(s, 2000)
if s == "" {
return `""`
}
for _, r := range s {
if r <= ' ' || r == '"' || r == '=' || r == '\\' || unicode.IsControl(r) || r == 0x7f {
return strconv.Quote(s)
}
}
return s
}
// logLimited logs a record at most once per `every` for a given key, so a server that stays
// offline for a day does not fill the journal with the same line.
var (
limitedMu sync.Mutex
limitedLast = map[string]time.Time{}
)
func logLimited(key string, every time.Duration, level slog.Level, msg string, args ...any) {
limitedMu.Lock()
last, seen := limitedLast[key]
now := time.Now()
if seen && now.Sub(last) < every {
limitedMu.Unlock()
return
}
if len(limitedLast) > 1000 {
limitedLast = map[string]time.Time{}
}
limitedLast[key] = now
limitedMu.Unlock()
slog.Log(context.Background(), level, msg, args...)
}
signatures.txt129 lines
# Heuristic signatures for the Approvalens server agent (embedded into the binary at build time).
# This is NOT an antivirus database: it names a few things that are almost never legitimate on a
# web server, so a match is a reason to look, not proof.
#
# name <word> a process whose name (comm, or the base name of argv[0]) equals <word> (case-insensitive)
# arg <text> a process whose command line contains <text> (case-insensitive)
# busy <word> a process name allowed to use a whole core for a long time (no "sustained CPU" signal)
#
# Crypto miners and the droppers that start them
name xmrig
name xmrig-notls
name xmr-stak
name xmr-stak-cpu
name xmr-stak-rx
name minerd
name cpuminer
name cpuminer-multi
name cgminer
name bfgminer
name ccminer
name nheqminer
name ethminer
name t-rex
name nbminer
name lolminer
name srbminer
name srbminer-multi
name teamredminer
name gminer
name kdevtmpfsi
name kinsing
name kthreaddi
name kthreaddk
name sysrv
name sysrv-hello
name sysupdate
name networkservice
name sysguard
name dbused
name watchbog
name xmrigdaemon
name xmrigminer
name moneroocean
name c3pool_miner
name tsm
name pnscan
# Command-line fragments of mining software and public pools
arg stratum+tcp://
arg stratum+ssl://
arg stratum+tls://
arg stratum2+tcp://
arg --donate-level
arg --randomx-
arg --cpu-max-threads-hint
arg --coin=monero
arg --algo=rx/0
arg -a rx/0
arg c3pool.com
arg c3pool.org
arg supportxmr.com
arg moneroocean.stream
arg nanopool.org
arg minexmr.com
arg xmrpool.eu
arg hashvault.pro
arg herominers.com
arg 2miners.com
arg f2pool.com
arg unmineable.com
arg monerohash.com
arg xmr.pool.minergate.com
arg pool.hashvault.pro
arg gulf.moneroocean.stream
arg auto.c3pool.org
# Programs that legitimately keep a core busy for a long time
busy mysqld
busy mariadbd
busy postgres
busy mongod
busy redis-server
busy clickhouse-server
busy elasticsearch
busy java
busy ffmpeg
busy x264
busy x265
busy handbrakecli
busy rsync
busy gzip
busy pigz
busy xz
busy zstd
busy bzip2
busy tar
busy borg
busy restic
busy duplicity
busy clamscan
busy clamd
busy freshclam
busy apt
busy apt-get
busy dpkg
busy unattended-upgr
busy yum
busy dnf
busy rpm
busy cc1
busy cc1plus
busy ld
busy go
busy rustc
busy cargo
busy make
busy webpack
busy esbuild
busy imagick
busy convert
busy magick
busy gs
busy ollama
busy updatedb
busy mlocate
busy plocate
busy find
busy fstrim
busy btrfs
busy mdadm
go.mod6 lines
module approvalens.com/agent
go 1.22
toolchain go1.27.1
LICENSE22 lines
MIT License
Copyright (c) 2026 Kaan Tokalı / Approvalens
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
Server monitoring plans start at $2.99 a month, and every site monitoring plan includes one server. Server monitoring