../

Linux sysadmin

Day-to-day administration of a modern Linux server: files, users, services, packages, disks, networking, firewalls, logs and triage. Targets Ubuntu 24.04/26.04 LTS, Debian 13 and Fedora / RHEL 10, with apt vs dnf shown where they differ. Keys and TLS live in SSH, keys & certificates.

Filesystem hierarchy

All four distros use a merged /usr: /bin, /sbin and /lib are symlinks into /usr.

/ at a glance
/boot/                # kernel, initramfs, bootloaderdev/                 # device nodes: nvme0n1, sda, ttyetc/                 # host config; back this upsystemd/system/  # your unit files and overridesssh/             # sshd_config, host keyssudoers.d/       # sudo rules, one file per rolehome/                # user homes; root's is /rootopt/                 # self-contained third-party appsproc/                # virtual: PIDs and kernel staterun/                 # sockets, PID files; wiped at bootsrv/                 # data served: sites, apps, repossys/                 # virtual: devices, drivers, cgroupstmp/                 # scratch; tmpfs on Debian 13/Fedorausr/bin/             # commands; Fedora merges sbin herelib/             # libraries, packaged unit fileslocal/           # software you installed by handshare/           # docs, man pages, shared datavar/cache/           # regenerable: apt, dnf cacheslib/             # persistent state: docker, postgreslog/             # logs; journal in var/log/journal
Put…In
Hand-installed binaries/usr/local/bin
A service's code/srv/app or /opt/app
A service's mutable state/var/lib/app
Config and secrets/etc/app/ (mode 0640, group = service)
Logs (if not journald)/var/log/app/
Runtime sockets/run/app/ (RuntimeDirectory=app in systemd)

/proc/PID/ is handy: cmdline, environ, cwd, exe, fd/, status, limits.

Users, groups & sudo

CommandDoes
adduser aliceDebian/Ubuntu friendly wrapper: home, shell, password prompt
useradd -m -s /bin/bash -G sudo alicelow-level; -m makes the home. Admin group: sudo (Debian/Ubuntu), wheel (Fedora/RHEL)
useradd -r -s /usr/sbin/nologin -d /var/lib/app appsystem account for a service, no login
passwd aliceset password; -l lock, -u unlock, -e force change
usermod -aG docker aliceappend a group; without -a it replaces all groups
gpasswd -d alice dockerremove from a group
groupadd deploynew group
userdel -r alicedelete user and home
id alice, groups aliceUID, GID, groups
getent passwd alicelookup through NSS (covers LDAP/SSSD users too)
sudo -iu postgreslogin shell as another user
chage -l alicepassword aging; chage -E 0 expires the account

Group changes apply at the next login (or newgrp docker in the current shell).

FileHolds
/etc/passwdusers: name, UID, GID, home, shell
/etc/shadowpassword hashes, aging (root-only)
/etc/groupgroups and members
/etc/sudoers, /etc/sudoers.d/sudo rules
/etc/skel/copied into every new home
/etc/login.defsUID ranges, default umask, hash method

sudo rules

# always edit via visudo: it syntax-checks before saving
sudo visudo -f /etc/sudoers.d/deploy
sudo -l              # what may I run?
sudo -k              # drop cached credentials
sudoedit /etc/hosts  # edit as root with your own $EDITOR
/etc/sudoers.d/deploy
# who    where=(as whom)   [tags:] commands
deploy   ALL=(root)        NOPASSWD: /usr/bin/systemctl restart app
%ops     ALL=(ALL:ALL)     ALL

Files in sudoers.d whose names contain a . or end in ~ are ignored; keep mode 0440. Ubuntu 26.04 ships sudo-rs (a Rust rewrite) as sudo; common sudoers syntax is unchanged.

Permissions

ls -l shows -rwxr-x---: type, then owner / group / other triplets. A trailing + means ACLs are set.

OctalBitsOn a fileOn a directory
4rread contentslist names
2wmodify contentscreate, rename, delete entries
1xexecuteenter / traverse
ModeSymbolicTypical use
644rw-r--r--regular files
755rwxr-xr-xdirectories, scripts, binaries
640rw-r-----config readable by a service group
600rw-------secrets, private keys
700rwx------private dirs (~/.ssh)
2775rwxrwsr-xshared team dir (setgid)
1777rwxrwxrwt/tmp (sticky)
Special bitOctalSet withEffect
setuid4000chmod u+sexecutable runs as the file's owner (passwd)
setgid2000chmod g+sfile runs as its group; on a dir, new entries inherit the dir's group
sticky1000chmod +ton a dir, only an entry's owner may delete it
chmod 640 app.env
chmod u+x,go-w deploy.sh
chmod -R u=rwX,go=rX site/   # X: exec only on dirs
chown -R app:app /srv/app
chown :www-data upload/      # group only
stat -c '%a %U:%G %n' file   # 640 app:app file
# audit setuid binaries
sudo find / -xdev -perm -4000 -type f

umask

New files start at 666, dirs at 777, minus the umask.

umaskFilesDirsWhen
022644755common default
002664775users with a private group (Debian/Ubuntu default)
027640750services; set UMask=0027 in the unit
077600700secrets-only dirs

ACLs

For access beyond one owner and one group. Needs the acl package.

setfacl -m u:alice:rwX /srv/shared     # grant a user
setfacl -R -m g:web:rX /srv/site       # recursive, group
setfacl -d -m g:web:rwX /srv/uploads   # default ACL
getfacl /srv/shared
setfacl -x u:alice /srv/shared         # remove one entry
setfacl -b /srv/shared                 # remove all ACLs

Processes

CommandShows / does
ps auxevery process, BSD style
ps -ef --forestwith parent/child tree
ps -eo pid,user,%cpu,%mem,etime,cmd --sort=-%cpu | headtop CPU users, custom columns
pgrep -af bunPIDs + full command lines matching
pstree -ptree with PIDs
toplive; keys P CPU, M memory, 1 per-core, c full cmd, k kill
htop, btopfriendlier live views (install them)
systemctl status PIDwhich unit owns a process
lsof -p PIDits open files and sockets
ls -l /proc/PID/cwd /proc/PID/exeworking dir and binary
cat /proc/PID/environ | tr '\0' '\n'its environment

Signals

SignalNo.DefaultUse
SIGHUP1terminateterminal closed; many daemons reload config on it
SIGINT2terminateCtrl-C
SIGQUIT3core dumpCtrl-\
SIGKILL9killcannot be caught; no cleanup, last resort
SIGUSR1, SIGUSR210, 12terminateapp-defined (reopen logs, dump state)
SIGTERM15terminatepolite stop; default for kill and systemctl stop
SIGCONT18continueresume a stopped process
SIGSTOP19stoppause; cannot be caught
SIGTSTP20stopCtrl-Z
kill 1234              # SIGTERM
kill -HUP 1234         # reload
kill -9 1234           # only after TERM failed
pkill -f 'bun run'     # match full command line
killall nginx          # by name
kill -l                # list signals

Priority and jobs

nice -n 10 ./reindex.sh     # -20 (greedy) .. 19 (polite)
renice -n 5 -p 1234
ionice -c3 tar czf b.tgz d  # idle I/O class
long-task &                 # background; Ctrl-Z then bg
jobs -l; fg %1; disown %1
nohup ./task.sh > task.log 2>&1 &
# better: a transient unit with logs + limits
sudo systemd-run --unit=reindex -p MemoryMax=2G \
  /srv/app/reindex.sh
journalctl -u reindex -f

systemd

CommandDoes
systemctl status appstate, PID, memory, last log lines
systemctl start / stop / restart appcontrol now
systemctl reload appre-read config without restart (if supported)
systemctl enable --now appstart at boot and start now
systemctl disable --now appthe reverse
systemctl is-active app, is-enabled appscriptable checks (exit code)
systemctl list-units --failedwhat's broken
systemctl list-unit-files --state=enabledwhat starts at boot
systemctl cat appthe unit plus all drop-ins
systemctl edit appcreate a drop-in override
systemctl edit --full appcopy the whole unit to /etc and edit
systemctl daemon-reloadafter changing unit files by hand
systemctl mask appmake it impossible to start
systemctl show app -p MainPID,Restartany property
systemctl reset-failed appclear failed state / start limit
systemctl --user …per-user units; loginctl enable-linger zach keeps them running after logout
systemd-analyze blameslow boot units
systemd-cgtopresource use per unit

Where units live

unit search path (first wins)
/etc/systemd/system/                # yours; overrides allapp.serviceapp.service.d/override.conf  # from systemctl edit appbackup.servicebackup.timermulti-user.target.wants/app.service    # symlink made by enabletimers.target.wants/backup.timer   # symlink made by enablerun/systemd/system/                # runtime, transient unitsusr/lib/systemd/system/            # packaged; don't editnginx.service

User units go in ~/.config/systemd/user/.

Unit file anatomy

SectionDirectiveMeaning
[Unit]Description=shown in status and logs
After=, Before=ordering only
Wants=, Requires=pull in another unit (soft / hard)
StartLimitBurst=, StartLimitIntervalSec=give up after N restarts in a window
[Service]Type=simple, exec, notify, oneshot, forking
ExecStart=absolute path + args; no shell unless you call one
ExecReload=e.g. kill -HUP $MAINPID
User=, Group=drop root
WorkingDirectory=cwd
Environment=, EnvironmentFile=env vars; file is KEY=value lines
Restart=no, on-failure, always
RestartSec=delay between restarts
TimeoutStopSec=SIGTERM → SIGKILL grace (default 90s)
MemoryMax=, CPUQuota=cgroup limits
[Install]WantedBy=target that pulls it in on enable
Type=Considered started when
simpleimmediately after fork (default)
execafter the binary is exec'd; bad paths fail loudly
notifythe process sends READY=1 via sd_notify
oneshotthe process exits (scripts; add RemainAfterExit=yes for "state" units)
forkingthe parent exits (legacy daemons; prefer not)
/etc/systemd/system/worker.service
[Unit]
Description=Queue worker
After=network-online.target postgresql.service
Wants=network-online.target
 
[Service]
Type=exec
User=worker
WorkingDirectory=/srv/worker
EnvironmentFile=/etc/worker/env
ExecStart=/usr/local/bin/bun run worker.ts
Restart=always
RestartSec=5
 
[Install]
WantedBy=multi-user.target

Drop-in overrides only change what they set. To replace a list setting such as ExecStart, clear it first:

/etc/systemd/system/worker.service.d/override.conf
[Service]
ExecStart=
ExecStart=/usr/local/bin/bun run --smol worker.ts
Environment=LOG_LEVEL=debug
MemoryMax=512M

Stopping cleanly

systemctl stop sends SIGTERM, waits TimeoutStopSec, then sends SIGKILL.

server.ts
const server = Bun.serve({
  port: Number(process.env.PORT ?? 3000),
  fetch: () => new Response("ok"),
});
 
process.on("SIGTERM", async () => {
  await server.stop(); // let in-flight requests finish
  process.exit(0);
});

journalctl

CommandShows
journalctl -u app -ffollow one unit
journalctl -u app -n 200 --no-pagerlast 200 lines
journalctl -u app --since "1 hour ago"time window; also --until, "2026-09-25 14:00", today
journalctl -b, -b -1this boot, previous boot
journalctl --list-bootsboot IDs
journalctl -p err -bpriority err and worse (emerg..debug)
journalctl -kkernel messages (like dmesg)
journalctl -g 'timeout|refused'regex grep over messages
journalctl _PID=1234, _COMM=sshd, _UID=1000field matches
journalctl -u app -o json-prettystructured output; also -o cat, -o short-iso
journalctl -xerecent errors with explanations
journalctl --disk-usagejournal size
journalctl --vacuum-time=2weeks, --vacuum-size=1Gprune

Scheduling: cron & timers

┌──────── minute        0-59
│ ┌────── hour          0-23
│ │ ┌──── day of month  1-31
│ │ │ ┌── month         1-12 or jan-dec
│ │ │ │ ┌ day of week   0-7 or sun-sat (0 and 7 = Sunday)
│ │ │ │ │
* * * * *  command
ExpressionRuns
*/5 * * * *every 5 minutes
0 * * * *top of every hour
0 3 * * *03:00 daily
30 9 * * 1-509:30 on weekdays
0 */6 * * *every 6 hours
0 0 1 * *midnight on the 1st
15 2 * * 0Sundays at 02:15
0 9 1,15 * *09:00 on the 1st and 15th
@rebootonce at boot
@hourly, @daily, @weekly, @monthlyshortcuts

When both day-of-month and day-of-week are restricted, cron fires if either matches.

WhereFormat
crontab -e / -l / -rper-user table (edit / list / remove)
/etc/cron.d/appsystem file; has a user column: 0 3 * * * app /srv/app/job.sh
/etc/cron.daily/ etc.executable scripts, no extension (run-parts)

Cron gotchas: minimal PATH (use absolute paths), % must be escaped as \%, output is mailed or lost (append >> /var/log/job.log 2>&1), and runs can overlap (wrap in flock -n /run/lock/job.lock cmd). Install cron (Debian/Ubuntu) or cronie (Fedora/RHEL) if missing.

systemd timers

OnCalendar=Fires
hourly, daily, weeklyat the start of each period
*-*-* 03:00:0003:00 daily
Mon..Fri 09:30weekdays at 09:30
*:0/15every 15 minutes
*-*-01 00:001st of the month
Sat *-*-1..7 04:00first Saturday of the month

Relative timers: OnBootSec=5min, OnUnitActiveSec=1h. Test expressions with systemd-analyze calendar 'Mon..Fri 09:30'. List with systemctl list-timers.

cronsystemd timer
Setupone line.service + .timer
Logsmail or manual redirectjournal (journalctl -u job)
Missed run while offskippedPersistent=true catches up
Overlappossiblenever; won't start while still running
Jittermanual sleepRandomizedDelaySec=
Limits, sandboxingnoneany [Service] directive
Run nowcopy the commandsystemctl start job.service

Packages

TaskDebian / UbuntuFedora / RHEL
Refresh indexapt updateautomatic (dnf makecache)
Upgrade allapt upgradednf upgrade
Upgrade, allow removalsapt full-upgradednf distro-sync
Install / removeapt install x / apt remove xdnf install x / dnf remove x
Remove + configapt purge x(config left as .rpmsave)
Drop unused depsapt autoremovednf autoremove
Search / detailsapt search x / apt show xdnf search x / dnf info x
Installed listapt list --installeddnf list --installed
Which pkg owns a filedpkg -S /usr/bin/digrpm -qf /usr/bin/dig
Files in a pkgdpkg -L xrpm -ql x
Which pkg provides a fileapt-file search bin/digdnf provides '*/bin/dig'
History/var/log/apt/history.logdnf history, dnf history undo N
Pin a versionapt-mark hold xdnf versionlock add x
Install a local fileapt install ./x.debdnf install ./x.rpm
Clean cacheapt cleandnf clean all
Reboot needed?/run/reboot-required exists (Ubuntu)dnf needs-restarting -r

Fedora uses dnf5; RHEL 10 still ships dnf4. Everyday commands match. Extra software on RHEL usually comes from EPEL. Flatpak and Snap are for desktop apps; avoid them for server daemons.

Adding a third-party repo

/etc/apt/sources.list.d/example.sources
Types: deb
URIs: https://pkg.example.com/apt
Suites: stable
Components: main
Signed-By: /etc/apt/keyrings/example.gpg
curl -fsSL https://pkg.example.com/key.asc \
  | sudo gpg --dearmor -o /etc/apt/keyrings/example.gpg
sudo apt modernize-sources   # Debian 13: old .list → deb822
# Fedora (dnf5)
sudo dnf config-manager addrepo \
  --from-repofile=https://pkg.example.com/x.repo
# RHEL 10 (dnf4)
sudo dnf config-manager --add-repo \
  https://pkg.example.com/x.repo

Disks & filesystems

CommandShows / does
lsblk -fblock devices, filesystems, UUIDs, mount points
blkidUUIDs and types
df -hTusage per mounted filesystem; df -i for inodes
du -sh * | sort -hsize of each entry here
du -xh -d1 / | sort -htop-level usage, one filesystem
ncdu -x /interactive explorer
findmntmount tree with options
mount /dev/sdb1 /mnt, umount /mntmanual mount
mkfs.ext4 -L data /dev/sdb1format (mkfs.xfs on RHEL)
growpart /dev/sda 1grow a partition (cloud disk resized)
resize2fs /dev/sda1, xfs_growfs /grow ext4 / XFS online
smartctl -a /dev/nvme0disk health (smartmontools)
wipefs -a /dev/sdberase signatures (destructive)

Defaults: ext4 on Debian/Ubuntu, XFS on LVM for RHEL and Fedora Server, Btrfs on Fedora Workstation.

/etc/fstab
# what                                      where  type  options                  dump pass
UUID=3f1c2a9e-0d4b-4c1e-9b7a-2e5f6d8c9a10  /data  ext4  defaults,noatime,nofail  0    2
/swapfile                                   none   swap  sw                       0    0

Use UUID= (from blkid), not /dev/sdX. nofail keeps boot going if the disk is missing. After editing: sudo systemctl daemon-reload && sudo mount -a and fix any error before rebooting.

LVM in brief

Physical volumes (disks) → volume group (pool) → logical volumes (resizable "partitions").

CommandDoes
pvs, vgs, lvssummaries
pvcreate /dev/sdbmark a disk for LVM
vgextend vg0 /dev/sdbadd it to the pool
lvcreate -n data -L 50G vg0new volume /dev/vg0/data
lvextend -r -L +20G vg0/datagrow volume and filesystem (-r)
lvextend -r -l +100%FREE ubuntu-vg/ubuntu-lvUbuntu installer leaves space unallocated; claim it
lvcreate -s -n snap -L 5G vg0/datasnapshot (copy-on-write)

Networking

CommandShows / does
ip -br ainterfaces and addresses, one line each
ip r, ip route get 1.1.1.1routes; which route a packet takes
ip link set eth0 upbring up (not persistent)
ss -tulpnlistening TCP/UDP sockets + process
ss -tnp state establishedopen connections
ss -ssocket summary
hostnamectl set-hostname web1set hostname
resolvectl statusDNS servers per link (systemd-resolved)
resolvectl query example.comresolve via the system resolver
resolvectl flush-cachesclear DNS cache
dig +short example.com Araw DNS answer (dnsutils / bind-utils)
dig @1.1.1.1 example.com MXask a specific server
dig -x 203.0.113.7reverse lookup
dig +trace example.comwalk from the root servers
curl -sSI https://example.comresponse headers only
curl -v https://example.comTLS handshake + headers
curl --resolve example.com:443:10.0.0.5 https://example.comtest a backend before DNS changes
curl -so /dev/null -w '%{http_code} %{time_total}\n' URLstatus and timing
nc -zv db.internal 5432is a TCP port reachable?
mtr example.comlive traceroute with loss per hop

Concepts (subnets, ports, DNS, TLS) are in Communication networks.

DistroPersistent configTool
Ubuntu Server/etc/netplan/*.yaml → systemd-networkdnetplan try, netplan apply
Fedora, RHELNetworkManager keyfiles in /etc/NetworkManager/system-connections/nmcli, nmtui
Debian 13 server/etc/network/interfaces (ifupdown)ifup, ifdown
/etc/netplan/60-static.yaml
network:
  version: 2
  ethernets:
    eth0:
      dhcp4: false
      addresses: [10.0.0.10/24]
      routes:
        - to: default
          via: 10.0.0.1
      nameservers:
        addresses: [1.1.1.1, 9.9.9.9]

netplan try rolls back after 120 s unless you confirm, so a typo won't lock you out. Netplan files should be mode 600.

nmcli device status
nmcli con show
nmcli con mod eth0 ipv4.method manual \
  ipv4.addresses 10.0.0.10/24 ipv4.gateway 10.0.0.1 \
  ipv4.dns "1.1.1.1 9.9.9.9"
nmcli con up eth0

Firewall

Front endDefault onBackend
ufwUbuntu (installed, inactive); apt install ufw on Debiannftables
firewalldFedora, RHEL (active)nftables
nfteverywherethe kernel's nftables directly

Use one front end per host. Always allow SSH before enabling a default-deny policy.

sudo ufw default deny incoming
sudo ufw default allow outgoing
sudo ufw allow OpenSSH           # or: ufw limit 22/tcp
sudo ufw allow 80,443/tcp
sudo ufw allow from 10.0.0.0/8 to any port 5432 proto tcp
sudo ufw enable
sudo ufw status numbered
sudo ufw delete 4
sudo firewall-cmd --get-active-zones
sudo firewall-cmd --list-all
sudo firewall-cmd --permanent --add-service=https
sudo firewall-cmd --permanent --add-port=3000/tcp
sudo firewall-cmd --permanent --zone=trusted \
  --add-source=10.0.0.0/8
sudo firewall-cmd --reload   # --permanent needs a reload
/etc/nftables.conf
#!/usr/sbin/nft -f
flush ruleset
 
table inet filter {
  chain input {
    type filter hook input priority 0; policy drop;
    ct state established,related accept
    iif lo accept
    meta l4proto { icmp, ipv6-icmp } accept
    tcp dport { 22, 80, 443 } accept
  }
  chain forward {
    type filter hook forward priority 0; policy drop;
  }
}

Check with sudo nft -c -f /etc/nftables.conf, load with sudo systemctl enable --now nftables, inspect with sudo nft list ruleset.

Logs & monitoring

WhereWhat
journalctleverything from systemd units, kernel, auth; the primary log on all four distros
/var/log/syslog, /var/log/auth.logUbuntu (rsyslog installed); Debian 12+ has no rsyslog by default
/var/log/messages, /var/log/secureRHEL with rsyslog
/var/log/apt/, /var/log/dnf5.logpackage operations
/var/log/nginx/, /var/log/postgresql/per-daemon files
dmesg -T --level=err,warnkernel ring buffer with timestamps
last -n 20, lastblogins, failed logins
who, wwho is on now and what they run

Cap the journal with a drop-in:

/etc/systemd/journald.conf.d/size.conf
[Journal]
SystemMaxUse=1G
MaxRetentionSec=1month

Then sudo systemctl restart systemd-journald.

NeedReach for
Metrics over timenode_exporter + Prometheus + Grafana
Quick dashboards on one boxbtop, netdata, Cockpit (Fedora/RHEL: systemctl enable --now cockpit.socket)
Historical CPU/mem/IOsysstat (sar), below
Failed unitssystemctl list-units --failed in a check
Disk healthsmartd from smartmontools
Uptimeexternal HTTP checks; the box can't report its own outage

Text & file toolkit

ToolCommon forms
grepgrep -rn 'TODO' src/, -i ignore case, -v invert, -E regex, -o only match, -l files only, -c count, -C3 context
rgripgrep: recursive, respects .gitignore, fast. rg -n 'fn \w+' -t ts, rg -l, rg --hidden
sedsed -n '10,20p' f, sed -i 's/old/new/g' f (GNU; macOS needs -i ''), sed '/^#/d'
awkawk '{print $1}', awk -F: '$3>=1000 {print $1}' /etc/passwd, awk '{s+=$5} END {print s}'
cutcut -d: -f1,7 /etc/passwd, cut -c1-80
sort-n numeric, -h human sizes, -r reverse, -k2,2 by column, -t, separator, -u unique
uniqneeds sorted input; -c count, -d dupes only
wc-l lines, -w words, -c bytes
head, tail-n 50; tail -F follows across rotation
trtr a-z A-Z, tr -d '\r', tr -s ' '
xargsxargs -0 with find -print0, -n1 one arg per run, -P4 parallel, -I{} placeholder
findfind . -name '*.log' -mtime +7 -delete, -type f -size +100M, -newer ref, -exec cmd {} +
jqjq '.items[].name', jq -r '.[] | select(.ok) | .id', jq -c, jq --arg k v
column -talign whitespace-separated output
diff -u a bunified diff; comm -12 lines in both sorted files
teecmd | sudo tee /etc/file writes as root
# top 10 client IPs in an access log
awk '{print $1}' access.log | sort | uniq -c \
  | sort -rn | head
# count HTTP status codes
awk '{print $9}' access.log | sort | uniq -c
# replace across files safely (NUL-separated)
rg -l -0 'old\.host' \
  | xargs -0 sed -i 's/old\.host/new.host/g'
# JSON logs: errors from the last hour
journalctl -u app --since -1h -o cat \
  | jq -c 'select(.level == "error")'

Archives

CommandDoes
tar czf a.tgz dir/create gzip tarball
tar --zstd -cf a.tar.zst dir/create zstd tarball (fast, small)
tar -I 'zstd -19 -T0' -cf a.tar.zst dir/max compression, all cores
tar xf a.tar.zst -C /destextract; format auto-detected
tar tf a.tgzlist contents
tar czf - dir | ssh host 'tar xzf - -C /srv'copy a tree over SSH
zstd -T0 big.sql, zstd -d big.sql.zstcompress / decompress one file
gzip -k f, gunzip f.gzclassic; -k keeps the original
xz -T0 fsmallest, slowest
zip -r a.zip dir/, unzip -l a.zipfor people on Windows

Performance triage

The first minute on a slow box, in order:

CommandLook at
uptimeload averages (1, 5, 15 min) vs nproc; above core count means queueing
dmesg -T | tail, journalctl -k -p warningOOM kills, disk or NIC errors
free -havailable column, not free; swap in use
vmstat 1 5r (runnable) above cores = CPU saturated; si/so swapping; wa I/O wait; st steal (noisy VM neighbor)
mpstat -P ALL 1one hot core = single-threaded bottleneck
pidstat 1which process uses CPU
iostat -xz 1%util near 100 and high r_await/w_await = disk bound
sar -n DEV 1network throughput per interface
ss -s, ss -tan | awk '{print $1}' | sort | uniq -cconnection counts and states
top / btopconfirm the culprit

mpstat, pidstat, iostat and sar come from sysstat. Enable history with sudo systemctl enable --now sysstat (Debian/Ubuntu also want ENABLED="true" in /etc/default/sysstat); then sar -u, sar -r, sar -q show the day so far.

# who got OOM-killed?
journalctl -k -g 'out of memory|oom-kill' -b
# memory per unit (cgroups)
systemd-cgtop -m
# biggest RSS processes
ps -eo pid,rss,cmd --sort=-rss | head

Hardening checklist

AreaDo
UpdatesDebian/Ubuntu: unattended-upgrades. Fedora: dnf5-plugin-automatic, dnf5-automatic.timer. RHEL 10: dnf-automatic, dnf-automatic.timer. On dnf, set apply_updates = yes in /etc/dnf/automatic.conf
SSHkeys only, no root login, AllowGroups; see sshd hardening
Firewalldefault deny inbound; open only what ss -tulpn should show
Brute forcefail2ban (or ufw limit), or keep SSH on a VPN / Tailscale only
Least privilegeone system user per service; no shared sudo accounts; sudoers.d rules scoped to commands
Service sandboxingNoNewPrivileges, ProtectSystem=strict, ProtectHome, PrivateTmp; score with systemd-analyze security app
MACkeep AppArmor (Debian/Ubuntu) or SELinux enforcing (Fedora/RHEL); fix labels with restorecon -Rv, debug with ausearch -m avc
Secrets/etc/app/env mode 0640 root:app, or systemd LoadCredential=
TimeNTP on (timedatectl); logs and TLS need a correct clock
Attack surfaceremove unused packages and services; systemctl list-unit-files --state=enabled
Backupsautomated, off-host, and restore-tested
/etc/apt/apt.conf.d/20auto-upgrades
APT::Periodic::Update-Package-Lists "1";
APT::Periodic::Unattended-Upgrade "1";
/etc/fail2ban/jail.local
[DEFAULT]
bantime  = 1h
findtime = 10m
maxretry = 5
backend  = systemd
 
[sshd]
enabled = true

sudo fail2ban-client status sshd shows bans; fail2ban-client set sshd unbanip IP lifts one.

Recipes

systemd service for a Bun app

Run a Bun server as an unprivileged, sandboxed, auto-restarting service.

/etc/systemd/system/app.service
[Unit]
Description=My Bun app
After=network-online.target
Wants=network-online.target
 
[Service]
Type=exec
User=app
WorkingDirectory=/srv/app
EnvironmentFile=/etc/app/env
Environment=NODE_ENV=production PORT=3000
ExecStart=/usr/local/bin/bun run src/index.ts
Restart=on-failure
RestartSec=2
TimeoutStopSec=20
NoNewPrivileges=yes
ProtectSystem=strict
ProtectHome=yes
PrivateTmp=yes
StateDirectory=app
 
[Install]
WantedBy=multi-user.target

Install Bun system-wide with curl -fsSL https://bun.sh/install | sudo BUN_INSTALL=/usr/local bash, create the user with useradd -r -s /usr/sbin/nologin app, then sudo systemctl daemon-reload && sudo systemctl enable --now app. StateDirectory=app gives a writable /var/lib/app; bun build --compile produces a single binary if you'd rather not ship Bun.

systemd timer

Replace a cron job with something that logs to the journal and catches up after downtime.

# /etc/systemd/system/backup.service
[Unit]
Description=Nightly backup
 
[Service]
Type=oneshot
User=backup
ExecStart=/usr/local/bin/backup.sh
 
# /etc/systemd/system/backup.timer
[Unit]
Description=Run backup nightly
 
[Timer]
OnCalendar=*-*-* 03:00:00
RandomizedDelaySec=15m
Persistent=true
 
[Install]
WantedBy=timers.target

Enable the timer, not the service: sudo systemctl enable --now backup.timer. Run once by hand with sudo systemctl start backup.service.

Find what's using a port

When EADDRINUSE says port 3000 is taken.

sudo ss -ltnp 'sport = :3000'
sudo lsof -nP -iTCP:3000 -sTCP:LISTEN
systemctl status 12345        # which unit owns the PID
ps -o pid,user,etime,cmd -p 12345
sudo fuser -k -TERM 3000/tcp  # stop whatever holds it

Find big files

When df says a disk is full.

df -hT -x tmpfs -x devtmpfs
sudo du -xh -d1 / 2>/dev/null | sort -h | tail -15
sudo ncdu -x /
sudo find / -xdev -type f -size +500M \
  -exec ls -lh {} + 2>/dev/null | sort -k5 -h
sudo journalctl --vacuum-size=500M
sudo apt clean              # or: sudo dnf clean all
docker system df            # then: docker system prune
# space not freed after rm? a process still holds it
sudo lsof -nP +L1

Rotate logs with logrotate

For apps that write their own files instead of stdout; logrotate runs daily from logrotate.timer.

/etc/logrotate.d/app
/var/log/app/*.log {
    daily
    rotate 14
    maxsize 100M
    compress
    delaycompress
    missingok
    notifempty
    create 0640 app app
    sharedscripts
    postrotate
        systemctl kill -s HUP app.service
    endscript
}

Dry-run with sudo logrotate -d /etc/logrotate.d/app; force with -f. If the app can't reopen its log on SIGHUP, use copytruncate instead of create + postrotate.

Add a swap file

Small VPS with no swap, to survive memory spikes instead of OOM-killing.

sudo fallocate -l 2G /swapfile
sudo chmod 600 /swapfile
sudo mkswap /swapfile
sudo swapon /swapfile
echo '/swapfile none swap sw 0 0' | sudo tee -a /etc/fstab
swapon --show && free -h
# prefer RAM; swap only under pressure
echo 'vm.swappiness=10' \
  | sudo tee /etc/sysctl.d/99-swappiness.conf
sudo sysctl --system

On Btrfs use sudo btrfs filesystem mkswapfile --size 2G /swapfile. Fedora already has compressed swap in RAM (zramctl).

New server bootstrap

First five minutes on a fresh Ubuntu/Debian VPS, as root. Keep your current session open and test a new login before logging out.

bootstrap.sh
#!/usr/bin/env bash
# usage (as root): ./bootstrap.sh zach ~/zach.pub
set -euo pipefail
U=${1:?user}; KEY=$(cat "${2:?pubkey file}")
 
apt-get update && apt-get -y full-upgrade
apt-get -y install ufw fail2ban unattended-upgrades \
  curl git htop ncdu jq sysstat
 
id "$U" &>/dev/null || useradd -m -s /bin/bash -G sudo "$U"
passwd "$U"   # sudo will ask for it
install -d -m 700 -o "$U" -g "$U" "/home/$U/.ssh"
install -m 600 -o "$U" -g "$U" /dev/null \
  "/home/$U/.ssh/authorized_keys"
echo "$KEY" >> "/home/$U/.ssh/authorized_keys"
 
cat > /etc/ssh/sshd_config.d/00-hardening.conf <<'EOF'
PermitRootLogin no
PasswordAuthentication no
KbdInteractiveAuthentication no
EOF
sshd -t && systemctl try-reload-or-restart ssh
ufw allow OpenSSH && ufw --force enable
dpkg-reconfigure -f noninteractive unattended-upgrades

References