1---2name: sysadmin3description: Manage Linux servers with user administration, process control, storage, and system maintenance.4---5
6# System Administration Rules
7
8## User Management
9- Create service accounts with `--system` flag — no home directory, no login shell
10- `sudo` with specific commands, not blanket ALL — principle of least privilege
11- Lock accounts instead of deleting: `usermod -L` — preserves audit trail and file ownership
12- SSH keys in `~/.ssh/authorized_keys` with restrictive permissions — 600 for file, 700 for directory
13- `visudo` to edit sudoers — catches syntax errors before saving, prevents lockout
14
15## Process Management
16- `systemctl` for services, not `service` — systemd is standard on modern distros
17- `journalctl -u service -f` for live logs — more powerful than tail on log files
18- `nice` and `ionice` for background tasks — don't compete with production workloads
19- Kill signals: SIGTERM (15) first, SIGKILL (9) last resort — SIGKILL doesn't allow cleanup
20- `nohup` or `screen`/`tmux` for long-running commands — SSH disconnect kills regular processes
21
22## File Systems and Storage
23- `df -h` for disk usage, `du -sh *` to find culprits — check before disk fills completely
24- `lsof +D /path` finds processes using a directory — needed before unmounting
25- `ncdu` for interactive disk usage — faster than repeated du commands
26- Mount options matter: `noexec`, `nosuid` for security on data partitions
27- Resize filesystems with care: grow is safe, shrink risks data loss — always backup first
28
29## Logs and Monitoring
30- `logrotate` prevents disk fill — configure size limits and retention
31- Centralize logs to external system — local logs lost if server dies
32- `/var/log/auth.log` or `/var/log/secure` for login attempts — watch for brute force
33- `dmesg` for kernel messages — hardware errors, OOM kills appear here
34- Monitor inode usage, not just disk space — many small files exhaust inodes
35
36## Permissions and Security
37- `chmod 600` for secrets, `640` for configs, `644` for public — world-writable is almost never correct
38- Sticky bit on shared directories (`chmod +t`) — users can only delete their own files
39- `setfacl` for complex permissions — when traditional owner/group/other isn't enough
40- `chattr +i` makes files immutable — even root can't modify without removing flag
41- SELinux/AppArmor in enforcing mode — permissive logs but doesn't protect
42
43## Package Management
44- `apt update` before `apt upgrade` — upgrade without update uses stale package lists
45- Unattended security updates: `unattended-upgrades` — critical patches shouldn't wait
46- Pin package versions in production — unexpected upgrades cause unexpected outages
47- Remove unused packages: `apt autoremove` — reduces attack surface and disk usage
48- Know your package manager: apt/yum/dnf/pacman — commands differ, concepts similar
49
50## Backups
51- Test restores regularly — backups that can't restore are worthless
52- Include package lists and configs, not just data — recreating environment is painful
53- Offsite backups mandatory — local backups don't survive disk failure or ransomware
54- Backup before any risky change — "I'll just quickly edit" famous last words
55- Document restore procedure — 3am disaster is wrong time to figure it out
56
57## Performance
58- `top`/`htop` for live view, `vmstat` for trends — understand baseline before diagnosing
59- `iotop` for disk I/O bottlenecks — slow disk often blamed on CPU
60- Load average: 1.0 per core is healthy — consistently higher means queuing
61- Swap usage isn't inherently bad — but consistent swapping indicates memory shortage
62- `sar` for historical data — retroactively diagnose what happened during incident
63
64## Networking Basics
65- `ss -tulpn` shows listening ports — `netstat` is deprecated
66- `ip addr` and `ip route` replace `ifconfig` and `route` — learn the new tools
67- Check both host firewall and cloud security groups — traffic blocked at either level fails
68- `/etc/hosts` for local overrides — quick testing without DNS changes
69- `curl -v` shows full connection details — headers, timing, TLS handshake
70
71## Common Mistakes
72- Running services as root — one exploit owns the system
73- No monitoring until something breaks — reactive is expensive
74- Editing config without backup — `cp file file.bak` takes two seconds
75- Rebooting to "fix" issues — masks the problem, it'll return
76- Ignoring disk space warnings — 100% full causes cascading failures
77- Forgetting timezone configuration — logs from different servers don't correlate