Linux for DevOps 2026 - Essential Skills & Shell Scripting for Automation
Why Linux is the foundation of DevOps, the essential commands and concepts to master (files, processes, permissions, networking), and how shell scripting automates real DevOps work.
If you learn one thing before everything else in DevOps, make it Linux. Servers, containers, CI runners and cloud instances almost all run Linux, and most automation is written against it.
This guide is the version I wish someone had given me. It has the commands, but more importantly it has what the output looks like and what you do with it, because knowing that journalctl exists and knowing how to read what it prints at 2am are different skills, and only the second one gets you hired.
Why is Linux mandatory for DevOps?
The overwhelming majority of production servers, cloud instances and containers run Linux. You cannot deploy, debug or automate systems you cannot navigate from the command line.
It is also the skill that separates candidates in interviews, because it is the one that cannot be memorised. When a container will not start or a service is refusing connections, the useful next move is reading logs, checking processes, ports and permissions - and that is Linux, not Kubernetes.
There is a version of this that people find hard to believe until they see it, so I will say it plainly. Every Docker container is a Linux process with some walls around it. Every Kubernetes pod is one or more of those processes on a Linux node. Every CI job runs in a Linux shell. Every cloud instance is a Linux machine you happen not to own. When any of those things breaks, the tool's own abstractions are the first thing to stop being useful, and what is left is the operating system underneath. The engineers who are calm during incidents are, almost without exception, the ones who are comfortable there.
Essential Linux skills to master
- Core commands: navigating files, viewing logs, searching text (grep), editing files.
- File & process management: permissions, ownership, running and inspecting processes.
- Permissions: users, groups, and read/write/execute - critical for security.
- Networking: ports, DNS, SSH, and basic troubleshooting (ping, curl, netstat).
That list is what every roadmap says, and it is correct but unhelpfully flat, because those four areas are not equally important and are not learned in that order. Here is how the weight actually falls:
Reading logs and searching text is the skill you use most, by a wide margin. Something breaks; you look at what it wrote before it broke. Learn tail -f, less, grep with -i, -v, -n and -R, and journalctl. Half of all debugging is these five commands.
Processes and services come next. What is running, what is not, what died and why, what is eating the CPU. ps, top or htop, systemctl, kill and knowing the difference between the signals.
Permissions and ownership matter because a surprising share of "it does not work" is "it does not have permission to work". Not complicated, but you need it cold.
Networking is the area where most DevOps candidates are weakest, and where the gap between adequate and good is widest. Is the port open. Is something listening. Does the name resolve. Can I reach it from here but not from there. This is what the section below with real output is for.
File navigation and editing - the thing tutorials spend the most time on - is the thing you will pick up fastest and think about least. Learn enough vim to edit a config file on a server that has nothing else installed, and move on.
The commands worth knowing cold
| Command | When you reach for it |
|---|---|
journalctl -u nginx -n 100 --no-pager | Last 100 log lines for a systemd service. Usually the first command of an incident. |
systemctl status nginx | Is it running, did it fail, and why. |
ss -tulpn | What is listening on which port, and which process owns it. Replaces the older netstat. |
df -h and du -sh * | Disk full - which is the cause of a startling share of production incidents. |
grep -R "timeout" /var/log/ | Find a string across log files. |
ps aux --sort=-%mem | head | What is eating the memory. |
curl -sS -o /dev/null -w "%{http_code}" localhost:3000 | Is the app actually answering, separate from whether the proxy is. |
Processes and services, with the output you will actually see
Commands are easy to list. What trips people up is the output, because nobody shows it to you until you are looking at it under pressure. Here is what the important ones print, and what to look at.
systemctl status on a service that has failed:
$ systemctl status myapp
x myapp.service - My Application
Loaded: loaded (/etc/systemd/system/myapp.service; enabled)
Active: failed (Result: exit-code) since Tue 2026-09-29 02:14:07 IST; 3min ago
Process: 18422 ExecStart=/opt/myapp/bin/server (code=exited, status=1/FAILURE)
Main PID: 18422 (code=exited, status=1/FAILURE)
Sep 29 02:14:07 web-01 server[18422]: Error: listen EADDRINUSE: address already in use :::3000
Sep 29 02:14:07 web-01 systemd[1]: myapp.service: Main process exited, code=exited, status=1/FAILURE
Sep 29 02:14:07 web-01 systemd[1]: myapp.service: Failed with result 'exit-code'.
Four lines matter. Active: failed tells you it is down. status=1/FAILURE tells you the process exited on its own rather than being killed. The timestamp tells you when - three minutes ago, which lines up with the page you just got. And the log tail at the bottom, which most people skip past to reach the red text, has the actual answer: EADDRINUSE, something else is already on port 3000. You have not needed a second command yet.
ss -tulpn to find out what that something is:
$ sudo ss -tulpn | grep 3000
tcp LISTEN 0 511 *:3000 *:* users:(("node",pid=17903,fd=19))
There is the culprit: a node process, PID 17903, that is not the one systemd just tried to start. Usually it is a previous instance that did not shut down cleanly, or a developer who ran the app by hand for a test and forgot. ps -p 17903 -o pid,user,etime,cmd tells you who started it and how long ago. Then you decide whether to kill it.
journalctl when you need more than the tail:
$ journalctl -u myapp --since "10 min ago" --no-pager
$ journalctl -u myapp -f # follow, like tail -f
$ journalctl -u myapp -p err -b # errors only, since last boot
$ journalctl --disk-usage # because the journal itself fills disks
The --since form is the one you want during an incident - you know roughly when it started, and you do not want to scroll through a week. -p err filters by priority and is how you find the one real error in ten thousand info lines.
ps and top for "the server is slow":
$ ps aux --sort=-%cpu | head -5
USER PID %CPU %MEM VSZ RSS TTY STAT START TIME COMMAND
app 21044 97.3 4.1 1.2G 330M ? Rl 01:58 14:22 /usr/bin/python3 /opt/jobs/reindex.py
postgres 1180 3.2 22.8 8.1G 1.8G ? Ss Aug14 402:11 postgres: checkpointer
One process at 97 percent CPU for fourteen minutes, started at 01:58. That is not a mystery; that is a batch job that should be running somewhere else, or at a different time, or with a nice value. The STAT column is worth learning: R is running, S is sleeping, D is waiting on disk and cannot be killed, and Z is a zombie that has already exited and is only still listed because its parent has not collected it. If you see a lot of D, the problem is the storage, not the process.
The signals, since the interview question is coming: kill sends SIGTERM by default, which asks the process to shut down cleanly. kill -9 sends SIGKILL, which does not ask. Always try the first one first; a database that gets -9 mid-write will make you regret it. systemctl stop does the polite version for you and waits.
Networking from the command line
"It cannot connect" is the most common incident description there is, and it has five different causes that need five different commands. Here is the order, with what each one tells you.
Does the name resolve?
$ dig +short api.payments.internal
10.40.2.17
$ dig +short api.payments.internal @8.8.8.8
# empty: internal name, public resolver - expected
No answer from your normal resolver means DNS, and nothing downstream matters until that is fixed. An answer that is a different IP from yesterday means someone changed a record, and that may be the whole incident.
Can I reach the host?
$ ping -c 3 10.40.2.17
$ traceroute -n 10.40.2.17 # where does it stop?
Cloud security groups often block ICMP, so a failed ping is not conclusive. Traceroute showing the path stopping at a particular hop is more useful: it tells you which network boundary is dropping you.
Is the port open, and is something listening?
$ nc -zv 10.40.2.17 8443
Connection to 10.40.2.17 8443 port [tcp/*] succeeded!
$ nc -zv 10.40.2.17 8443
nc: connect to 10.40.2.17 port 8443 (tcp) failed: Connection refused
These two failures are different. Connection refused means the packet arrived and the host actively said no - nothing is listening on that port. Connection timed out means the packet never got an answer - a firewall, security group or route is eating it. Refused is a service problem; go to the host and run ss -tulpn. Timed out is a network problem; go to the security groups. Knowing that distinction saves an hour on almost every connectivity incident.
Does the application answer, and what does it say?
$ curl -sS -o /dev/null -w "%{http_code} %{time_total}s\n" https://api.payments.internal:8443/health
200 0.043s
$ curl -sS -v https://api.payments.internal:8443/health 2>&1 | grep -E "^(<|\*)" | head
* Trying 10.40.2.17:8443...
* Connected to api.payments.internal (10.40.2.17) port 8443
* SSL certificate problem: certificate has expired
The first form is the one for scripts and health checks: status code and time, nothing else. The second, -v, shows the conversation, and it is how you find out that the connection is fine and the problem is a certificate that expired at midnight - which, if you have been doing this for a while, you will have already guessed from the timing.
What does this machine's network look like?
$ ip addr show # interfaces and their addresses (replaces ifconfig)
$ ip route # where packets go; look for the default route
$ cat /etc/resolv.conf # which DNS server this box asks
You need these less often, but when a freshly provisioned instance cannot reach anything, a missing default route or a wrong resolver is usually why.
Permissions, properly
Every file has an owner, a group, and three permission sets - owner, group, everyone else. Read is 4, write is 2, execute is 1, and you add them up. So chmod 644 file means owner reads and writes, everyone else only reads. chmod 755 script.sh adds execute, which is what makes a script runnable.
The number that should make you stop is 777. It means anyone on the system can modify the file, and it is almost never the real fix - it is what people try when the actual problem is that the file belongs to the wrong user. chown is usually the correct answer, and ls -l tells you which.
On a private SSH key the rule is stricter: it must be 600. SSH refuses to use a key that others can read, which is the cause of the classic "permissions 0644 for key are too open" error.
What ls -l shows, and how to read it:
$ ls -l /opt/myapp/
-rw-r--r-- 1 root root 1284 Sep 12 10:02 config.yaml
drwxr-x--- 2 app app 4096 Sep 29 01:58 data
-rwxr-xr-x 1 app app 22140 Sep 12 10:02 server
The first character is the type - - for a file, d for a directory. Then three groups of three: owner, group, other. Then the owner and group names. So config.yaml is owned by root and readable by everyone; data is owned by app and nobody else can even list it; server is executable by everyone. If myapp runs as user app and needs to write to config.yaml, it cannot, and that is your bug. The fix is chown app:app config.yaml, not chmod 777.
Directories have a wrinkle: execute on a directory means "can enter it". A directory with read but not execute lets you list the names but not open anything inside. This one catches people during deploys, when a directory gets created with the wrong mode and the app can see its config file but not read it.
Debugging a service that will not start
A repeatable order that resolves most cases:
systemctl status <service>- read the failure line, not just the red text.journalctl -u <service> -n 50- the actual error is almost always here.ss -tulpn | grep <port>- is something else already holding the port?df -h- a full disk causes failures that look like anything but a full disk.ls -lon the config and data paths - wrong owner after a deploy is a frequent cause.
Being able to describe that sequence is worth more in an interview than naming twenty commands.
Step four deserves its own paragraph, because it is the one people do last and should do second. A full disk does not produce an error that says "disk full". It produces a database that cannot write its journal and reports corruption, a web server that cannot write its access log and stops accepting connections, a container runtime that cannot pull an image and reports a network error, and a logging daemon that dies silently so that the very errors you need are not being recorded. An experienced team can easily spend ninety minutes on a "database corruption" incident that turns out to be df -h showing 100 percent on /var. Check it early. It takes one second.
$ df -h
Filesystem Size Used Avail Use% Mounted on
/dev/nvme0n1p1 80G 80G 0 100% /
tmpfs 3.9G 0 3.9G 0% /dev/shm
$ sudo du -sh /var/* 2>/dev/null | sort -rh | head -3
41G /var/lib/docker
22G /var/log
9.1G /var/cache
And there it is. docker system prune and a log rotation policy, and the "corruption" resolves itself.
Shell scripting for automation
Shell scripting (mostly Bash) is how DevOps engineers turn repetitive manual steps into reliable, repeatable automation - backups, deployments, log rotation, health checks and glue between tools. A useful script chains commands together, uses variables and conditionals, loops over items, and exits cleanly on errors. Even basic scripting dramatically multiplies what one engineer can manage, and it is a common interview topic.
#!/usr/bin/env bash
set -euo pipefail
APP_URL=http://localhost:3000/healthz
RETRIES=5
for i in $(seq 1 $RETRIES); do
code=$(curl -sS -o /dev/null -w "%{http_code}" "$APP_URL" || true)
if [ "$code" = "200" ]; then
echo "healthy after $i attempt(s)"
exit 0
fi
echo "attempt $i returned $code, retrying"
sleep 3
done
echo "still unhealthy after $RETRIES attempts" >&2
exit 1
The first two lines matter more than the rest. set -e stops on the first failing command, -u makes an undefined variable an error rather than an empty string, and -o pipefail makes a pipeline fail if any stage fails. Without them a broken script carries on cheerfully and exits 0, and your pipeline reports success for a deploy that did not happen.
Here is a second one that does a job you will actually be asked to do - a deploy wrapper that refuses to run under the wrong conditions, records what it did, and can undo itself:
#!/usr/bin/env bash
set -euo pipefail
ENV="$1" # staging or production
TAG="$2" # image tag, e.g. a1b2c3d4e5f6
LOG=/var/log/deploys.log
[ "$ENV" = "staging" ] || [ "$ENV" = "production" ] || {
echo "usage: $0 <staging|production> <tag>" >&2; exit 2; }
[ "$ENV" = "production" ] && [ "$(date +%u)" -ge 6 ] && {
echo "refusing production deploy on a weekend" >&2; exit 3; }
previous=$(kubectl -n "$ENV" get deploy myapp \
-o jsonpath='{.spec.template.spec.containers[0].image}' | cut -d: -f2)
echo "$(date -Is) $ENV $USER $previous -> $TAG" | tee -a "$LOG"
kubectl -n "$ENV" set image deploy/myapp myapp="ghcr.io/acme/myapp:$TAG"
kubectl -n "$ENV" rollout status deploy/myapp --timeout=120s || {
echo "rollout failed, rolling back to $previous" | tee -a "$LOG"
kubectl -n "$ENV" rollout undo deploy/myapp
exit 1
}
What makes this a script someone would trust rather than a script someone wrote: it checks its arguments and refuses bad ones, with a usage line. It has a guard rail that encodes a team rule. It records the previous version before changing anything, so rollback is possible. It logs who did what, when, to a file that survives the terminal closing. And when the rollout fails, it does not just exit - it puts things back. That last part is the difference between a script and automation.
Two small things in it that are worth adopting everywhere: "$VAR" in double quotes, always, because a variable containing a space will otherwise split into two arguments and do something surprising; and tee -a to write to the log and the screen at once, so the person running it sees what was recorded.
Real-world usage
DevOps engineers use Linux daily to manage servers, debug incidents, inspect logs, and run automation - it underpins every other tool you will learn.
To make that specific, here are three situations from ordinary weeks where the Linux was the whole job, and the tool on top of it was almost incidental:
A Kubernetes pod in CrashLoopBackOff. kubectl logs shows the app starting, printing "permission denied" on /data/cache, and exiting. That is not a Kubernetes problem. The volume is mounted with root ownership and the container runs as UID 1000. The fix is a securityContext with the right fsGroup - a Kubernetes setting - but you only knew which setting because you understood what "permission denied" on a directory means and how ownership works. Everything in the Kubernetes guide sits on top of this.
A CI build that passes locally and fails on the runner. The log says command not found: jq. The developer's laptop has it; the runner image does not. That is a Linux packaging problem in a CI/CD pipeline costume. The fix is one apt-get install line in the workflow - or, better, a Dockerfile that declares every tool the build needs so this cannot happen again.
A Docker image that is 1.4 GB. docker history shows one layer at 900 MB: the build toolchain, installed to compile a dependency and never removed. Multi-stage builds fix it, but the reason you knew to look is that you understand layers as filesystem snapshots, and that rm in a later layer does not shrink an earlier one. That is the Linux filesystem model, applied. The Docker guide goes through multi-stage builds properly.
Notice what these have in common. In each case the tool reported a symptom in its own vocabulary, the actual cause was an operating system concept, and the fix went back into the tool. That round trip - tool, Linux, tool - is what a DevOps engineer does forty times a week.
Career impact
Strong Linux and scripting skills are non-negotiable for DevOps interviews and jobs - they are assumed from day one.
A practical way to build them: run a small Linux virtual machine or cloud instance, install and configure a web server on it by hand, break it deliberately, and fix it. That single exercise covers permissions, services, ports, logs and networking, and it is far more useful than working through a command list.
If you want a structure for the breaking, do it in this order over a weekend: change the port in the config to one already in use and watch it fail to start; chown the config to root and watch it fail differently; fill the disk with fallocate -l 20G /tmp/big and watch what breaks first; block the port with a firewall rule and confirm you can tell the difference between refused and timed out; then delete the default route and see how much stops working. Fix each one before moving to the next. By Sunday evening you will have seen the five most common production failures in a place where they cost nothing, and the commands above will have become reflexes instead of a list.
Linux is the first stop on the DevOps roadmap for a reason, and Linux with Git is week one of the DevOps certification course. Everything after it - Docker, the pipeline, Kubernetes, the cloud - assumes this is already in place. If you are working out where you stand, the DevOps engineer job role guide shows how Linux questions actually get asked in interviews, and the interview preparation strategy covers how to answer them.
Frequently Asked Questions
Why is Linux important for DevOps?
Because almost everything runs on it. Every container is a Linux process, every Kubernetes node is a Linux machine, every CI job runs in a Linux shell, every cloud instance is Linux you do not own. When a tool breaks, its abstractions stop being useful and what is left is the operating system. The engineers who are calm in incidents are the ones comfortable there.
How much Linux do I need to know for DevOps?
Enough to debug a service that will not start without looking anything up: read logs with journalctl and grep, check what is listening with ss, understand permissions and ownership from ls -l, know the difference between connection refused and timed out, and write a shell script that exits properly on errors. That is a few weeks of deliberate practice, not a certification.
Which Linux commands are most used in DevOps?
journalctl, systemctl, grep, tail, ss, ps, df, du, curl, ls -l, chmod and chown. Reading logs and searching text is the biggest category by far. Networking commands - ss, curl, dig, nc - are where most candidates are weakest and where the gap between adequate and good is widest.
Do I need to learn shell scripting for DevOps?
Yes. Bash is how you turn a thing you did by hand three times into something that runs itself. You do not need to be clever with it - you need set -euo pipefail at the top, quoted variables, argument checks and a rollback path. A script that handles errors properly is worth more than one that is elegant.
Should I learn Linux before Docker and Kubernetes?
Yes, and in that order. Docker is Linux processes with namespaces around them; Kubernetes schedules those processes across Linux nodes. A pod stuck in CrashLoopBackOff because of a permission denied error is a Linux problem in a Kubernetes costume. Learning the tools first means learning to recognise symptoms without understanding causes.
Which Linux distribution should I learn for DevOps?
Ubuntu or Debian for learning, because most container images and CI runners are Debian-based and most tutorials assume it. Know that Amazon Linux, RHEL and Rocky use dnf instead of apt and have some path differences, because you will meet them in enterprise and AWS environments. The commands that matter - systemd, journalctl, ss, permissions - are identical everywhere.
What is the difference between connection refused and connection timed out?
Refused means the packet reached the host and it actively said no - nothing is listening on that port. Timed out means the packet never got an answer - a firewall, security group or route is dropping it. Refused is a service problem, so go to the host and run ss -tulpn. Timed out is a network problem, so go to the security groups.
What does chmod 777 do and why is it bad?
It lets every user on the system read, write and execute the file. People reach for it when something says permission denied, but the real problem is almost always that the file belongs to the wrong user, and the fix is chown. 777 does not fix the cause and creates a security hole. On a private SSH key the permission must be 600 or SSH refuses to use it.
How do I check what is using a port in Linux?
sudo ss -tulpn | grep followed by the port number. It shows the listening socket and, with sudo, the process name and PID that owns it. netstat did the same job and still works on older systems, but ss is the current tool. Once you have the PID, ps -p with that PID shows who started it and when.
What is the fastest way to learn Linux for DevOps?
Run a small VM or cloud instance, install a web server by hand, then break it deliberately in five ways: wrong port, wrong file owner, full disk, blocked port, deleted default route. Fix each one. That weekend covers permissions, services, logs, networking and disk - the five most common production failures - somewhere it costs nothing.
Ready to Start Your DevOps Career?
Join our comprehensive DevOps + GenAI course with hands-on projects, live mentorship, and placement support