Linux splits root's power into about 40 capabilities, such as CAP_CHOWN or CAP_NET_RAW. Docker 514 gives a container 14 of them by default, enough for most root-run software; --cap-drop removes some or ALL, and --cap-add gives back single ones. The effective set is a bitmask in /proc/self/status, which capsh decodes:
docker run --rm alpine:3 grep CapEff /proc/self/status
capsh --decode=a80425fb | cut -d= -f2 | tr ',' ' ' | fold -sw 90
docker run -d --name l3-web-nocap --cap-drop ALL l3-booknest-web:latest >/dev/null; sleep 2
docker logs l3-web-nocap 2>&1 | grep -m1 emerg | cut -d' ' -f3-
docker rm -f l3-web-nocap >/dev/null
docker run -d --name l3-web-caps --cap-drop ALL --cap-add CHOWN --cap-add SETUID \
--cap-add SETGID l3-booknest-web:latest >/dev/null; sleep 4
docker inspect l3-web-caps --format '{{.State.Health.Status}}'
docker rm -f l3-web-caps >/dev/nullCapEff: 00000000a80425fb
cap_chown cap_dac_override cap_fowner cap_fsetid cap_kill cap_setgid cap_setuid
cap_setpcap cap_net_bind_service cap_net_raw cap_sys_chroot cap_mknod cap_audit_write
cap_setfcap
[emerg] 1#1: chown("/var/cache/nginx/client_temp", 101) failed (1: Operation not permitted)
healthyWithout capabilities Nginx 75 's root master could not hand its cache directory to the nginx user; with CHOWN, SETUID and SETGID it starts and is healthy. It needs no NET_BIND_SERVICE for port 80, because Docker sets net.ipv4.ip_unprivileged_port_start=0 inside containers. The API needs nothing at all: a process running as a non-root user gets an empty effective set, so cap_drop: [ALL] costs it nothing and also empties the bounding set that a setuid binary could use to regain them. Start from ALL and add back what the logs ask for; never use --privileged, which grants everything.