7 Common Mistakes When Deploying Docker Containers to Production

7 Common Mistakes When Deploying Docker Containers to Production

A container that behaves perfectly on your laptop can fall over within minutes in production. The difference is rarely the application code. It is the handful of small, unglamorous decisions — which user the process runs as, what happens when memory runs out, whether the orchestrator knows the app is actually responding — that decide whether a release is boring or eventful. Here are seven mistakes worth fixing before your next deployment.

1. Running the process as root

Most base images start you off as root. Unless you say otherwise, your Node, Python or Java process runs as UID 0 inside the container. A remote code execution bug then doesn't just compromise your app; it hands the attacker root inside the namespace, and a much shorter path to the host if anything sensitive is mounted carelessly.

Create a dedicated user in the Dockerfile, copy the application files with the right ownership, and switch with USER before the final CMD. Bind to a high port such as 8080 and let a reverse proxy or load balancer handle 80 and 443, because non-root users cannot bind below 1024. Then tighten the runtime: drop all Linux capabilities and add back only what is needed, set no-new-privileges, run with a read-only root filesystem and mount a small writable volume for temporary files. Never reach for --privileged just to make something start. If a container needs access to the Docker socket, pause and reconsider the design.

2. Deploying the latest tag

latest is a label, not a version. It moves under your feet, which means two deployments of the same compose file a week apart can run entirely different code. Rollbacks become guesswork, and nobody can say with confidence what is running in production right now.

Pin images to an immutable tag, or better still to a digest. Build once, tag with the Git commit SHA, and promote that same artefact from staging to production by retagging rather than rebuilding. Rebuilding between environments is how "it worked in staging" quietly becomes a lie: a different base image, a different dependency resolution, a different binary. Keep a short record of which SHA is live in each environment, and make the rollback a one-line change to a previously known-good tag.

3. Treating "running" as "healthy"

A container can be up while the application inside it is useless. The process is alive, but the database connection pool is exhausted, the migration failed, or the event loop is blocked. Without a health check, your load balancer or orchestrator keeps sending traffic to it anyway.

A container that reports itself healthy when it isn't causes more damage than one that crashes outright.

Add a real endpoint such as /healthz that checks what actually matters, and wire it up with HEALTHCHECK in the Dockerfile or with probes in your orchestrator. Be precise about the difference:

  • Liveness answers "should this container be restarted?" It should be cheap, local, and almost never depend on a downstream service, or a brief database blip will restart your whole fleet at once.
  • Readiness answers "should traffic be sent here right now?" This one can and should check dependencies such as the database and queue.
  • Startup gives slow-booting apps time to warm up before liveness checks begin, which prevents restart loops during deployment.

Set sensible timeouts and failure thresholds too. Three failures thirty seconds apart is very different from three failures two seconds apart.

4. Leaving resource limits unset

An unlimited container is a container that can take the host down with it. A memory leak will consume everything available until the kernel OOM killer intervenes, often choosing a victim you care about more. A busy loop or a runaway build can starve every other workload on the same machine.

Set memory and CPU requests and limits for every service, not just the ones that have misbehaved before. Remember that runtimes need to be told about those limits: a JVM sized for the host rather than the container, or a Node process with a default heap larger than its limit, will be killed long before it needs to be. Set the limit slightly above steady-state usage, watch your metrics for a couple of weeks, and adjust. If you see exit code 137, that is usually the OOM killer, not a random crash. Also watch ephemeral storage — a container that writes logs or temp files to its writable layer can fill a disk node and pull unrelated services down with it.

5. Baking secrets into images

Copying a .env file into the image, or passing credentials as build arguments, leaves them in the layer history for anyone who can pull the image. Environment variables passed at runtime are better, but they are still visible to anyone who can inspect the container or read a crash dump.

Use build-time secret mounts for anything needed during the build, and a secrets manager for runtime values — Vault, AWS Secrets Manager, Kubernetes Secrets with encryption at rest, or whatever your platform provides. Keep a strict .dockerignore so keys, certificates and local configuration never enter the build context in the first place. Rotate credentials on a schedule, give each service its own credentials, and make sure your logging pipeline is not quietly capturing tokens in request headers.

6. Writing logs inside the container

Log files written to the container's filesystem disappear when the container is replaced, which happens constantly. They also grow until the disk fills, and they are invisible to anyone debugging from outside the host.

Log to stdout and stderr and let the platform collect it. Emit structured lines, one event per line, so they can be queried rather than grepped. Configure your log driver with rotation limits such as a maximum file size and file count, and include a request or trace identifier so you can follow a single request across services. Treat PII carefully — production logs are often the least protected copy of your data.

7. Ignoring shutdown signals

When a deployment or scale-down happens, the orchestrator sends SIGTERM and waits. If your process ignores it, or never receives it because the shell is swallowing signals, the platform waits out the grace period and then kills the container mid-request. In-flight users see errors, and half-written data can be left behind.

Use the exec form of CMD or ENTRYPOINT so your application is PID 1 and receives signals directly. Add an init process such as tini if your app spawns children and does not reap them. Handle SIGTERM properly: stop accepting new connections, finish in-flight requests, close database pools and flush buffers, then exit. Set the grace period longer than your slowest realistic request.

A ten-minute check before you deploy

None of this requires a platform rewrite. Run through these before the next release:

  1. Confirm the process runs as a non-root user and that the root filesystem is read-only where possible.
  2. Check that every image is pinned to an immutable tag or digest.
  3. Test the health endpoint by breaking a dependency and watching what the orchestrator does.
  4. Verify memory and CPU limits are set, and that the runtime knows about them.
  5. Scan the image history for anything that looks like a credential.
  6. Send SIGTERM to a running container and watch the logs to see whether it shuts down cleanly.

Each fix is small. Together they turn deployments from an anxious event into routine housekeeping, which is exactly what you want.

Photo: panumas nikhomkhai / Pexels