Open source software gives you enormous leverage. Databases, web servers, message queues, dashboards, and monitoring tools that would cost a fortune to build in-house are available with source you can read and communities you can join. The friction is rarely licensing — it is operational. Docker solved a large part of that friction by turning "install this stack on a server and hope the versions line up" into "run this image and get the same result everywhere."
This guide walks through how to containerize the most common open source components in a modern application stack, which decisions matter when you choose images, how to handle state, and where teams usually go wrong. It is written for developers and small platform teams who want repeatable environments without hiring an entire Kubernetes department.
Why container-first thinking wins for open source stacks
The dependency problem containers actually solve
Traditional installation instructions for open source software assume a specific operating system, a specific compiler version, specific system libraries, and a specific directory layout. A PostgreSQL setup guide written for one distribution may fail on another because of a changed OpenSSL version or a different default locale. Multiply that by five services and you get an environment that only one person on the team can rebuild.
Containers change the unit of delivery. Instead of shipping instructions, you ship a filesystem and a process definition. The host only needs a container runtime. Everything above that — library versions, environment variables, user IDs, entrypoint scripts — travels with the image.
What containers are not
Containers are not virtual machines. They share the host kernel, which makes them fast to start and cheap to pack densely. That also means a kernel-level exploit or a misconfigured privileged container is a real risk. Containers are also not a backup strategy. If you run a database in a container without a tested backup path, you have simply moved the risk.
Finally, containers are not automatically reproducible. An image built from a floating tag such as latest is not reproducible at all. Pin what you ship.
When a container is the wrong answer
For a single long-running service on a single host with a stable environment, a system package plus a configuration management tool is often simpler. Containers add a build step, an image registry, and a runtime. If your team has no CI pipeline and no registry, start there before you start containerizing everything.
Docker building blocks you actually need
Images, layers, and the build cache
A Docker image is a stack of read-only layers. Each instruction in a Dockerfile creates a layer. The build cache reuses a layer when the instruction and its inputs are unchanged, which is why ordering matters: copy dependency manifests and install dependencies before copying application source. That single habit turns a two-minute rebuild into a two-second rebuild.
Use multi-stage builds for anything compiled. Build in a heavy image with toolchains, then copy only the artifact into a slim runtime image. A Go or Rust service that ships as a 20 MB image instead of a 900 MB one starts faster and has a smaller attack surface.
Volumes, bind mounts, and where data lives
Container filesystems are ephemeral by design. Two mechanisms persist data:
- Named volumes are managed by Docker and are the right default for database data directories.
- Bind mounts map a host path into the container and are convenient for configuration files and local development.
The classic mistake is mounting a named volume over a directory that the image populated at build time, then wondering why the application's default config vanished. Understand that mounting is a shadowing operation, not a merge.
Networks and service discovery
A user-defined bridge network gives every container a DNS name. Two services on the same network reach each other by container name, not by IP. That is the entire basis of the docker compose experience: a db service is reachable at hostname db from the api service.
Keep databases on an internal network with no published ports. Only the reverse proxy should be exposed to the outside world.
Compose as a contract
A compose.yaml file is documentation that executes. It should contain pinned image tags, health checks, restart policies, resource limits where relevant, and explicit environment variable sources. If a new developer can clone the repository and run one command to get a working stack, your Compose file is doing its job.
Choosing base images and building lean containers
Distroless, Alpine, and slim variants
The choice of base image affects size, security, and debugging comfort.
| Base | Size | Debugging | Notes |
|---|---|---|---|
| Full distribution image | Large | Easy | Shell and package manager included |
| Slim variant | Medium | Good | Fewer preinstalled tools |
| Alpine | Small | Moderate | musl libc can break native binaries |
| Distroless | Smallest | Hard | No shell, minimal CVE surface |
Alpine is popular, but musl libc occasionally causes issues with precompiled binaries and DNS resolution behavior. Test before committing.
Non-root users and file ownership
Run your process as a non-root user. Create the user in the Dockerfile and chown the application directory. On Kubernetes, matching runAsUser to the image's user avoids permission surprises with mounted volumes.
Image hygiene checklist
- Pin base images by digest for production builds.
- Keep secrets out of layers; use runtime environment variables or mounted secret files.
- Add a
.dockerignoreso node_modules and .git never enter the build context. - Set
HEALTHCHECKor define readiness probes. - Scan images for known vulnerabilities in CI and fail the build on critical findings.
Databases and stateful services in containers
PostgreSQL, MySQL, and MariaDB
The official PostgreSQL image supports initialization scripts placed in /docker-entrypoint-initdb.d, which run once when the data directory is empty. This is a clean way to create schemas, extensions, and read-only users on first boot.
Key configuration points:
- Set
POSTGRES_PASSWORDor, better,POSTGRES_PASSWORD_FILEpointing at a mounted secret. - Mount a named volume at the data directory path used by the image.
- Tune
shared_buffers,work_mem, and connection limits viapostgresql.confoverrides rather than guessing defaults. - Use
pg_dumpwith a cron container or a managed backup tool rather than copying the data directory of a running server.
MySQL and MariaDB follow the same pattern. Pay attention to character set and collation settings at initialization, because changing them later requires a dump and reload.
Redis and caching layers
Redis in a container is easy, which is exactly why people forget persistence. If the data matters, enable append-only file persistence and mount a volume. If it does not, set a memory limit and an eviction policy so the cache cannot starve the host.
Connection pooling and the container mindset
Containers scale horizontally, which multiplies database connections. Put a pooler such as PgBouncer in front of PostgreSQL, or use application-level pooling with sane maximums. A stack that worked with three app instances can fall over at thirty.
Migration strategy
Run schema migrations as a separate job, not as part of application startup. In Compose this can be a one-off service with depends_on and a completion condition. In Kubernetes it is a Job. This prevents several replicas from racing to apply the same migration.
Web servers, proxies, and application runtimes
NGINX, Caddy, and Traefik as entry points
NGINX remains the most common reverse proxy and static file server. Its container image supports templated configuration, and it can act as a TLS terminator, rate limiter, and cache.
Caddy is attractive when automatic HTTPS is a priority and configuration simplicity beats raw feature depth. Traefik shines in dynamic environments where services appear and disappear, because it discovers routes from container labels instead of a static config file.
Decision criteria:
- Static configuration, maximum familiarity, and tunability: NGINX.
- Automatic certificates with minimal config: Caddy.
- Frequent service churn and label-based routing: Traefik.
Node.js, Python, and PHP runtime patterns
For Node.js, use npm ci with a committed lockfile, run the build stage separately, and start with a non-root user. Keep the runtime image free of dev dependencies.
For Python, install into a virtual environment or use a tool that produces a lockfile, and avoid copying the entire project before dependency installation or you will invalidate the cache on every code change.
For PHP, separate the web server container from the PHP-FPM container and share application code through a read-only volume or bake it into both images. Baking it in is more predictable.
Connecting runtime to proxy
Use an internal network and never bind application ports to the host in production. The proxy is the only component with a published port. This single rule eliminates a surprising number of accidental exposures.
Automation: CI/CD pipelines with Docker
Building and pushing images reliably
A typical pipeline has four stages: lint and test, build and tag, scan, and deploy. Tag images with the commit SHA for traceability and add a human-readable tag for releases.
Use build caching in CI. Docker BuildKit with a registry-backed cache makes repeated builds dramatically faster, especially for compiled languages.
Compose in CI for integration tests
Integration tests often need a real database, a queue, and a cache. Instead of installing them on the CI runner, start them with Compose, wait for health checks, run the test suite, then tear down. This keeps CI behavior close to local behavior.
Deployment patterns
- Pull-based: the host pulls a new image and restarts the service.
- Push-based: CI connects to the runtime and updates workloads.
- GitOps: a repository declares desired state and an agent reconciles it.
GitOps tends to win once more than one environment exists, because the deployed state is always visible in version control.
Rollbacks that actually work
If you cannot roll back in one command, your deployment is not finished. Keep the previous image tag available, avoid irreversible migrations in the same release as risky code changes, and separate schema changes from behavior changes when possible.
Observability and logging stacks
Metrics with Prometheus and Grafana
Prometheus scrapes metrics endpoints and stores time series. Grafana visualizes them. Run both in containers, mount configuration and dashboards as code, and alert on symptoms users feel rather than on every threshold breach.
Expose application metrics on a separate port or path that is not publicly routable. Scrape only from inside the cluster network.
Logs with Loki or the ELK stack
Container logs go to standard output. Collect them centrally; do not rely on docker logs for anything beyond debugging. Loki is lighter and pairs naturally with Grafana. Elasticsearch, Logstash, and Kibana offer richer search at a higher operational cost.
Add structured logging with request identifiers so a single trace can be followed across the proxy, application, and database pooler.
Traces and profiling
OpenTelemetry has become the neutral way to emit traces, metrics, and logs. Instrument once and change backends without touching application code. Even sampling a small percentage of requests reveals latency sources that metrics alone hide.
Orchestration: Compose, Swarm, and Kubernetes
Decision criteria that hold up
Ask three questions:
- How many hosts will this run on within the next year?
- Does the team already operate a scheduler?
- How much downtime is acceptable during deploys?
A single host with a handful of services is a Compose problem. Multiple hosts with self-healing requirements is a Kubernetes problem. Do not adopt Kubernetes to avoid learning how your application behaves under load.
Kubernetes without the overwhelm
Start with Deployments, Services, ConfigMaps, Secrets, and a single Ingress. Add namespaces and resource requests early, because they are hard to retrofit. Leave operators, service meshes, and custom controllers until a concrete problem demands them.
Health checks and graceful shutdown
Readiness probes decide whether traffic reaches a pod. Liveness probes decide whether it restarts. Getting these wrong causes restart loops under load: a liveness probe that depends on a saturated dependency will kill healthy pods precisely when you need them. Handle SIGTERM, drain connections, and give termination a realistic grace period.
Security, backup, and troubleshooting checklist
Hardening without slowing delivery
- Run as non-root and drop unnecessary Linux capabilities.
- Make the root filesystem read-only where possible and mount a writable temp directory.
- Keep images small and rebuild regularly to pick up patched base images.
- Rotate secrets and never bake them into images.
- Limit egress from containers that only need internal access.
Backups and restore drills
A backup that has never been restored is a hypothesis. Schedule logical dumps, store them off-host, and run a restore drill into a scratch environment on a fixed cadence. Document the recovery time you actually measured, not the one you hoped for.
Troubleshooting common failures
Container exits immediately. Check the entrypoint and the logs from the failed run. Often a configuration file is missing or a required environment variable is unset.
Permission denied on a mounted volume. The container user's ID must match the ownership of the host path or volume.
Service cannot reach another service. Confirm they share a network and that you are using the service name, not localhost. Inside a container, localhost is the container itself.
Disk fills up unexpectedly. Old images, stopped containers, and build cache accumulate. Prune regularly and set log rotation limits.
DNS resolution is slow or flaky. Check the container's DNS configuration and, on Alpine-based images, verify that musl behavior is not the cause.
Frequently asked questions
Should I run databases in containers in production?
Yes, if you treat them as stateful workloads with volumes, backups, monitoring, and resource limits. Many teams run PostgreSQL and Redis in containers successfully. If you would rather not own replication, failover, and point-in-time recovery, a managed database is usually the better trade.
How do I keep my Compose file from becoming unmanageable?
Split configuration with override files for environment-specific values, use .env files for non-secret defaults, and move long-lived infrastructure to its own Compose project. If a file has more than roughly a dozen services, consider whether those services belong to the same lifecycle.
Do I need Kubernetes to be "production ready"?
No. Production readiness means automated deploys, health checks, monitoring, backups, and a rollback path. You can achieve all of that with Compose and a disciplined pipeline on one or two hosts.
How often should I rebuild images?
At minimum, rebuild when base images receive security patches, and on a scheduled cadence so you are not only rebuilding during incidents. Automate a weekly rebuild and let your scan stage catch problems early.
What is the biggest mistake teams make with containerized open source software?
Treating containers as magic. Every open source component still needs configuration, tuning, monitoring, and a backup plan. Containers make packaging consistent; they do not make operations disappear.
How do I choose between self-hosting and a managed service?
Compare total cost of ownership rather than license price. Self-hosting wins when you need deep control, data locality, or unusual configuration. Managed services win when your team's time is better spent on your product than on replication failover and storage tuning.
Bringing it together
The practical path looks like this: containerize one service at a time, pin every image, keep state on volumes with tested backups, put a single reverse proxy in front of everything, and automate build, scan, and deploy before you add more components. Add observability early, because you cannot tune what you cannot see. Add orchestration only when a second host or a hard availability requirement forces it.
Open source components reward teams that read the documentation and respect the operational details. Docker removes the installation lottery, and that is its real value. What remains — configuration, capacity, recovery, and security — is still your job, and it is where the difference between a stack that merely starts and a stack that survives its first bad day is decided.



