Exit code 137 alone does not prove an out-of-memory kill: Docker returns the container command’s exit status, and several paths can end at 137. Before restarting or raising memory, save the container ID, inspect state, Docker events, configured limits, and matching kernel log. Correlate State.OOMKilled with the same time window; never treat it as the whole diagnosis.
Contents
By LK Wood IV · 2026-10-03 · ~8 min read · St. Louis County, MO
Docker exit code 137 is a clue, not an OOM verdict. Under GNU Bash’s exit-status rules, a command terminated by fatal signal number N is reported as 128 + N. Linux lists SIGKILL as signal 9, which is why a shell-observed SIGKILL commonly appears as 137. That convention is not proof of the sender or cause: Docker returns the container command’s exit status, so a command can also return 137 itself. Docker documents at least two non-memory ways to send SIGKILL: docker kill uses it by default, and docker stop escalates to it when the process does not exit before its timeout.
Do not restart, recreate, or raise the memory limit yet. Those changes can erase the state and timing that separate a container limit from host pressure, a stop-timeout escalation, an operator kill, or an application-selected exit code.
The job is to bind every observation to the same container ID and time window. Capture first. Change one thing only after the evidence names a cause.
The four records to save first
Set the container name or ID, plus a bounded window that covers a few minutes before and after the failure. Use timestamps with an explicit timezone offset or Z; Docker says an unqualified timestamp uses the client machine’s local timezone.
CONTAINER='replace-with-name-or-id'
SINCE='2026-10-03T18:20:00-05:00'
UNTIL='2026-10-03T18:30:00-05:00'
First, save the stopped container’s identity, state and configured memory controls. docker inspect is read-only and works on stopped containers.
docker inspect --type=container \
--format 'id={{.Id}} name={{.Name}} image={{.Config.Image}} exit={{.State.ExitCode}} oom_killed={{.State.OOMKilled}} started={{.State.StartedAt}} finished={{.State.FinishedAt}} error={{json .State.Error}} memory={{.HostConfig.Memory}} reservation={{.HostConfig.MemoryReservation}} memory_swap={{.HostConfig.MemorySwap}} oom_kill_disable={{json .HostConfig.OomKillDisable}} restart_count={{.RestartCount}}' \
"$CONTAINER"
Second, capture Docker’s event history for that exact container and interval. Docker exposes separate oom, kill, die, stop and restart events. Its retained event history is limited to the last 256 events, which is another reason to collect it before doing more work on a busy host.
docker events \
--since "$SINCE" \
--until "$UNTIL" \
--filter type=container \
--filter container="$CONTAINER" \
--format '{{json .}}'
Third, save application output from the same interval. This can show an orderly shutdown request, an application-level allocation failure, or the last completed operation. Absence of a useful final line is not proof of a kernel kill.
docker logs --timestamps --since "$SINCE" --until "$UNTIL" "$CONTAINER"
Fourth, check the host kernel and Docker daemon around those timestamps. journalctl accepts space-separated date and time values for --since and --until; the separate UTC values below describe the same interval as the Docker variables above and work on systemd releases that do not accept that RFC 3339 form. These read-only queries explicitly select the current boot:
JOURNAL_SINCE='2026-10-03 23:20:00 UTC'
JOURNAL_UNTIL='2026-10-03 23:30:00 UTC'
journalctl -k --boot=0 --since "$JOURNAL_SINCE" --until "$JOURNAL_UNTIL" --no-pager \
| grep -Ei 'out of memory|oom-kill|killed process|memory cgroup'
journalctl -u docker.service --boot=0 --since "$JOURNAL_SINCE" --until "$JOURNAL_UNTIL" --no-pager
If the incident occurred before the current boot, run journalctl --list-boots, choose the boot ID whose recorded interval contains the incident, and replace --boot=0 in both commands with --boot=THE_32_CHARACTER_BOOT_ID. If the host does not use systemd, use its actual kernel and daemon log mechanism. Do not substitute an unbounded dmesg | tail and then assume the last message belongs to this incident.
Read the evidence as a set
| What you captured | What it establishes | What it does not establish |
|---|---|---|
Exit code 137 only | The container command reported 137 | Who or what chose that value |
State.OOMKilled=true | Docker recorded an OOM-killed state | Which memory boundary caused it, or how much memory is safe |
Docker oom event in the same window | Docker emitted an OOM event for that container | Whether the configured limit is simply too small, the host is overcommitted, or the workload is leaking |
kill followed by die, with the signal and matching time | A signal was sent in Docker’s event sequence | Who initiated it unless the surrounding daemon/orchestrator record says so |
| Kernel OOM record naming the victim/cgroup | The kernel selected a process under memory pressure | That blindly raising this container’s limit protects the host |
The exact cgroup’s local oom counter increased | An allocation in that cgroup was about to fail because of that cgroup’s memory limit | Which process was selected, or how much memory is safe |
cgroup oom_kill counter increase | An OOM killer killed a process in that cgroup | Which cgroup supplied the limiting boundary, or that the killed process was the container’s PID 1 |
State.OOMKilled=false and no memory evidence | No OOM cause has been demonstrated | That memory pressure is impossible |
That last row prevents a reasoning error. A false flag is not permission to declare “not memory,” just as a 137 is not permission to declare “memory.” Keep the cause unknown until another record closes the gap.
Separate a container boundary from host pressure
The memory value in the inspect record is the hard memory limit in bytes. A zero means Docker did not configure a per-container hard limit. It does not mean the host, VM, systemd unit, parent cgroup, or orchestrator is unlimited.
Docker’s resource documentation distinguishes a hard --memory limit from a softer --memory-reservation. It also warns that without constraints a container can use as much memory as the host kernel scheduler allows, and that host OOM can select processes when the machine cannot perform essential work.
Use this split:
- A nonzero container limit and matching Docker
oomevent do not identify the limiting boundary by themselves. A kernel OOM record naming that cgroup as the limiting boundary, or a same-window increase inoomfrom that exact cgroup’smemory.events.local, corroborates the container boundary. Anoom_killincrease alone only says a process in the cgroup was killed by an OOM killer. Once the boundary is identified, measure the workload and decide whether the limit, the application’s own memory configuration, or a leak is wrong. - No container hard limit plus a host kernel OOM record points to host or parent-boundary pressure. Giving one container more memory is not a fix; it has no private boundary to raise.
- A
killor stop-timeout sequence with no correlated memory evidence points away from an OOM diagnosis. Docker documents thatdocker stopsends the configured stop signal first and forcibly sendsSIGKILLafter the timeout. - No correlated record stays unresolved. Preserve the worksheet, increase observability, and wait for a second incident rather than editing limits by guesswork.
Read cgroup counters only where the host exposes them
Do not paste a path from someone else’s machine. Docker’s cgroup driver, systemd integration, parent cgroup and cgroup version determine the location. Locate the container’s cgroup through that host’s runtime configuration, replace the placeholder assignment below with the path you verified, then read the files that actually exist there. These commands only read those files.
For cgroup v2, read the following Linux kernel memory-interface files with the read-only commands below:
CGROUP_PATH='/replace/with/verified/container-cgroup-path'
cat "$CGROUP_PATH/memory.max"
cat "$CGROUP_PATH/memory.high"
cat "$CGROUP_PATH/memory.events"
cat "$CGROUP_PATH/memory.events.local"
memory.max is the hard limit. memory.high applies reclaim pressure but does not itself invoke the OOM killer. In memory.events, oom counts occasions when allocation was about to fail at the limit, while oom_kill counts processes in the cgroup killed by an OOM killer. The regular file is hierarchical; memory.events.local is local to that cgroup.
For a host still using cgroup v1, the kernel documents the older files below. The same documentation marks the v1 soft-limit and OOM-control interfaces deprecated, so read them as legacy evidence rather than advice for new configuration.
CGROUP_PATH='/replace/with/verified/container-cgroup-path'
cat "$CGROUP_PATH/memory.limit_in_bytes"
cat "$CGROUP_PATH/memory.usage_in_bytes"
cat "$CGROUP_PATH/memory.failcnt"
cat "$CGROUP_PATH/memory.oom_control"
If the container was removed or its cgroup disappeared, say that the counters were unavailable. Do not replace missing evidence with a guessed path or a counter from a new container.
Copy this incident worksheet before changing anything
Docker exit-137 incident
Observed at (timestamp + timezone):
Host name / VM:
Docker Engine version:
Container name:
Full container ID:
Image + immutable digest if available:
Compose project / service or launcher:
State
ExitCode:
OOMKilled:
StartedAt:
FinishedAt:
RestartCount:
State.Error:
Configured boundary (bytes as reported by inspect)
Memory:
MemoryReservation:
MemorySwap:
OomKillDisable:
Parent/orchestrator limit, if separately proven:
Same-window evidence
Docker oom event: yes / no / unavailable
Docker kill event + signal:
Docker die event + exitCode:
Stop/restart event immediately before death:
Kernel OOM record naming process/cgroup: yes / no / unavailable
cgroup version and exact path:
memory.events(.local) before change, if retained:
Application's last timestamped log line:
Classification
[ ] container/parent OOM is correlated
[ ] host OOM is correlated
[ ] stop timeout or external signal is correlated
[ ] application returned 137 without proven SIGKILL
[ ] unknown — preserve evidence and observe again
One change to test next:
Result and timestamp:
The empty boxes matter. “Unavailable” is evidence quality; it is not a “no.”
Choose the next action from the proven branch
Container or parent memory boundary proven: compare the configured boundary with the application’s measured peak and memory model. A Java heap, Node heap, database cache, shared memory, native allocation and filesystem cache do not all appear under one application setting. Fix the thing consuming memory or set a tested boundary with headroom. The Docker RAM planner estimates future host capacity for a known stack; it does not diagnose this incident and should not replace the capture above.
Host pressure proven: inventory every workload and host reserve before changing one container. Docker warns that the Linux kernel can kill processes, including Docker or other important applications, when the host cannot satisfy essential memory work. Moving the failure from one container to the whole host is regression, not remediation.
Stop-timeout escalation proven: find why PID 1 did not handle the configured stop signal or finish within the timeout. Docker notes that shell-form ENTRYPOINT and CMD run below /bin/sh -c, which does not pass signals to the executable. Fix signal handling or set a justified grace period; do not change memory because the final number happened to be 137.
Manual/orchestrator kill proven: trace the actor from the daemon, service, scheduler or automation log. The container behaved as instructed.
Cause still unknown: keep the worksheet and improve capture for the next occurrence. A restart may restore service, but it does not turn an unknown cause into an OOM diagnosis.
What I verified and what I did not
I verified the command syntax and field semantics against Docker’s current documentation and the cgroup file meanings against current Linux kernel documentation on 2026-10-03. I did not run an OOM experiment on a Docker host for this article, so none of the worksheet values or branches is presented as a TechFuelHQ measurement.
The article also does not assume Kubernetes. Pod eviction, Kubernetes OOMKilled status, pod limits and node-pressure decisions add another control plane and need their own evidence path.
Sources
- Docker Engine: container exit status — Docker returns the command’s exit status when the command runs and exits.
- GNU Bash: Exit Status — Bash reports a command terminated by fatal signal N as 128 + N; read 2026-10-03.
- Linux
signal(7)— standard signal behavior and the signal-number table identifyingSIGKILLas 9; read 2026-10-03. - Docker inspect — low-level object data, JSON output and Go-template field selection; read 2026-10-03.
- Docker system events — event types, time filters, container filters, JSON Lines output and 256-event retention; read 2026-10-03.
- Docker container kill — default
SIGKILLbehavior; read 2026-10-03. - Docker container stop — initial stop signal and forced
SIGKILLafter timeout; read 2026-10-03. - Docker resource constraints — no-limit default, hard and soft controls, host OOM risk, swap behavior and OOM-killer warning; read 2026-10-03.
- Docker Compose service attributes — current
mem_limit,mem_reservation,memswap_limitand OOM-related keys; read 2026-10-03. - systemd
journalctl— accepted time forms, current/selected boot filtering and boot inventory; read 2026-10-03. - systemd time syntax — calendar and UTC time syntax referenced by
journalctl; read 2026-10-03. - Linux kernel cgroup v2 memory controller —
memory.max,memory.high,memory.eventsandmemory.events.local; read 2026-10-03. - Linux kernel cgroup v1 memory controller — legacy limit, usage, failure and OOM-control files plus deprecation status; read 2026-10-03.
If the incident is actually disk pressure rather than memory pressure, the Docker cleanup guide separates disposable cache and images from volumes that can contain the only copy of data.
Frequently asked questions
Does Docker exit code 137 always mean out of memory?
Does State.OOMKilled true prove the container memory limit was too low?
Can a container be killed by memory pressure when State.OOMKilled is false?
Should I increase the memory limit after an exit 137?
Should I use oom_kill_disable to stop Docker OOM kills?
Sources and corrections
- Last updated
- Methodology
- See our methodology for research and review standards.
- Update log
- 2026-10-03 — Created a read-only incident-capture path for Docker exit 137. The article deliberately treats 137, State.OOMKilled, Docker oom events, kernel logs and cgroup counters as separate evidence instead of turning any one value into a universal diagnosis. Docker Engine CLI/resource documentation and Linux cgroup v1/v2 documentation were opened on 2026-10-03. No Docker daemon was available here, so no example output is presented as a TechFuelHQ measurement.
- Corrections
- Spotted an error or a stale number? Email contact@techfuelhq.com. Confirmed corrections are recorded here with the date and change.