Loki WAL at 100%: it wasn't the storage
The ingesters' WAL disks were full. It looked like a known object storage issue, but it was a checkpoint stuck in a disk-full loop.
Context
OpenShift logging stack (LokiStack) whose ingesters write a Write-Ahead Log to 150 GiB volumes before flushing to S3-compatible object storage.
Symptom
Both ingesters were Ready but their WAL volumes were at 100%. Months earlier a full S3 bucket had caused something similar, so that was the first suspect.
level=error caller=checkpoint.go:613 msg="error checkpointing series"
err="write /tmp/wal/00080651: no space left on device"
Diagnosis
- Ruled out storage: logs showed S3 flushes working, dozens of streams per minute and zero quota errors.
- The ingester logs showed the real pattern: every 5 minutes Loki tried to checkpoint and failed with “no space left on device”.
- Without a successful checkpoint Loki cannot truncate old WAL segments, so the disk never freed up. It had been looping for 4 days.
- There was also an interrupted checkpoint (.tmp) of more than 7 GB per ingester.
Root cause
A loop: full disk → checkpoint fails → no segment truncation → disk stays full. The object storage was healthy.
Fix
- Deleted the interrupted .tmp checkpoint, which is never used for replay.
- Not enough. df showed 95%, but the process runs as non-root and ext4 reserves ~5% for root: stat -f showed the space actually available.
- Also deleted the previous complete checkpoint after confirming S3 flushes were still active. A checkpoint is only a replay optimization; the WAL segments are the source of truth.
- The next automatic checkpoint completed and Loki truncated about 130 old segments. WAL went from 100% to 3%, with no pod restarts and no data loss.
Takeaways
- The same symptom as a past incident doesn’t mean the same cause: rule things out again from evidence.
- To measure the space available to a non-root process, use stat -f (Available), not just df.
- Knowing which state is an optimization and which is the source of truth lets you fix things without downtime.