An APM webhook that broke every pipeline
Overnight, pipelines in dozens of namespaces started failing at the first step. The cause was a mutating webhook installed by a monitoring agent.
Context
OpenShift cluster with hundreds of Tekton pipelines. CI pods run as UID 65532 under a pipelines-specific SCC.
Symptom
PipelineRuns failing at the code clone step with two errors: PodAdmissionFailed for an out-of-range runAsUser, and Permission denied on the git user’s home.
PodAdmissionFailed: runAsUser: Invalid value: 65532:
must be in the ranges: [1000660000, 1000669999]
Diagnosis
- Confirmed it was cluster-wide: dozens of namespaces, always the same step, starting at a specific time.
- UID 65532 wasn’t new: every catalog ClusterTask uses it and the pipelines SCC allows it.
- Listing mutating webhooks revealed a new one from an APM agent installed the previous afternoon, intercepting every pod in the cluster.
- The webhook injected an init-container with a fixed UID and seccomp profile. Once mutated, SCC evaluation fell through to restricted-v2, which requires a UID within the namespace range.
- Used the apiserver audit log to rebuild the timeline: install, first failure, uninstall and first successful pipeline.
Root cause
A cluster-wide mutating webhook that also modified ephemeral CI pods, pushing them out of their intended SCC.
Fix
- The mitigation was uninstalling the agent’s operator; pipelines recovered within minutes.
- Documented a side effect: a manifest-update task without HOME set failed to copy SSH credentials. Proposed setting HOME=/tekton/home.
- Left a recommendation for any reinstall: exclude CI namespaces from agent injection.
Takeaways
- For mass admission failures, list mutating and validating webhooks first.
- The apiserver audit log lets you rebuild who changed what and when.
- Sidecar-injecting tools need an explicit scope: CI shouldn’t be included by default.