<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Troubleshooting case studies · Gastón Mardones</title>
    <link>https://gastonmardones.dev/en/cases/</link>
    <atom:link href="https://gastonmardones.dev/en/cases/rss.xml" rel="self" type="application/rss+xml" />
    <description>Real production incidents I diagnosed and solved. I publish a new one every month; internal details are anonymized.</description>
    <language>en-us</language>
    <item>
      <title>Console log view broken by stale managedFields</title>
      <link>https://gastonmardones.dev/en/cases/console-logs-managedfields/</link>
      <guid>https://gastonmardones.dev/en/cases/console-logs-managedfields/</guid>
      <pubDate>Mon, 05 Oct 2026 00:00:00 GMT</pubDate>
      <description>Several teams couldn't see their QA namespaces in Observe → Logs. It looked like RBAC; it was a Server-Side Apply conflict that had kept the observability operator degraded for months.</description>
      <category>OpenShift</category>
      <category>Operators</category>
      <category>Server-Side Apply</category>
      <category>RBAC</category>
    </item>
    <item>
      <title>Security audit across ~1,000 repositories, fully local</title>
      <link>https://gastonmardones.dev/en/cases/security-audit-at-scale/</link>
      <guid>https://gastonmardones.dev/en/cases/security-audit-at-scale/</guid>
      <pubDate>Fri, 25 Sep 2026 00:00:00 GMT</pubDate>
      <description>The goal was to find vulnerabilities and exposed secrets across the organization's entire codebase: not just source code, but also scripts, config and docs.</description>
      <category>DevSecOps</category>
      <category>Semgrep</category>
      <category>gitleaks</category>
      <category>Python</category>
    </item>
    <item>
      <title>An APM webhook that broke every pipeline</title>
      <link>https://gastonmardones.dev/en/cases/mutating-webhook-pipelines/</link>
      <guid>https://gastonmardones.dev/en/cases/mutating-webhook-pipelines/</guid>
      <pubDate>Thu, 10 Sep 2026 00:00:00 GMT</pubDate>
      <description>Overnight, pipelines in dozens of namespaces started failing at the first step. The cause was a mutating webhook installed by a monitoring agent.</description>
      <category>Tekton</category>
      <category>SCC</category>
      <category>Admission webhooks</category>
      <category>CI/CD</category>
    </item>
    <item>
      <title>Loki WAL at 100%: it wasn't the storage</title>
      <link>https://gastonmardones.dev/en/cases/loki-wal-checkpoint/</link>
      <guid>https://gastonmardones.dev/en/cases/loki-wal-checkpoint/</guid>
      <pubDate>Fri, 28 Aug 2026 00:00:00 GMT</pubDate>
      <description>The ingesters' WAL disks were full. It looked like a known object storage issue, but it was a checkpoint stuck in a disk-full loop.</description>
      <category>Loki</category>
      <category>OpenShift Logging</category>
      <category>Storage</category>
      <category>Troubleshooting</category>
    </item>
  </channel>
</rss>
