Tag
Kubernetes
Running production workloads on K3s — ArgoCD GitOps, Velero backups, network policies, and troubleshooting.
24 articles
- How a Single Volume Was 65% of My Velero Backup and What I Almost Excluded Instead26 Jul 2026The nfs-provisioner's root-mount volume was backing up the entire shared NFS export as one 33.4GB blob every night — redundant because every app's PVC was already backed up separately. The investigation almost went wrong when the first hypothesis pointed at the wrong disk.
- Zero NetworkPolicies on Vault: How I Found the Biggest Gap in My Cluster and a GitOps Tracking Bug That Hid It19 Jul 2026Vault is the trust root every ExternalSecret reads from. Its namespace had zero NetworkPolicies — any pod in the cluster could reach it. The investigation also uncovered a class of ArgoCD drift that silently ignores merged changes.
- 310 Restarts in 21 Days: CNPG's Silent PodMonitor Failure and the Leader-Election Trap17 Jul 2026CloudNativePG's auto-generated PodMonitor was missing a single label — Prometheus never scraped it. The same I/O fragility that causes etcd timeouts was triggering leader-election failures, restarting the operator 310 times in 21 days. Here's how I traced both root causes.
- Beszel: Lightweight Host Monitoring That Doesn't Deserve Its Own Server16 Jul 2026Why I replaced a heavyweight monitoring stack for host-level metrics with a single container, how Kyverno caught my first deploy before it hit production, and why internal-only dashboards don't need a public DNS record.
- Kubernetes Health Probes: The Host Header Trap That Restarts Healthy Pods15 Jul 2026Adding health probes to 20+ workloads taught me that kubelet sends the Pod IP as the Host header — and apps with host-validation reject it. Here's the full sweep, the gotcha that caught me, and the probe patterns that actually work.
- Migrating Atlantis to an LXC Accidentally Made It Fully Public14 Jul 2026Moving Atlantis from Kubernetes to a dedicated LXC involved repointing the Cloudflare Tunnel. The new tunnel pointed straight at the LXC's IP, bypassing Traefik and Authelia entirely. Atlantis has no auth of its own. For roughly 18 hours, anyone with the URL had full plan/apply access to the infrastructure repo.
- Renovate OOMKilled Three Times: Why the Fix Wasn't More Memory14 Jul 2026Two GiB wasn't enough, so I bumped to 3 GiB. Still OOMKilled. Bumped to 4 GiB. Still OOMKilled. The real fix wasn't memory at all — it was Terraform hash concurrency. Here's how I isolated the actual spike.
- Zero NetworkPolicies on the Database Namespace: The Gap That Let Any Pod Reach Authelia's Postgres13 Jul 2026The database namespace holding Authelia's session storage had no network restrictions. Any pod in the cluster could reach it. Here's the audit that found it, the traffic pattern that shaped the fix, and the namespace I deliberately left unrestricted.
- Full Observability on k3s: kube-prometheus-stack + Loki + Grafana OIDC04 Jul 2026Deploy a production-grade monitoring stack on bare-metal k3s: Prometheus, Loki with Garage S3 storage, Promtail on edge nodes via Ansible, SNMP monitoring for MikroTik, and Grafana SSO via Authelia OIDC - all GitOps-managed.
- Redis Killed Nextcloud and Nobody Noticed for Hours01 Jul 2026Redis running without a PVC still has persistence enabled by default. When it can't write RDB snapshots to a read-only rootfs, it doesn't crash - it silently refuses all writes. Here's how that turned into a full Nextcloud outage and why the logs pointed at the wrong thing first.
- I Hardened Pod securityContext and Broke 9 Containers in Production25 Jun 2026capabilities.drop: [ALL] and runAsNonRoot: true passed schema validation cleanly. Within minutes of merge, nine containers - including both Postgres instances backing Paperless and Nextcloud - were down. Here's the failure analysis, why a manual kubectl fix got silently undone, and the lesson for any blanket securityContext change.
- kubectl Said Everything Was Correct. Traefik 404'd Anyway.25 Jun 2026Migrating Jellyfin off k3s onto a GPU-passthrough LXC meant pointing a Service at an external IP. The EndpointSlice looked completely correct via kubectl - Service existed, endpoint listed right - but Traefik 404'd every request. A second, unrelated gotcha surfaced in the same migration: a PVC silently shared by reference across two unrelated files.
- ArgoCD Gotchas: Cache Staleness and the SharedResourceWarning Nobody Explains22 Jun 2026kubectl apply succeeds, the field reverts within seconds, and there's no error anywhere. Two ArgoCD debugging patterns that hit the same homelab three times in one day: repo-server cache staleness reverting live edits, and two Applications silently fighting over the same resource.
- I Ran Gitleaks Against My Own Repo and Found 12 Real Secrets22 Jun 2026A full-history gitleaks scan of a homelab repo that had been running for months turned up 12 distinct plaintext secrets - including an OIDC signing key. Here's the scanning setup, the baseline strategy that doesn't block on pre-existing leaks, and the remediation plan.
- k3s Backup Without the Complexity: Velero + Garage S3 on Longhorn20 Jun 2026Replace MinIO with Garage - a single 50MB binary - as the Velero backup target. Full daily cluster backups with Longhorn volume snapshots, deployed via ArgoCD.
- External Secrets Operator + HashiCorp Vault: GitOps Secret Lifecycle in Kubernetes18 Jun 2026Kubernetes Secrets are base64-encoded, not encrypted. Moving secrets out of the cluster into Vault - and syncing them back via External Secrets Operator - gives you rotation, audit logging, and compliance without changing how applications consume secrets.
- How a 1 GiB Memory Limit Took Down My Entire k3s Cluster18 Jun 2026A single misconfigured resource limit triggered a cascade: OOMKill on the control-plane, load average of 90, 1.2M DNS queries per day, and kubelet reporting the wrong allocatable memory. Here's the full post-mortem.
- Kyverno: Supply Chain Security as Admission Control on Kubernetes18 Jun 2026Most Kubernetes clusters accept any container image, any privilege level, and any resource configuration by default. Kyverno lets you enforce policies at admission time - before anything runs. Here's how to build a supply chain security baseline with Audit-first rollout.
- SLO Burn-Rate Alerting with Prometheus: Beyond Threshold Alerts18 Jun 2026Most teams alert when availability drops below a threshold. Burn-rate alerting tells you how fast you're spending your error budget - so you page on trajectory, not just current state. Here's how to implement the Google SRE Workbook approach on a bare-metal k3s cluster.
- Self-Hosted Tailscale Control Plane: Headscale on k3s with Authelia OIDC13 Jun 2026Deploy Headscale on a bare-metal k3s cluster with Longhorn persistence, Traefik ingress, and Authelia OIDC authentication - fully GitOps-managed via ArgoCD.
- Wildcard TLS Certificates on K3s with cert-manager and Cloudflare DNS22 May 2026How to automate wildcard Let's Encrypt certificates on a bare-metal K3s cluster using cert-manager's DNS-01 challenge with Cloudflare - and why HTTP-01 won't work for internal services.
- GitOps on K3s: Managing a Complete Homelab with ArgoCD20 May 2026How to manage an entire Kubernetes homelab - MetalLB, Traefik, Longhorn, Authelia, and more - as a Git repository using ArgoCD's App-of-Apps pattern.
- Bare-Metal LoadBalancer on K3s: MetalLB + Traefik with ArgoCD18 May 2026How to get a real external IP on a bare-metal Kubernetes cluster using MetalLB L2 mode, and wire it up with Traefik for automatic HTTPS - fully GitOps-managed with ArgoCD.
- Enterprise Homelab: K3s, Authelia & Longhorn on Proxmox with Terraform16 May 2026How to build a production-grade Kubernetes homelab with K3s, Authelia SSO, Longhorn storage, and ArgoCD - and the five painful mistakes that will cost you hours if you don't know about them.