Commit graph

11 commits

Author SHA1 Message Date
sync-test
bc13ea25d1 fix(forgejo-runner-prune): actually reach buildx cache, bring under gitops management
All checks were successful
sync-to-github / sync (push) Successful in 42s
docker system prune --volumes can never touch the buildx builder's
cache volume - it's attached to a running container, and 'volumes'
prune only removes *unattached* ones. This CronJob ran every 30
minutes for 59 days reporting 'Total reclaimed space: 0B' every
single time while each runner's buildx cache silently grew to
57-98GB (discovered chasing a disk-pressure report on nxtgauge-2/3/4).

Fix: attach the real named builder ('nxtgauge-builder', matching
what CI actually uses) first, then call 'docker buildx prune'
directly, capped at --keep-storage 20GB so it doesn't just regrow
unbounded. Verified live: manually pruned gwsh7 (80%->16% node disk)
and ktst5 (84%->47%), then confirmed the patched CronJob is a true
no-op on an already-clean cache.

Also found and cleaned up (host-level, not gitops - out of band from
k8s): 18 orphaned buildx_buildkit_* containers on nxtgauge-1 from 2
months of ad-hoc 'docker buildx create' calls with no --name reuse,
totally invisible to buildx CLI and unrelated to CI (39.63GB, node
went 76%->50%). Added a daily cron job on that host
(~/.local/bin/docker-cleanup.sh) to keep it from reaccumulating,
since that's the host's standalone Docker daemon, outside k8s/gitops
entirely.

This CronJob + its RBAC (docker-prune-sa/-role/-binding) previously
existed only as manually-applied live objects, not tracked in git -
same pattern as nxtgauge-containerd-cleanup. Adding manifests and
wiring into clusters/production so Flux manages it going forward.
2026-08-16 04:11:21 +05:30
sync-test
c9fd51d5bc fix(containerd-cleanup): use k3s-bundled ctr, bring under gitops management
All checks were successful
sync-to-github / sync (push) Successful in 39s
The prune step called bare 'ctr', which has never existed as a
standalone binary on these k3s nodes - only the k3s binary itself
(bundling 'k3s ctr'/'k3s crictl' subcommands) lives at
/usr/local/bin. The failure was silently swallowed by '|| true',
so the CronJob has been reporting Complete every 10 minutes for
68 days while doing nothing; node disk sat at 83% used.

Verified live: patched the CronJob's command to invoke
/usr/local/bin/k3s ctr directly, manually triggered a run, confirmed
it actually prunes now (deleted ~75 dangling images).

This CronJob and its ServiceAccount previously existed only as a
manually-applied live object, not tracked in git. Adding the
manifests here and wiring them into clusters/production so Flux
manages it going forward instead of it silently drifting again.
2026-08-15 23:41:59 +05:30
Ashwin Kumar Sivakumar
fce1da5b3f chore: remove forgejo and registry dependencies 2026-06-15 01:52:43 +05:30
Ashwin Kumar Sivakumar
ad686f6075 fix: use Docker Hub for docker-cli image instead of private registry 2026-06-12 04:34:37 +05:30
Ashwin Kumar Sivakumar
c8fa8be29e fix: disable OpenObserve Telegram alerts 2026-06-12 04:14:28 +05:30
Ashwin Kumar Sivakumar
c4a7e1e330 chore: remove argocd and standardize on flux 2026-06-11 17:17:42 +05:30
Ashwin Kumar Sivakumar
37a589fa87 fix(backend): add PORT env to all rust deployments (was crashing on boot)
16 of 20 rust services had no PORT env var set; their main.rs calls
std::env::var('PORT').expect('PORT must be a valid u16') which panicked
on startup. This commit adds env.PORT matching the existing containerPort
for each service. Service ports: gateway=9100 users=9101 companies=9102
jobs=9103 job_seekers=9104 customers=9105 employees=9106 photographers=9107
tutors=9108 makeup_artists=9109 developers=9110 video_editors=9111
graphic_designers=9112 social_media_managers=9113 fitness_trainers=9114
catering_services=9115 payments=9116 ugc_content_creators=9117 leads=9118
2026-06-11 01:17:15 +05:30
Ashwin Kumar Sivakumar
a97cbf1743 fix: resolve registry.nxtgauge.com to traefik service for in-cluster pushes 2026-04-18 01:31:38 +05:30
Ashwin Kumar Sivakumar
1b4ef92083 fix: run db migrations as Argo PreSync hook + add openobserve collector/alerts 2026-04-17 17:10:37 +05:30
Ashwin Kumar Sivakumar
75acea11eb fix: registry ingress + woodpecker pulls + registry dns overrides 2026-04-17 05:25:04 +05:30
Ashwin Kumar
96bc5aa42a fix(registry): use node-resolvable backend registry endpoint and add k3s registries runbook 2026-04-11 21:59:50 +02:00