nxtgauge-backend-rust/.forgejo/workflows/build.yaml
Ashwin Kumar Sivakumar ba1b0ebcc0
Some checks failed
build-and-release / build (cron) (push) Successful in 3m8s
build-and-release / build (gateway) (push) Successful in 5m7s
build-and-release / build (jobs) (push) Successful in 6m39s
build-and-release / build (payments) (push) Failing after 2m34s
build-and-release / build (graphic-designers) (push) Failing after 7m40s
build-and-release / build (job-seekers) (push) Failing after 7m38s
build-and-release / build (employees) (push) Failing after 8m8s
build-and-release / build (developers) (push) Failing after 8m10s
build-and-release / build (fitness-trainers) (push) Failing after 8m6s
build-and-release / build (leads) (push) Failing after 7m54s
build-and-release / build (photographers) (push) Failing after 1m17s
build-and-release / build (makeup-artists) (push) Failing after 7m52s
build-and-release / build (customers) (push) Failing after 13m56s
build-and-release / build (companies) (push) Failing after 17m4s
build-and-release / build (catering-services) (push) Failing after 17m5s
build-and-release / build (tutors) (push) Successful in 9m19s
build-and-release / build (social-media-managers) (push) Successful in 9m26s
build-and-release / build (ugc-content-creators) (push) Successful in 9m36s
build-and-release / build (video-editors) (push) Successful in 9m37s
build-and-release / build (users) (push) Successful in 12m26s
fix(ci): raise max-parallel to 9, make gitops push failures actually fail the job
Bumped runner.capacity from 1 to 3 on all 3 runner pods (9 total
concurrent slots - nodes were sitting at 7-11% CPU during builds, so
plenty of headroom), matched here with max-parallel: 9.

Also fixed a real bug: the gitops-push retry loop had no check after
exhausting all attempts, so a job whose every push attempt failed
would still exit 0 and report "success" - which is exactly what
happened on the previous run (verified: all 20 services built and
pushed their images correctly, but the actual GITOPS_PAT secret was
invalid/expired, and the retry loop silently swallowed the resulting
failure across all 20 jobs). Fixed the secret itself (confirmed the
existing admin-scoped Forgejo token has valid push access to
ashwin/nxtgauge-gitops and rotated GITOPS_PAT to it), and added an
explicit exit 1 if the retry loop exhausts without a successful push.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-08 02:30:36 +05:30

210 lines
8.3 KiB
YAML

name: build-and-release
on:
push:
branches:
- main
- high-performance
concurrency:
group: backend-build-${{ github.ref }}
cancel-in-progress: true
jobs:
build:
runs-on: docker-ready
strategy:
fail-fast: false
max-parallel: 9
matrix:
service:
- gateway
- users
- companies
- jobs
- leads
- job-seekers
- customers
- payments
- employees
- photographers
- makeup-artists
- tutors
- developers
- video-editors
- graphic-designers
- social-media-managers
- fitness-trainers
- catering-services
- ugc-content-creators
- cron
env:
DOCKER_BUILDKIT: "1"
steps:
- name: Checkout
uses: actions/checkout@v4
with:
fetch-depth: 0
# Static matrix (all 20 services every push) instead of a single job
# looping sequentially - up to 3 run concurrently (one per runner pod;
# each runner's capacity is 1). This step is each job's own quick
# "do I actually need to do anything" check, so unaffected services
# skip in a couple of seconds rather than sitting in a shared queue.
- name: Check if this service needs building
id: check
run: |
set -euo pipefail
service="${{ matrix.service }}"
svc_dir="$(echo "$service" | tr '-' '_')"
if git rev-parse --verify HEAD^ >/dev/null 2>&1; then
CHANGED_FILES="$(git diff --name-only HEAD^ HEAD)"
else
CHANGED_FILES="$(git ls-files)"
fi
LAST_COMMIT_MSG="$(git log -1 --pretty=%B | tr '\n' ' ')"
build=false
if echo "$LAST_COMMIT_MSG" | grep -Eiq 'trigger build|force build|rebuild all'; then
build=true
elif echo "$CHANGED_FILES" | grep -Eq '^(\.forgejo/workflows/|Dockerfile|Cargo\.toml|Cargo\.lock|crates/|scripts/)'; then
build=true
elif echo "$CHANGED_FILES" | grep -q "^apps/${svc_dir}/"; then
build=true
fi
echo "build=$build" >> "$GITHUB_OUTPUT"
if [ "$build" = "false" ]; then
echo "No changes relevant to $service - skipping."
fi
- name: Point DOCKER_HOST at this container's own gateway
if: steps.check.outputs.build == 'true'
run: |
set -euo pipefail
# 127.0.0.1 doesn't work: the job container is nested one level
# inside the runner pod's dind sidecar, so its own loopback isn't
# the sidecar's. --add-host=host.docker.internal:host-gateway is
# not being honored by this runner (tried via both global
# container.options and a per-job container: block — neither
# resolved), so read the container's actual default-route gateway
# directly instead, which is the dind engine that spawned it.
# Read directly from /proc/net/route instead of relying on the `ip`
# or `route` CLI tools being installed in the job image.
GATEWAY="$(awk '$2 == "00000000" {print $3}' /proc/net/route | head -1 | \
sed -E 's/(..)(..)(..)(..)/0x\4 0x\3 0x\2 0x\1/' | \
{ read -r a b c d; printf '%d.%d.%d.%d' "$a" "$b" "$c" "$d"; })"
echo "Detected docker host gateway: $GATEWAY"
echo "DOCKER_HOST=tcp://$GATEWAY:2375" >> "$GITHUB_ENV"
- name: Set up Docker Buildx
if: steps.check.outputs.build == 'true'
run: |
set -euo pipefail
docker version
# Same builder name on every node is safe: each runner has
# capacity 1, so only one job ever touches a given node's dind
# engine at a time - reusing the name lets its build cache persist
# (and warm up) across services scheduled on that node over time.
docker buildx create --use --name nxtgauge-builder || docker buildx use nxtgauge-builder
docker buildx inspect --bootstrap
- name: Login to registry
if: steps.check.outputs.build == 'true'
env:
REGISTRY_HOST: ${{ secrets.REGISTRY_HOST || 'ci.nxtgauge.com' }}
REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }}
REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }}
run: |
set -euo pipefail
printf '%s' "$REGISTRY_PASSWORD" | docker login "$REGISTRY_HOST" -u "$REGISTRY_USERNAME" --password-stdin
- name: Build and push
if: steps.check.outputs.build == 'true'
env:
REGISTRY_HOST: ${{ secrets.REGISTRY_HOST || 'ci.nxtgauge.com' }}
REGISTRY_NAMESPACE: ${{ secrets.REGISTRY_NAMESPACE || 'ashwin' }}
SHA: ${{ github.sha }}
run: |
set -euo pipefail
service="${{ matrix.service }}"
metadata_file="/tmp/${service}-metadata.json"
image_ref="$REGISTRY_HOST/$REGISTRY_NAMESPACE/nxtgauge-rust-${service}:${SHA}"
docker buildx build --push \
--metadata-file "$metadata_file" \
-f Dockerfile.simple \
--build-arg SERVICE_NAME="$service" \
-t "$image_ref" \
.
# buildx writes the metadata file pretty-printed (space after the
# colon), which the old compact-JSON-only pattern never matched -
# tolerate optional whitespace and extract the digest directly.
digest="$(grep -o '"containerimage\.digest"[[:space:]]*:[[:space:]]*"sha256:[^"]*"' "$metadata_file" | grep -o 'sha256:[^"]*')"
if [ -z "$digest" ]; then
echo "Failed to determine digest for $service" >&2
exit 1
fi
echo "DIGEST=$digest" >> "$GITHUB_ENV"
- name: Update GitOps release state
if: steps.check.outputs.build == 'true'
env:
GITOPS_SERVER: ${{ secrets.GITOPS_SERVER || 'ci.nxtgauge.com' }}
GITOPS_OWNER: ${{ secrets.GITOPS_OWNER || 'ashwin' }}
GITOPS_REPO: ${{ secrets.GITOPS_REPO || 'nxtgauge-gitops' }}
GITOPS_BRANCH: ${{ secrets.GITOPS_BRANCH || 'main' }}
# The gitops repo now lives on Forgejo (ci.nxtgauge.com), not
# GitHub - GITOPS_PUSH_USERNAME/GITOPS_PUSH_TOKEN were never
# actually configured as repo secrets (only the leftover
# GITOPS_GITHUB_* pair and GITOPS_PAT exist). Use GITOPS_PAT,
# which matches the current Forgejo-hosted repo.
GITOPS_PAT: ${{ secrets.GITOPS_PAT }}
SHA: ${{ github.sha }}
run: |
set -euo pipefail
test -n "${GITOPS_PAT:-}" || { echo "GITOPS_PAT is empty"; exit 1; }
service="${{ matrix.service }}"
git clone "https://forgejo-actions:${GITOPS_PAT}@${GITOPS_SERVER}/${GITOPS_OWNER}/${GITOPS_REPO}.git" /tmp/nxtgauge-gitops
cd /tmp/nxtgauge-gitops
# Up to 9 of these jobs can be pushing to the same gitops branch at
# once now - retry with a fresh pull+rebase on a non-fast-forward
# rejection instead of assuming we're the only writer.
pushed=false
for attempt in 1 2 3 4 5 6 7 8; do
git checkout "$GITOPS_BRANCH"
./scripts/set-backend-rust-release.sh "$service" "$DIGEST"
if git diff --quiet; then
echo "GitOps repo already up to date for $service."
pushed=true
break
fi
git config user.name "forgejo-actions[bot]"
git config user.email "forgejo-actions@ci.nxtgauge.com"
git add \
apps/nxtgauge-backend-rust/overlays/prod/backend-release-state.tsv \
apps/nxtgauge-backend-rust/overlays/prod/release-patches.yaml \
apps/nxtgauge-backend-rust/overlays/prod/disabled-deployments.yaml
git commit -m "chore(gitops): update ${service} image for ${SHA}"
if git push origin "HEAD:${GITOPS_BRANCH}"; then
pushed=true
break
fi
echo "Push rejected (attempt $attempt/8), pulling latest and retrying..."
git fetch origin "$GITOPS_BRANCH"
git reset --hard "origin/${GITOPS_BRANCH}"
sleep $((attempt * 2))
done
if [ "$pushed" != true ]; then
echo "Failed to push gitops update for $service after all retries" >&2
exit 1
fi