Some checks failed
build-and-release / build (companies) (push) Failing after 8s
build-and-release / build (customers) (push) Failing after 7s
build-and-release / build (developers) (push) Failing after 7s
build-and-release / build (employees) (push) Failing after 7s
build-and-release / build (fitness-trainers) (push) Failing after 8s
build-and-release / build (gateway) (push) Failing after 7s
build-and-release / build (graphic-designers) (push) Failing after 7s
build-and-release / build (job-seekers) (push) Failing after 8s
build-and-release / build (jobs) (push) Failing after 7s
build-and-release / build (leads) (push) Failing after 6s
build-and-release / build (makeup-artists) (push) Failing after 7s
build-and-release / build (payments) (push) Failing after 8s
build-and-release / build (photographers) (push) Failing after 7s
build-and-release / build (social-media-managers) (push) Failing after 7s
build-and-release / build (tutors) (push) Failing after 7s
build-and-release / build (ugc-content-creators) (push) Failing after 8s
build-and-release / build (users) (push) Failing after 7s
build-and-release / build (video-editors) (push) Failing after 7s
build-and-release / build (cron) (push) Failing after 3m19s
build-and-release / build (catering-services) (push) Failing after 5m54s
The build was structured as a single job looping through all ~20 services sequentially, so only 1 of the 3 deployed runner pods (one per worker node) was ever used - the other 2 sat idle for the entire build. Switched to a static per-service matrix (max-parallel: 3, matching the 3 runners at capacity 1 each) so independent services build concurrently. Each matrix job does its own quick "does this service need building" check up front (same change-detection logic, now per-job) rather than relying on a shared job output, to avoid needing cross-job artifact/output passing. GitOps updates also move into each matrix job (previously a single step at the end) since there's no longer one job aggregating all results - added a fetch/reset/retry loop since multiple jobs can now push to the same gitops branch concurrently. Tried adding cargo registry/target cache mounts to Dockerfile.simple for a bigger per-build win too, but measured it directly (local A/B: cold build 2m37s vs a second, supposedly-cached build 7m18s) and it made things slower here, likely cargo's own cache-verification pass outweighing the benefit for this dependency set - reverted that part, Dockerfile.simple is unchanged. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
198 lines
7.8 KiB
YAML
198 lines
7.8 KiB
YAML
name: build-and-release
|
|
|
|
on:
|
|
push:
|
|
branches:
|
|
- main
|
|
- high-performance
|
|
|
|
concurrency:
|
|
group: backend-build-${{ github.ref }}
|
|
cancel-in-progress: true
|
|
|
|
jobs:
|
|
build:
|
|
runs-on: docker-ready
|
|
strategy:
|
|
fail-fast: false
|
|
max-parallel: 3
|
|
matrix:
|
|
service:
|
|
- gateway
|
|
- users
|
|
- companies
|
|
- jobs
|
|
- leads
|
|
- job-seekers
|
|
- customers
|
|
- payments
|
|
- employees
|
|
- photographers
|
|
- makeup-artists
|
|
- tutors
|
|
- developers
|
|
- video-editors
|
|
- graphic-designers
|
|
- social-media-managers
|
|
- fitness-trainers
|
|
- catering-services
|
|
- ugc-content-creators
|
|
- cron
|
|
env:
|
|
DOCKER_BUILDKIT: "1"
|
|
steps:
|
|
- name: Checkout
|
|
uses: actions/checkout@v4
|
|
with:
|
|
fetch-depth: 0
|
|
|
|
# Static matrix (all 20 services every push) instead of a single job
|
|
# looping sequentially - up to 3 run concurrently (one per runner pod;
|
|
# each runner's capacity is 1). This step is each job's own quick
|
|
# "do I actually need to do anything" check, so unaffected services
|
|
# skip in a couple of seconds rather than sitting in a shared queue.
|
|
- name: Check if this service needs building
|
|
id: check
|
|
run: |
|
|
set -euo pipefail
|
|
service="${{ matrix.service }}"
|
|
svc_dir="$(echo "$service" | tr '-' '_')"
|
|
|
|
if git rev-parse --verify HEAD^ >/dev/null 2>&1; then
|
|
CHANGED_FILES="$(git diff --name-only HEAD^ HEAD)"
|
|
else
|
|
CHANGED_FILES="$(git ls-files)"
|
|
fi
|
|
LAST_COMMIT_MSG="$(git log -1 --pretty=%B | tr '\n' ' ')"
|
|
|
|
build=false
|
|
if echo "$LAST_COMMIT_MSG" | grep -Eiq 'trigger build|force build|rebuild all'; then
|
|
build=true
|
|
elif echo "$CHANGED_FILES" | grep -Eq '^(\.forgejo/workflows/|Dockerfile|Cargo\.toml|Cargo\.lock|crates/|scripts/)'; then
|
|
build=true
|
|
elif echo "$CHANGED_FILES" | grep -q "^apps/${svc_dir}/"; then
|
|
build=true
|
|
fi
|
|
|
|
echo "build=$build" >> "$GITHUB_OUTPUT"
|
|
if [ "$build" = "false" ]; then
|
|
echo "No changes relevant to $service - skipping."
|
|
fi
|
|
|
|
- name: Point DOCKER_HOST at this container's own gateway
|
|
if: steps.check.outputs.build == 'true'
|
|
run: |
|
|
set -euo pipefail
|
|
# 127.0.0.1 doesn't work: the job container is nested one level
|
|
# inside the runner pod's dind sidecar, so its own loopback isn't
|
|
# the sidecar's. --add-host=host.docker.internal:host-gateway is
|
|
# not being honored by this runner (tried via both global
|
|
# container.options and a per-job container: block — neither
|
|
# resolved), so read the container's actual default-route gateway
|
|
# directly instead, which is the dind engine that spawned it.
|
|
# Read directly from /proc/net/route instead of relying on the `ip`
|
|
# or `route` CLI tools being installed in the job image.
|
|
GATEWAY="$(awk '$2 == "00000000" {print $3}' /proc/net/route | head -1 | \
|
|
sed -E 's/(..)(..)(..)(..)/0x\4 0x\3 0x\2 0x\1/' | \
|
|
{ read -r a b c d; printf '%d.%d.%d.%d' "$a" "$b" "$c" "$d"; })"
|
|
echo "Detected docker host gateway: $GATEWAY"
|
|
echo "DOCKER_HOST=tcp://$GATEWAY:2375" >> "$GITHUB_ENV"
|
|
|
|
- name: Set up Docker Buildx
|
|
if: steps.check.outputs.build == 'true'
|
|
run: |
|
|
set -euo pipefail
|
|
docker version
|
|
# Same builder name on every node is safe: each runner has
|
|
# capacity 1, so only one job ever touches a given node's dind
|
|
# engine at a time - reusing the name lets its build cache persist
|
|
# (and warm up) across services scheduled on that node over time.
|
|
docker buildx create --use --name nxtgauge-builder || docker buildx use nxtgauge-builder
|
|
docker buildx inspect --bootstrap
|
|
|
|
- name: Login to registry
|
|
if: steps.check.outputs.build == 'true'
|
|
env:
|
|
REGISTRY_HOST: ${{ secrets.REGISTRY_HOST || 'ci.nxtgauge.com' }}
|
|
REGISTRY_USERNAME: ${{ secrets.REGISTRY_USERNAME }}
|
|
REGISTRY_PASSWORD: ${{ secrets.REGISTRY_PASSWORD }}
|
|
run: |
|
|
set -euo pipefail
|
|
printf '%s' "$REGISTRY_PASSWORD" | docker login "$REGISTRY_HOST" -u "$REGISTRY_USERNAME" --password-stdin
|
|
|
|
- name: Build and push
|
|
if: steps.check.outputs.build == 'true'
|
|
env:
|
|
REGISTRY_HOST: ${{ secrets.REGISTRY_HOST || 'ci.nxtgauge.com' }}
|
|
REGISTRY_NAMESPACE: ${{ secrets.REGISTRY_NAMESPACE || 'ashwin' }}
|
|
SHA: ${{ github.sha }}
|
|
run: |
|
|
set -euo pipefail
|
|
service="${{ matrix.service }}"
|
|
metadata_file="/tmp/${service}-metadata.json"
|
|
image_ref="$REGISTRY_HOST/$REGISTRY_NAMESPACE/nxtgauge-rust-${service}:${SHA}"
|
|
|
|
docker buildx build --push \
|
|
--metadata-file "$metadata_file" \
|
|
-f Dockerfile.simple \
|
|
--build-arg SERVICE_NAME="$service" \
|
|
-t "$image_ref" \
|
|
.
|
|
|
|
# buildx writes the metadata file pretty-printed (space after the
|
|
# colon), which the old compact-JSON-only pattern never matched -
|
|
# tolerate optional whitespace and extract the digest directly.
|
|
digest="$(grep -o '"containerimage\.digest"[[:space:]]*:[[:space:]]*"sha256:[^"]*"' "$metadata_file" | grep -o 'sha256:[^"]*')"
|
|
if [ -z "$digest" ]; then
|
|
echo "Failed to determine digest for $service" >&2
|
|
exit 1
|
|
fi
|
|
echo "DIGEST=$digest" >> "$GITHUB_ENV"
|
|
|
|
- name: Update GitOps release state
|
|
if: steps.check.outputs.build == 'true'
|
|
env:
|
|
GITOPS_SERVER: ${{ secrets.GITOPS_SERVER || 'ci.nxtgauge.com' }}
|
|
GITOPS_OWNER: ${{ secrets.GITOPS_OWNER || 'ashwin' }}
|
|
GITOPS_REPO: ${{ secrets.GITOPS_REPO || 'nxtgauge-gitops' }}
|
|
GITOPS_BRANCH: ${{ secrets.GITOPS_BRANCH || 'main' }}
|
|
GITOPS_PUSH_USERNAME: ${{ secrets.GITOPS_PUSH_USERNAME }}
|
|
GITOPS_PUSH_TOKEN: ${{ secrets.GITOPS_PUSH_TOKEN }}
|
|
SHA: ${{ github.sha }}
|
|
run: |
|
|
set -euo pipefail
|
|
test -n "${GITOPS_PUSH_TOKEN:-}" || { echo "GITOPS_PUSH_TOKEN is empty"; exit 1; }
|
|
service="${{ matrix.service }}"
|
|
|
|
git clone "https://${GITOPS_PUSH_USERNAME}:${GITOPS_PUSH_TOKEN}@${GITOPS_SERVER}/${GITOPS_OWNER}/${GITOPS_REPO}.git" /tmp/nxtgauge-gitops
|
|
cd /tmp/nxtgauge-gitops
|
|
|
|
# Up to 3 of these jobs can be pushing to the same gitops branch at
|
|
# once now - retry with a fresh pull+rebase on a non-fast-forward
|
|
# rejection instead of assuming we're the only writer.
|
|
for attempt in 1 2 3 4 5; do
|
|
git checkout "$GITOPS_BRANCH"
|
|
./scripts/set-backend-rust-release.sh "$service" "$DIGEST"
|
|
|
|
if git diff --quiet; then
|
|
echo "GitOps repo already up to date for $service."
|
|
break
|
|
fi
|
|
|
|
git config user.name "forgejo-actions[bot]"
|
|
git config user.email "forgejo-actions@ci.nxtgauge.com"
|
|
git add \
|
|
apps/nxtgauge-backend-rust/overlays/prod/backend-release-state.tsv \
|
|
apps/nxtgauge-backend-rust/overlays/prod/release-patches.yaml \
|
|
apps/nxtgauge-backend-rust/overlays/prod/disabled-deployments.yaml
|
|
git commit -m "chore(gitops): update ${service} image for ${SHA}"
|
|
|
|
if git push origin "HEAD:${GITOPS_BRANCH}"; then
|
|
break
|
|
fi
|
|
|
|
echo "Push rejected (attempt $attempt/5), pulling latest and retrying..."
|
|
git fetch origin "$GITOPS_BRANCH"
|
|
git reset --hard "origin/${GITOPS_BRANCH}"
|
|
sleep $((attempt * 2))
|
|
done
|