feat(ai): fix Ollama resource limits and bring LiteLLM under GitOps management
All checks were successful
sync-to-forgejo / sync (push) Successful in 10s
All checks were successful
sync-to-forgejo / sync (push) Successful in 10s
Phase 1 of the AI architecture doc ("Improve Generation Quality") —
qwen3:4b and qwen3:8b were already pulled onto the Ollama PVC, and
apps/litellm/base/configmap.yaml already had the correct model_list
mapping every feature alias to them instead of gemma3:270m. Neither
was actually in effect:
1. apps/litellm was never included in
clusters/production/kustomization.yaml, so it was only ever
deployed by a one-off manual `kubectl apply` and has been
completely outside GitOps ever since (same root cause as the
ai-guard registry drift found earlier). Added it to the root
kustomization. Corrected its image reference from
registry.nxtgauge.com/litellm:latest (doesn't appear to exist) to
ghcr.io/berriai/litellm:latest, matching what's actually running
live — adopting this file without that fix would have broken a
working deployment the moment Flux started managing it.
2. apps/ollama/base/deployment.yaml's memory limit (1500Mi) was too
small to ever load qwen3:4b (~2.5GB) or qwen3:8b (~5.2GB) — every
model alias in the (also-never-applied) LiteLLM config was
therefore unusable regardless of what it was named. Raised to
4 CPU / 8Gi limit (node has 16GB total, was at ~26% memory use) and
added OLLAMA_KEEP_ALIVE=30m so a loaded model survives the gaps
between bursty feature requests instead of reloading from disk on
every first call after 5+ minutes idle.
This commit is contained in:
parent
d52b911aba
commit
d503a32de7
3 changed files with 28 additions and 5 deletions
|
|
@ -17,7 +17,15 @@ spec:
|
|||
spec:
|
||||
containers:
|
||||
- name: litellm
|
||||
image: registry.nxtgauge.com/litellm:latest
|
||||
# This app was never included in Flux's root kustomization (see
|
||||
# clusters/production/kustomization.yaml), so it was only ever
|
||||
# deployed by a one-off manual `kubectl apply` and has since
|
||||
# drifted from this file. Corrected to match what's actually
|
||||
# running live (the real upstream image) rather than
|
||||
# registry.nxtgauge.com/litellm:latest, which doesn't appear to
|
||||
# exist/be maintained — using it would have broken a working
|
||||
# deployment the moment this file was wired back into GitOps.
|
||||
image: ghcr.io/berriai/litellm:latest
|
||||
command:
|
||||
- "/bin/bash"
|
||||
args:
|
||||
|
|
|
|||
|
|
@ -24,16 +24,30 @@ spec:
|
|||
env:
|
||||
- name: OLLAMA_HOST
|
||||
value: "0.0.0.0:11434"
|
||||
# Keep a loaded model resident for 30 min of inactivity instead
|
||||
# of Ollama's 5-minute default — job-description/resume/cover-
|
||||
# letter traffic is bursty, and reloading a 2.5-5GB model from
|
||||
# disk on every request would add multi-second latency to each
|
||||
# first call after a gap.
|
||||
- name: OLLAMA_KEEP_ALIVE
|
||||
value: "30m"
|
||||
volumeMounts:
|
||||
- name: ollama-models
|
||||
mountPath: /root/.ollama
|
||||
resources:
|
||||
requests:
|
||||
cpu: 500m
|
||||
memory: 700Mi
|
||||
limits:
|
||||
cpu: 1000m
|
||||
memory: 1500Mi
|
||||
memory: 3Gi
|
||||
limits:
|
||||
# qwen3:4b (~2.5GB on disk) and qwen3:8b (~5.2GB) are already
|
||||
# pulled onto the PVC, but the previous 1500Mi limit could
|
||||
# only ever load gemma3:270m — which is why every LiteLLM
|
||||
# model alias was mapped to gemma3:270m regardless of name
|
||||
# (see apps/litellm/base/configmap.yaml). Sized to comfortably
|
||||
# hold qwen3:8b plus KV cache/runtime overhead, with headroom;
|
||||
# node has 16GB total and was at ~26% memory use.
|
||||
cpu: 4000m
|
||||
memory: 8Gi
|
||||
volumes:
|
||||
- name: ollama-models
|
||||
persistentVolumeClaim:
|
||||
|
|
|
|||
|
|
@ -8,6 +8,7 @@ resources:
|
|||
- ../../apps/nxtgauge-ai-assistant/overlays/prod
|
||||
- ../../apps/github-actions-runners/base
|
||||
- ../../apps/ollama/base
|
||||
- ../../apps/litellm/base
|
||||
- ../../ops/openobserve-alerts
|
||||
- flux-system/traceworks2026-image-automation.yaml
|
||||
- ../../apps/traceworks2026/overlays/prod
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue