- Add comprehensive AI plans implementation documentation - Add LiteLLM gateway Kubernetes manifests - Update PostgreSQL and Forgejo deployment configs - Add build-from-binaries script
260 lines
7.6 KiB
Markdown
260 lines
7.6 KiB
Markdown
# AI Plans Implementation Plan for Nxtgauge
|
|
|
|
## 1. Goal
|
|
Add per-user AI plans with API keys and usage tracking for Ask Ash.
|
|
|
|
## 2. User Flow
|
|
1. User registers → gets a default Free plan
|
|
2. System generates one API key per user
|
|
3. User sends AI requests with their API key
|
|
4. Backend validates key, checks plan limits, forwards to LiteLLM
|
|
5. Usage is logged and credits are deducted
|
|
|
|
## 3. Plan Tiers
|
|
|
|
| Plan | Monthly Credits | Models | RPM | Max Tokens |
|
|
|------|-----------------|--------|-----|------------|
|
|
| Free | 100 | askash-fast, help-assistant, messenger | 10 | 1000 |
|
|
| Pro | 1000 | all 4B models | 60 | 4000 |
|
|
| Business | 5000 | all models (4B + 8B) | 120 | 8000 |
|
|
| Enterprise | 50000 | all + priority | unlimited | 32000 |
|
|
|
|
## 4. Database Tables
|
|
|
|
```sql
|
|
CREATE TABLE ai_plans (
|
|
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
|
|
name VARCHAR(50) UNIQUE NOT NULL,
|
|
display_name VARCHAR(100) NOT NULL,
|
|
monthly_credits INTEGER NOT NULL,
|
|
rate_limit_rpm INTEGER NOT NULL,
|
|
max_tokens_per_request INTEGER NOT NULL,
|
|
price_monthly DECIMAL(10,2) NOT NULL,
|
|
allowed_models JSONB NOT NULL,
|
|
is_active BOOLEAN DEFAULT true
|
|
);
|
|
|
|
CREATE TABLE user_ai_subscriptions (
|
|
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
|
|
user_id UUID UNIQUE NOT NULL REFERENCES users(id) ON DELETE CASCADE,
|
|
plan_id UUID NOT NULL REFERENCES ai_plans(id),
|
|
credits_remaining INTEGER NOT NULL,
|
|
credits_used_this_month INTEGER DEFAULT 0,
|
|
status VARCHAR(50) DEFAULT 'active',
|
|
period_start TIMESTAMP NOT NULL,
|
|
period_end TIMESTAMP NOT NULL,
|
|
created_at TIMESTAMP DEFAULT NOW(),
|
|
updated_at TIMESTAMP DEFAULT NOW()
|
|
);
|
|
|
|
CREATE TABLE api_keys (
|
|
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
|
|
user_id UUID NOT NULL REFERENCES users(id) ON DELETE CASCADE,
|
|
key_hash VARCHAR(255) UNIQUE NOT NULL,
|
|
key_prefix VARCHAR(50) NOT NULL,
|
|
is_active BOOLEAN DEFAULT true,
|
|
created_at TIMESTAMP DEFAULT NOW()
|
|
);
|
|
|
|
CREATE TABLE ai_usage_logs (
|
|
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
|
|
user_id UUID NOT NULL REFERENCES users(id),
|
|
api_key_id UUID REFERENCES api_keys(id),
|
|
model VARCHAR(100) NOT NULL,
|
|
request_type VARCHAR(100) NOT NULL,
|
|
tokens_input INTEGER NOT NULL,
|
|
tokens_output INTEGER NOT NULL,
|
|
tokens_total INTEGER NOT NULL,
|
|
credits_deducted INTEGER NOT NULL,
|
|
duration_ms INTEGER,
|
|
was_successful BOOLEAN DEFAULT true,
|
|
created_at TIMESTAMP DEFAULT NOW()
|
|
);
|
|
```
|
|
|
|
## 5. Models Available
|
|
|
|
| Model Alias | Ollama Model | Use Case |
|
|
|-------------|--------------|----------|
|
|
| askash-fast | qwen3:4b | Quick help, forms, chat |
|
|
| help-assistant | qwen3:4b | Help articles, platform guidance |
|
|
| messenger | qwen3:4b | Notifications, short messages |
|
|
| recommender | qwen3:4b | Job/professional matching |
|
|
| safety-check | qwen3:4b | Spam/abuse detection |
|
|
| askash-main | qwen3:8b | Resume, cover letters |
|
|
| jd-generator | qwen3:8b | Job descriptions |
|
|
| profile-writer | qwen3:8b | Profile completion |
|
|
| service-writer | qwen3:8b | Service descriptions |
|
|
| requirement-writer | qwen3:8b | Customer requirements |
|
|
| support-drafter | qwen3:8b | Support tickets |
|
|
| decision-support | qwen3:8b | Admin approvals |
|
|
| ultra-fast | gemma3:270m | Ultra-quick fallback |
|
|
|
|
## 6. Backend Services to Build
|
|
|
|
### 6.1 ApiKeyService
|
|
- `generate_key(user_id)` → returns `(full_key, hash, prefix)`
|
|
- `validate_key(provided_key, stored_hash)` → bool
|
|
- `save_key(user_id, hash, prefix)` → store in DB
|
|
- `get_active_key_for_user(user_id)` → Option<ApiKey>
|
|
|
|
### 6.2 PlanService
|
|
- `get_plan_by_name(name)` → AiPlan
|
|
- `create_subscription(user_id, plan_name)` → setup initial subscription
|
|
- `can_use_model(user_id, model)` → bool
|
|
- `get_subscription(user_id)` → UserAiSubscription
|
|
|
|
### 6.3 UsageService
|
|
- `can_make_request(user_id, model)` → checks credits + rate limit + model access
|
|
- `log_usage(user_id, api_key_id, model, request_type, tokens_input, tokens_output, duration_ms, success)` → deducts credits
|
|
- `get_usage_summary(user_id, start, end)` → usage stats
|
|
- `reset_monthly_credits()` → cron job at start of billing period
|
|
|
|
### 6.4 LiteLLM Client
|
|
- `chat_completion(model, messages)` → calls internal LiteLLM service
|
|
- returns tokens used + response
|
|
|
|
## 7. API Endpoints
|
|
|
|
### Public (requires API key)
|
|
```
|
|
POST /api/v1/ai/chat
|
|
Headers: Authorization: Bearer sk-nxtgauge-{user_id}-{random}
|
|
Body: { model, messages, request_type }
|
|
Response: { choices, usage, credits_remaining }
|
|
```
|
|
|
|
### Authenticated (requires user JWT)
|
|
```
|
|
GET /api/v1/ai/usage
|
|
POST /api/v1/ai/keys
|
|
GET /api/v1/ai/keys
|
|
DELETE /api/v1/ai/keys/{id}
|
|
POST /api/v1/ai/upgrade
|
|
GET /api/v1/ai/plans
|
|
```
|
|
|
|
## 8. Request Flow
|
|
|
|
```
|
|
User Request
|
|
↓
|
|
Nginx/Traefik
|
|
↓
|
|
API Gateway
|
|
↓
|
|
Extract API Key
|
|
↓
|
|
Validate API Key (lookup hash)
|
|
↓
|
|
Get User Subscription + Plan
|
|
↓
|
|
Check:
|
|
- Subscription active?
|
|
- Credits > 0?
|
|
- Model allowed?
|
|
- Rate limit OK?
|
|
↓
|
|
Forward to LiteLLM (internal master key)
|
|
↓
|
|
Parse response tokens
|
|
↓
|
|
Log usage + deduct credits
|
|
↓
|
|
Return response + X-Credits-Remaining header
|
|
```
|
|
|
|
## 9. Model Selection Helper
|
|
|
|
```rust
|
|
fn select_model(request_type: &str, user_plan: &str) -> &str {
|
|
match request_type {
|
|
"help" | "form_fill" | "validation" | "notification" => "askash-fast",
|
|
"job_recommendation" | "professional_match" => "recommender",
|
|
"safety_check" | "spam" => "safety-check",
|
|
"resume" | "cover_letter" | "profile_completion" => "askash-main",
|
|
"jd_generation" => "jd-generator",
|
|
"service_description" | "proposal" => "service-writer",
|
|
"requirement" => "requirement-writer",
|
|
"support_ticket" => "support-drafter",
|
|
"admin_approval" => "decision-support",
|
|
_ => "askash-fast",
|
|
}
|
|
}
|
|
```
|
|
|
|
## 10. Frontend Integration
|
|
|
|
- Display current plan in user dashboard
|
|
- Show credits remaining
|
|
- Show usage chart (daily/weekly/monthly)
|
|
- Upgrade plan button
|
|
- Reveal/regenerate API key button
|
|
|
|
## 11. Cron Jobs
|
|
|
|
- `reset_monthly_credits`: Run at start of each user's billing period
|
|
- `cleanup_old_usage_logs`: Archive logs older than 90 days
|
|
- `notify_low_credits`: Send email when credits below 20%
|
|
|
|
## 12. Testing Plan
|
|
|
|
- Unit tests for key generation and validation
|
|
- Unit tests for credit deduction
|
|
- Integration tests for rate limiting
|
|
- Load tests for concurrent requests
|
|
- Security tests (invalid keys, expired keys, plan downgrade)
|
|
|
|
## 13. Deployment Steps
|
|
|
|
1. Add database migrations
|
|
2. Deploy new backend version
|
|
3. Seed default plans
|
|
4. Generate API keys for existing users
|
|
5. Update frontend to show AI usage
|
|
6. Monitor for errors
|
|
|
|
## 14. MVP Scope (First Version)
|
|
|
|
- Free and Pro plans only
|
|
- Single API key per user
|
|
- Basic usage tracking
|
|
- Token-based credit deduction
|
|
- Monthly credit reset
|
|
|
|
## 15. Files to Create/Modify
|
|
|
|
### New Files
|
|
- `migrations/001_add_ai_plans.sql`
|
|
- `src/models/ai_plan.rs`
|
|
- `src/models/api_key.rs`
|
|
- `src/models/ai_usage.rs`
|
|
- `src/services/api_key_service.rs`
|
|
- `src/services/plan_service.rs`
|
|
- `src/services/usage_service.rs`
|
|
- `src/services/litellm_client.rs`
|
|
- `src/controllers/ai_controller.rs`
|
|
- `src/middleware/ai_auth.rs`
|
|
|
|
### Modified Files
|
|
- `src/main.rs` → add routes and services
|
|
- `src/routes.rs` → register AI routes
|
|
- existing user model → add plan relationship
|
|
|
|
## 16. Timeline
|
|
|
|
| Phase | Duration | Deliverable |
|
|
|-------|----------|-------------|
|
|
| 1 | 2-3 days | Database + models |
|
|
| 2 | 3-4 days | Services (keys, plans, usage) |
|
|
| 3 | 2-3 days | API endpoints + middleware |
|
|
| 4 | 2-3 days | Frontend usage UI |
|
|
| 5 | 2 days | Testing + deployment |
|
|
|
|
**Total: ~2 weeks for MVP**
|
|
|
|
## 17. Next Immediate Step
|
|
|
|
Create database migration and Rust models.
|
|
|
|
Approve this plan and I'll start Phase 1.
|