# AI Plans API Key Management for Nxtgauge ## Overview Each user gets a unique API key to track usage and enforce plan limits. ## API Key Structure ``` User ID: user_12345 Plan: free | pro | enterprise API Key: sk-nxtgauge-user_12345-abc123xyz ``` ## Implementation Approach Since LiteLLM Community Edition has limited virtual key features, we'll implement a custom middleware/proxy approach: ### Backend Implementation (Rust/Node.js) 1. **User Registration** → Generate API key 2. **API Key Validation** → Check against database 3. **Usage Tracking** → Increment counters per request 4. **Rate Limiting** → Enforce plan limits ### Database Schema ```sql -- Users table CREATE TABLE users ( id UUID PRIMARY KEY, email VARCHAR(255) UNIQUE, plan_type VARCHAR(50), -- 'free', 'pro', 'enterprise' api_key VARCHAR(255) UNIQUE, ai_credits_remaining INTEGER DEFAULT 100, monthly_usage_tokens INTEGER DEFAULT 0, created_at TIMESTAMP ); -- AI Usage tracking CREATE TABLE ai_usage ( id UUID PRIMARY KEY, user_id UUID REFERENCES users(id), model VARCHAR(100), -- 'askash-fast', 'askash-main', etc. tokens_input INTEGER, tokens_output INTEGER, request_type VARCHAR(100), -- 'help', 'resume', 'jd', etc. created_at TIMESTAMP ); -- API Keys table (for rotation) CREATE TABLE api_keys ( id UUID PRIMARY KEY, user_id UUID REFERENCES users(id), key_hash VARCHAR(255), is_active BOOLEAN DEFAULT true, created_at TIMESTAMP, expires_at TIMESTAMP ); ``` ### Plan Tiers | Plan | Price | Credits | Models Available | Rate Limit | |------|-------|---------|------------------|------------| | **Free** | $0 | 100/month | askash-fast only | 10 req/min | | **Pro** | $9/mo | 1000/month | All 4B models | 60 req/min | | **Business** | $29/mo | 5000/month | All models incl 8B | 120 req/min | | **Enterprise** | Custom | Unlimited | All + Priority | Unlimited | ### API Key Generation (Example in Rust) ```rust use uuid::Uuid; use rand::{distributions::Alphanumeric, Rng}; pub fn generate_api_key(user_id: &str) -> String { let random_suffix: String = rand::thread_rng() .sample_iter(&Alphanumeric) .take(24) .map(char::from) .collect(); format!("sk-nxtgauge-{}-{}", user_id, random_suffix) } // Example: sk-nxtgauge-user_12345-aBc3xYz9mNqP7rStUvWxYz12 ``` ### API Middleware Flow ``` User Request (with API key) ↓ [Nginx/Traefik Ingress] ↓ [API Gateway - Validate Key] ↓ [Check Plan & Limits] ├─ Check credits remaining ├─ Check rate limit (Redis) └─ Check model access ↓ [Route to LiteLLM] ↓ [Track Usage] ├─ Decrement credits ├─ Log usage to DB └─ Update metrics ↓ [Return Response] ``` ### Model Access by Plan ```yaml free_tier: models: - askash-fast # qwen3:4b - help-assistant # qwen3:4b - messenger # qwen3:4b max_tokens_per_request: 1000 pro_tier: models: - askash-fast - askash-main # qwen3:8b - help-assistant - jd-generator # qwen3:8b - profile-writer # qwen3:8b - recommender # qwen3:4b max_tokens_per_request: 4000 business_tier: models: - ALL_MODELS max_tokens_per_request: 8000 enterprise_tier: models: - ALL_MODELS - PRIORITY_QUEUE max_tokens_per_request: 32000 ``` ### Cost Calculation (Per 1K tokens) Since we're running local Ollama: - Cost is compute-based, not API-based - Track by GPU time or request duration - Alternative: flat rate per request type ```rust // Example pricing (based on compute cost) const PRICING: &[(str, f64)] = &[ ("askash-fast", 0.001), // $0.001 per 1K tokens ("askash-main", 0.005), // $0.005 per 1K tokens ("jd-generator", 0.008), // $0.008 per 1K tokens ("profile-writer", 0.008), // $0.008 per 1K tokens ]; ``` ### Usage Endpoints for Frontend ```javascript // Get user's current usage GET /api/v1/ai/usage Headers: Authorization: Bearer sk-nxtgauge-user_12345-... Response: { "plan": "pro", "credits_remaining": 750, "credits_used_this_month": 250, "requests_today": 45, "rate_limit": { "requests_per_minute": 60, "current_window": "58/60" } } ``` ### Implementation in Existing Backend Add to your Rust backend: 1. **Migration**: Add `api_key`, `ai_plan`, `ai_credits` columns to users table 2. **Middleware**: Create `AiAuthMiddleware` to validate keys 3. **Service**: Create `AiUsageService` to track and limit 4. **Endpoints**: - POST /api/v1/ai/chat (with API key auth) - GET /api/v1/ai/usage - POST /api/v1/ai/upgrade (change plan) ### Quick Start Commands ```bash # Generate API key for user curl -X POST https://api.nxtgauge.com/v1/ai/keys \ -H "Authorization: Bearer $USER_JWT" \ -d '{"plan": "pro"}' # Use API key curl https://llm.nxtgauge.com/v1/chat/completions \ -H "Authorization: Bearer sk-nxtgauge-user_12345-abc123" \ -H "Content-Type: application/json" \ -d '{ "model": "askash-main", "messages": [{"role": "user", "content": "Help with my resume"}] }' ``` ## Files to Create 1. `/src/services/ai_usage.rs` - Usage tracking service 2. `/src/middleware/ai_auth.rs` - API key validation 3. `/src/models/ai_plan.rs` - Plan definitions 4. Database migrations for API keys and usage tables ## Next Steps 1. Choose: Build custom middleware OR use LiteLLM Enterprise 2. Create database migrations 3. Implement API key generation 4. Add usage tracking middleware 5. Create billing integration ## LiteLLM Alternative For simpler setup, LiteLLM Enterprise ($500/mo) provides: - Built-in virtual keys - Usage dashboards - Team management - Budget controls - SSO/SAML But custom implementation gives more control and lower cost.