AI
Architecture, multi-provider model loading, Copilot guardrails and mock sandbox.
VibeKit treats AI as a solved foundation layer rather than scattered vendor SDK calls.
The packages/ai workspace owns provider selection, model initialization, cancellation timeouts, and a unified facade. Product features consume this layer instead of directly instantiating vendor clients.
VibeKit maintains a strict distinction between two concepts:
- The Coding Agent: The AI agent operating this repository (using skills, rules, and tests).
- Application AI: User-facing inference features built into the running product (such as the Copilot widget or domain-specific generation).
Supported Model Providers
VibeKit includes native adapters for 10 model providers plus an offline mock engine. Each provider specifies an authoritative environment key and default model.
| Provider | Environment Variable | Default Model | Base URL / Notes |
|---|---|---|---|
| OpenRouter | OPENROUTER_API_KEY | deepseek/deepseek-v4-flash-0731 | https://openrouter.ai/api/v1 |
| OpenAI | OPENAI_API_KEY | gpt-5.6-sol | Supports OPENAI_BASE_URL |
| Anthropic | ANTHROPIC_API_KEY | claude-fable-5-1 | Anthropic SDK native |
| Google Gemini | GEMINI_API_KEY / GOOGLE_API_KEY | gemini-3.8-flash | Google Generative AI |
| Groq | GROQ_API_KEY | openai/gpt-oss-120b | High-speed inference |
| Mistral | MISTRAL_API_KEY | mistral-large-latest | Mistral AI platform |
| DeepSeek | DEEPSEEK_API_KEY | deepseek-v4-pro | Direct DeepSeek API |
| xAI | XAI_API_KEY | grok-4.6 | xAI platform |
| Cohere | COHERE_API_KEY | command-r-plus | Cohere platform |
| Together AI | TOGETHER_API_KEY | openai/gpt-oss-120b | https://api.together.ai/v1 |
| Mock AI | None (sandbox) | mock-model | Built-in offline responses |
Provider Resolution Hierarchy
When an application feature requests an AI model or completion, resolveConfiguredAiProvider evaluates credentials in strict order of precedence:
1. Explicit caller arguments (customProvider, custom credentials)
↓
2. Sandbox enforcement (MOCK_SERVICES=true or AI_PROVIDER='mock')
↓
3. Saved admin database settings (Admin UI service integrations)
↓
4. Environment variable override (AI_PROVIDER)
↓
5. Inferred active provider (first provider with an active API key)
↓
6. Fallback to mock AI (safe zero-credential default)Real-provider failures never silently fall back to mock AI. The mock sandbox is an explicit configuration choice, not an error recovery mechanism.
Application AI Facade
1. Simple Request/Response (chatCompletion)
For standard prompt/response tasks, use chatCompletion. It automatically applies a 30-second default deadline (DEFAULT_CHAT_COMPLETION_TIMEOUT_MS) and composes it with any caller-supplied AbortSignal:
import { chatCompletion } from "ai";
const response = await chatCompletion([
{ role: "system", content: "You are a concise assistant." },
{ role: "user", content: "Summarize the project status in 3 bullet points." },
]);
console.log(response.text);2. Streaming & Structured Outputs (getLanguageModel)
Advanced features that require streaming, token usage metadata, or structured schemas use getLanguageModel with Vercel AI SDK primitives:
import { getLanguageModel, streamText, generateObject } from "ai";
import { z } from "zod";
// Streaming responses
const model = await getLanguageModel();
const result = streamText({
model,
prompt: "Draft an announcement for our upcoming product release.",
});
// Structured JSON generation
const { object } = await generateObject({
model,
schema: z.object({
sentiment: z.enum(["positive", "neutral", "negative"]),
tags: z.array(z.string()),
urgency: z.number().min(1).max(5),
}),
prompt: "Analyze this customer feedback message: ...",
});Application Copilot
VibeKit includes an embedded, floating Copilot widget (CopilotChat.tsx) for in-app conversational assistance.
Admin Controls
Platform administrators can configure Copilot behavior directly from the Admin Settings dashboard without code changes:
aiCopilot: Global kill switch to enable or disable the Copilot widget across the application.aiAnonymous: Controls whether unauthenticated visitors may query the Copilot.aiAgentName: Custom name displayed in the chat header (default:Product AI).suggestedPrompts: Starter prompts displayed on empty chat sessions.
Production Guardrails & Ceilings
To prevent runaway billing and denial-of-wallet attacks, the server-side Copilot procedure enforces strict admission limits:
| Ceiling Variable | Default | Purpose |
|---|---|---|
AI_COPILOT_MAX_INPUT_CHARS | 8000 | Maximum allowed prompt character length per request |
AI_COPILOT_USER_DAILY_LIMIT | 200 | Maximum requests per authenticated user per UTC day |
AI_COPILOT_ANON_DAILY_LIMIT | 30 | Maximum requests per anonymous IP per UTC day |
AI_COPILOT_GLOBAL_DAILY_LIMIT | 2000 | Global requests accepted per UTC day across the instance |
AI_COPILOT_GLOBAL_PER_MINUTE | 60 | Sliding window rate limit across all callers |
AI_COPILOT_MAX_CONCURRENT | 8 | Maximum active concurrent completions in flight |
AI_COPILOT_GLOBAL_DAILY_TOKENS | 20000000 | Hard conservative daily token budget cap |
Privacy and Data Boundary
The Copilot minimizes server-added context before a provider call. Account email, raw backend exceptions and hidden authorization data are not added to the model prompt. Retrieved product context is bounded and sanitized before use.
A user message is still model input. Do not type passwords, API keys or other secrets into the chat.
Application code, not the model, owns identity and authorization. Team-scoped application data may only be read through server paths that already authorize the requesting user.
Grounding and Source Authority
Copilot does not treat every retrieved result as equally authoritative. The server resolves evidence in this order:
- Current canonical runtime state.
- Administrator-verified facts.
- Administrator-approved knowledge.
- Product documentation and published releases.
- Public product data such as roadmap or public feedback.
Higher-authority evidence wins when the existing source contract resolves a disagreement. When current evidence does not establish a product fact, Copilot should say what is unknown instead of guessing.
Semantic Actions and Approval
The site agent has a narrow semantic action set. It can currently propose navigation, waitlist signup, feedback submission and supported administrator setting changes.
The model is only the planner. Server code decides which actions are available for the current user, modules and route. Every write is validated and executed through the canonical application API.
- Navigation is limited to approved same-site routes.
- Joining the waitlist and submitting feedback use the final labeled button as the user's explicit confirmation. Behind that click, the application still mints a one-time approval capability before executing the write.
- Administrator module and site-setting changes show exact before and after values before approval.
- Approval capabilities are bound to the actor, action, normalized arguments, target, expiry and revision when a revision applies.
- A successful action is shown only after the canonical API returns a receipt. Replaying the same completed approval can return that stored receipt without executing the write twice.
Copilot does not have generic database, HTTP, browser or arbitrary tRPC write access.
Conversations, Traces and Corrections
Copilot conversations, messages, runs and approvals are durable so a session can be restored and an action can be audited. Run traces keep bounded diagnostic metadata such as prompt version, provider/model, evidence references, grounding status, selected tools, action lifecycle, timing, token usage, estimated cost and failure class. Raw provider/backend exceptions and credentials do not belong in traces.
Users can rate an answer. A negative rating can enter the Admin AI Operations review queue. When an administrator creates a correction from a reviewed run, the rejected answer is reference material only. The correction starts empty and must be written and explicitly published before it can become administrator-approved knowledge.
Model Context Protocol (MCP)
VibeKit provides a native Model Context Protocol server at /api/mcp.
This enables external developer agents (such as Claude Desktop or Cursor) to securely access product context, database schemas, and approved procedures using user-scoped OAuth consent grants.
Refer to the MCP Documentation for scope boundaries and client configuration.
Zero-Credential Mock Sandbox
For local development, CI pipelines, and review testing, VibeKit runs completely offline without third-party API keys:
# Enable mock sandbox in your .env or command line
MOCK_SERVICES=true
# Or select mock provider specifically
AI_PROVIDER=mockIn mock mode, chatCompletion returns deterministic fixture text immediately without network requests, token consumption, or external latency.
Agent-First Instructions
For a normal AI feature, keep using the shared provider facade in packages/ai. Do not import vendor SDKs directly into feature code.
When extending the site agent itself, keep the action contract narrow. Add a semantic tool only when a real user outcome requires it. Let server-owned capability state decide whether the tool exists for the current actor. Validate typed arguments, call the canonical application API, keep approval proportional to the write risk and return a real receipt. Add the smallest focused regression and include the behavior in bun run test:ai-evals.
Give your coding agent this prompt:
Extend the VibeKit site agent for this user outcome.
Reuse packages/api/features/ai instead of creating another agent runtime.
If a new action is required, make it a narrow typed semantic tool.
Keep identity, capability filtering, authorization, validation, approval and writes in application code.
Execute through the canonical application API and return a real receipt.
Do not add generic database, HTTP, browser or arbitrary write tools.
Add focused tests and update the existing AI eval corpus.
Verify the completed change with $verify-changes.