VibeKit

AI

Architecture, multi-provider model loading, Copilot guardrails and mock sandbox.

VibeKit treats AI as a solved foundation layer rather than scattered vendor SDK calls.

The packages/ai workspace owns provider selection, model initialization, cancellation timeouts, and a unified facade. Product features consume this layer instead of directly instantiating vendor clients.

VibeKit maintains a strict distinction between two concepts:

  • The Coding Agent: The AI agent operating this repository (using skills, rules, and tests).
  • Application AI: User-facing inference features built into the running product (such as the Copilot widget or domain-specific generation).

Supported Model Providers

VibeKit includes native adapters for 10 model providers plus an offline mock engine. Each provider specifies an authoritative environment key and default model.

ProviderEnvironment VariableDefault ModelBase URL / Notes
OpenRouterOPENROUTER_API_KEYdeepseek/deepseek-v4-flash-0731https://openrouter.ai/api/v1
OpenAIOPENAI_API_KEYgpt-5.6-solSupports OPENAI_BASE_URL
AnthropicANTHROPIC_API_KEYclaude-fable-5-1Anthropic SDK native
Google GeminiGEMINI_API_KEY / GOOGLE_API_KEYgemini-3.8-flashGoogle Generative AI
GroqGROQ_API_KEYopenai/gpt-oss-120bHigh-speed inference
MistralMISTRAL_API_KEYmistral-large-latestMistral AI platform
DeepSeekDEEPSEEK_API_KEYdeepseek-v4-proDirect DeepSeek API
xAIXAI_API_KEYgrok-4.6xAI platform
CohereCOHERE_API_KEYcommand-r-plusCohere platform
Together AITOGETHER_API_KEYopenai/gpt-oss-120bhttps://api.together.ai/v1
Mock AINone (sandbox)mock-modelBuilt-in offline responses

Provider Resolution Hierarchy

When an application feature requests an AI model or completion, resolveConfiguredAiProvider evaluates credentials in strict order of precedence:

1. Explicit caller arguments (customProvider, custom credentials)
   ↓
2. Sandbox enforcement (MOCK_SERVICES=true or AI_PROVIDER='mock')
   ↓
3. Saved admin database settings (Admin UI service integrations)
   ↓
4. Environment variable override (AI_PROVIDER)
   ↓
5. Inferred active provider (first provider with an active API key)
   ↓
6. Fallback to mock AI (safe zero-credential default)

Real-provider failures never silently fall back to mock AI. The mock sandbox is an explicit configuration choice, not an error recovery mechanism.


Application AI Facade

1. Simple Request/Response (chatCompletion)

For standard prompt/response tasks, use chatCompletion. It automatically applies a 30-second default deadline (DEFAULT_CHAT_COMPLETION_TIMEOUT_MS) and composes it with any caller-supplied AbortSignal:

import { chatCompletion } from "ai";

const response = await chatCompletion([
  { role: "system", content: "You are a concise assistant." },
  { role: "user", content: "Summarize the project status in 3 bullet points." },
]);

console.log(response.text);

2. Streaming & Structured Outputs (getLanguageModel)

Advanced features that require streaming, token usage metadata, or structured schemas use getLanguageModel with Vercel AI SDK primitives:

import { getLanguageModel, streamText, generateObject } from "ai";
import { z } from "zod";

// Streaming responses
const model = await getLanguageModel();
const result = streamText({
  model,
  prompt: "Draft an announcement for our upcoming product release.",
});

// Structured JSON generation
const { object } = await generateObject({
  model,
  schema: z.object({
    sentiment: z.enum(["positive", "neutral", "negative"]),
    tags: z.array(z.string()),
    urgency: z.number().min(1).max(5),
  }),
  prompt: "Analyze this customer feedback message: ...",
});

Application Copilot

VibeKit includes an embedded, floating Copilot widget (CopilotChat.tsx) for in-app conversational assistance.

Admin Controls

Platform administrators can configure Copilot behavior directly from the Admin Settings dashboard without code changes:

  • aiCopilot: Global kill switch to enable or disable the Copilot widget across the application.
  • aiAnonymous: Controls whether unauthenticated visitors may query the Copilot.
  • aiAgentName: Custom name displayed in the chat header (default: Product AI).
  • suggestedPrompts: Starter prompts displayed on empty chat sessions.

Production Guardrails & Ceilings

To prevent runaway billing and denial-of-wallet attacks, the server-side Copilot procedure enforces strict admission limits:

Ceiling VariableDefaultPurpose
AI_COPILOT_MAX_INPUT_CHARS8000Maximum allowed prompt character length per request
AI_COPILOT_USER_DAILY_LIMIT200Maximum requests per authenticated user per UTC day
AI_COPILOT_ANON_DAILY_LIMIT30Maximum requests per anonymous IP per UTC day
AI_COPILOT_GLOBAL_DAILY_LIMIT2000Global requests accepted per UTC day across the instance
AI_COPILOT_GLOBAL_PER_MINUTE60Sliding window rate limit across all callers
AI_COPILOT_MAX_CONCURRENT8Maximum active concurrent completions in flight
AI_COPILOT_GLOBAL_DAILY_TOKENS20000000Hard conservative daily token budget cap

Privacy and Data Boundary

The Copilot minimizes server-added context before a provider call. Account email, raw backend exceptions and hidden authorization data are not added to the model prompt. Retrieved product context is bounded and sanitized before use.

A user message is still model input. Do not type passwords, API keys or other secrets into the chat.

Application code, not the model, owns identity and authorization. Team-scoped application data may only be read through server paths that already authorize the requesting user.

Grounding and Source Authority

Copilot does not treat every retrieved result as equally authoritative. The server resolves evidence in this order:

  1. Current canonical runtime state.
  2. Administrator-verified facts.
  3. Administrator-approved knowledge.
  4. Product documentation and published releases.
  5. Public product data such as roadmap or public feedback.

Higher-authority evidence wins when the existing source contract resolves a disagreement. When current evidence does not establish a product fact, Copilot should say what is unknown instead of guessing.

Semantic Actions and Approval

The site agent has a narrow semantic action set. It can currently propose navigation, waitlist signup, feedback submission and supported administrator setting changes.

The model is only the planner. Server code decides which actions are available for the current user, modules and route. Every write is validated and executed through the canonical application API.

  • Navigation is limited to approved same-site routes.
  • Joining the waitlist and submitting feedback use the final labeled button as the user's explicit confirmation. Behind that click, the application still mints a one-time approval capability before executing the write.
  • Administrator module and site-setting changes show exact before and after values before approval.
  • Approval capabilities are bound to the actor, action, normalized arguments, target, expiry and revision when a revision applies.
  • A successful action is shown only after the canonical API returns a receipt. Replaying the same completed approval can return that stored receipt without executing the write twice.

Copilot does not have generic database, HTTP, browser or arbitrary tRPC write access.

Conversations, Traces and Corrections

Copilot conversations, messages, runs and approvals are durable so a session can be restored and an action can be audited. Run traces keep bounded diagnostic metadata such as prompt version, provider/model, evidence references, grounding status, selected tools, action lifecycle, timing, token usage, estimated cost and failure class. Raw provider/backend exceptions and credentials do not belong in traces.

Users can rate an answer. A negative rating can enter the Admin AI Operations review queue. When an administrator creates a correction from a reviewed run, the rejected answer is reference material only. The correction starts empty and must be written and explicitly published before it can become administrator-approved knowledge.


Model Context Protocol (MCP)

VibeKit provides a native Model Context Protocol server at /api/mcp.

This enables external developer agents (such as Claude Desktop or Cursor) to securely access product context, database schemas, and approved procedures using user-scoped OAuth consent grants.

Refer to the MCP Documentation for scope boundaries and client configuration.


Zero-Credential Mock Sandbox

For local development, CI pipelines, and review testing, VibeKit runs completely offline without third-party API keys:

# Enable mock sandbox in your .env or command line
MOCK_SERVICES=true
# Or select mock provider specifically
AI_PROVIDER=mock

In mock mode, chatCompletion returns deterministic fixture text immediately without network requests, token consumption, or external latency.


Agent-First Instructions

For a normal AI feature, keep using the shared provider facade in packages/ai. Do not import vendor SDKs directly into feature code.

When extending the site agent itself, keep the action contract narrow. Add a semantic tool only when a real user outcome requires it. Let server-owned capability state decide whether the tool exists for the current actor. Validate typed arguments, call the canonical application API, keep approval proportional to the write risk and return a real receipt. Add the smallest focused regression and include the behavior in bun run test:ai-evals.

Give your coding agent this prompt:

Extend the VibeKit site agent for this user outcome.
Reuse packages/api/features/ai instead of creating another agent runtime.
If a new action is required, make it a narrow typed semantic tool.
Keep identity, capability filtering, authorization, validation, approval and writes in application code.
Execute through the canonical application API and return a real receipt.
Do not add generic database, HTTP, browser or arbitrary write tools.
Add focused tests and update the existing AI eval corpus.
Verify the completed change with $verify-changes.

On this page