Build with Your Agent
Supervised QA
Watch a fresh agent use VibeKit and turn its friction into product improvements.
A supervised QA run uses a second agent as a realistic VibeKit user.
The goal is not to help that agent finish at any cost. The goal is to see whether a capable fresh agent can understand VibeKit, use the normal workflow and reach a correct result without unnecessary setup, guessing or hand-holding.
This is useful after major workflow, architecture, documentation or Agent Skill changes.
Run the evaluation
- Give the fresh agent a realistic product brief and a clean starting point.
- Let it discover AGENTS.md and use vibekit-deploy as the normal full product entry point.
- Watch how it handles checkout state, local-environment-setup, product decisions, reusable building blocks, feature implementation and verify-changes.
- Do not answer routine questions too quickly. A fresh agent should discover repository facts from VibeKit itself.
- Let it fix its own normal mistakes. Step in only when it is truly blocked or when continuing would make the run invalid.
- Do not call something a VibeKit defect because the agent failed. Reproduce or verify the issue against the original product or repository contract first.
- Keep a short QA log. Record only findings that can improve the product, workflow, docs, skills or verification.
- After the run, use test-and-qa-pipeline for independent proof where needed. Use first-principles-review to decide what should be removed, simplified or automated before adding more instructions.
- If you fix a confirmed issue, use verify-changes and keep the regression proof with the owning code.
What to watch for
- Real product defects
- Builder mistakes
- Unclear ownership or instructions
- Slow setup or decision loops
- Places where the agent guesses instead of discovering
- Places where the user is asked a question VibeKit could answer itself
- Missing or weak verification
- Repeated mistakes that should become a test, guard, skill or clearer scoped instruction
Supervisor prompt
You are supervising another AI agent as it uses VibeKit from start to finish.
Let the second agent act like a capable real user seeing the product for the first time. Do not do the work for it.
Give it the normal VibeKit path. It should discover AGENTS.md, use vibekit-deploy for a full product build, establish local setup with local-environment-setup when needed, use focused skills for implementation and finish with verify-changes.
Watch for:
- Real VibeKit defects
- Mistakes caused by the second agent
- Unclear or slow workflows
- Places where the agent needs too much setup, guessing or hand-holding
- Questions sent to the user that the repository should have answered
- Missing proof or false completion claims
Do not call something a VibeKit defect just because the second agent fails. Verify the issue against the original product, current source, tests or documented contract first.
Let the second agent find and fix its own normal mistakes. Only step in if it is truly blocked or the evaluation would become invalid.
Keep a short QA log with only meaningful findings. Separate confirmed VibeKit defects, workflow friction and second-agent mistakes.
When a finding needs broader proof, use test-and-qa-pipeline. Before adding more process or documentation, use first-principles-review to decide whether the better fix is to delete, simplify, accelerate or automate the problem away.
At the end, report the confirmed defects, workflow problems, second-agent mistakes, what worked and the highest-impact fixes.Optional tooling note
Pi is one optional harness for the second agent. Its RPC mode exposes JSON responses and event frames over stdin and stdout, which can make direct observation easier for a supervisor. The QA method does not depend on Pi.