Corrige el conocimiento desactualizado del LLM sobre la plataforma Vercel e introduce sus productos nuevos. Se inyecta al inicio de la sesión.
- Costo de contexto al activarse
- 1.8k tok
- Tamaño del paquete
- 1 archivo
- Última actualización
- el mes pasado
Suite de benchmark end-to-end para vercel-plugin: ejecuta proyectos reales con inyección de skills, verifica servidores dev, analiza logs y genera un reporte de mejora.
en todo el repo
0–100, la ruta de este skill
último commit aquí
últimos 90 días
61 tok en reposo
5 KB
Funciona con cualquier agente que lea SKILL.md
npx -y skills add vercel/vercel-plugin --skill benchmark-e2e --agent claude-codeSe instala solo en este repositorio.
Di cualquiera de estas frases y el agente debería cargar este skill.
Single-command pipeline that creates projects, exercises skill injection via claude --print, launches dev servers, verifies they work, analyzes conversation logs, and generates actionable improvement reports.
# Full suite (9 projects, ~2-3 hours)
bun run scripts/benchmark-e2e.ts
# Quick mode (first 3 projects, ~30-45 min)
bun run scripts/benchmark-e2e.ts --quick
Options:
| Flag | Description | Default |
|---|---|---|
--quick |
Run only first 3 projects | false |
--base <path> |
Override base directory | ~/dev/vercel-plugin-testing |
--timeout <ms> |
Per-project timeout (forwarded to runner) | 900000 (15 min) |
The orchestrator chains four stages sequentially, aborting on failure:
claude --print with VERCEL_PLUGIN_LOG_LEVEL=tracerun-manifest.json, extracts metricsreport.md and report.json with scorecards and recommendationsrun-manifest.jsonWritten by the runner at <base>/results/run-manifest.json. Links all downstream stages to the same run.
interface BenchmarkRunManifest {
runId: string; // UUID for this pipeline run
timestamp: string; // ISO 8601
baseDir: string; // Absolute path to base directory
projects: Array<{
slug: string; // e.g. "01-recipe-platform"
cwd: string; // Absolute path to project dir
promptHash: string; // SHA hash of the prompt text
expectedSkills: string[];
}>;
}
The analyzer and verifier read this manifest to correlate sessions precisely instead of guessing from directory listings.
events.jsonlThe orchestrator writes NDJSON events to <base>/results/events.jsonl tracking pipeline lifecycle:
// Each line is one JSON object:
{ "stage": "pipeline", "event": "start", "timestamp": "...", "data": { "baseDir": "...", "quick": false } }
{ "stage": "runner", "event": "start", "timestamp": "...", "data": { "script": "...", "args": [...] } }
{ "stage": "runner", "event": "complete", "timestamp": "...", "data": { "exitCode": 0, "durationMs": 120000 } }
// On failure:
{ "stage": "verify", "event": "error", "timestamp": "...", "data": { "exitCode": 1, "durationMs": 5000, "slug": "04-conference-tickets" } }
{ "stage": "pipeline", "event": "abort", "timestamp": "...", "data": { "failedStage": "verify", "exitCode": 1, "slug": "04-conference-tickets" } }
report.jsonMachine-readable report at <base>/results/report.json for programmatic consumption:
interface ReportJson {
runId: string | null;
timestamp: string;
verdict: "pass" | "partial" | "fail";
gaps: Array<{
slug: string;
expected: string[];
actual: string[];
missing: string[];
}>;
recommendations: string[];
suggestedPatterns: Array<{
skill: string; // Skill that was expected but not injected
glob: string; // Suggested pathPattern glob
tool: string; // Tool name that should trigger injection
}>;
}
Run the pipeline repeatedly with a cooldown between iterations:
while true; do
bun run scripts/benchmark-e2e.ts
sleep 3600
done
Each run produces timestamped report.json and report.md files. Compare across runs to track improvement.
The pipeline enables a closed feedback loop:
bun run scripts/benchmark-e2e.ts exercises the plugin against realistic projectsreport.json lists which skills were expected but never injected, with exact slugssuggestedPatterns entries (copy-pasteable YAML) to add missing frontmatter patterns; use recommendations to fix hook logicreport.json across runs: verdict should trend from "fail" → "partial" → "pass"For overnight automation, combine with the loop above. Wake up to reports showing exactly what improved and what still needs work.
Prompts never name specific technologies — they describe the product and features, letting the plugin infer which skills to inject.
| # | Slug | Expected Skills |
|---|---|---|
| 01 | recipe-platform | auth, vercel-storage, nextjs |
| 02 | trivia-game | vercel-storage, nextjs |
| 03 | code-review-bot | ai-sdk, nextjs |
| 04 | conference-tickets | payments, email, auth |
| 05 | content-aggregator | cron-jobs, ai-sdk |
| 06 | finance-tracker | cron-jobs, email |
| 07 | multi-tenant-blog | routing-middleware, cms, auth |
| 08 | status-page | cron-jobs, vercel-storage, observability |
| 09 | dog-walking-saas | payments, auth, vercel-storage, env-vars |
rm -rf ~/dev/vercel-plugin-testing
Reproducido de vercel/vercel-plugin bajo licencia NOASSERTION. Leer esta página en markdown.
1 archivo en el paquete. Solo se lee SKILL.md al activarse — las referencias se cargan si el skill decide que las necesita.
Requiere bun, el CLI `claude`, y un directorio base como ~/dev/vercel-plugin-testing (configurable con --base).
Necesita en el PATH:bun
Este repo incluye 43 skills. Si instalas uno, normalmente ya tienes los demás. Ver el pack vercel-plugin entero y su comando de instalación
Corrige el conocimiento desactualizado del LLM sobre la plataforma Vercel e introduce sus productos nuevos. Se inyecta al inicio de la sesión.
Guía experta de Vercel Connect: obtén tokens OAuth con permisos limitados para servicios de terceros (Slack, GitHub, servidores MCP, OAuth, Snowflake) en nombre de apps o usuarios vía Vercel OIDC.
Depura el caching de la CDN de Vercel: tasa de aciertos, contenido obsoleto, revalidación, ISR + PPR, cacheReason, ppr_state y costos.
Guía experta sobre Vercel Functions: Serverless Functions, Edge Functions, Fluid Compute, streaming, Cron Jobs y configuración de runtime para código server-side en Vercel.
Guía de arquitectura backend: úsala para planear, construir o migrar una API o backend, elegir entre Functions, Services, contenedores, Workflow, Queues y bases de datos de Marketplace, o seleccionar framework y runtime.
Guía de eve, el framework para agentes de IA duraderos: runtime basado en filesystem, sesiones, tools, skills, canales, sandboxes, subagentes, schedules, evals y observabilidad con Agent Runs.
Escenarios avanzados de benchmark para agentes de IA que ponen a prueba Workflow SDK, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox y orquestación multi-agente de Vercel.
Corre escenarios de eval de vercel-plugin en Vercel Sandboxes en vez de paneles locales de WezTerm: aprovisiona microVMs con Claude Code, ejecuta prompts y genera reportes de cobertura.
Crea y lanza proyectos de prueba de benchmark para ejercitar la inyección de skills de vercel-plugin en escenarios realistas, usando directorios aislados y paneles de WezTerm con Claude Code.