Skills Agentes

Test Flakiness

Detecta tests no deterministas (flaky) leyendo logs de CI o histórico de resultados; agrega tasas de aprobación, recomienda cuarentena o fix y mantiene un registro de tests flaky.

Solicitareadglobgrepwriteeditbash
Estrellas
24.4k

en todo el repo

Actividad
43

0–100, la ruta de este skill

Actualizado
hace 3 meses

último commit aquí

Commits
0

últimos 90 días

Contexto
2k tok

69 tok en reposo

Paquete
1 archivo

8 KB

Instalar

Funciona con cualquier agente que lea SKILL.md

npx -y skills add Donchitos/Claude-Code-Game-Studios --skill test-flakiness --agent claude-code

Se instala solo en este repositorio.

Este skill runs shell commands, writes to your files.

Qué hace

  • Parsea logs de CI o históricos de resultados de tests para detectar tests inestables (flaky)
  • Calcula tasas de fallo por test y clasifica la causa probable (timing, orden, seed, fuga de recursos, etc.)
  • Recomienda cuarentena, corrección o monitoreo según el nivel de flakiness
  • Actualiza la sección de cuarentena en tests/regression-suite.md
  • Genera un informe opcional en production/qa/flakiness-report-[date].md

Úsalo cuando

  • En la fase de Polish, cuando ya hay muchas ejecuciones de CI con señal estadística fiable
  • Cuando los desarrolladores empiezan a descartar fallos de CI como 'probablemente flaky'
  • Después de que /regression-suite identifique tests en cuarentena que necesitan diagnóstico

No lo uses cuando

    Qué lo activa

    Di cualquiera de estas frases y el agente debería cargar este skill.

    • Analiza este log de CI para detectar tests flaky
    • /test-flakiness scan
    • /test-flakiness registry
    • Revisa los tests en cuarentena de regression-suite.md

    SKILL.md

    En inglés

    Test Flakiness Detection

    A flaky test is one that sometimes passes and sometimes fails without any code change. Flaky tests are worse than no tests in some ways — they train the team to ignore red CI runs, masking genuine failures. This skill identifies them, explains likely causes, and recommends whether to quarantine or fix each one.

    Output: Updated tests/regression-suite.md quarantine section + optional production/qa/flakiness-report-[date].md

    When to run:

    • Polish phase (tests have had many runs; statistical signal is reliable)
    • When developers start dismissing CI failures as "probably flaky"
    • After /regression-suite identifies quarantined tests that need diagnosis

    1. Parse Arguments

    Modes:

    • /test-flakiness [ci-log-path] — analyse a specific CI run log file
    • /test-flakiness scan — scan all available CI logs in .github/ or standard log output directories
    • /test-flakiness registry — read existing regression-suite.md quarantine section and provide remediation guidance for already-known flaky tests
    • No argument — auto-detect: run scan if CI logs are accessible, else registry

    2. Locate CI Log Data

    Option A — GitHub Actions (preferred)

    Check for test result artifacts:

    ls -t .github/ 2>/dev/null
    ls -t test-results/ 2>/dev/null
    

    For Godot projects: GdUnit4 outputs XML results compatible with JUnit format. Check test-results/ for .xml files.

    For Unity projects: game-ci test runner outputs NUnit XML to test-results/ by default.

    For Unreal projects: automation logs go to Saved/Logs/. Grep for Result: Success and Result: Fail patterns.

    Option B — Local log files

    If a path argument is provided, read that file directly.

    Option C — No log data available

    If no logs found:

    "No CI log data found. To detect flaky tests, this skill needs test result history from multiple runs. Options:

    1. Run the test suite at least 3 times and collect the output logs
    2. Check CI pipeline output and save a log to test-results/
    3. Run /test-flakiness registry to review tests already flagged as flaky in tests/regression-suite.md"

    Stop and ask the user which option to pursue.


    3. Parse Test Results

    For each CI log or result file found, parse:

    JUnit XML format (GdUnit4 / Unity):

    • Grep for <testcase name= to get test names
    • Grep for <failure or <error to identify failures
    • Parse classname and name attributes for full test identifiers

    Plain text logs:

    • Grep for pass/fail patterns:
      • Godot: PASSED / FAILED adjacent to test names
      • Unreal: Result: Success / Result: Fail
      • Unity: Test passed / Test failed

    Build a table: test_id → [run1_result, run2_result, run3_result, ...]


    4. Identify Flaky Tests

    A test is flaky if it appears in the result history with both PASS and FAIL outcomes across runs with no code changes between them.

    Flakiness thresholds:

    • High flakiness: Fails in >25% of runs — quarantine immediately
    • Moderate flakiness: Fails in 5–25% of runs — investigate and fix soon
    • Low/suspected flakiness: Fails in 1–5% of runs — monitor; may be genuinely rare failure

    For each flaky test, classify the likely cause:

    Cause classification

    Cause Symptoms Fix direction
    Timing / async Fails after awaiting signals or timers; pass rate correlates with system load Add explicit await/synchronisation; avoid time-based delays
    Order dependency Fails when run after specific other tests; passes in isolation Add proper setup/teardown; ensure test isolation
    Random seed Fails intermittently with no pattern; involves RNG Pass explicit seed; don't use randf() in tests
    Resource leak Fails more often later in a test run Fix cleanup in teardown; check orphan nodes (Godot) or object disposal (Unity)
    External state Fails when a file, scene, or global exists from a prior test Isolate test from file system; use in-memory mocks
    Floating point Fails on comparisons like == 0.5 Use epsilon comparison (is_equal_approx, Assert.AreApproximately)
    Scene/prefab load race Fails when scenes are not yet ready Await one frame after instantiation; use await get_tree().process_frame

    Use Grep to check the test file for timing calls, randf, global state access, or equality comparisons on floats to narrow down the cause.


    5. Recommend Action

    For each flaky test:

    Quarantine (High flakiness):

    "Quarantine this test immediately. Disable it in CI by adding @pytest.mark.skip / [Ignore] / GdUnitSkip annotation. Log it in tests/regression-suite.md quarantine section. The test is now opt-in only. Fix the root cause before removing quarantine."

    Investigate and fix soon (Moderate):

    "This test is intermittently unreliable. Root cause appears to be [cause]. Suggested fix: [specific fix based on cause classification]. Do not quarantine yet — fix the test directly."

    Monitor (Low/suspected):

    "This test shows suspected flakiness. Collect more run data before quarantining. Note it as 'suspected' in the regression suite."


    6. Generate Reports

    In-conversation summary

    ## Flakiness Detection Results
    
    **Runs analysed**: [N]
    **Tests tracked**: [N]
    
    ### Flaky Tests Found
    
    | Test | System | Fail Rate | Likely Cause | Recommendation |
    |------|--------|-----------|--------------|----------------|
    | [test_name] | [system] | [N]% | Timing | Quarantine + fix async |
    | [test_name] | [system] | [N]% | Float comparison | Fix: use epsilon compare |
    | [test_name] | [system] | [N]% | Order dependency | Investigate teardown |
    
    ### Clean Tests (no flakiness detected)
    
    [N] tests ran across [N] runs with consistent results — no flakiness detected.
    
    ### Data Limitations
    
    [Note if fewer than 5 runs were available — fewer runs = less statistical confidence]
    

    7. Update Regression Suite + Optional Report File

    Ask: "May I update the quarantine section of tests/regression-suite.md with the flaky tests found?"

    If yes: use Edit to append entries to the Quarantined Tests table. Never remove existing quarantine entries — only add new ones.

    Ask (separately): "May I write a full flakiness report to production/qa/flakiness-report-[date].md?"

    The full report includes per-test analysis with cause details and engine-specific fix snippets.

    After writing:

    • For each quarantined test: "Add the engine-specific skip annotation to disable this test in CI. Re-enable after the root cause is fixed."
    • For fix-eligible tests: "The fix for [test] is straightforward — change the equality comparison on line [N] to use is_equal_approx."
    • Summary: "Once all quarantine annotations are applied, CI should run green. Schedule fix work for the [N] quarantined tests before the release gate."

    Collaborative Protocol

    • Never delete test files — quarantine means annotate + list, not remove
    • Statistical confidence matters — with < 3 runs, flag findings as "suspected" not "confirmed"; ask if more run data is available
    • Fix is always the goal — quarantine is temporary; surface the fix direction even when recommending quarantine
    • Ask before writing — both the regression-suite update and the report file require explicit approval. On write: Verdict: COMPLETE — flakiness report written. On decline: Verdict: BLOCKED — user declined write.
    • Flakiness in CI is a team problem — surface the list and recommended actions clearly; do not just silently quarantine without the team knowing

    Reproducido de Donchitos/Claude-Code-Game-Studios bajo licencia MIT. Leer esta página en markdown.

    Archivos

    1 archivo en el paquete. Solo se lee SKILL.md al activarse — las referencias se cargan si el skill decide que las necesita.

    Antes de instalar

    Necesita historial de resultados de al menos varias ejecuciones de test (logs de CI en .github/, test-results/ o Saved/Logs/).

    Detalles

    Creador
    Donchitos
    Categoría
    Testing y QA
    Licencia
    MIT
    Recursos incluidos
    Solo SKILL.md
    Código fuente
    Ver SKILL.md

    Etiquetas

    Más de Donchitos/Claude-Code-Game-Studios

    Este repo incluye 73 skills. Si instalas uno, normalmente ya tienes los demás.

    Adopt

    24.4k

    Onboarding brownfield: audita el cumplimiento de formato de los artefactos existentes, clasifica los vacíos por impacto y genera un plan de migración numerado.

    Costo de contexto al activarse
    4.5k tok
    Tamaño del paquete
    1 archivo
    Última actualización
    hace 3 meses
    herramientas desarrollo

    Crea un Registro de Decisión de Arquitectura (ADR) que documenta una decisión técnica importante, su contexto, alternativas consideradas y consecuencias.

    Costo de contexto al activarse
    4.8k tok
    Tamaño del paquete
    1 archivo
    Última actualización
    hace 3 meses
    documentos

    Valida que la arquitectura del proyecto cubra por completo los GDD: cruza requisitos con ADR, detecta conflictos entre decisiones y compatibilidad de motor, y da un veredicto PASS/CONCERNS/FAIL.

    Costo de contexto al activarse
    6.7k tok
    Tamaño del paquete
    1 archivo
    Última actualización
    hace 3 meses
    herramientas desarrollo

    Autoría guiada, sección por sección, del Art Bible. Crea la especificación de identidad visual que condiciona toda la producción de assets. Se ejecuta tras aprobar /brainstorm y antes de /map-systems o de redactar cualquier GDD.

    Costo de contexto al activarse
    3.7k tok
    Tamaño del paquete
    1 archivo
    Última actualización
    hace 3 meses
    documentos

    Audita los assets del juego según convenciones de nombres, presupuestos de tamaño, formatos estándar y requisitos de pipeline. Identifica assets huérfanos, referencias faltantes e infracciones de estándares.

    Costo de contexto al activarse
    697 tok
    Tamaño del paquete
    1 archivo
    Última actualización
    hace 3 meses
    testing qa

    Genera especificaciones visuales por asset y prompts de generación IA a partir de GDDs, docs de nivel o perfiles de personaje. Produce archivos de spec y actualiza el manifiesto maestro.

    Costo de contexto al activarse
    4.1k tok
    Tamaño del paquete
    1 archivo
    Última actualización
    hace 3 meses
    documentos

    Skills relacionados

    Audita los assets del juego según convenciones de nombres, presupuestos de tamaño, formatos estándar y requisitos de pipeline. Identifica assets huérfanos, referencias faltantes e infracciones de estándares.

    Costo de contexto al activarse
    697 tok
    Tamaño del paquete
    1 archivo
    Última actualización
    hace 3 meses
    testing qa

    Crea un informe de bug estructurado a partir de una descripción o analiza código para identificar bugs potenciales, con pasos de reproducción, severidad y contexto completos.

    Costo de contexto al activarse
    1.5k tok
    Tamaño del paquete
    1 archivo
    Última actualización
    hace 3 meses
    testing qa

    Lee los bugs abiertos en production/qa/bugs/, reevalúa prioridad frente a severidad, los asigna a sprints, detecta tendencias sistémicas y genera un informe de triage.

    Costo de contexto al activarse
    2k tok
    Tamaño del paquete
    1 archivo
    Última actualización
    hace 3 meses
    testing qa