# Scholar Evaluation > Ofrece revisión formativa de trabajos académicos, cualitativa y trazable en evidencia, y audita rúbricas de evaluación de investigación de bajo riesgo. Nunca la uses para rankear personas ni decisiones consecuentes. Fuente: https://skillsagentes.com/skills/k-dense-ai/scientific-agent-skills/scholar-evaluation Markdown: https://skillsagentes.com/skills/k-dense-ai/scientific-agent-skills/scholar-evaluation.md Repositorio: https://github.com/K-Dense-AI/scientific-agent-skills Autor: K-Dense-AI Licencia: MIT Actualizado: el mes pasado Coste de contexto: 57 tok instalada, 2.7k tok al activarse, 39.2k tok con todos los archivos del bundle Bundle: 19 archivos, 153 KB Permisos que pide: read write bash glob python ## Instalación Un skill son archivos markdown: los mismos archivos valen para cualquier agente y lo único que cambia es el directorio de destino, es decir la bandera `--agent`. Añade `-g` para instalarlo en todos los proyectos de la máquina. ```bash # Claude Code npx -y skills add K-Dense-AI/scientific-agent-skills --skill scholar-evaluation --agent claude-code # Cursor npx -y skills add K-Dense-AI/scientific-agent-skills --skill scholar-evaluation --agent cursor # Codex npx -y skills add K-Dense-AI/scientific-agent-skills --skill scholar-evaluation --agent codex # Gemini CLI npx -y skills add K-Dense-AI/scientific-agent-skills --skill scholar-evaluation --agent gemini # Windsurf npx -y skills add K-Dense-AI/scientific-agent-skills --skill scholar-evaluation --agent windsurf # Cline npx -y skills add K-Dense-AI/scientific-agent-skills --skill scholar-evaluation --agent cline ``` ## Qué hace - Da retroalimentación formativa y trazable en evidencia sobre un trabajo académico (paper, borrador, protocolo, idea de investigación) - Audita si un proceso de evaluación de investigación de bajo riesgo documenta constructo, procedencia, calidad de evaluadores e incertidumbre - Valida rúbricas acotadas y calcula puntuaciones locales sin recomendar una decisión - Comprueba trazabilidad de evidencia, acuerdo entre evaluadores y sensibilidad de pesos con scripts locales - Bloquea explícitamente su uso para contratación, admisión, becas, premios o sanciones ## Cuándo usarla - Se necesita revisión de desarrollo de un paper, borrador, protocolo o idea de investigación - Se audita si un proceso de evaluación de investigación de bajo riesgo sigue buenas prácticas de rigor y equidad - Se valida o calcula una rúbrica de evaluación acotada con evidencia trazable ## Cuándo no - Se pide rankear personas o influir en contratación, promoción, admisión, becas, premios o sanciones - Se pide un juicio de listo para publicar, aceptar/rechazar o top-tier ## Qué la activa - "Dame retroalimentación de desarrollo sobre este borrador de paper" - "Audita si esta rúbrica de evaluación de investigación documenta bien la incertidumbre" - "Valida esta rúbrica acotada y calcula las puntuaciones locales" - "Revisa la trazabilidad de evidencia de esta evaluación de investigación" ## Antes de instalar - Requiere Python 3.11+ para las CLIs incluidas; todo el procesamiento es local, sin red, credenciales ni modelos externos. - makes network requests ## Archivos - SKILL.md — 10 KB - assets/evaluation_template.json — 1 KB - assets/evidence_manifest_template.json — 2 KB - assets/process_checklist_template.json — 2 KB - assets/ratings_template.csv — 2 KB - assets/rubric_template.json — 12 KB - references/evaluation_framework.md — 10 KB - references/local_tooling.md — 8 KB - references/responsible_assessment.md — 9 KB - references/security_validation.md — 4 KB - references/source_ledger.md — 12 KB - scripts/_common.py — 35 KB - scripts/calculate_scores.py — 2 KB - scripts/check_process.py — 8 KB - scripts/check_traceability.py — 9 KB - scripts/generate_report_scaffold.py — 9 KB - scripts/summarize_agreement.py — 9 KB - scripts/validate_rubric.py — 2 KB - scripts/weight_sensitivity.py — 9 KB ## SKILL.md Reproducido tal cual desde K-Dense-AI/scientific-agent-skills bajo MIT. Esta sección es el documento original y está en inglés. # Scholar Evaluation ## Purpose Provide developmental, evidence-traceable feedback on a **scholarly work**: paper, draft, protocol, literature synthesis, or research idea. Use qualitative judgment first. Optional scores only describe how submitted evidence maps to a predeclared bounded rubric. This skill also audits whether a low-stakes assessment process documents its construct, provenance, rater quality, uncertainty, traceability, sensitivity, fairness, accessibility, privacy, and human governance. ## Hard safety boundary Never use this skill to automate, recommend, materially influence, or score: - hiring, promotion, or tenure; - admissions; - grants or other funding; - prizes, honors, or awards; - discipline, dismissal, or sanctions; or - any other high-impact personnel decision. Never rank people. Never reduce a person to a composite score. Never infer ability, character, integrity, protected traits, future performance, or worth. A nominal human-in-the-loop does not remove this boundary. If asked for a prohibited use, stop. Offer developmental comments on a scholarly work or a process-only audit that does not process applications, compare people, recommend an outcome, or advise a decision. Do not issue publication-readiness, accept/reject, or “top-tier” judgments. Read `references/responsible_assessment.md` before any organizational use. ## ScholarEval status The referenced ScholarEval project is an **experimental literature-grounded research-idea evaluation framework**, not validated psychometrics. The verified primary record is Moussa et al., *ScholarEval: Research Idea Evaluation Grounded in Literature*, arXiv:2510.16234v2, revised 2026-02-28. It reports a retrieval-augmented soundness/contribution framework, a 117-idea four-discipline dataset, coverage experiments, and a user study. Do not generalize those results to person assessment, consequential decisions, all disciplines, or this skill's rubric. No peer-reviewed publication status was verified during the dated review. See `references/source_ledger.md`. ## Metric and prestige policy Do not score or infer quality from: - Journal Impact Factor or other journal measures; - h-index, publication counts, or citation counts; - altmetrics or attention; - journal, conference, venue, institution, employer, or geographic prestige; - author affiliation, reputation, network, or career path. The rubric validator rejects common proxy-measure criteria. If a qualified reviewer mentions an indicator descriptively outside the scoring tools, record its exact purpose, source, coverage, field and time effects, uncertainty, missingness, biases, gaming risk, and why it does not directly measure quality. Never hide indicators inside an opaque composite. ## Data boundary Bundled scripts accept only strict local JSON/CSV containing pseudonymous IDs, bounded ratings, statuses, uncertainty, and local references. Do not put raw private applications, CVs, letters, reviewer identities, contact details, protected attributes, or source-document text in inputs, outputs, logs, examples, or prompts. Keep source content in the authorized records system and use opaque local references. Allowed classifications are: - `synthetic` - `public_scholarly_work` - `deidentified_low_stakes` No script searches the web, loads environment files, reads credentials, calls a model, executes supplied text, deserializes executable objects, or launches a process. Use Bash only to invoke the documented local `python3` commands. ## Workflow ### 1. Confirm allowed use and authorization Record: - developmental purpose; - unit of assessment: `scholarly_work`; - work type, stage, discipline, language, and audience; - authorized source location and data classification; - accountable committee owner; - conflicts and recusals; - accessibility and accommodation process; - appeal or correction route; and - data purpose, access, retention, and deletion. Stop on a prohibited decision context or unnecessary private data. ### 2. Define the construct before criteria State: - what quality or support is being examined; - excluded constructs; - intended interpretation; - contexts where the interpretation does not travel; - evidence requirements; and - known limitations. Start with values and disciplinary context, not available metrics. ### 3. Adapt and validate the rubric Begin with `assets/rubric_template.json`, then obtain qualified disciplinary, assessment-methods, stakeholder, accessibility, privacy, and fairness review. The template deliberately records content validity as `not_established`. Do not change that status without documented evidence for the exact intended use. Validate structure: ```bash PYTHONDONTWRITEBYTECODE=1 python3 scripts/validate_rubric.py \ --rubric assets/rubric_template.json ``` Read `references/evaluation_framework.md` for construct, anchor, validity, and rater guidance. ### 4. Build traceable evidence records Reviewers may read an authorized work outside the scripts. Record only stable local locators and claim references in `assets/evidence_manifest_template.json`. For every criterion, distinguish: - observed evidence from interpretation; - supporting from contrary evidence; - available from unavailable evidence; - `missing` from `not_applicable`; and - uncertainty from absence. Failure to find prior work does not prove novelty. ### 5. Rate independently Use `assets/evaluation_template.json`. Each criterion must be: - `rated` with an anchor score, bounded uncertainty, evidence IDs, and a local rationale reference; - `missing` with null score/uncertainty and a rationale reference; or - `not_applicable` with null score/uncertainty and a rationale reference. Do not encode missing or not-applicable as zero. Raters should train, calibrate, disclose conflicts, rate independently, and document disagreement. ### 6. Run local quality checks Bounded scoring, without labels or recommendation: ```bash PYTHONDONTWRITEBYTECODE=1 python3 scripts/calculate_scores.py \ --rubric assets/rubric_template.json \ --evaluation assets/evaluation_template.json ``` Evidence traceability: ```bash PYTHONDONTWRITEBYTECODE=1 python3 scripts/check_traceability.py \ --rubric assets/rubric_template.json \ --evaluation assets/evaluation_template.json \ --evidence assets/evidence_manifest_template.json ``` Inter-rater agreement: ```bash PYTHONDONTWRITEBYTECODE=1 python3 scripts/summarize_agreement.py \ --rubric assets/rubric_template.json \ --ratings assets/ratings_template.csv ``` Weight sensitivity requires two or more distinct scholarly-work evaluation files: ```bash PYTHONDONTWRITEBYTECODE=1 python3 scripts/weight_sensitivity.py \ --rubric assets/rubric_template.json \ --evaluation /tmp/work-a-evaluation.json \ --evaluation /tmp/work-b-evaluation.json ``` Process controls: ```bash PYTHONDONTWRITEBYTECODE=1 python3 scripts/check_process.py \ --process assets/process_checklist_template.json ``` The checklist template is intentionally unconfirmed and fails closed. Instructions and exact schemas are in `references/local_tooling.md`. ### 7. Synthesize qualitative findings Lead with criterion-level evidence, not the composite. For each criterion: 1. cite evidence references; 2. state `rated`, `missing`, or `not_applicable`; 3. explain the anchor interpretation; 4. report score and uncertainty only if rated; 5. note disagreements and context; 6. identify strengths and limitations; and 7. offer non-prescriptive improvement options. Generate an empty-reference scaffold if useful: ```bash PYTHONDONTWRITEBYTECODE=1 python3 scripts/generate_report_scaffold.py \ --rubric assets/rubric_template.json \ --evaluation assets/evaluation_template.json \ --output /tmp/developmental-report-scaffold.json ``` The scaffold does not read source documents or draft findings. ### 8. Human review and release Before releasing an organizational report, a qualified accountable human committee must verify: - construct and rubric provenance; - content-validity evidence and limits; - rater training, agreement, inter-rater reliability evidence, and drift; - evidence traceability and source access; - missingness, not-applicable rationales, and uncertainty; - weight sensitivity and order instability; - disciplinary and subgroup bias review; - conflicts and recusals; - accessibility and accommodations; - privacy, minimization, retention, and output controls; and - correction or appeal information. Document dissent. Do not imply consensus, validity, or precision beyond the evidence. Periodically evaluate the evaluation and retire harmful criteria. ## Interpretation rules - A score is an ordinal rubric summary, not a natural measurement. - Normalization does not repair incomplete evidence. - The bundled uncertainty range is not a confidence interval. - Agreement does not establish reliability, validity, fairness, or correctness. - Stable results under tested weights do not establish validity. - The overall score never overrides criterion evidence or qualified judgment. - No output is a decision recommendation. ## Bundled resources - `references/responsible_assessment.md` — safety, metrics, governance, accessibility, privacy, and bias. - `references/evaluation_framework.md` — ScholarEval boundary, construct, criteria, anchors, validity, and interpretation. - `references/local_tooling.md` — strict schemas, formulas, commands, and output behavior. - `references/source_ledger.md` — authoritative sources and publication-status verification dated 2026-07-23. - `references/security_validation.md` — baseline remediation, validation, and residual security-scan record. - `assets/rubric_template.json` — bounded rubric template. - `assets/evaluation_template.json` — rating template. - `assets/evidence_manifest_template.json` — traceability template. - `assets/process_checklist_template.json` — fail-closed process checklist. - `assets/ratings_template.csv` — synthetic agreement data. ## Dónde encaja - Categoría: [Investigación](https://skillsagentes.com/categorias/investigacion.md) — Investigación estructurada, búsqueda de fuentes y síntesis. - Creador: [K-Dense-AI](https://skillsagentes.com/creators/k-dense-ai.md) — 163 skills en el directorio - [Todas las skills](https://skillsagentes.com/skills.md) - [Ranking de instalaciones](https://skillsagentes.com/ranking.md) ## Otras skills del mismo repositorio - [Citation Management](https://skillsagentes.com/skills/k-dense-ai/scientific-agent-skills/citation-management.md): Gestión integral de citas académicas: busca en OpenAlex, PubMed y Google Scholar, extrae metadatos precisos, valida citas y genera entradas BibTeX correctamente formateadas. - [Scientific Slides](https://skillsagentes.com/skills/k-dense-ai/scientific-agent-skills/scientific-slides.md): Crea decks de diapositivas y presentaciones para charlas de investigación: PowerPoint, presentaciones de conferencia, seminarios, defensas de tesis. Da estructura, plantillas, guía de tiempos y validación visual. - [Literature Review](https://skillsagentes.com/skills/k-dense-ai/scientific-agent-skills/literature-review.md): Realiza revisiones bibliográficas sistemáticas y completas usando varias bases académicas (PubMed, arXiv, bioRxiv, Semantic Scholar). Genera markdown y PDF con citas verificadas en varios estilos (APA, Nature, Vancouver). - [Infographics](https://skillsagentes.com/skills/k-dense-ai/scientific-agent-skills/infographics.md): Crea infografías profesionales con Nano Banana Pro AI y refinamiento iterativo inteligente. Usa Gemini 3.6 Flash para revisar la calidad e integra investigación con Perplexity Sonar. Soporta 10 tipos, 8 estilos y paletas para daltonismo. - [Latex Posters](https://skillsagentes.com/skills/k-dense-ai/scientific-agent-skills/latex-posters.md): Crea pósteres de investigación profesionales en LaTeX con beamerposter, tikzposter o baposter, para conferencias y comunicación científica: layout, colores, columnas múltiples e integración de figuras. --- Skills Agentes · [Índice de páginas en markdown](https://skillsagentes.com/sitemap.md) · [Inicio](https://skillsagentes.com/index.md)