Token导航 LogoToken导航TokenDH.com
研究检索只读github未标认证来源可访问许可证需确认审计通过

skill-evaluator技能评估员

Agent Skill

skill-evaluator 用于查找、检索和筛选相关信息,适合在 Codex、Claude、Cursor、Gemini CLI 中需要根据关键词、任务场景或来源线索快速定位候选结果时使用。可结合来源仓库、安装命令和原始 README 继续核验具体用法。安装前建议确认权限范围、维护状态,以及是否会触发联网、命令执行或文件读写。

总安装

517

周安装

22

GitHub Stars

217

下载量

181
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

unknown

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:skill-evaluator(技能评估员)
来源仓库:https://github.com/mathews-tom/praxis-skills
仓库路径:skills/skill-evaluator
安装命令:
npx skills add https://github.com/mathews-tom/praxis-skills --skill skill-evaluator
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 npx skills 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

skills.shnpx skills
npx skills add https://github.com/mathews-tom/praxis-skills --skill skill-evaluator

简介

skill-evaluator 用于查找、检索和筛选相关信息,适合在 Codex、Claude、Cursor、Gemini CLI 中需要根据关键词快速定位候选结果时使用。

  • 适用于技能评估和性能分析等研究检索场景。
  • 通过 npx skills add 命令从 GitHub 仓库安装,需确认权限范围和文件读写操作。
  • 建议结合原始 README 核验具体用法,注意维护状态和网络访问限制。
  • 适用宿主包括 Codex、Claude、Cursor、Gemini CLI,接入前应确认版本、权限和运行环境要求。

SKILL.md

Skill Evaluator

Skills that do not activate on relevant queries waste the entire investment in writing them. A skill can have deep, well-structured content and still deliver zero value if its frontmatter description lacks the trigger phrases users actually type. Quality evaluation catches trigger gaps, missing sections, and shallow content before deployment — turning a skill from a static document into a reliable tool.

Reference Files

FileContents
references/evaluation-rubric.mdDetailed 1-5 scoring criteria per dimension, weight justifications, worked examples for calibration

Audit Modes

Two modes, selected by input:

  • Quick Audit: Evaluate a single skill. Produces a full per-dimension scored report with findings, severity classifications, and recommendations.
  • Full Audit: Evaluate all skills in the repository. Produces a comparative ranking table sorted by overall score, plus condensed per-skill summaries.

Mode Selection

InputMode
Path to a specific skill directory or SKILL.mdQuick Audit
"all", "every skill", no path specified, or repo-level requestFull Audit
Multiple specific pathsQuick Audit for each, then comparative summary

Evaluation Dimensions

Six dimensions, each scored 1-5. Weighted sum determines overall percentage.

D1: Frontmatter Quality (20%)

Evaluates the YAML frontmatter block for completeness and discoverability.

Signals:

  • name field present and non-empty
  • description field present and non-empty
  • Description length between 200-800 characters (sweet spot for keyword density without bloat)
  • Description contains explicit trigger phrases users would type
  • Description includes a "Use this skill when..." clause or equivalent
  • Description is keyword-dense, not generic filler

Scoring constraints: A description under 100 characters caps this dimension at 2/5. A missing name or description field caps at 1/5.

D2: Trigger Coverage (18%)

Evaluates whether the skill activates on the queries users actually type.

Signals:

  • Synonym breadth — multiple phrasings for the same intent (e.g., "review", "audit", "critique", "evaluate", "assess", "check")
  • Implied contexts — situations where the skill applies even without explicit keywords (e.g., "user provides a design doc and asks for feedback")
  • Domain-specific terms relevant to the skill's function
  • Explicit trigger phrase list in the description frontmatter
  • Coverage of both imperative ("review this") and interrogative ("is this good?") forms

Scoring constraints: Fewer than 3 distinct trigger phrases caps at 2/5. Zero trigger phrases in the description caps at 1/5.

D3: Structural Completeness (20%)

Evaluates whether the skill contains the sections needed to function reliably.

Signals:

  • Prerequisites or setup instructions (if applicable)
  • Multi-phase workflow or step-by-step procedure
  • Error handling guidance or edge case documentation
  • Output format specification (template, example, or schema)
  • Limitations or scope boundaries stated
  • Reference file table (if references/ directory exists)
  • Calibration rules or quality gates

Scoring constraints: A skill with no workflow section caps at 2/5. A skill with a workflow but no error handling or output format caps at 3/5.

D4: Content Depth (22%)

Evaluates the substantive quality of the skill's guidance — whether it provides enough detail for an agent to execute well without human intervention.

Signals:

  • Multi-step workflows with decision points, not bare command lists
  • Error cases documented with recovery actions
  • Decision frameworks (when to do X vs Y, mode selection tables)
  • Verbatim output examples or templates
  • Severity classifications or scoring rubrics (where applicable)
  • Cross-cutting analysis or synthesis steps beyond simple checklists

Scoring constraints: A skill consisting only of bare commands with no explanatory context caps at 2/5. Reference files count toward this dimension only if they contain substantive guidance (checklists, rubrics, criteria), not just link collections.

D5: Consistency and Integrity (12%)

Evaluates internal consistency and structural integrity.

Signals:

  • Directory name matches the name field in frontmatter exactly
  • All files referenced in SKILL.md exist on disk (reference files, scripts, assets)
  • Description content aligns with body content (description does not promise features the body does not deliver)
  • Consistent terminology throughout (same concept uses same term)
  • No broken internal links or dangling references
  • Self-containment: No cross-skill references (../other-skill/) in SKILL.md or reference files. Skills must be standalone packages — all referenced files must live within the skill's own directory. Shared content should use the _templates/ sync system to maintain local copies.

Scoring constraints: A name mismatch between directory and frontmatter is a CRITICAL finding and caps at 1/5. Cross-skill ../ references are a CRITICAL finding and cap at 1/5 — they break standalone packaging. Missing referenced files cap at 2/5.

D6: CONTRIBUTING.md Compliance (8%)

Evaluates adherence to the repository's contribution guidelines.

Signals:

  • Skill name is kebab-case
  • Skill name is 64 characters or fewer
  • Description is 1024 characters or fewer
  • No angle brackets in description
  • No pushy trigger language in description ("always use", "you must", "never do")
  • Valid YAML frontmatter syntax

Scoring constraints: Any single violation caps at 3/5. Multiple violations cap at 2/5. Invalid YAML that prevents parsing caps at 1/5.


Severity Classification

SeverityCriteriaScore Impact
CRITICALSkill cannot activate or breaks on load — missing frontmatter, name mismatch, invalid YAMLCaps overall score at 40%
HIGHSignificant trigger gap or missing core section — no workflow, no error handling, zero trigger phrasesCaps affected dimension at 3/5
MEDIUMWeak coverage, shallow content, few trigger synonymsDimension needs improvement but functions
LOWMinor polish — formatting inconsistencies, slightly short description, missing calibration rulesFix when convenient

Workflow

Phase 1: Input

  1. Determine audit mode from user input (see Mode Selection table above).
  2. For Quick Audit: validate the skill directory exists and contains a SKILL.md file. If the path points to a SKILL.md directly, use its parent directory.
  3. For Full Audit: enumerate all directories under skills/ that contain a SKILL.md.
  4. For each skill to evaluate, note the directory name for D5 consistency checks.

Phase 2: Analysis

For each skill under evaluation:

  1. Read SKILL.md in full.
  2. Parse YAML frontmatter — extract name and description fields. If YAML parsing fails, record a CRITICAL finding and score D1 and D6 as 1/5.
  3. Check the references/ directory for existence and contents. Verify every file referenced in the SKILL.md body exists on disk.
  4. Scan SKILL.md and all reference files for cross-skill path references (../ patterns pointing outside the skill directory). Flag any as CRITICAL D5 findings.
  5. Evaluate each of the 6 dimensions using the criteria above and the detailed rubric in references/evaluation-rubric.md.
  6. Record findings with severity, dimension tag, description, and recommendation.

Phase 3: Scoring

  1. Score each dimension 1-5 using references/evaluation-rubric.md.
  2. Apply severity caps: if any CRITICAL finding exists, cap overall at 40% regardless of dimension scores.
  3. Compute weighted score: Overall% = (sum of dimension_score x weight) / 5 x 100.
  4. Determine verdict from the scale below.
RangeVerdict
90-100%Exemplary
80-89%Strong
70-79%Adequate
60-69%Needs Work
Below 60%Deficient

Phase 4: Report

Generate the structured output using the appropriate template below.


Output Format

Quick Audit Template

## Skill Audit: {skill-name}

| Dimension | Score | Weight | Weighted | Key Finding |
|-----------|-------|--------|----------|-------------|
| D1: Frontmatter Quality | X/5 | 20% | X.XXX | ... |
| D2: Trigger Coverage | X/5 | 18% | X.XXX | ... |
| D3: Structural Completeness | X/5 | 20% | X.XXX | ... |
| D4: Content Depth | X/5 | 22% | X.XXX | ... |
| D5: Consistency & Integrity | X/5 | 12% | X.XXX | ... |
| D6: CONTRIBUTING Compliance | X/5 | 8% | X.XXX | ... |

**Overall: XX% — {Verdict}**

### Findings

[Severity-sorted list. Each entry includes dimension tag, severity, description,
evidence, and recommendation.]

- **[CRITICAL] D5:** ...
- **[HIGH] D2:** ...
- **[MEDIUM] D4:** ...
- **[LOW] D3:** ...

### Score Calculation

D1: {score} x 0.20 = {result}
D2: {score} x 0.18 = {result}
D3: {score} x 0.20 = {result}
D4: {score} x 0.22 = {result}
D5: {score} x 0.12 = {result}
D6: {score} x 0.08 = {result}
Sum = {weighted_sum}
Overall = {weighted_sum} / 5 x 100 = {percentage}% — {Verdict}

Full Audit Template

## Skill Repository Audit

| Skill | Overall | Verdict | Worst Dimension | Top Issue |
|-------|---------|---------|-----------------|-----------|
| {name} | XX% | {verdict} | {dimension} | {issue} |
| ... | ... | ... | ... | ... |

### Per-Skill Summaries

[Condensed Quick Audit for each skill: scorecard table, overall score, top 3 findings.
Omit the full Score Calculation section in condensed mode.]

Error Handling

ProblemCauseFix
SKILL.md not found in directoryPath incorrect or file missingReport as CRITICAL; do not attempt evaluation; surface the path and stop
YAML frontmatter parse failureInvalid YAML syntax (unclosed quotes, bad indentation)Report as CRITICAL finding; score D1 and D6 as 1/5; continue evaluating the body content where parseable
references/ directory missingSkill has no reference filesNot an error — score D5 normally; check only that any files referenced in the SKILL.md body actually exist on disk
references/ exists but referenced file is absentFile path in SKILL.md body doesn't resolveRecord as a CRITICAL D5 finding; missing referenced files cap D5 at 2/5
Empty SKILL.md (zero bytes or whitespace only)File created but never populatedTreat as CRITICAL; score all dimensions 1/5; overall verdict: Deficient
references/evaluation-rubric.md not foundSkill's own reference file missingNote the irony; evaluate using the criteria inline in this SKILL.md; flag D5 as a CRITICAL finding

Calibration Rules

  1. Score what exists, not what could exist — evaluate the skill as-is, not its potential.
  2. Weight trigger coverage heavily for skills targeting broad domains (e.g., a GitHub skill covers issues, PRs, CI, releases, and API — it needs proportionally more trigger synonyms).
  3. A skill with strong triggers but shallow content scores higher than deep content with poor triggers — activation is prerequisite to utility.
  4. Reference files count toward Content Depth only if they contain substantive guidance (checklists, rubrics, criteria), not link lists or stub files.
  5. When evaluating the skill-evaluator itself, apply identical standards — no self-inflation.
  6. Frontmatter description quality is the single highest-leverage improvement for any skill.

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Codex

36.52%
按下载量换算66

Claude

31.26%
按下载量换算57

Cursor

19.24%
按下载量换算35

Gemini CLI

8.14%
按下载量换算15

安全审计

Gen Agent Trust Hub

通过

Socket

通过

Snyk

通过

权限和风险

只读

该 Skill 主要提供规则、说明或参考内容,本身偏只读;真正读写文件、联网或执行命令仍取决于宿主 Agent 的任务。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills