Token导航 LogoToken导航TokenDH.com
研究检索操作浏览器clawhub未标认证来源可访问clear审计提醒

llm-benchmark-analystLLM benchmark analyst 搜索

Agent Skill

llm-benchmark-analyst 用于查找、检索和筛选相关信息,适合在 OpenClaw 中需要根据关键词、任务场景或来源线索快速定位候选结果时使用。可结合来源仓库、安装命令和原始 README 继续核验具体用法。安装前建议确认权限范围、维护状态,以及是否会触发联网、命令执行或文件读写。

总安装

9,376

周安装

383

GitHub Stars

公开资料未说明

下载量

3,003
OpenClaw

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

MIT-0

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:llm-benchmark-analyst(LLM benchmark analyst 搜索)
来源仓库:https://github.com/chekhovin/llm-benchmark-analyst
安装命令:
openclaw skills install llm-benchmark-analyst
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 OpenClaw 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

ClawHubOpenClaw
openclaw skills install llm-benchmark-analyst

简介

系统化分析 LLM 基准测试结果并生成竞争力报告的工具。

  • 覆盖主流评测数据集,识别模型优势领域与潜在短板。
  • 输出基于证据的领域领导者建议,辅助技术选型决策。
  • 依赖公开基准数据库,对私有或定制化测试支持有限。
  • 报告仅供参考,实际应用效果可能受任务特性影响。llm-benchmark-analyst 属于研究检索类 Skill,可作为该场景下的辅助能力补充。

SKILL.md

name
llm-benchmark-analyst
description
search and analyze llm benchmark results within a fixed benchmark universe, then produce evidence-based model strength and weakness reports or domain-leader summaries. use when comparing a model across benchmarks, ranking the best models by domain, explaining what a benchmark measures, checking predecessor-vs-current progress, or writing benchmark reports that must prioritize exact model version, evaluation date, benchmark variant, score semantics, sub-scores, and benchmark defect warnings. works with browser, web, and multimodal extraction for text, table, canvas, or image-only leaderboards.

LLM Benchmark Analyst

Overview

Use this skill to research benchmark evidence and write structured reports about:

  1. a single model's strengths and weaknesses
  2. best models in a capability domain
  3. what a benchmark measures and how trustworthy it is
  4. predecessor vs current-model progress

Default to the user's language. Never invent scores, ranks, dates, benchmark variants, or missing table values.

Core constraints

  • Restrict the benchmark universe to references/benchmark-source.md. If a benchmark is not in that file, exclude it.
  • Use references/core-dimensions.md to collapse scattered benchmarks into a small set of report dimensions.
  • Follow references/search-playbook.md for routing, overlap expansion, evidence gathering, and comparison anchors.
  • Follow references/report-template.md for output structure.
  • Apply references/data-defect-warnings.md benchmark by benchmark, inline and again in the limitations section.
  • Prefer official benchmark or benchmark-author pages. Use aggregators mainly to discover links and context.
  • Record the evaluation mode exactly: benchmark version, split, difficulty, public/private, verified/original, with-tools/without-tools, pass@k, and any visible sub-score names.
  • Keep score units exact. Do not average incompatible metrics into a fake composite.

Required workflow

  1. Normalize the model identity before searching

- Resolve exact provider, family, generation, version suffix, and release label. - Put time and version first. Reject ambiguous aliases like claude, gemini pro, gpt latest, or qwen max until you have the exact currently relevant model string for the searched leaderboard rows. - Capture the evaluation time point or access date for every key score.

  1. Route the request through core dimensions before web crawling

- Start with references/core-dimensions.md to select the primary dimension(s). - Then list candidate benchmarks inside those dimensions. - Only then start website-by-website retrieval. - Keep the first pass narrow and token-efficient: start from the best 3-6 benchmarks for the asked domain, then expand only if needed.

  1. Expand beyond section labels

- Do not let the source document's headings blind you. - After selecting the primary dimension, inspect benchmark descriptions and overlap tags to find relevant benchmarks that live in other sections. - Example: a coding analysis may need coding benchmarks, agentic coding benchmarks, general benchmarks with coding components, and research/math benchmarks with strong code components. - Example: a multimodal analysis may need vision benchmarks, OCR, GUI/computer-use, multimodal deep-research, and omni/video/audio benchmarks.

  1. Collect evidence in this order

- official leaderboard or benchmark site - benchmark paper or benchmark README - benchmark-author blog or release note - trusted aggregator - vendor blog only as secondary evidence, clearly labeled as vendor-reported if no independent leaderboard row exists

  1. Use multimodal extraction when the leaderboard is not machine-readable

- If the page uses images, canvas, screenshots, or chart-only rendering and plain text extraction misses the table, inspect screenshots or page images. - Extract only values that are clearly visible. - Mark the provenance as image-extracted. - If the image is unreadable or partially occluded, say so instead of guessing.

  1. Apply anchor comparisons

- For code or agentic coding, compare against the latest available Claude Opus, latest Claude Sonnet, and latest GPT family model. - For multimodal analysis, compare against the latest available Gemini model. Add the latest GPT multimodal model if relevant. - For intelligence or reasoning analysis, compare against the latest available GPT family model. - Never assume which model is currently latest. Search that first.

  1. Apply predecessor comparison

- If data exists, compare the target model with its immediate predecessor or last broadly comparable prior generation from the same provider/family. - Only compare like-for-like benchmark variants. If the predecessor only appears under a different benchmark mode, say the comparison is not clean.

  1. Attach defect warnings

- Any benchmark with a known quality or methodology issue must carry an inline warning from references/data-defect-warnings.md. - If the report's conclusion depends heavily on warned benchmarks, lower confidence and say so explicitly.

Decision rules

  • When the user asks for best models in a domain, do not use only one benchmark. Use a cluster of relevant benchmarks and explain why each one matters.
  • When the user asks for what is this model good or bad at, synthesize at the core-dimension level first, then support with benchmark evidence.
  • When benchmark scores conflict, prefer freshness, exact version match, official source quality, and the number of agreeing benchmarks over one standout score.
  • Treat very small gaps as non-decisive when the benchmark is noisy, image-extracted, or known to be unstable.
  • Always include one short clause describing what each benchmark actually tests.

Minimum evidence to capture

For every benchmark you cite, capture:

  • benchmark name
  • what it tests in one short phrase
  • exact model row name
  • exact score and unit
  • rank or relative placement if visible
  • benchmark variant, split, or mode
  • date or access time point
  • source quality note if not official
  • data warning if applicable

Output expectations

Use the matching template in references/report-template.md.

At minimum, every substantive report must include:

  • a scope and identity section
  • a short executive summary
  • strengths
  • weaknesses or gaps
  • evidence table
  • comparison section
  • data-defect warnings and confidence
  • methodology or exclusions

Resource map

  • references/core-dimensions.md: benchmark routing and de-fragmentation map
  • references/search-playbook.md: token-efficient search order, overlap expansion, and comparison rules
  • references/data-defect-warnings.md: warning catalog and ready-to-use caution language
  • references/report-template.md: output structures for single-model, domain-leader, and benchmark-explainer tasks
  • references/benchmark-source.md: full allowed benchmark universe copied from the user's benchmark document

Example tasks

  • analyze gpt-5's coding and agentic coding strengths and weaknesses, and compare it with the latest claude opus, claude sonnet, and gpt model
  • find the best multimodal models right now using only the approved benchmark list and explain each benchmark briefly
  • write a report on qwen's reasoning strengths, benchmark gaps, predecessor comparison, and all data-quality caveats
  • tell me which models lead in deep research and search, with benchmark-specific warnings and freshness notes

适合场景

01

OpenClaw 用户查找和安装 Skill 时

02

用户想查找某类 Agent Skill 时

03

需要根据任务场景推荐可安装能力包时

04

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

补充不同宿主或平台的使用分布数据

能力 5

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

OpenClaw

93.56%
按下载量换算2,810

安全审计

VirusTotal

通过

ClawScan

可疑

Static analysis

通过

权限和风险

操作浏览器

该 Skill 可能涉及浏览器控制能力,使用时可能读取或操作网页内容,需要在受控环境中确认权限边界。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。来源安全扫描存在 warning/failed 结果,不能写成本站确认安全。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills