Token导航 LogoToken导航TokenDH.com
研究检索只读github未标认证来源可访问clear审计通过

dedupe-rank重复数据删除等级

Agent Skill

dedupe-rank 用于查找、检索和筛选相关信息,适合在 Codex、Claude、Cursor、Gemini CLI 中需要根据关键词、任务场景或来源线索快速定位候选结果时使用。可结合来源仓库、安装命令和原始 README 继续核验具体用法。安装前建议确认权限范围、维护状态,以及是否会触发联网、命令执行或文件读写。

总安装

776

周安装

32

GitHub Stars

422

下载量

253
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

3

许可证

MIT

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:dedupe-rank(重复数据删除等级)
来源仓库:https://github.com/willoscar/research-units-pipeline-skills
仓库路径:skills/dedupe-rank
安装命令:
npx skills add https://github.com/willoscar/research-units-pipeline-skills --skill dedupe-rank
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。不同来源提供的安装方式可能略有差异;本站展示可直接复制的安装命令,安装前请核对来源页面。

skills.shnpx skills
npx skills add https://github.com/willoscar/research-units-pipeline-skills --skill dedupe-rank

简介

dedupe-rank 用于查找、检索和筛选相关信息,适合在 Codex、Claude、Cursor、Gemini CLI 中根据关键词快速定位候选结果。

  • 适用于需要去重和排序检索集的场景,尤其适合 LLM 代理主题分类和 outline 构建。
  • 通过 npx skills add 命令从 research-units-pipeline-skills 仓库安装,支持 CSV 生成和稳定 paper_id。
  • 使用时需遵循 domain packs 加载顺序,并仅用 run.py 处理标题归一化和评分逻辑。
  • 建议结合原始 README 核验具体流程,确保输出可重复且无缓存依赖。

SKILL.md

Dedupe + Rank

Turn a broad retrieved set into a smaller core set for taxonomy/outline building.

This is a deterministic “curation” step: it should be stable and repeatable.

Load Order

Always read:

  • references/domain_pack_overview.md — how domain packs drive topic-specific behavior

Domain packs (loaded by topic match):

  • assets/domain_packs/llm_agents.json — pinned classics, survey detection, ranking signals for LLM agent topics

Script Boundary

Use scripts/run.py only for:

  • title normalization and deduplication logic
  • relevance scoring from query tokens
  • core set CSV generation with stable paper_id values

Do not treat run.py as the place for:

  • hardcoded pinned paper IDs (use domain packs)
  • hardcoded survey detection rules (use domain packs)
  • domain-specific topic detection logic (use domain packs)

Input

  • papers/papers_raw.jsonl

Outputs

  • papers/papers_dedup.jsonl
  • papers/core_set.csv

Workflow (high level)

  1. Dedupe by normalized (title, year) and keep the richest metadata per duplicate cluster.
  2. Rank by relevance/recency signals (and optionally pin known classics for certain topics). For LLM-agent topics, also ensure a small quota of prior surveys/reviews is present to support a paper-like Related Work section.
  3. Write papers/core_set.csv with stable paper_id values and useful metadata columns (arxiv_id, pdf_url, categories).

Quality checklist

  • papers/papers_dedup.jsonl exists and is valid JSONL.
  • papers/core_set.csv exists and has a header row.

Script

Quick Start

  • python.codex/skills/dedupe-rank/scripts/run.py --help
  • python.codex/skills/dedupe-rank/scripts/run.py --workspace <workspace_dir> --core-size 300

All Options

  • --core-size <n>: target size for papers/core_set.csv
  • queries.md also supports core_size / core_set_size / dedupe_core_size (overrides default when present)

Examples

  • Smaller core set for fast iteration (non-A150++):

- python.codex/skills/dedupe-rank/scripts/run.py --workspace <ws> --core-size 25

Notes

  • This step may annotate papers/core_set.csv:reason with tags such as pinned_classic and prior_survey (deterministic, topic-aware guards for survey writing).
  • Evidence-review default: if the active pipeline is evidence-review (or legacy systematic-review) and core_size is not specified, the script keeps the full deduped pool in papers/core_set.csv so screening does not silently drop candidates.
  • This step is deterministic; reruns should be stable for the same inputs.

Troubleshooting

Common Issues

Issue: papers/core_set.csv is too small / empty

Symptom:

  • Core set has very few rows.

Causes:

  • Input papers/papers_raw.jsonl is small, or many rows are missing required fields.

Solutions:

  • Broaden retrieval (or provide a richer offline export) and rerun.
  • Lower --core-size only if you intentionally want a small core set.

Issue: Duplicates still appear after dedupe

Symptom:

  • Near-identical titles remain.

Causes:

  • Title normalization is defeated by noisy exports.

Solutions:

  • Clean title fields in the export (strip prefixes/suffixes, fix encoding) and rerun.

Recovery Checklist

  • papers/papers_raw.jsonl lines contain title/year/url.
  • papers/core_set.csv has stable paper_id values.

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

04

需要参考平台分布和安装热度时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

补充不同宿主或平台的使用分布数据

能力 5

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Claude Code

26.58%
按下载量换算67

Gemini CLI

21.34%
按下载量换算54

Cursor

20.11%
按下载量换算51

Codex

13.2%
按下载量换算33

OpenCode

8.22%
按下载量换算21

Antigravity

3.18%
按下载量换算8

安全审计

Gen Agent Trust Hub

通过

Socket

通过

Snyk

通过

权限和风险

只读

该 Skill 主要提供规则、说明或参考内容,本身偏只读;真正读写文件、联网或执行命令仍取决于宿主 Agent 的任务。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。

来源信息

继续浏览同类 Skills