Token导航 LogoToken导航TokenDH.com
研究检索需要联网github未标认证来源可访问许可证需确认审计提醒

github-researchGitHub 研究

Agent Skill

用于围绕 GitHub 仓库、Issue、Pull Request、分支、提交和代码协作流程提供辅助能力。它适合让 Agent 查询项目状态、整理变更、辅助创建或检查协作事项,并把仓库中的信息转成可执行的下一步。使用时需要区分只读查询和写入操作;涉及创建 PR、修改 Issue、推送分支或访问私有仓库时,应确认 token 权限、目标仓库范围和用户授权。

总安装

3,168

周安装

132

GitHub Stars

40

下载量

1,056
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

unknown

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:github-research(GitHub 研究)
来源仓库:https://github.com/lingzhi227/agent-research-skills
仓库路径:skills/github-research
安装命令:
npx skills add https://github.com/lingzhi227/agent-research-skills --skill github-research
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 npx skills 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

skills.shnpx skills
npx skills add https://github.com/lingzhi227/agent-research-skills --skill github-research

简介

专为 GitHub 项目调研设计的智能信息提取工具。

  • 可快速定位技术方案与实现细节。适用宿主包括 Codex、Claude、Cursor、Gemini CLI,接入前应确认版本、权限和运行环境要求。
  • 集成于 Claude、Codex 等 AI 编程助手生态。
  • 建议限制高频请求以避免触发 API 速率限制。
  • github-research 属于研究检索类 Skill,可作为该场景下的辅助能力补充。

SKILL.md

GitHub Research Skill

Trigger

Activate this skill when the user wants to:

  • "Find repos for [topic]", "GitHub research on [topic]"
  • "Analyze open-source code for [topic]"
  • "Find implementations of [paper/technique]"
  • "Which repos implement [algorithm]?"
  • Uses /github-research <deep-research-output-dir> slash command

Overview

This skill systematically discovers, evaluates, and deeply analyzes GitHub repositories related to a research topic. It reads deep-research output (paper database, phase reports, code references) and produces an actionable integration blueprint for reusing open-source code.

Installation: ~/.claude/skills/github-research/ — scripts, references, and this skill definition. Output: ./github-research-output/{slug}/ relative to the current working directory. Input: A deep-research output directory (containing paper_db.jsonl, phase reports, code_repos.md, etc.)

6-Phase Pipeline

Phase 1: Intake     → Extract refs, URLs, keywords from deep-research output
Phase 2: Discovery  → Multi-source broad GitHub search (50-200 repos)
Phase 3: Filtering  → Score & rank → select top 15-30 repos
Phase 4: Deep Dive  → Clone & deeply analyze top 8-15 repos (code reading)
Phase 5: Analysis   → Per-repo reports + cross-repo comparison
Phase 6: Blueprint  → Integration/reuse plan for research topic

Output Directory Structure

github-research-output/{slug}/
├── repo_db.jsonl                     # Master repo database
├── phase1_intake/
│   ├── extracted_refs.jsonl          # URLs, keywords, paper-repo links
│   └── intake_summary.md
├── phase2_discovery/
│   ├── search_results/               # Raw JSONL from each search
│   └── discovery_log.md
├── phase3_filtering/
│   ├── ranked_repos.jsonl            # Scored & ranked subset
│   └── filtering_report.md
├── phase4_deep_dive/
│   ├── repos/                        # Cloned repos (shallow)
│   ├── analyses/                     # Per-repo analysis .md files
│   └── deep_dive_summary.md
├── phase5_analysis/
│   ├── comparison_matrix.md          # Cross-repo comparison
│   ├── technique_map.md              # Paper concept → code mapping
│   └── analysis_report.md
└── phase6_blueprint/
    ├── integration_plan.md           # How to combine repos
    ├── reuse_catalog.md              # Reusable components catalog
    ├── final_report.md               # Complete compiled report
    └── blueprint_summary.md

Scripts Reference

All scripts are Python 3, stdlib-only, located in ~/.claude/skills/github-research/scripts/.

ScriptPurposeKey Flags
extract_research_refs.pyParse deep-research output for GitHub URLs, paper refs, keywords--research-dir, --output
search_github.pySearch GitHub repos via gh api--query, --language, --min-stars, --sort, --max-results, --topic, --output
search_github_code.pySearch GitHub code for implementations--query, --language, --filename, --max-results, --output
search_paperswithcode.pySearch Papers With Code for paper→repo mappings--paper-title, --arxiv-id, --query, --output
repo_db.pyJSONL repo database managementsubcommands: merge, filter, score, search, tag, stats, export, rank
repo_metadata.pyFetch detailed metadata via gh api--repos, --input, --output, --delay
clone_repo.pyShallow-clone repos for analysis--repo, --output-dir, --depth, --branch
analyze_repo_structure.pyMap file tree, key files, LOC stats--repo-dir, --output
extract_dependencies.pyExtract and parse dependency files--repo-dir, --output
find_implementations.pySearch cloned repo for specific code patterns--repo-dir, --patterns, --output
repo_readme_fetch.pyFetch README without cloning--repos, --input, --output, --max-chars
compare_repos.pyGenerate comparison matrix across repos--input, --output
compile_github_report.pyAssemble final report from all phases--topic-dir

Phase 1: Intake

Goal: Extract all relevant references, URLs, and keywords from the deep-research output.

Steps

  1. Create output directory structure: SLUG=$(echo "$TOPIC" | tr '[:upper:]' '[:lower:]' | tr ' ' '-' | tr -cd 'a-z0-9-') mkdir -p github-research-output/$SLUG/{phase1_intake,phase2_discovery/search_results,phase3_filtering,phase4_deep_dive/{repos,analyses},phase5_analysis,phase6_blueprint}
  2. Extract references from deep-research output: python ~/.claude/skills/github-research/scripts/extract_research_refs.py \ --research-dir <deep-research-output-dir> \ --output github-research-output/$SLUG/phase1_intake/extracted_refs.jsonl
  3. Review extracted refs: Read the generated JSONL. Note:

- GitHub URLs found directly in reports - Paper titles and arxiv IDs (for Papers With Code lookup) - Research keywords and themes (for GitHub search queries)

  1. Write intake summary: Create phase1_intake/intake_summary.md with:

- Number of direct GitHub URLs found - Number of papers with potential code links - Key research themes extracted - Planned search queries for Phase 2

Checkpoint

  • extracted_refs.jsonl exists with entries
  • intake_summary.md written
  • Search strategy documented

Phase 2: Discovery

Goal: Cast a wide net to find 50-200 candidate repos from multiple sources.

Steps

  1. Search by direct URLs: Any GitHub URLs from Phase 1 → fetch metadata: python ~/.claude/skills/github-research/scripts/repo_metadata.py \ --repos owner1/name1 owner2/name2... \ --output github-research-output/$SLUG/phase2_discovery/search_results/direct_urls.jsonl
  2. Search Papers With Code: For each paper with an arxiv ID: python ~/.claude/skills/github-research/scripts/search_paperswithcode.py \ --arxiv-id 2401.12345 \ --output github-research-output/$SLUG/phase2_discovery/search_results/pwc_2401.12345.jsonl
  3. Search GitHub by keywords (3-8 queries based on research themes): python ~/.claude/skills/github-research/scripts/search_github.py \ --query "multi-agent LLM coordination" \ --min-stars 10 --sort stars --max-results 50 \ --output github-research-output/$SLUG/phase2_discovery/search_results/gh_query1.jsonl
  4. Search GitHub code (for specific implementations): python ~/.claude/skills/github-research/scripts/search_github_code.py \ --query "class MultiAgentOrchestrator" \ --language python --max-results 30 \ --output github-research-output/$SLUG/phase2_discovery/search_results/code_query1.jsonl
  5. Fetch READMEs for repos that lack descriptions: python ~/.claude/skills/github-research/scripts/repo_readme_fetch.py \ --input <repos.jsonl> \ --output github-research-output/$SLUG/phase2_discovery/search_results/readmes.jsonl
  6. Merge all results into master database: python ~/.claude/skills/github-research/scripts/repo_db.py merge \ --inputs github-research-output/$SLUG/phase2_discovery/search_results/*.jsonl \ --output github-research-output/$SLUG/repo_db.jsonl
  7. Write discovery log: Create phase2_discovery/discovery_log.md with search queries used, results per source, total unique repos found.

Rate Limits

  • GitHub search API: 30 requests/minute (authenticated)
  • Papers With Code API: No strict limit but be respectful (1 req/sec)
  • Add --delay 1.0 to batch operations when needed

Checkpoint

  • repo_db.jsonl populated with 50-200 repos
  • discovery_log.md with search details

Phase 3: Filtering

Goal: Score and rank repos, select top 15-30 for deeper analysis.

Steps

  1. Enrich metadata for all repos: python ~/.claude/skills/github-research/scripts/repo_metadata.py \ --input github-research-output/$SLUG/repo_db.jsonl \ --output github-research-output/$SLUG/repo_db.jsonl \ --delay 0.5
  2. Score repos (quality + activity scores): python ~/.claude/skills/github-research/scripts/repo_db.py score \ --input github-research-output/$SLUG/repo_db.jsonl \ --output github-research-output/$SLUG/repo_db.jsonl
  3. LLM relevance scoring: Read through the top ~50 repos (by quality_score) and assign relevance_score (0.0-1.0) based on: python ~/.claude/skills/github-research/scripts/repo_db.py tag \ --input github-research-output/$SLUG/repo_db.jsonl \ --ids owner/name --tags "relevance:0.85"

- Direct relevance to research topic - Implementation completeness - Code quality signals (from README, description) - Update the relevance scores:

  1. Compute composite scores and rank: python ~/.claude/skills/github-research/scripts/repo_db.py score \ --input github-research-output/$SLUG/repo_db.jsonl \ --output github-research-output/$SLUG/repo_db.jsonl python ~/.claude/skills/github-research/scripts/repo_db.py rank \ --input github-research-output/$SLUG/repo_db.jsonl \ --output github-research-output/$SLUG/phase3_filtering/ranked_repos.jsonl \ --by composite_score
  2. Select top repos: Filter to top 15-30: python ~/.claude/skills/github-research/scripts/repo_db.py filter \ --input github-research-output/$SLUG/phase3_filtering/ranked_repos.jsonl \ --output github-research-output/$SLUG/phase3_filtering/ranked_repos.jsonl \ --max-repos 30 --not-archived
  3. Write filtering report: Create phase3_filtering/filtering_report.md:

- Stats before/after filtering - Score distributions - Top 30 repos with scores and rationale

Scoring Formula

activity_score = sigmoid((days_since_push < 90) * 0.4 + has_recent_commits * 0.3 + open_issues_ratio * 0.3)
quality_score  = normalize(log(stars+1) * 0.3 + log(forks+1) * 0.2 + has_license * 0.15 + has_readme * 0.15 + not_archived * 0.2)
composite_score = relevance * 0.4 + quality * 0.35 + activity * 0.25

Checkpoint

  • ranked_repos.jsonl with 15-30 repos
  • filtering_report.md with scoring details

Phase 4: Deep Dive

Goal: Clone and deeply analyze the top 8-15 repos.

Steps

  1. Select repos for deep dive: Take top 8-15 from ranked list.
  2. Clone each repo (shallow): python ~/.claude/skills/github-research/scripts/clone_repo.py \ --repo owner/name \ --output-dir github-research-output/$SLUG/phase4_deep_dive/repos/
  3. Analyze structure for each cloned repo: python ~/.claude/skills/github-research/scripts/analyze_repo_structure.py \ --repo-dir github-research-output/$SLUG/phase4_deep_dive/repos/name/ \ --output github-research-output/$SLUG/phase4_deep_dive/analyses/name_structure.json
  4. Extract dependencies: python ~/.claude/skills/github-research/scripts/extract_dependencies.py \ --repo-dir github-research-output/$SLUG/phase4_deep_dive/repos/name/ \ --output github-research-output/$SLUG/phase4_deep_dive/analyses/name_deps.json
  5. Find implementations: Search for key algorithms/concepts from research: python ~/.claude/skills/github-research/scripts/find_implementations.py \ --repo-dir github-research-output/$SLUG/phase4_deep_dive/repos/name/ \ --patterns "class Transformer" "def forward" "attention" \ --output github-research-output/$SLUG/phase4_deep_dive/analyses/name_impls.jsonl
  6. Deep code reading: For each repo, READ the key source files identified by structure analysis. Write a per-repo analysis in phase4_deep_dive/analyses/{name}_analysis.md:

- Architecture overview - Key algorithms implemented - Code quality assessment - API / interface design - Dependencies and requirements - Strengths and limitations - Reusability assessment (how easy to extract components)

  1. Write deep dive summary: phase4_deep_dive/deep_dive_summary.md

IMPORTANT: Actually Read Code

Do NOT just summarize READMEs. You must:

  • Read the main source files (entry points, core modules)
  • Understand the actual implementation approach
  • Identify specific functions/classes that implement research concepts
  • Note code patterns, design decisions, and trade-offs

Checkpoint

  • Repos cloned in repos/
  • Per-repo analysis files in analyses/
  • deep_dive_summary.md written

Phase 5: Analysis

Goal: Cross-repo comparison and technique-to-code mapping.

Steps

  1. Generate comparison matrix: python ~/.claude/skills/github-research/scripts/compare_repos.py \ --input github-research-output/$SLUG/phase4_deep_dive/analyses/ \ --output github-research-output/$SLUG/phase5_analysis/comparison.json
  2. Write comparison matrix: Create phase5_analysis/comparison_matrix.md:

- Table comparing repos across dimensions (language, LOC, stars, framework, license, tests) - Dependency overlap analysis - Strengths/weaknesses per repo

  1. Write technique map: Create phase5_analysis/technique_map.md:

- Map each paper concept / research technique → specific repo + file + function - Identify gaps (techniques with no implementation found) - Note alternative implementations of the same concept

  1. Write analysis report: phase5_analysis/analysis_report.md:

- Executive summary of findings - Key insights from code analysis - Recommendations for which repos to use for which purposes

Checkpoint

  • comparison_matrix.md with repo comparison table
  • technique_map.md mapping concepts to code
  • analysis_report.md with findings

Phase 6: Blueprint

Goal: Produce an actionable integration and reuse plan.

Steps

  1. Write integration plan: phase6_blueprint/integration_plan.md:

- Recommended architecture for combining repos - Step-by-step integration approach - Dependency resolution strategy - Potential conflicts and how to resolve them

  1. Write reuse catalog: phase6_blueprint/reuse_catalog.md:

- For each reusable component: source repo, file path, function/class, what it does, how to extract it - License compatibility matrix - Effort estimates (easy/medium/hard to integrate)

  1. Compile final report: python ~/.claude/skills/github-research/scripts/compile_github_report.py \ --topic-dir github-research-output/$SLUG/
  2. Write blueprint summary: phase6_blueprint/blueprint_summary.md:

- One-page executive summary - Top 5 repos and why - Recommended next steps

Checkpoint

  • integration_plan.md complete
  • reuse_catalog.md with component catalog
  • final_report.md compiled
  • blueprint_summary.md as executive summary

Quality Conventions

  1. Repos are ranked by composite score: relevance × 0.4 + quality × 0.35 + activity × 0.25
  2. Deep dive requires reading actual code, not just READMEs
  3. Integration blueprint must map paper concepts → specific code files/functions
  4. Incremental saves: Each phase writes to disk immediately
  5. Checkpoint recovery: Can resume from any phase by checking what outputs exist
  6. All scripts are stdlib-only Python — no pip installs needed
  7. gh CLI is required for GitHub API access (must be authenticated)
  8. Deduplication by repo_id (owner/name) across all searches
  9. Rate limit awareness: Respect GitHub search API limits (30 req/min)

Error Handling

  • If gh is not installed: warn user and provide installation instructions
  • If a repo is archived/deleted: skip gracefully, note in log
  • If clone fails: skip, note in log, continue with remaining repos
  • If Papers With Code API is down: skip, rely on GitHub search only
  • Always write partial progress to disk so work is not lost

References

  • See references/phase-guide.md for detailed phase execution guidance
  • Deep-research skill: ~/.claude/skills/deep-research/SKILL.md
  • Paper database pattern: ~/.claude/skills/deep-research/scripts/paper_db.py

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Codex

35.29%
按下载量换算373

Claude

29.32%
按下载量换算310

Cursor

16.95%
按下载量换算179

Gemini CLI

9.24%
按下载量换算98

安全审计

Gen Agent Trust Hub

通过

Socket

通过

Snyk

可疑

权限和风险

需要联网

该 Skill 可能需要联网访问来源站点、仓库或外部 API;具体网络访问范围需要结合源码和 README 复核。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。来源安全扫描存在 warning/failed 结果,不能写成本站确认安全。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills