Token导航 LogoToken导航TokenDH.com
研究检索需要联网github未标认证来源可访问许可证需确认审计通过

fault-diagnosis故障诊断

Agent Skill

fault-diagnosis 用于查找、检索和筛选相关信息,适合在 Codex、Claude、Cursor、Gemini CLI 中需要根据关键词、任务场景或来源线索快速定位候选结果时使用。可结合来源仓库、安装命令和原始 README 继续核验具体用法。安装前建议确认权限范围、维护状态,以及是否会触发联网、命令执行或文件读写。

总安装

312

周安装

13

GitHub Stars

24

下载量

104
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

unknown

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:fault-diagnosis(故障诊断)
来源仓库:https://github.com/noobygains/godmode
仓库路径:skills/fault-diagnosis
安装命令:
npx skills add https://github.com/noobygains/godmode --skill fault-diagnosis
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 npx skills 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

skills.shnpx skills
npx skills add https://github.com/noobygains/godmode --skill fault-diagnosis

简介

fault-diagnosis 提供严格的故障诊断流程,禁止在未查明根因前实施任何修复。

  • 适用于在 Codex、Claude、Cursor、Gemini CLI 中对测试失败、性能下降等异常进行处理。
  • 划分五个分析阶段:现象观察、假设提出、实验验证、结论确认与预防措施制定。
  • 特别强调生产环境问题必须优先保障稳定性,避免激进修复引入新缺陷。
  • 所有诊断过程需留下可追溯的记录,便于团队协作与知识沉淀。

SKILL.md

Fault Diagnosis

Overview

Guessing at fixes wastes time and introduces new defects. Quick patches mask underlying problems.

Core principle: ALWAYS identify root cause before attempting any fix. Treating symptoms is failure.

No exceptions. No workarounds. No shortcuts.

The Prime Directive

NO FIXES WITHOUT ROOT CAUSE INVESTIGATION FIRST

If you have not completed Phase 1, you are not authorized to propose fixes.

When to Use

Apply to ANY technical issue:

  • Test failures
  • Production bugs
  • Unexpected behavior
  • Performance degradation
  • Build failures
  • Integration breakdowns

Especially important when:

  • Under time pressure (urgency makes guessing tempting)
  • "Just one quick fix" seems obvious
  • You have already attempted multiple fixes
  • A previous fix did not resolve the issue
  • You do not fully understand the problem

Do not skip when:

  • The issue appears simple (simple bugs have root causes too)
  • You are in a hurry (systematic investigation is faster than flailing)
  • Someone wants it resolved NOW (methodical work is faster than thrashing)

The Four Phases

You MUST complete each phase before advancing to the next.

Phase 1: Root Cause Investigation

BEFORE attempting ANY fix:

  1. Read Error Messages Thoroughly

- Do not skip past errors or warnings - They frequently contain the exact answer - Read stack traces completely - Note line numbers, file paths, error codes

  1. Reproduce Reliably

- Can you trigger it consistently? - What are the exact reproduction steps? - Does it happen every time? - If not reproducible, gather more data -- do not guess

  1. Examine Recent Changes

- What changed that could cause this? - Git diff, recent commits - New dependencies, configuration changes - Environmental differences

  1. Gather Evidence in Multi-Component Systems WHEN the system has multiple components (CI -> build -> signing, API -> service -> database): BEFORE proposing fixes, add diagnostic instrumentation: For EACH component boundary: - Log what data enters the component - Log what data exits the component - Verify environment/config propagation - Check state at each layer Run once to collect evidence showing WHERE it breaks THEN analyze evidence to identify the failing component THEN investigate that specific component Example (multi-layer system): # Layer 1: Orchestrator echo "=== Orchestrator state: ===" echo "TOKEN: ${TOKEN:+SET}${TOKEN:-UNSET}" # Layer 2: Build script echo "=== Build environment: ===" env | grep TOKEN || echo "TOKEN not in environment" # Layer 3: Signing module echo "=== Certificate state: ===" security list-keychains security find-identity -v # Layer 4: Actual operation codesign --sign "$IDENTITY" --verbose=4 "$ARTIFACT" This reveals: Which layer fails (secrets -> orchestrator OK, orchestrator -> build FAIL)
  2. Trace Data Flow WHEN the error is deep in the call stack: See root-cause-tracing.md in this directory for the complete backward tracing method. Short version:

- Where does the bad value originate? - What called this function with the bad value? - Keep tracing upward until you find the source - Fix at the source, not at the symptom

Phase 2: Pattern Analysis

Find the pattern before fixing:

  1. Locate Working Examples

- Find similar working code in the same codebase - What works that resembles what is broken?

  1. Compare Against References

- If implementing a pattern, read the reference implementation COMPLETELY - Do not skim -- read every line - Understand the pattern fully before applying

  1. Identify Differences

- What differs between working and broken? - List every difference, no matter how small - Do not assume "that cannot matter"

  1. Understand Dependencies

- What other components does this require? - What settings, configuration, environment? - What assumptions does it make?

Phase 3: Hypothesis and Testing

Scientific method:

  1. Form a Single Hypothesis

- State clearly: "I believe X is the root cause because Y" - Write it down - Be specific, not vague

  1. Test Minimally

- Make the SMALLEST possible change to test the hypothesis - One variable at a time - Do not fix multiple things simultaneously

  1. Verify Before Continuing

- Did it work? Yes -> Phase 4 - Did not work? Form a NEW hypothesis - DO NOT pile additional fixes on top

  1. When You Do Not Know

- Say "I do not understand X" - Do not pretend to know - Ask for help - Research further

Phase 4: Implementation

Fix the root cause, not the symptom:

  1. Create a Failing Test Case

- Simplest possible reproduction - Automated test if possible - One-off test script if no framework available - MUST exist before fixing - Use the godmode:test-first skill for writing proper failing tests

  1. Implement a Single Fix

- Address the root cause identified - ONE change at a time - No "while I'm here" improvements - No bundled refactoring

  1. Verify the Fix

- Test passes now? - No other tests broken? - Issue actually resolved?

  1. If the Fix Does Not Work

- STOP - Count: How many fixes have you attempted? - If < 3: Return to Phase 1, re-analyze with new information - If >= 3: STOP and question the architecture (step 5 below) - DO NOT attempt fix #4 without architectural discussion

  1. If 3+ Fixes Failed: Question Architecture Pattern indicating an architectural problem: STOP and question fundamentals: Discuss with your human partner before attempting more fixes This is NOT a failed hypothesis -- this is a flawed architecture.

- Each fix reveals new shared state/coupling/problems in different locations - Fixes require "massive refactoring" to implement - Each fix creates new symptoms elsewhere - Is this pattern fundamentally sound? - Are we persisting through sheer inertia? - Should we refactor the architecture vs. continue fixing symptoms?

Guardrails - STOP and Follow Process

If you catch yourself thinking:

  • "Quick fix for now, investigate later"
  • "Just try changing X and see what happens"
  • "Apply multiple changes, run tests"
  • "Skip the test, I'll verify manually"
  • "It's probably X, let me fix that"
  • "I don't fully understand but this might work"
  • "Pattern says X but I'll adapt differently"
  • "Here are the main problems: [lists fixes without investigation]"
  • Proposing solutions before tracing data flow
  • "One more fix attempt" (when already tried 2+)
  • Each fix reveals new problems in different places

ALL of these mean: STOP. Return to Phase 1.

If 3+ fixes failed: Question the architecture (see Phase 4, step 5)

Human Partner Signals You Are Off Track

Watch for these redirections:

  • "Is that not happening?" - You assumed without verifying
  • "Will it show us...?" - You should have added evidence gathering
  • "Stop guessing" - You are proposing fixes without understanding
  • "Think deeper" - Question fundamentals, not just symptoms
  • "We're stuck?" (frustrated) - Your approach is not working

When you see these: STOP. Return to Phase 1.

Cognitive Traps

RationalizationWhat Is Actually True
"Issue is simple, process not needed"Simple issues have root causes too. The process is fast for simple bugs.
"Emergency, no time for process"Systematic diagnosis is FASTER than guess-and-check flailing.
"Just try this first, then investigate"The first fix sets the pattern. Do it right from the start.
"I'll write the test after confirming the fix works"Untested fixes do not hold. Test first proves it.
"Multiple fixes at once saves time"Cannot isolate what worked. Creates new bugs.
"Reference too long, I'll adapt the pattern"Partial understanding guarantees bugs. Read it completely.
"I see the problem, let me fix it"Seeing symptoms is not the same as understanding root cause.
"One more fix attempt" (after 2+ failures)3+ failures = architectural problem. Question the pattern, do not fix again.

Quick Reference

PhaseKey ActivitiesSuccess Criteria
1. Root CauseRead errors, reproduce, check changes, gather evidenceUnderstand WHAT and WHY
2. PatternFind working examples, compareIdentify differences
3. HypothesisForm theory, test minimallyConfirmed or new hypothesis
4. ImplementationCreate test, fix, verifyBug resolved, tests pass

When Investigation Reveals No Root Cause

If systematic investigation reveals the issue is truly environmental, timing-dependent, or external:

  1. You have completed the process
  2. Document what you investigated
  3. Implement appropriate handling (retry, timeout, error message)
  4. Add monitoring/logging for future investigation

But: 95% of "no root cause" cases are incomplete investigation.

Supporting Methods

These methods are part of fault diagnosis and available in this directory:

  • root-cause-tracing.md - Trace bugs backward through the call stack to find the original trigger
  • defense-in-depth.md - Add validation at multiple layers after finding root cause
  • condition-based-waiting.md - Replace arbitrary timeouts with condition polling

Related skills:

  • godmode:test-first - For creating failing test case (Phase 4, Step 1)
  • godmode:completion-gate - Verify fix worked before declaring success

Real-World Impact

From diagnosis sessions:

  • Systematic approach: 15-30 minutes to resolution
  • Random fix approach: 2-3 hours of flailing
  • First-attempt fix rate: 95% vs 40%
  • New bugs introduced: Near zero vs common

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Codex

33.76%
按下载量换算35

Claude

31.96%
按下载量换算33

Cursor

17.28%
按下载量换算18

Gemini CLI

9.2%
按下载量换算10

安全审计

Gen Agent Trust Hub

通过

Socket

通过

Snyk

通过

权限和风险

需要联网

该 Skill 可能需要联网访问来源站点、仓库或外部 API;具体网络访问范围需要结合源码和 README 复核。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills