Token导航 LogoToken导航TokenDH.com
待分类需要联网github未标认证来源可访问许可证需确认审计通过

pytorch-fsdp2pytorch fsdp2 搜索

Agent Skill

pytorch-fsdp2 用于处理 GitHub 仓库、Issue、Pull Request 和代码协作信息,适合在 Codex、Claude、Cursor、Gemini CLI 中需要围绕仓库状态、代码变更或协作事项进行整理时使用。可结合来源仓库、安装命令和原始 README 继续核验具体用法。安装前建议确认权限范围、维护状态,以及是否会触发联网、命令执行或文件读写。

总安装

552

周安装

23

GitHub Stars

1

下载量

184
CodexClaudeCursorGemini CLI

安装说明

本站只整理中文说明和来源信息,不托管安装包,也不代用户安装。

GitHub

来源数

2

许可证

unknown

最后核验

2026-05-01

来源状态

来源可访问

安装方式

通过对话安装

复制提示词发给支持本地命令或 Skills 的 AI 助手,先确认命令和权限,再让它执行。

请帮我安装这个 Agent Skill:pytorch-fsdp2(pytorch fsdp2 搜索)
来源仓库:https://github.com/kiterlin/intelligent-detection-system
仓库路径:skills/pytorch-fsdp2
安装命令:
npx skills add https://github.com/kiterlin/intelligent-detection-system --skill pytorch-fsdp2
安装前请先检查当前环境是否支持对应 CLI,并向我确认将要执行的命令、安装目录、联网范围和文件读写权限;确认后再执行。

命令行安装

复制命令到本机终端执行。该命令会通过 npx skills 从第三方来源获取 Skill;本站只展示命令,不托管安装包,也不自动执行。

skills.shnpx skills
npx skills add https://github.com/kiterlin/intelligent-detection-system --skill pytorch-fsdp2

简介

用于处理 GitHub 仓库、Issue、Pull Request 和代码协作信息。

  • 适合围绕仓库状态、代码变更或协作事项进行整理。
  • 通过 GitHub 安装,结合原始 README 核验具体用法。
  • 安装前建议确认权限范围和维护状态,避免触发敏感操作。
  • pytorch-fsdp2 属于待分类类 Skill,可作为该场景下的辅助能力补充。

SKILL.md

Skill: Use PyTorch FSDP2 (fully_shard) correctly in a training script

This skill teaches a coding agent how to add PyTorch FSDP2 to a training loop with correct initialization, sharding, mixed precision/offload configuration, and checkpointing.

FSDP2 in PyTorch is exposed primarily via torch.distributed.fsdp.fully_shard and the FSDPModule methods it adds in-place to modules. See: references/pytorch_fully_shard_api.md, references/pytorch_fsdp2_tutorial.md.

When to use this skill

Use FSDP2 when:

  • Your model doesn’t fit on one GPU (parameters + gradients + optimizer state).
  • You want an eager-mode sharding approach that is DTensor-based per-parameter sharding (more inspectable, simpler sharded state dicts) than FSDP1.
  • You may later compose DP with Tensor Parallel using DeviceMesh.

Avoid (or be careful) if:

  • You need strict backwards-compatible checkpoints across PyTorch versions (DCP warns against this).
  • You’re forced onto older PyTorch versions without the FSDP2 stack.

Alternatives (when FSDP2 is not the best fit)

  • DistributedDataParallel (DDP): Use the standard data-parallel wrapper when you want classic distributed data parallel training.
  • FullyShardedDataParallel (FSDP1): Use the original FSDP wrapper for parameter sharding across data-parallel workers.

Reference: references/pytorch_ddp_notes.md, references/pytorch_fsdp1_api.md.


Contract the agent must follow

  1. Launch with torchrun and set the CUDA device per process (usually via LOCAL_RANK).
  2. Apply fully_shard() bottom-up, i.e., shard submodules (e.g., Transformer blocks) before the root module.
  3. Call model(input), not model.forward(input), so the FSDP2 hooks run (unless you explicitly unshard() or register the forward method).
  4. Create the optimizer after sharding and make sure it is built on the DTensor parameters (post-fully_shard).
  5. Checkpoint using Distributed Checkpoint (DCP) or the distributed-state-dict helpers, not naïve torch.save(model.state_dict()) unless you deliberately gather to full tensors.

(Each of these rules is directly described in the official API docs/tutorial; see references.)


Step-by-step procedure

0) Version & environment sanity

  • Prefer a recent stable PyTorch where the docs show FSDP2 and DCP updated recently.
  • Use torchrun --nproc_per_node <gpus_per_node>... and ensure RANK, WORLD_SIZE, LOCAL_RANK are visible.

Reference: references/pytorch_fsdp2_tutorial.md (launch commands and setup), references/pytorch_fully_shard_api.md (user contract).


1) Initialize distributed and set device

Minimal, correct pattern:

  • dist.init_process_group(backend="nccl")
  • torch.cuda.set_device(int(os.environ["LOCAL_RANK"]))
  • Optionally create a DeviceMesh to describe the data-parallel group(s)

Reference: references/pytorch_device_mesh_tutorial.md (why DeviceMesh exists & how it manages process groups).


2) Build model on meta device (recommended for very large models)

For big models, initialize on meta, apply sharding, then materialize weights on GPU:

  • with torch.device("meta"): model =...
  • apply fully_shard(...) on submodules, then fully_shard(model)
  • model.to_empty(device="cuda")
  • model.reset_parameters() (or your init routine)

Reference: references/pytorch_fsdp2_tutorial.md (migration guide shows this flow explicitly).


3) Apply fully_shard() bottom-up (wrapping policy = “apply where needed”)

Do not only call fully_shard on the topmost module.

Recommended sharding pattern for transformer-like models:

  • iterate modules, if isinstance(m, TransformerBlock): fully_shard(m,...)
  • then fully_shard(model,...)

Why:

  • fully_shard forms “parameter groups” for collective efficiency and excludes params already grouped by earlier calls. Bottom-up gives better overlap and lower peak memory.

Reference: references/pytorch_fully_shard_api.md (bottom-up requirement and why).


4) Configure reshard_after_forward for memory/perf trade-offs

Default behavior:

  • None means True for non-root modules and False for root modules (good default).

Heuristics:

  • If you’re memory-bound: keep defaults or force True on many blocks.
  • If you’re throughput-bound and can afford memory: consider keeping unsharded params longer (root often False).
  • Advanced: use an int to reshard to a smaller mesh after forward (e.g., intra-node) if it’s a meaningful divisor.

Reference: references/pytorch_fully_shard_api.md (full semantics).


5) Mixed precision & offload (optional but common)

FSDP2 uses:

  • mp_policy=MixedPrecisionPolicy(param_dtype=..., reduce_dtype=..., output_dtype=..., cast_forward_inputs=...)
  • offload_policy=CPUOffloadPolicy() if you want CPU offload

Rules of thumb:

  • Start with BF16 parameters/reductions on H100/A100-class GPUs (if numerically stable for your model).
  • Keep reduce_dtype aligned with your gradient reduction expectations.
  • If you use CPU offload, budget for PCIe/NVLink traffic and runtime overhead.

Reference: references/pytorch_fully_shard_api.md (MixedPrecisionPolicy / OffloadPolicy classes).


6) Optimizer, gradient clipping, accumulation

  • Create the optimizer after sharding so it holds DTensor params.
  • If you need gradient accumulation / no_sync:

- use the FSDP2 mechanism (set_requires_gradient_sync) instead of FSDP1’s no_sync().

Gradient clipping:

  • Use the approach shown in the FSDP2 tutorial (“Gradient Clipping and Optimizer with DTensor”), because parameters/gradients are DTensors.

Reference: references/pytorch_fsdp2_tutorial.md.


7) Checkpointing: prefer DCP or distributed state dict helpers

Two recommended approaches:

A) Distributed Checkpoint (DCP) — best default

  • DCP saves/loads from multiple ranks in parallel and supports load-time resharding.
  • DCP produces multiple files (often at least one per rank) and operates “in place”.

B) Distributed state dict helpers

  • get_model_state_dict / set_model_state_dict with StateDictOptions(full_state_dict=True, cpu_offload=True, broadcast_from_rank0=True,...)
  • For optimizer: get_optimizer_state_dict / set_optimizer_state_dict

Avoid:

  • Saving DTensor state dicts with plain torch.save unless you intentionally convert with DTensor.full_tensor() and manage memory carefully.

References:

  • references/pytorch_dcp_overview.md (DCP behavior and caveats)
  • references/pytorch_dcp_recipe.md and references/pytorch_dcp_async_recipe.md (end-to-end usage)
  • references/pytorch_fsdp2_tutorial.md (DTensor vs DCP state-dict flows)
  • references/pytorch_examples_fsdp2.md (working checkpoint scripts)

Workflow checklists (copy-paste friendly)

Workflow A: Retrofit FSDP2 into an existing training script

  • Launch with torchrun and initialize the process group.
  • Set the CUDA device from LOCAL_RANK; create a DeviceMesh if you need multi-dim parallelism.
  • Build the model (use meta if needed), apply fully_shard bottom-up, then fully_shard(model).
  • Create the optimizer after sharding so it captures DTensor parameters.
  • Use model(inputs) so hooks run; use set_requires_gradient_sync for accumulation.
  • Add DCP save/load via torch.distributed.checkpoint helpers.

Reference: references/pytorch_fsdp2_tutorial.md, references/pytorch_fully_shard_api.md, references/pytorch_device_mesh_tutorial.md, references/pytorch_dcp_recipe.md.

Workflow B: Add DCP save/load (minimal pattern)

  • Wrap state in Stateful or assemble state via get_state_dict.
  • Call dcp.save(...) from all ranks to a shared path.
  • Call dcp.load(...) and restore with set_state_dict.
  • Validate any resharding assumptions when loading into a different mesh.

Reference: references/pytorch_dcp_recipe.md.

Debug checklist (what the agent should check first)

  1. All ranks on distinct GPUs? If not, verify torch.cuda.set_device(LOCAL_RANK) and your torchrun flags.
  2. Did you accidentally call forward() directly? Use model(input) or explicitly unshard() / register forward.
  3. Is fully_shard() applied bottom-up? If only root is sharded, expect worse memory/perf and possible confusion.
  4. Optimizer created at the right time? Must be built on DTensor parameters *after* sharding.
  5. Checkpointing path consistent?

- If using DCP, don’t mix with ad-hoc torch.save unless you understand conversions. - Be mindful of PyTorch-version compatibility warnings for DCP.


Common issues and fixes

  • Forward hooks not running → Call model(inputs) (or unshard() explicitly) instead of model.forward(...).
  • Optimizer sees non-DTensor params → Create optimizer after all fully_shard calls.
  • Only root module sharded → Apply fully_shard bottom-up on submodules before the root.
  • Memory spikes after forward → Set reshard_after_forward=True for more modules.
  • Gradient accumulation desync → Use set_requires_gradient_sync instead of FSDP1’s no_sync().

Reference: references/pytorch_fully_shard_api.md, references/pytorch_fsdp2_tutorial.md.


Minimal reference implementation outline (agent-friendly)

The coding agent should implement a script with these labeled blocks:

  • init_distributed(): init process group, set device
  • build_model_meta(): model on meta, apply fully_shard, materialize weights
  • build_optimizer(): optimizer created after sharding
  • train_step(): forward/backward/step with model(inputs) and DTensor-aware patterns
  • checkpoint_save/load(): DCP or distributed state dict helpers

Concrete examples live in references/pytorch_examples_fsdp2.md and the official tutorial reference.


References

  • references/pytorch_fsdp2_tutorial.md
  • references/pytorch_fully_shard_api.md
  • references/pytorch_ddp_notes.md
  • references/pytorch_fsdp1_api.md
  • references/pytorch_device_mesh_tutorial.md
  • references/pytorch_tp_tutorial.md
  • references/pytorch_dcp_overview.md
  • references/pytorch_dcp_recipe.md
  • references/pytorch_dcp_async_recipe.md
  • references/pytorch_examples_fsdp2.md
  • references/torchtitan_fsdp_notes.md (optional, production notes)
  • references/ray_train_fsdp2_example.md (optional, integration example)

适合场景

01

用户想查找某类 Agent Skill 时

02

需要根据任务场景推荐可安装能力包时

03

需要对比不同来源的安装命令和来源信息时

能力概览

能力 1

按任务关键词查找相关 Skills

能力 2

展示可复制的安装命令

能力 3

保留来源站点、仓库和原始说明,方便继续核验

能力 4

展示第三方安全扫描或审计结果

安装后应在对应宿主中按原始 README 的触发条件使用;具体调用方式请以来源页面和 README 为准。

平台分布

Codex

34.19%
按下载量换算63

Claude

32.92%
按下载量换算61

Cursor

20.14%
按下载量换算37

Gemini CLI

9.05%
按下载量换算17

安全审计

Gen Agent Trust Hub

通过

Socket

通过

Snyk

通过

权限和风险

需要联网

该 Skill 可能需要联网访问来源站点、仓库或外部 API;具体网络访问范围需要结合源码和 README 复核。

安装前确认

本站仅展示第三方公开信息,不托管安装包,不提供自动安装或运行环境。安装前应自行审查源码、依赖和命令行为。当前只有一个来源,正式发布前建议补源仓库或其他目录站核验。

来源信息

继续浏览同类 Skills