skillme
BrowsePacksCreative OSConnect
skillme

Agent skills for Claude, Codex, and Cursor. Connect once, use them anywhere.

Explore

  • Browse skills
  • Skill packs
  • Trending
  • Live demo

Get started

  • Connect your AI
  • For teams
  • Submit a skill
  • Media guide

Company

  • About
  • News
  • Changelog
  • GitHub
  • Threads

Legal

  • Privacy
  • Terms

© 2026 Skill Me. Built in the open by Alexander Ouellet.

Skill Me is an independent project. It is not affiliated with, endorsed by, or sponsored by Anthropic. “Claude” is a trademark of Anthropic.

  • Home
  • Browse
  • Packs
  • Connect
Browse/Data/Pytorch Fsdp2
DataVerified

Pytorch Fsdp2

by Orchestra Research

Adds PyTorch FSDP2 (fully_shard) to training scripts with correct init, sharding, mixed precision/offload config, and distributed checkpointing. Use when models exceed single-GPU memory or when you need DTensor-based …

mlllmresearchfine-tuning

Read exactly what Claude will be told

SKILL.md
View source on GitHub →
---
name: pytorch-fsdp2
description: Adds PyTorch FSDP2 (fully_shard) to training scripts with correct init, sharding, mixed precision/offload config, and distributed checkpointing. Use when models exceed single-GPU memory or when you need DTensor-based sharding with DeviceMesh.
version: 1.0.0
author: Orchestra Research
license: MIT
tags: [PyTorch, FSDP2, Fully Sharded Data Parallel, Distributed Training, DTensor, Device Mesh, Sharded Checkpointing, Mixed Precision, Offload, Torch Distributed]
dependencies: [torch]
---

# Skill: Use PyTorch FSDP2 (`fully_shard`) correctly in a training script

This skill teaches a coding agent how to **add PyTorch FSDP2** to a training loop with correct initialization, sharding, mixed precision/offload configuration, and checkpointing.

> FSDP2 in PyTorch is exposed primarily via `torch.distributed.fsdp.fully_shard` and the `FSDPModule` methods it adds in-place to modules. See: `references/pytorch_fully_shard_api.md`, `references/pytorch_fsdp2_tutorial.md`.

---

## When to use this skill

Use FSDP2 when:
- Your model **doesn’t fit** on one GPU (parameters + gradients + optimizer state).
- You want an eager-mode sharding approach that is **DTensor-based per-parameter sharding** (more inspectable, simpler sharded state dicts) than FSDP1.  
- You may later compose DP with **Tensor Parallel** using **DeviceMesh**.
… load the full skill through Skill Me

Ask Claude to install

In any Claude conversation, say:

Install the Pytorch Fsdp2 skill

If full content is available, it applies in this conversation and stays installed for future sessions.

Not connected yet? Connect your AI first →

MCP endpoint

https://skillme.dev/api/mcp
Post on X
Source on GitHub →