huggingface/accelerate
🚀 A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (including fp8), and easy-to-configure FSDP and DeepSpeed support
https://skillcdn.ai/gh/huggingface/accelerateConnect this address to your AI to use these skills. How to connect
- Unverified
- Default branch
main - Commit
01c73fb - License
Apache-2.0
Explore this repository
Searches the original text of skills and documents; display translations are not searched.
Documents
- Loading big models into memoryWhen loading a pre-trained model in PyTorch, the usual workflow looks like this:docs/source/concept_guides/big_model_inference.md
- Context Parallel in 🤗`accelerate`This guide will cover basics of using context parallelism in 🤗accelerate, for the more curious readers, we will also cover some technicalities in the later sections.docs/source/concept_guides/context_parallelism.md
- Executing and deferring jobsWhen you run your usual script, instructions are executed in order. Using Accelerate to deploy your script on several GPUs at the same time introduces a complication: while each process executes all…docs/source/concept_guides/deferring_execution.md
- FSDP1 vs FSDP2This guide explains the key differences between FSDP1 and FSDP2 and helps you migrate your existing code to use FSDP2 with minimal changes.docs/source/concept_guides/fsdp1_vs_fsdp2.md
- FSDP vs DeepSpeedAccelerate offers flexibility of training frameworks, by integrating two extremely powerful tools for distributed training, namely Pytorch FSDP and Microsoft DeepSpeed. The aim of this tutorial is to…docs/source/concept_guides/fsdp_and_deepspeed.md
- Gradient synchronizationPyTorch's distributed module operates by communicating back and forth between all of the GPUs in your system. This communication takes time, and ensuring all processes know the states of each other h…docs/source/concept_guides/gradient_synchronization.md
- Accelerate's internal mechanismsInternally, Accelerate works by first analyzing the environment in which the script is launched to determine which kind of distributed setup is used, how many different processes there are and which…docs/source/concept_guides/internal_mechanism.md
- Low precision training methodsThe release of new kinds of hardware led to the emergence of new training paradigms that better utilize them. Currently, this is in the form of training in 8-bit precision using packages such as Tran…docs/source/concept_guides/low_precision_training.md
- Comparing performance across distributed setupsEvaluating and comparing the performance from different setups can be quite tricky if you don't know what to look for. For example, you cannot run the same script with the same batch size across TPU,…docs/source/concept_guides/performance.md
- Sequence parallel in 🤗`accelerate`This guide will cover basics of using sequence parallelism in 🤗accelerate.docs/source/concept_guides/sequence_parallelism.md
- Training on TPUsTraining on TPUs can be slightly different from training on multi-gpu, even with Accelerate. This guide aims to show you where you should be careful and why, as well as the best practices in general.docs/source/concept_guides/training_tpu.md