huggingface/accelerate
🚀 A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (including fp8), and easy-to-configure FSDP and DeepSpeed support
https://skillcdn.ai/gh/huggingface/accelerateConnect this address to your AI to use these skills. How to connect
- Unverified
- Default branch
main - Commit
01c73fb - License
Apache-2.0
Explore this repository
Searches the original text of skills and documents; display translations are not searched.
Documents
- Big Model InferenceOne of the biggest advancements Accelerate provides is Big Model Inference, which allows you to perform inference with models that don't fully fit on your graphics card.docs/source/usage_guides/big_modeling.md
- CheckpointingWhen training a PyTorch model with Accelerate, you may often want to save and continue a state of training. Doing so requires saving and loading the model, optimizer, RNG generators, and the GradScal…docs/source/usage_guides/checkpoint.md
- CompilationPytorch 2.0 introduced torch.compile, a powerful feature that makes PyTorch code run faster by JIT-compiling PyTorch code into optimized kernels. Key features of torch.compile include:docs/source/usage_guides/compilation.md
- DDP Communication HooksDistributed Data Parallel (DDP) communication hooks provide a generic interface to control how gradients are communicated across workers by overriding the vanilla allreduce in DistributedDataParallel…docs/source/usage_guides/ddp_comm_hook.md
- DeepSpeedDeepSpeed implements everything described in the ZeRO paper. Some of the salient optimizations are:docs/source/usage_guides/deepspeed.md
- Using multiple models with DeepSpeedThis guide assumes that you have read and understood the DeepSpeed usage guide.docs/source/usage_guides/deepspeed_multiple_model.md
- Distributed inferenceDistributed inference can fall into three brackets:docs/source/usage_guides/distributed_inference.md
- Start Here!Please use the interactive tool below to help you get started with learning about a particular feature of Accelerate and how to utilize it! It will provide you with a code diff, an explanation toward…docs/source/usage_guides/explore.md
- Fully Sharded Data ParallelTo accelerate training huge models on larger batch sizes, we can use a fully sharded data parallel model. This type of data parallel paradigm enables fitting more data and larger models by sharding t…docs/source/usage_guides/fsdp.md
- Intel GaudiUsers can take advantage of Intel Gaudi AI accelerators for significantly faster and cost-effective model training and inference. The Intel Gaudi AI accelerator family currently includes three produc…docs/source/usage_guides/gaudi.md
- Performing gradient accumulation with AccelerateGradient accumulation is a technique where you can train on bigger batch sizes than your machine would normally be able to fit into memory. This is done by accumulating gradients over several batches…docs/source/usage_guides/gradient_accumulation.md
- Training on Intel CPUAccelerate has full support for Intel CPU, all you need to do is enabling it through the config.docs/source/usage_guides/intel_cpu.md
- Using Local SGD with AccelerateLocal SGD is a technique for distributed training where gradients are not synchronized every step. Thus, each process updates its own version of the model weights and after a given number of steps th…docs/source/usage_guides/local_sgd.md
- Low Precision Training MethodsAccelerate provides integrations to train on lower precision methods using specified supported hardware through the TransformersEngine, MS-AMP, and torchao packages. This documentation will help guid…docs/source/usage_guides/low_precision_training.md
- Megatron-LMMegatron-LM enables training large transformer language models at scale. It provides efficient tensor, pipeline and sequence based model parallelism for pre-training transformer based Language Models…docs/source/usage_guides/megatron_lm.md
- Model memory estimatorOne very difficult aspect when exploring potential models to use on your machine is knowing just how big of a model will fit into memory with your current device (such as loading the model onto CUDA…docs/source/usage_guides/model_size_estimator.md
- Accelerated PyTorch Training on MacWith PyTorch v1.12 release, developers and researchers can take advantage of Apple silicon GPUs for significantly faster model training. This unlocks the ability to perform machine learning workflows…docs/source/usage_guides/mps.md
- ProfilerProfiler is a tool that allows the collection of performance metrics during training and inference. Profiler’s context manager API can be used to better understand what model operators are the most e…docs/source/usage_guides/profiler.md
- Model quantizationAccelerate brings bitsandbytes quantization to your model. You can now load any pytorch model in 8-bit or 4-bit with a few lines of code.docs/source/usage_guides/quantization.md
- Amazon SageMakerHugging Face and Amazon introduced new Hugging Face Deep Learning Containers (DLCs) to make it easier than ever to train Hugging Face Transformer models in Amazon SageMaker.docs/source/usage_guides/sagemaker.md
- Experiment trackersThere are a large number of experiment tracking APIs available, however getting them all to work in a multi-processing environment can oftentimes be complex. Accelerate provides a general tracking AP…docs/source/usage_guides/tracking.md
- Example ZooBelow contains a non-exhaustive list of tutorials and scripts showcasing Accelerate.docs/source/usage_guides/training_zoo.md