Spaces:

thenativefox
/

RAG

Sleeping

RAG / sentence-transformers_all-MiniLM-L6-v2 /recursive_chunks /_accelerate.txt_chunk_0.txt

thenativefox

Added split files and tables

939262b 11 months ago

2.42 kB

	Distributed training with 🤗 Accelerate
	As models get bigger, parallelism has emerged as a strategy for training larger models on limited hardware and accelerating training speed by several orders of magnitude. At Hugging Face, we created the 🤗 Accelerate library to help users easily train a 🤗 Transformers model on any type of distributed setup, whether it is multiple GPU's on one machine or multiple GPU's across several machines. In this tutorial, learn how to customize your native PyTorch training loop to enable training in a distributed environment.
	Setup
	Get started by installing 🤗 Accelerate:

	pip install accelerate
	Then import and create an [~accelerate.Accelerator] object. The [~accelerate.Accelerator] will automatically detect your type of distributed setup and initialize all the necessary components for training. You don't need to explicitly place your model on a device.

	from accelerate import Accelerator
	accelerator = Accelerator()

	Prepare to accelerate
	The next step is to pass all the relevant training objects to the [~accelerate.Accelerator.prepare] method. This includes your training and evaluation DataLoaders, a model and an optimizer:

	train_dataloader, eval_dataloader, model, optimizer = accelerator.prepare(
	train_dataloader, eval_dataloader, model, optimizer
	)

	Backward
	The last addition is to replace the typical loss.backward() in your training loop with 🤗 Accelerate's [~accelerate.Accelerator.backward]method:

	for epoch in range(num_epochs):
	for batch in train_dataloader:
	outputs = model(**batch)
	loss = outputs.loss
	accelerator.backward(loss)

	optimizer.step()
	lr_scheduler.step()
	optimizer.zero_grad()
	progress_bar.update(1)

	As you can see in the following code, you only need to add four additional lines of code to your training loop to enable distributed training!

	+ from accelerate import Accelerator
	from transformers import AdamW, AutoModelForSequenceClassification, get_scheduler

	accelerator = Accelerator()

	model = AutoModelForSequenceClassification.from_pretrained(checkpoint, num_labels=2)
	optimizer = AdamW(model.parameters(), lr=3e-5)

	device = torch.device("cuda") if torch.cuda.is_available() else torch.device("cpu")

	model.to(device)

	train_dataloader, eval_dataloader, model, optimizer = accelerator.prepare(

	train_dataloader, eval_dataloader, model, optimizer
	)