← Back to Articles
GPU & AI Solutions • 8 min read

GPU & AI Solutions

Navigating the Volatility of Interruptible GPU Instances

In the rapidly evolving landscape of artificial intelligence and machine learning, optimizing computational costs is paramount. Interruptible GPU instances, often referred to as spot instances or preemptible VMs, offer a compelling solution by providing significant cost savings compared to on-demand alternatives. These instances leverage surplus capacity from cloud providers such as AWS Canada (Central), Azure Canada East, or Google Cloud Platform Canada (Montreal/Toronto), often at discounts ranging from 70% to 90% below standard rates. The trade-off, however, is their ephemeral nature: providers can reclaim these resources with minimal notice, typically 30 seconds to 2 minutes, posing a substantial risk to long-running, compute-intensive tasks like deep learning model training.

Successfully harnessing the economic advantages of interruptible GPUs requires robust strategies for safeguarding training progress. Losing hours or even days of computation due to an instance preemption is not only frustrating but can negate any initial cost savings. The core challenge lies in efficiently and reliably saving the model's state, optimizers, and other critical metadata, known as a 'checkpoint', so that training can seamlessly resume from the last saved point on a new instance. This article delves into several prominent strategies for achieving this, detailing their technical mechanics and practical trade-offs.

Periodic Checkpoints to Local Disk: The Foundational Approach

The most straightforward method for state preservation involves saving checkpoints to the local disk attached to the GPU instance. Deep learning frameworks like PyTorch and TensorFlow provide native capabilities for this:

These operations typically write a file to a directory on the local storage, which might be a solid-state drive (SSD) for performance. The process is usually triggered at regular intervals, such as every 'N' training steps or epochs, or after a specific performance metric improves.

Advantages:

Disadvantages:

Synchronizing to Object Storage for Durability

To overcome the ephemeral nature of local disks, a common and effective strategy is to periodically transfer checkpoints from the local disk to a durable object storage service. These services, such as S3-compatible object storage (e.g., offered by major cloud providers with Canadian regions like AWS S3 Canada (Central) or Azure Blob Storage Canada East), are designed for high durability and availability.

The process usually involves:

  1. Saving the checkpoint to local disk first.
  2. Uploading the checkpoint file(s) to the object storage bucket using cloud SDKs (e.g., boto3 for AWS S3, Azure SDK for Blob Storage) or command-line tools like gsutil for Google Cloud Storage.
  3. Implementing versioning on the object storage bucket to retain multiple checkpoints and prevent accidental overwrites.

Advantages:

Disadvantages:

Leveraging Network-Attached Volumes

Network-attached volumes provide persistent storage that can be detached from one instance and reattached to another, or even mounted by multiple instances simultaneously. This approach allows checkpoints to be written directly to durable storage without an explicit upload step.

Examples include:

When an interruptible instance receives a preemption notice, the application can attempt to perform a final, rapid checkpoint to the mounted network volume. Upon restarting on a new instance, the same volume is mounted, and training resumes.

Advantages:

Disadvantages:

Continuous File Synchronization Services

This strategy employs dedicated tools or services that continuously monitor a local directory for changes and synchronize them to a remote, durable storage location. Instead of explicit calls within the training script, the synchronization happens in the background.

Examples include:

Advantages:

Disadvantages:

Choosing the Right Strategy: A Comparative View

The optimal checkpointing strategy for interruptible GPUs is not universal; it depends heavily on the specific requirements of the training workload, acceptable recovery time objectives (RTO), and budget constraints.

A hybrid approach is often the most robust. For instance, saving frequent lightweight checkpoints locally for fast recovery from minor errors, while periodically (e.g., every few hours or epochs) synchronizing a full, consistent checkpoint to durable object storage. This balances speed with long-term safety.

Implementing Robust Checkpointing Best Practices

Regardless of the chosen strategy, several best practices enhance the reliability of checkpointing on interruptible GPUs:

By carefully considering these strategies and adopting best practices, AI and ML practitioners in Canada can confidently leverage the cost efficiencies of interruptible GPUs without compromising the integrity and progress of their valuable training workloads.

Seamless GPU Job Recovery

Move your interrupted training jobs across cloud providers.

pip install spotwarp
← Return to GPU-Action Main Portal