AWS Parallel Computing Service (PCS) now allows you to reboot compute nodes using Slurm commands without triggering instance replacement. With this feature, you can reboot nodes for operational reasons such as troubleshooting, resource cleanup, and recovery from degraded states before requiring full node replacement, enabling you to efficiently maintain cluster health at lower costs.
This feature is available in all AWS Regions where PCS is available. You can use the ‘scontrol reboot’ command with options to schedule immediate or deferred reboots, while reboots through other methods will continue to trigger instance replacement. To learn more, refer to Rebooting compute nodes with Slurm in AWS PCS.
PCS is a managed service that simplifies running and scaling high performance computing (HPC) workloads on AWS using Slurm. To learn more about PCS, refer to the service documentation.
Categories: marketing:marchitecture/management-tools,general:products/aws-govcloud-us,marketing:marchitecture/compute
Source: Amazon Web Services
Latest Posts
- MC796790: Microsoft 365 Unifies App and Agent Management Across Teams, Outlook, and Admin Centers

- Claude Opus 5 is now available in AWS GovCloud (US)

- Amazon Quick Microsoft 365 extensions are now generally available

- MC796790: Microsoft 365 Unifies App and Agent Availability and Installation Management Across Teams, Outlook, and Admin Centers







