There will be a scheduled downtime of all the HPC systems of NHR@FAU starting on Monday, August 31 at 6:00, and in parts lasting until Friday, September 04.
As usual, Jobs that would collide with the downtime will automatically be postponed until the downtime is over. Frontends and most fileservers will be available through most of the downtime with some short interruptions.
Reason for the downtime is an assortment of maintenance work, which is also why the duration of the downtime will vary a lot depending on the system.
All systems
On Monday, August 31 and Tuesday, September 01 we will perform maintenance on central systems. On these days, all our clusters will be unavailable, and, while frontends and fileservers will be available most of the time, there will be some service interruptions. All systems except the ones mentioned below should be back to normal operation by the evening of September 01.
Helma
On Helma, there will be maintenance of the internal cooling loops for the GPU part. This is expected to take the whole week, so Helma-GPU will be unavailable until Friday, September 04. Note that this replaces the emergency maintenance we had originally announced for August 24 – 28, which will now NOT take place, so Helma will be unavailable from August 31 to September 04 but will be available from August 24 to August 28.
The infrastructure maintenance could not be finished in time with success!
Helma remains unavailable at least until Monday, Sept. 7th (including).
Alex
On Alex, there is maintenance of central servers. This is also expected to take the whole week, so Alex will be unavailable until Friday, September 04. This will also include the login nodes, which will not permit login for extended periods of time.
Updates
Check back here for regular updates.
- 2026-09-01T11:00 – Batch processing on Woody, TinyGPU and TinyFAT has been resumed
- 2026-09-01T13:45 – Batch processing on Testcluter has been resumed
- 2026-09-01T22:00 – Batch processing on parts of Fritz has been resumed
Access toanvmeworkspaces is not yet possible and will only become available when Alex is also up and running hopefully by the end of the week. - 2026-09-03T17:30 – Batch processing on Alex has been temporarily resumed. We are not fully done with the maintenance on Alex yet, but the system should be stable enough to process jobs, so we have made the decision to resume batch processing a day earlier than planned. There will be another short interruption (i.e. less than one hour) on Alex at the beginning of next week to finish the maintenance.
- 2026-09-04 – infrastructure work on Helma could not yet be finished with success. Helma remains down.
Login nodes are available and access to thehnvmeworkspaces is possible. - 2026-09-09 – The downtime has been finished, all clusters are operational again.
