Key Responsibilities
Design, install, and maintain HPC clusters, including compute nodes, storage, and networking components.
Configure and optimize SLURM workloads, ensuring efficient job scheduling and resource allocation.
Develop and maintain Python scripts for automation, monitoring, and performance analysis.
Collaborate with researchers to troubleshoot performance bottlenecks and implement tuning strategies.
Integrate hybrid cloud resources (AWS/GCP) to extend capacity and provide elas