Software
Are you passionate about Linux, automation, cloud technologies, and High-Performance Computing (HPC)? Join us as a Junior Site Reliability Engineer and gain hands-on experience in building, maintaining, and optimizing next-generation computing infrastructure.
Your role
Key responsibilities in your new role
- Monitor and maintain HPC clusters to ensure system reliability and uptime.
- Assist in troubleshooting hardware, software, storage, and networking issues.
- Develop and support automation scripts using Python and Bash.
- Help implement and optimize monitoring tools such as Prometheus and Grafana.
- Participate in incident management, root cause analysis, and issue resolution.
- Collaborate with senior engineers to support HPC users and workloads.
- Create and maintain technical documentation, runbooks, and knowledge articles.
- Continuously learn and develop expertise in HPC, cloud technologies, and Site Reliability Engineering practices.
Your profile
Qualifications and skills to help you succeed
- Bachelor’s degree in Computer Science, IT, Engineering, or a related field.
- Basic knowledge of Linux (RHEL, CentOS, Ubuntu).
- Familiarity with Python, Bash, C++, or similar scripting/programming languages.
- Understanding of networking fundamentals (TCP/IP, DNS) and system administration.
- Interest in HPC, cloud infrastructure, or distributed systems.
- Strong problem-solving and learning mindset.
- Good communication and teamwork skills.
- Ability to work effectively with and learn from senior team members.
Bengaluru, Karnataka, India
On-site
Full-time
1 hour ago