Research Computing Systems Engineer III, Cloud Technologies IS&T Research Computing
- Location
- Boston, MA, United States
- Workplace
- Hybrid
- Compensation
- $100k – $125k
About this role
Boston University Information Services & Technology (IS&T) is seeking applicants with diverse skills and experiences to join our innovative and inclusive community. We're looking for a Research Computing Systems Engineer to help operate and grow the high-performance computing (HPC) environment that powers research across every discipline at BU. In this role, you will keep our production HPC cluster running reliably, help deploy and operate a new NIST 800-171–compliant secure research computing environment built on tiCrypt, and modernize the way we monitor, automate, and deliver services. This is a full-stack, operations- and automation-focused position for an engineer who enjoys working on systems end to end — from writing infrastructure as code, to troubleshooting hardware, to planning physical rack deployments. You'll join a small, deeply technical team where everyone stays hands-on and has a real voice in shaping how we operate our infrastructure. You'll work closely with senior engineers, collaborate with other HPC professionals, support researchers, and also work independently. This position is partially remote.
You Will:
- Operate, maintain, and optimize the University's production HPC cluster, including a 1,000+ node bare-metal Linux deployment, the Grid Engine (SGE) scheduler, and GPFS (IBM Storage Scale) parallel file system, to keep research computing services reliable and performant.
- Contribute to the build, deployment, and operations of a new NIST 800-171–compliant secure research computing environment built on tiCrypt, including a Slurm workload manager.
- Help improve monitoring, observability, and configuration management using CI/CD pipelines, Ansible, Git, and related automation tools.
- Automate provisioning, deployment, and routine operations across the full stack, from bare metal through services.
- Identify, diagnose, and resolve complex system, storage, network, and performance issues.
- Contribute to technical documentation and share your expertise with colleagues and the research community.
Tired of cold applications?
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
Know someone who'd be great for this?