Harvard University

Bioinformatics Software Engineer

Full-time · Boston, MA
✓ Verified live on the employer's own system · added 167 days ago
Save search
Mid-level · 5+ yrs exp

Requirements

Education: Doctorate or related field

Experience: 5+ years

Skills & tools

Data AnalysisProgrammingQuality AssuranceGitManagementCloud PlatformsDevopsResearch
Apply on company site ↗ See your fit → free

Full job description

Participate in the design of software that supports and enriches research productivity and reliability; implement software solutions. Develop software and data services with researchers to ensure that modern standards of reproducible code are kept.

We are looking for a highly skilled Bioinformatics Software Engineer who specializes in designing, developing, deploying, and maintaining scalable bioinformatics pipelines on cloud-based infrastructure. The candidate will be responsible for the code base supporting the large-scale genomic processing and analysis pipelines at the SMaHT Data Analysis Center that manages multi-omic data (e.g., Illumina/PacBio/ONT Whole Genome Sequencing (WGS), RNA-Seq).

The ideal candidate will have a deep understanding of next-generation sequencing (NGS) data analysis, workflow automation, cloud computing, and cloud software engineering best practices. This role will support research and production environments where reproducibility, scalability, and performance are critical.

- Design, implement, and maintain bioinformatics pipelines for high-throughput sequencing data (e.g., alignment, QC, variant calling from WGS and RNA-seq) similar to those in existing repositories: https://github.com/smaht-dac/main-pipelines . - Build reproducible, well-tested, and automated workflows using workflow management systems (particularly CWL). - Architect and manage AWS-based compute infrastructure to support pipeline execution, including automated deployment, scaling, and monitoring. - Containerize workflows using Docker or similar tools for managed execution and portability. - Integrate CI/CD tooling to automate testing, deployment, and version control to ensure data integrity and correct execution of the pipeline. - Develop utility tools for metadata management, file integrity checks or conversion (e.g., VCF, BAM to CRAM), and integration with the SMaHT Data Portal. - Collaborate cross-functionally with research scientists, engineers, and IT teams to refine requirements and deliver high-quality solutions. - Document code, workflows, and infrastructure configurations clearly.

- Minimum of five years' post-secondary education or relevant work experience.

- PhD in computational biology/bioinformatics/statistics/CS or another quantitative field is strongly preferred. - Superb programming skills, especially in Python and shell scripting, and communication skills are strongly preferred. - Extensive experience with analysis of high-throughput sequencing data and knowledge of bioinformatics tools for sequence alignment, variant calling, sequence data QC, etc. - Proficiency in Docker for creating a reproducible execution environment and Workflow Description Language for orchestrating complex tasks. - Strong understanding of AWS services (EC2, S3) or similar cloud platforms for compute and storage. - Version Control & CI/CD: Git, automated testing, deployment workflows. - Experience with Linux systems, HPC, and distributed computing environments. - Knowledge of optimizing pipelines for large-scale genomic projects.

- Appointment End Date: This is a one-year term position from the date of hire, with the possibility of extension, contingent upon work performance and continued funding to support the position. - Standard Hours/Schedule: 35 hours per week - Visa Sponsorship Information: Harvard University is unable to provide visa sponsorship for this position. - Pre-Employment Screening: Identity - Other Information: Please note that we are currently conducting a majority of interviews and onboarding remotely and virtually.

We appreciate your understanding. - Staying Informed About Your Application : Due to the high volume of applications, we may not always be able to reach out right away, but you can track your status anytime through the Careers@Harvard portal.

This position is salary grade level 057. Please visit Harvard's Salary Ranges to view the corresponding salary range and related information.

Harvard has an equal employment opportunity policy that outlines our commitment to prohibiting discrimination on the basis of race, ethnicity, color, national origin, sex, sexual orientation, gender identity, veteran status, religion, disability, or any other characteristic protected by law or identified in the university's non-discrimination policy .

Harvard's equal employment opportunity policy and non-discrimination policy help all community members participate fully in work and campus life free from harassment and discrimination. Basic

  • Minimum of five years’ post-secondary education or relevant work experience.
  • PhD in computational biology/bioinformatics/statistics/CS or another quantitative field is strongly preferred.
  • Superb programming skills, especially in Python and shell scripting, and communication skills are strongly preferred.
  • Extensive experience with analysis of high-throughput sequencing data and knowledge of bioinformatics tools for sequence alignment, variant calling, sequence data QC, etc.
  • Proficiency in Docker for creating a reproducible execution environment and Workflow Description Language for orchestrating complex tasks.
  • Strong understanding of AWS services (EC2, S3) or similar cloud platforms for compute and storage.
  • Version Control & CI/CD: Git, automated testing, deployment workflows.
  • Experience with Linux systems, HPC, and distributed computing environments.
  • Knowledge of optimizing pipelines for large-scale genomic projects.

Job Summary:

Participate in the design of software that supports and enriches research productivity and reliability; implement software solutions. Develop software and data services with researchers to ensure that modern standards of reproducible code are kept.

We are looking for a highly skilled Bioinformatics Software Engineer who specializes in designing, developing, deploying, and maintaining scalable bioinformatics pipelines on cloud-based infrastructure. The candidate will be responsible for the code base supporting the large-scale genomic processing and analysis pipelines at the SMaHT Data Analysis Center that manages multi-omic data (e.g., Illumina/PacBio/ONT Whole Genome Sequencing (WGS), RNA-Seq).

The ideal candidate will have a deep understanding of next-generation sequencing (NGS) data analysis, workflow automation, cloud computing, and cloud software engineering best practices. This role will support research and production environments where reproducibility, scalability, and performance are critical.

  • Design, implement, and maintain bioinformatics pipelines for high-throughput sequencing data (e.g., alignment, QC, variant calling from WGS and RNA-seq) similar to those in existing repositories: https://github.com/smaht-dac/main-pipelines .
  • Build reproducible, well-tested, and automated workflows using workflow management systems (particularly CWL).
  • Architect and manage AWS-based compute infrastructure to support pipeline execution, including automated deployment, scaling, and monitoring.
  • Containerize workflows using Docker or similar tools for managed execution and portability.
  • Integrate CI/CD tooling to automate testing, deployment, and version control to ensure data integrity and correct execution of the pipeline.
  • Develop utility tools for metadata management, file integrity checks or conversion (e.g., VCF, BAM to CRAM), and integration with the SMaHT Data Portal.
  • Collaborate cross-functionally with research scientists, engineers, and IT teams to refine requirements and deliver high-quality solutions.
  • Document code, workflows, and infrastructure configurations clearly.

More jobs at Harvard University

Similar jobs near Boston, MA

Tell me when more Software Engineer, Quality Integration jobs post near Boston, MA We re-check every listing against the employer’s own board — no résumé needed.

Search Bioinformatics Software Engineer jobs near Boston, MA → Browse all live jobs

This posting was published by Harvard University on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.