- Salary
- $5k – $6k
- Location
- Garvan / TKCC Sydney, Australia
- Type
- Full-time
- Source
- Workday
Description
THE OPPORTUNITY
The Summer Scholarship Program, provides an exceptional opportunity for currently enrolled undergraduate students to engage in research projects during the summer of 2026/2027.
This program is designed to immerse highly talented undergraduates, particularly those in Science or related disciplines, in the research process. It aims to enrich your educational journey and inspire a deeper interest in research or related fields.
Participants will gain invaluable experience by working alongside our esteemed supervisors on meaningful research projects, providing a fantastic insight into the research process and helping you determine if a research career is right for you.
The program spans 8 -10 weeks, offering 9 scholarship positions. Each scholarship is valued up to $5,000 - $6,250, with funding allocated at a rate of $625 per week based on the program's duration.
We encourage you to seize this opportunity to enhance your academic and professional development.
WHAT YOU WILL DO
Our range of projects for the Summer of 2026/2027:
1. Inferring cell-type-specific gene regulatory networks from single-cell multi-omics and CRISPR perturbation data using deep learning
Over 90% of disease-associated genetic variants sit in noncoding regions and act by changing gene regulation, often through distal trans-regulatory networks that remain poorly resolved. This project will contribute to developing a deep learning method to infer cell-type-specific gene regulatory networks by integrating TenK10K Phase 1 data (scRNA-seq and scATAC-seq of 10 million cells from ~2,000 donors, matched with whole-genome sequencing), one of the largest single-cell genomics datasets in the world, along with externally curated CRISPR perturbation datasets. First, the student will curate and harmonise multi-omics and perturbation data for blood immune cell types. Second, they will get involved in developing a deep learning approach that combines multi-omics evidence with perturbation signal. Third, they will apply the method to characterise trans-regulatory interactions linking noncoding variants to downstream immune functions and regulatory pathways..
What you will learn
New skills relevant to single-cell cohort curation, multi-omics integration, gene regulatory network inference and deep learning techniques.
Prerequisites
- Have basic coding skills and a foundational understanding of biology and genetics.
- Experience in bioinformatics, coding agents, or single-cell data analysis is highly desirable.
2. Investigating the reactivation of developmental programs in cancer using machine learning
It has been long hypothesised that malignant tumours are capable of undergoing "oncofetal reprogramming" - a phenomena, which refers to cancerous tissues reactivating molecular programs normally restricted to embryo and fetal development. This has been shown to occur in certain cancers and associated with worse therapy response, however, many general questions remain unanswered. For example, what is the biological nature of such a signal and how robust is it? During this summer program, we are looking for a student who would be interested in conducting computational biology research and building relevant machine learning models that would help us shine light on these problems.
What you will learn
The candidate will learn how to handle high-dimensional molecular datasets and use them to build machine learning models
Prerequisites
- Coding skills, basic knowledge of probability theory/statistics; interest in biology would help
3. Building a protein annotation database for spatial mass spectrometry imaging of the tumour matrisome
MALDI mass spectrometry imaging (MSI) underway within the ACRF MATRIX Centre at Garvan generates thousands of unannotated peptide and protein mass features per tissue section, capturing the spatial architecture of the tumour extracellular matrix as it evolves during cancer progression and response to treatment. Currently, annotating these features against tissue type, disease stage and protein identity is manual and unscalable, presenting a major bottleneck to the teams goal of precision matrix medicine. Building on our in house computational pipelines, this project will design and populate a curated peptide/protein annotation database purpose-built for matrisome peptides. The student will develop database architecture linking detected mass features, putative IDs, and tissue/cancer context, enabling rapid, reproducible annotation across studies and supporting discovery of matrisomal biomarkers relevant to stromal activity and therapy response. The project will be cross-supervised by researchers from the Cancer Ecosystems Program and the Data Science Platform.
What you will learn
The student will gain hands-on experience in database design and curation, scripting in Python and/or R, and the handling of large spatial mass spectrometry imaging datasets. They will also develop a working understanding of matrisome biology and peptide/protein annotation, and how curated reference resources underpin biomarker discovery in spatial proteomics.
Prerequisites
-Proficiency in R and/or Python, with confidence working from the command line
- Familiarity with relational databases (such as SQL) and database design principles
- Strong attention to detail and an interest in structured data curation and accessibility
- No prior wet-lab or proteomics experience required, but genuine curiosity about mass spectrometry and large 'omics datasets is essential
4. How Diverse Is HLA Gene Regulation Across Human Ancestries?
The HLA genes play a crucial role in the immune system, allowing immune cells recognise and respond to threats such as pathogens and cancer. They are among the most genetically diverse genes in the human genome, meaning that two individuals are likely to carry different HLA variants, with some of this genetic diversity shaped by human ancestry. While HLA genetic diversity has been extensively studied, much less is known about how the regulation of these genes differs between individuals. Gene regulation refers to the processes that control how much a gene is expressed, allowing cells to adjust their activity in response to different signals. In this project, we will use large-scale sequencing datasets from individuals with diverse genetic ancestries to investigate how HLA genes are regulated. We will identify differences in HLA gene regulation between individuals and populations, providing new insights into the diversity of immune function across human populations.
Prerequisites
- Some coding experience is required.
- No prior experience in immunology is necessary, but a keen interest in the immune system is essential.
5. ASTRA-CYTE: A Machine Learning Pipeline for Cell Type Annotation in Spatial Transcriptomics
Accurate cell type annotation remains a bottleneck in spatial transcriptomics, where sparse transcript counts, ambient RNA contamination, and imperfect cell segmentation degrade classification performance. We have developed ASTRA-CYTE, an annotation framework for our 2,000-gene spatial panel, and now aim to improve its accuracy and robustness. This 10-week project has two phases. First, the student will curate a comprehensive, panel-specific marker gene list from the literature and use it to generate high-confidence manual annotations of cell populations in our spatial data. This ground-truth reference will define the cell types the panel can realistically resolve and expose where current annotations fail. Second, the student will adapt ASTRA-CYTE against these manual annotations, evaluating and mitigating the technical artefacts that drive misclassification such as ambient RNA spillover, low transcript-per-cell counts, and segmentation errors that mix neighbouring cells.
What you will learn
The student will gain experience in spatial transcriptomics, marker-based cell typing, and benchmarking computational methods against expert-curated ground truth.
Prerequisites
- Basic coding in Python or R
6. Synthetic patients: evaluating agentic AI in rare disease diagnosis
About half of families with a rare genetic disease never receive a molecular diagnosis. Agentic AI is poised to rapidly improve this, although our ability to evaluate, benchmark, and train these tools with real data is limited by privacy, training data contamination, and small sample sizes. To help overcome this, you will develop a pipeline to generate a corpus of synthetic rare disease case data. This will include both genomic data (read-level data and annotated variant calls) and paired clinical data (structured phenotypes and unstructured clinical notes), with datasets calibrated against real cases in the published literature and our own patient cohorts (n>3,000). Time permitting, you will also use this dataset for targeted evaluations and benchmarks.
What you will learn
The student will gain experience in spatial transcriptomics, marker-based cell typing, and benchmarking computational methods against expert-curated ground truth.
Prerequisites
- Strong experience in software engineering and agentic coding, and familiarity with genomic data (e.g. BAM/CRAM/VCF).
- Direct experience with unstructured clinical data, evaluations, and benchmarking is valuable but not essential.
7. Genetic analysis of plasma proteins and single-cell genomics to understand disease mechanisms.
Integration of plasma protein quantitative trait loci (pQTL) mapping with genome-wide association studies (GWAS) have emerged as popular tools to identify molecular aetiology of diseases. However, plasma pQTL are limited to proteins detectable in plasma and are unable to pinpoint the causal cell types. This project aims to bridge this gap by integrating plasma pQTL data from the UK Biobank (including ~3000 proteins derived from ~50,000 participants). with QTL mapping of molecular traits at cell type resolution from the TenK10K phase 1 project, including gene expression (eQTL), isoform expression (iso-eQTL), and chromatin accessibility (caQTL), derived from up to ~5 million cells from 2,000 donors. The student will perform statistical genetic analyses using pQTL and GWAS, and compare the results to those generated using other molecular data. There will be opportunities to develop new methods and integration strategies, as well as computational tools and bioinformatics pipelines to optimise existing methods for high-throughput data analysis.
What you will learn
The student will learn technical skills to integrate multi-modal biomedical data and perform statistical genetics analysis by working in an active translational research environment with expertise in both computational and molecular experiment methods.
Prerequisites
- Candidates should have a basic understanding of human biology and genetics.
- Proficiency in coding and data analysis using shell scripting, python, R, or Rust is preferable.
- Experience in statistical genetics, data engineering, and software development is highly desirable.
8. Developing evaluation toolchain for large language models and agentic AI systems in biology.
Large language models (LLM) and agentic artificial intelligence (AI) systems have been increasingly used to perform various tasks in biological research. Nonetheless, evaluating the accuracy and efficiency of these systems remains a challenge due to scarcity of ground truth data that reflects causal biology. This project aims to bridge this gap by developing a dataset and computational toolchain to perform a fair evaluation of multiple LLMs and agentic AI systems. First, the student will conduct a literature search to curate datasets that can serve as ground truth for evaluation. Second, the student will develop a computational toolchain to extract information from the curated datasets and integrate this into a platform for evaluating and benchmarking different LLMs and agentic AI systems. Third, the platform will be used to compare the performance of basic LLM, generic agentic AI system, and an in-house specialised AI system. Results from these projects are expected to provide insights into developing an efficient AI agent for biology, with a focus on single-cell genetics.
What you will learn
The student will learn about the use of LLMs and AI agents in biology and develop skills to conduct a fair evaluation of these systems grounded in fundamental causal biology.
Prerequisites
- Should have foundational skills and knowledge in coding and basic understanding of human biology and genetics.
- Familiarity with LLM and agentic AI frameworks (e.g., LangChain or LangGraph) and LLM APIs (e.g., OpenAI or Anthropic APIs) would be advantageous.
- Experience in AI evaluations and software development is highly desirable.
9. Reading the immune landscape from routine histology: multi-layer foundation models for breast cancer immunotyping.
Every breast cancer diagnosis begins with an H&E slide, yet the only immune information routinely extracted from it is a single lymphocyte score read from one section. We ask how much more is recoverable. Using pathology foundation models applied to whole-slide images from over 200 breast cancers with matched single-cell RNA sequencing and spatial transcriptomics from the same tissue, we will build a layered prediction framework: learned image features are used first to impute immune gene expression, then combined with predicted cellular composition and tissue architecture to assign an immunotype to every region of a slide. Because imaging and transcriptomics come from the same tissue, every layer is validated directly against ground truth. The project will establish which immune features morphology can and cannot reveal, whether explicit morphological features add information beyond learned representations, and, clinically most important, how much an immunotype call depends on where the tumour was sampled.
What you will learn
Practical skills in applying modern AI to biomedical imaging data — working with pretrained deep learning models, processing large image datasets, and evaluating model performance in a HPC environment. Experience integrating different types of biological data, and an understanding of how computational predictions are validated against real measurements — transferable skills for any data-driven role in cancer research
Prerequisites
- Working proficiency in Python (essential), exposure to PyTorch, image analysis, or single-cell/spatial transcriptomics is an advantage, command line and HPC/GPU skills will be developed during the project.
ABOUT GARVAN
Garvan Institute of Medical Research is an independent Medical Research Institute (MRI) in Sydney, delivering scientific and clinical impact on a global basis and in partnership with organisations that share our vision. We are proud to be one of Australia’s largest and most highly regarded MRI’s. Our vision is global leadership in discoveries to impact and our enduring purpose is to impact human health, by harnessing information encoded in our genome. We seek to see our world-class discovery research achieve life-changing impacts, not only for individual patients with rare diseases, but for the many thousands affected by complex, common disease. Garvan promotes a diverse workplace and is committed to the principles of equity, diversity, inclusion and belonging. We are always looking for culture ‘add’, not culture ‘fit’ and are building diverse teams with great sets of complementary styles and skills to help deliver our important work effectively.
HOW TO APPLY
To apply you must complete both Parts as below:
Part One:
Your application via the Garvan Careers Site/Workday should include:
- Copy of your CV/resume [no more than five (5) pages]
- Cover letter outlining which project(s) you are applying for [one page only]
- Copy of your academic transcript/s
[Note - Our system requires these documents to be compiled into one PDF document]
Part Two:In addition to submitting your application via Workday, please complete the Student Applicant Form at: https://forms.gle/zJATVNKTtraGYiN37
Note:
All applications must be submitted via the Garvan Careers site [Workday]. Applications from other sites/channels will not be considered.
Incomplete applications or applications without all supporting documents will not be assessed.
CLOSING DATE
The position will remain open until filled. We will be reviewing applications as they are received, and so we encourage you to submit your application as soon as possible. We aim to have positions filled by mid-October 2026 for project commencement in Mid-November 2026. All applicants will be notified of the outcome of their application by mid-October 2026.