Finnsh Society for Bionformatics Finnsh Society for Bionformatics
  • About us
  • Events
  • Jobs
  • Collaborators
  • Research Highlights
  • Bionformatics Webinars
  • Join the Society
  • Bioinformatics Day 2026

Research Highlights

Here you can find recipients of the Best PhD Thesis Award and research highlights from our community.

To suggest outstanding research for inclusion on this page, email us at fsfbioinf@gmail.com.

Best Bioinformatics Thesis 2024–2025

Heli Julkunen

Machine Learning for Precision Medicine

NoteAbstract

Precision medicine is an emerging approach to healthcare that tailors prevention and treatment strategies by accounting for individual patient variability. Its implementation is becoming more feasible due to advances in the scalability and cost-effectiveness of various molecular profiling technologies, such as genomics, transcriptomics, proteomics, and metabolomics. These advances have expanded not only the amount of molecular data measurable from individuals but also the availability of large-scale datasets for research, creating opportunities to discover more effective treatments, identify disease biomarkers, and develop models for predicting disease risk. However, the sheer volume and complexity of these data necessitate advanced computational methods to extract meaningful and actionable insights for precision medicine. This dissertation develops and applies computational frameworks to address various aspects of precision medicine, including predicting the effects of drug combination treatments, utilizing metabolomic biomarkers in disease risk assessment, and improving the methodological aspects of disease risk prediction. The first publication presents a machine learning framework designed to predict the effects of drug combinations across varying doses, providing an improvement over existing methods by enabling precise dose-specific predictions. This method achieved highly accurate predictions and identified novel drug combination synergies, which were subsequently experimentally validated. This framework provides an efficient tool for systematic pre-screening of drug combinations, particularly to advance cancer treatments. The second set of publications expands the current understanding of blood biomarkers in disease risk prediction through the analysis of population-scale metabolomic data. These studies identified novel metabolomic biomarkers and highlighted their potential in predicting the risks of various diseases, including diseases where metabolomics had not previously been studied at scale. The final publication proposes a machine learning method aimed at improving time-toevent disease risk prediction by incorporating comprehensive interaction effects among predictor variables. This method demonstrated improved accuracy in risk prediction compared to standard methods across multiple diseases and different data sources, thereby supporting the development of more accurate tools for risk assessment. Taken together, the novel methods and biological insights presented in this dissertation advance the translation of molecular data into prevention and treatment strategies in precision medicine.

Read the thesis

Best Bioinformatics Thesis 2023–2024

Tuomo Hartonen

Novel computational methods for studying the role and interactions of transcription factors in gene regulation

NoteAbstract

Regulation of which genes are expressed and when enables the existence of different cell types sharing the same genetic code in their DNA. Erroneously functioning gene regulation can lead to diseases such as cancer. Gene regulatory programs can malfunction in several ways. Often if a disease is caused by a defective protein, the cause is a mutation in the gene coding for the protein rendering the protein unable to perform its functions properly. However, protein-coding genes make up only about 1.5% of the human genome, and majority of all disease-associated mutations discovered reside outside protein-coding genes. The mechanisms of action of these non-coding disease-associated mutations are far more incompletely understood. Binding of transcription factors (TFs) to DNA controls the rate of transcribing genetic information from the coding DNA sequence to RNA. Binding affinities of TFs to DNA have been extensively measured in vitro, ligands by exponential enrichment) and Protein Binding Microarrays (PBMs), and the genome-wide binding locations and patterns of TFs have been mapped in dozens of cell types. Despite this, our understanding of how TF binding to regulatory regions of the genome, promoters and enhancers, leads to gene expression is not at the level where gene expression could be reliably predicted based on DNA sequence only. In this work, we develop and apply computational tools to analyze and model the effects of TF-DNA binding. We also develop new methods for interpreting and understanding deep learning-based models trained on biological sequence data. In biological applications, the ability to understand how machine learning models make predictions is as, or even more important as raw predictive performance. This has created a demand for approaches helping researchers extract biologically meaningful information from deep learning model predictions. We develop a novel computational method for determining TF binding sites genome-wide from recently developed high-resolution ChIP-exo and ChIP-nexus experiments. We demonstrate that our method performs similarly or better than previously published methods while making less assumptions about the data. We also describe an improved algorithm for calling allele-specific TF-DNA binding. We utilize deep learning methods to learn features predicting transcriptional activity of human promoters and enhancers. The deep learning models are trained on massively parallel reporter gene assay (MPRA) data from human genomic regulatory elements, designed regulatory elements and promoters and enhancers selected from totally random pool of synthetic input DNA. This unprecedentedly large set of measurements of human gene regulatory element activities, in total more than 100 times the size of the human genome, allowed us to train models that were able to predict genomic transcription start site positions more accurately than models trained on genomic promoters, and to correctly predict effects of disease-associated promoter variants. We also found that interactions between promoters and local classical enhancers are non-specific in nature. The MPRA data integrated with extensive epigenetic measurements supports existence of three different classes of enhancers: classical enhancers, closed chromatin enhancers and chromatin-dependent enhancers. We also show that TFs can be divided into four different, non-exclusive classes based on their activities: chromatin opening, enhancing, promoting and TSS determining TFs. Interpreting the deep learning models of human gene regulatory elements required application of several existing model interpretation tools as well as developing new approaches. Here, we describe two new methods for visualizing features and interactions learned by deep learning models. Firstly, we describe an algorithm for testing if a deep learning model has learned an existing binding motif of a TF. Secondly, we visualize mutual information between pairwise k-mer distributions in sample inputs selected according to predictions by a machine learning model. This method highlights pairwise, and positional dependencies learned by a machine learning model. We demonstrate the use of this model-agnostic approach with classification and regression models trained on DNA, RNA and amino acid sequences.

Read the thesis

We would love to hear from you! Send us an email: fsfbioinf@gmail.com

Copyright © 2026 The Finnish Society for Bioinformatics