Skip to content

Research

Our primary research goal is to unlock transformative applications in medicine and health by substantially improving the scalability and accessibility of biological data analysis (e.g., omics data). To achieve this goal, we design algorithms and co-design hardware and software for more accurate, high-performing, and energy-efficient omics analysis.

Real-Time and Portable Nanopore Sequencing

Nanopore sequencing is a commonly used technology to sequence biological molecules such as DNA, RNA, and proteins. Translating their initial raw data, electrical signals, into human-readable sequences of characters (e.g., DNA characters of A, C, T, G), a process called basecalling, is costly and ineffective. We design solutions that can directly analyze nanopore electrical signals without basecalling. These solutions help us build solutions that better utilize resource constrained devices (e.g., mobile devices or drones) to enable in-the-field and real-time biological data analysis. Additionally, we explore integrating our signal analysis solutions into standard genomics pipelines further to improve their accuracy and speed.

Key publications

RawHash: enabling fast and accurate real-time analysis of raw nanopore signals for large genomes

Can Firtina, Nika Mansouri Ghiasi, Joel Lindegger, Gagandeep Singh, Meryem Banu Cavlak, Haiyu Mao, Onur Mutlu

Proceedings of the 31st Annual Conference on Intelligent Systems for Molecular Biology and the 22nd European Conference on Computational Biology (ISMB/ECCB 2023), Lyon, France, July 2023.

[PDF][Code][DOI]

Conference talk[Video][Slides (pptx)][Slides (pdf)]

Preprint (bioRxiv)[Link][PDF]

Social media thread[Twitter (X)]

RawHash2: Mapping Raw Nanopore Signals Using Hash-Based Seeding and Adaptive Quantization

Can Firtina, Melina Soysal, Joël Lindegger, Onur Mutlu

Bioinformatics, July 2024.

[PDF][Code][DOI]

RawAlign: Accurate, Fast, and Scalable Raw Nanopore Signal Mapping via Combining Seeding and Alignment

Joël Lindegger, Can Firtina, Nika Mansouri Ghiasi, Mohammad Sadrosadati, Mohammed Alser, Onur Mutlu

IEEE Access, December 2024.

[PDF][Code][DOI]

Algorithms and Machine Learning for Genome Analysis

Analyzing genomic data is challenging as solutions must analyze very large volumes of data quickly and accurately. Such an analysis usually requires designing effective algorithmic and machine learning solutions in the genome analysis pipeline (e.g., read mapping, de novo genome assembly, error correction, metagenomics, and basecalling). We explore improving the accuracy and speed of analyzing genomic data to better generate insights from them.

Key publications

BLEND: a fast, memory-efficient and accurate mechanism to find fuzzy seed matches in genome analysis

Can Firtina, Jisung Park, Mohammed Alser, Jeremie S. Kim, Damla Senol Cali, Taha Shahroodi, Nika Mansouri Ghiasi, Gagandeep Singh, Konstantinos Kanellopoulos, Can Alkan, Onur Mutlu

NAR Genomics and Bioinformatics (NARGAB), March 2023.

[PDF][Code][DOI]

Apollo: a sequencing-technology-independent, scalable and accurate assembly polishing algorithm

Can Firtina, Jeremie S. Kim, Mohammed Alser, Damla Senol Cali, A Ercument Cicek, Can Alkan, Onur Mutlu

Bioinformatics, June 2020.

[PDF][Code][Link]

TargetCall: Eliminating the Wasted Computation in Basecalling via Pre-Basecalling Filtering

Meryem Banu Cavlak, Gagandeep Singh, Mohammed Alser, Can Firtina, Joel Lindegger, Mohammad Sadrosadati, Nika Mansouri Ghiasi, Can Alkan, Onur Mutlu

Frontiers in Genetics, September 2024.

[PDF][Code][DOI]

RUBICON: a framework for designing efficient deep learning-based genomic basecallers

Gagandeep Singh, Mohammed Alser, Kristof Denolf, Can Firtina, Alireza Khodamoradi, Meryem Banu Cavlak, Henk Corporaal, Onur Mutlu

Genome Biology, February 2024.

[PDF][Code][DOI]

Hardware Acceleration for Bioinformatics

To substantially improve speed and energy-efficiency of the computational approaches in genomics, we explore hardware acceleration of the underlying analysis. To this end, we explore designing solutions for GPUs, FPGAs, as well as emerging technologies such as processing in-/near-memory (i.e., data-centric computing), analog computing, and neuromorphic computing.

Key publications