Our primary research goal is to unlock transformative applications in medicine and health by substantially improving the scalability and accessibility of biological data analysis (e.g., omics data). To achieve this goal, we design algorithms and co-design hardware and software for more accurate, high-performing, and energy-efficient omics analysis.
Real-Time and Portable Nanopore Sequencing
Nanopore sequencing is a commonly used technology to sequence biological molecules such as DNA,
RNA, and proteins. Translating their initial raw data, electrical signals, into human-readable
sequences of characters (e.g., DNA characters of A, C, T, G), a process called basecalling, is
costly and ineffective. We design solutions that can directly analyze nanopore electrical signals
without basecalling. These solutions help us build solutions that better utilize resource
constrained devices (e.g., mobile devices or drones) to enable in-the-field and real-time
biological data analysis. Additionally, we explore integrating our signal analysis solutions into
standard genomics pipelines further to improve their accuracy and speed.
Key publications
RawHash: enabling fast and accurate real-time analysis of raw nanopore signals for large genomes
Can Firtina, Nika Mansouri Ghiasi, Joel Lindegger, Gagandeep Singh, Meryem Banu Cavlak, Haiyu Mao, Onur Mutlu
Algorithms and Machine Learning for Genome Analysis
Analyzing genomic data is challenging as solutions must analyze very large volumes of data quickly
and accurately. Such an analysis usually requires designing effective algorithmic and machine
learning solutions in the genome analysis pipeline (e.g., read mapping, de novo genome assembly,
error correction, metagenomics, and basecalling). We explore improving the accuracy and speed of
analyzing genomic data to better generate insights from them.
Key publications
BLEND: a fast, memory-efficient and accurate mechanism to find fuzzy seed matches in genome analysis
Can Firtina, Jisung Park, Mohammed Alser, Jeremie S. Kim, Damla Senol Cali, Taha Shahroodi, Nika Mansouri Ghiasi, Gagandeep Singh, Konstantinos Kanellopoulos, Can Alkan, Onur Mutlu
To substantially improve speed and energy-efficiency of the computational approaches in genomics,
we explore hardware acceleration of the underlying analysis. To this end, we explore designing
solutions for GPUs, FPGAs, as well as emerging technologies such as processing in-/near-memory
(i.e., data-centric computing), analog computing, and neuromorphic computing.
Key publications
ApHMM: Accelerating Profile Hidden Markov Models for Fast and Energy-efficient Genome Analysis
Can Firtina, Kamlesh Pillai, Gurpreet S. Kalsi, Bharathwaj Suresh, Damla Senol Cali, Jeremie S. Kim, Taha Shahroodi, Meryem Banu Cavlak, Joël Lindegger, Mohammed Alser, Juan Gómez Luna, Sreenivas Subramoney, Onur Mutlu