TY - GEN
T1 - Efficient 3D protein structure alignment on large hadoop clusters in microsoft azure cloud
AU - Małysiak-Mrozek, Bożena
AU - Daniłowicz, Paweł
AU - Mrozek, Dariusz
N1 - Publisher Copyright:
© Springer Nature Switzerland AG 2018.
PY - 2018
Y1 - 2018
N2 - Exploration of 3D protein structures provides a broad potential for possible applications of its results in medical diagnostics, drug design, and treatment of patients. 3D protein structure similarity searching is one of the important exploration processes performed in structural bioinformatics. However, the process is time-consuming and requires increased computational resources when performed against large repositories. In this paper, we show that 3D protein structure similarity searching can be significantly accelerated by using modern processing techniques and computer architectures. Results of our experiments prove that by distributing computations on large Hadoop/HBase (HDInsight) clusters and scaling them out and up in the Microsoft Azure public cloud we can reduce the execution times of similarity search processes from hundred of hours to minutes. We will also show that the utilization of public clouds to perform scientific computations is very beneficial and can be successfully applied when scaling time-consuming computations over a mass of biological data.
AB - Exploration of 3D protein structures provides a broad potential for possible applications of its results in medical diagnostics, drug design, and treatment of patients. 3D protein structure similarity searching is one of the important exploration processes performed in structural bioinformatics. However, the process is time-consuming and requires increased computational resources when performed against large repositories. In this paper, we show that 3D protein structure similarity searching can be significantly accelerated by using modern processing techniques and computer architectures. Results of our experiments prove that by distributing computations on large Hadoop/HBase (HDInsight) clusters and scaling them out and up in the Microsoft Azure public cloud we can reduce the execution times of similarity search processes from hundred of hours to minutes. We will also show that the utilization of public clouds to perform scientific computations is very beneficial and can be successfully applied when scaling time-consuming computations over a mass of biological data.
KW - 3D protein structure
KW - Cloud computing
KW - Hadoop
KW - MapReduce
KW - Microsoft Azure
KW - Proteins
KW - Similarity searching
KW - Structural alignment
KW - Structural bioinformatics
UR - https://www.scopus.com/pages/publications/85053858905
U2 - 10.1007/978-3-319-99987-6_3
DO - 10.1007/978-3-319-99987-6_3
M3 - Conference contribution
AN - SCOPUS:85053858905
SN - 9783319999869
T3 - Communications in Computer and Information Science
SP - 33
EP - 46
BT - Beyond Databases, Architectures and Structures. Facing the Challenges of Data Proliferation and Growing Variety - 14th International Conference, BDAS 2018, Held at the 24th IFIP World Computer Congress, WCC 2018, Proceedings
A2 - Kozielski, Stanislaw
A2 - Mrozek, Dariusz
A2 - Kasprowski, Pawel
A2 - Malysiak-Mrozek, Bozena
A2 - Kostrzewa, Daniel
PB - Springer Verlag
T2 - 14th International Conference on Beyond Databases, Architectures and Structures, BDAS 2018 Held at the 24th IFIP World Computer Congress, WCC 2018
Y2 - 18 September 2018 through 20 September 2018
ER -