Scalable Cloud-Based Data Analysis Software Systems for Big Data from Next Generation Sequencing
Monika Szczerba , Marek Wiewiórka , Michał Okoniewski , Henryk Rybiński
AbstractNext generation sequencing (NGS) technology has become a serious computational challenge since its commercial introduction in 2008. Currently, thousands of machines worldwide produce daily billions of sequenced nucleotide base pairs of data. Due to continuous development of faster and economical sequencing technologies, processing the large amounts of data produced by high throughput sequencing technologies became the main challenge in bioinformatics. It can be solved by the new generation of software tools based on the paradigms and principles developed within the Hadoop ecosystem. This chapter presents the overall perspective for data analysis software for genomics and prospects for the emerging applications. To show genomic big data analysis in practice, a case study of the SparkSeq system that delivers tool for biological sequence analysis is presented.
* presented citation count is obtained through Internet information analysis and it is close to the number calculated by the Publish or Perish system.