Scalable Cloud-Based Data Analysis Software Systems for Big Data from Next Generation Sequencing

Monika Szczerba , Marek Wiewiórka , Michał Okoniewski , Henryk Rybiński

Abstract

Next generation sequencing (NGS) technology has become a serious computational challenge since its commercial introduction in 2008. Currently, thousands of machines worldwide produce daily billions of sequenced nucleotide base pairs of data. Due to continuous development of faster and economical sequencing technologies, processing the large amounts of data produced by high throughput sequencing technologies became the main challenge in bioinformatics. It can be solved by the new generation of software tools based on the paradigms and principles developed within the Hadoop ecosystem. This chapter presents the overall perspective for data analysis software for genomics and prospects for the emerging applications. To show genomic big data analysis in practice, a case study of the SparkSeq system that delivers tool for biological sequence analysis is presented.
Author Monika Szczerba II
Monika Szczerba,,
- The Institute of Computer Science
, Marek Wiewiórka II
Marek Wiewiórka,,
- The Institute of Computer Science
, Michał Okoniewski II - [Research Informatics, Scientific IT Services, ETH Zürich, Zurich, Switzerland]
Michał Okoniewski,,
- The Institute of Computer Science
- Research Informatics, Scientific IT Services, ETH Zürich, Zurich, Switzerland
, Henryk Rybiński II
Henryk Rybiński,,
- The Institute of Computer Science
Pages263-283
Publication size in sheets1
Book Japkowicz Nathalie, Stefanowski Jerzy (eds.): Big Data Analysis: New Algorithms for a New Society, Studies in Big Data, vol. 16, 2016, Springer International Publishing, ISBN 978-3-319-26987-0, [978-3-319-26989-4], 329 p., DOI:10.1007/978-3-319-26989-4
Keywords in EnglishGenomics – Big data – RNA – DNA – Next-generation sequencing – Biobanking
DOIDOI:10.1007/978-3-319-26989-4_11
URL http://link.springer.com/chapter/10.1007/978-3-319-26989-4_11
projectDevelopment of new algorithms in the areas of software and computer architecture, artificial intelligence and information systems and computer graphics . Project leader: Rybiński Henryk, , Phone: +48 22 234 7731, start date 18-05-2015, end date 30-11-2016, II/2015/DS/1, Completed
WEiTI Działalność statutowa
Languageen angielski
File
genomics.pdf 301.63 KB
Score (nominal)5
ScoreMinisterial score = 5.0, 27-03-2017, BookChapterNotSeriesMainLanguages
Citation count*3 (2018-07-06)
Cite
Share Share

Get link to the record
msginfo.png


* presented citation count is obtained through Internet information analysis and it is close to the number calculated by the Publish or Perish system.
Back