Fast elastic peak detection for mass spectrometry data mining

Xin Zhang, Dennis E. Shasha, Yang Song, Jason T.L. Wang

Research output: Contribution to journalArticlepeer-review

4 Scopus citations

Abstract

We study a data mining problem concerning the elastic peak detection in 2D liquid chromatography-mass spectrometry (LC-MS) data. These data can be modeled as time series, in which the X-axis represents time points and the Y-axis represents intensity values. A peak occurs in a set of 2D LC-MS data when the sum of the intensity values in a sliding time window exceeds a user-determined threshold. The elastic peak detection problem is to locate all peaks across multiple window sizes of interest in the data set. We propose a new data structure, called a Shifted Aggregation Tree or AggTree for short, and use the data structure to find the different peaks. Our method, called PeakID, solves the elastic peak detection problem in 2D LC-MS data yielding neither false positives nor false negatives. The method works by first constructing an AggTree in a bottom-up manner from the given data set, and then searching the AggTree for the peaks in a top-down manner. We describe a state-space algorithm for finding the topology and structure of an efficient AggTree to be used by PeakID. Our experimental results demonstrate the superiority of the proposed method over other methods on both synthetic and real-world data.

Original languageEnglish (US)
Article number5645627
Pages (from-to)634-648
Number of pages15
JournalIEEE Transactions on Knowledge and Data Engineering
Volume24
Issue number4
DOIs
StatePublished - 2012
Externally publishedYes

All Science Journal Classification (ASJC) codes

  • Information Systems
  • Computer Science Applications
  • Computational Theory and Mathematics

Keywords

  • Knowledge discovery from LC-MS data
  • algorithms and data structures
  • bioinformatics
  • computational proteomics
  • time series data mining

Fingerprint

Dive into the research topics of 'Fast elastic peak detection for mass spectrometry data mining'. Together they form a unique fingerprint.

Cite this