Arpeggio: Harmonic compression of ChIP-seq data reveals protein-chromatin interaction signatures

Kelly Patrick Stanton, Fabio Parisi, Francesco Strino, Neta Rabin, Patrik Asp, Yuval Kluger

Research output: Contribution to journalArticlepeer-review

Abstract

Researchers generating new genome-wide data in an exploratory sequencing study can gain biological insights by comparing their data with well-annotated data sets possessing similar genomic patterns. Data compression techniques are needed for efficient comparisons of a new genomic experiment with large repositories of publicly available profiles. Furthermore, data representations that allow comparisons of genomic signals from different platforms and across species enhance our ability to leverage these large repositories. Here, we present a signal processing approach that characterizes protein-chromatin interaction patterns at length scales of several kilobases. This allows us to efficiently compare numerous chromatin-immunoprecipitation sequencing (ChIP-seq) data sets consisting of many types of DNA-binding proteins collected from a variety of cells, conditions and organisms. Importantly, these interaction patterns broadly reflect the biological properties of the binding events. To generate these profiles, termed Arpeggio profiles, we applied harmonic deconvolution techniques to the autocorrelation profiles of the ChIP-seq signals. We used 806 publicly available ChIP-seq experiments and showed that Arpeggio profiles with similar spectral densities shared biological properties. Arpeggio profiles of ChIP-seq data sets revealed characteristics that are not easily detected by standard peak finders. They also allowed us to relate sequencing data sets from different genomes, experimental platforms and protocols. Arpeggio is freely available at http://sourceforge.net/p/arpeggio/ wiki/Home/.

Original languageEnglish
Pages (from-to)e161
JournalNucleic acids research
Volume41
Issue number16
DOIs
StatePublished - Sep 2013
Externally publishedYes

All Science Journal Classification (ASJC) codes

  • Genetics

Fingerprint

Dive into the research topics of 'Arpeggio: Harmonic compression of ChIP-seq data reveals protein-chromatin interaction signatures'. Together they form a unique fingerprint.

Cite this