Exploiting hidden information interleaved in the redundancy of the genetic code without prior knowledge

Hadas Zur, Tamir Tuller

Research output: Contribution to journalArticlepeer-review

Abstract

Motivation: Dozens of studies in recent years have demonstrated that codon usage encodes various aspects related to all stages of gene expression regulation. When relevant high-quality large-scale gene expression data are available, it is possible to statistically infer and model these signals, enabling analysing and engineering gene expression. However, when these data are not available, it is impossible to infer and validate such models. Results: In this current study, we suggest Chimera-an unsupervised computationally efficient approach for exploiting hidden high-dimensional information related to the way gene expression is encoded in the open reading frame (ORF), based solely on the genome of the analysed organism. One version of the approach, named Chimera Average Repetitive Substring (ChimeraARS), estimates the adaptability of an ORF to the intracellular gene expression machinery of a genome (host), by computing its tendency to include long substrings that appear in its coding sequences; the second version, named ChimeraMap, engineers the codons of a protein such that it will include long substrings of codons that appear in the host coding sequences, improving its adaptation to a new hostâ (tm)s gene expression machinery. We demonstrate the applicability of the new approach for analysing and engineering heterologous genes and for analysing endogenous genes. Specifically, focusing on Escherichia coli, we show that it can exploit information that cannot be detected by conventional approaches (e.g. the CAI-Codon Adaptation Index), which only consider single codon distributions; for example, we report correlations of up to 0.67 for the ChimeraARS measure with heterologous gene expression, when the CAI yielded no correlation. Availability and implementation: For non-commercial purposes, the code of the Chimera approach can be downloaded from http://www.cs.tau.ac.il/tamirtul/Chimera/download.htm.

Original languageEnglish
Pages (from-to)1161-1168
Number of pages8
JournalBioinformatics
Volume31
Issue number8
DOIs
StatePublished - 15 Apr 2015

All Science Journal Classification (ASJC) codes

  • Statistics and Probability
  • Biochemistry
  • Molecular Biology
  • Computer Science Applications
  • Computational Theory and Mathematics
  • Computational Mathematics

Fingerprint

Dive into the research topics of 'Exploiting hidden information interleaved in the redundancy of the genetic code without prior knowledge'. Together they form a unique fingerprint.

Cite this