Skip to content
Yuki Moriya edited this page Dec 14, 2018 · 22 revisions

Category: Omics

Lead: Susumu Goto

Members: Susumu Goto, Yuki Moriya, Shin Kawano, Akiyasu C. Yoshizawa, Tsuyoshi Tabata, Yu Watanabe, Tomoyo Takami, Satoshi Tanaka, Masaki Murase

Analysis protocol

Human reference protein (peptide) database construction

  • Swiss-Prot database with all variations (1). -> Replace with UniProt?
    • Swiss-Prot-all (20,410 canonical proteins and 22,016 isoforms) may not be enough.
    • Consideration of variation will be necessary to cover possible variants.
  • Human lung cancer cell line data (KERO): Create protein sequence data from KERO vcf and Ensembl (2).
    • Currently by AnnoVar -> probably better be replaced by VEP (in correspondence with TogoVar)
  • Japanese and ExAC variation data (TogoVar): Create protein sequence data from TogoVar vcf and Ensembl by VEP (3).
    • more than 4,577,000 genomic variants related to protein altering variants, in togoVar.
  • Merge 1 & 2 & 3.
  • Create peptide sequence data using several cleavage enzymes: Trypsin, Trypsin/P, Lys-C, ...
    • Trypsin: [KR]|[^P], ([KR]|[KRIFL] sometimes not cleaved, but we do not consider this)
    • Trypsin/P: [KR]|*
    • Lys-C: [K]|*
    • Arg-N: *|[R]

Yeast and other species

  • Create a reference protein database sample by sample using genome sequence.
  • UniProt in case of yeast?

Data structure:

Viewer

Links to KERO / ChIP-Atlas?

Clone this wiki locally