Non-Cellulosomal Cohesin- and Dockerin-Like Modules in the Three Domains of Life Ayelet Peera, Steven P
Total Page:16
File Type:pdf, Size:1020Kb
1 Non-cellulosomal Cohesin- and Dockerin-like Modules in the Three Domains of Life Ayelet Peera, Steven P. Smithb, Edward A. Bayerc,*, Raphael Lameda and Ilya Borovoka aDepartment of Molecular Microbiology and Biotechnology, Tel Aviv University, Ramat Aviv 69978 Israel bDepartment of Biochemistry, Queen’s University Kingston Ontario Canada K7L 3N6 cDepartment of Biological Sciences, Weizmann Institute of Science, Rehovot 76100 Israel *Corresponding author: Edward A. Bayer Tel: (+972) -8-934-2373 Fax: (+972)-8-946-8256 Email: [email protected] Supplementary Table S1. Compendium of cohesins and dockerins in the three domains of life. In order to discover new putative cohesin/dockerin-containing proteins, we used sequences of all the classical cohesin and dockerin modules from C. thermocellum, C. cellulovorans, C. cellulolyticum, B. cellulosolvens and Acetivibrio cellulolyticus as well as cohesins and dockerins recently discovered in rumen bacteria, Ruminococcus albus and R. flavefaciens as BlastP queries for the main NCBI Blast server against all non- redundant protein sequences deposited in GenBank/EMBL/DDBJ databases. We also performed extensive searches using the TblastN algorithm through all publicly available microbial genome databases including those attached to the NCBI BLAST server for bacterial genomes (http://www.ncbi.nlm.nih.gov/sutils/genom_table.cgi?), as well as several additional microbial genome databases – Microbial Genomics at the DOE Joint Genome Institute (http://genome.jgi-psf.org/mic_home.html), the Rumenomics database at TIGR/JCVI (http://tigrblast.tigr.org/rumenomics/index.cgi) and Bacterial Genomes at the Sanger Centre (http://www.sanger.ac.uk/Projects/Microbes/). Once a putative cohesin or dockerin-encoding gene product was identified, gene-walking techniques were employed to analyze and locate possible cellulosome-like gene clusters. The date of the final search of the database, used for construction of this table, was September 23, 2008. 2 ARCHAEA Euryarchaeota (No cohesins or dockerins are currently predicted in Crenarchaeota) Gene Protein Putative Putative Annotation and comments1 size (no. cohesin dockerin residues) Archaeoglobus fulgidus DSM 4304 / ATCC 49558 (Archaeoglobi; Archaeoglobales) AF2375 506 279..418 436..488 No obvious signal peptide (SP). TMD (69..91) + Cohesin + (operon) Dockerin Cohesin AF2375 TTIIAGSAEAPQGSDIQVPVKIENADKVGSINLILSYPNVLEVEDV LQGSLTQNSLFDYNVEGNQIKVGIADSNGISGDGSLFYVKFRVTGN EKAEQAENVKGKLRGLGQQLSEITLRNSHALTLQGIEIYDIDGNSV KV Dockerin AF23751 GDVNGDGEINSLDALLALQMSIG KVEPNPV ADMDGDGKVLAKDATEIMKMATD AF2376 190 33..156 - SP, Hypothetical protein, +Cohesin, a 'strong' TMD (167..185) (operon) Cohesin AF2376 MVVKIPDTSGAVGGTVEVPVEVENAQNLGSMDIVIVYDPTILKVQN VGKGELNKGLLSSNTGEGMVAISLADSKGINGKGSVAVITFQVLKA GSTDLTIQSVKAYDVNTHVDIPVKADNGKFEA 1In dockerin sequences, calcium-coordinating amino acid residues in positions 1, 3, 5, 9 and 12 are highlighted (those fitting the consensus are in blue, atypical residues – in gray; an asterisks follow the C-terminal residue in some proteins. 3 Haloarcula marismortui / Halobacterium marismortui ATCC 43049 (Halobacteria; Halobacteriales) phoD or 741 - 688..741 Alkaline phosphatase D, EC 3.1.3.1; a putative dockerin is similar rrnAC0273 to that in Mhun_3046, Methanospirillum hungatei, (chromosome I) Putative dockerin: EDVNGNGRVDYDDIQLLFDNFDD DSVALNKTA YDFNENGKLEYDDIVTLYSEVN* rrnAC0766 892 - 839..892 Putative cell surface protein, (chromosome I) Putative dockerin: EDLNGNGRLDYEDVQVLFSNMDS DSVQLNTGA YDFNENGKLDFADVTALYEEVN* rrnAC2380; 615 - 559-613 putative serine/threonine phosphatase (PF00149; Metallophos) operon? Putative dockerin: (chromosome I) GDADGDGDVDTDDVDRIQRSVAG EDSDIDRDA ADIDGDGDVDIGDAVSARNLAGD rrnAC2381; 442 42-175 - Hypothetical protein; Cohesin domain predicted in both, Pfam operon? (PF00963) and InterPro (IPR002102). (chromosome I) Putative cohesin: PEASGDALSVAPGETVTVQVWANATAVRGYQTNVTFDPTIAEIDSV AGSDDFDSPVTNVNNEAGWVAFNQLRSRETSDPVLAEITLTVPADA AGSTQLAFIEADTKLSNSDGETVAPAAFNSIELRVNNGSTTP rrnB0167 1562 - 126..180 Cbp, calcium-binding protein-like: approx. 8 TSP_3 like domains; (chromosome II) PF00404; dockerin_1 EDVNGDGAVNIVDVDALSRHLES TAAGANWSA YDYTGDNRTDVGDIQWLFVATRS Halobacterium salinarium NRC-1 / Halobacterium halobium NRC-1 (Halobacteria; Halobacteriales) VNG1951G 585 - 532..585 Subtilisin homolog or Sub; PF00404; Dockerin_1 EDVNGDGVVDMADVELLYQLALR NAVPESTTA FDVTGDGKFDMHDVQALTRQVE* 4 Haloquadratum walsbyi DSM 16790 (Halobacteria; Halobacteriales) HQ1205A, 1626 1285-1420 1571-1626 Probable cell surface adhesin, operon? Putative cohesin: VEGVELQSGNTTTLATNFTLREDTTADLTATVEFSNETAIENTDVT FSGSGGVPSSNDASLPVTQTTNSSGIVQFEIPASDTNDAYTITADS ANFSNASTGVSPAPGDSVSRDFSLEEAQAQEPSVSVNFTDQTV Putative dockerin: EDVNGDNATNFDDVVALAFVLTE KDDLTQQQRDA FDINDDGVVNFDDVVRLAFPGQ* HQ1208A, 619 - 558-619 Subtilisin-like serine protease Sub, operon? Putative dockerin: EDVNGDNGYTFDDVIALAFADTE RLTNQQRDA LDFDHDGNIDFDDVVELAFER* HQ1210A, 125 71-125 Hypothetical protein, operon? Putative dockerin: EDVNGDAESTFDDVIALAFVISQ TGQLTEQQRDA LNIDGEDDVDFDNVIALTFNI* HQ3467A, 2079 - 1952-2006; Probable cell surface glycoprotein Hmu3 operon? 2027-2079 Putative dockerin: EDINGDGEFNFDDVVALAFNTGT DPIQQNSDL FDINNDGVVNFDDVVALAFESPS Putative dockerin: EDINGDGELNFDDVVALAFNIES DAVQQNTDQ FDFDNDGNVDFDDVVELAFET* HQ3469A, 506 - 443-506 Probable cell surface glycoprotein, operon? Putative dockerin: EDINGDGKANFDDVVTFAFKITT NAVQENADL FDFNDDNDVNFDDVIELAFTLPT Bolhuis et al., 2006. 5 Methanococcoides burtonii DSM 6242 (Methanomicrobia; Methanosarcinales) Mbur_0125 418 35-167 - Hypothetical protein precursor; Cohesin domain, Pfam (PF00963) and InterPro (IPR002102): VSISPPDQLISIGSNVTVIIYVDPDVPVSGAQFDLDFDGDILSVIS ISEGDIFTNGVSTLFSGGTIDNLEGNIMNVYGALFGKNEVIKSGTF ATIVFETKSSGSSSLQLSNVVVSNSIGAAVPIAVINGIVSINGD Mbur_0728*, 1126 846-979 1072-1123 Putative Ig precursor, operon with Cohesin, Pfam (PF00963) and InterPro (IPR002102): Mbur_0729? VVMSSSAQIVSPGQPFSIGISIDPSEPVSGAQLDLLFDEQLVSAEG TIEGDLFNLNGASTLFNSGTIGDSGVITDIYCSIIGAAPVSSKGTM * see its ortholog, ATIAFTAGNTAGVAEFSLSDVVISNARSEGTPYTVTDTMVVI MM_2098 Putative dockerin: WDINRDCAVNILDITLVSQKIGT DGAGTD QDVNQDGVVNIQDLTLVAQHFGE Mbur_0729, 226 53-178 - SP, annotated as "Cellulosome anchoring protein, precursor"; operon with Cohesin, Pfam (PF00963) and InterPro (IPR002102): Mbur_0728? FTVSSGEVLNVAPGETFNVDVFVSPDVDIAGMQFDLLFDSSKFQID SVTEGNLFKQSGMGTFFVAASVTPGLLDNTYGCILGNAAVLTPANF ATITITANEHASGRSAFILKDVIVSDPKGNAVNV Methanococcus aeolicus Nankai-3 / ATCC BAA-1280 (Methanococci; Methanococcales), isolated from deep marine sediment from the Nankai Trough off the coast of Japan Maeo_0111 421 - 24-73 Periplasmic binding protein precursor, Putative N-terminal dockerin: GDVNEDGSINIIDVVYLFKNRNL DYEV GDINCDNSINIIDVVYLFKHYDQ Maeo_0115 703 25-155 SP, PKD domain containing protein precursor, Pfam PF00963; InterPro IPR002102; Cohesin: NITVELTPNNTNTIVGDTINLDLFVKNIPNDTKCSGFETTLSYDSN LLNLTDIKLSDISNNADLKEADLSTGSIKIAWFSNEPSGNITIASL SFKTLNEATSPPATISLSSKTAVSNINGVGYNNITIINSNITIN 6 Maeo_0117 182 22-150 SP, hypothetical protein, Pfam PF00963; InterPro IPR002102; Cohesin: ASNVEVNCPSEVNVGDTFDIVITIKNNESLSGFECQVNVPDNLKII NVLENDDIRKKASNYYKNDFSDRNGIISFMLINEPMQSDFYLATIK VKVLKDSSNTTIHITPKAADKDANEIKMNNINLKLNV Maeo_0118 454 - 36-85 SP, ABC-type Fe3+-hydroxamate transport system periplasmic component-like protein precursor, Putative N-terminal dockerin: GDVNADGSINIIDVVYLFKNRNL DYED GDINADNSINIIDVVYLFKHYDQ Maeo_0119 976 23-160 - SP, APHP domain protein, Pfam PF07705; CARDB (x2); PF04473; DUF553; Pfam PF00963; InterPro IPR002102; Cohesin: NVSVELTPSNVNAIVGDTISLDLVVKNIPNDTKCSGFDTTLSYDSN LLNLTDIKLSDISNNANLKRANLSTGFISLSWSSDEPSGNITIASL SFKALNEATSPTTISLTQTIVSDIEPKGYKITLNNSNIIITKPKAD Methanococcus vannielii SB / ATCC 35089 / DSM 1224 (+ two other strains) (Methanococci; Methanococcales) Mevan_0131 411 - 14-63 Periplasmic binding protein, precursor Putative N-terminal dockerin: GDVNGDGNINIADVVYLFKNRNV PIDV GDLNCDGNVNVADVVYLFRNYDK Mevan_0132 695 22-145 SP, Putative uncharacterized protein, PKD domain, Pfam PF00963; InterPro IPR002102; Cohesin: SVVLNSSSDSFFIGESFDLELNLLNIPSDKKCGGFETRISYDKTQL QLSSITLSKDLNSDIGDVSLSNGRISIVWFSTPPTGDINIAKIKFN VLKEGNSTIELLGTTVSDSNGFTYNNMTLSNLKVKLSPLGYTVT Mevan_0135 754 18-145 SP, Putative uncharacterized protein, Pfam PF00963; InterPro IPR002102; Cohesin: IELSPNVINIQKGNNFTVDVLLRDIPEIAGYDSKLNYDKTILKLNS ISLSTDVNSASLKDTNLNNDGSVSLLWMGNSVTGNLTVATFTFEAI NEGDTKISLKNAALSDIEGQGINNFDTLDSEVLVFE 7 Methanococcus voltae A3 (Methanococci; Methanococcales) MvolDRAFT_0511, 453 - 26-82 Putative iron transport system substrate-binding protein, operon? precursor, Putative dockerin: GNYDYRLGDVNCDGKISVSDVVY LFNNRNMDIED GDVNADNKISVSDVVYLFNYYDK MvolDRAFT_0512, 1103 36-170 - SP, APHP domain protein precursor, operon? Putative cohesin: TNDNIVVSLAPQTQQVNNGDQFVLELWVNNSNITISAFESNFVYDS SNVNLTAIQLSGIANTAGLKTVDVSDNSISLAWMGGGITESNVNIA NLTFETFNVSSNEIRLRNLDITNQGGISTMPTDEHILDASVEV MvolDRAFT_0578 789 49-186 - Hypothetical protein. Pfam PF00963; Cohesin; InterPro IPR002102; Cohesin: NYSMTLSPTVKSAGVGDSFELTLWGNTPKVSAIQSKFNYDTAMLEL TDIELGTIADTASSKRAEVSSSLITMLWFDPSTAPEGYYKIATLSF KVLKGGNTSIVPRNYEFSDNTGVSITPVINKANISVSEGMKAELII MvolDRAFT_0734 451 - 25-81 Putative iron transport system substrate-binding protein, precursor, Putative dockerin: GNYDYRLGDINEDGKISVSDVVY LFNNRDLDIEDADVN ADVNADNKISVSDVVYLFNYYDE Methanoculleus marisnigri JR1 / ATCC 35101 / DSM 1498 (Methanomicrobia; Methanomicrobiales) Memar_0673 857 - 800-852 2 CARDB (PF07705), Aminotrans_II_pyridoxalP_BS (IPR001917), Putative dockerin: GDLNGDGNVDWADVTIAAGMAQK TAPSDPA ADVNGDGTVDWKDVALLADFFFG