Structural Alignment based Kernels for Protein Structure Classification Sourangshu Bhattacharya
[email protected] Chiranjib Bhattacharyya
[email protected] Dept. of Computer Science and Automation, Indian Institute of Science, Bangalore - 560 012, India. Nagasuma Chandra
[email protected] Bioinformatics Center, Indian Institute of Science, Bangalore - 560012, India. Abstract acid types. This view is taken by most structure align- Structural alignments are the most widely ment algorithms (Eidhammer et al., 2000), which are used tools for comparing proteins with low the most widely accepted methods for protein struc- sequence similarity. The main contribution ture comparison. of this paper is to derive various kernels on In this paper, we explore the idea of designing clas- proteins from structural alignments, which sifiers based on Support vector machines(SVMs) for do not use sequence information. Central proteins having low sequence similarity. More explic- to the kernels is a novel alignment algorithm itly our goal is to derive positive semi-definite ker- which matches substructures of fixed size us- nels motivated from structural alignments without us- ing spectral graph matching techniques. We ing sequence information. The hypothesis is that any derive positive semi-definite kernels which structure classification technique should capture the capture the notion of similarity between sub- notion of similarity defined by structural alignment structures. Using these as base more sophis- algorithms. Though there is lot of work on design- ticated kernels on protein structures are pro- ing kernels on protein sequences (Jaakkola et al., 1999; posed. To empirically evaluate the kernels we J.-P. Vert, 2004; Leslie & Kwang, 2004), to the best used a 40% sequence non-redundant struc- of our knowledge, there is relatively less work on ker- tures from 15 different SCOP superfamilies.