Evaluating Rugby Team Performance with Dynamic Graph Analysis
Total Page:16
File Type:pdf, Size:1020Kb
The Haka Network: Evaluating Rugby Team Performance with Dynamic Graph Analysis Paolo Cintia1, Michele Coscia2 and Luca Pappalardo1 1 KDDLab - ISTI CNR, Via G. Moruzzi 1 - 56124 Pisa, Italy Email: [email protected] 2 Naxys - University of Namur, Rempart de la Vierge 8, 5000 Namur Belgium Email: [email protected] Abstract—Real world events are intrinsically dynamic and 1 5 analytic techniques have to take into account this dynamism. This 1 4 aspect is particularly important on complex network analysis 1 3 when relations are channels for interaction events between 1 2 actors. Sensing technologies open the possibility of doing so 1 1 1 0 for sport networks, enabling the analysis of team performance 9 in a standard environment and rules. Useful applications are directly related for improving playing quality, but can also shed 8 light on all forms of team efforts that are relevant for work 6 7 teams, large firms with coordination and collaboration issues 4 5 and, as a consequence, economic development. In this paper, we 1 2 3 consider dynamics over networks representing the interaction between rugby players during a match. We build a pass network 1 2 3 and we introduce the concept of disruption network, building a multilayer structure. We perform both a global and a micro- 4 5 level analysis on game sequences. When deploying our dynamic 6 8 7 graph analysis framework on data from 18 rugby matches, we discover that structural features that make networks resilient to 9 disruptions are a good predictor of a team’s performance, both at 1 0 the global and at the local level. Using our features, we are able 1 2 to predict the outcome of the match with a precision comparable 1 3 to state of the art bookmaking. 1 1 1 4 1 5 I. INTRODUCTION Mining dynamics on graphs is a challenging, complex Fig. 1: The pass and tackles networks imposed on the same and useful problem [4]. Many networks are representation structure for one match in our dataset: Italy (blue nodes, home of evolving phenomena, thus understanding graph dynamics team) versus New Zealand (black nodes, away team). The brings us a step closer for an accurate description of reality. number of the node refers to the role of the player. Green Sensing technologies open the possibility of performing dy- edges are the pass network and red edges are the tackle namic graph analysis in an ever expanding set of contexts. network. Node layout is determined by the classical rugby One of them is competition events. Sports analysis is par- player positioning on the pitch during a scrum. ticularly interesting because it involves a setting where both environment and rules are standardized, thus providing us an paper. If we are able to shed light on how group dynamics objective measure of players’ contributions. affect sport teams, we can try to universalize collaboration best These possibilities power a number of potentially useful practices to enhance productivity in many different scenarios. applications. The direct one is related to the improvement It comes as no surprise that analyzing sport events is a fast of a team’s performances. Once the most important factors growing literature in data science. Previous works involve the contributing or preventing victories are identified, the team analysis of different sports: from soccer [14] [9] to basketball can work on strategies to regulate their collective effort to- [27], from American football [20] to cycling [8]. In this paper ward the best practices. However, there are non-sport related we focus on a sport that was not analyzed before: rugby. externalities too. A sport team is nothing more than a group of The reason is that rugby has some distinctive features that people with different skills that is trying to achieve a goal in makes it an ideal sport to consider. Like American football, the most efficient way possible. This description can be applied it is a sport where there is a very clear and simple success to any other form of team [15]: a start-up enterprise, a large measure: the number of meters gained in territory. Again like manufacturing firm, a group of scientists writing a scientific American football, it contains a disruption network: tackles. IEEE/ACM ASONAM 2016, August 18-21, 2016, San Francisco, CA, USA Unlike American football, performance is also related to how 978-1-5090-2846-7/16/$31.00 c 2016 IEEE well a team can weave its own pass structure, involving all the players in the field in the construction of uninterrupted Barabasi´ develop a predictive model that relies on a tennis sequences that can lead to a score. player’s performance in tournaments to predict her popularity One of the main contributions of the paper is to use both [28]. In NBA basketball league the performance efficiency the pass and disruption networks at the same time, creating a rating introduced by Hollinger [12] is a stable and widely multilayer network analysis [13] [3] of dynamics on graphs. used measure to assess players’ performance by combining Figure 1 depicts an example of such structure. To the best the manifold type of data gathered during every game (pass of our knowledge, there has not been such attempt elsewhere completed, shots achieved, etc.). Vaz de Melo et al. [27] in the literature. We also do a multiscale analysis: we first introduced network analysis to the mix. In baseball, Rosales focus on the multilayer network as a whole – as a result of and Spratt propose a new methodology to quantify the credit all interactions during a match like in [9] – and we perform for whether a pitch is called a ball or strike among the catcher, also a micro-level analysis on each match sequence. the pitcher, the batter, and the umpire involved [21]. Smith et Our data comes from the 2012 Tri-Nations championship, al. propose a Bayesian classifier to predict baseball awards 2012 New Zealand Europe tour, 2012 Irish tour to New in the US Major League Baseball, reaching an accuracy of Zealand and 2011 Churchill Cup (only USA matches), for 80% in the predictions [25]. In soccer, networks are a widely a total of 18 matches. The data was collected by Opta1, using used tool to determine the interactions between players on field positioning and semi-automatically annotated events. the field, where soccer players are nodes of a network and We find that there are some features of the pass networks a pass between two players represents a link between the that make teams more successful in their quest for territorial respective nodes. For example, Cintia et al. [9], [7] exploit gain. In particular, the connectivity of the network seems to passing networks to detect the winner of a game based on the play an important role. Being able avoid structurally crucial passing behavior of the teams. [14] shared the aim, without players, that would make part of the team isolated if they were using networks. They discovered that, while the strategy of removed from the network by a tackle, is associated with the the majority of successful teams is based on maximizing highest amounts of meters made. This is also confirmed if ball possession, another successful strategy is to maximize a we do not analyze only the global network as the result of defense/attack efficiency score. Dynamic graph analysis has the entire match, but also mining the patterns of each single also been applied for ranking purposes [24], [20], [18]. Other sequence of the game. In the latter case, the signal is harder examples of wider applicability of team-based success research to disentangle from noise, and most analyzed features did not comes from analysis of citation [22] and social [15] networks. yield any significant result. This highlights how much rugby Methodologically, our paper is indebted with the vast field is a game dependent on a grand match strategy, rather than on of dynamics on networks analysis [4]. This field is usually ap- just a sequence-by-sequence short term tactical one. proached both from a statistical [26] and a mining perspective When applied to a real world prediction task, our framework [5]. We adopt the latter approach. In mining network dynam- fares well in comparison with state of the art attempts. The ics, the aim is to find regularities in the evolution of networks power in predicting the winner of a match is comparable with [2]. There are many applications for these techniques, beyond the one of bookmakers, who have access to the full history sports analytics: epidemiology [17], mobility [10] and genetics of teams and of the players actually performing on the pitch. [16]. In this work, we focus on a narrower part of this field, In particular, our framework is less susceptible to reputation since our networks are multilayer: nodes can be connected bias: the algorithm is not afraid to design New Zealand as a with edges belonging to multiple types. A good survey on loser in what bookmakers saw as a great upset result – its loss multilayer networks, both modeling and analysis, can be found to England in the December 1st, 2012 Twickenham clash. in [13]. The specific multilayer model we adopt is the one of multidimensional networks firstly presented in [3]. II. RELATED WORK In the last decade, sports analytics has increased its per- III. DATA vasiveness as large-scale performance data became available The data has been collected by Opta and made available in [23]. Researchers from different disciplines started to analyze 2013 as part of the AIG Rugby Innovation Challenge2. Opta massive datasets of players’ and teams’ performance collected performs a semi-automatic data collection that happens as the from monitoring devices.