Discovery with Data: Leveraging Statistics with Computer Science to Transform Science and Society
Total Page:16
File Type:pdf, Size:1020Kb
Load more
Recommended publications
-
Stable Treemaps Via Local Moves
Stable treemaps via local moves Citation for published version (APA): Sondag, M., Speckmann, B., & Verbeek, K. A. B. (2018). Stable treemaps via local moves. IEEE Transactions on Visualization and Computer Graphics, 24(1), 729-738. [8019841]. https://doi.org/10.1109/TVCG.2017.2745140 DOI: 10.1109/TVCG.2017.2745140 Document status and date: Published: 01/01/2018 Document Version: Accepted manuscript including changes made at the peer-review stage Please check the document version of this publication: • A submitted manuscript is the version of the article upon submission and before peer-review. There can be important differences between the submitted version and the official published version of record. People interested in the research are advised to contact the author for the final version of the publication, or visit the DOI to the publisher's website. • The final author version and the galley proof are versions of the publication after peer review. • The final published version features the final layout of the paper including the volume, issue and page numbers. Link to publication General rights Copyright and moral rights for the publications made accessible in the public portal are retained by the authors and/or other copyright owners and it is a condition of accessing publications that users recognise and abide by the legal requirements associated with these rights. • Users may download and print one copy of any publication from the public portal for the purpose of private study or research. • You may not further distribute the material or use it for any profit-making activity or commercial gain • You may freely distribute the URL identifying the publication in the public portal. -
Arxiv:1705.10225V8 [Stat.ML] 6 Feb 2020
Bayesian stochastic blockmodelinga Tiago P. Peixotoy Department of Mathematical Sciences and Centre for Networks and Collective Behaviour, University of Bath, United Kingdom and ISI Foundation, Turin, Italy This chapter provides a self-contained introduction to the use of Bayesian inference to ex- tract large-scale modular structures from network data, based on the stochastic blockmodel (SBM), as well as its degree-corrected and overlapping generalizations. We focus on non- parametric formulations that allow their inference in a manner that prevents overfitting, and enables model selection. We discuss aspects of the choice of priors, in particular how to avoid underfitting via increased Bayesian hierarchies, and we contrast the task of sampling network partitions from the posterior distribution with finding the single point estimate that maximizes it, while describing efficient algorithms to perform either one. We also show how inferring the SBM can be used to predict missing and spurious links, and shed light on the fundamental limitations of the detectability of modular structures in networks. arXiv:1705.10225v8 [stat.ML] 6 Feb 2020 aTo appear in “Advances in Network Clustering and Blockmodeling,” edited by P. Doreian, V. Batagelj, A. Ferligoj, (Wiley, New York, 2019 [forthcoming]). y [email protected] 2 CONTENTS I. Introduction 3 II. Structure versus randomness in networks 3 III. The stochastic blockmodel (SBM) 5 IV. Bayesian inference: the posterior probability of partitions 7 V. Microcanonical models and the minimum description length principle (MDL) 11 VI. The “resolution limit” underfitting problem, and the nested SBM 13 VII. Model variations 17 A. Model selection 18 B. Degree correction 18 C. -
Facing the Future: European Research Infrastructures for the Humanities and Social Sciences
Facing the Future: European Research Infrastructures for the Humanities and Social Sciences Adrian Duşa, Dietrich Nelle, Günter Stock, and Gert G. Wagner (Eds.) Facing the Future: European Research Infrastructures for the Humanities and Social Sciences E d i t o r s : Adrian Duşa (SCI-SWG), Dietrich Nelle (BMBF), Günter Stock (ALLEA), and Gert. G. Wagner (RatSWD) ISBN 978-3-944417-03-5 1st edition © 2014 SCIVERO Verlag, Berlin SCIVERO is a trademark of GWI Verwaltungsgesellschaft für Wissenschaftspoli- tik und Infrastrukturentwicklung Berlin UG (haftungsbeschränkt). This book documents the results of the conference Facing the Future: European Research Infrastructure for Humanities and Social Sciences (November 21/22 2013, Berlin), initiated by the Social and Cultural Innovation Strategy Work- ing Group of ESFRI (SCI-SWG) and the German Federal Ministry of Education and Research (BMBF), and hosted by the European Federation of Academies of Sciences and Humanities (ALLEA) and the German Data Forum (RatSWD). Thanks and appreciation are due to all authors, speakers and participants of the conference, and all involved institutions, in particular the German Federal Ministry of Education and Research (BMBF). The ministry funded the conference and this subsequent publication as part of the Union of the German Academies of Sciences and Humanities’ project “Survey and Analysis of Basic Humanities and Social Science Research at the Science Academies Related Research Insti- tutes of Europe”. The views expressed in this publication are exclusively the opinions of the authors and not those of the German Federal Ministry of Education and Research. Editing: Dominik Adrian, Camilla Leathem, Thomas Runge, Simon Wolff Layout and graphic design: Thomas Runge Contents Preface . -
Extracting Insights from Differences: Analyzing Node-Aligned Social
Extracting Insights from Differences: Analyzing Node-aligned Social Graphs by Srayan Datta A dissertation submitted in partial fulfillment of the requirements for the degree of Doctor of Philosophy (Computer Science and Engineering) in The University of Michigan 2019 Doctoral Committee: Associate Professor Eytan Adar, Chair Associate Professor Mike Cafarella Assistant Professor Danai Koutra Associate Professor Clifford Lampe Srayan Datta [email protected] ORCID iD: 0000-0002-5800-830X c Srayan Datta 2019 To my family and friends ii ACKNOWLEDGEMENTS There are several people who made this dissertation possible, first among this long list is my adviser, Eytan Adar. Pursuing a doctoral program after just finishing un- dergraduate studies can be a daunting task but Eytan made it easy with his patience, kindness, and guidance. I learned a lot from our collaborations and idle conversations and I am very grateful for that. I would like to extend my thanks to the rest of my thesis committee, Mike Ca- farella, Danai Koutra and Cliff Lampe for their suggestions and constructive feed- back. I would also like to thank the following faculty members, Daniel Romero, Ceren Budak, Eric Gilbert and David Jurgens for their long insightful conversations and suggestions about some of my projects. I would like to thank all of friends and colleagues who helped (as a co-author or through critique) or supported me through this process. This is an enormous list but I am especially thankful to Chanda Phelan, Eshwar Chandrasekharan, Sam Carton, Cristina Garbacea, Shiyan Yan, Hari Subramonyam, Bikash Kanungo, and Ram Srivatasa. I would like to thank my parents for their unwavering support and faith in me. -
Network Analysis with Nodexl
Social data: Advanced Methods – Social (Media) Network Analysis with NodeXL A project from the Social Media Research Foundation: http://www.smrfoundation.org About Me Introductions Marc A. Smith Chief Social Scientist Connected Action Consulting Group [email protected] http://www.connectedaction.net http://www.codeplex.com/nodexl http://www.twitter.com/marc_smith http://delicious.com/marc_smith/Paper http://www.flickr.com/photos/marc_smith http://www.facebook.com/marc.smith.sociologist http://www.linkedin.com/in/marcasmith http://www.slideshare.net/Marc_A_Smith http://www.smrfoundation.org http://www.flickr.com/photos/library_of_congress/3295494976/sizes/o/in/photostream/ http://www.flickr.com/photos/amycgx/3119640267/ Collaboration networks are social networks SNA 101 • Node A – “actor” on which relationships act; 1-mode versus 2-mode networks • Edge B – Relationship connecting nodes; can be directional C • Cohesive Sub-Group – Well-connected group; clique; cluster A B D E • Key Metrics – Centrality (group or individual measure) D • Number of direct connections that individuals have with others in the group (usually look at incoming connections only) E • Measure at the individual node or group level – Cohesion (group measure) • Ease with which a network can connect • Aggregate measure of shortest path between each node pair at network level reflects average distance – Density (group measure) • Robustness of the network • Number of connections that exist in the group out of 100% possible – Betweenness (individual measure) F G • -
On Community Structure in Complex Networks: Challenges and Opportunities
On community structure in complex networks: challenges and opportunities Hocine Cherifi · Gergely Palla · Boleslaw K. Szymanski · Xiaoyan Lu Received: November 5, 2019 Abstract Community structure is one of the most relevant features encoun- tered in numerous real-world applications of networked systems. Despite the tremendous effort of a large interdisciplinary community of scientists working on this subject over the past few decades to characterize, model, and analyze communities, more investigations are needed in order to better understand the impact of community structure and its dynamics on networked systems. Here, we first focus on generative models of communities in complex networks and their role in developing strong foundation for community detection algorithms. We discuss modularity and the use of modularity maximization as the basis for community detection. Then, we follow with an overview of the Stochastic Block Model and its different variants as well as inference of community structures from such models. Next, we focus on time evolving networks, where existing nodes and links can disappear, and in parallel new nodes and links may be introduced. The extraction of communities under such circumstances poses an interesting and non-trivial problem that has gained considerable interest over Hocine Cherifi LIB EA 7534 University of Burgundy, Esplanade Erasme, Dijon, France E-mail: hocine.cherifi@u-bourgogne.fr Gergely Palla MTA-ELTE Statistical and Biological Physics Research Group P´azm´any P. stny. 1/A, Budapest, H-1117, Hungary E-mail: [email protected] Boleslaw K. Szymanski Department of Computer Science & Network Science and Technology Center Rensselaer Polytechnic Institute 110 8th Street, Troy, NY 12180, USA E-mail: [email protected] Xiaoyan Lu Department of Computer Science & Network Science and Technology Center Rensselaer Polytechnic Institute 110 8th Street, Troy, NY 12180, USA E-mail: [email protected] arXiv:1908.04901v3 [physics.soc-ph] 6 Nov 2019 2 Cherifi, Palla, Szymanski, Lu the last decade. -
Treemap User Guide
TreeMap User Guide Macrofocus GmbH Version 2019.8.0 Table of Contents Introduction. 1 Getting started . 2 Load and filter the data . 2 Set-up the visualization . 5 View and analyze the data. 7 Fine-tune the visualization . 12 Export the result. 15 Treemapping . 16 User interface . 20 Menu and toolbars. 20 Status bar . 24 Loading data . 25 File-based data sources. 25 Directory-based data sources . 31 Database connectivity. 32 On-line data sources . 33 Automatic default configuration . 33 Data types . 33 Configuration panel . 36 Layout . 38 Group by. 53 Size. 56 Color . 56 Height . 61 Labels . 61 Tooltip. 63 Rendering . 66 Legend . 67 TreeMap view . 69 Zooming . 69 Drilling . 70 Probing and selection . 70 TreePlot view. 71 Configuration . 71 Zooming . 72 Drilling . 72 Probing and selection . 72 TreeTable view . 73 Sorting . 73 Probing and selection . 74 Filter on a subset . 75 Search . 76 Filter . 76 See details. 77 Configure variables . 78 Formatting patterns. 78 Expression. -
Community Detection with Dependent Connectivity∗
Submitted to the Annals of Statistics arXiv: arXiv:0000.0000 COMMUNITY DETECTION WITH DEPENDENT CONNECTIVITY∗ BY YUBAI YUANy AND ANNIE QUy University of Illinois at Urbana-Champaigny In network analysis, within-community members are more likely to be connected than between-community members, which is reflected in that the edges within a community are intercorrelated. However, existing probabilistic models for community detection such as the stochastic block model (SBM) are not designed to capture the dependence among edges. In this paper, we propose a new community detection approach to incorporate intra-community dependence of connectivities through the Bahadur representation. The pro- posed method does not require specifying the likelihood function, which could be intractable for correlated binary connectivities. In addition, the pro- posed method allows for heterogeneity among edges between different com- munities. In theory, we show that incorporating correlation information can achieve a faster convergence rate compared to the independent SBM, and the proposed algorithm has a lower estimation bias and accelerated convergence compared to the variational EM. Our simulation studies show that the pro- posed algorithm outperforms the popular variational EM algorithm assuming conditional independence among edges. We also demonstrate the application of the proposed method to agricultural product trading networks from differ- ent countries. 1. Introduction. Network data has arisen as one of the most common forms of information collection. This is due to the fact that the scope of study not only focuses on subjects alone, but also on the relationships among subjects. Networks consist of two components: (1) nodes or ver- tices corresponding to basic units of a system, and (2) edges representing connections between nodes. -
Quick Notes on Nodexl
Quick notes on NodeXL Programme organisation, 3 components: • NodeXL command ‘ribbon’ – access with menu-like item at top of the screen, contains all network data functions • Data sheets – 5 separate sheets for edges, vertices, metrics etc, switching between them using tabs on the bottom. NB because the screen gets quite crowded the sheet titles sometimes disappear off to one side so use the small arrows (bottom left) to scroll across the tabs if something seems to have gone missing. • Graph viewer – separate window pane labelled ‘Document Actions’ with graphical layout functions. Can be closed while working with data – to open it again click the ‘Show Graph’ button near the left-hand side of the ribbon. Working with data, basic process: 1. Start in ‘Edges’ data sheet (leftmost tab at the bottom); each row represents one connection of the form A links to B (directed graph) OR A and B are linked (undirected). (Directed/Undirected changed in ‘Type’ option on ribbon. In directed graph this sheet might include one row for A->B and one for B->A; in an undirected graph this would be tautologous.) No other data in this sheet is required, although optionally: • Could include text under ‘label’ if the visible label on graphs should be different from the name used in columns A or B. NB These might be overwritten when using ‘autofill columns’. • Could add an extra column at column L to identify weight of links if (as is the case with IssueCrawler data) each connection might represent more than 1 actually existing link between two vertices, hence inclusion of ‘Edge Weight’ in blogs data. -
Developing a Credit Scoring Model Using Social Network Analysis
Developing a Credit Scoring Model Using Social Network Analysis Ahmad Abd Rabuh UP821483 This thesis is submitted in fulfilment of the requirements for the award of the degree of Doctor of Philosophy in the School of Business and Law at the University of Portsmouth November 2020 1 Author’s Declaration Whilst registered as a candidate for the above degree, I have not been registered for any other research award. The results and conclusions embodied in this thesis are the work of the named candidate and have not been submitted for any other academic award. Signature: Name: Ahmad Abd Rabuh Date: the 30th of April 2020 Word Count: 63,809 2 Acknowledgements As I witness this global pandemic, my greatest gratitude goes to my mother, Hanan, for her moral support and encouragement through the good, the bad and the ugly. During the period of my PhD revision, I received mentoring and help from my future lifetime partner, Kawthar. Also, I have been blessed with the help of wonderful people from Scholars at Risk, Rose Anderson and Sarina Rosenthal, who assisted me in the PhD application and scholarship processes. Additionally, I cannot be thankful enough to the Associate Dean of Research, Andy Thorpe, who supported me in every possible way during my time at the University of Portsmouth. Finally, many thanks to my supervisors, Mark Xu and Renatas Kizys for their continuous guidance. 3 Developing a Credit Scoring Model Using Social Network Analysis Table of Contents Abstract ......................................................................................................................................... 10 CHAPTER 1: INTRODUCTION ................................................................................................. 12 1.1. Research Questions ........................................................................................................ 16 1.2. Aims and Objectives ..................................................................................................... -
Livro De Resumos EDITORA DA UNIVERSIDADE FEDERAL DE SERGIPE
Organizadores: Carlos Alexandre Borges Garcia Marcus Eugênio Oliveira Lima Livro de Resumos EDITORA DA UNIVERSIDADE FEDERAL DE SERGIPE COORDENADORA DO PROGRAMA EDITORIAL Messiluce da Rocha Hansen COORDENADOR GRÁFICO DA EDITORA UFS Germana Gonçalves de Araújo PROJETO GRÁFICO E EDITORAÇÃO ELETRÔNICA Alisson Vitório de Lima FOTOGRAFIAS Disponibilizadas sob licença Creative Commons, ou de domínio público. Adilson Andrade - Foto da página X; FICHA CATALOGRÁFICA ELABORADA PELA BIBLIOTECA CENTRAL UNIVERSIDADE FEDERAL DE SERGIPE Encontro de Pós-Graduação (8. : 2016 : São Cristóvão, SE) Livro de resumos [recurso eletrônico] : VIII Encontro de Pós-Graduação / organizadores: Carlos Alexandre Borges Garcia, Marcus Eugênio Oliveira Lima. – São Cristóvão : Editora UFS : Universidade E56l Federal de Sergipe, Programa de Pós-Graduação, 2016. 353 p. ISBN 978-85-7822-550-6 1. Pós-graduação – Congresso. I. Universidade Federal de Sergipe. II. Garcia, Carlos Alexandre Borges. III. Lima, Marcus Eugênio Oliveira. CDU 378.046-021.68 Cidade Universitária “Prof. José Aloísio de Campos” CEP 49.100-000 – São Cristóvão - SE. Telefone: 3194 - 6922/6923. e-mail: [email protected] http://livraria.ufs.br/ Este portfólio, ou parte dele, não pode ser reproduzido por qualquer meio sem autorização escrita da Editora. Organizadores: Carlos Alexandre Borges Garcia Marcus Eugênio Oliveira Lima Livro de Resumos UFS São Cristóvão/SE - 2016 Ciências Agrárias A (des)realização da estratégia democrático-popular: Uma análise a partir da realidade do movimento dos trabalhadores rurais sem terra (MST) e do Partido dos Trabalhadores (PT) Autor: SOUSA, Ronilson Barboza de. Orientador: RAMOS FILHO, Eraldo da Silva. A referente tese de doutorado, que vem sendo desenvolvida, tem como objetivo analisar o processo de realização da estratégia democrático-popular, adotada pelo Movimento dos Trabalhadores Rurais Sem Terra (MST) e pelo Partido dos Trabalhadores (PT), especial- mente na luta pela terra e pela reforma agrária. -
NODEXL for Beginners Nasri Messarra, 2013‐2014
NODEXL for Beginners Nasri Messarra, 2013‐2014 http://nasri.messarra.com Why do we study social networks? Definition from: http://en.wikipedia.org/wiki/Social_network: A social network is a social structure made up of a set of social actors (such as individuals or organizations) and a set of the dyadic ties between these actors. Social networks and the analysis of them is an inherently interdisciplinary academic field which emerged from social psychology, sociology, statistics, and graph theory. From http://en.wikipedia.org/wiki/Sociometry: "Sociometric explorations reveal the hidden structures that give a group its form: the alliances, the subgroups, the hidden beliefs, the forbidden agendas, the ideological agreements, the ‘stars’ of the show". In social networks (like Facebook and Twitter), sociometry can help us understand the diffusion of information and how word‐of‐mouth works (virality). Installing NODEXL (Microsoft Excel required) NodeXL Template 2014 ‐ Visit http://nodexl.codeplex.com ‐ Download the latest version of NodeXL ‐ Double‐click, follow the instructions The SocialNetImporter extends the capabilities of NodeXL mainly with extracting data from the Facebook network. To install: ‐ Download the latest version of the social importer plugins from http://socialnetimporter.codeplex.com ‐ Open the Zip file and save the files into a directory you choose, e.g. c:\social ‐ Open the NodeXL template (you can click on the Windows Start button and type its name to search for it) ‐ Open the NodeXL tab, Import, Import Options (see screenshot below) 1 | Page ‐ In the import dialog, type or browse for the directory where you saved your social importer files (screenshot below): ‐ Close and open NodeXL again For older Versions: ‐ Visit http://nodexl.codeplex.com ‐ Download the latest version of NodeXL ‐ Unzip the files to a temporary folder ‐ Close Excel if it’s open ‐ Run setup.exe ‐ Visit http://socialnetimporter.codeplex.com ‐ Download the latest version of the socialnetimporter plug in 2 | Page ‐ Extract the files and copy them to the NodeXL plugin direction.