T U M INSTITUT FÜR INFORMATIK Finite Automata, Digraph Connectivity, and Regular Expression Size Hermann Gruber Markus Holzer TUM-I0725 Dezember 07 TECHNISCHE UNIVERSITÄT MÜNCHEN TUM-INFO-12-I0725-0/1.-FI Alle Rechte vorbehalten Nachdruck auch auszugsweise verboten c 2007 Druck: Institut für Informatik der Technischen Universität München Finite Automata, Digraph Connectivity, and Regular Expression Size Hermann Gruber Institut für Informatik, Ludwig-Maximilians-Universität München, Oettingstraße 67, 80538 München, Germany
[email protected] Markus Holzer Institut für Informatik, Technische Universität München, Boltzmannstraße 3, D-85748 Garching, Germany
[email protected] Abstract We study some descriptional complexity aspects of regular expressions. In particular, we give lower bounds on the minimum required size for the conversion of deterministic finite automata into regular expressions and on the required size of regular expressions resulting from applying some basic language operations on them, namely intersection, shuffle, and complement. Some of the lower bounds obtained are asymptotically tight, and, notably, the examples we present are over an alphabet of size two. To this end, we develop a new lower bound technique that is based on the star height of regular languages. It is known that for a restricted class of regular languages, the star height can be determined from the digraph underlying the transition structure of the minimal finite automaton accepting that language. In this way, star height is tied to cycle rank, a structural complexity measure for digraphs proposed by Eggan and Büchi, which measures the degree of connectivity of directed graphs. This measure, although a robust graph theoretic notion, appears to have remained widely unnoticed outside formal language theory.