A Performance Evaluation of the Preprocessing Phase of Multiple Keyword Matching Algorithms

Size: px

Start display at page:

Download "A Performance Evaluation of the Preprocessing Phase of Multiple Keyword Matching Algorithms"

Stewart Owen
5 years ago
Views:

1 A Performance Evaluation of the Preprocessing Phase of Multiple Keyword Matching Algorithms Charalampos S. Kouzinopoulos and Konstantinos G. Margaritis Parallel and Distributed Processing Laboratory Department of Applied Informatics, University of Macedonia 156 Egnatia str., P.O. Box 1591, 546 Thessaloniki, Greece Abstract Multiple keyword matching is an important problem in text processing that involves the location of all the positions of an input string where one or more keywords from a finite set occur. Modern multiple keyword matching algorithms can scan the input string in a single pass by preprocessing the keyword set, an essential phase that affects the overall performance of each algorithm. This paper presents a performance evaluation in terms of preprocessing of the well known Commentz-Walter, Wu-Manber, Set Backward Oracle Matching and Salmela-Tarhio-Kytöjoki multiple keyword matching algorithms for different types of keywords and for several problem parameters. Keywords-Algorithms, Performance Evaluation, Preprocessing, Multiple Keyword Matching, Multiple Pattern Matching I. INTRODUCTION Multiple keyword matching is an important problem in text processing and is commonly used to locate all the positions of an input string (the so called text ) where one or more keywords (the so called patterns ) from a finite set of keywords occur. It is the computationally intensive kernel of many security and network applications including information retrieval, intrusion detection systems, web filtering, virus scanners and spam filters while it is also used as a powerful tool in locating nucleotide or Amino Acid sequence keywords in biological sequence databases. The multiple keyword matching problem can be defined as: Definition 1 Given an input string T = t 1 t 2...t n of length n and a finite set of r keywords P = p 1,p 2,...,p r, where each p i is a string p i = p i 1p i 2...p i m of length m over a finite character set Σ and the total size of all keywords is denoted as P, the task is to find all occurrences of any of the keywords in the input string. Preprocessing is an important phase of multiple keyword matching algorithms that is used so that they can scan the input string T in a single pass to locate all occurrences of any keyword from a finite keyword set. This is achieved by processing the keyword set and constructing some necessary data structures, usually automatons or tables of hashed keywords. An efficient preprocessing phase is therefore crucial for the overall performance of the algorithms in terms of running time and memory usage. For the experiments of this paper, the Commentz-Walter (CW) [2], Wu-Manber (WM) [15], Set Backward Oracle Matching (SBOM) [8] and the Salmela-Tarhio-Kytöjoki [12] family of the Horspool with q-grams (HG), Shift-Or with q-grams and BNDM with q-grams (BG) algorithms were used, algorithms that are simple, efficient and widely used. Commentz-Walter combines the filtering functions of the single keyword matching Boyer-Moore algorithm and a suffix automaton to search for the occurrence of multiple keywords in an input string. During the preprocessing phase, the algorithm creates a trie structure from the reversed keywords of the keyword set where each node corresponds to a single character, constructs the two shift functions of the Boyer-Moore algorithm extended to multiple keywords and specifies the exit nodes that indicate that a complete match is found in O( P ) time. Wu-Manber is a generalization of the Horspool [3] algorithm for multiple keyword matching. To achieve a good performance as P increases, the algorithm essentially enlarges the alphabet size by considering the text as blocks of size B instead of single characters. During the preprocessing phase three tables are built, the SHIFT, HASH and PREFIX tables. SHIFT is used to determine the number of characters that can be safely skipped based on the previous B characters on each text position, PREFIX stores a hashed value of the B-characters prefix of each keyword while HASH contains a list of all keywords with the same prefix. As recommended in [15], usually B could be equal to 2 for a small keyword set size or to 3 otherwise. Since the experiments of this paper involve large keyword set sizes, Wu-Manber was implemented with a block size of B =3. The Set Backward Oracle Matching algorithm uses a factor oracle, an acyclic automaton that was first introduced in [1], with at most P +1 states and a linear in P number of transitions. During the preprocessing phase, the factor oracle is created from the set of the reversed keywords in O( P ) time. Apart from the transitions that link the nodes of each keyword, a set of at most P external transitions must be built so that the oracle can recognize at least any

2 factor of a keyword. The external transitions associate each state i of the oracle to a previous state j that is called the supply state of i, such as i>j. The Salmela-Tarhio-Kytöjoki algorithms are character class filters; they essentially construct a generalized keyword with a length of m characters in O( P ) time for BG and SOG and O( P m) time for the HG algorithm that simultaneously matches all the keywords. As P increases, the efficiency of the filters should decrease since a candidate match would occur in almost every position [11]. To solve this problem, the algorithms are using a similar technique to the Wu-Manber algorithm. They treat the input string and the keywords in groups of B characters, effectively enlarging the alphabet size to Σ B characters. To reduce the required memory space to 2 21 bytes a hashing technique can be applied. For the experiments of this paper, the HG, SOG and BG algorithms were implemented using hashed 3-grams. The Commentz-Walter algorithm is substantially faster in practice than the Aho-Corasick algorithm, particularly when long keywords are involved [14][15]. Wu-Manber is considered to be a practical, simple and efficient algorithm for multiple keyword matching [8]. The Set Backward Oracle Matching algorithm has the same performance as Set Backward Dawg Matching but uses a much simpler automaton while at the same time appears to be very efficient when used on large keyword sets [8]. Finally, Salmela- Tarhio-Kytöjoki is a recently introduced family of algorithms that has a reportedly good performance on specific types of data [5]. Table I KNOWN THEORETICAL PREPROCESSING, WORST AND AVERAGE TIME COMPLEXITY OF THE MULTIPLE KEYWORD MATCHING ALGORITHMS Algorithm Preprocessing Worst case Average case CW P n m n SBOM P n P n HG P m n P nlog Σ ( P )/m SOG P n P n BG P n P nlog Σ ( P )/m Table I summarizes the known theoretical preprocessing, worst and average time complexity of the presented algorithms. The worst case complexity noted for the HG, SOG and BG algorithms is for keywords that all have the same hash value. If all keywords have different hash values, the worst case time complexity is O(n(logr + m)) instead. It was impossible to calculate the theoretical complexity of the Wu-Manber algorithm and thus was omitted, as the original paper does not specify the best size of the HASH and SHIFT tables and the hash functions, parameters that affect the complexity [8]. Several experiments on multiple keyword matching algorithms have already been reported in [4], [5], [6], [9], [11], [13] for alphabet input strings, randomly generated data sets, live network traffic and biological sequence databases. Most of that work though, concentrated on the performance evaluation of the search phase of the algorithms. The performance of the preprocessing phase of different multiple keyword matching algorithms has not been studied extensively in the past and usually focuses on different implementations of the same algorithm (i.e. [1]). The aim of this paper is to evaluate the performance of the preprocessing phase in terms of running time of the presented multiple keyword matching algorithms. The algorithms are compared for different types of keywords including randomly generated keywords, alphabet keywords and biological sequence databases and for several problem parameters such as the total size of the keyword set and the length and alphabet size of the keywords. II. EXPERIMENTAL METHODOLOGY The parameters that describe the performance of the preprocessing phase of multiple keyword matching algorithms are the size of the keyword set r, the length of the keywords m and the alphabet size Σ used. The data set was similar to the sets used in [4], [7], [13]. It consisted of randomly generated texts of size n =4.. with a binary alphabet and an alphabet of size 8, the CIA World Fact Book from the Large Canterbury Corpus with a size of n = and an alphabet of size 94, the genome of Escherichia coli from the Large Canterbury Corpus with a size of n = and an alphabet of size Σ=4, the SWISS-PROT Amino Acid sequence database with a size of n = and an alphabet of size Σ=2, the FASTA Amino Acid () of the A-thaliana genome with a size of n = and an alphabet of size Σ=2and the FASTA Nucleidic Acid () sequences of the A-thaliana genome with a size of n = and an alphabet of size Σ=4.The keyword set used consisted of 1. and 1. keywords where each keyword had a length of m =8and m =32 characters. The experiments were executed locally on an Intel Core 2 Duo CPU with a 3.GHz clock speed and 2 Gb of memory, 64 KB L1 cache and 6 MB L2 cache. The Ubuntu Linux operating system was used and during the experiments only the typical background processes ran. To decrease random variation, the time results were averages of 1 runs. All algorithms were implemented using the ANSI C programming language and were compiled using the GCC compiler with the -O2 and -funroll-loops optimization flags. III. ANALYSIS Since the performance of the preprocessing phase of different multiple keyword matching algorithms has not been studied extensively in the past, it is not possible to compare the results with other published metrics. The running time results presented in Tables II to V generally agree with previous work as discussed in [5], [6]. Based on the theoretical time complexity of the algorithms as reported in the original papers and summarized in Table I,

3 Table II PREPROCESSING AND RUNNING TIME OF THE ALGORITHMS FOR 1. KEYWORDS, FOR ALL TYPES OF DATA WITH m =8(SEC) Random Σ=2 Random Σ=8 Swiss Prot. CW WM SBOM HG SOG BG Table III PREPROCESSING AND RUNNING TIME OF THE ALGORITHMS FOR 1. KEYWORDS, FOR ALL TYPES OF DATA WITH m =32(SEC) Random Σ=2 Random Σ=8 Swiss Prot. CW WM SBOM HG SOG BG Table IV PREPROCESSING AND RUNNING TIME OF THE ALGORITHMS FOR 1. KEYWORDS, FOR ALL TYPES OF DATA WITH m =8(SEC) Random Σ=2 Random Σ=8 Swiss Prot. CW WM SBOM HG SOG BG Table V PREPROCESSING AND RUNNING TIME OF THE ALGORITHMS FOR 1. KEYWORDS, FOR ALL TYPES OF DATA WITH m =32(SEC) Random Σ=2 Random Σ=8 Swiss Prot. CW WM SBOM HG SOG BG the time required by an algorithm to construct the necessary data structures and process the keyword set during the preprocessing phase is expected to be based on the size r of the keyword set and the length m of the keywords. In practice though it can be seen that the alphabet Σ of the keywords is an important factor that also affects the preprocessing time. Tables II to V present the time spent by the algorithms during the preprocessing phase along the running time for keyword sets consisting of 1. and 1. keywords with lengths of m =8and m =32and for all types of data while Figure 1 depicts the preprocessing time for the same data set as a percent of the running time of the algorithms. It is clear that the size r of the keyword set was the primary factor that affected the performance of the algorithms in terms of preprocessing time. As r increased to 1. keywords, the preprocessing time of each algorithm increased linear in r as expected from the theoretical complexity presented in Table I. The preprocessing time to compute the trie of the reversed keywords and the two shift functions of the Commentz-Walter algorithm increased up to 5 times when larger keyword sets were used although the running time of the algorithm was not affected as much. The Set Backward Oracle Matching algorithm had the slowest preprocessing time comparing to the rest of the algorithms when 1. keywords were used. The factor oracle constructed during preprocessing though is the reason why SBOM is one of the fastest algorithms in terms of

4 Algorithm used (m=8, 1. keywords) Algorithm used (m=32, 1. keywords) Algorithm used (m=8, 1. keywords) Algorithm used (m=32, 1. keywords) Figure 1. Percentage of preprocessing on running time for different types of data running time, especially on randomly generated binary data sets and biological databases when a keyword of length m =8was used. As depicted in Figure 1, the preprocessing of Set Backward Oracle Matching accounted for 97% of the running time of the algorithm for some types of data. HG had a 2 to 8 times better performance in terms of preprocessing time than the SOG and BG algorithms when used on sets with 1. keywords and in general its preprocessing phase was faster than the rest of the presented algorithms for most types of data. On sets with 1. keywords though, the performance of HG in terms of preprocessing time was affected more from the increase in the size of the keyword set than the SOG and BG algorithms as can be explained by the P m comparing to P theoretical preprocessing time of the algorithm. The exact theoretical preprocessing time of the Wu-Manber algorithm is not clear from the original analysis by Wu and Manber but it can be concluded from the experimental results that the time spent by the algorithm to construct the necessary hash tables was generally linear in r and independent of the alphabet size. From Tables II to V it can also be seen that together with the Salmela-Tarhio-Kytöjoki algorithms, the preprocessing time of Wu-Manber increased less than that of the Commentz-Walter and the Set Backward Oracle Matching algorithms on sets with 1. keywords. The preprocessing time of the algorithms was affected in different ways by the length m of the keywords. When m increased from 8 to 32 characters and for all types of data, the preprocessing time of the Commentz-Walter and the Set Backward Oracle Matching algorithms generally increased linear in m while the time to complete the preprocessing phase of the Wu-Manber and the Salmela- Tarhio-Kytöjoki algorithms was unaffected, as presented in Tables II to V. For most types of data, the increase in the preprocessing time of the Commentz-Walter and the Set Backward Oracle Matching algorithms affected negatively their total performance. A notable exception to that was the use of the Commentz-Walter algorithm on randomly generated keyword sets with a binary alphabet and on the and keywords, keywords with an alphabet of

5 Σ = 4. In these cases were a small alphabet size was used, the increase in the preprocessing time resulted in a significant decrease in the running time of the algorithm, of up to 17 times. Although not expected from their theoretical complexity, the preprocessing time of the algorithms also depended on the alphabet size Σ of the keyword set. As can be concluded from Figure 1, the preprocessing time of the algorithms increased on keyword sets with larger alphabets. When language keywords were used, with an alphabet size Σ of 94 characters, the preprocessing to running time ratio of all algorithms drastically increased. It is interesting that for randomly generated keyword sets where a binary alphabet was used, the performance of the Set Backward Oracle Matching algorithm in terms of preprocessing time also decreased, indicative of the sensitivity of the algorithm to the alphabet size. Finally it should be noted that the preprocessing time of the Salmela-Tarhio-Kytöjoki algorithms was roughly constant in the alphabet size while at the same time their running time decreased for data sets with a larger alphabet size, outperforming the rest of the algorithms on language data sets in terms of running time. IV. CONCLUSIONS This paper presented a performance evaluation in terms of preprocessing time of the well known Commentz- Walter, Wu-Manber, Set Backward Oracle Matching and the Salmela-Tarhio-Kytöjoki multiple keyword matching algorithms for alphabet keyword sets, randomly generated keyword sets of a binary alphabet and an alphabet of size 8, the genome, the SWISS-PROT Amino Acid sequence database and the FASTA Amino Acid () and FASTA Nucleidic Acid () sequences of the A-thaliana genome. The keyword sets used consisted of 1. and 1. keywords with a length of m =8and m =32. It was shown that the time required by an algorithm to construct the necessary data structures and process the keyword set during the preprocessing phase is based on the size r of the keyword set and the length m of the keywords. It was also discussed that the alphabet Σ of the keywords is an important factor that also affects the preprocessing time of the algorithms. More specifically it was concluded that the preprocessing time of each algorithm increased generally linear in r as expected from their theoretical complexity. Additionally, the preprocessing time of the Commentz- Walter and Set Backward Oracle Matching algorithms generally increased linear in m while the time to complete the preprocessing phase of the Wu-Manber and the Salmela- Tarhio-Kytöjoki algorithms was unaffected by the keyword length. Finally, the performance of all algorithms in terms of preprocessing time decreased when used on keyword sets with a large alphabet size, although not expected by their theoretical preprocessing complexity. The work presented in this paper could be extended with a performance evaluation of the preprocessing phase of additional families of pattern matching algorithms including two dimensional and approximate pattern matching algorithms. A study of the preprocessing phase of the presented algorithms in terms of memory usage would also be interesting. REFERENCES [1] Allauzen, C., Crochemore, M., Raffinot, M.: Factor oracle: A new structure for pattern matching 1725, (1999) [2] Commentz-Walter, B.: A string matching algorithm fast on the average. Proceedings of the 6th Colloquium, on Automata, Languages and Programming pp (1979) [3] Horspool, R.: Practical fast searching in strings. Software: Practice and Experience 1(6), (198) [4] Kalsi, P., Peltola, H., Tarhio, T.: Comparison of exact string matching algorithms for biological sequences. Communications in Computer and Information Science pp (28) [5] Kouzinopoulos, C., Margaritis, K.: Experimental Results on Algorithms for Multiple Keyword Matching. In: IADIS International Conference on Informatics (21) [6] Kouzinopoulos, C., Margaritis, K.: Experimental Results On Multiple Pattern Matching Algorithms For Biological Sequences. In: International Conference on Bioinformatics - Models, Methods and Algorithms (211) [7] Lecroq, T.: Fast exact string matching algorithms. Information Processing Letters 12(6), (27) [8] Navarro, G., Raffinot, M.: Flexible pattern matching in strings: practical on-line search algorithms for texts and biological sequences. Cambridge University Press (22) [9] Navarro, G., Tarhio, J.: LZgrep: A Boyer-Moore string matching tool for Ziv-Lempel compressed text. Software-Practice and Experience 35(12), (25) [1] Nieminen, J., Kilpeläinen, P.: Efficient implementation of aho corasick pattern matching automata using unicode. Software: Practice and Experience 37(6), (27) [11] Salmela, L.: Improved Algorithms for String Searching Problems. Ph.D. thesis, Helsinki University of Technology (29) [12] Salmela, L., Tarhio, J., Kytöjoki, J.: Multipattern string matching with q -grams. Journal of Experimental Algorithmics 11, 1 19 (26) [13] Sheik, S., Aggarwal, S., Poddar, A., Sathiyabhama, B., Balakrishna, N., Sekar, K.: Analysis of string-searching algorithms on biological sequence databases. Current Science 89(2), (25) [14] Watson, B.: Taxonomies and toolkits of regular language algorithms. Ph.D. thesis, Eindhoven University of Technology (1995) [15] Wu, S., Manber, U.: A fast algorithm for multi-pattern searching pp (24), technical report TR-94-17

Multi-Pattern String Matching with Very Large Pattern Sets

Multi-Pattern String Matching with Very Large Pattern Sets Leena Salmela L. Salmela, J. Tarhio and J. Kytöjoki: Multi-pattern string matching with q-grams. ACM Journal of Experimental Algorithmics, Volume