A Novel Modified Apriori Approach for Web Document Clustering

Roul, Rajendra Kumar; Varshneya, Saransh; Kalra, Ashu; Sahay, Sanjay Kumar

doi:10.1007/978-81-322-2202-6_14

Rajendra Kumar Roul⁷,
Saransh Varshneya⁷,
Ashu Kalra⁷ &
…
Sanjay Kumar Sahay⁷

Part of the book series: Smart Innovation, Systems and Technologies ((SIST,volume 33))

1467 Accesses
8 Citations

Abstract

The Traditional apriori algorithm can be used for clustering the web documents based on the association technique of data mining. But this algorithm has several limitations due to repeated database scans and its weak association rule analysis. In modern world of large databases, efficiency of traditional apriori algorithm would reduce manifolds. In this paper, we proposed a new modified apriori approach by cutting down the repeated database scans and improving association analysis of traditional apriori algorithm to cluster the web documents. Further we improve those clusters by applying Fuzzy C-Means (FCM), K-Means and Vector Space Model (VSM) techniques separately. We use Classic3 and Classic4 datasets of Cornell University having more than 10,000 documents and run both traditional apriori and our modified apriori approach on it. Experimental results show that our approach outperforms the traditional apriori algorithm in terms of database scan and improvement on association of analysis.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Log in via an institution

Chapter: USD 29.95; Price excludes VAT (USA)

eBook: USD 169.00; Price excludes VAT (USA)

Softcover Book: USD 219.99; Price excludes VAT (USA)

Hardcover Book: USD 219.99; Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

References

Barsagade, N.: Web usage mining and pattern discovery: asurvey paper. In: CSE8331, Dec 2003
Google Scholar
Kumar, T.S.: Introduction to Data Mining. Pearson Education, Upper Saddle River (2006)
Google Scholar
Agrawal, R., Mannila, H., Srikant, R., Toivonen, H., InkeriVerkamo, A.: Fast discovery of association rules in large databases. In: Fayyad, U.M., Piatetsky-Shapiro, G., Smyth, P. (eds.) Advances in Knowledge Discovery and Data Mining, pp. 307–328. AAAI press, Menlo Park (1996)
Google Scholar
Wang, P., Shi, L., Bai, J., Zhao, Y.: Mining association rules based on apriori algorithm and application. In: International Forum on Computer Science-Technology and Applications, IFCSTA’09, vol. 1 (2009)
Google Scholar
Bodon, F.: A trie-based APRIORI implementation for mining frequent itemset. In: ACM 1-59593-210-0/05/08 (2005)
Google Scholar
Cao, X.: An algorithm of mining association rules based on granular computing. In: International Conference on Medical Physics and Biomedical Engineering (2012)
Google Scholar
Wang, Y., Jin, Y., Li, Y., Geng, K.: Data mining based on improved apriori algorithm. Commun. Comput. Inf. Sci. 392, 354–363 (2013)
Article Google Scholar
Li, X., Shang, J.: A novel apriori algorithm based on cross linker. In: Proceedings of the International Conference on Information Engineering and Applications (IEA) (2012)
Google Scholar
Tomanová, I., Kupka, J.: Implementation of background knowledge and properties induced by fuzzy confirmation measures in apriori algorithm. In: International Joint Conference CISIS’12-ICEUTE’12-SOCO’, vol. 189, pp. 533–542 (2013)
Google Scholar
Li, Y., Xing, J., Wu, R., Zheng, F.: Web clustering using a two-layer approach. LNCS, vol. 6988, pp. 211–218. Springer, Heidelberg (2011)
Google Scholar
Huang, F., Zhang, S., He, M., Wu, X.: Clustering web documents using hierarchical representation with multi-granularity. World Wide Web 17, 105–126 (2014). doi:10.1007/s11280-012-0197-x
Article Google Scholar
Roul, R.K., Devanand O.R., Sahay S.K.: Web document clustering and ranking using Tf-Idf based Apriori Approach. In: IJCA Proceedings on International Conference on Advances in Computer Engineering and Applications, pp. 34–39 (2014)
Google Scholar
Lee, I., On, B.W.: An effective web document clustering algorithm based on bisection and merge. Artif. Intell. 36, 69–85 (2011). doi:10.1007/s10462-011-9203-4
Article Google Scholar
http://www.nlp.fi.muni.cz/projekty/gensim/intro.html
Orlando, S., Palmerini, P., Perego, R.: Enhancing the apriori algorithm for frequent set counting. LNCS, vol. 2114, pp. 71–82. Springer, Heidelberg (2001)
Google Scholar
Singh, J., Ram, H., Sodhi, J.S.: Improving efficiency of apriori algorithm using transaction reduction. Int. J. Sci. Res. Publ. 3(1), 1–4 (2013)
Google Scholar
http://www.dataminingresearch.com/index.php/2010/09/classic3-classic4-datasets

Download references

Author information

Authors and Affiliations

BITS Pilani K. K. Birla Goa Campus, Zuarinagar, 403726, Goa, India
Rajendra Kumar Roul, Saransh Varshneya, Ashu Kalra & Sanjay Kumar Sahay

Authors

Rajendra Kumar Roul
View author publications
You can also search for this author in PubMed Google Scholar
Saransh Varshneya
View author publications
You can also search for this author in PubMed Google Scholar
Ashu Kalra
View author publications
You can also search for this author in PubMed Google Scholar
Sanjay Kumar Sahay
View author publications
You can also search for this author in PubMed Google Scholar

Corresponding author

Correspondence to Rajendra Kumar Roul .

Editor information

Editors and Affiliations

School of Electrical and Information Engineering, University of South Australia, South Australia, Australia
Lakhmi C. Jain
Computer Science and Engineering, Veer Surendra Sai University of Technolo, Sambalpur, Odisha, India
Himansu Sekhar Behera
Computer Science & Engineering, Kalyani University, Nadia, West Bengal, India
Jyotsna Kumar Mandal
Dept. of Computer Science and Eng., National Institute of Technology Rourkela, Rourkela, India
Durga Prasad Mohapatra

Rights and permissions

Reprints and permissions

Copyright information

About this paper

Cite this paper

Roul, R.K., Varshneya, S., Kalra, A., Sahay, S.K. (2015). A Novel Modified Apriori Approach for Web Document Clustering. In: Jain, L., Behera, H., Mandal, J., Mohapatra, D. (eds) Computational Intelligence in Data Mining - Volume 3. Smart Innovation, Systems and Technologies, vol 33. Springer, New Delhi. https://doi.org/10.1007/978-81-322-2202-6_14

Download citation

DOI: https://doi.org/10.1007/978-81-322-2202-6_14
Published: 12 December 2014
Publisher Name: Springer, New Delhi
Print ISBN: 978-81-322-2201-9
Online ISBN: 978-81-322-2202-6
eBook Packages: EngineeringEngineering (R0)

Publish with us

Policies and ethics