AN OPTIMAL UNSUPERVISED TEXT DATA SEGMENTATION USING GENETIC ALGORITHM

Authors

  • N. Silpa Principal Investigator,Shri Vishnu Engineering College for Women, Bhimavaram, AP, India Author
  • V. V. R. Maheswara Rao Scientist Mentor ,Shri Vishnu Engineering College for Women, Bhimavaram, AP, India Author

Keywords:

Text Mining, Clustering, Genetic Algorithm, Optimization techniques

Abstract

The popularity of information available in electronic forms has been rapidly growing in the last decade, and turn into a golden mount containing extremely unstructured data for the researchers. Extracting interesting information and knowledge from such data creates promising future path into the era of text mining. The roots of text mining lie in most related research areas clustering, classification, information retrieval, machine learning and soft computing paradigms. Among all, Clustering is an unsupervised methodology having the ability to form meaningful natural groups of objects from given unlabeled data. A large number of clustering algorithms based on K-Means have been proposed on variety of domains for different types of applications none of these algorithms is suitable for all kinds of applications. This motivated and find a room for new clustering algorithm that is more efficient and optimal with computationally feasible. Genetic algorithms are randomized optimization techniques guided by the principles of evolution and natural genetics, having a large amount of implicit parallelism. To improve the text data segmentation accuracy the authors proposed An Optimal Unsupervised Text Data Segmentation Model using Genetic Algorithm so called OUTDSM. The encoding strategy, fitness function and operators of proposed OUTDSM works together and achieve high accuracy rated optimal clusters. Additionally, the nature of biological diversity of OUTDSM prevents the population from stagnating at any local optima and promises to arrive at global optima. The experimental results proving this claim are given in this paper.

References

Deepankar Bharadwaj, “Text Mining Technique using Genetic Algorithm”,

International Conference on Advances in Computer Application (ICACA - 2013),

pp:7-10, 2013.

Rashmi Agrawal, Mridula Batra, “A Detailed Study on Text Mining Techniques”,

International Journal of Soft Computing and Engineering (IJSCE) ISSN: 2231-

, Volume-2, Issue-6, January 2013.

V.V.R. Maheswara Rao, Dr. V. Valli Kumari, “An Intelligent Optimal Genetic

Model to Investigate the User Usage Behaviour on World Wide Web”, International

Journal of Data Mining & Knowledge Management Process Vol.3, No.2, 2013.

Divya Nasa, “Text Mining Techniques- A Survey”, International Journal of

Advanced Research in Computer Science and Software Engineering, Vol. 2, 2012.

Dr. A.V. Senthil Kumar, S.Mythili, “Parallel Implementation of Genetic Algorithm

using K-Means Clustering”, Int. J. Advanced Networking and Applications,

Volume:03 Issue:06 Pages:1450-1455, 2012.

K.Arun Prabha a, R.Saranya, “Refinement of K-Means Clustering Using Genetic

Algorithm”, Journal of Computer Applications (JCA) ISSN: 0974-1925, Volume

IV, Issue 2, pp: 40-44, 2011.

N. El-Bathy, C. Gloster, I. Kateeb, G. Stein, “Intelligent Extended Clustering

Genetic Algorithm for Information Retrieval Using BPEL”, American Journal of

Intelligent Systems, Vol 1(1): pp: 10-15, 2011.

Dharmendra K Roy and Lokesh K Sharma, “ Genetic k-Means Clustering

Algorithm for Mixed Numeric and Categoorical Data Sets”, International Journal of

Artificial Intelligence & Applications ( IJAIA), Vol.1, No.2, pp:23-28, 2010.

Vidhya. K. A & G. Aghila,” Text Mining Process, Techniques and Tools : An

Overview”, International Journal of Information Technology and Knowledge

Management, Volume 2, No. 2, pp. 613-622, 2010.

Kuan C. Chen, “Text Mining e-Complaints Data from e-Auction Store with

Implications for Internet Marketing Research”, Journal of Business & Economics

Research, Volume 7, Number 5, pp: 15-24, 2009.

Mahesh T R, Suresh M B, M Vinayababu, “Text Mining: Advancements,

Challenges and Future Directions”, International Journal of Reviews in Computing,

pp: 61-65, ISSN: 2076-3328, 2009.

Anna Huang, “Similarity Measures for Text Document Clustering”, New Zealand

Computer Science Research Student Conference, pp:49-56, 2008.

Milos Radovanovic, Mirjana Ivanovic “Text mining: Approaches and

Applications”, Novi Sad J. Math, Vol. 38, No. 3, pp: 227-234, 2008.

Anna Stavrianou, Periklis Andritsos, Nicolas Nicoloyannis, “Overview and

Semantic Issues of Text Mining”, SIGMOD, Vol. 36, No. 3, pp: 23-34, 2007.1. Deepankar Bharadwaj, “Text Mining Technique using Genetic Algorithm”,

International Conference on Advances in Computer Application (ICACA - 2013),

pp:7-10, 2013.

Rashmi Agrawal, Mridula Batra, “A Detailed Study on Text Mining Techniques”,

International Journal of Soft Computing and Engineering (IJSCE) ISSN: 2231-

, Volume-2, Issue-6, January 2013.

V.V.R. Maheswara Rao, Dr. V. Valli Kumari, “An Intelligent Optimal Genetic

Model to Investigate the User Usage Behaviour on World Wide Web”, International

Journal of Data Mining & Knowledge Management Process Vol.3, No.2, 2013.

Divya Nasa, “Text Mining Techniques- A Survey”, International Journal of

Advanced Research in Computer Science and Software Engineering, Vol. 2, 2012.

Dr. A.V. Senthil Kumar, S.Mythili, “Parallel Implementation of Genetic Algorithm

using K-Means Clustering”, Int. J. Advanced Networking and Applications,

Volume:03 Issue:06 Pages:1450-1455, 2012.

K.Arun Prabha a, R.Saranya, “Refinement of K-Means Clustering Using Genetic

Algorithm”, Journal of Computer Applications (JCA) ISSN: 0974-1925, Volume

IV, Issue 2, pp: 40-44, 2011.

N. El-Bathy, C. Gloster, I. Kateeb, G. Stein, “Intelligent Extended Clustering

Genetic Algorithm for Information Retrieval Using BPEL”, American Journal of

Intelligent Systems, Vol 1(1): pp: 10-15, 2011.

Dharmendra K Roy and Lokesh K Sharma, “ Genetic k-Means Clustering

Algorithm for Mixed Numeric and Categoorical Data Sets”, International Journal of

Artificial Intelligence & Applications ( IJAIA), Vol.1, No.2, pp:23-28, 2010.

Vidhya. K. A & G. Aghila,” Text Mining Process, Techniques and Tools : An

Overview”, International Journal of Information Technology and Knowledge

Management, Volume 2, No. 2, pp. 613-622, 2010.

Kuan C. Chen, “Text Mining e-Complaints Data from e-Auction Store with

Implications for Internet Marketing Research”, Journal of Business & Economics

Research, Volume 7, Number 5, pp: 15-24, 2009.

Mahesh T R, Suresh M B, M Vinayababu, “Text Mining: Advancements,

Challenges and Future Directions”, International Journal of Reviews in Computing,

pp: 61-65, ISSN: 2076-3328, 2009.

Anna Huang, “Similarity Measures for Text Document Clustering”, New Zealand

Computer Science Research Student Conference, pp:49-56, 2008.

Milos Radovanovic, Mirjana Ivanovic “Text mining: Approaches and

Applications”, Novi Sad J. Math, Vol. 38, No. 3, pp: 227-234, 2008.

Anna Stavrianou, Periklis Andritsos, Nicolas Nicoloyannis, “Overview and

Semantic Issues of Text Mining”, SIGMOD, Vol. 36, No. 3, pp: 23-34, 2007.

Lipika Dey, Muhammad Abulaish, Jahiruddin, Gaurav Sharma, “Text Mining through Entity-Relationship Based Information Extraction”, IEEE/WIC/ACM International Conferences on Web Intelligence and Intelligent Agent Technology – Workshops, pp:177-180, 2007.

M. Castellano, G. Mastronardi, A. Aprile, and G. Tarricone, “A Web Text Mining Flexible Architecture”, World Academy of Science, Engineering and Technology, pp: 78-85, 2007.

Elizabeth Leon, Olfa Nasraoui, Jonatan Gomez, “ECSAGO: Evolutionary Clustering with Self Adaptive Genetic Operators”, IEEE Congress on Evolutionary Computation Sheraton Vancouver Wall Centre Hotel, pp: 1768-1175, 2006.

Louis A., Francis, FCAS, MAAA, “Taming Text: An introduction to Text mining”, Casualty Actuarial Society Forum, pp: 51-88, 2006.

Joel D. Martin, “Fast and Furious Text Mining”, and IEEE Computer Society Technical Committee on Data Engineering, pp: 1-10, 2005.

Published

2013-05-13

How to Cite

N. Silpa, & V. V. R. Maheswara Rao. (2013). AN OPTIMAL UNSUPERVISED TEXT DATA SEGMENTATION USING GENETIC ALGORITHM. International Journal of Computer Science and Engineering Research and Development (IJCSERD), 3(2), 46-49. https://ijcserd.in/index.php/home/article/view/IJCSERD_03_02_005