AN OPTIMAL UNSUPERVISED TEXT DATA SEGMENTATION USING GENETIC ALGORITHM
Keywords:
Text Mining, Clustering, Genetic Algorithm, Optimization techniquesAbstract
The popularity of information available in electronic forms has been rapidly growing in the last decade, and turn into a golden mount containing extremely unstructured data for the researchers. Extracting interesting information and knowledge from such data creates promising future path into the era of text mining. The roots of text mining lie in most related research areas clustering, classification, information retrieval, machine learning and soft computing paradigms. Among all, Clustering is an unsupervised methodology having the ability to form meaningful natural groups of objects from given unlabeled data. A large number of clustering algorithms based on K-Means have been proposed on variety of domains for different types of applications none of these algorithms is suitable for all kinds of applications. This motivated and find a room for new clustering algorithm that is more efficient and optimal with computationally feasible. Genetic algorithms are randomized optimization techniques guided by the principles of evolution and natural genetics, having a large amount of implicit parallelism. To improve the text data segmentation accuracy the authors proposed An Optimal Unsupervised Text Data Segmentation Model using Genetic Algorithm so called OUTDSM. The encoding strategy, fitness function and operators of proposed OUTDSM works together and achieve high accuracy rated optimal clusters. Additionally, the nature of biological diversity of OUTDSM prevents the population from stagnating at any local optima and promises to arrive at global optima. The experimental results proving this claim are given in this paper.
References
Deepankar Bharadwaj, “Text Mining Technique using Genetic Algorithm”,
International Conference on Advances in Computer Application (ICACA - 2013),
pp:7-10, 2013.
Rashmi Agrawal, Mridula Batra, “A Detailed Study on Text Mining Techniques”,
International Journal of Soft Computing and Engineering (IJSCE) ISSN: 2231-
, Volume-2, Issue-6, January 2013.
V.V.R. Maheswara Rao, Dr. V. Valli Kumari, “An Intelligent Optimal Genetic
Model to Investigate the User Usage Behaviour on World Wide Web”, International
Journal of Data Mining & Knowledge Management Process Vol.3, No.2, 2013.
Divya Nasa, “Text Mining Techniques- A Survey”, International Journal of
Advanced Research in Computer Science and Software Engineering, Vol. 2, 2012.
Dr. A.V. Senthil Kumar, S.Mythili, “Parallel Implementation of Genetic Algorithm
using K-Means Clustering”, Int. J. Advanced Networking and Applications,
Volume:03 Issue:06 Pages:1450-1455, 2012.
K.Arun Prabha a, R.Saranya, “Refinement of K-Means Clustering Using Genetic
Algorithm”, Journal of Computer Applications (JCA) ISSN: 0974-1925, Volume
IV, Issue 2, pp: 40-44, 2011.
N. El-Bathy, C. Gloster, I. Kateeb, G. Stein, “Intelligent Extended Clustering
Genetic Algorithm for Information Retrieval Using BPEL”, American Journal of
Intelligent Systems, Vol 1(1): pp: 10-15, 2011.
Dharmendra K Roy and Lokesh K Sharma, “ Genetic k-Means Clustering
Algorithm for Mixed Numeric and Categoorical Data Sets”, International Journal of
Artificial Intelligence & Applications ( IJAIA), Vol.1, No.2, pp:23-28, 2010.
Vidhya. K. A & G. Aghila,” Text Mining Process, Techniques and Tools : An
Overview”, International Journal of Information Technology and Knowledge
Management, Volume 2, No. 2, pp. 613-622, 2010.
Kuan C. Chen, “Text Mining e-Complaints Data from e-Auction Store with
Implications for Internet Marketing Research”, Journal of Business & Economics
Research, Volume 7, Number 5, pp: 15-24, 2009.
Mahesh T R, Suresh M B, M Vinayababu, “Text Mining: Advancements,
Challenges and Future Directions”, International Journal of Reviews in Computing,
pp: 61-65, ISSN: 2076-3328, 2009.
Anna Huang, “Similarity Measures for Text Document Clustering”, New Zealand
Computer Science Research Student Conference, pp:49-56, 2008.
Milos Radovanovic, Mirjana Ivanovic “Text mining: Approaches and
Applications”, Novi Sad J. Math, Vol. 38, No. 3, pp: 227-234, 2008.
Anna Stavrianou, Periklis Andritsos, Nicolas Nicoloyannis, “Overview and
Semantic Issues of Text Mining”, SIGMOD, Vol. 36, No. 3, pp: 23-34, 2007.1. Deepankar Bharadwaj, “Text Mining Technique using Genetic Algorithm”,
International Conference on Advances in Computer Application (ICACA - 2013),
pp:7-10, 2013.
Rashmi Agrawal, Mridula Batra, “A Detailed Study on Text Mining Techniques”,
International Journal of Soft Computing and Engineering (IJSCE) ISSN: 2231-
, Volume-2, Issue-6, January 2013.
V.V.R. Maheswara Rao, Dr. V. Valli Kumari, “An Intelligent Optimal Genetic
Model to Investigate the User Usage Behaviour on World Wide Web”, International
Journal of Data Mining & Knowledge Management Process Vol.3, No.2, 2013.
Divya Nasa, “Text Mining Techniques- A Survey”, International Journal of
Advanced Research in Computer Science and Software Engineering, Vol. 2, 2012.
Dr. A.V. Senthil Kumar, S.Mythili, “Parallel Implementation of Genetic Algorithm
using K-Means Clustering”, Int. J. Advanced Networking and Applications,
Volume:03 Issue:06 Pages:1450-1455, 2012.
K.Arun Prabha a, R.Saranya, “Refinement of K-Means Clustering Using Genetic
Algorithm”, Journal of Computer Applications (JCA) ISSN: 0974-1925, Volume
IV, Issue 2, pp: 40-44, 2011.
N. El-Bathy, C. Gloster, I. Kateeb, G. Stein, “Intelligent Extended Clustering
Genetic Algorithm for Information Retrieval Using BPEL”, American Journal of
Intelligent Systems, Vol 1(1): pp: 10-15, 2011.
Dharmendra K Roy and Lokesh K Sharma, “ Genetic k-Means Clustering
Algorithm for Mixed Numeric and Categoorical Data Sets”, International Journal of
Artificial Intelligence & Applications ( IJAIA), Vol.1, No.2, pp:23-28, 2010.
Vidhya. K. A & G. Aghila,” Text Mining Process, Techniques and Tools : An
Overview”, International Journal of Information Technology and Knowledge
Management, Volume 2, No. 2, pp. 613-622, 2010.
Kuan C. Chen, “Text Mining e-Complaints Data from e-Auction Store with
Implications for Internet Marketing Research”, Journal of Business & Economics
Research, Volume 7, Number 5, pp: 15-24, 2009.
Mahesh T R, Suresh M B, M Vinayababu, “Text Mining: Advancements,
Challenges and Future Directions”, International Journal of Reviews in Computing,
pp: 61-65, ISSN: 2076-3328, 2009.
Anna Huang, “Similarity Measures for Text Document Clustering”, New Zealand
Computer Science Research Student Conference, pp:49-56, 2008.
Milos Radovanovic, Mirjana Ivanovic “Text mining: Approaches and
Applications”, Novi Sad J. Math, Vol. 38, No. 3, pp: 227-234, 2008.
Anna Stavrianou, Periklis Andritsos, Nicolas Nicoloyannis, “Overview and
Semantic Issues of Text Mining”, SIGMOD, Vol. 36, No. 3, pp: 23-34, 2007.
Lipika Dey, Muhammad Abulaish, Jahiruddin, Gaurav Sharma, “Text Mining through Entity-Relationship Based Information Extraction”, IEEE/WIC/ACM International Conferences on Web Intelligence and Intelligent Agent Technology – Workshops, pp:177-180, 2007.
M. Castellano, G. Mastronardi, A. Aprile, and G. Tarricone, “A Web Text Mining Flexible Architecture”, World Academy of Science, Engineering and Technology, pp: 78-85, 2007.
Elizabeth Leon, Olfa Nasraoui, Jonatan Gomez, “ECSAGO: Evolutionary Clustering with Self Adaptive Genetic Operators”, IEEE Congress on Evolutionary Computation Sheraton Vancouver Wall Centre Hotel, pp: 1768-1175, 2006.
Louis A., Francis, FCAS, MAAA, “Taming Text: An introduction to Text mining”, Casualty Actuarial Society Forum, pp: 51-88, 2006.
Joel D. Martin, “Fast and Furious Text Mining”, and IEEE Computer Society Technical Committee on Data Engineering, pp: 1-10, 2005.
Downloads
Published
Issue
Section
License
Copyright (c) 2013 International Journal of Computer Science and Engineering Research and Development (IJCSERD)

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.




