Publication: Tree-based text stream clustering with application to spam mail classification
Issued Date
2018-01-01
Resource Type
ISSN
17591171
17591163
17591163
Other identifier(s)
2-s2.0-85054534251
Rights
Mahidol University
Rights Holder(s)
SCOPUS
Bibliographic Citation
International Journal of Data Mining, Modelling and Management. Vol.10, No.4 (2018), 353-370
Suggested Citation
Phimphaka Taninpong, Sudsanguan Ngamsuriyaroj Tree-based text stream clustering with application to spam mail classification. International Journal of Data Mining, Modelling and Management. Vol.10, No.4 (2018), 353-370. doi:10.1504/IJDMMM.2018.095354 Retrieved from: https://repository.li.mahidol.ac.th/handle/20.500.14594/45390
Research Projects
Organizational Units
Authors
Journal Issue
Thesis
Title
Tree-based text stream clustering with application to spam mail classification
Author(s)
Other Contributor(s)
Abstract
Copyright © 2018 Inderscience Enterprises Ltd. This paper proposes a new text clustering algorithm based on a tree structure. The main idea of the clustering algorithm is a sub-tree at a specific node represents a document cluster. Our clustering algorithm is a single pass scanning algorithm which traverses down the tree to search for all clusters without having to predefine the number of clusters. Thus, it fits our objectives to produce document clusters having high cohesion, and to keep the minimum number of clusters. Moreover, an incremental learning process will perform after a new document is inserted into the tree, and the clusters will be rebuilt to accommodate the new information. In addition, we applied the proposed clustering algorithm to spam mail classification and the experimental results show that tree-based text clustering spam filter gives higher accuracy and specificity than the cobweb clustering, naïve Bayes and KNN.