Table based single pass algorithm for clustering electronic documents in 20NewsGroups

  • JO GEUN SIK

초록

This research proposes a modified version of single pass algorithm which is specialized for text clustering. Encoding documents into numerical vectors for using the traditional version of single pass algorithm causes the two main problems: huge dimensionality and sparse distribution. Therefore, in order to address the two problems, this research modifies the single pass algorithm into its version where documents are encoded into not numerical vectors but alternative forms. In the proposed version, documents are mapped into tables and a similarity of two documents is computed by comparing their tables with each other. The goal of this research is to improve the performance of single pass algorithm for text clustering by modifying it into the specialized version. ? 2008 IEEE.

제목
Table based single pass algorithm for clustering electronic documents in 20NewsGroups
저자
JO GEUN SIK
학회명
Proceedings -IEEE International Workshop on Semantic Computing and Applications 2008
개최지
인천
학회 개최일
2008-07-10 ~ 2008-07-11