Self-Switching Classification Framework for Titled Documents

Guo, Hang; Zhou, Li-Zhu; Feng, Ling

doi:10.1007/s11390-009-9262-z

Self-Switching Classification Framework for Titled Documents

Regular Paper
Published: 11 September 2009

Volume 24, pages 615–625, (2009)
Cite this article

Journal of Computer Science and Technology Aims and scope Submit manuscript

Hang Guo¹,
Li-Zhu Zhou² &
Ling Feng²

44 Accesses
2 Citations
Explore all metrics

Abstract

Ambiguous words refer to words that have multiple meanings such as apple, window. In text classification they are usually removed by feature reduction methods like Information Gain. Sometimes there are too many ambiguous words in the corpus, which makes throwing away all of them not a viable option, as in the case when classifying documents from the Web. In this paper we look for a method to classify Titled documents with the help of ambiguous words. Titled documents are a kind of documents that have a simple structure containing a title and an excerpt. News, messages, and paper abstracts with titles are examples of titled documents. Instead of introducing another feature reduction method, we describe a framework to make the best use of ambiguous words in the titled documents. The framework improves the performance of a traditional bag-of-words classifier with the help of a bag-of-word-pairs classifier. The framework is implemented using one of the most popular classifiers, Multinomial NaiveBayes (MNB) as an example. The experiments with three real life datasets show that in our framework the MNB model performs much better than traditional MNB classifier and a naive weighted algorithm, which simply puts more weight on words in the title.

This is a preview of subscription content, log in via an institution to check access.

Access this article

Log in via an institution

Price excludes VAT (USA)
Tax calculation will be finalised during checkout.

Instant access to the full article PDF.

Institutional subscriptions

Bayesian Multinomial Naïve Bayes Classifier to Text Classification

Evaluation of the Document Classification Approaches

Using Discriminative Phrases for Text Categorization

References

McCallum A, Nigam K. A comparison of event models for naive Bayes text classification. In Proc. AAAI Workshop on Learning for Text Categorization, Madison, Wisconsin, USA, July 26–27, 1998, pp.41–48.
Joachims T. Text categorization with support vector machines: Learning with many relevant features. In Proc. 10th Euro. Conf. Machine Learning, Chemnitz, Germany, April 21–23, 1998, pp.137–142.
Keerthi S, Shevade S, Bhattacharyya C, Murthy K. Improvements to Platt's SMO algorithm for SVM classifier design. Neural Computation, 2001, 13: 637–649.
Article MATH Google Scholar
Caropreso M, Matwin S, Sebastiani F. A Learner Independent Evaluation of the Usefulness of Statistical Phrases for Automated Text Categorization. Text Databases and Document Management: Theory and Practice, IGI Publishing, 2001, pp.78–102.
Larkey S. Automatic essay grading using text categorization techniques. In Proc. the 21st Int. ACM. SIG Conf. Info. Retrieval, Melbourne, Australia, August 24–28, 1998, pp.90–95.
Sebastiani F. Machine learning in automated text categorization. ACM Comput. Surv., 2002, 34(1): 1–47.
Article Google Scholar
Jin R, Hauptmann A G, Zhai C. Title language model for information retrieval. In Proc. the 25th Int. ACM SIG. Conf. Info. Retrieval, Tampere, Finland, 2002, pp.42–48.
Clark J, Koprinska I, Poon J. A neural network based approach to automated e-mail classification. In Proc. Int. Conf. IEEE/WIC, Halifax, Canada, October 13–16, 2003, pp.702–705.
Schütze H, Hull D, Pedersen J. A comparison of classifiers and document representations for the routing problem. In Proc. the 18th Int. ACM SIG. Conf. Info. Retrieval, Seattle, Washington, USA, July 9–13, 1995, pp.229–237.
Tzeras K, Hartmann S. Automatic indexing based on Bayesian inference networks. In Proc. the 16th Int. ACM SIG Conf. Info. Retrieval, Pittsburgh, USA, June 27–July 1, 1993, pp.22–34.
Lewis D. An evaluation of phrasal and clustered representations on text categorization task. In Proc. the 15th Int. ACM SIG. Conf. Info. Retrieval, Copenhagen, Denmark, June 21–24, 1992, pp.37–50.
Dumais T, Platt J, Heckerman D, Sahami M. Inductive learning algorithms and representations for text categorization. In Proc. the 7th Int. Conf. Information and Knowledge Management, Bethesda, Spain, November 3–7, 1998, pp.148–155.
Fung B, Wang K, Ester M. Large hierarchical document clustering using frequent itemsets. In Proc. the 3rd Int. Conf. Data Mining, Melbourne, Florida, USA, November 19–22, 2003, pp.59–70.
Beil F, Ester M, Xu X. Frequent term-based text clustering. In Proc. 8th Int. Conf. Knowledge Discovery and Data Mining, Alberta, Canada, July 23–26, 2002, pp.436–442.
Li Y, Chung S, Holt J. Text document clustering based on frequent word meaning sequences. Data & Knowledge Engineering, 2008, 64(1): 381–404.
Article Google Scholar
Slonim N, Tishby N. The power of word clusters for text classification. In Proc. the 23rd European Colloquium on Information Retrieval Research, Darmstadt, Germany, April 2001, pp.1–11.
Lin J. Divergence measures based on the Shannon entropy. IEEE Trans. Info. Theory, 1991, 37(1): 145–151.
Article MATH Google Scholar
Yang Y, Liu X. A re-examination of text categorization methods. In Proc. the 23rd Int. ACM SIG. Conf. Info. Retrieval, Berkeley, CA, USA, August 15–19, 1999, pp.42–49.

Download references

Author information

Authors and Affiliations

EMC Research China, Tsinghua Science Park, Beijing, 100084, China
Hang Guo
Department of Computer Science and Technology, Tsinghua University, Beijing, 100084, China
Li-Zhu Zhou (Member, ACM) & Ling Feng (Member, CCF, ACM, IEEE)

Authors

Hang Guo
View author publications
You can also search for this author in PubMed Google Scholar
Li-Zhu Zhou
View author publications
You can also search for this author in PubMed Google Scholar
Ling Feng
View author publications
You can also search for this author in PubMed Google Scholar

Corresponding author

Correspondence to Hang Guo.

Additional information

This work was done when the first author was studying in Tsinghua University, China. It is supported by the National Natural Science Foundation of China under Grant Nos. 60833003 and 60773156.

Electronic Supplementary Material

Below is the link to the electronic supplementary material.

(PDF 80.2 kb)

Rights and permissions

Reprints and permissions

About this article

Cite this article

Guo, H., Zhou, LZ. & Feng, L. Self-Switching Classification Framework for Titled Documents. J. Comput. Sci. Technol. 24, 615–625 (2009). https://doi.org/10.1007/s11390-009-9262-z

Download citation

Revised: 26 February 2009
Published: 11 September 2009
Issue Date: July 2009
DOI: https://doi.org/10.1007/s11390-009-9262-z

Keywords

Access this article

Log in via an institution

Price excludes VAT (USA)
Tax calculation will be finalised during checkout.

Instant access to the full article PDF.

Institutional subscriptions

Self-Switching Classification Framework for Titled Documents

Abstract

Access this article

Similar content being viewed by others

Bayesian Multinomial Naïve Bayes Classifier to Text Classification

Evaluation of the Document Classification Approaches

Using Discriminative Phrases for Text Categorization

References

Author information

Authors and Affiliations

Corresponding author

Additional information

Electronic Supplementary Material

(PDF 80.2 kb)

Rights and permissions

About this article

Cite this article

Keywords

Navigation

Self-Switching Classification Framework for Titled Documents

Abstract

Access this article

Similar content being viewed by others

Bayesian Multinomial Naïve Bayes Classifier to Text Classification

Evaluation of the Document Classification Approaches

Using Discriminative Phrases for Text Categorization

References

Author information

Authors and Affiliations

Corresponding author

Additional information

Electronic Supplementary Material

(PDF 80.2 kb)

Rights and permissions

About this article

Cite this article

Share this article

Keywords

Search

Navigation