Improving the accuracy of text classification using stemming method, a case of non-formal Indonesian conversation

Rianto, Rianto and Mutiara, Achmad Benny and Wibowo, Eri Prasetyo and Santosa, Paulus Insap (2021) Improving the accuracy of text classification using stemming method, a case of non-formal Indonesian conversation. JOURNAL OF BIG DATA, 8 (1).

Full text not available from this repository. (Request a copy)

Abstract

Background: Stemming has long been used in data pre-processing to retrieve information by tracking affixed words back into their root. In an Indonesian setting, existing stemming methods have been observed, and the existing stemming methods are proven to result in high accuracy level. However, there are not many stemming methods for non-formal Indonesian text processing. This study introduces a new stemming method to solve problems in the non-formal Indonesian text data pre-processing. Furthermore, this study aims to improve the accuracy of text classifier models by strengthening stemming method. Using the Support Vector Machine algorithm, a text classifier model is developed, and its accuracy is checked. The experimental evaluation was done by testing 550 datasets in Indonesian using two different stemming methods. Findings: The results show that using the proposed stemming method, the text classifier model has higher accuracy than the existing methods with a score of 0.85 and 0.73, respectively. These results indicate that the proposed stemming methods produces a classifier model with a small error rate, so it will be more accurate to predict a class of objects. Conclusion: The existing Indonesian stemming methods are still oriented towards Indonesian formal sentences, therefore the method has limitations to be used in Indonesian non-formal sentences. This phenomenon underlies the suggestion of developing a corpus by normalizing Indonesian non-formal into formal to be used as a better stemming method. The impact of using the corpus as a stemming method is that it can improve the accuracy of the classifier model. In the future, the proposed corpus and stemming methods can be used for various purposes including text clustering, summarizing, detecting hate speech, and other text processing applications in Indonesian.

Item Type: Article
Uncontrolled Keywords: Accuracy; Classification; Indonesian; Stemming; Text processing
Subjects: T Technology > TK Electrical engineering. Electronics Nuclear engineering
Divisions: Faculty of Engineering > Electronics Engineering Department
Depositing User: Sri JUNANDI
Date Deposited: 19 Oct 2024 07:07
Last Modified: 19 Oct 2024 07:07
URI: https://ir.lib.ugm.ac.id/id/eprint/9191

Actions (login required)

View Item
View Item