Pengaruh Tahapan Preprocessing terhadap Performa Model Indobert dan Indobertweet pada Sentimen Banjir Aceh 2025
Keywords:
IndoBERT, IndoBERTweet, Preprocessing, Sentiment Analysis, Social MediaAbstract
Indonesia is highly vulnerable to hydrometeorological disasters such as floods, which often generate public responses on social media. These responses provide valuable insights for sentiment analysis but are characterized by unstructured and noisy text. This study aims to analyze the effect of preprocessing stages on the performance of IndoBERT and IndoBERTweet models for sentiment classification of tweets related to the Aceh flood in 2025. The dataset was collected from platform X and manually labeled into three sentiment classes: positive, negative, and neutral. Three preprocessing scenarios were applied, starting from case folding as minimal preprocessing and extending to more complex stages such as cleaning, normalization, stopword removal, and stemming. The models were evaluated using accuracy, precision, recall, and F1-score with a macro approach. The results show that preprocessing affects model performance, although the differences are relatively small. IndoBERT achieved its best performance under minimal preprocessing, while IndoBERTweet showed improved performance with more complex preprocessing and achieved the highest F1-score of 0.8435. Overall, IndoBERTweet demonstrated slightly better and more stable performance compared to IndoBERT. These findings indicate that the effectiveness of preprocessing depends on the compatibility between the model and data characteristics, particularly for social media text.
