Decoding Market Emotions in Cryptocurrency Tweets via Predictive-Statement Classification with Machine Learning and Transformers
Abstract
Cryptocurrency discussions on social media combine emotional reactions with statements about future market behavior, but these two signals are not equivalent. This study presents a two-stage framework for identifying predictive statements in English-language cryptocurrency posts concerning Cardano (ADA), Binance Coin (BNB), Polygon (MATIC), XRP, and Fantom (FTM). Task 1 distinguishes Predictive from Non-Predictive posts; Task 2 assigns Predictive posts to Incremental, Decremental, or Neutral categories. The original corpus contains 3,116 human-reviewed posts. Two fixed, stratified 80/20 splits were created independently before augmentation, one for the binary Task 1 labels and one for the three-way Task 2 direction labels, leaving 624 original posts for Task 1 testing and 224 original predictive posts for Task 2 testing. GPT-generated paraphrases were added only to minority classes in the training partition. We compared TF-IDF-based traditional classifiers, CNN and BiLSTM models with GloVe or FastText embeddings, and five transformer checkpoints. Macro-F1 was the primary metric. Under the augmented-training condition, the XLM-RoBERTa language-identification checkpoint achieved the highest Task 1 macro-F1 (0.7011), whereas Random Forest achieved the highest Task 2 macro-F1 (0.6488), narrowly exceeding SVM-RBF (0.6478). SenticNet was used only after classification to describe non-exclusive emotion distributions; it was not used as a classifier input. Positive and negative emotions appeared across directional classes, showing why emotional polarity alone cannot reliably represent market expectations. The results support task-dependent model selection and a strict separation between directional prediction and post-hoc emotion analysis.
Keywords
Cryptocurrency, predictive statements, social media, text classification, data augmentation, emotion analysis, SenticNet