A BERT-Based Approach for Hedges Detection in Scientific Texts with Cue-Based Labeling
DOI:
https://doi.org/10.65204/djes.v3i3.838Keywords:
Natural Language Processing (NLP) Text Classification Hedge Detection Scientific Text Analysis Deep Learning BERTAbstract
Hedging expressions play a crucial role in scientific writing by conveying uncertainty and limiting the strength of claims. Accurately detecting such expressions is essential for reliable information extraction and decision-making in domains such as biomedical research. In this study, we propose a BERT-based approach for the automatic detection of hedging in scientific texts. A domain-specific dataset is constructed from XML-formatted scientific documents, including abstracts and full-text articles, where sentences are automatically labeled based on the presence of hedge cues. The proposed model is fine-tuned for binary classification to distinguish between hedged and non-hedged sentences. By leveraging deep contextual representations, the model effectively captures the semantic and syntactic characteristics of uncertainty in text. Experimental results demonstrate strong performance, achieving an accuracy of approximately 96.9% and an F1-score of 94.7%, indicating the robustness and effectiveness of the proposed approach. The findings highlight the potential of transformer-based models in handling complex linguistic phenomena such as hedging. The focus of future study will be on extending the proposed approach into a multi-task framework that integrates additional components, including context understanding, intent classification, and attitude analysis, in order to provide a more thorough understanding of scientific discourse.