English Benchmark Hate Speech and Offensive Content Detection Based on Layer-by-Layer Classification
DOI:
https://doi.org/10.65563/jeaai.v2i3.104Keywords:
Text Content Analysis, Text Classification, Machine Learning (ML), LSTM, DistilBERT, TF-IDFAbstract
The ability to understand English text snippets content is an essential component of human-like artificial intelligence, as English text snippets content greatly influence human cognition, decision making, and social interactions. In addition to intention recognition in diffusion of harmful content on online web, the task of identifying the potential hate and offensive categories behind an individual’s text state in English text snippets is of great importance in many application scenarios. The main content
of our research is content classification and labeling from English text snippets of hate and offensive language in various social media comments content, which aims at assign a content classification label to English text snippets taken from social media content. Each snippet contains 1–3 sentences and corresponds to one of the predefined categories based on the English text comments. Our main task is Subtask 1 divided into two hierarchical subtasks on the English HASOC2021 dataset. Subtask 1A:
Identifying hate, offensive and profane content from the post; and Subtask 1B: Discrimination between hate, profane and offensive posts; We used various machine learning (ML) and deep learning (DL) techniques and their results were compared. Subsequently, the pre-trained language model DistilBERT was used to identify hate, offensive and profane content in the posts. This is done in order to understand which prediction model framework is the most effective when predicting hate and offensive categories. We obtained the best prediction model through the layer-by-layer classification method in the Macro F1 scores of Subtask 1A and Subtask 1B were 76.00% and 58.80% respectively. The project code is available from https://github.com/WangKongQiang/DHOW.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Kongqiang Wang, Qingli Tan, Peng Zhang

This work is licensed under a Creative Commons Attribution 4.0 International License.