Источник
ACL
Дата публикации
27.07.2025
Авторы
Naquee Rizwan Seid Muhie Yimam Дарина Дементьева Florian Skupin Tim Fischer Даниил Московский Aarushi Ajay Borkar Robert Geislinger Punyajoy Saha Sarthak Roy Martin Semmann Александр Панченко Chris Biemann Animesh Mukherjee
Поделиться

HatePRISM: Policies, Platforms, and Research Integration. Advancing NLP for Hate Speech Proactive Mitigation

Аннотация

Despite regulations imposed by nations and social media platforms (Government of India, 2021; European Parliament and Council of the European Union, 2022), hateful content persists as a significant challenge. Existing ap proaches primarily rely on reactive measures such as blocking or suspending offensive messages, with emerging strategies focusing on proactive measurements like detoxification and counterspeech. In this work, we conduct a comprehensive examination of hate speech regulations and strategies from multiple perspectives: country regulations, social platform policies, and NLP research datasets. Our findings reveal significant inconsistencies in hate speech definitions and moderation practices across jurisdictions and platforms, alongside a lack of alignment with research efforts. Based on these insights, we suggest ideas and research direction for further exploration of a unified framework for automated hate speech moderation incorporating diverse strategies.

Присоединяйтесь к AIRI в соцсетях