AfriSenti: A Twitter Sentiment  Analysis Benchmark for African Languages

Shamsuddeen Hassan Muhammad; Idris Abdulmumin; Abinew Ali Ayele; Nedjma OUSIDHOUM; David Ifeoluwa Adelani; Seid Muhie Yimam; Ibrahim Said Ahmad; Meriem Beloucif; Saif M. Mohammad; Sebastian Ruder; Oumaima Hourrane; Alipio Jorge; Pavel Brazdil; Felermino D. M. A. Ali; Davis David; Salomey Osei; Bello Shehu-Bello; Falalu Ibrahim Lawan; Tajuddeen Gwadabe; Samuel Rutunda; Tadesse Destaw Belay; Wendimu Baye Messelle; Hailu Beshada Balcha; Sisay Adugna Chala; Hagos Tesfahun Gebremichael; Bernard Opoku; Stephen Arthur

AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages

Published: 07 Oct 2023, Last Modified: 01 Dec 2023EMNLP 2023 MainEveryoneRevisionsBibTeX

Submission Type: Regular Long Paper

Submission Track: Resources and Evaluation

Submission Track 2: Sentiment Analysis, Stylistic Analysis, and Argument Mining

Keywords: Africa, Sentiment, Dataset, NLP

Abstract: Africa is home to over 2,000 languages from over six language families and has the highest linguistic diversity among all continents. This includes 75 languages with at least one million speakers each. Yet, there is little NLP research conducted on African languages. Crucial in enabling such research is the availability of high-quality annotated datasets. In this paper, we introduce AfriSenti, a sentiment analysis benchmark that contains a total of >110,000 tweets in 14 African languages (Amharic, Algerian Arabic, Hausa, Igbo, Kinyarwanda, Moroccan Arabic, Mozambican Portuguese, Nigerian Pidgin, Oromo, Swahili, Tigrinya, Twi, Xitsonga, and Yoruba) from four language families. The tweets were annotated by native speakers and used in the AfriSenti-SemEval shared task (with over 200 participants, see website: https://afrisenti-semeval.github.io). We describe the data collection methodology, annotation process, and the challenges we dealt with when curating each dataset. We further report baseline experiments conducted on the AfriSenti datasets and discuss their usefulness.

Submission Number: 11

Loading