Comparing Human-Centered Language Modeling: Is it Better to Model Groups, Individual Traits, or Both?
Abstract: Natural language processing has made progress in incorporating human context into its models, but whether it is more effective to use group-wise attributes (e.g., \textit{over-45-year-olds}) or model individuals remains open. Group attributes are technically easier but coarse: not all 45-year-olds write the same way. In contrast, modeling individuals captures the complexity of each person's identity. It allows for a more personalized representation, but we may have to model an infinite number of users and require data that may be impossible to get.
We compare modeling human context via group attributes, individual users, and combined approaches. Combining group and individual features significantly benefits \textit{user}-level regression tasks like age estimation or personality assessment from a user's documents. Modeling individual users significantly improves the performance of single \textit{document}-level classification tasks like stance and topic detection. We also find that individual-user modeling does well even without user's historical data.
Paper Type: long
Research Area: Computational Social Science and Cultural Analytics
Contribution Types: Model analysis & interpretability
Languages Studied: English
0 Replies
Loading