Mapping cultural differences in word meanings using cultural knowledge graphs
In our interconnected world, effective cross-cultural communication is crucial. However, even when words have equivalent translations, their meanings and associations can vary significantly across languages and cultures. One of the consequences is that speakers from different cultures unknowingly understand similar concepts in different ways.
This discrepancy can lead to misunderstandings and miscommunications, affecting everything from international business dealings to diplomatic relations. As AI and large language models (LLMs) become more prevalent in global communications, understanding these nuanced differences is critical to ensure accurate and culturally sensitive interactions.
This project aims to quantify and map these cross-cultural differences in concept meaning, addressing a fundamental question in cross-lingual semantics. By leveraging the Small World of Words (SWOW) dataset, which contains native speakers' mental associations for over 12,000 common words in 19 languages, we can capture nuanced concept semantics from a cultural perspective.
Our research team will develop novel approaches to automatically annotate and map meaning across languages using graph neural networks and LLMs. We aim to quantify semantic equivalence for words in English, Dutch, Chinese, and Spanish, including semantic relationships and sense mapping, to highlight and explain ambiguity and differences in the cross-lingual mappings.
We will then use this data to develop a public web API, where users can explore meaning differences between two languages by inputting pairs of translationally equivalent words and examining sources of similarity and dissimilarity. This tool will benefit researchers, language learners, and the public by providing a new way to explore language and culture-specific meanings in an accessible format.
MDAP's expertise in data analytics, research methodology, App development, LLMs and natural language processing (NLP) are crucial to this project. Their input will be particularly valuable in building a scalable open-source NLP pipeline and in API development to power interactive search and visualisations. MDAP's involvement will ensure the project efficiently leverages the latest developments in LLMs and compute infrastructure.
This research will contribute to new knowledge about the relationship between language, culture, and cognition, address cultural understanding as a core frontier for language learners, and help identify variations in human and artificial language biases to improve the cultural alignment of LLMs.
Who's involved
Chief Investigator
Dr Lea Frermann, Senior Lecturer and ARC Decra Fellow, School of Computing and Information Systems, University of Melbourne
Co investigators
Dr Simon De Deyne, Senior Research Fellow, Melbourne School of Psychological Sciences, Complex Human Data Hub, University of Melbourne
Dr Chunhua Liu, Postdoctoral Research Fellow and Associate Lecturer in NLP, School of Computing and Information Systems, University of Melbourne
Prof Andrew Perfors, Director, Complex Human Data Hub, Melbourne School of Psychological Sciences, University of Melbourne
A/Prof Álvaro Cabana, Faculty of Psychology, Universidad de la República, Uruguay