Current alignment techniques, such as RLHF (Reinforcement Learning from Human Feedback), typically rely on a centralized set of "safety guidelines" that reflect the cultural norms of the model's developers. This often leads to cultural misalignment, where the LLM may inadvertently erase local nuances, apply inappropriate social taboos, or fail to recognize regional linguistic etiquette.
This thesis aims to move beyond "cultural awareness" toward Cultural Verification. The student will design a systematic framework to evaluate and verify whether an LLM’s outputs adhere to the specific socio-cultural norms, value systems, and historical contexts of a target demographic.
Key Responsibilities & Tasks:
The research will be structured around four primary pillars, beginning with the development of a comprehensive Taxonomy of Cultural Correctness to define the specific dimensions—such as social hierarchy, religious sensitivities, and gender roles—necessary for nuanced LLM evaluation. Building upon this foundation, the study will propose a Verification Methodology centered on a "Cultural Guardrail" architecture, utilizing both programmatic checks and human-in-the-loop oversight to validate model responses. To test the efficacy of this framework, a specialized Benchmark will be curated, featuring a localized "red-teaming" suite designed to expose cultural hallucinations and latent insensitivities. Finally, a Comparative Analysis will be conducted to measure the performance of leading models, including GPT-4, Llama 3, and Claude, against these proposed designs across two or more distinct cultural contexts.
Requirements:
- Experience with LLMs, HuggingFace
- LLM fine-tuning
- Python
- Knowledge of German culture and language is a plus.
Related Work:
- Atari et al. (2023): "Which Humans? Probabilistic Expectations and Social Values in LLMs."
- Wang, Jiahao, Songkai Xue, Jinghui Li, and Xiaozhen Wang. "Diverse Human Value Alignment for Large Language Models via Ethical Reasoning." In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, vol. 8, no. 3, pp. 2637-2648. 2025.
- Venktesh, V., Mandeep Rathee, and Avishek Anand. "Trust but verify! a survey on verification design for test-time scaling." arXiv preprint arXiv:2508.16665 (2025).
Project/Thesis language:
English
Contact: