2026 · NLP & AI Safety · Paris-Saclay (T4) · run 20260208_200540_385119c8
LLM Biases for Political Conspiracies — Constrained Prompt Evaluation
A controlled Entity × Narrative × Template protocol for measuring how frontier LLMs handle political conspiracy-framed prompts — summarized by the Narrative Susceptibility Score (NSS), not by anecdotal jailbreaks.
What this project does
Team project at Paris-Saclay (Sukhiot, Oudoum, Yuze, Nguyen, Frederic). We built a reproducible evaluation pipeline that crosses 20 political entities with 10 sensitive narrative topics (elections, corruption, immigration, media, …) and three constrained prompt styles — Direct, Creative, and Socratic — for 600 prompts total. Model outputs are scored into an NSS that makes entity-level and style-level bias comparable across models.
- Full factorial coverage: 200 entity–narrative pairs × 3 templates
- Multi-model inference (Mistral Small, GPT-4o Mini, DeepSeek Chat, Llama 3.1 70B)
- Safety / hallucination labeling plus entity-level fairness analysis with confidence intervals
- Ablations on prompt style and key political figures
Experimental design
Results
My contributions
- Pipeline architecture and inference execution
- Data visualization for the figure suite above
- LaTeX report formatting