Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems (Q7284615)
From MaRDI portal
!
This is the item page for this Wikibase entity, intended for internal use and editing purposes. Please use the normal view instead:
No description defined
| Language | Label | Description | Also known as |
|---|---|---|---|
| default for all languages | No label defined |
||
| English | Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems |
No description defined |
Statements
Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems (English)
0 references
28 July 2026
0 references
cs.AI
0 references
Marylou Fauchard
0 references
Florian Carichon
0 references
Margarida Carvalho
0 references
Golnoosh Farnadi
0 references
The paper introduces a novel framework to evaluate objective misalignment in large language model (LLM) multi-agent systems using the social deduction game Werewolf and finds that even subtle objective misalignment can significantly impact collective decision-making outcomes, with compromised agents developing distinct reasoning strategies but keeping them hidden from public behavior.
0 references
Objective Misalignment
0 references
Multi-Agent Systems
0 references
Large Language Models
0 references
Social Influence
0 references
Deception
0 references
1 reference
1 reference