Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems (Q7284615)

From MaRDI portal

!

This is the item page for this Wikibase entity, intended for internal use and editing purposes. Please use the normal view instead:

No description defined
Language Label Description Also known as
default for all languages
No label defined
    English
    Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems
    No description defined

      Statements

      Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems (English)
      0 references
      28 July 2026
      0 references
      cs.AI
      0 references
      Marylou Fauchard
      0 references
      Florian Carichon
      0 references
      Margarida Carvalho
      0 references
      Golnoosh Farnadi
      0 references
      The paper introduces a novel framework to evaluate objective misalignment in large language model (LLM) multi-agent systems using the social deduction game Werewolf and finds that even subtle objective misalignment can significantly impact collective decision-making outcomes, with compromised agents developing distinct reasoning strategies but keeping them hidden from public behavior.
      0 references
      Objective Misalignment
      0 references
      Multi-Agent Systems
      0 references
      Large Language Models
      0 references
      Social Influence
      0 references
      Deception
      0 references

      Identifiers

      0 references