Safety of Latent Communication in Multi-Agent Systems (Q7361999)

From MaRDI portal

!

This is the item page for this Wikibase entity, intended for internal use and editing purposes. Please use the normal view instead:

No description defined
Language Label Description Also known as
default for all languages
No label defined
    English
    Safety of Latent Communication in Multi-Agent Systems
    No description defined

      Statements

      Safety of Latent Communication in Multi-Agent Systems (English)
      0 references
      30 September 2026
      0 references
      cs.AI
      0 references
      cs.LG
      0 references
      cs.MA
      0 references
      Muhammad Huzaifa
      0 references
      Sina Mavali
      0 references
      Thorsten Eisenhofer
      0 references
      We study latent communication in multi-agent systems and find that benignly trained links can increase harmful compliance compared to text-based communication. A reinforcement-learning attack that rewards both harmless response accuracy and harmful compliance further amplifies this effect, raising the mean harmful-compliance score from 27.9 to 76.9 across three topologies and four safety benchmarks.
      0 references
      Latent Communication
      0 references
      Safety Alignment
      0 references
      Multi-Agent Systems
      0 references

      Identifiers

      0 references