Optimizing Against Safety Representations: Activation-Guided Adversarial Suffixes and the Geometry of Refusal

Chronological Source Flow
Back

AI Fusion Summary

Recent research examines the geometry of representations in language models. Studies on Gemma-2-2B, Gemma-2-9B, and Qwen3-4B reveal that 1D manifolds with place-cell feature tiling emerge for locally computable ordinal tasks, while complex tasks produce higher-dimensional representations. Simultaneously, another study explores behavioral alignment and safety representations. By introducing Activation-Guided GCG, researchers use adversarial suffix attacks to target internal refusal directions, discovering that suppressing refusal globally across all layers and positions is more effective than localized targeting.
Community Comments
Loading updates...
0