Restricting What Each Module Can See Beats Full Visibility by 20+ Points in 9 of 10 Matched LLM Society Pairs
Four-cell societies share one frozen pretrained model and one low-rank adapter, communicating only through two model-width continuous vectors in a fixed relay, with ten matched restricted/global pairs identical in initialization bytes, training order, token layout, parameters, and computation so only the attention mask differs. Restricted societies beat their globally visible twins by at least 20 points at both composition depths in 9 of 10 pairs, with median paired advantages of 0.7648 and 0.6050, and the depth-three advantage holds at 0.558 on composite functions never seen in training. The authors report that their own preregistered battery formally fails because restricted-arm median depth-three accuracy is 0.6988, just under the 0.70 floor they set in advance.
↳ Follow the thread