Located but Not Releasable: Silent Gate Inversion and Bounded Linear Release
Di cosa parla
Hanno testato se una conoscenza che un modello linguistico conserva ma non usa può essere trovata e trasformata in comportamento osservabile. In un esperimento preregistrato su un modello addestrato a capire quando una prova indica una causa, hanno cercato il punto interno dove quella conoscenza è immagazzinata, hanno provato a modificarla per farla emergere e hanno introdotto un meccanismo che decide quando intervenire. Risultato: trovare e stimolare la rappresentazione funziona e ristabilisce il comportamento atteso; però il controllore sbaglia contesti e rende il sistema inoperoso, e una correzione semplice applicata sempre migliora ma si blocca prima di raggiungere la piena efficacia.
Cosa permette di osservare
Permette di esplorare se individuare una conoscenza interna basta a cambiarne l’uso reale, quanto sia affidabile un meccanismo che decide quando intervenire e se soluzioni semplici e fisse possono davvero sbloccare comportamenti non usati dal modello.
Dalla fonte
A growing body of work reports that language models represent task-relevant latent structure that they fail to use. Whether such structure, once located, can be converted into behavior is a separate question that is rarely tested end to end. We submit the complete pipeline -- detect, localize, and release -- to a fully preregistered stress test on a 25.7M transformer trained on causal-evidence discrimination, where a known suppression phenomenon (latent causal structure present but behaviorally unused) has previously been documented. Every threshold, claim template, and decision-tree branch was hashed and archived before any corresponding data existed. Three findings. (i) Localization succeeds: interventions at observation-evidence channels of mid layers restore target behavior on otherwise-suppressed worlds (paired release advantages $0.563$ and $0.854$, 97.5% CIs excluding zero; best-…