Top News
Decomposing Language Models Into Understandable Components
Anthropic published a paper that details a new approach to understanding the complex behaviors of language models, by decomposing them into more understandable components called features. These features, unlike individual neurons, have consistent relationships to network behavior and repr…



