Forbes reports that Anthropic and AE Studio are testing, a method called Gradient Routed Auxiliary Modules that tries to corral dangerous knowledge, like weapon-making instructions, into isolated modules during initial LLM training rather than letting it diffuse through the entire model.
The idea is that once isolated, that knowledge could be switched off after the fact, unlike today's monolithic models where harmful content is smeared across the whole numeric structure and hard to excise.
The article flags real unresolved problems: whether this scales past small research models, whether isolating knowledge fragments the AI's coherence, and the uncomfortable possibility that concentrating dangerous data into one module makes it an easier, not harder, target for extraction.
The full dispatch is available from the source below.