A CANONICAL ANSWER · FROM THE VINDEX KNOWLEDGE GRAPH
Small decisive weights stay resident
The control class — token embeddings, normalisation weights, the LM head, and the routers — rides every token, so it is always resident and preserved at source precision, never approximated. Shared feed-forward weights stay resident too. The routed expert banks are the opposite case: paged, trimmed, or left on disk as a profile decides. The split rule cuts exactly where those decisions differ.