Inside how AI alignment really rewires models
Researchers found that preference tuning reorganizes model weights in a structured way, with a compact 'head' driving visible behavior change and a subtle 'tail' needed for full coverage.
Researchers found that preference tuning reorganizes model weights in a structured way, with a compact 'head' driving visible behavior change and a subtle 'tail' needed for full coverage.