
A technical note on long-context training and positional extrapolation for large language models, covering techniques such as rotary position embedding scaling and attention modifications used to extend context windows beyond their original training length.
