Accepted · EMNLP 2026 BabyLM WorkshopFirst author
Halved CLM exposure: keeping small recurrent models from losing grammar late in training
- Question
- Why do small recurrent language models get worse at grammar late in training?
- Model
- RWKV-7, a 27.4M-parameter O(T) recurrent model, trained locally on a MacBook in MLX
- Result
- Halving how often the model is updated curbs a 2.6 to 3.9 point late drop in grammar score
- Benchmark
- Beats the official GPT-2 baseline on BLiMP with about 28% of its parameters
Bar length = parameter count
