Motivation

We need to understand how language models develop complex skills. Previous works on probing and evaluating language models gives evidence that LMs can learn abstractions and composition abilities grows with scale. We aim to understand further in two axis:

RQ1: How “skewed” can the training length distribution be?

RQ2: How does skill abstraction help with OOD generalization?

RQ3: Composition scaling law with compute

RQ4: Tracking skill learning on a data-point level