7.1 TinyStories

7.1.1 Hyperparameter tuning

Total tokens processed 327,680,000 (your batch size × total step count × context length should equal roughly this value).

  1. learning rate warmup 2%

  2. betas (0.9, 0.95)

  3. eps 1e-8

  4. weight_decay 0.1

learning rate(1e-3 5e-4 3e-4 1e-4)

 

min_lr = 0.1 * max_lr, only search max_lr

0.1

image-20260728112218920

0.02

image-20260728113119278

0.01

image-20260728112429466

 

1e-3

image-20260728103840421

image-20260728103855261

and later the gradient is less than 1 so chart didn't record it.

5e-4

image-20260728104904529

image-20260728104911915

strategy:

first try 1e-3, find it not diverge, so try bigger lr, finnaly found

260728 13:26 我们尝试使用学习率1e-3进行一次完整训练,观察其曲线

see diagrams above.

lr=1e-3

image-20260729154030263

image-20260729154146573

 

batch_size variations

image-20260731143209990

we can see that with batch_size increasing, training stability increases. and the loss of batch_size = 50 is 1.27 which is the least

Problem (generate): Generate text (1 point)

Once upon a time, there was a little bird named Timmy. Timmy was an obedient bird who always listened to his mom. One day, Timmy wanted to fly high in the sky. He told his mom, "I want to fly high, but I am too little to fly." His mom said, "Don't worry, Timmy. Just believe in yourself. I will encourage you to fly high in the sky." Timmy started to fly up, up, up. He was so happy. But then, he saw a big, shiny thing in the sky. It was a big, red balloon! Timmy was so excited. He took a deep breath and went closer to the balloon. He held it high, just like his mom said. Timmy was so happy. He thanked his mom and flew back down to the ground. From that day on, Timmy always believed in himself. He was a happy little bird who loved to fly high in the sky.

comment: I think it's fluent. factors: sampling parameters, input prompts.

Problem (layer_norm_ablation): Remove RMSNorm and train (0.5 B200 hrs) (1 point)

image-20260801084334879

Without RMSNorm, It has loss explosion.

So RMSNorm helps stability of training.

Problem (pre_norm_ablation): Implement post-norm and train (0.5 B200 hrs) (1 point)

image-20260801172850372

postnorm

 

image-20260801172925746

prenorm

we can see postnorm has slightly higher loss of evaluation curve.

and the end loss also.

Problem (no_pos_emb): Implement NoPE (0.5 B200 hrs) (1 point)

image-20260801173657925

Problem (swiglu_ablation): SwiGLU vs. SiLU (0.5 B200 hrs) (1 point)

image-20260801173748738

 

Problem (main_experiment): Experiment on OWT (2 B200 hrs) (2 points)

ski[p]

maybe more difficult. maybe it is large

Problem (leaderboard): Leaderboard (10 B200 hrs) (6 points)

skip.

image-20260801175529938

image-20260801175538814