Pre-Training My First Base Language Model From Scratch
Took a model from random initialization through 2B tokens of FineWeb-Edu on a single H100 and watched it learn next-token…
Showing 2 of 18 posts tagged Pre-Training
Took a model from random initialization through 2B tokens of FineWeb-Edu on a single H100 and watched it learn next-token…
I didn't come close to the Parameter Golf leaderboard, but I still had a lot of fun running scattered H100 experiments on Modal…