The report is unusually detailed on the parts that are normally omitted — data curation, the annealing schedule, and the infrastructure failure modes encountered across a 16k-GPU run. For anyone reasoning about what frontier training actually costs, that operational detail is worth more than the benchmark tables.
Llama 3.1 405B released with open weights
Meta published the Llama 3 technical report alongside open weights for models up to 405B parameters. The significance for general-purpose work is distributional rather than architectural: capability near the contemporary frontier became something a research group could run and modify locally rather than only query through an API.
reactions
no reactions yet
react on GitHub