Is pre-training in FP8 new? Also, 10M input token context is insane! EDIT: https...

		nattaylor 5 months ago \| parent \| context \| favorite \| on: The Llama 4 herd Is pre-training in FP8 new? Also, 10M input token context is insane! EDIT: https://huggingface.co/meta-llama/Llama-3.1-405B is BF16 so yes, it seems training in FP8 is new.

Deepseek v3 was FP8