K2 Horizon 各模型预训练约使用 20T token,语料混合网页、代码、数学、科学、多语言与合成数据,其中约 17% 的预训练语料包含显式推理轨迹。IFM 称预训练阶段约用了 10T 合成 token,后训练生成 1 亿+ 唯一任务,管线通过监督微调、模型合并、强化学习和专用智能体训练发布。已开源的预训练数据集之一 TxT360-v2(5.3 TB)已上线 Hugging Face。
K2 Horizon 披露预训练用约 20T token 且 17% 含显式推理轨迹,读者可借此对照模型的训练数据构成。
The training data is another big part of K2 Horizon: each model was pretrained on roughly 20T tokens, with web, code, math, science, multilingual and synthetic data in the mixture.
Around 17% of the pre-training corpus contains explicit reasoning trajectories, and IFM says roughly 10T synthetic tokens were used during pre-training.
For post-training, IFM generated 100M+ unique tasks and is releasing the pipeline through supervised fine-tuning, model merging, reinforcement learning and specialized agent training.
One of the released pre-training datasets, TxT360-v2 (5.3 TB), is already up on Hugging Face:
https://t.co/xBl8hgGSLS
来源:@rohanpaul_ai · x.com