跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 25 天前精选AI 评分70
AI 导读

K2 Horizon 各模型预训练约使用 20T token,语料混合网页、代码、数学、科学、多语言与合成数据,其中约 17% 的预训练语料包含显式推理轨迹。IFM 称预训练阶段约用了 10T 合成 token,后训练生成 1 亿+ 唯一任务,管线通过监督微调、模型合并、强化学习和专用智能体训练发布。已开源的预训练数据集之一 TxT360-v2(5.3 TB)已上线 Hugging Face。

推荐理由

K2 Horizon 披露预训练用约 20T token 且 17% 含显式推理轨迹,读者可借此对照模型的训练数据构成。

正文

The training data is another big part of K2 Horizon: each model was pretrained on roughly 20T tokens, with web, code, math, science, multilingual and synthetic data in the mixture.

Around 17% of the pre-training corpus contains explicit reasoning trajectories, and IFM says roughly 10T synthetic tokens were used during pre-training.

For post-training, IFM generated 100M+ unique tasks and is releasing the pipeline through supervised fine-tuning, model merging, reinforcement learning and specialized agent training.

One of the released pre-training datasets, TxT360-v2 (5.3 TB), is already up on Hugging Face:

https://t.co/xBl8hgGSLS

来源:@rohanpaul_ai · x.com