Hi authors,
Thank you for the great work!
I have a question regarding the SFT configuration used in the experiments. Appendix C.4 states that SFT uses a maximum sequence length of 2,048 tokens with sequence packing enabled.
For DeepMath-103K in particular, a substantial number of examples contain long chain-of-thought responses that exceed 2,048 tokens. Could you provide the exact sequence-length and packing configuration used for both DeepMath-103K and OpenCodeInstruct?
It would also be helpful to understand how overlength examples were handled and what packing-related settings were used during training. If available, any additional details on the SFT training configuration would be greatly appreciated for reproduction purposes.
Thank you.
Hi authors,
Thank you for the great work!
I have a question regarding the SFT configuration used in the experiments. Appendix C.4 states that SFT uses a maximum sequence length of 2,048 tokens with sequence packing enabled.
For DeepMath-103K in particular, a substantial number of examples contain long chain-of-thought responses that exceed 2,048 tokens. Could you provide the exact sequence-length and packing configuration used for both DeepMath-103K and OpenCodeInstruct?
It would also be helpful to understand how overlength examples were handled and what packing-related settings were used during training. If available, any additional details on the SFT training configuration would be greatly appreciated for reproduction purposes.
Thank you.