Skip to content

SFT sequence length and training configuration #21

Description

@Heinz217

Hi authors,

Thank you for the great work!

I have a question regarding the SFT configuration used in the experiments. Appendix C.4 states that SFT uses a maximum sequence length of 2,048 tokens with sequence packing enabled.

For DeepMath-103K in particular, a substantial number of examples contain long chain-of-thought responses that exceed 2,048 tokens. Could you provide the exact sequence-length and packing configuration used for both DeepMath-103K and OpenCodeInstruct?

It would also be helpful to understand how overlength examples were handled and what packing-related settings were used during training. If available, any additional details on the SFT training configuration would be greatly appreciated for reproduction purposes.

Thank you.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions