AIbase
Product LibraryTool Navigation

ChainForge-R1-SuperCoT

Public

A multi-stage pipeline that enhances Qwen2.5 language models with DeepSeek Reasoner's chain-of-thought capabilities. Implements the DeepSeek-R1 methodology through cold-start SFT, reasoning-oriented RL, rejection sampling, and optional model distillation.

Creat2025-01-25T03:13:53
Update2025-02-24T17:02:19
9
Stars
0
Stars Increase

Related projects