🧠 Qwen3-1.7B-Thinking-Distil

Extended reasoning distilled from Qwen3-30B-A3B-Thinking into a 1.7B student (~2.03B params, BF16, 40K context) via SFT on longwriter-6k. The model deliberates before answering — its chain of thought streams into the collapsible Reasoning panel, with the final answer below it.

Model card · Apache-2.0 · running on ZeroGPU

256 4096
0 1.5
0.05 1
1 100
1 1.5