you are viewing a single comment's thread
view the rest of the comments
[–] 11 points 2 years ago (1 child)

Seems like the model you mentioned is more like a fine tuned Llama?

Specifically, these are fine-tuned versions of Qwen and Llama, on a dataset of 800k samples generated by DeepSeek R1.

https://github.com/Emericen/deepseek-r1-distilled

  • source
  • parent
  • hideshow 2 child comments