Npbstats
Add a review FollowOverview
-
Founded Date September 28, 1949
-
Sectors Private Duty Nurse
-
Posted Jobs 0
-
Viewed 13
Company Description
DeepSeek’s First-generation Reasoning Models
DeepSeek’s first-generation reasoning designs, attaining performance equivalent to OpenAI-o1 throughout mathematics, code, and reasoning tasks.
Models
DeepSeek-R1
Distilled models
DeepSeek group has actually demonstrated that the reasoning patterns of larger designs can be distilled into smaller sized models, leading to better efficiency compared to the reasoning patterns found through RL on small models.
Below are the designs produced via fine-tuning against a number of thick designs widely utilized in the research study community using generated by DeepSeek-R1. The examination results demonstrate that the distilled smaller sized thick models carry out incredibly well on standards.
DeepSeek-R1-Distill-Qwen-1.5 B
DeepSeek-R1-Distill-Qwen-7B
DeepSeek-R1-Distill-Llama-8B
DeepSeek-R1-Distill-Qwen-14B

DeepSeek-R1-Distill-Qwen-32B

DeepSeek-R1-Distill-Llama-70B
License
The model weights are licensed under the MIT License. DeepSeek-R1 series support industrial usage, allow for any adjustments and acquired works, including, however not limited to, distillation for training other LLMs.
