All posts / AI models · January 28, 2025 · 6 min read
Is DeepSeek Worth the Hype?
DeepSeek R1 activates only 37 billion of its 671 billion parameters for any given task. What this efficiency breakthrough actually means for businesses building with AI, and why it changes the economics entirely.
Written by SIEL AI engineering team · Published January 28, 2025

A few months ago, we wrote that the next gains in LLMs would come from smarter models, not bigger ones. DeepSeek R1 has now emerged as the perfect validation of that perspective. Their breakthrough is not just impressive. It changes how we think about efficiency.
Efficiency Meets Intelligence
At first glance, DeepSeek R1's 671 billion parameters might make it seem like just another massive model. But here is the twist: it only activates 37 billion parameters for any given task. This is made possible through its Mixture-of-Experts system. Instead of throwing the entire model at every problem, it operates like a highly efficient team of specialists, dynamically routing tasks to the most relevant expert within the model.
The result is exceptional cost efficiency. Input tokens are priced at $0.55 per million, output tokens at $2.19 per million. DeepSeek's operational costs are roughly 95% lower than OpenAI's current offerings. This is not a minor improvement. It fundamentally changes the economics of AI deployment.
Open and Accessible
What truly sets DeepSeek apart is its commitment to openness. The model's code is open-source under an MIT license. Anyone can access, modify, and build upon it. In an industry often criticised for opacity, this is a bold move that will accelerate innovation.
How They Built It
DeepSeek R1's training process is a masterclass in optimisation, reportedly achieved for under $6 million. Their four-stage approach: start from a pre-trained base, use reinforcement learning to develop reasoning skills, clean up outputs with supervised fine-tuning, then apply another RL round focused on math, logic, and coding. Finally, they distil these capabilities into compact versions ranging from 1.5B to 70B parameters, making advanced reasoning accessible on everyday hardware.
What This Means for Businesses
For startups and smaller companies this changes the arithmetic. Competitive performance at roughly 5% of the usual model cost puts work within reach that was simply not fundable before. For larger organisations the saving matters less as access and more as room to try several things before one of them works.
The practical consequence we see in client work is that model choice stops being the expensive decision. Once inference is cheap, the cost and the risk both move to everything around the model: the data you ground it in, the guardrails you put on it, and the workflow you wire it into.
The Challenges
It is not all smooth sailing. Two key hurdles stand out: ensuring trust and safety in their models, and navigating hardware access constraints due to export controls. Censorship complexities add another layer. But these constraints could actually fuel greater innovation. DeepSeek has already proven that challenges can inspire breakthroughs.
Worth the Hype?
From our hands-on assessment, DeepSeek R1 is the real deal. But perhaps not for the reasons making headlines. Its true significance lies not in benchmark scores but in how it challenges fundamental assumptions about LLM training. The question is not whether this approach works. DeepSeek has already proven it does. The question is how quickly other providers and businesses will adapt.
Fixed fee, agreed before we start. A senior engineer replies within two business days.