Read More
Taxi stuck sideways in narrow Central alley
24-08-2026 15:55 HKT
DeepSeek is working with Tsinghua University on reducing the training its artificial intelligence models need in an effort to lower operational costs.
The new method aims to help artificial intelligence models better adhere to human preferences by offering rewards for more accurate and understandable responses, the researchers wrote. Reinforcement learning has proven effective in speeding up AI tasks in narrow applications and spheres. However, expanding it to more general applications has proven challenging - and that's the problem that DeepSeek's team is trying to solve with something it calls self-principled critique tuning.
DeepSeek is calling these new models DeepSeek-GRM - short for "generalist reward modeling" - and will release them on an open-source basis, the company said. Other AI developers, including Chinese tech giant Alibaba (9988) and San Francisco-based OpenAI, are also pushing into a new frontier of improving reasoning and self-refining capabilities while an AI model is performing tasks in real time.
Menlo Park, California-based Meta Platforms released its latest family of AI models, Llama 4, over the weekend and marked them as its first to use the Mixture of Experts architecture. DeepSeek's models rely significantly on MoE to make more efficient use of resources, and Meta benchmarked its new release against the Hangzhou-based startup.DeepSeek hasn't specified when it might release its next flagship model.