{"id":864,"date":"2025-12-04T15:19:09","date_gmt":"2025-12-04T07:19:09","guid":{"rendered":"https:\/\/www.chain258.com\/?p=864"},"modified":"2025-12-04T15:19:09","modified_gmt":"2025-12-04T07:19:09","slug":"bug-found-in-deepseek-v3-2-an-old-issue-token-waste","status":"publish","type":"post","link":"https:\/\/www.chain258.com\/index.php\/2025\/12\/04\/bug-found-in-deepseek-v3-2-an-old-issue-token-waste\/","title":{"rendered":"Bug Found in DeepSeek-V3.2: An Old Issue \u2013 Token Waste"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">Many users have pointed out that while\u00a0<strong>DeepSeek-V3.2&#8217;s long-thinking enhanced version, Speciale<\/strong>, has indeed put pressure on top closed-source models as an open-source solution, it also comes with a clear issue:When tackling complex tasks, it consumes an unusually high number of tokens \u2014 sometimes producing answers that are\u00a0<strong>long but incorrect<\/strong>.For example, when solving the same problem:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Gemini<\/strong>\u200b used just\u00a0<strong>20,000 tokens<\/strong><\/li>\n\n\n\n<li><strong>Speciale<\/strong>\u200b used up to\u00a0<strong>77,000 tokens<\/strong><\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">So, what\u2019s going on?<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Unresolved \u201cLength Bias\u201d<\/strong>Some researchers have noted that this is actually an old \u201cbug\u201d that has persisted in the\u00a0<strong>DeepSeek series<\/strong>\u200b since\u00a0<strong>DeepSeek-R1-Zero<\/strong>.In short, the issue lies in the\u00a0<strong>GRPO algorithm<\/strong>\u200b (Group Relative Policy Optimization).Researchers from institutions such as Sea AI Lab and the National University of Singapore have pointed out that GRPO contains\u00a0<strong>two hidden biases<\/strong>:<\/p>\n\n\n\n<h4 class=\"wp-block-heading\"><strong>1. Length Bias: The Longer the Wrong Answer, the Lighter the Penalty<\/strong><\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">When calculating rewards, GRPO takes&nbsp;<strong>answer length<\/strong>\u200b into account \u2014 and as a result,&nbsp;<strong>shorter wrong answers are penalized more harshly<\/strong>\u200b than longer ones.This leads to a counterintuitive behavior:The model tends to generate&nbsp;<strong>longer, incorrect answers<\/strong>\u200b that may look like it\u2019s \u201cthinking deeply\u201d or \u201creasoning step-by-step,\u201d but is actually&nbsp;<strong>padding its response to avoid penalties<\/strong>.<\/p>\n\n\n\n<h4 class=\"wp-block-heading\"><strong>2. Difficulty Bias: Overemphasis on Extremely Easy or Hard Questions<\/strong><\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">GRPO adjusts the weight of questions based on the&nbsp;<strong>score standard deviation within a batch<\/strong>.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>If everyone gets a question right (low standard deviation), or everyone gets it wrong (also low standard deviation), that question is treated as a \u201cfocus point\u201d and gets repeated training.<\/li>\n\n\n\n<li>Meanwhile,\u00a0<strong>medium-difficulty questions<\/strong>\u200b \u2014 where some get it right and some don\u2019t (high standard deviation) \u2014 are often ignored.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">But in reality,&nbsp;<strong>medium-difficulty questions are the most valuable for improving a model\u2019s capabilities<\/strong>.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><strong>Progress Made, But the Bias Remains<\/strong><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Zichen Liu, the lead author of the study, pointed out that&nbsp;<strong>DeepSeek-V3.2 has already fixed the \u201cdifficulty bias\u201d<\/strong>\u200b by introducing a new advantage value calculation method (as highlighted in the red box in the diagram below).However, the&nbsp;<strong>biased length normalization term still remains<\/strong>\u200b (blue box in the diagram).That means:&nbsp;<strong>the length bias is still there.<\/strong><\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><strong>Official Acknowledgement from DeepSeek<\/strong><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Interestingly, this issue has also been mentioned in&nbsp;<strong>DeepSeek\u2019s own technical report<\/strong>.The researchers admitted that&nbsp;<strong>token efficiency remains a challenge for DeepSeek-V3.2<\/strong>:In general, the two newly released models need to generate&nbsp;<strong>longer response trajectories<\/strong>\u200b to match the output quality of&nbsp;<strong>Gemini-3.0-Pro<\/strong>.Speciale, in particular, was designed with&nbsp;<strong>relaxed RL length limits<\/strong>, allowing the model to produce&nbsp;<strong>extremely long reasoning chains<\/strong>. This approach enables deep self-correction and exploration \u2014 but at the cost of&nbsp;<strong>burning a lot of tokens<\/strong>.In essence, DeepSeek is taking a path of&nbsp;<strong>\u201ccontinuously extending reinforcement learning under ultra-long contexts.\u201d<\/strong>That said, considering the&nbsp;<strong>cost per million tokens<\/strong>,&nbsp;<strong>DeepSeek-V3.2 is priced at just 1\/24th of GPT-5<\/strong>, which may be seen as a reasonable trade-off.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><strong>Also Worth Noting: 128K Context Limit<\/strong><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Additionally, some users have pointed out that&nbsp;<strong>DeepSeek\u2019s 128K context window hasn\u2019t been updated in a long time<\/strong>\u200b \u2014 which may also be related to&nbsp;<strong>limited GPU resources<\/strong>.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Many users have pointed out th&hellip;<\/p>\n","protected":false},"author":2,"featured_media":865,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[9,8,6],"tags":[228,51,229],"class_list":["post-864","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-companies","category-deep-tech","category-people","tag-bug","tag-deepseek","tag-token"],"_links":{"self":[{"href":"https:\/\/www.chain258.com\/index.php\/wp-json\/wp\/v2\/posts\/864","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.chain258.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.chain258.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.chain258.com\/index.php\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.chain258.com\/index.php\/wp-json\/wp\/v2\/comments?post=864"}],"version-history":[{"count":1,"href":"https:\/\/www.chain258.com\/index.php\/wp-json\/wp\/v2\/posts\/864\/revisions"}],"predecessor-version":[{"id":866,"href":"https:\/\/www.chain258.com\/index.php\/wp-json\/wp\/v2\/posts\/864\/revisions\/866"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.chain258.com\/index.php\/wp-json\/wp\/v2\/media\/865"}],"wp:attachment":[{"href":"https:\/\/www.chain258.com\/index.php\/wp-json\/wp\/v2\/media?parent=864"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.chain258.com\/index.php\/wp-json\/wp\/v2\/categories?post=864"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.chain258.com\/index.php\/wp-json\/wp\/v2\/tags?post=864"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}