Enhancing LLM Reasoning: Chain of Draft with Semantically Diverse Thinking Tokens Using GRPO
Enhancing LLM Reasoning: Chain of Draft with Semantically Diverse Thinking Tokens Using GRPO The Challenge: Efficient Reasoning in LLMs Large Language Models (LLMs) have become remarkably capable at complex reasoning tasks, but this often comes at a cost: verbose outputs that consume significant computational resources. The Chain of Thought (CoT) prompting technique, while effective for accuracy, generates lengthy reasoning steps that increase token usage and latency. Enter Chain of Draft (CoD), a promising alternative introduced by Xu et al....