WRAVAL — WRiting Assist eVALuation
arXiv:2601.03268v1 Announce Type: new Abstract: The emergence of Large Language Models (LLMs) has shifted language model evaluation toward reasoning and problem-solving tasks as measures of general intelligence. Small Language Models (SLMs) — defined here as models under 10B parameters — typically score 3-4 times lower than LLMs on these metrics. However, we demonstrate that these evaluations fail to capture SLMs’ effectiveness in common industrial applications, such as tone modification tasks (e.g., funny, serious, professional). We propose an evaluation […]