Evaluating the Instruction-following Abilities of Language Models using Knowledge Tasks

Rudra Murthy, Prince Kumar, Praveen Venkateswaran, Danish Contractor

Published: 2024, Last Modified: 28 Jul 2025CoRR 2024EveryoneRevisionsBibTeXCC BY-SA 4.0

Abstract: LLM evaluation benchmarks have traditionally separated the testing of knowledge/reasoning capabilities from instruction following. In this work, we study the interaction between knowledge and instruction following, and observe that LLMs struggle to follow simple answer modifying instructions, and are also distracted by instructions that should have no bearing on the original knowledge task answer. We leverage existing multiple-choice answer based knowledge benchmarks and apply a set of simple instructions which include manipulating text (eg.: change case), numeric quantities (eg.: increase value, change formatting), operate on lists (eg.: sort answer candidates) and distractor instructions (eg.: change case of numeric answers).