Chain-of-Thought Control Tests May Underestimate Model Capabilities

The CoTControl eval tests whether models can follow formatting constraints in their reasoning. Recent models score poorly, but the author finds that prompt engineering can significantly improve performance. This suggests current evals may not fully capture models' ability to control their chain-of-thought.
The CoTControl evaluation probes whether language models can adhere to explicit formatting constraints within their internal reasoning traces. Recent model versions have shown weak performance on this benchmark, yet the author demonstrates that simple prompt adjustments can markedly boost scores. This discrepancy implies that the eval may be measuring instruction-following sensitivity rather than fundamental reasoning control. In the broader field of safety-alignment, such tests aim to verify that models can keep their hidden deliberations aligned with user directives. If prompt engineering alone can alter outcomes, current assessments might not reliably distinguish between models that genuinely lack control and those that merely need better elicitation. This raises questions about how to design more robust evaluations that separate capability from compliance.
If current evals underestimate models' chain-of-thought control, developers and auditors may overestimate safety risks, potentially slowing deployment of otherwise capable systems. Conversely, over-reliance on prompt engineering could mask underlying fragility, leading to unexpected failures in real-world contexts where prompts are not optimized. Users and regulators could be misled about model reliability, affecting trust and oversight decisions. The finding suggests that evaluation methodology itself needs scrutiny, as both under- and over-estimation carry societal consequences for AI accountability.