Contrastive-LM Introduces CLM-8B, an Open Action-Scoring System One Model

Contrastive-LM has released CLM-8B, an open System One model that scores candidate actions against a state instead of generating text. It attaches two small projection heads to a frozen Qwen3-8B encoder and trains them with a contrastive InfoNCE objective. In zero-shot tests it runs up to 9× faster than TypeSafe's Jev, and with fine-tuned heads as a verifier it reaches 81.6% on held-out DeepSWE tasks and 87.6% on held-out Terminal-Bench 2.1 tasks.
CLM-8B is an Apache-2.0 release whose trainable head is only 75 MB. It runs on a single NVIDIA GPU under Linux, with vLLM serving the frozen Qwen3-8B encoder. Each encoder adds a 20M-parameter projection head, trained with a bidirectional InfoNCE loss.
The system exposes three query types: Noul for truth probability, Choice for selecting among declared options, and Score for ordered rubric levels. Its pipeline used roughly 60M Nemotron DQA pairs, 30M Gemini 2.5 Flash-Lite synthetic hard negatives, and about 1M agent trajectories. On held-out questions, pretraining reached 52.1% top-1, mid-training 69.2%.
The release could affect agent developers, coding-tool teams, and infrastructure providers by making action scoring and verifier selection cheaper and faster to deploy. Smaller organizations may experiment with System One-style decision layers without proprietary access. If open CLMs gain traction, proprietary vendors could face pressure to differentiate on accuracy, reliability, or ecosystem support. Users of coding agents may see quicker verification, though gains depend on task fit and fine-tuning. The broader impact remains uncertain because benchmark results cover limited held-out subsets.