BottleCap AI Cuts Reasoning Tokens with ThinkingCap-Qwen3.8-27B Fine-Tune

BottleCap AI has released ThinkingCap-Qwen3.8-27B, a fine-tune of Qwen3.8-27B designed to shorten reasoning traces. Across 12 benchmarks, it uses 37.2% fewer thinking tokens on average while macro accuracy slips 0.86 percentage points, from 86.65% to 85.79%. The model is a drop-in replacement on vLLM and SGLang and is offered in FP8, NVFP4, GGUF, and MLX formats.
ThinkingCap-Qwen3.8-27B is BottleCap AI’s second entry in its ThinkingCap line. It adapts Qwen3.8-27B to shorten reasoning traces without adding knowledge or changing answer style. The earlier release targeted Qwen3.6-27B. Testing spanned 12 benchmarks, with per-benchmark token reductions ranging from 10.7% to 65.5%.
The model ships in FP8, NVFP4, GGUF, and MLX builds and can replace Qwen3.8-27B directly in vLLM or SGLang deployments. Its repository is access-controlled; commercial use past a small-business license threshold needs a BottleCap agreement. Evaluation ran on one NVIDIA H200 with vLLM 0.29.0 and fixed sampling settings.
If shorter reasoning traces hold across deployments, organizations running reasoning models may see lower inference costs, faster responses, and reduced compute demand. Developers and researchers could benefit from easier local or edge use via GGUF and MLX builds, while accuracy-sensitive users, such as those in math or agentic tasks, may need to weigh token savings against benchmark regressions. The gated license and commercial terms could shape who can adopt it.