Google Cloud AI Research released RRSI, a framework allowing language model agents to autonomously optimize their prompts, tools, and memory structures without modifying underlying model weights. The system uses regulari…
#Fine-Tuning & Adaptation
Fine-tuning approaches are advancing toward greater efficiency and autonomous optimization. Google's RRSI framework enables language models to self-improve their prompts and tools through regularization techniques that prevent overfitting to training data, achieving measurable gains on held-out benchmarks. Meanwhile, BottleCap AI demonstrated how targeted fine-tuning can reduce computational costs by streamlining reasoning processes, though with modest accuracy trade-offs. Perplexity took a different angle, leveraging real user sessions—including failures—combined with hint-guided distillation to improve agent reliability. These developments reflect a shift toward fine-tuning methods that enhance model capabilities while managing computational trade-offs and generalization challenges.
BottleCap AI has released ThinkingCap-Qwen3.8-27B, a fine-tune of Qwen3.8-27B designed to shorten reasoning traces. Across 12 benchmarks, it uses 37.2% fewer thinking tokens on average while macro accuracy slips 0.86 per…
Perplexity Research described a post-training approach for its computer agent that learns from real user sessions, including unsuccessful ones. The method combines rejection-sampling fine-tuning with hint-guided self-dis…