Perplexity Uses Failed Sessions and Hint-Guided Distillation to Improve Its Computer Agent

Perplexity Research described a post-training approach for its computer agent that learns from real user sessions, including unsuccessful ones. The method combines rejection-sampling fine-tuning with hint-guided self-distillation, assigning imitation, correction, or context-only treatment to assistant turns. In an A/B test, tool-call failures dropped from 2.24% to 1.77%, a reported 21.2% relative reduction, though weights and code were not released.
Perplexity Research described a post-training method for a model inside Perplexity Computer. It learns from real user sessions, including failed ones, by combining rejection-sampling fine-tuning with hint-guided self-distillation. Assistant turns receive one of three treatments: imitation, correction, or context-only. In an A/B test, tool-call failures dropped from 2.24% to 1.77%, a reported 21.2% relative reduction.
The weights and training code were not released. The model runs only as an option in Perplexity Computer, though its base model, GLM 5.2, is openly available on Hugging Face. Hints are short corrections grounded in information the model already had; on-policy self-distillation uses a teacher pass with the hint and a student pass without it.
If such training methods prove robust, people who rely on computer agents for tasks like search, scheduling, or form-filling could see fewer tool errors and smoother interactions. Developers and researchers may benefit from the described techniques, though the lack of released weights or code could slow independent verification and adaptation. Organizations deploying agents may weigh reliability gains against privacy concerns, since the pipeline uses real sessions while excluding opted-out users and personally identifiable information. The broader impact may depend on whether similar approaches become reproducible and widely audited.