RealityHackerOpen in RealityHacker ⇢
Models · Fine-Tuning & Adaptation · published 2026-09-25T00:00:00+00:00 · via MarkTechPost

Perplexity Uses Failed Sessions and Hint-Guided Distillation to Improve Its Computer Agent

Image via MarkTechPost
Image via MarkTechPost

Perplexity Research described a post-training approach for its computer agent that learns from real user sessions, including unsuccessful ones. The method combines rejection-sampling fine-tuning with hint-guided self-distillation, assigning imitation, correction, or context-only treatment to assistant turns. In an A/B test, tool-call failures dropped from 2.24% to 1.77%, a reported 21.2% relative reduction, though weights and code were not released.

Expanded Detail

Perplexity Research described a post-training method for a model inside Perplexity Computer. It learns from real user sessions, including failed ones, by combining rejection-sampling fine-tuning with hint-guided self-distillation. Assistant turns receive one of three treatments: imitation, correction, or context-only. In an A/B test, tool-call failures dropped from 2.24% to 1.77%, a reported 21.2% relative reduction.

The weights and training code were not released. The model runs only as an option in Perplexity Computer, though its base model, GLM 5.2, is openly available on Hugging Face. Hints are short corrections grounded in information the model already had; on-policy self-distillation uses a teacher pass with the hint and a student pass without it.

Context

If such training methods prove robust, people who rely on computer agents for tasks like search, scheduling, or form-filling could see fewer tool errors and smoother interactions. Developers and researchers may benefit from the described techniques, though the lack of released weights or code could slow independent verification and adaptation. Organizations deploying agents may weigh reliability gains against privacy concerns, since the pipeline uses real sessions while excluding opted-out users and personally identifiable information. The broader impact may depend on whether similar approaches become reproducible and widely audited.

Expanded detail and Context are AI-generated analysis; the linked article remains the authoritative source.
Read the full article at MarkTechPost →
This summary is Al-enhanced to contain extended analysis and broader social context. The original is {NAME); the linked article is the authoritative source. Original headline: “Perplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation.” Browse more stories.