The author revises their previous work on endogenous alignment by incorporating feedback about the importance of imitation learning and the parent-child dependency relationship. Drawing from human development, they argue…
#Alignment Research
Two critical vulnerabilities in AI alignment have drawn recent attention. Researchers highlight how fixed-parameter systems inevitably create imprecise boundaries around abstract concepts, potentially causing severe misclassifications under optimization pressure. Separately, emerging architectural designs that enable latent reasoning—operating in non-transparent internal states rather than explicit text—pose challenges for human oversight mechanisms. Both concerns center on fundamental limitations: concept imprecision may prove difficult to resolve within current approaches, while latent reasoning channels could substantially degrade transparency and human supervision capabilities as AI systems become more capable and numerous.
The author argues that AI systems with fixed parameters must inevitably draw imperfect boundaries around abstract concepts like consciousness and suffering in their internal models, creating unavoidable vulnerabilities. …
A new analysis argues that emerging AI architectures allowing models to reason in latent states rather than explicit text would severely degrade the usefulness of chain-of-thought oversight. The authors note that current…
Researchers introduced WorkspaceBench, a benchmark of 3,356 questions across 27 families to evaluate activation-to-text interpretability tools on reading a model's global workspace. It includes a hallucination-focused su…
The author argues that reinforcement learning (RL) acts as a black-box source of agency, raising classic misalignment risks compared to more transparent scaffolding-based approaches. Recent incidents and personal experie…
A new paper extends prior work on sequence prediction by introducing a restricted complexity measure based on layered zipline programs. The proposed algorithm runs in quasilinear time and polylog space for highly structu…
Tests indicate Astra outperforms other models in solving reasoning tasks without producing chain-of-thought, with significantly higher odds and more serial arithmetic steps per forward pass. The findings raise concerns a…
The article proposes a concrete measure to estimate how much unspoken sequential thinking an AI can perform. It aims to help companies disclose the degree to which their architectures enable latent reasoning, aiding over…
The article tests synthetic document finetuning as a method to inoculate models against reward hacking. Despite models expressing the intended beliefs, they still exhibited misalignment generalization. This suggests beli…