Aller directement aux détails du preprintAller directement aux PREreviews

PREreviews de What Does an LLM Learn from Reinforcement Learning? A Mechanistic Interpretability Perspective with Fixed-SAE Track

1 PREreview

  1. PREreview par Sneh Pankajbhai Vora

    Summary

    The authors tackle a real obstacle in interpretability: you can't compare SAE features across training checkpoints, because separately trained dictionaries assign feature indices arbitrarily. Their fix is simple and sensible — pool activations from the base model and every RL checkpoint,…

    Lire la PREreview de Sneh Pankajbhai Vora