Saltar a detalles del preprintSaltar a PREreviews

PREreviews de What Does an LLM Learn from Reinforcement Learning? A Mechanistic Interpretability Perspective with Fixed-SAE Track

1 PREreview

  1. PREreview de Sneh Pankajbhai Vora

    Summary

    The authors tackle a real obstacle in interpretability: you can't compare SAE features across training checkpoints, because separately trained dictionaries assign feature indices arbitrarily. Their fix is simple and sensible — pool activations from the base model and every RL checkpoint,…

    Leer la PREreview de Sneh Pankajbhai Vora