Rajan Agarwal

Research

Models, training, and evaluation.

  1. Post-training a small model to follow a live stream of typing and learn when to act or stay quiet.

  2. An ultra long-horizon benchmark for coding agents doing engineering and research work.

  3. Reinforcement-learning work for the browser-use model as a research intern.

    Research internship at Amazon AGI, Mentioned in: RL training recipe
  4. An RL setup in which language models learn their own compact codebooks.

  5. Hidden-information games for studying deceptive behavior learned from reward.

  6. Cross-lingual alignment through encoder injection for low-resource languages.