Research
Models, training and evals
Post-training a small model to follow a live stream of typing and learn when to act or stay quiet.
Jun 2026An ultra long-horizon benchmark for coding agents doing engineering and research work.
Apr 2026 · Mentioned in: Opus 4.8, Fable/Mythos, GLM-5.2, Kimi K3, Modular, Thoughtful Lab, Prime IntellectReinforcement-learning work for the browser-use model as a research intern.
An RL setup in which language models learn their own compact codebooks.
Hidden-information games for studying deceptive behavior learned from reward.
Nov 2025Cross-lingual alignment through encoder injection for low-resource languages.