I am a second-year CS PhD student at Princeton University, where I'm fortunate to be advised by Prof. Sanjeev Arora. I am gratefully supported by the Gordon Wu Fellowship and the Ezoe Memorial Recruit Foundation.
I previously finished my B.S./M.S. at Columbia University where I was part of the Egleston Scholars Program. I am indebted to Professors Kathleen McKeown, Daniel Hsu, and Nakul Verma, by whom I have had the great fortune of being advised.

Research Interests: I'm broadly interested in AI safety (e.g., developing a scientific understanding of failure modes in current models) and misalignment (e.g., understanding and mitigating reward/goal misgeneralization).

Link to CV.

Contact

Google Scholar · GitHub · Twitter · Email

Prospective collaborators: for information.

Thanks for your interest! I'm always happy to chat and to work with new people.

For collaborators: Email me about any research topics or ideas you'd like to explore together. The papers below give a sense of the areas I can help with.

For Princeton undergraduate/master's students: If you're at Princeton and interested in research opportunities, email me your CV and a brief description of your research interests.

For non-Princeton students: Same. My bandwidth is limited, but I'm still happy to chat! Arranging campus compute access is also difficult, so remote collaboration works best if you already have the compute your project needs.

Papers

Do Thinking Tokens Help with Safety?
Narutatsu Ri, Abhishek Panigrahi, Sanjeev Arora
NeurIPS 2026 (Spotlight); ICML 2026 AI4GOOD Workshop (trophy Best Paper Award, Oral)

Measuring the Limits of Continual Learning for LLMs
Nimit Kalra*, Narutatsu Ri*, Zerzar Bukhari, Ang Li, Sanae Lotfi, Liam Fowl, Micah Goldblum
ICML 2026 CompLearn Workshop

Rethinking On-Policy Self-Distillation for Thinking Models
Simran Kaur, Narutatsu Ri, Yinghui He, Liam Fowl, Sanjeev Arora
NeurIPS 2026; ICML 2026 FoGen Workshop

Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision
Yinghui He, Simran Kaur, Adithya Bhaskar, Yongjin Yang, Jiarui Liu, Narutatsu Ri, Liam Fowl, Abhishek Panigrahi, Danqi Chen, Sanjeev Arora
ICML 2026 RLxF Workshop (Oral)

Reranking-based Generation for Unbiased Perspective Summarization
Narutatsu Ri, Nicholas Deas, Kathleen McKeown
ACL 2025 (Findings)

Speak Easy: Eliciting Harmful Instructions from LLMs with Simple Interactions
Yik Siu Chan*, Narutatsu Ri*, Yuxin Xiao*, Marzyeh Ghassemi
ICML 2025

Latent Space Interpretation for Stylistic Analysis and Explainable Authorship Attribution
Milad Alshomary, Narutatsu Ri, Marianna Apidianaki, Ajay Patel, Smaranda Muresan, Kathleen McKeown
COLING 2025

Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations
Yanda Chen, Ruiqi Zhong, Narutatsu Ri, Chen Zhao, He He, Jacob Steinhardt, Zhou Yu, Kathleen McKeown
ICML 2024 (Spotlight)

Enhancing Few-shot Text-to-SQL Capabilities of Large Language Models: A Study on Prompt Design Strategies
Linyong Nan, Yilun Zhao, Weijin Zou, Narutatsu Ri, Jaesung Tae, Ellen Zhang, Arman Cohan, Dragomir Radev
EMNLP 2023 (Findings)

The Effect of Model Capacity on the Emergence of In-Context Learning in Transformers
Berkan Ottlik*, Narutatsu Ri*, Daniel Hsu, Clayton Sanford
ICLR 2024 (ME-FoMo Workshop)

Contrastive Loss is All You Need to Recover Analogies as Parallel Lines
Narutatsu Ri, Fei-Tzin Lee, Nakul Verma
ACL 2023 (RepL4NLP Workshop)