My name is Narutatsu, but please feel free to call me Edward or Tatsu. I am a PhD student at Princeton University in the Computer Science Department and Princeton Language + Intelligence (PLI), where I am fortunate to be advised by Professor Sanjeev Arora. I also work closely with Professor Micah Goldblum at Columbia University. I am gratefully supported by the Gordon Wu Fellowship and the Ezoe Memorial Recruit Foundation. Please find my CV here.
Research Interests: I'm broadly interested in AI safety (e.g., developing a scientific understanding of failure modes in current models), misalignment (e.g., how scaling capabilities enables undesirable behavior in frontier models), and self-improvement, along with the safety implications that come with it.
I previously finished my B.S./M.S. at Columbia University where I was part of the Egleston Scholars Program. I am indebted to Professors Kathleen McKeown, Daniel Hsu, and Nakul Verma, by whom I have had the great fortune of being advised.
Google Scholar · GitHub · Twitter · Email
Dear prospective collaborators: for information.
Thanks for your interest! I'm always happy to chat and to work with new people.
For collaborators: Email me about any research topics or ideas you'd like to explore together. The papers below give a sense of the areas I can help with.
For Princeton undergraduate/master's students: If you're at Princeton and interested in research opportunities, email me your CV and a brief description of your research interests.
For non-Princeton students: Same. My bandwidth is limited, but I'm still happy to chat! Arranging campus compute access is also difficult, so remote collaboration works best if you already have the compute your project needs.
Do Thinking Tokens Help with Safety?
, Abhishek Panigrahi, Sanjeev Arora
ICML 2026 AI4GOOD Workshop (
General Track Best Paper Award, Oral Presentation)
Measuring the Limits of Continual Learning for LLMs
Nimit Kalra*, *, Zerzar Bukhari, Ang Li, Sanae
Lotfi, Liam Fowl, Micah Goldblum
ICML 2026 CompLearn Workshop
Rethinking On-Policy Self-Distillation for Thinking Models
Simran Kaur, , Yinghui He, Liam Fowl, Sanjeev
Arora
ICML 2026 FoGen Workshop
Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision
Yinghui He, Simran Kaur, Adithya Bhaskar, Yongjin Yang, Jiarui Liu, , Liam Fowl, Abhishek Panigrahi, Danqi Chen,
Sanjeev Arora
ICML 2026 RLxF Workshop (Oral Presentation)
Reranking-based Generation for Unbiased Perspective Summarization
, Nicholas Deas, Kathleen McKeown
ACL 2025 (Findings)
Speak Easy: Eliciting Harmful Instructions from LLMs with Simple Interactions
Yik Siu Chan*, *, Yuxin Xiao*, Marzyeh Ghassemi
ICML 2025
Latent Space Interpretation for Stylistic Analysis and Explainable Authorship
Attribution
Milad Alshomary, , Marianna Apidianaki, Ajay
Patel,
Smaranda Muresan, Kathleen McKeown
COLING 2025
Do Models Explain Themselves? Counterfactual Simulatability of Natural Language
Explanations
Yanda Chen, Ruiqi Zhong, , Chen Zhao, He He,
Jacob Steinhardt,
Zhou Yu, Kathleen McKeown
ICML 2024 (Spotlight)
Enhancing Few-shot Text-to-SQL Capabilities of Large Language Models: A Study on Prompt
Design
Strategies
Linyong Nan, Yilun Zhao, Weijin Zou, , Jaesung
Tae, Ellen
Zhang, Arman Cohan, Dragomir Radev
EMNLP 2023 (Findings)
The Effect of Model Capacity on the Emergence of In-Context Learning in Transformers
Berkan Ottlik*, *, Daniel Hsu, Clayton Sanford
ICLR 2024 (ME-FoMo Workshop)
Contrastive Loss is All You Need to Recover Analogies as Parallel Lines
, Fei-Tzin Lee, Nakul Verma
ACL 2023 (RepL4NLP Workshop)