Xiangning Lin
I am actively looking for 27 Fall CS PhD position. I am always open to collaborate, feel free to drop me an email.
My research interests lie in AI accountability, trustworthy agents, and agent evaluation.
I am fortunate to collaborate with Dr. Jiaxin Pei at Stanford HAI and Prof. Alex Pentland on system prompt auditing, and to contribute to the Terminal-Bench community on agent benchmarking.
I am the co-founder of systempromptindex.ai, the largest system prompt library in the world.
Feel free to reach out if you have any ideas for potential collaboration, or just feel like having a casual chat!
🔥 News
- 2026.09: Harbor-Index was accepted to NeurIPS 2026!
- 2026.09: We released Harbor-Index, a curated set of 82 hard, diverse agentic tasks drawn from 29 benchmarks, together with Harbor Adapters, a unified evaluation infrastructure for 80+ agentic benchmarks.
- 2026.09: Gave an invited talk at BAAI on AISPA and systempromptindex.ai, the world’s largest open system prompt database.
- 2026.08: AISPA and systempromptindex.ai were reported by 机器之心.
📝 Publications
-
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Mike A Merrill, Alexander Glenn Shaw, Nicholas Carlini, et al. (including Xiangning Lin), Ludwig Schmidt
Accepted by ICLR 2026
Website | Paper | Code -
Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation
Lin Shi, Haowei Lin, Zixuan Zhu, Xiaoyue Zhou, Xiang Li, Xiangning Lin, et al.
Accepted by NeurIPS 2026
Website | Paper | Code -
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
Xiangning Lin*, Shenzhe Zhu*, Shu Yang, et al., Alex Pentland, Jiaxin Pei
Under review, ACL Rolling Review 2026
Website | Paper | Code