Publications
(* indicates equal contribution)
GRADE: Probing Knowledge Gaps in LLMs through Gradient Subspace Dynamics
Y. Wang, Y. Liang, Y. Li, H. Zhang, H. Yan
COLM26 | Paper
alignment Interpretability
GraphMind: Unveiling Scientific Reasoning through Contextual Graphs for Novelty Assessment
I. Silva, H. Yan, L. Gui, Y. He
KDD26 | Paper
AI4science
Stop the Flip-Flop: Context-Preserving Verification for Fast Revocable Diffusion Decoding
Y. Xiang, W. Lan, …, H. Yan, …, Y. He
ICML26 | Paper)
reasoning
AutoMonitor-Bench: Evaluating the Reliability of LLM-Based Misbehavior Monitor
S. Yang, J. Hu, T. Li, H. Yan, W. Wang, D. Wang
ACL26-findings | Paper
alignment
Multi-Faceted Multimodal Monosemanticity
H. Yan, X. Cui*, Y. Lu, P Liang, J. Gu, Y. He, Y. Wang
ACL26-findings | Paper
Interpretability
When Thinking Backfires: Mechanistic intepretability into reason-induced misalignment
H. Yan, H, Xu, S. Qi, S. Yang, Y. He
ICLR26 | Paper
Interpretability alignment
Spectrum Projection Score: Aligning Retrieved Summaries with Reader Models in Retrieval-Augmented Generation
Z.Hu, Q.Zhu, S. Qi, Y. He, H. Yan, L. Gui
AAAI25 Oral | Paper
reasoning
GraphMind: Interactive Novelty Assessment System for Accelerating Scientific Discovery
I. Silva, H. Yan, L. Gui, Y. He
EMNLP25 Demo | Paper
AI4science
CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation
Z. Shen, H. Yan, L. Zhang, Y. Du, Y. He
EMNLP25 | Paper
reasoning
Position: LLMs Need a Bayesian Meta-Reasoning Framework for More Robust and Generalizable Reasoning
H. Yan, L. Zhang, J. Li, Z. S, Y. He
ICML25, Position Track | Paper
reasoning alignment
Drift: Enhancing LLM Faithfulness in Rationale Generation via Dual-Reward Probabilistic Inference
J. Li, H. Yan, Y. He
ACL25, Main | Paper
reasoning Interpretability alignment
Navigating Solution Spaces in Large Language Models through Controlled Embedding Exploration
Q. Zhu, R. Zhao. H. Yan, Y. He, Y. Chen, L. Gui
ICML25, Spotlight | Paper
reasoning
Direct preference optimization using sparse feature-level constraints
Q. Yin, C. Leong, H. Zhang, M. Zhu, H. Yan, Q. Zhang, Y. He, W. Li, J. Wang, Y. Zhang, L. Yang
ICML25 | Paper
reasoning
SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers
Y. Xiang, H. Yan, S. Ouyang, L. Gui, Y. He
COLM25 | Paper
AI4science
Encourage or Inhibit Monosemanticity? Revisit Monosemanticity from a Feature Decorrelation Perspective
H. Yan, Y. Xiang, G Chen, Y. Wang, L. Gui, Y. He
EMNLP24, main | Paper
Interpretability reasoning
Weak Reward Model Transforms Generative Models into Robust Causal Event Extraction Systems
I. Silva, H. Yan, L. Gui, Y. He
EMNLP24, main | Paper
reasoning
The Mystery and Fascination of LLMs: A Comprehensive Survey on the Interpretation and Analysis of Emergent Abilities
Y. Zhou, J. Li, Y.Xiang, H.Yan, L. Gui, Y. He
EMNLP24, main | Paper
Interpretability reasoning
Mirror: A Multiple-perspective Self-Reflection Method for Knowledge-rich Reasoning
H. Yan, Q. Zhu, X. Wang, L. Gui, Y. He
ACL24, main | Paper
reasoning
Addressing Order Sensitivity of In-Context Demonstration Examples in Causal Language Models.
Y. Xiang, H. Yan, L. Gui, Y. He
ACL24, findings | Paper
reasoning Interpretability
Counterfactual Generation with Identifiability Guarantee
H. Yan, L. Kong, L. Gui, Y. Chi, Eric. Xing, Y. He, K. Zhang
Neurips23, main | Paper
Interpretability reasoning
Explainable Recommender with Geometric Information Bottleneck
H. Yan, L. Gui, M. Wang, K. Zhang and Y. He
TKDE | Paper
Interpretability reasoning
Hierarchical Interpretation of Neural Text Classification
H. Yan, L. Gui and Y. He
Computational Linguistics, Present at EMNLP23 | Paper
Interpretability reasoning
Addressing Token Uniformity in Transformers via Singular Value Transformation
H. Yan, L. Gui, W. Li and Y. He
UAI22, spotlight | Paper
Interpretability reasoning
Distinguishability Calibration to In-Context Learning
H. Li, H. Yan, L. Gui, W. Li and Y. He
EACL23, findings | Paper
Interpretability reasoning
A Knowledge-Aware Graph Model for Emotion Cause Extraction
H. Yan, L. Gui and Y. He
ACL21, Oral | Paper
reasoning