Abstract
Machine learning models increasingly inform research funding and hiring decisions by predicting future citation impact from early-career publication patterns, raising stakes for model interpretability. We compare four post-hoc explainability methods (SHAP, LIME, integrated gradients, and attention-weight inspection) applied to a gradient-boosted and a transformer-based research-impact predictor trained on a corpus of 42,000 publication records. We evaluate methods on fidelity (agreement with model behaviour under perturbation), stability (consistency across similar inputs), and interpretability to domain-expert evaluators in a blinded assessment. SHAP achieves the highest fidelity and stability scores across both model architectures, but domain experts rated integrated-gradients explanations as most actionable for the transformer model specifically. We find a persistent fidelity-comprehensibility trade-off and argue that explainability method selection for research-evaluation contexts should be use-case-specific rather than defaulting to a single "best" method.