Shortcut-Aware Reasoning Training

Gradient-aware training that detects and mitigates shortcut reasoning in language models

This project targets shortcut reasoning in language models, where predictions rely on surface pattern matching, memorization, or keyword correlations rather than logical inference. The method, Shortcut-Aware Reasoning Training, identifies shortcut-promoting training samples through gradient misalignment with the reasoning objective and through answer-token concentration, then applies gradient surgery to adjust the training dynamics. The repository provides the data-centric training pipeline and the evaluation protocol on controlled reasoning benchmarks. (Cao et al., 2026)

References

2026

  1. arXiv
    Mitigating Shortcut Reasoning in Language Models: A Gradient-Aware Training Approach
    Hongyu Cao, Kunpeng Liu, Dongjie Wang, and 1 more author
    arXiv preprint arXiv:2603.20899, 2026