Invited to serve as an Area Chair for ARR Rolling Review 2026.
PhD Student in Computer Science
Jinhe Bi
My name is . I am a PhD student at LMU Munich advised by Prof. Volker Tresp, working on post-training and agentic reasoning. I am also a visiting researcher at NUS NExT++ with Prof. Chua Tat-Seng, studying latent knowledge in post-training, and a research intern at Huawei 2012 working on efficient post-training.
Open to active collaborations across research, projects, and mentoring; open-source by default, with all current works released publicly. Discussions, issues, and collaboration ideas are welcome by email or WeChat.
News
Updates
EchoRL and The Geometry of Reasoning were accepted to ICML 2026.
PRISM was included in ACL 2026 Best Paper consideration.
Started visiting research at NUS NExT++ with Prof. Chua Tat-Seng.
LLaVA Steering was accepted to ACL 2025 Main Conference.
Joined Huawei 2012 for research on efficient post-training.
Started the PhD program in Computer Science at LMU Munich.
Research
Efficient AI
Efficiency as a scaling axis: reducing trainable parameters, supervision, and adaptation cost while preserving generalization.
Post-Training Optimization
Stable and sample-efficient post-training through RLVR, OPD, and SFT as controllable optimization stages.
Agentic Reasoning
Long-horizon task solving under extended interaction, tool-use, and delayed-feedback settings.
Background
Experience
Mar 2026 - Present
Visiting Student Researcher, NUS NExT++
Latent knowledge in LLM post-training, supervised by Prof. Chua Tat-Seng.
May 2024 - Present
Doctoral Researcher, Huawei 2012
Efficient post-training, multimodal learning, and reasoning-oriented optimization.
Nov 2022 - Mar 2024
Student Research Assistant, LMU Munich
Research on state-of-the-art methods in video understanding.
Education
Apr 2024 - Apr 2027
PhD in Computer Science
Ludwig Maximilian University of Munich, Germany.
Apr 2022 - Dec 2023
MSc in Computer Science
Ludwig Maximilian University of Munich, Germany.
Sep 2016 - Aug 2020
BSc in Electrical and Electronics Engineering
Tianjin University, China.
Selected Work
Publications
2026
ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning
Under Review · First Author
2026
OPD-V: Visual On-Policy Self-Distillation with Modality Balance
Under Review · Project Leader
2026
MAE I Trust Myself? Self-Evaluating VLA Action Generation with Markov Attention Entropy
Under Review · Project Leader
2026
2026
2026
2025
Academic Service
Service
ARR Rolling Review 2026 (ACL, EMNLP, ...)
NeurIPS, ICML, ICLR, AAAI, CVPR, ECCV, ACM MM, ACL Rolling Review (ARR)
Open Source
Contributions
evolvinglmms-lab/lmms-eval
Supported effective evaluation of models on text tasks.
camel-ai/loong
Built verified environments and data for mathematical reasoning tasks.
Recognition
Grants and Awards
- EuroHPC AI and Data-Intensive grant
- NHR@FAU GPU computing grants
- GCS-JUPITER supercomputing grant
- Allianz China Scholarship
- Deutschlandstipendium
- MCM/ICM First Prize
- RoboMaster National Finals Third Prize
Visitors
Visitor Map
Live country-level visitor map powered by Flag Counter. Click the map to inspect pageviews and country statistics.
View visitor stats