PhD Student in Computer Science

Jinhe Bi

My name is . I am a PhD student at LMU Munich advised by Prof. Volker Tresp, working on post-training and agentic reasoning. I am also a visiting researcher at NUS NExT++ with Prof. Chua Tat-Seng, studying latent knowledge in post-training, and a research intern at Huawei 2012 working on efficient post-training.

Open to active collaborations across research, projects, and mentoring; open-source by default, with all current works released publicly. Discussions, issues, and collaboration ideas are welcome by email or WeChat.

LMU Munich Huawei 2012 NUS NExT++

News

Updates

Invited to serve as an Area Chair for ARR Rolling Review 2026.

EchoRL and The Geometry of Reasoning were accepted to ICML 2026.

PRISM was included in ACL 2026 Best Paper consideration.

Started visiting research at NUS NExT++ with Prof. Chua Tat-Seng.

LLaVA Steering was accepted to ACL 2025 Main Conference.

Joined Huawei 2012 for research on efficient post-training.

Started the PhD program in Computer Science at LMU Munich.

Research

01

Efficient AI

Efficiency as a scaling axis: reducing trainable parameters, supervision, and adaptation cost while preserving generalization.

02

Post-Training Optimization

Stable and sample-efficient post-training through RLVR, OPD, and SFT as controllable optimization stages.

03

Agentic Reasoning

Long-horizon task solving under extended interaction, tool-use, and delayed-feedback settings.

Background

Experience

Mar 2026 - Present

Visiting Student Researcher, NUS NExT++

Latent knowledge in LLM post-training, supervised by Prof. Chua Tat-Seng.

May 2024 - Present

Doctoral Researcher, Huawei 2012

Efficient post-training, multimodal learning, and reasoning-oriented optimization.

Nov 2022 - Mar 2024

Student Research Assistant, LMU Munich

Research on state-of-the-art methods in video understanding.

Education

Apr 2024 - Apr 2027

PhD in Computer Science

Ludwig Maximilian University of Munich, Germany.

Apr 2022 - Dec 2023

MSc in Computer Science

Ludwig Maximilian University of Munich, Germany.

Sep 2016 - Aug 2020

BSc in Electrical and Electronics Engineering

Tianjin University, China.

Selected Work

Publications

2026

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning

Jinhe Bi, Chennan Zhou, Zengjie Jin, Aniri, Shuo Lu, Wenke Huang, Hu Cao, Xun Xiao, Zhihong Zhu, Volker Tresp, Fei Shen, Yunpu Ma, Tat-Seng Chua

Under Review · First Author

2026

OPD-V: Visual On-Policy Self-Distillation with Modality Balance

Aniri, Jinhe Bi, Peng Liao, Zengjie Jin, Volker Tresp, Fei Shen, Yunpu Ma, Tat-Seng Chua

Under Review · Project Leader

2026

MAE I Trust Myself? Self-Evaluating VLA Action Generation with Markov Attention Entropy

Aniri, Chen Yilin, Jinhe Bi, Zengjie Jin, Yujun Wang, Yijun Tian, Volker Tresp, Fei Shen, Tat-Seng Chua, Yunpu Ma

Under Review · Project Leader

2026

EchoRL: Reinforcement Learning via Rollout Echoing

Jinhe Bi, Aniri, Minglai Yang, Xingcheng Zhou, Wenke Huang, Sikuan Yan, Yujun Wang, Zixuan Cao, Michael Färber, Xun Xiao, Volker Tresp, Yunpu Ma

ICML 2026 · Paper · Code

2026

The Geometry of Reasoning: Self-Evaluation via Layerwise Trajectory Evolution

Jinhe Bi, Danqi Yan, Yifan Wang, Wenke Huang, Haokun Chen, Guancheng Wan, Mang Ye, Xun Xiao, Hinrich Schütze, Volker Tresp, Yunpu Ma

ICML 2026 · Paper · Code

2026

PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection

Jinhe Bi, Aniri, Yifan Wang, Danqi Yan, Wenke Huang, Zengjie Jin, Xiaowen Ma, Sikuan Yan, Artur Hecker, Mang Ye, Xun Xiao, Hinrich Schütze, Volker Tresp, Yunpu Ma

ACL 2026 Best Paper Candidate · Paper · Code

2025

LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering

Jinhe Bi, Yujun Wang, Haokun Chen, Xun Xiao, Artur Hecker, Volker Tresp, Yunpu Ma

ACL 2025 Main Conference · Paper · Code

Academic Service

Service

Area Chair

ARR Rolling Review 2026 (ACL, EMNLP, ...)

Reviewer

NeurIPS, ICML, ICLR, AAAI, CVPR, ECCV, ACM MM, ACL Rolling Review (ARR)

Open Source

Contributions

camel-ai/loong

Built verified environments and data for mathematical reasoning tasks.

504 stars

Recognition

Grants and Awards

Compute Grants
  • EuroHPC AI and Data-Intensive grant
  • NHR@FAU GPU computing grants
  • GCS-JUPITER supercomputing grant
Scholarships
  • Allianz China Scholarship
  • Deutschlandstipendium
Competitions
  • MCM/ICM First Prize
  • RoboMaster National Finals Third Prize

Visitors

Visitor Map

Live country-level visitor map powered by Flag Counter. Click the map to inspect pageviews and country statistics.

View visitor stats
Live visitor map by country from Flag Counter