Homepage About Me News Technical Reports Publications Tutorials Honors and Awards Educations Experience

Hello! I'm Shicheng Xu (徐士成), a Post-training Researcher at DeepSeek, focusing on reinforcement learning and agentic post-training for large language models.

I received my Ph.D. from the Institute of Computing Technology, Chinese Academy of Sciences (State Key Laboratory of AI Safety), under the supervision of Prof. Liang Pang and Prof. Xueqi Cheng, and my B.S. from Harbin Institute of Technology. Before joining DeepSeek, I was an intern at ByteDance Seed (TopSeed Talent Program), where I worked for RL scaling in the post-training of Seed 1.8, Seed 2.0 and Seed 2.1.

My research interests include Reinforcement Learning (RL), Agent, and Retrieval-Augmented Generation (RAG).

🔥 News

🚀 Foundation Model Technical Reports

Frontier foundation models I contributed to as a core member of the RL scaling / post-training team at ByteDance Seed.

Seed2.1 Model Card: Agentic Intelligence for Productivity Technical Report

Worked for RL scaling algorithms in post-training (Core Contributor).

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity Technical Report

Worked for RL scaling in post-training (Core Contributor).

Seed1.8 Model Card: Towards Generalized Real-World Agency Technical Report

Worked for RL scaling in post-training (Core Contributor).

📝 Publications

First-author publications. Full list on Google Scholar.

DxDirector: An Agentic Large Language Model Driving the Full-process Clinical Diagnosis Nature Communications

Shicheng Xu, Xin Huang, Zihao Wei, Liang Pang, Huawei Shen, Xueqi Cheng

A 7B clinical agentic LLM that outperforms much more expensive models (DeepSeek, o3-mini, o1, Gemini-2.0-flash) on dynamic, complex, end-to-end full-process clinical reasoning, while interacting with real-world medical workflows efficiently and robustly.

A Theory for Token-Level Harmonization in Retrieval-Augmented Generation ICLR 2025

Shicheng Xu, Liang Pang, Huawei Shen, Xueqi Cheng

The first theoretical framework for the fusion between LLMs' parametric knowledge and retrieved external knowledge, explaining the underlying mechanism and enabling a robust collaborative generation framework.

Cross-Modal Safety Mechanism Transfer in Large Vision-Language Models ICLR 2025

Shicheng Xu, Liang Pang, Yunchang Zhu, Huawei Shen, Xueqi Cheng

Identifies and fixes the hidden alignment gap between visual and textual modalities in LVLMs, improving the robustness of multimodal agents against harmful visual inputs.

Search-in-the-Chain: Interactively Enhancing Large Language Models with Search for Knowledge-intensive Tasks WWW 2024

Shicheng Xu, Liang Pang, Huawei Shen, Xueqi Cheng, Tat-Seng Chua

Reshapes the interaction paradigm between LLMs and external tools: keeps the LLM's deep reasoning ability while fully exploiting external information, with robustness to noisy environments.

Unsupervised Information Refinement Training of Large Language Models for Retrieval-Augmented Generation ACL 2024

Shicheng Xu, Liang Pang, Mo Yu, Fandong Meng, Huawei Shen, Xueqi Cheng, Jie Zhou

An unsupervised training-data construction method that aligns LLMs with knowledge fed back from external environments, achieving efficient and robust fusion of internal and external knowledge.

List-aware Reranking-Truncation Joint Model for Search and Retrieval-augmented Generation WWW 2024 Oral

Shicheng Xu, Liang Pang, Jun Xu, Huawei Shen, Xueqi Cheng

A generative joint reranking-truncation model that provides high-quality retrieval lists for RAG and keeps RAG robust to retrieval results.

Invisible Relevance Bias: Text-Image Retrieval Models Prefer AI-Generated Images SIGIR 2024 Oral

Shicheng Xu, Danyang Hou, Liang Pang, Jingcheng Deng, Jun Xu, Huawei Shen, Xueqi Cheng

Reveals the hidden fairness impact of AI-generated images on cross-modal retrieval models.

BERM: Training the Balanced and Extractable Representation for Matching to Improve Generalization Ability of Dense Retrieval ACL 2023 Oral

Shicheng Xu, Liang Pang, Huawei Shen, Xueqi Cheng

Improves multi-task, multi-domain generalization of dual-encoder retrieval models by explicitly modeling relevance-scoring behavior during contrastive learning.

Match-Prompt: Improving Multi-task Generalization Ability for Neural Text Matching via Prompt Learning CIKM 2022 Oral

Shicheng Xu, Liang Pang, Huawei Shen, Xueqi Cheng

Improves multi-task, multi-domain generalization of text matching models via prompt learning in continuous vector space.

NIR-Prompt: A Multi-task Generalized Neural Information Retrieval Training Framework TOIS

Shicheng Xu, Liang Pang, Huawei Shen, Xueqi Cheng

A full-pipeline training framework for generalizable neural information retrieval, from recall to ranking.

RLKD: Distilling LLMs' Reasoning via Reinforcement Learning AAAI 2026

Shicheng Xu, Liang Pang, Huawei Shen, Xueqi Cheng

Effective distillation of LLMs' reasoning ability via reinforcement learning.

📚 Tutorials

🎖 Honors and Awards

📖 Educations

💻 Experience