|
|
Zichao Li (李梓超)
I am a RedStar researcher on the post-training team at dots studio, Xiaohongshu (RedNote). My work focuses on reasoning, agents, and reinforcement learning for foundation-model post-training.
Before joining Xiaohongshu, I was a research intern with the AMD AI Group, where I worked on efficient attention.
I received my Master's degree from the Chinese Information Processing Laboratory at the Institute of Software, Chinese Academy of Sciences, where I was advised by Professor Le Sun and Professor Xianpei Han. Before that, I received my Bachelor's degree from Beihang University in 2023.
Contact: lazyc81 [at] gmail [dot] com
Google Scholar
/
GitHub
|
Current Focus
-
Foundation models.
I am a core contributor to the post-training of RedNote's in-house dots.llm family, spanning dots.llm1, dots.llm2, and dots3. We recently released and open-sourced dots3-note preview, the first open-weight model in the dots3 family (Hugging Face).
-
Agent self-evolution.
For the dots team's IMO 2026 effort, I was responsible for the end-to-end agent harness and agent training. Following official evaluation, dots-note-3.0 achieved a perfect-score gold medal (42/42).
-
Reinforcement learning.
I study robust reward modeling (Cheems), efficient RL post-training (Tackling Length Inflation), and agentic RL (MemSearcher).
-
Multimodal intelligence.
My work has evolved from unified OCR (Seg2Act, READoc), through multimodal alignment (The Devil Is in the Details), toward computer-use agents (ongoing).
|
Selected Publications
Tackling Length Inflation Without Trade-offs: Group Relative Reward Rescaling for Reinforcement Learning
Zichao Li, Jie Lou, Fangchen Dong, Zhiyuan Fan, Mengjie Ren, Hongyu Lin, Xianpei Han, Debing Zhang, Le Sun, Yaojie Lu, Xing Yu
International Conference on Machine Learning (ICML), 2026
MemSearcher: Iterative Memory Integration for Search Agent via End-to-End Reinforcement Learning
Qianhao Yuan, Jie Lou, Zichao Li, Jiawei Chen, Yaojie Lu, Hongyu Lin, Le Sun, Debing Zhang, Xianpei Han
Findings of the Association for Computational Linguistics: ACL, 2026
The Devil Is in the Details: Tackling Unimodal Spurious Correlations for Generalizable Multimodal Reward Models
Zichao Li, Xueru Wen, Jie Lou, Yuqiu Ji, Yaojie Lu, Xianpei Han, Debing Zhang, Le Sun
International Conference on Machine Learning (ICML), 2025
READoc: A Unified Benchmark for Realistic Document Structured Extraction
Zichao Li*, Aizier Abulaiti*, Yaojie Lu, Xuanang Chen, Jia Zheng, Hongyu Lin, Xianpei Han, Shanshan Jiang, Bin Dong, Le Sun
Findings of the Association for Computational Linguistics: ACL, 2025
Cheems: A Practical Guidance for Building and Evaluating Chinese Reward Models from Scratch
Xueru Wen*, Jie Lou*, Zichao Li*, Yaojie Lu, Xing Yu, Yuqiu Ji, Guohai Xu, Hongyu Lin, Ben He, Xianpei Han, Le Sun, Debing Zhang
Annual Meeting of the Association for Computational Linguistics (ACL), 2025
Seg2Act: Global Context-aware Action Generation for Document Logical Structuring
Zichao Li*, Shaojie He*, Meng Liao, Xuanang Chen, Yaojie Lu, Hongyu Lin, Yanxiong Lu, Xianpei Han, Le Sun
Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024
* Equal contribution.
|
Beyond Work
I enjoy hiking. On a recent solo trip to Jeju Island, I hiked along the Jeju Olle Trail. During my time in Beijing, I often went hiking around Xiangbala.
I also enjoy films, TV series, anime, books, and rock and indie music.
|
|