Kaichen Huang

Kaichen Huang 黄楷宸

Ph.D. Student, Nanyang Technological University

I received my B.Sc. (2022) and M.Sc. (2025) from the School of Artificial Intelligence, Nanjing University, where I was a member of the LAMDA Group, advised by Prof. De-Chuan Zhan. I am now a Ph.D. student at the College of Computing and Data Science, Nanyang Technological University, advised by Prof. Bo An.

News

Research

My research centers on world models and interactive / streaming video generation — learning generative simulators of dynamic environments — together with the reinforcement- and imitation-learning methods that make agents act well inside them: unsupervised RL, and learning from imperfect or third-person demonstrations.

Selected Publications [ Full List · Google Scholar ]

Matrix-Game 3.5 teaser
Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory
Runjia Qian*, Zile Wang*, Jihai Zhang*, Kai Zou*, Wei Yu*, Jiaxing Li*, Zexiang Liu*, Yaokun Li, Fei Kang, Kaichen Huang, Mengyin An, Haobo Zhang, Biao Jiang, Jiahua Wang, Haofeng Sun, Yang Liu, Yangguang Li (*equal)
Preprint, 2026
Interactive world models extend video generation from offline clip synthesis toward persistent simulation of interactive virtual worlds, enabling applications in games, robotics, embodied agents, and XR. Achieving stable long-horizon interactive generation, however, remains challenging, as the model must simultaneously preserve scene geometry, dynamic consistency, and camera control while supporting real-time autoregressive generation. Building upon Matrix-Game 3.0, we present Matrix-Game 3.5, as shown in Figure 1, which advances real-time interactive world generation toward geometry-aware and long-horizon consistent simulation through three key improvements. First, we propose a unified geometry-aware memory framework, whose patch-memory and tiled-PRoPE components introduce no additional learnable parameters, combining explicit 3D patch retrieval with projective camera conditioning to enable geometry-consistent camera control and faithful long-horizon scene recall. Second, we introduce a static-dynamic disentangled world representation that separately models static scene geometry and dynamic subjects, preserving both geometric consistency and subject identity throughout long-horizon generation. Third, we develop a two-stage progressive real-time distillation framework that converts a bidirectional diffusion model into a few-step causal generator through Perceptual Flow Matching and curriculum based Self-Rollout DMD, enabling minute-long real-time interactive generation. Extensive experiments demonstrate that, with a unified training corpus spanning Unreal simulation environments, open-world games, and internet videos, MatrixGame 3.5 achieves strong performance in long-horizon scene recall, precise camera control, subject consistency, prompt-driven world generation, and stable real-time open-world interaction.
DistillAlign init+DMD pipelines
DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation
Jiaxing Li*, Kai Zou*, Cindy Zhou, Kaichen Huang, Junyao Gao, Zile Wang, Yang Liu, Bin Liu, Bo An, Yangguang Li (*equal)
Preprint, 2026
Existing autoregressive video distillation methods commonly adopt a Distribution Matching Distillation (DMD)-based multi-stage pipeline. However, they typically decouple the initialization and DMD stages -- which then pursue different target distributions -- and judge the intermediate student mainly by visual scores such as VBench. In this paper, we revisit this design from a distributional perspective. Given the mode-seeking nature of the distribution matching loss, a good initialization should match the mode coverage of the target DMD teacher, rather than merely pursuing high quality. To analyze this, we introduce a distributional evaluation protocol that measures precision and coverage between student and teacher distributions in a shared latent space. It exposes differences hidden by visual scores: some initializations reach high precision but low coverage, leading to suboptimal refinement, while mode-covering ones preserve broader support. Furthermore, even when the target distributions are aligned, DMD's reverse-KL objective can still drive the student toward high-probability teacher regions in late training, reducing coverage and diversity. To address this, we propose joint distillation, which combines DMD's mode-seeking objective with a Consistency Distillation-based mode-covering constraint. Experiments show that our method improves generation quality, coverage, and diversity; notably, even with a Wan-1.3B DMD teacher, it outperforms baselines refined with Wan-14B, underscoring the importance of distributional alignment in autoregressive video distillation.
Matrix-Game 3.0 gameplay
Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
Zile Wang*, Zexiang Liu*, Jiaxing Li*, Kaichen Huang*, Baixin Xu*, Fei Kang, Mengyin An, Peiyu Wang, Biao Jiang, Yichen Wei, Yidan Xietian, Jiangbo Pei, Liang Hu, Boyi Jiang, Hua Xue, Zidong Wang, Haofeng Sun, Wei Li, Wanli Ouyang, Xianglong He, Yang Liu, Yangguang Li, Yahui Zhou (*equal)
Preprint, 2026
With the advancement of interactive video generation, diffusion models have increasingly demonstrated their potential as world models. However, existing approaches still struggle to simultaneously achieve memory-enabled long-term temporal consistency and high-resolution real-time generation, limiting their applicability in real-world scenarios. To address this, we present Matrix-Game 3.0, a memory-augmented interactive world model designed for 720p real-time long-form video generation. Building upon Matrix-Game 2.0, we introduce systematic improvements across data, model, and inference. First, we develop an upgraded industrial-scale infinite data engine that integrates Unreal Engine–based synthetic data, large-scale automated collection from AAA games, and real-world video augmentation to produce high-quality Video–Pose–Action–Prompt quadruplet data at scale. Second, we propose a training framework for long-horizon consistency: by modeling prediction residuals and re-injecting imperfect generated frames during training, the base model learns self-correction; meanwhile, camera-aware memory retrieval and injection enable the base model to achieve long horizon spatiotemporal consistency. Third, we design a multi-segment autoregressive distillation strategy based on Distribution Matching Distillation (DMD), combined with model quantization and VAE decoder pruning, to achieve efficient real-time inference. Experimental results show that Matrix-Game 3.0 achieves up to 40 FPS real-time generation at 720p resolution with a 5B model, while maintaining stable memory consistency over minute-long sequences. Scaling up to a 2×14B model further improves generation quality, dynamics, and generalization. Our approach provides a practical pathway toward industrial-scale deployable world models.
Error Forcing framework: robust causal pre-training + distillation
Direct Autoregressive Diffusion Distillation via Error-Aware Causal Pretraining
Jiaxing Li*, Kaichen Huang*, Baixin Xu*, Zexiang Liu, Xianglong He, Zile Wang, Junyao Gao, Yang Liu, Ying He, Bo An, Yangguang Li (*equal)
ECCV, 2026
Real-time autoregressive (AR) video diffusion has progressed rapidly. Existing approaches typically distill pretrained bidirectional video foundation models into few-step causal students; however, naive distillation often collapses due to architectural mismatch. To obtain stable initialization, prior methods rely on ordinary differential equation (ODE) trajectory matching, which incurs substantial computation and can introduce errors by forcing the student to imitate global trajectories. In this paper, we bypass trajectory matching stage by pretraining a robust causal AR model that equips the student for direct few-step distillation. Moreover, to bridge the gap between pretraining and distillation, we propose Error Forcing, an error-aware and parallelizable AR video diffusion framework. During training, Error Forcing injects controlled residual errors into the conditioning context, approximating the inference-time history distribution. This enables the causal model to learn from degraded yet clean contexts, improving robustness to accumulated errors without sacrificing parallel training throughput. With this initialization, we can perform direct distillation, and the student and teacher/critic naturally operate under consistent conditioning distributions. Extensive experiments show that our method outperforms existing baselines in generation quality during both the pre-training and distillation stages.
SafeWork-R1 safety vs capability benchmarks
SafeWork-R1: Coevolving Safety and Intelligence under the AI-45° Law
Shanghai AI Lab · Kaichen Huang (contributor)
Preprint, 2025
We introduce SafeWork-R1, a cutting-edge multimodal reasoning model that demonstrates the coevolution of capabilities and safety. It is developed by our proposed SafeLadder framework, which incorporates large-scale, progressive, safety-oriented reinforcement learning post-training, supported by a suite of multi-principled verifiers. Unlike previous alignment methods such as RLHF that simply learn human preferences, SafeLadder enables SafeWork-R1 to develop intrinsic safety reasoning and self-reflection abilities, giving rise to safety `aha' moments. Notably, SafeWork-R1 achieves an average improvement of 46.54% over its base model Qwen2.5-VL-72B on safety-related benchmarks without compromising general capabilities, and delivers state-of-the-art safety performance compared to leading proprietary models such as GPT-4.1 and Claude Opus 4. To further bolster its reliability, we implement two distinct inference-time intervention methods and a deliberative search mechanism, enforcing step-level verification. Finally, we further develop SafeWork-R1-InternVL3-78B, SafeWork-R1-DeepSeek-70B, and SafeWork-R1-Qwen2.5VL-7B. All resulting models demonstrate that safety and capability can co-evolve synergistically, highlighting the generalizability of our framework in building robust, reliable, and trustworthy general-purpose AI.
Explainable MLLM survey framework
Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey
Yunkai Dang*, Kaichen Huang*, Jiahao Huo*, Yibo Yan, Sirui Huang, Dongrui Liu, Mengxi Gao, Jie Zhang, Chen Qian, Kun Wang, Yong Liu, Jing Shao, Hui Xiong, Xuming Hu (*equal)
Preprint, 2024
The rapid development of Artificial Intelligence (AI) has revolutionized numerous fields, with large language models (LLMs) and computer vision (CV) systems driving advancements in natural language understanding and visual processing, respectively. The convergence of these technologies has catalyzed the rise of multimodal AI, enabling richer, cross-modal understanding that spans text, vision, audio, and video modalities. Multimodal large language models (MLLMs), in particular, have emerged as a powerful framework, demonstrating impressive capabilities in tasks like image-text generation, visual question answering, and cross-modal retrieval. Despite these advancements, the complexity and scale of MLLMs introduce significant challenges in interpretability and explainability, essential for establishing transparency, trustworthiness, and reliability in high-stakes applications. This paper provides a comprehensive survey on the interpretability and explainability of MLLMs, proposing a novel framework that categorizes existing research across three perspectives: (I) Data, (II) Model, (III) Training & Inference. We systematically analyze interpretability from token-level to embedding-level representations, assess approaches related to both architecture analysis and design, and explore training and inference strategies that enhance transparency. By comparing various methodologies, we identify their strengths and limitations and propose future research directions to address unresolved challenges in multimodal explainability. This survey offers a foundational resource for advancing interpretability and transparency in MLLMs, guiding researchers and practitioners toward developing more accountable and robust multimodal AI systems.
Separated world model (SeeX) framework
Leveraging Separated World Model for Exploration in Visually Distracted Environments
Kaichen Huang*, Shenghua Wan*, Minghao Shao, Shuai Feng, Le Gan, De-Chuan Zhan (*equal)
NeurIPS, 2024
Model-based unsupervised reinforcement learning (URL) has gained prominence for reducing environment interactions and learning general skills using intrinsic rewards. However, distractors in observations can severely affect intrinsic reward estimation, leading to a biased exploration process, especially in environments with visual inputs like images or videos. To address this challenge, we propose a bi-level optimization framework named Separation-assisted eXplorer (SeeX). In the inner optimization, SeeX trains a separated world model to extract exogenous and endogenous information, minimizing uncertainty to ensure task relevance. In the outer optimization, it learns a policy on imaginary trajectories generated within the endogenous state space to maximize task-relevant uncertainty. Evaluations on multiple locomotion and manipulation tasks demonstrate SeeX's effectiveness.
MINER modality-specific neurons concept
MINER: Mining the Underlying Pattern of Modality-Specific Neurons in Multimodal Large Language Models
Kaichen Huang, Jiahao Huo, Yibo Yan, Kun Wang, Yutao Yue, Xuming Hu
Preprint, 2024
In recent years, multimodal large language models (MLLMs) have significantly advanced, integrating more modalities into diverse applications. However, the lack of explainability remains a major barrier to their use in scenarios requiring decision transparency. Current neuron-level explanation paradigms mainly focus on knowledge localization or language- and domain-specific analyses, leaving the exploration of multimodality largely unaddressed. To tackle these challenges, we propose MINER, a transferable framework for mining modality-specific neurons (MSNs) in MLLMs, which comprises four stages: (1) modality separation, (2) importance score calculation, (3) importance score aggregation, (4) modality-specific neuron selection. Extensive experiments across six benchmarks and two representative MLLMs show that (I) deactivating ONLY 2% of MSNs significantly reduces MLLMs performance (0.56 to 0.24 for Qwen2-VL, 0.69 to 0.31 for Qwen2-Audio), (II) different modalities mainly converge in the lower layers, (III) MSNs influence how key information from various modalities converges to the last token, (IV) two intriguing phenomena worth further investigation, i.e., semantic probing and semantic telomeres. The source code is available at this URL.
SENSOR active vision concept
SENSOR: Imitate Third-Person Expert's Behaviors via Active Sensoring
Kaichen Huang*, Minghao Shao*, Shenghua Wan, Hai-Hang Sun, Shuai Feng, Le Gan, De-Chuan Zhan (*equal)
Preprint, 2024
In many real-world visual Imitation Learning (IL) scenarios, there is a misalignment between the agent's and the expert's perspectives, which might lead to the failure of imitation. Previous methods have generally solved this problem by domain alignment, which incurs extra computation and storage costs, and these methods fail to handle the hard cases where the viewpoint gap is too large. To alleviate the above problems, we introduce active sensoring in the visual IL setting and propose a model-based SENSory imitatOR (SENSOR) to automatically change the agent's perspective to match the expert's. SENSOR jointly learns a world model to capture the dynamics of latent states, a sensor policy to control the camera, and a motor policy to control the agent. Experiments on visual locomotion tasks show that SENSOR can efficiently simulate the expert's perspective and strategy, and outperforms most baseline methods.
DIDA method framework
DIDA: Denoised Imitation Learning based on Domain Adaptation
Kaichen Huang*, Hai-Hang Sun*, Shenghua Wan, Minghao Shao, Shuai Feng, Le Gan, De-Chuan Zhan (*equal)
Preprint, 2024
Imitating skills from low-quality datasets, such as sub-optimal demonstrations and observations with distractors, is common in real-world applications. In this work, we focus on the problem of Learning from Noisy Demonstrations (LND), where the imitator is required to learn from data with noise that often occurs during the processes of data collection or transmission. Previous IL methods improve the robustness of learned policies by injecting an adversarially learned Gaussian noise into pure expert data or utilizing additional ranking information, but they may fail in the LND setting. To alleviate the above problems, we propose Denoised Imitation learning based on Domain Adaptation (DIDA), which designs two discriminators to distinguish the noise level and expertise level of data, facilitating a feature encoder to learn task-related but domain-agnostic representations. Experiment results on MuJoCo demonstrate that DIDA can successfully handle challenging imitation tasks from demonstrations with various types of noise, outperforming most baseline methods.
IMF energy-spectrum error analysis
An Error Analysis Method for Externally Measured Data
Yihan Liu, Kaichen Huang, Yechao Bai
Measurement and Control Technology, 2023
We propose a method for analyzing external ballistic measurement errors caused by various conditions, which result in unobservable random errors and latent systematic errors. Using Intrinsic Mode Function (IMF) energy inflection points, this method categorizes IMFs into high-frequency errors, mixed information, and useful information, effectively compensating for systematic errors and improving positioning accuracy.

Experience

2025.08 – 2026.10
Skywork AI · Riemann Dynamics
Joint-Training Ph.D. Student · host: Yangguang Li, Wei Yu
Real-time interactive world models and video generation — Matrix-Game 3.0 / 3.5, DistillAlign, Error Forcing (ECCV 2026).
2024.11 – 2025.03
Shanghai AI Laboratory
Research Intern · host: Jing Shao
Multimodal safety and reasoning — SafeWork-R1.
2024.06 – 2024.09
HKUST (Guangzhou)
Research Assistant · host: Xuming Hu
Interpretability of multimodal LLMs — MINER, and a survey on explainable & interpretable MLLMs.
2023.06 – 2023.12
Ministry of Industry and Information Technology (MIIT)
Data Processing Engineer
Automated analysis system for industrial economic-trend data; resulted in a granted software patent.
2022.06 – 2023.03
Pengcheng Laboratory
Assistant Engineer · host: Tong Zhang
Medical imaging for congenital heart disease — a ViT / MAE-based ellipse-detection method for robust fetal echocardiography measurement.

Honors & Awards

Academic Service

Reviewer for NeurIPS, ICML, ICLR, and AAAI.