Zhifei Xie Ph.D. Student, LV-Lab, Nanyang Technological University

Biography

I am a Ph.D. student at LV-Lab, Nanyang Technological University, advised by Prof. Shuicheng Yan and Prof. Chunyan Miao. I am also a two-time founder and CEO: I built and led two startups before and during my Ph.D., shipping DreamFactory as CEO of my first company and later founding Pask.ai, where our proactive agent work grew out of real product needs. Before that I received my bachelor's degree from Beijing Institute of Technology, studied at Tsinghua University, and interned at the Qwen Team under the supervision of Dr. Jin Xu.

My research focuses on large audio language models and real-time speech interaction — building models that hear, think and speak at the same time. I led the Mini-Omni series, the first open-source end-to-end speech-to-speech conversational models, whose architecture has since been adopted as a starting point by a generation of omni models, together with Audio-Reasoner, Mini-Omni-Reasoner, Mega-ASR, the Audio Interaction Model and VoiceMem. I hold research to one bar: it has to be worth doing. Going forward I keep pushing on always-on multimodal interaction, with a focus on interactivity, robustness, reasoning and proactivity.

Outside the lab I advise industry teams on speech technology, have played in symphony orchestras for 16 years, and love animals — many of my projects use a Shiba Inu avatar in memory of Guagua, who grew up with me.

605 Citations
GitHub Stars
Founder & CEO
#1 Paper of the Day on Hugging Face

News

Publications * equal contribution | full list on Google Scholar

VoiceMemNew
VoiceMem: A Streaming Dual-Brain Memory Architecture for Voice Agents
Zhifei Xie, et al.
Technical Report, 2026 (Streaming memory for real-time voice agents)

VoiceMem makes long-term memory native to the real-time voice loop. Its factual “left brain” and emotion/persona “right brain” retrieve compact memories while the user is still speaking, combining personalization with near-zero interaction overhead. The result is a practical path from stateless voice assistants to agents that remember, adapt and build continuity across conversations.

Audio Interaction Model
Audio Interaction Model
Zhifei Xie, Zihang Liu, Ze An, Xiaobin Hu, Yue Liao, Ziyang Ma, Dongchao Yang, Mingbao Lin, Deheng Ye, Shuicheng Yan, Chunyan Miao
arXiv, 2026 (First fully audio-driven real-time model)

Most voice assistants only react after a user speaks. Audio Interaction Model turns continuous sound into an always-on perceive–decide–respond loop, so a model can decide when to stay silent or proactively intervene. SoundFlow, StreamAudio-2M and Proactive-Sound-Bench make real-time proactive audio interaction a trainable and measurable research problem.

MMAE
MMAE: A Massive Multitask Audio Editing Benchmark
Ziyang Ma*, Ruiqi Yan*, Ruiyang Xu*, Jie Fang*, Zhikang Niu*, Yi-Wen Chao*, Wenming Tu*, Tianrui Wang*, …, Zhifei Xie, …, Kai Yu, Liefeng Bo, Eng-Siong Chng, Xie Chen
arXiv, 2026

MMAE provides the missing evaluation foundation for general-purpose instruction-based audio editing: 7 modalities, 6 complexity levels, 8 edit types and 2,000 curated samples. With current systems still below 5% exact-match on its hardest setting, the benchmark exposes how far audio editing remains from reliable, compositional control.

Mega-ASR
Mega-ASR: Towards In-the-wild² Speech Recognition via Scaling up Real-world Acoustic Simulation
Zhifei Xie, Kaiyu Pang, Haobin Zhang, Deheng Ye, Xiaobin Hu, Shuicheng Yan, Chunyan Miao
arXiv, 2026 (#1 on the FFASR Leaderboard)

Mega-ASR targets the gap between clean ASR benchmarks and deployment in noisy, reverberant, moving environments. By scaling physically grounded compound acoustic simulation, it cuts WER by more than 30% versus strong baselines in complex scenes and tops the FFASR far-field leaderboard — a step toward ASR that remains dependable in the real world.

Deep-Reporter
Deep-Reporter: Deep Research for Grounded Multimodal Long-Form Generation
Fangda Ye*, Kuicai Dong*, Zhifei Xie*, Yuxin Hu*, Yihang Yin, Shurui Huang, Shikai Dong, Chen Zhang, Jianzhu Bao, Shuicheng Yan
ACL, 2026

Deep-Reporter addresses a core weakness of research agents: long reports drift from evidence and usually treat figures as afterthoughts. Its multimodal search, checklist-guided synthesis and recurrent context management keep text, images and citations grounded together. M²LongBench makes this harder, more realistic form of long-form generation reproducibly measurable.

PASK
PASK: Toward Intent-Aware Proactive Agents with Long-Term Memory
Zhifei Xie (project leader), Zongzheng Hu, Fangda Ye, Xin Zhang, Haobo Chai, Zihang Liu, Pengcheng Wu, Guibin Zhang, Yue Liao, Xiaobin Hu, Deheng Ye, Chunyan Miao, Shuicheng Yan
arXiv, 2026 (Project leader · work done as CEO at Pask.ai)

PASK moves assistants from answering explicit requests toward anticipating latent needs under real latency constraints. Its demand detection, long-term memory and streaming IntentFlow components form an end-to-end proactive-agent stack, while LatentNeeds-Bench evaluates the setting on user-consented data. The work reframes proactivity as a deployable systems problem rather than a demo feature.

Mini-Omni-Reasoner
Mini-Omni-Reasoner: Token-Level Thinking-in-Speaking in Large Speech Models
Zhifei Xie, Ziyang Ma, Zihang Liu, Kaiyu Pang, Hongyu Li, Jialin Zhang, Yue Liao, Deheng Ye, Chunyan Miao, Shuicheng Yan
arXiv, 2025 (First model that reasons while speaking)

Mini-Omni-Reasoner removes the usual trade-off between reasoning quality and spoken latency by interleaving silent reasoning tokens with speech tokens in one stream. It improves Spoken-MQA arithmetic by +19.1% with no extra decoding latency, establishing thinking-in-speaking as a practical paradigm for reasoning-capable real-time voice models.

Audio-Reasoner
Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models
Zhifei Xie, Mingbao Lin, Zihang Liu, Pengcheng Wu, Shuicheng Yan, Chunyan Miao
EMNLP, 2025 (First audio reasoning model · Silver medal, MMAU)

Audio-Reasoner brought deliberate reasoning to large audio-language models, showing that structured reasoning training transfers beyond text to sound, music and speech. With the 1.2M-sample CoTA dataset, it achieved strong gains across MMAU-mini, AIR-Bench and MELD, helping establish audio reasoning as a distinct and reproducible research direction.

Cited by Qwen3-Omni · Audio Flamingo 3 · MMAR (NeurIPS 2025) · MMAU-Pro · SonicBench · referenced by the Interspeech 2026 Audio Reasoning Challenge

Mini-Omni2
Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities
Zhifei Xie, Changqiao Wu
arXiv, 2024 (First real-time audio-visual duplex model · 1.5k+ stars)

Mini-Omni2 was an early open-source reference for GPT-4o-style interaction: a single model that can see, hear and speak in real time, including mid-sentence interruption. Its duplex, end-to-end design showed that natural multimodal conversation did not have to remain exclusive to closed frontier systems.

Cited by Qwen2.5-Omni · Baichuan-Omni-1.5 · MiniCPM-o · VITA-1.5 · Step-Audio · EMOVA

Mini-Omni
Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
Zhifei Xie, Changqiao Wu
arXiv, 2024 (#1 Paper of the Day on Hugging Face · 3k+ stars)

Mini-Omni made direct end-to-end speech-to-speech interaction openly reproducible. Parallel speech generation lets the model hear, think and talk in one streaming pipeline instead of chaining ASR, an LLM and TTS. The work became a widely used architectural reference for the open omni-model wave and was later cited as a basis for Qwen2.5-Omni.

Cited by Qwen2.5-Omni · Kimi-Audio · Step-Audio · GLM-4-Voice · Baichuan-Omni-1.5 · MiniCPM-o · LLaMA-Omni2 · Freeze-Omni

DreamFactory
DreamFactory: Pioneering Multi-Scene Long Video Generation with a Multi-Agent Framework
Zhifei Xie, Daniel Tang, Dingwei Tan, Jacques Klein, Tegawendé F. Bissyandé, Saad Ezzini
Technical report, 2024 (Work done as founder & CEO)

DreamFactory reframed long-video generation as coordinated production rather than isolated clip sampling. Its multi-agent “film crew” separates scripting, storyboarding, direction and continuity so multiple scenes can stay coherent over time. The work was an early demonstration that agent orchestration could extend generative video from short clips toward structured, multi-scene narratives.

Other Publications

Activities

Awards & Honors
  • Gold Award, NTU Entrepreneurship Competition
  • National Scholarship, Ministry of Education of China
  • Tsinghua University Scholarship
  • #1 on the FFASR Leaderboard — Mega-ASR, 2026
  • Silver medal on the MMAU Leaderboard — Audio-Reasoner, 2025
  • #1 Paper of the Day on Hugging Face — Mini-Omni (2024) and Mega-ASR (2026)
Industry Advisory
Entrepreneurship
  • Founder & CEO, Pask — proactive, intent-aware AI agents with long-term memory
  • Founder & CEO, first startup — multi-agent long video generation (DreamFactory)
  • Research intern, Qwen Team — supervised by Dr. Jin Xu
  • Research intern, Inspirai — where the Mini-Omni series was started
Media Coverage
Academic Service
  • Reviewer for NeurIPS, ICLR, ICML, ACL, EMNLP, ACM MM, CVPR, ICCV
  • Reviewer for IEEE/ACM Transactions on Audio, Speech and Language Processing
Open Source
  • Maintainer of the Mini-Omni series, Audio-Reasoner, Mega-ASR and Audio-Interaction — thousands of GitHub stars and widely used as baselines in the speech-LM community
  • Open datasets: CoTA, Spoken-Math-Problems-3M, Voices-in-the-Wild-2M, StreamAudio-2M
Beyond Research
  • Symphonic musician for 16 years — orchestral performance is my other long-running project
  • A lifelong dog person; my projects carry a Shiba Inu avatar in memory of Guagua