Is this you? Claim this profile to correct it, add a bio and choose the work people see first.

Claim this profile

Academic lineage

Possible advisorsa guess from early papers, not confirmed

  • Erik Cambria

    Possible advisor · last author on 8 of their early first-author papers, 2021–2023

    Suggested from co-authorship

Is this you? Claim this profile to confirm or dismiss it.

Works26 from public data

TitleCited by
  • Recent advances in deep learning based dialogue systems: a systematic survey

    Jinjie Ni, Tom Young, Vlad Pandelea, Fuzhao Xue, Erik Cambria

    Artificial Intelligence Review · 2022

    This survey is the most comprehensive and up-to-date one at present for deep learning based dialogue systems, extensively covering the popular techniques.

    360
  • OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models

    Fuzhao Xue, Zian Zheng, Yao Fu, Jinjie Ni, Zangwei Zheng, Wangchunshu Zhou, Yang You

    arXiv · 2024

    This investigation confirms that MoE-based LLMs can offer a more favorable cost-effectiveness trade-off than dense LLMs, highlighting the potential effectiveness for future LLM development and proposes potential strategies for mitigating the issues found and further improving off-the-shelf MoE LLM designs.

    232
  • Fusing Task-Oriented and Open-Domain Dialogues in Conversational Agents

    Tom Young, Frank Xing, Vlad Pandelea, Jinjie Ni, Erik Cambria

    Proceedings of the AAAI Conference on Artificial Intelligence · 2022

    A new dataset, based on the popular TOD dataset MultiWOZ, is built, by rewriting the existing TOD turns and adding new ODD turns, and it features inter-mode contextual dependency, i.e., the dialogue turns from the two modes depend on each other.

    66
  • Diffusion Language Models are Super Data Learners

    Jinjie Ni, Qian Liu, Longxu Dou, Du, Chao, Zili Wang, Hang Yan, Tianyu Pang, Michael Shieh

    arXiv · 2025

    A Crossover: when unique data is limited, diffusion language models (DLMs) consistently surpass autoregressive (AR) models by training for more epochs by attributing the gains to three compounding factors: any-order modeling, super-dense compute from iterative bidirectional denoising, and built-in Monte Carlo augmentation.

    54
  • A survey on semantic processing techniques

    Rui Mao, Kai He, Xulang Zhang, Guanyi Chen, Jinjie Ni, Zonglin Yang, Erik Cambria

    Information Fusion · 2023

    This survey analyzed five semantic processing tasks, e.g., word sense disambiguation, anaphora resolution, named entity recognition, concept extraction, and subjectivity detection, to compare the different semantic processing techniques and summarize their technical trends, application trends, and future directions.

    54
  • Logical Reasoning over Natural Language as Knowledge Representation: A Survey

    Zonglin Yang, Xinya Du, Rui Mao, Jinjie Ni, Erik Cambria

    arXiv · 2023

    This paper provides a comprehensive overview on a new paradigm of logical reasoning, which uses natural language as knowledge representation and pretrained language models as reasoners, including philosophical definition and categorization of logical reasoning.

    41
  • MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP Use

    Zijian Wu, Liu, Xiangyan, Xinyuan Zhang, Chen, Lingjun, Fanqing Meng, Du, Lingxiao, Yiran Zhao, Zhang, Fanshi, +7 more

    arXiv · 2025

    A comprehensive evaluation of cutting-edge LLMs using a minimal agent framework that operates in a tool-calling loop, significantly surpassing those in previous MCP benchmarks and highlighting the stress-testing nature of MCPMark.

    40
  • HiTKG: Towards Goal-Oriented Conversations via Multi-Hierarchy Learning

    Jinjie Ni, Vlad Pandelea, Tom Young, Haicang Zhou, Erik Cambria

    Proceedings of the AAAI Conference on Artificial Intelligence · 2022

    This work presents HiTKG, a hierarchical transformer-based graph walker that leverages multiscale inputs to make precise and flexible predictions on KG paths and proposes MetaPath as the backbone method for KG path representation to exploit the entity and relation information concurrently.

    35
  • An Embarrassingly Simple Model for Dialogue Relation Extraction

    Fuzhao Xue, Aixin Sun, Hao Zhang, Jinjie Ni, Eng Siong Chng

    IEEE International Conference on Acoustics Speech and Signal Processing · 2022

    A simple yet effective model named SimpleRE is proposed for the RE task, which captures the interrelations among multiple relations in a dialogue through a novel input format named BERT Relation Token Sequence.

    32
  • Training Optimal Large Diffusion Language Models

    Jinjie Ni, Qian Liu, Chao Du, Longxu Dou, Hang Yan, Zili Wang, Tianyu Pang, Michael Shieh

    arXiv · 2025

    24
  • De’hubert: Disentangling Noise in a Self-Supervised Model for Robust Speech Recognition

    Dianwen Ng, Ruixi Zhang, Jia Qi Yip, Zhao Yang, Jinjie Ni, Chong Zhang, Yukun Ma, Chongjia Ni, +2 more

    IEEE International Conference on Acoustics Speech and Signal Processing · 2023

    A novel training framework, called deHuBERT, is proposed for noise reduction encoding inspired by H. Barlow’s redundancy-reduction principle, which improves the HuBERT training algorithm by introducing auxiliary losses that drive the self- and cross-correlation matrix between pairwise noise-distorted embeddings towards identity matrix.

    22
  • Finding the Pillars of Strength for Multi-Head Attention

    Jinjie Ni, Rui Mao, Zonglin Yang, Han Lei, Erik Cambria

    Annual Meeting of the Association for Computational Linguistics (ACL) · 2023

    12
  • Unnatural Languages Are Not Bugs but Features for LLMs

    Keyu Duan, Yiran Zhao, Zhili Feng, Jinjie Ni, Tianyu Pang, Qian Liu, Tianle Cai, Longxu Dou, +4 more

    arXiv · 2025

    This work demonstrates that unnatural languages - strings that appear incomprehensible to humans but maintain semantic meanings for LLMs - contain latent features usable by models, and demonstrates that models fine-tuned on unnatural versions of instruction datasets perform on-par with those trained on natural language.

    8
  • MixEval-X: Any-to-Any Evaluations from Real-World Data Mixtures

    Jinjie Ni, Yifan Song, Deepanway Ghosal, Bo Li, D. Zhang, Yue, Xiang, Fuzhao Xue, Zheng, Zian, +5 more

    arXiv · 2024

    This work introduces MixEval-X, the first any-to-any, real-world benchmark designed to optimize and standardize evaluations across diverse input and output modalities, and proposes multi-modal benchmark mixture and adaptation-rectification pipelines to reconstruct real-world task distributions.

    6
  • A class-aware representation refinement framework for graph classification

    Jiaxing Xu, Jinjie Ni, Yiping Ke

    Information Sciences · 2024

    6
  • Adaptive Knowledge Distillation Between Text and Speech Pre-Trained Models

    Jinjie Ni, Yukun Ma, Wen Wang, Qian Chen, Dianwen Ng, Lei Han, Trung Hieu Nguyen, Chong Zhang, +2 more

    IEEE International Conference on Acoustics Speech and Signal Processing · 2023

    This paper proposes the Prior-informed Adaptive knowledge Distillation (PAD) that adaptively leverages text/speech units of variable granularity and prior distributions to achieve better global and local alignments between text and speech pre-trained models.

    6
  • 3
  • Recent Advances in Deep Learning-based Dialogue Systems

    Jinjie Ni, Tom Young, Vlad Pandelea, Fuzhao Xue, V. Ananth Krishna Adiga, Erik Cambria

    2021

    3
  • On Development of An Autonomous Ball Collecting Wheeled Mobile Robot

    Hongyi Nie, Guodong Cai, Baorong Liu, Xinyu Hu, Feng Yang, Jinjie Ni, Xiaolei Hou

    2019 3rd Conference on Vehicle Control and Intelligence (CVCI) · 2019

    3
  • ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

    Yujie Liu, Zonglin Yang, Xie Tong, Jinjie Ni, Ben Gao, Yuqiang Li, Shixiang Tang, Wanli Ouyang, +2 more

    Findings of the Association for Computational Linguistics: ACL 2026 · 2026

    2
  • Boosting LLM via Learning from Data Iteratively and Selectively

    Jia, Qi, Siyu Ren, Ziheng Qin, Fuzhao Xue, Jinjie Ni, Yang You

    arXiv · 2024

    This work proposes to perform instruction tuning by iterative data selection by iteratively updating the complexity score for the top-ranked samples and greedily selecting the ones with the highest complexity-diversity score.

    1
  • Auxiliary Pooling Layer For Spoken Language Understanding

    Yukun Ma, Trung Hieu Nguyen, Jinjie Ni, Wen Wang, Qian Chen, Chong Zhang, Bin Ma

    IEEE International Conference on Acoustics Speech and Signal Processing · 2023

    1
  • SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis

    Zijian Wu, Jinjie Ni, Xiangyan Liu, Zichen Liu, Hang Yan, Michael Shieh

    Findings of the Association for Computational Linguistics: ACL 2026 · 2026

    –
  • NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation

    Xiangyan Liu, Jinjie Ni, Zijian Wu, Chao Du, Longxu Dou, Haonan Wang, Tianyu Pang, Michael Shieh

    neural information processing systems · 2025

    –
  • MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures

    Jinjie Ni, Fuzhao Xue, Xiang Yue, Yuntian Deng, Mahir Shah, Kabir Jain, Graham Neubig, Yang You

    neural information processing systems · 2024

    –

Show all 26 works

Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.

Report an error

Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.