Qianru Sun lists you as their PhD student. Claim this profile to confirm it and keep the rest of your record right.
Claim this profileAcademic lineage
Traced back 16 generations →Advisors
Works12 from public data
- Benchmarking Foundation Models with Language-Model-as-an-Examiner249
A novel benchmarking framework where the LM serves as a knowledgeable examiner that formulates questions based on its knowledge and evaluates responses in a reference-free manner is proposed, allowing for effortless extensibility as various LMs can be adopted as the examiner and the questions can be constantly updated given more diverse trigger topics.
- Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks39
This survey probes the core challenges that the rise of LLMs poses for evaluation, and identifies and analyzes two pivotal transitions: from task-specific to capability-based evaluation and from manual to automated evaluation.
- Babel: Open Multilingual Large Language Models Serving Over 90% of Global Speakers22
Babel is introduced, an open multilingual LLM that covers the top 25 languages by number of speakers, supports over 90% of the global population, and includes many languages neglected by other open multilingual LLMs and sets a new standard for open multilingual LLMs.
- OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents6
The results show that skill availability does not guarantee effective skill usage, that the benefit of skill augmentation depends strongly on both the underlying model and the agent framework, and that many publicly popular skills do not consistently outperform base agents without skills.
- Automating Dataset Updates Towards Reliable and Timely Evaluation of Large Language Models5
This paper proposes to automate dataset updates for reliable and timely evaluation to generate unseen and high-quality testing samples based on existing ones to mitigate leakage issues and designs an extending strategy that adjusts the difficulty of the generated samples according to varying cognitive levels.
- 3
- 3
- 3
- 2
- 1
- –
- –
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.