Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
View as a treeAdvisors
- Sun AixinFrom a thesis record ↗
Works35 from public data
- A Unified Framework for Multi-Domain CTR Prediction via Large Language Models43
Uni-CTR leverages Large Language Model (LLM) to extract layer-wise semantic representations that capture domain commonalities, mitigating the seesaw phenomenon and enhancing generalization, and incorporates a pluggable domain-specific network to capture domain characteristics, ensuring scalability to dynamic domain growth.
- Collaborative Cross-modal Fusion with Large Language Model for Recommendation30
This work proposes a framework of Collaborative Cross-modal Fusion with Large Language Models, termed CCF-LLM, for recommendation that outperforms existing methods by effectively utilizing semantic and collaborative signals in the LLM4Rec context.
- Aligning Crowd Feedback via Distributional Preference Reward Modeling26
The Distributional Preference Reward Model (DPRM) is proposed, a simple yet effective framework to align large language models with diverse human preferences and introduces a Bayesian updater to accommodate shifted or new preferences.
- CtrlA: Adaptive Retrieval-Augmented Generation via Inherent Control22
This work presents the first attempts to solve adaptive RAG from a representation perspective and develops an inherent control-based framework, termed \name, which is superior to existing adaptive RAG methods on a diverse set of tasks.
- DocOIE: A Document-level Context-Aware Dataset for OpenIE19
This work manually annotates 800 sentences from 80 documents in two domains to form a DocOIE dataset and proposes DocIE, a novel document-level context-aware OpenIE model, to demonstrate that incorporating document- level context is helpful in improving OpenIE performance.
- 17
- Doc-Researcher: A Unified System for Multimodal Document Parsing and Deep Research16
Doc-Researcher is presented, a unified system that bridges the gap between deep parsing that preserves layout structure and visual semantics while creating multi-granular representations from chunk to document level and systematic retrieval architecture supporting text-only, vision-only, and hybrid paradigms with dynamic granularity selection.
- Reinforcement Learning Foundations for Deep Research Systems: A Survey13
This survey is, to the authors' knowledge, the first dedicated to the RL foundations of deep research systems, systematizes recent work along three axes, and distill recurring patterns, surface infrastructure bottlenecks, and offer practical guidance for training robust, transparent deep research agents with RL.
- To Search or Not to Search: Aligning the Decision Boundary of Deep Search Agents via Causal Intervention8
This work introduces causal intervention-based diagnosis that identifies boundary errors by comparing factual and counterfactual trajectories at each decision point, and develops Decision Boundary Alignment for Deep Search agents (DAS), which constructs preference datasets from causal feedback and aligns policies via preference optimization.
- 8
- 8
- Exploring Recommender System Evaluation: A Multi-Modal LLM Agent Framework for A/B Testing7
A multi-modal user agent for A/B testing (A/B Agent) is introduced, enabling multimodal and multi-page interactions that align with real user behavior on online platforms and finding that the data generated by A/B Agent can effectively enhance the capabilities of recommendation models.
- 7
- 6
- LLMTreeRec: Unleashing the Power of Large Language Models for Cold-Start Recommendations6
This work introduces an LLM-based end-to-end recommendation framework that boasts high efficiency, reducing the input token need by 86% compared to existing LLM-based models, and proposes a novel strategy to structure all items into an item tree, which can be dynamically updated and effectively retrieved.
- Multi-view Content-aware Indexing for Long Document Retrieval5
The Multi-view Content-aware indexing (MC-indexing) is proposed for more effective long DocQA via (i) segment structured document into content chunks, and (ii) represent each content chunk in raw-text, keywords, and summary views.
- SRR-Judge: Step-Level Rating and Refinement for Enhancing Search-Integrated Reasoning in Search Agents4
SRR-Judge is introduced, a framework for reliable step-level assessment of reasoning and search actions and delivers more reliable step-level evaluations than much larger models such as DeepSeek-V3.1, with its ratings showing strong correlation with final answer correctness.
- 4
- UniASA: A Unified Generative Framework for Argument Structure Analysis4
A unified generative framework for argument structure analysis (UniASA) that can uniformly address multiple argument structure analysis tasks in a sequence-to-sequence manner and can be effectively integrated with large language models, such as Llama, through fine-tuning or in-context learning.
- 4
- 3
- Shall We Trust All Relational Tuples by Open Information Extraction? A Study on Speculation Detection2
This paper formally defines the research problem of tuple-level speculation detection and conducts a detailed data analysis on the LSOIE dataset which contains labels for speculative tuples to determine whether an extracted tuple is speculative.
- FollowTable: A Benchmark for Instruction-Following Table Retrieval1
The results indicate that existing retrieval models struggle to follow fine-grained instructions over tabular data, and exhibit systematic biases toward surface-level semantic cues and remain limited in handling schema-grounded constraints, highlighting substantial room for future improvements.
- Learning More from Less: Exploiting Counterfactuals for Data-Efficient Chart Understanding1
ChartCF is introduced, a data-efficient training framework designed to enhance counterfactual sensitivity in Vision-Language Models and achieves superior or comparable performance to strong chart-specific VLMs while using significantly less training data.
- When Hard Negatives Hurt: Bridging the Generative Discriminative Gap in Hard Negative Synthesis for Retrieval1
Experiments show that CausalNeg outperforms mining-only and naïve generation baselines, validating causally grounded synthesis and entropy-regularized training as complementary solutions to the generative–discriminative gap.
- 1
- DOCoR: Document-level OpenIE with Coreference Resolution1
This demonstration presents a system which refines the semantic tuples generated by OpenIE with the aid of a coreference resolution tool, and is able to resolve both anaphoric and cataphoric references, to achieve Document-levelopenIE with Coreference Resolution (DOCoR).
- 1
- –
- –
- –
- –
- –
- –
- –
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-10. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.