Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks11 from public data
- Let's Think Outside the Box: Exploring Leap-of-Thought in Large Language Models with Creative Humor Generation101
A creative Leap-of-Thought (CLoT) paradigm is introduced to improve LLM's LoT ability and boosts creative abilities in various tasks like “cloud guessing game” and “divergent association task”.
- Consistent3D: Towards Consistent High-Fidelity Text-to-3D Generation with Deterministic Sampling Prior64
A novel and effective “Consistent3D” method that explores the ODE deterministic sampling prior for text-to-3D generation and designs a consistency distillation sampling loss which samples along the ODE trajectory to generate two adjacent samples and uses the less noisy sample to guide another more noisy one for distilling the deterministic prior into the 3D model.
- A Survey on Post-training of Large Language Models59
This paper presents the first comprehensive survey of PoLMs, systematically tracing their evolution across five core paradigms: Fine-tuning, which enhances task-specific accuracy; Alignment, which ensures ethical coherence and alignment with human preferences; Reasoning, which advances multi-step inference despite challenges in reward design; Efficiency, which optimizes resource utilization amidst increasing complexity; Integration and Adaptation, which extend capabilities across diverse modalities while addressing coherence issues.
- Diffusion Time-step Curriculum for One Image to 3D Generation26
The Diffusion Time-step Curriculum one-image-to-3D pipeline (DTC123), which involves both the teacher and student models collaborating with the time-step curriculum in a coarse-to-fine manner, and can produce multiview consistent, high-quality, and diverse 3D assets.
- Fit the Distribution: Cross-Image/Prompt Adversarial Attacks on Multimodal Large Language Models21
A novel adversarial attack on MLLMs is proposed based on distribution approximation theory, which models the potential image-prompt input distribution and adds the same distribution-fitting adversarial perturbation on multimodal input pairs to achieve effective cross-image/prompt transfer attacks.
- Towards Building Model/Prompt-Transferable Attackers against Large Vision-Language Models8
A new perspective of information theory is introduced to investigate LVLMs’ transferable characteristics by exploring the relative dependence between outputs of the LVLM model and input adversarial samples and formulate the complicated calculation of information gain as an estimation problem and incorporate such informative constraints into the adversarial learning process.
- 7
- 5
- 1
- 1
- –
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.