Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
Possible advisorsa guess from early papers, not confirmed
- Tianwei ZhangSuggested from co-authorship
Is this you? Claim this profile to confirm or dismiss it.
Works13 from public data
- Instruction Tuning for Large Language Models: A Survey955
This work makes a systematic review of the literature, including the general methodology of IT, the construction of IT datasets, the training of IT models, and applications to different modalities, domains and application, along with analysis of aspects that influence the outcome of IT.
- VideoShield: Regulating Diffusion-based Video Generation Models via Watermarking39
VideoShield is a novel watermarking framework specifically designed for popular diffusion-based video generation models that effectively extracts watermarks and detects tamper without compromising video quality and is applicable to image generation models, enabling tamper detection in generated images as well.
- 14
- When Memory Becomes a Vulnerability: Towards Multi-turn Jailbreak Attacks against Text-to-Image Generation Systems8
This paper proposes Inception, the first multi-turn jailbreak attack against the memory mechanism in real-world text-to-image generation systems, which embeds the malice at the inception of the chat session turn by turn, leveraging the mechanism that T2I generation systems retrieve key information in their memory.
- Inference-time Alignment via Sparse Junction Steering3
This work shows that dense intervention is unnecessary and proposes Sparse Inference time Alignment (SIA), which performs sparse junction steering by intervening only at critical decision points along the generation trajectory, and reduces computational cost by up to 6x.
- 3
- 2
- Unifying Watermarking via Dimension-Aware Mapping1
DiM, a new multi-dimensional watermarking framework that formulates watermarking as a dimension-aware mapping problem, thereby unifying existing watermarking methods at the functional level is proposed.
- Emage: Non-Autoregressive Text-to-Image Generation1
This work explores non-autoregressive text-to-image models that efficiently generate hundreds of image tokens in parallel and runs 16 times to generate images of competitive quality with an order of magnitude lower inference latency.
- –
- –
- TRACE: Structure-Aware Character Encoding for Robust and Generalizable Document Watermarking–
TRACE achieves broad generalizability across multiple languages and fonts, making it particularly suitable for practical document security applications, and superior performance over state-of-the-art methods.
- Visible Yet Unreadable: A Systematic Blind Spot of Vision Language Models Across Writing Systems–
This paper constructs two psychophysics inspired benchmarks across distinct writing systems, Chinese logographs and English alphabetic words, by splicing, recombining, and overlaying glyphs to yield visible but unreadablestimuli for models while remaining legible to humans.
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.