Yiping Ke lists you as their PhD student. Claim this profile to confirm it and keep the rest of your record right.
Claim this profileAcademic lineage
View as a treeAdvisors
Works27 from public data
- Text4Seg: Reimagining Image Segmentation as Text Generation70
Text4Seg is introduced, a novel text-as-mask paradigm that casts image segmentation as a text generation problem, eliminating the need for additional decoders and significantly simplifying the segmentation process.
- GeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual Grounding59
GeoGround is a novel framework that unifies support for HBB, OBB, and mask RS visual grounding tasks, allowing flexible output selection through the Text-Mask technique, and defines prompt-assisted and geometry-guided learning to enhance consistency across different signals.
- 52
- 51
- 48
- 45
- MIMO Is All You Need:A Strong Multi-in-Multi-Out Baseline for Video Prediction44
Surprisingly, empirical studies reveal that a simple MIMO model can outperform the state-of-the-art work with a large margin much more than expected, especially in dealing with long-term error accumulation.
- Contrasformer: A Brain Network Contrastive Transformer for Neurodegenerative Condition Identification25
Evaluated on 4 functional brain network datasets over 4 different diseases, Contrasformer outperforms the state-of-the-art methods for brain networks by achieving up to 10.8% improvement in accuracy, which demonstrates its efficacy in neurological disorder identification.
- 24
- 24
- 22
- Multi-Atlas Brain Network Classification Through Consistency Distillation and Complementary Information Fusion15
The Atlas-Integrated Distillation and Fusion network (AIDFusion), a novel framework designed to enhance brain network classification using fMRI data, is proposed, which demonstrates its superior classification performance and computational efficiency compared to state-of-the-art methods.
- Text4Seg++: Advancing Image Segmentation via Generative Language Modeling11
A novel text-as-mask paradigm that casts image segmentation as a text generation problem, eliminating the need for additional decoders and significantly simplifying the segmentation process is proposed, highlighting the effectiveness, scalability, and generalizability of text-driven image segmentation within the MLLM framework.
- Learning to Discover Knowledge: A Weakly-Supervised Partial Domain Adaptation Approach9
Self-paced transfer classifier learning (SP-TCL) learns to discover faithful knowledge via a carefully designed prudent loss function and simultaneously adapts the learned knowledge to the target domain by iteratively excluding source examples from training under the self-paced fashion.
- 5
- Multimodal mathematical reasoning embedded in aerial vehicle imagery: Benchmarking, analysis, and exploration3
This paper benchmarks 14 prominent VLMs through a comprehensive evaluation and demonstrates that, despite their success on previous multimodal benchmarks, these models struggle with the reasoning tasks in AVI-Math, the first benchmark to rigorously evaluate multimodal mathematical reasoning in aerial vehicle imagery.
- 2
- 2
- 1
- –
- –
- –
- –
- –
- NOCL: Node-Oriented Conceptualization LLM for Graph Tasks without Message Passing–
The Node-Oriented Conceptualization LLM (NOCL), a novel framework that leverages two core techniques that converts heterogeneous node attributes into structured natural language, extending LLM from TAGs to non-TAGs and significantly reducing token lengths, is proposed.
Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.