Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
Possible advisorsa guess from early papers, not confirmed
- Rui ZhaoSuggested from co-authorship
Is this you? Claim this profile to confirm or dismiss it.
Works24 from public data
- TimeCMA: Towards LLM-Empowered Multivariate Time Series Forecasting via Cross-Modality Alignment245
This work introduces TimeCMA, an intuitive yet effective framework for MTSF via cross-modality alignment that retrieves both disentangled and robust time series embeddings, ``the best of two worlds'', from the prompt embeddings based on time series and prompt modality similarities.
- Spatial-Temporal Large Language Model for Traffic Prediction234
Comprehensive experiments on real traffic datasets offer evidence that ST-LLM is a powerful spatial-temporal learner that outperforms state-of-the-art models and exhibits robust performance in both few-shot and zero-shot prediction scenarios.
- Self-Calibrated Cross Attention Network for Few-Shot Segmentation95
A self-calibrated cross attention (SCCA) block that takes a query patch as Q, and groups the patches from the same query image and the aligned patches from the support image as K&V for efficient patch-based attention, and a scaled-cosine mechanism to better utilize the support features for similarity calculation.
- 86
- Hybrid Mamba for Few-Shot Segmentation50
A cross (attention-like) Mamba is devised to capture inter-sequence dependencies for FSS, including a hybrid Mamba network (HMNet), including a support recapped Mamba to periodically recap the support features when scanning query, so the hidden state can always contain rich support information.
- Efficient Multivariate Time Series Forecasting via Calibrated Language Models with Privileged Knowledge Distillation45
TimeKD is introduced, an efficient MTSF framework that leverages the calibrated language models and privileged knowledge distillation to generate high-quality future representations from the proposed cross-modality teacher model and cultivate an effective student model.
- Road Extraction With Satellite Images and Partial Road Maps42
This article proposes a two-branch partial-to-complete network (P2CNet) for the road extraction task, which has two prominent components: gated self-attention module (GSAM) and missing part (MP) loss.
- POPEN: Preference-Based Optimization and Ensemble for LVLM-Based Reasoning Segmentation33
POPEN includes a preference-based optimization method to finetune the LVLM, aligning it more closely with human preferences and thereby generating better text responses and segmentation results.
- 21
- HHGT: Hierarchical Heterogeneous Graph Transformer for Heterogeneous Graph Representation Learning16
A novel Hierarchical Heterogeneous Graph Transformer (HHGT) model is proposed, which seamlessly integrates a Type-level Transformer for aggregating nodes of different types within each k-ring neighborhood, followed by a Ring-level Transformer for aggregating different k-ring neighborhoods in a hierarchical manner.
- Unlocking the Power of SAM 2 for Few-Shot Segmentation15
Pseudo Prompt Generator is designed to encode pseudo query memory, matching with query features in a compatible way, and Iterative Memory Refinement is designed to fuse more query FG features into the memory, and a Support-Calibrated Memory Attention to suppress the unexpected query BG features in memory.
- Traffic Speed Imputation with Spatio-Temporal Attentions and Cycle-Perceptual Training15
STCPA captures complex traffic correlations among the spatial and temporal dimensions via the attention mechanism, which helps mitigate the data sparsity issue and adopts an imputation cycle consistency constraint for providing reliable supervisions on unobserved entries, which improves the training.
- 11
- 7
- CamGeo: Sparse Camera-Conditioned Image-to-Video Generation with 3D Geometry Priors5
CamGeo is introduced, a novel framework that distills rich 3D geometric knowledge from a pre-trained video-to-3D model (VGGT) directly into the diffusion backbone, and progressively scales geometric complexity, from global structure coherence to fine-grained refinement, achieving stable optimization.
- 5
- HierPromptLM: A Pure PLM-based Framework for Representation Learning on Heterogeneous Text-rich Networks3
HierPromptLM, a novel pure PLM-based framework that seamlessly models both text data and graph structures without the need for separate processing, is proposed and two innovative HTRN-tailored pretraining tasks to fine-tune PLMs for representation learning are introduced.
- A Survey of Image Super Resolution Based on CNN2
A brief survey on the task of SISR is given, introducing the SR problem, some recent SR methods, public benchmark datasets and evaluation metrics, and denoting some points that could be further improved in the future.
- 1
- Explicit Intent-Enhanced Knowledge Distillation for Trip Recommendation1
EKD-Trip is proposed, an explicit intent-enhanced knowledge distillation framework that outperforms all baselines over three metrics, with a particularly notable improvement of 13.70% in pairs-F1.
- Language-Instructed Reasoning for Group Activity Detection via Multimodal Large Language Model1
This work proposes LIR-GAD, a novel framework of language-instructed reasoning for GAD via Multimodal Large Language Model (MLLM), and designs a Multimodal Dual-Alignment Fusion (MDAF) module that integrates MLLM's hidden embeddings corresponding to the designed tokens with visual features, significantly enhancing the performance of GAD.
- 1
- –
- –
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.