Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks9 from public data
- MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual Representations24
The proposed MedUnifier framework seamlessly integrates text-grounded image generation capabilities with multi-modal learning strategies, including image-text contrastive alignment, image-text matching and image-grounded text generation, and employs visual vector quantization, which enhances multi-modal generation quality.
- PVChat: Personalized Video Chat with One-Shot Learning14
PVChat is proposed, the first personalized ViLLM that enables subject-aware question answering from a single video for each subject, and adopts a two-stage training strategy, transitioning from image pre-training to video fine-tuning, enabling a gradual learning process from static attributes to dynamic representations.
- Edge-Guided and Cross-Scale Feature Fusion Network for Efficient Multi-contrast MRI Super-Resolution7
A novel edge-guided and cross-scale feature fusion network, namely ECFNet is proposed that achieves state-of-the-art performance compared to other multi-contrast MRI super-resolution methods, and the method is robust in terms of different super-resolution scales.
- 4
- VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine1
This work proposes a novel vision-language pre-training (VLP) framework, specifically designed for limited volumetric data such as 3D CT and associated radiology reports, and incorporates uni-modal self-supervised learning into VLP framework, which are often underexplored in the existing literature.
- A machine learning approach to conglomerate multi-domain features of cardiac aging–
Multi-domain ML identified clinically interpretable signals associated with impaired myocardial relaxation in ageing and with clinical events associated with death-or-admission events.
- –
- –
- Exploiting Low-Dimensional Manifold of Features for Few-Shot Whole Slide Image Classification–
A Squeeze-and-Recalibrate (SR) block is proposed, a drop-in replacement for linear layers in MIL models to address few-shot multiple instance learning challenges and provides theoretical guarantees that the SR block can approximate any linear mapping to arbitrary precision.
Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.