Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks14 from public data
- Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels664
The proposed Q-Align achieves state-of-the-art performance on image quality assessment (IQA), image aesthetic assessment (IAA), as well as video quality assessment (VQA) tasks under the original LMM structure and unify the three tasks into one model, termed the OneAlign.
- Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical Perspectives463
The Disentangled Objective Video Quality Evaluator (DOVER) is proposed, the first approach to provide reliable clear-cut quality evaluations from a single aesthetic or technical perspective, and the first approach to provide reliable clear-cut quality evaluations from a single aesthetic or technical perspective.
- Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision309
Q-Bench, a holistic benchmark crafted to systematically evaluate potential abilities of MLLMs on three realms: low-level visual perception, low-level visual description, and overall visual quality assessment, confirms that MLLMs possess preliminary low-level visual skills.
- Q-Instruct: Improving Low-Level Visual Abilities for Multi-Modality Foundation Models218
The first dataset consisting of human natural language feedback on low-level vision, and a GPT-participated transformation to convert these feedbacks into a rich set of 200K instruction-response pairs, termed Q-Instruct are collected.
- Towards Explainable In-the-Wild Video Quality Assessment: A Database and a Language-Prompted Approach80
The MaxVQA is proposed, a language-prompted VQA approach that modifies vision-language foundation model CLIP to better capture important quality issues as observed in the authors' analyses.
- Surgical SAM 2: Real-time Segment Anything in Surgical Video by Efficient Frame Pruning67
This work introduces Surgical SAM 2 (SurgSAM2), an advanced model to utilize SAM2 with an Efficient Frame Pruning (EFP) mechanism, to facilitate real-time surgical video segmentation in resource-constrained environments, and significantly improves both efficiency and segmentation accuracy.
- SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence44
This work proposes SurgVLM, one of the first large vision-language foundation models for surgical intelligence, where this single universal model can tackle versatile surgical tasks and builds upon Qwen2.5-VL, which is built upon Qwen2.5-VL and undergoes instruction tuning to 10+ surgical tasks.
- 31
- Q-Boost: On Visual Quality Assessment Ability of Low-Level Multi-Modality Foundation Models30
Q-Boost is introduced, a novel strategy designed to enhance low-level MLLMs in image quality assessment (IQA) and video quality assessment (VQA) tasks, which is structured around two pivotal components: Triadic-Tone Integration and Multi-Prompt Ensemble.
- Q-Bench+: A Benchmark for Multi-Modal Foundation Models on Low-Level Vision From Single Images to Pairs28
It is demonstrated that several MLLMs have decent low-level visual competencies on single images, but only GPT-4V exhibits higher accuracy on pairwise comparisons than single image evaluations than single image evaluations like humans.
- 26
- A JND Guided Foveation Video Coding2
A novel just noticeable distortion (JND) guided foveation video coding method, by which the foveated region can be adaptively selected according to the video content, is presented.
- Abstract PO4-07-03: Machine learning breast cancer risk prediction using sequential past mammograms - a pilot study1
The study showed that deep learning breast cancer risk prediction can be further improved by using sequential past mammograms instead of random mammograms for AI model training and can also potentially enhance other AI risk prediction models that employ combined mammogram and traditional breast cancer clinical risk factors.
- –
Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.