Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks89 from public data
- Annotated Point Clouds and Images from a Spatial-Semantic Perception Pipeline: Robotized Deconstruction at the Reference Construction Site, Aachen5,204
An open-set object detector, called Grounding DINO, is presented by marrying Transformer-based detector DINO with grounded pre-training, which can detect arbitrary objects with human inputs such as category names or referring expressions, and performs remarkably well on all three settings.
- Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks2,301
This paper proposes a new learning method Oscar (Object-Semantics Aligned Pre-training), which uses object tags detected in images as anchor points to significantly ease the learning of alignments.
- Grounded Language-Image Pre-training1,831
A grounded language-image pretraining model for learning object-level, language-aware, and semantic-rich visual representations that unifies object detection and phrase grounding for pre-training and can leverage massive image-text pairs by generating grounding boxes in a self-training fashion.
- Training Language Models to Self-Correct via Reinforcement Learning483
A multi-turn online reinforcement learning (RL) approach that significantly improves an LLM's self-correction ability using entirely self-generated data, and uses appropriate regularization to steer the learning process into learning a self-correction behavior that is effective at test time as opposed to fitting high-reward responses for a given prompt.
- Semi-supervised Left Atrium Segmentation with Mutual Consistency Training363
A novel Mutual Consistency Network (MC-Net) for semi-supervised left atrium segmentation from 3D MR images that consists of one encoder and two slightly different decoders, and the prediction discrepancies of two decoder are transformed as an unsupervised loss by a designed cycled pseudo label scheme to encourage mutual consistency.
- A Simple Framework for Open-Vocabulary Segmentation and Detection280
OpenSeeD is the first to explore the potential of joint training on segmentation and detection, and hope it can be received as a strong baseline for developing a single model for both tasks in the open world.
- SwinFuse: A Residual Swin Transformer Fusion Network for Infrared and Visible Images208
This work builds a fully attentional feature encoding backbone to model the long-range dependencies, which is a pure transformer network and has a stronger representation ability compared with the CNNs and designs a novel feature fusion strategy based on the $L_{1}$ -norm for sequence matrices.
- Tag2Text: Guiding Vision-Language Model via Image Tagging121
This paper presents Tag2Text, a vision language pre-training (VLP) framework, which introduces image tagging into vision-language models to guide the learning of visual-linguistic features and effectively enhances the performance of vision-language models on both generation-based and aligning tasks.
- TIGEr: Text-to-Image Grounding for Image Caption Evaluation86
The empirical tests show that TIGEr has a higher consistency with human judgments than alternative existing metrics, and comprehensively assess the metric’s effectiveness in caption evaluation by measuring the correlation between human judgments and metric scores.
- Vision-Language Models as a Source of Rewards76
This work investigates the feasibility of using off-the-shelf vision-language models, or VLMs, as sources of rewards for reinforcement learning agents and presents a scaling trend showing how larger VLMs lead to more accurate rewards for visual goal achievement, which in turn produces more capable RL agents.
- A Comprehensive Study of Multimodal Large Language Models for Image Quality Assessment73
A difficult sample selection procedure is presented, taking into account sample diversity and uncertainty, to further challenge MLLMs equipped with the respective optimal prompting systems for IQA.
- 73
- Neighbor discovery for wireless networks via compressed sensing69
A novel paradigm, called compressed neighbor discovery is proposed, which enables all nodes to simultaneously discover their respective neighborhoods with a single frame of transmission, which is in principle scalable to networks with billions of nodes with 48-bit IEEE 802.11 MAC addresses.
- VIVO: Visual Vocabulary Pre-Training for Novel Object Captioning64
The results show that the VIsual VOcabulary pre-training model can not only generate fluent image captions that describe novel objects, but also identify the locations of these objects.
- 62
- 57
- Vision-Language Intelligence: Tasks, Representation Learning, and Large Models44
The development in this field is summarized into three time periods, namely task-specific methods, vision-language pre-training (VLP) methods, and larger models empowered by large-scale weakly-labeled data.
- 41
- Boosting Human-Object Interaction Detection with Text-to-Image Diffusion Model37
DiffHOI is introduced, a novel HOI detection scheme grounded on a pre-trained text-image diffusion model, which enhances the detector's performance via improved data diversity and HOI representation, and SynHOI, a class-balance, large-scale, and high-diversity synthetic dataset containing over 140K HOI images with fully triplet annotations.
- 32
- 29
- Perceive, Understand and Restore: Real-World Image Super-Resolution with Autoregressive Multimodal Generative Models27
Experimental results demonstrate that PURE preserves image content while generating realistic details, especially in complex scenes with multiple objects, showcasing the potential of autoregressive multimodal generative models for robust Real-ISR.
- 26
- 23
- 20
- A fast indexing approach for protein structure comparison20
IR Tableau, a fast protein comparison algorithm, which leverages the tableau representation to compare protein tertiary structures, shows that it is possible to obtain very significant speedups for the protein structure comparison problem, by employing an information retrieval style approach for indexing proteins.
- 19
- 19
- Imputing Trip Purpose Based on GPS Travel Survey Data and Machine Learning Methods19
This paper explores the feasibility of automating trip purpose detection employing machine learning method with geospatial location data, the land use data, and the in-practice GPS-based survey conducted by University of Minnesota, and evaluates the impacts of different land use coding methods.
- X-Pose: Detecting Any Keypoints17
This work makes the first attempt to develop an end-to-end prompt-based keypoint detection framework called UniPose to detect keypoints of any objects and shows that UniPose has strong finegrained localization and generalization abilities across image styles, categories, and poses.
- 17
- 17
- 16
- 16
- 15
- MMESGBench: Pioneering Multimodal Understanding and Complex Reasoning Benchmark for ESG Tasks12
MMESGBench is a first-of-its-kind benchmark dataset targeted to evaluate multimodal understanding and reasoning across multi-source ESG documents, and initial experiments validate that multimodal and retrieval-augmented models substantially outperform text-only baselines.
- 12
- 11
- Language Guided Local Infiltration for Interactive Image Retrieval10
This work proposes a Language Guided Local Infiltration (LGLI) system, which fully utilizes the text information and penetrates text features into image features as much as possible and outperforms most state-of-the-art IIR approaches.
- Package Arrival Time Prediction via Knowledge Distillation Graph Neural Network9
A novel Knowledge Distillation Graph neural network-based package ETA prediction model, which uses knowledge distillation in the training phase to distill the knowledge of historical trajectories into OD pair embeddings and consistently outperforms existing state-of-the-art OD-based ETA prediction methods on three real-world Alibaba datasets.
- 9
- Breast cancer X-ray image staging: based on efficient net with multi-scale fusion and cbam attention9
Compared with other existing image classification algorithms, the proposed Efficientnet model based on the cbam attention mechanism has the highest accuracy, thus the researchers conclude that EfficientNet with CBAM and multi-scale fusion will improve the classification performance.
- 8
- High-quality blind defocus deblurring of multispectral images with optics and gradient prior8
This paper presents a blind defocus deblurring method that produces high-quality deblurred multispectral images with very good restoration on the obscure details with the advantages of the method over state-of-the-artdeblurring methods.
- Research on frequency inertia response control strategy of SCESS‐DFIG system considering variable wind speed8
In this study, the inertia of a doubly fed wind generator with a super capacitor energy storage system (SCESS-DFIG) is defined from the standpoint of a single machine and an inertia response control strategy of multi energy coordinated control is proposed at the wind field level.
- 7
- 7
- Optimal Resource Allocation in URLLC for Real-Time Wireless Control Systems7
This paper investigates the resource allocation for URLLC uplink in real-time wireless control systems and proposes an iteration algorithm to obtain the optimal wireless resource allocation.
- 6
- 6
- 6
- 5
- 5
- 5
- A Novel Network Protocol Syntax Extracting Method for Grammar-Based Fuzzing4
A novel method for extracting syntax information from Wireshark protocol dissector files is proposed, which innovatively extracts network protocol syntax from Wireshark protocol dissector files and ensures the comprehensiveness and accuracy of the extracted syntax information.
- Double Circulation Wear Leveling for PCM-Based Embedded Systems4
A simple, novel, and effective wear leveling technique to evenly distribute write activities across the PCM chips, called Double Circulation Wear Leveling (DCWL), is proposed.
- 4
- 4
- 3
- 3
- 3
- 2
- 2
- Hyperspectral image super-resolution based on attention ConvBiLSTM network2
A hyperspectral (HS) image super-resolution (SR) approach based on attention convolutional bi-long short-term memory (ConvBiLSTM) network is proposed, aiming to explore the collaborative spatial and spectral attention characteristics, thereby enhancing the spatial resolution of HS image.
- 2
- 2
- 2
- 2
- 1
- Guided Image Restoration via Simultaneous Feature and Image Guided Fusion1
A Simultaneous Feature and Image Guided Fusion (SFIGF) network, that simultaneously considers feature and image-level guided fusion following the guided filter (GF) mechanism, is proposed.
- Review of Tibetan Medical Data Mining1
A novel decision support framework of Tibetan medicine treatment innovatively based on the main problems of data mining is put forward, which aims to realize the personalized treatment of Tibetan Medicine and provide effective support for the scientific treatment of common diseases of Tibetan plateau further.
- 1
- 1
- –
- A Review of the Research Progress and Application of Skin Care Ointments in Chronic Eczema Treatment–
- –
- –
- –
- –
- 2D-VMD Embedded Fusion of Infrared Polarization and Intensity Images Using Muitiple-Algorithms Based on Their Complementary Relation–
A new multiple-algorithm embedded fusion of infrared polarization and intensity images based on the complementary relation of the algorithms is proposed, which can clearly improve the fusion performance of multiple embedded infrared polarized images and generate a better image fusion.
- –
- –
- A Prediction Method of Protein Disulfide Bond Based on Hybrid Strategy–
The result of simulation experiment shows that the prediction model based on support vector ma-chine and sample selection can increase the prediction accuracy of protein disulfide bond.
- Statistics-based reconstruction method with high random-error tolerance for integral imaging–
It can be verified that the proposed method can effectively reduce the impacts of random errors on 3D reconstruction of integral imaging and has relatively stable and better reconstruction accuracy than the conventional reconstruction method.
- Design and Implementation of the Simulated Taxi Scheduling System Based on GPS–
The purpose of this thesis is the introduction of a higher degree of automation of the taxi scheduling system to reduce no-load operation, traffic congestion and waste of resources, improve the efficiency of taxi operators and design a taxi real-time scheduling system.
- –
- Exclusive Communication Mechanism of Hybrid Architecture Indoor Location System for WSN–
Combining advantages of centralized and decentralized architecture helps to resolve the problem among signals interferences and collisions of different beacons effectively.
- –
- A system to enable relational persistence and semantic web style access simultaneously for objects–
A method and system is presented which elegantly generates relational schema, OWL ontology, and semantic mapping between them, for any given OO model, and enables relational persistence and Semantic Web style access simultaneously for objects.
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-10. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.