Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks15 from public data
- 12
- Situational Scene Graph for Structured Human-Centric Situation Understanding8
Experimental results show that the unified representation called SSG can not only benefit predicate classification and semantic role-value classification, but also benefit reasoning tasks on human-centric situation understanding.
- Image Retrieval Using Dominant Color Descriptor.6
A new method used to calculate the similarity of Dominant Color Descriptor using Earth Mover’s Distance (EMD) is used, and the lower bound is easier to implement and more efficient than the M-tree.
- Know-Show: Benchmarking Video-Language Models on Spatio-Temporal Grounded Reasoning5
Know-Show is presented, a new benchmark designed to evaluate spatio-temporal grounded reasoning, the ability of a model to reason about actions and their semantics while simultaneously grounding its inferences in visual and temporal evidence, establishes a unified standard for assessing grounded reasoning in video-language understanding.
- 3
- Next-Frame Feature Prediction for Multimodal Deepfake Detection and Temporal Localization2
This model introduces a window-level attention mechanism to capture discrepancies between predicted and actual frames, enabling the model to detect local artifacts around every frame, which is crucial for accurately classifying fully manipulated videos and effectively localizing deepfake segments in partially spoofed samples.
- 1
- 1
- –
- Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens–
This work examines the generation process and categorizes text tokens into three groups: image-positive, invariant, and negative, based on their visual dependence on input image tokens, and reveals that most generated tokens are minimally influenced by the image information.
- –
- –
- –
- –
- –
Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.