Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
Traced back 37 generations →Advisors
- Mohit BansalFrom Academic Family Tree ↗
- Sung Ju HwangFrom Academic Family Tree ↗
Works52 from public data
- Analyzing and Mitigating Object Hallucination in Large Vision-Language Models381
This work proposes a simple yet powerful algorithm, LVLM Hallucination Revisor (LURE), to post-hoc rectify object hallucination in LVLMs by reconstructing less hallucinatory descriptions and consistently ranks at the top in both GPT and human evaluations.
- VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos319
VideoTree is proposed, a training-free framework which builds a query-adaptive and hierarchical video representation for LLM reasoning over long-form videos that outperforms existing training-free approaches on EgoSchema and NExT-QA with less inference time.
- Federated Semi-Supervised Learning with Inter-Client Consistency.290
FedMatch improves upon naive federated semi-supervised learning approaches with a new inter-client consistency loss and decomposition of the parameters into parameters for labeled and unlabeled data.
- 178
- Multimodal Representation Learning by Alternating Unimodal Adaptation155
MLA reframes the conventional joint multimodal learning process by transforming it into an al-ternating unimodal learning process, thereby minimizing interference between modalities and demonstrating the superiority of MLA over competing prior approaches.
- Combined Group and Exclusive Sparsity for Deep Neural Networks154
This work proposes an exclusive sparsity regularization based on (1, 2)-norm, which promotes competition for features between different weights, thus enforcing them to fit to disjoint sets of features, and combines theexclusive sparsity with the group sparsity, to promote both sharing and competition for Features in training of a deep neural network.
- SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation145
This work proposes SAFREE, a novel, training-free approach for safe T2I and T2V, that does not alter the model's weights and incorporates a novel self-validating filtering mechanism that dynamically adjusts the denoising steps when applying the filtered embeddings.
- 63
- EnvGen: Generating and Adapting Environments via LLMs for Training Embodied Agents52
It is found that a small RL agent trained with EnvGen can outperform SOTA methods, including a GPT-4 agent, and learns long-horizon tasks significantly faster and how the environments are adapted to help improve RL agents' weaker skills over time is shown.
- BECoTTA: Input-dependent Online Blending of Experts for Continual Test-time Adaptation40
This paper proposes BECoTTA, an input-dependent and efficient modular framework for CTTA that contains Mixture-of Domain Low-rank Experts (MoDE) that contains two core components: Domain-Adaptive Routing, which helps to selectively capture the domain adaptive knowledge with multiple domain routers, and Domain-Expert Synergy Loss to maximize the dependency between each domain and expert.
- 39
- ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models27
Efficient Coarse-to-Fine LayerWise LayerWise Pruning (ECoFLaP), a two-stage coarse-to-fine weight pruning approach for LVLMs that validate the proposed method across various multimodal and unimodal models and datasets, demonstrating significant performance improvements over prevalent pruning techniques in the high-sparsity regime.
- EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance23
EPiC is introduced, an efficient and precise camera control learning framework that constructs well-aligned training anchor videos without the need for camera pose or point cloud estimation, and generalizes robustly to anchor videos made with point clouds at test time, enabling precise 3D-informed camera control.
- 23
- Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation20
Av Avatar Forcing is proposed, a new framework for interactive head avatar generation that models real-time user-avatar interactions through diffusion forcing and introduces a direct preference optimization method that leverages synthetic losing samples constructed by dropping user conditions, enabling label-free learning of expressive interaction.
- Federated Continual Learning with Adaptive Parameter Communication.18
This work proposes a novel federated continual learning framework, Federated continualLearning with Adaptive Parameter Communication, which additively decomposes the network weights into global shared parameters and sparse task-specific parameters and allows inter-client knowledge transfer by communicating the sparse Task Specific parameters.
- 17
- EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding13
EgoMemReason is introduced, a comprehensive benchmark for week-long egocentric video understanding through memory-driven reasoning that evaluates three complementary memory types: entity memory, tracking how object states evolve and change across days; event memory, recalling and ordering activities separated by hours or days; and behavior memory, abstracting recurring patterns from sparse, repeated observations over the whole week period.
- 13
- Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement12
VideoRepair is introduced, the first self-correcting, training-free, and model-agnostic video refinement framework that automatically detects fine-grained text-video misalignments and performs targeted, localized corrections.
- DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation12
This work proposes DreamRunner, a novel story-to-video generation method that presents retrieval-augmented test-time adaptation to capture target motion priors for objects in each scene, supporting diverse motion customization based on retrieved videos, thus facilitating the generation of new videos with complex, scripted motions.
- Frame Guidance: Training-Free Guidance for Frame-Level Control in Video Diffusion Models11
This work proposes a simple latent processing method that dramatically reduces memory usage, and applies a novel latent optimization strategy designed for globally coherent video generation, called Frame Guidance, a training-free guidance for controllable video generation based on frame-level signals.
- Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization10
Video-MSG, a training-free Guidance method for T2V generation based on Multimodal planning and Structured noise initialization, demonstrates its effectiveness in enhancing text alignment with multiple T2V backbones on popular T2V generation benchmarks (T2VCompBench and VBench).
- ORACLE: Order Robust Adaptive Continual LEarning.9
The ORACLE method is validated on multiple benchmark datasets against state-of-the-art continual learning methods, and the results show that it largely outperforms those strong baselines with significantly less increase in capacity and training time, as well as obtains smaller performance disparity for each task with different order sequences.
- Adaptive Network Sparsification via Dependent Variational Beta-Bernoulli Dropout7
Adaptive variational dropout whose probabilities are drawn from sparsity-inducing beta Bernoulli prior allows the resulting network to tolerate larger degree of sparsity without losing its expressive power by removing redundancies among features.
- 6
- 6
- Planning with Sketch-Guided Verification for Physics-Aware Video Generation5
SketchVerify is proposed, a training-free, sketch-verification-based planning framework that improves motion planning quality with more dynamically coherent trajectories prior to full video generation by introducing a test-time sampling and verification loop.
- 5
- RSQ: Learning from Important Tokens Leads to Better Quantized LLMs3
RSQ (Rotate, Scale, then Quantize), which applies rotations (orthogonal transformation) to the model to mitigate outliers (those with exceptionally large magnitude), and scales the token feature based on its importance, and quantizes the model using the GPTQ framework with the second-order statistics computed by scaled tokens.
- Adapt-$\infty$: Scalable Continual Multimodal Instruction Tuning via Dynamic Data Selection3
The effectiveness and efficiency of Adapt-∞ are validated over a sequence of various multimodal instruction tuning datasets with various tasks, including (Knowledge) VQA, multilingual, grounding, reasoning, language-only, and multi-image comprehension tasks.
- Enhanced Thermal-Only Object Detection via LoRA-Guided Thermal-to-Visible Translation and Cross-Modal Distillation2
A novel thermal-only object detection framework that bridges the modality gap via LoRA-guided Thermal-to-Visible (T2V) translation and cross-modal knowledge distillation is proposed, significantly outperforming existing thermal-only methods while requiring only a single thermal sensor at inference.
- STELLA: Continual Audio-Video Pre-training with Spatio-Temporal Localized Alignment2
A new continual audio-video pre-training method with two novel ideas: localized Patch Importance Scoring and Replay-guided Correlation Assessment, which proposes to assess the correlation of the current patches on the past steps to identify the patches exhibiting high correlations with the past steps.
- MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs1
Musebench is introduced, a comprehensive benchmark designed to evaluate MLLMs on nuanced artistic understanding that reveals that even the best-performing model achieves only 48.29% accuracy, substantially below human expert performance, exposing a significant gap in current models' creative domain expertise.
- PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation1
This work proposes PhyMotion, a structured, fine-grained motion reward that grounds recovered 3D human trajectories in a physics simulator and evaluates motion quality along multiple dimensions of physical feasibility, and shows that PhyMotion achieves stronger correlation with human judgments than existing reward formulations.
- 1
- 1
- 1
- 1
- Hierarchy-Aware Multimodal Unlearning for Medical AI1
This work introduces MedForget, a Hierarchy-Aware Multi-modal Unlearning Testbed with explicit retain and forget splits and evaluation sets containing rephrased variants, and introduces a reconstruction attack that progressively adds hierarchical level context to prompts.
- 1
- 1
- 1
- –
- –
- VGGT-Edit: Feed-forward Native 3D Scene Editing with Residual Field Prediction–
Experiments show that VGGT-Edit substantially outperforms 2D-lifting baselines, producing sharper object details, stronger multi-view consistency, and near-instant inference speed.
- –
- –
- –
- –
- –
- –
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-10. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.