Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks11 from public data
- Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities432
The first multi-view action dataset, with si-multaneous static and egocentric recordings, and a novel task of detecting mistakes is proposed, to investigate generalization to new toys, cross-view transfer, long-tailed distributions, and pose vs. appearance.
- 14
- 13
- 6
- Don't Pause! Every prediction matters in a streaming video4
AsynKV, a training-free streaming adaptation of offline models, that retains their event perception while improving their streaming behavior, serves as a strong baseline on SPOT-Bench, outperforming existing streaming models, and achieves state-of-the-art on retrospective benchmarks.
- 4
- 4
- Decouple and Cache: KV Cache Construction for Streaming Video Understanding2
Decoupled Streaming Cache is proposed, a training-free cache construction mechanism that adapts pretrained offline models to streaming settings that maintains a cumulative past KV cache while constructing a separate instant cache on-demand, decoupled from past caches to preserve the informativeness of recent inputs.
- On Discriminative vs. Generative classifiers: Rethinking MLLMs for Action Understanding1
GAD improves both accuracy and efficiency over generative methods, achieving state-of-the-art results on four tasks across five datasets, including an average 2.5% accuracy gain and 3x faster inference on the largest COIN benchmark.
- –
- –
Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.