Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks11 from public data
- 64
- IoT-LLM: A framework for enhancing large language model reasoning from real-world sensor data55
IoT-LLM is proposed, a unified framework that significantly improves the performance of IoT-sensory task reasoning of LLMs, with models such as GPT-4o-mini showing a 49.4% average improvement over previous methods.
- TENT: Connect Language Models With IoT Sensors for Zero-Shot Activity Recognition41
TENT not only achieves robust recognition of both seen and unseen activities but also significantly outperforms existing vision–language and sensor-language baselines, surpassing them by over 20% on zero-shot HAR tasks, establishing TENT as a new paradigm for generalizable IoT representation learning.
- Data Pyramid for Embodied Manipulation: A Survey8
This work organizes the embodied data ecosystem as a pyramidspanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-language data, and further characterize each source in terms of data quality, diversity, reusability, and physical fidelity.
- HiCoGen: Hierarchical Compositional Text-to-Image Generation in Diffusion Models via Reinforcement Learning4
This work proposes a Hierarchical Compositional Generative framework (HiCoGen) built upon a novel Chain of Synthesis (CoS) paradigm, and introduces a reinforcement learning (RL) framework that significantly outperforms existing methods in both concept coverage and compositional accuracy.
- Physics filtering favors the generalization of robot learning2
It is shown that robots can generalize effectively under dynamics uncertainties even with limited training data by leveraging a feedback mechanism, namely PhyFilter, that corrects learning outputs with physics-filtered learning residuals, and shows that physics-filtered feedback can serve as a powerful alternative to massive data scaling.
- SkeFi: Cross-Modal Knowledge Transfer for Wireless Skeleton-Based Action Recognition2
This work proposes the enhanced temporal correlation adaptive graph convolution (TC-AGC) with frame interactive enhancement to overcome the noise from missing or noncontinuous frames and underscores the effectiveness of enhancing multiscale temporal modeling (MST) through dual temporal convolution.
- 2
- 1
- –
- –
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.