Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks16 from public data
- Robust and Resource-Efficient Data-Free Knowledge Distillation by Generative Pseudo Replay63
A Variational Autoencoder with a training objective that is customized to learn the synthetic data representations optimally and optimizes the expected value of the distilled model accuracy while eliminating the large memory overhead incurred by the sample-storing methods.
- Identification and Detection of Phishing Emails Using Natural Language Processing Techniques48
This scheme is aimed at detecting phishing mails which do not contain any links but bank on the victim's curiosity by luring them into replying with sensitive information and is far better than the existing Phishing Email Detection techniques.
- 30
- DAOP: Data-Aware Offloading and Predictive Pre-Calculation for Efficient MoE Inference25
DAOP dynamically allocates experts between CPU and GPU based on per-sequence activation patterns, and selectively pre-calculates predicted experts on CPUs to minimize transfer latency and enable efficient resource utilization across various expert cache ratios while maintaining model accuracy through a novel graceful degradation mechanism.
- Chameleon: Dual Memory Replay for Online Continual Learning on Edge Devices20
This work proposes Chameleon, a hardware-friendly continual learning framework for user-centric training with dual replay buffers that leverages the hierarchical memory structure available on most edge devices, introducing a short-term replay store in the on-chip memory and a long-term replays in the off- chip memory to acquire new information while retaining past knowledge.
- 17
- TerEffic: Highly Efficient Ternary LLM Inference on FPGA16
TerEffic is introduced, an FPGA-based architecture tailored for ternary-quantized LLM inference that offers flexibility through reconfigurable hardware to meet various system requirements and demonstrates significant performance and energy efficiency improvements.
- Shedding the Bits: Pushing the Boundaries of Quantization with Minifloats on FPGAs10
This work presents minifloats, which are reduced-precision floating-point formats capable of further reducing the memory footprint, latency, and energy cost of a model while approaching full-precision model accuracy.
- 8
- 7
- 5
- 1
- Condensed Data Expansion Using Model Inversion for Knowledge Distillation1
This work proposes a method that expands condensed datasets using model inversion, a technique for generating synthetic data based on the impressions of a pre-trained model on its training data, which demonstrates significant gains in KD accuracy compared to using condensed datasets alone and outperforms standard model inversion-based KD methods.
- 1
- –
- Efficient Text-to-Image Generation: An Adaptive Step Schedule Controller for Diffusion Models–
This work proposes an adaptive diffusion controller that dynamically adjusts the number of steps to generate high-quality images efficiently, without additional model training, by leveraging a mixture of step schedules with varying step sizes.
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.