Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks16 from public data
- Explicit Interaction Model towards Text Classification121
A novel framework, EXplicit interAction Model (dubbed as EXAM), equipped with the interaction mechanism to incorporate word-level matching signals into the text classification task is designed.
- GliDe with a CaPE: A Low-Hassle Method to Accelerate Speculative Decoding84
GliDe is a modified draft model architecture that reuses the cached keys and values from the target LLM, while CaPE is a proposal expansion method that uses the draft model's confidence scores to help select additional candidate tokens for verification.
- SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference Acceleration81
SWIFT is introduced, an on-the-fly self-speculative decoding algorithm that adaptively selects intermediate layers of LLMs to skip during inference, making it a plug-and-play solution for accelerating LLM inference across diverse input data streams.
- 46
- Efficient Inference for Large Language Model-based Generative Recommendation32
An alignment framework named AtSpeed is proposed, which presents the AtSpeed-S optimization objective for top-K alignment under the strict top-K verification, and introduces a relaxed sampling verification strategy that allows high-probability non-top-K drafted sequences to be accepted, significantly reducing LLM calls.
- ngram-OAXE: Phrase-Based Order-Agnostic Cross Entropy for Non-Autoregressive Machine Translation12
This work extends oaxe by only allowing reordering between ngram phrases and still requiring a strict match of word order within the phrases, and shows that ngram noaxe indeed improves the translation of ngram words, and produces more fluent translation with a better modeling of sentence structure.
- ToolSpec: Accelerating Tool Calling via Schema-Aware and Retrieval-Augmented Speculative Decoding4
This work proposes ToolSpec, a schema-aware, retrieval-augmented speculative decoding method for accelerating tool calling that achieves up to a 4.2x speedup, substantially outperforming existing training-free speculative decoding methods.
- Tutorial Proposal: Speculative Decoding for Efficient LLM Inference4
This tutorial delves into the latest techniques in SD, including draft model architectures and verification strategies, and explores the acceleration potential and future research directions in this promising field of Speculative Decoding.
- Demystifying the Slash Pattern in Attention: The Role of RoPE3
This paper analyzes the training dynamics of a shallow Transformer equipped with RoPE under these conditions, and proves that models trained via gradient descent exhibit SDHs, which are intrinsic to models and generalize to out-of-distribution prompts.
- Revisiting the Markov Property for Machine Translation2
It is indicated that MAT with an order larger than 4 can generate translations with quality on par with that of conventional autoregressive transformers.
- 1
- Merlin’s Whisper: Enabling Efficient Reasoning in Large Language Models via Black-box Persuasive Prompting–
This work presents a new approach to mitigating overthinking in LRMs via black-box persuasive prompting via Whisper, an iterative refinement framework that generates high-quality persuasive prompts from diverse perspectives and reveals the broad applicability of Whisper across data domains, model scales, and families.
- –
- –
- –
- –
Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.