Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
Possible advisorsa guess from early papers, not confirmed
- Anthony K. H. TungSuggested from co-authorship
Is this you? Claim this profile to confirm or dismiss it.
Works16 from public data
- InvestLM: A Large Language Model for Investment using Financial Domain Instruction Tuning129
This work suggests that a high-quality domain specific LLM can be tuned using a small set of carefully curated instructions on a well-trained foundation model, which is consistent with the Superficial Alignment Hypothesis.
- Do Multi-Hop Question Answering Systems Know How to Answer the Single-Hop Sub-Questions?48
A neural decomposition model is adopted to generate sub-questions for a multi-hop question, followed by extracting the corresponding sub-answers in order to shed some light on explaining the reasoning process of QA systems in answering complex questions.
- Detecting adverse drug reactions in discharge summaries of electronic medical records using Readpeer27
A natural language processing framework to detect drug-AE relations from unstructured hospital discharge summaries is developed and the ease of reviewing and correcting the results of the algorithm as part of an iterative machine learning process is an important step towards use of hospital discharged summaries for an active pharmacovigilance program.
- QALink: Enriching Text Documents with Relevant Q&A Site Contents19
This paper designs an end-to-end system named QALink which assigns the most relevant Q&A contents to the corresponding section of the document, and presents a new segmentation approach to model each document with a hierarchical structure.
- 16
- Combining Machine Learning with a Rule-Based Algorithm to Detect and Identify Related Entities of Documented Adverse Drug Reactions on Hospital Discharge Summaries16
A hybrid model is developed combining rule-based and machine learning algorithms using discharge summaries with the aim of maximising capture of related drug-adverse event pairs and demonstrates reasonable generalisability on external validation.
- 14
- SQuAD-SRC: A Dataset for Multi-Accent Spoken Reading Comprehension9
A large-scale multi-accent human spoken dataset SQuAD-SRC is constructed and various adaption strategies to improve the SRC performance are explored, especially for multi-ACcent spoken questions.
- GAPrune: Gradient-Alignment Pruning for Domain-Aware Embeddings2
This work proposes GAPrune, a pruning framework that addresses this challenge by considering both domain importance and preserving general linguistic foundation, and demonstrates that principled pruning strategies can achieve model compression and enhanced domain specialization.
- Dual Graph Disambiguation for Multi-Instance Partial-Label Learning1
DualG, a novel framework that simultaneously addresses feature learning and label disambiguation through dual-level graph propagation, is proposed that outperforms existing MIPL and partial label learning methods, validating its effectiveness and superiority.
- 1
- –
- BAT: Target-Instance-Free Data Preparation Synthesis via LLM-Driven Tree Search–
This work proposes BAT, the first end-to-end ADP framework that enables the synthesis of training-free data preparation pipelines without requiring any instances from target tables, and introduces EPO, which invokes pipeline execution results from sources to targets to evaluate the reliability of the generated pipelines in FPG.
- SR-GRPO: Stable Rank as an Intrinsic Geometric Reward for Large Language Model Alignment–
Stable rank is proposed, an intrinsic, annotation-free quality signal derived from model representations that measures the effective dimensionality of hidden states by computing the ratio of total variance to dominant-direction variance, capturing quality through how information distributes across representation dimensions.
- A Review of Multimodal Brain Language Decoding–
This work outlines a systematic framework and technical roadmap for advancing brain-language decoding in both fundamental research and real-world applications and reflects on ethical concerns related to neurotechnology.
- –
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.