Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks91 from public data
- A phrase-based statistical model for SMS text normalization286
This paper views the task of SMS normalization as a translation problem from the SMS language to the English language and proposes to adapt a phrase-based statistical MT model for the task, which can largely boost SMS translation performance.
- A Tree Sequence Alignment-based Tree-to-Tree Translation Model125
A translation model that is based on tree sequence alignment, where a tree sequence refers to a single sequence of subtrees that covers a phrase, that statistically significantly outperforms the baseline systems and supports multi-level structure reordering of tree typology with larger span.
- SeaEval for Multilingual Foundation Models: From Cross-Lingual Alignment to Cultural Reasoning117
This work characterizes how these models understand and reason with natural language, and investigates how well they comprehend cultural practices, nuances, and values, leading to empirical results across classic NLP tasks, reasoning, and cultural comprehension.
- Topic-Aware Pointer-Generator Networks for Summarizing Spoken Conversations112
This work proposes a topic-aware architecture to exploit the inherent hierarchical structure in conversations to further adapt the pointer-generator model, which significantly outperforms competitive baselines, achieves more efficient learning outcomes, and attains more robust performance.
- Exploring syntactic structured features over parse trees for relation extraction using kernel methods78
A composite kernel for relation extraction by combining the convolution tree kernel with a simple linear kernel is proposed that can effectively capture both flat and structured features without extensive feature engineering, and easily scale to include more features.
- Data Diversification: A Simple Strategy For Neural Machine Translation76
Data Diversification: a simple but effective strategy to boost neural machine translation (NMT) performance by using the predictions of multiple forward and backward models and then merging them with the original dataset on which the final NMT model is trained is introduced.
- Term Extraction Through Unithood and Termhood Unification53
The state-of-the-art C/NCValue term weighting method is refined by considering both termhood and unithood measures, and use the former extracted terms to direct the local term extraction for each document.
- 44
- Feature-based method for document alignment in comparable news corpora40
Experimental results on the English-Chinese and English-Malay comparable news corpora show that the proposed Discrete Fourier Transform-based term frequency distribution feature is very effective and contributes 4.1% and 8% to performance improvement over Pearson's correlation method on the two comparable corpora.
- 37
- Forest-based tree sequence to string translation model34
A forest-based tree sequence to string translation model for syntax-based statistical machine translation, which automatically learns tree sequenceto string translation rules from word-aligned source-side-parsed bilingual texts, which statistically significantly outperforms the four baseline systems.
- Named-Entity Tagging and Domain adaptation for Better Customized Translation33
It is demonstrated that adding named-entity (NE) feature withnamed-entity recognition (NER) into the source language produces better translation with NMT, and that adding NE tags using NER and applying in-domain adaptation can be combined to further improve customized machine translation.
- A Grammar-driven Convolution Tree Kernel for Semantic Role Classification32
A grammardriven convolution tree kernel for semantic role classification by introducing more linguistic knowledge into the standard tree kernel is proposed and a composite kernel to integrate feature-based and tree kernel-based methods is presented.
- Automatic True/False Question Generation for Educational Purpose30
An unsupervised True/False Question Generation approach (TF-QG) that automatically generates true/false questions from a given passage for reading comprehension test and challenges the state-of-the-art inference models from NLI, QA, and fact verification tasks.
- A syntax-driven bracketing model for phrase-based translation29
A new model is presented that automatically learns syntactic constraints, including but not limited to constituent matching/violation, from training corpus and achieves a substantial improvement over the baseline which is not syntactically informed.
- Exploiting N-best hypotheses for SMT self-enhancement24
The SMT system is self-enhanced with the posterior knowledge learned from N-best hypotheses in a re-decoding framework and the combination of the three strategies achieves further improvements and outperforms the baseline by 0.67 BLEU score on NIST-2003 set, and 0.64 onNIST-2005 set.
- EM-based Hybrid Model for Bilingual Terminology Extraction from Comparable Corpora22
An unsupervised hybrid model which combines statistical, lexical, linguistic, contextual, and temporal features in a generic EM-based framework to harvest bilingual terminology from comparable corpora through comparable document alignment constraint is presented.
- Grammar comparison study for translational equivalence modeling and statistical machine translation22
This paper presents a general platform, namely synchronous tree sequence substitution grammar (STSSG), for the grammar comparison study in Translational Equivalence Modeling (TEM) and Statistical Machine Translation (SMT).
- 21
- Twitter corpus creation: The case of a Malay Chat-style-text Corpus (MCC)21
The provided criteria are used to demonstrate the process of constructing a Twitter corpus known as the Malay Chat-style Corpus (MCC), which has 1 million twitter messages and reveals characteristics of the corpus including the most frequent terms and collocations, Zipf law diagram, Twitter peak hours, and percentages of message types.
- Sentiment Aware Neural Machine Translation19
Empirical evaluations show that the valence-sensitive embedding (VSE) method significantly outperforms a sequence-to-sequence (seq2seq) baseline, both in terms of BLEU score and ambiguous word translation accuracy in test, given non-sentiment bearing contexts.
- 18
- MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models17
The results demonstrate improvements in both speech recognition and task-specific understanding, positioning MERaLiON-AudioLLM as a pioneering solution for region specific AI applications and to set a precedent for future models designed to address localised linguistic and cultural contexts in a global framework.
- Using a Hybrid Convolution Tree Kernel for Semantic Role Labeling17
A composite kernel is presented through combining the hybrid convolution tree kernel method with a feature-based method extended by the polynomial kernel, which achieves better performance than each of the individual methods and outperforms the best reported system on the CoNLL-2005 corpus.
- The TALP & I2R SMT Systems for IWSLT 200817
This paper gives a description of the statistical machine translation (SMT) systems developed at the TALP Research Center of the UPC (Universitat Politde Catalunya) for their participation in the IWSLT'08 evaluation campaign.
- Data Diversification: An Elegant Strategy For Neural Machine Translation.16
Data Diversification: a simple strategy to boost neural machine translation (NMT) performance by using the predictions of multiple forward and backward models and then merging them with the original dataset on which the final NMT model is trained is introduced.
- CoHS-CQG: Context and History Selection for Conversational Question Generation15
CoHS-CQG is proposed, a two-stage CQG framework, which adopts a novel CoHS module to shorten the context and history of the input, which achieves state-of-the-art performances on CoQA in both the answer-aware and answer-unaware settings.
- GCDST: A Graph-based and Copy-augmented Multi-domain Dialogue State Tracking15
This work first construct a dialogue state graph to transfer structured features among related domain-slot pairs across domains and encode the graph information of dialogue states by graph convolutional networks and utilize a hard copy mechanism to directly copy historical states from the previous conversation.
- AdaMCoT: Rethinking Cross-Lingual Factual Reasoning Through Adaptive Multilingual Chain-of-Thought14
This paper introduces AdaMCoT (Adaptive Multilingual Chain-of-Thought), a framework that enhances multilingual factual reasoning by dynamically routing thought processes in intermediary “thinking languages” before generating target-language responses and suggests that adaptive reasoning paths can effectively bridge the performance gap between high- and low-resource languages while maintaining cultural and linguistic nuances.
- MERaLiON-AudioLLM: Advancing Speech and Language Understanding for Singapore14
The first general-purpose multitask audio-based large language model designed to understand Singlish, a colloquial and code-switched variety of English spoken in Singapore is introduced, underscoring its utility for region-specific AI applications.
- Linguistically annotated BTG for statistical machine translation14
This paper proposes a Linguistically Annotated BTG (LABTG) that conveys linguistic knowledge of source-side syntax structures to BTG hierarchical structures through linguistic annotation and presents an annotation algorithm that captures syntactic information for BTG nodes.
- Toward Tweets Normalization Using Maximum Entropy13
A new approach using the maximum entropy model for normalizing Tweets, which addresses words that are unseen in the training phase and significantly outperforms previous well-known normalization approaches.
- A comparative study of hypothesis alignment and its improvement for machine translation system combination13
A method to build the confusion network from intersection word alignment, which utilizes both direct and inverse word alignment between the backbone and hypothesis to improve the reliability of hypothesis alignment is proposed.
- Refinements in BTG-based Statistical Machine Translation13
Two refinements for BTG-based SMT to achieve better reordering and higher-speed decoding are introduced, which include (1) reordering heuristics to prevent incorrect swapping and reduce search space, and (2) special phrases with tags to indicate sentence beginning and ending.
- A linguistically annotated reordering model for BTG-based statistical machine translation13
A linguistically annotated reordering model for BTG-based statistical machine translation that incorporates linguistic knowledge to predict orders for both syntactic and non-syntactic phrases is proposed.
- A Word Labeling Approach to Thai Sentence Boundary Detection and POS Tagging12
A word labeling approach which treats space as a normal word, and detects SB between any two words is proposed, which removes the restriction for SB to be oc-curred only at space and makes the system more robust for modern Thai writing.
- Addressing the Vulnerability of NMT in Input Perturbations11
This paper improves the robustness of NMT models by reducing the effect of noisy words through a Context-Enhanced Reconstruction (CER) approach and demonstrates robustness improvement on both news and social media text.
- 11
- Personalized Normalization for a Multilingual Chat System10
A personalized chat normalizer for English is developed and integrated with a multilingual chat system, allowing user to create and use personalized short-forms in mult bilingual chat.
- Winnowing Knowledge for Multi-choice Question Answering9
A novel encoding method is proposed which is able to conduct interception and soft filtering which contributes to the harvesting and absorption of representative information with less interference from noises.
- Revisit Automatic Error Detection for Wrong and Missing Translation – A Supervised Approach9
A supervised alignment model for translation error detection (AlignDet) is built based on a simple Alignment Triangle strategy to set the benchmark for automatic error detection task and discusses the difficulties and benefits of this task for existing evaluation metrics.
- Negative Focus Detection via Contextual Attention Mechanism9
An attention-based neural network to model contextual information is introduced which consists of a Bidirectional Long Short-Term Memory (BiLSTM) neural network and a Conditional Random Fields (CRF) layer to effectively encode the order information and the long-range context dependency in a sentence.
- Regenerating hypotheses for statistical machine translation9
Three techniques that improve the quality of N-best hypotheses through additional regeneration process: redecoding, n-gram expansion, and confusion network-based regeneration are studied.
- I 2 R Chinese-English Translation System for IWSLT 20079
This paper reports the effort on new translation entry generation and system combination and the system was ranked first with respect to the BLEU measure in Chinese-to-English open data track.
- MoWE-Audio: Multitask AudioLLMs with Mixture of Weak Encoders7
This work proposes to incorporate mixtures of ‘weak’ encoders (MoWE) into the AudioLLM framework, and demonstrates that MoWE effectively improves multi-task performance, broadening the applicability of AudioLLMs to more diverse audio tasks.
- Capturing Conversational Interaction for Question Answering via Global History Reasoning7
This paper further strengthens the ConvQA encoder by establishing long-distance dependency among global utterances in multi-turn conversation by using multi-layer transformers to resolve long- distance relationships, which potentially contribute to the reweighting of attentive information in historical utterances.
- Incorporating Contextual Paralinguistic Understanding in Large Speech-Language Models6
This work proposes two approaches to incorporate contextual paralinguistic information into model training: an explicit method that provides paralinguistic metadata directly to the LLM, and an implicit method that automatically generates novel training question-answer pairs using both categorical and dimensional emotion annotations alongside speech transcriptions.
- MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond6
The MERaLiON-SpeechEncoder was pre-trained from scratch on 200,000 hours of unlabelled speech data using a self-supervised learning approach based on masked language modelling, demonstrating improvements to spontaneous and Singapore speech benchmarks for speech recognition, while remaining competitive to other state-of-the-art speech encoders across ten other speech tasks.
- TaLAPi — A Thai Linguistically Annotated Corpus for Language Processing6
The POS tags used in ORCHID are adapted and proposed and a framework to tag Thai text is proposed and also addresses the tagging of loan and foreign words based on the proposed segmentation strategy.
- An Unsupervised and Data-Driven Approach for Spell Checking in Vietnamese OCR-scanned Texts6
This paper proposes a fully automatic approach combining both error detection and correction phases within a unique scheme designed in an unsupervised & data-driven manner, suitable for resource-poor languages.
- 5
- Coherent and Concise Radiology Report Generation via Context Specific Image Representations and Orthogonal Sentence States5
This work proposes a method to compute image representations specific to each sentential context and eliminate redundant content by exploiting diverse sentence states in neural models to ensure topical continuity, informativeness and content diversity of generated radiology reports.
- 5
- MERaLiON-TextLLM: Cross-Lingual Understanding of Large Language Models in Chinese, Indonesian, Malay, and Singlish4
This report presents MERaLiON-TextLLM, a series of open-source language models specifically tailored to improve understanding and generation in Chinese, Indonesian, Malay, and Singlish, exceeding the capabilities of the official Llama-3 models.
- 4
- Decomposed Prompting for Machine Translation Between Related Languages using Large Language Models4
DecoMT is introduced, a novel approach of few-shot prompting that decomposes the translation process into a sequence of word chunk translations that outperforms the strong few- shot prompting BLOOM model with an average improvement of 8 chrF++ scores across the examined languages.
- Refining Low-Resource Unsupervised Translation by Language Disentanglement of Multilingual Model4
This work proposes a simple refinement procedure to separate languages from a pre-trained multilingual UMT model for it to focus on only the target low-resource task.
- 4
- Linguistically Annotated Reordering: Evaluation and Analysis4
A comparative analysis is conducted that not only provides the insight into how linguistic knowledge affects phrase movement but also reveals new challenges in phrase reordering.
- 3
- Making Pre-trained Language Models Better Learn Few-Shot Spoken Language Understanding in More Practical Scenarios3
A prompt-based intent detection model in few-shot settings is developed, which leverages the BERT original pre-training next sentence prediction task and the prompt template to detect the user’s intent.
- DSPM-NLG: A Dual Supervised Pre-trained Model for Few-shot Natural Language Generation in Task-oriented Dialogue System3
A novel Dual Supervised Pre-trained Model for a few-shot Natural Language Generation (DSPM-NLG) is proposed to regularize the pre-training process and adopt a joint model with a dual supervised framework to learn the dual correlation between NLG and SLU from a probabilistic perspective.
- MARS: multilingual access and retrieval system with enhanced query translation and document retrieval3
A multilingual access and retrieval system with enhanced query translation and multilingual document retrieval, by mining bilingual terminologies and aligned document directly from the set of comparable corpora which are to be searched upon by users is introduced.
- Name Origin Recognition Using Maximum Entropy Model and Diverse Features3
This paper casts the name origin recognition as a multi-class classification problem and approach the problem using Maximum Entropy method, and investigates the use of different features, including phonetic rules, ngram statistics and character position information forName origin recognition.
- Evidence-based Interpretable Open-domain Fact-checking with Large Language Models2
Experimental results show that the Open-domain Explainable Fact-checking (OE-Fact) system outperforms general fact-checking baseline systems in both closed- and open-domain scenarios, ensuring stable and accurate verdicts while providing concise and convincing real-time explanations for fact- checking decisions.
- 2
- Two-Stage Hypotheses Generation for Spoken Language Translation2
This article proposes a strategy based on multiple hypotheses generation in a two-stage framework for spoken language translation and applies state-of-the-art, phrase-based, and syntax-based methods to generate basic translation hypotheses.
- Bridging Linguistic Gaps: Cross-Lingual Mapping in Pre-Training and Dataset for Enhanced Multilingual LLM Performance1
This work introduces a Cross-Lingual Mapping Task in the pre-training phase, which enhances cross-lingual alignment without compromising monolingual fluency, and introduces a Language Alignment Coefficient to robustly quantify cross-lingual consistency, even in limited-data scenarios.
- 1
- 1
- 1
- 1
- 1
- 1
- 1
- 1
- Uncertainty Modeling for Machine Comprehension Systems using Efficient Bayesian Neural Networks1
This work proposes a hybrid neural architecture to quantify model uncertainty using Bayesian weight approximation but boosts up the inference speed by 80% relative at test time, and applies it for a clinical dialogue comprehension task.
- 1
- 1
- 1
- Fast Computing Grammar-driven Convolution Tree Kernel for Semantic Role Labeling.1
Two fast grammar-driven convolution tree kernel algorithms are designed, which can compute the GTK in polynomial time, and Experimental results on the CoNLL-2005 SRL data show that the two FGTK algorithms are much faster than theGTK.
- –
- –
- –
- –
- –
- Improving Multilingual Social Media Insights: Aspect-based Comment Analysis–
This paper proposes a granular level of identifying and generating aspect terms from individual comments to guide model attention, and leverage multilingual large language models with supervised fine-tuning for comment aspect term generation (CAT-G), further aligning the model's predictions with human expectations through DPO.
- –
- Multi-Agent Cross-Translated Diversification for Unsupervised Machine Translation.–
Multi-Agent Cross-translated Diversification (MACD) is introduced, a method that trains multiple UMT agents and then translates monolingual data back and forth using non-duplicative agents to acquire synthetic parallel data for supervised MT.
- –
- A Rule-Augmented Statistical Phrase-based Translation System–
This work introduces an adaptive MT framework with a Rule Definition Language (RDL) for users to amend MT results through translation rules or patterns and acknowledges user feedback via RDL which improves the translations of the baseline system on three test sets for Vietnamese to English translation.
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-10. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.