Jinfeng Li, Yuliang Li, Xiaolan Wang, Wang-Chiew Tan
Semantic tagging, which has extensive applications in text mining, predicts whether a given piece of text conveys the meaning of a given semantic tag. The problem of semantic tagging is largely solved with supervised learning and today, deep learning models are widely perceived to be better for semantic tagging. However, there is no comprehensive study supporting the popular belief. Practitioners often have to train different types of models for each semantic tagging task to identify the best model. This process is both expensive and inefficient.
We embark on a systematic study to investigate the following question: Are deep models the best performing model for all semantic tagging tasks? To answer this question, we compare deep models against “simple models” over datasets with varying characteristics. Specifically, we select three prevalent deep models (i.e. CNN, LSTM, and BERT) and two simple models (i.e. LR and SVM), and compare their performance on the semantic tagging task over 21 datasets. Results show that the size, the label ratio, and the label cleanliness of a dataset significantly impact the quality of semantic tagging. Simple models achieve similar tagging quality to deep models on large datasets, but the runtime of simple models is much shorter. Moreover, simple models can achieve better tagging quality than deep models when targeting datasets show worse label cleanliness and/or more severe imbalance. Based on these findings, our study can systematically guide practitioners in selecting the right learning model for their semantic tagging task.
Yoshihiko Suhara, Xiaolan Wang, Stefanos Angelidis, Wang-Chiew Tan
We present OPINIONDIGEST, an abstractive opinion summarization framework, which
does not rely on gold-standard summaries for
training. The framework uses an Aspect-based
Sentiment Analysis model to extract opinion
phrases from reviews, and trains a Transformer
model to reconstruct the original reviews from
these extractions. At summarization time, we
merge extractions from multiple reviews and
select the most popular ones. The selected
opinions are used as input to the trained Transformer model, which verbalizes them into an
opinion summary. OPINIONDIGEST can also
generate customized summaries, tailored to
specific user needs, by filtering the selected
opinions according to their aspect and/or sentiment. Automatic evaluation on YELP data
shows that our framework outperforms competitive baselines. Human studies on two corpora verify that OPINIONDIGEST produces
informative summaries and shows promising
customization capabilities1
.
Dialog systems capable of filling slots with
numerical values have wide applicability to
many task-oriented applications. In this paper, we perform a particular case study on
the number of guests slot-filling in hotel
reservation domain, and propose two methods to improve current dialog system model
on 1. numerical reasoning performance by
training the model to predict arithmetic expressions, and 2. multi-turn question generation
by introducing additional context slots. Furthermore, because the proposed methods are
all based on an end-to-end trainable sequenceto-sequence (seq2seq) neural model, it is possible to achieve further performance improvement on growing dialog logs in the future.
We perform the textual entailment (TE) corpus construction for the Japanese Language with the following three characteristics: First, the
corpus consists of realistic sentences; that is, all sentences are spontaneous or almost equivalent. It does not need manual writing which
causes hidden biases. Second, the corpus contains adversarial examples. We collect challenging examples that can not be solved by a
recent pre-trained language model. Third, the corpus contains explanations for a part of non-entailment labels. We perform the reasoning
annotation where annotators are asked to check which tokens in hypotheses are the reason why the relations are labeled. It makes easy
to validate the annotation and analyze system errors. The resulting corpus consists of 48,000 realistic Japanese examples. It is the largest
among publicly available Japanese TE corpora. Additionally, it is the first Japanese TE corpus that includes reasons for the annotation as
we know. We are planning to distribute this corpus to the NLP community at the time of publication.
Knowledge bases (KBs) have long been the backbone of many real-world applications and
services. There are many KB construction (KBC) methods that can extract factual information,
where relationships between entities are explicitly stated in text. However, they cannot model
implications between opinions which are abundant in user-generated text such as reviews and often
have to be mined. Our goal is to develop a technique to build KBs that can capture both opinions and
their implications. Since it can be expensive to obtain training data to learn to extract implications
for each new domain of reviews, we propose an unsupervised KBC system, SAMPO, that is based
on matrix factorization techniques. Specifically, SAMPO is tailored to build KBs for domains where
many reviews on the same domain are available. We generate KBs for 20 different domains using
SAMPO and manually evaluate KBs for 6 domains. Our experiments show that KBs generated
using SAMPO capture information otherwise missed by other KBC methods. Specifically, we show
that our KBs can provide additional training data to fine-tune language models that are used for
downstream tasks such as review comprehension.
Xiong Zhang, Jonathan Engel, Sara Evensen, Yuliang Li, Çağatay Demiralp, Wang-Chiew Tan
Reviews are integral to e-commerce services and products. They contain a wealth of information about the opinions and experiences of users, which can help better understand consumer decisions and improve user experience with products and services. Today, data scientists analyze reviews by developing rules and models to extract, aggregate, and understand information embedded in the review text. However, working with thousands of reviews, which are typically noisy incomplete text, can be daunting without proper tools. Here we first contribute results from an interview study that we conducted with fifteen data scientists who work with review text, providing insights into their practices and challenges. Results suggest data scientists need interactive systems for many review analysis tasks. Towards a solution, we then introduce Teddy, an interactive system that enables data scientists to quickly obtain insights from reviews and improve their extraction and modeling pipelines.
Xiaolan Wang, Yoshihiko Suhara, Natalie Nuno, Yuliang Li, Jinfeng Li, Nofar Carmeli, Stefanos Angelidis, Eser Kindogan, Wang-Chiew Tan
Building summarization systems have become a necessity due to the extensive volume and growth of online reviews. Despite extensive research on this topic, existing summarization systems generally fall short on two aspects. First, existing techniques generate static summaries which cannot be tailored to specific user needs. Second, most existing systems generate extractive summaries which selects only certain salient aspects from the summaries. Hence, they do not completely depict the overall opinion of the reviews. In this paper, we demonstrate a novel summarization system, ExtremeReader, that overcomes the limitations of existing summarization systems described above. ExtremeReader allows summaries to be tailored and explored interactively so that users can quickly find the desired information. In addition, ExtremeReader generates abstractive summaries with an underlying structure that helps users understand, explore, and seek explanations to the generated summaries.
Zhengjie Miao, Yuliang Li, Xiaolan Wang, Wang-Chiew Tan
Online services are interested in solutions to opinion mining, which is the problem of extracting aspects, opinions, and sentiments from text. One method to mine opinions is to leverage the recent success of pre-trained language models which can be fine-tuned to obtain high-quality extractions from reviews. However, fine-tuning language models still requires a non-trivial amount of training data.
In this paper, we study the problem of how to significantly reduce the amount of labeled training data required in fine-tuning language models for opinion mining. We describe , an opinion mining system developed over a language model that is fine-tuned through semi-supervised learning with augmented data. A novelty of is its clever use of a two-prong approach to achieve state-of-the-art (SOTA) performance with little labeled training data through: (1) data augmentation to automatically generate more labeled training data from existing ones, and (2) a semi-supervised learning technique to leverage the massive amount of unlabeled data in addition to the (limited amount of) labeled data. We show with extensive experiments that performs comparably and can even exceed previous SOTA results on several opinion mining tasks with only half the training data required. Furthermore, it achieves new SOTA results when all training data are leveraged. By comparison to a baseline pipeline, we found that extracts significantly more fine-grained opinions which enable new opportunities of downstream applications.
Wataru Hirota, Yoshihiko Suhara, Behzad Golshan, Wang-Chiew Tan
We present Emu, a system that semantically enhances multilingual sentence embeddings. Our framework fine-tunes pre-trained multilingual sentence embeddings using two main components: a semantic classifier and a language discriminator. The semantic classifier improves the semantic similarity of related sentences, whereas the language discriminator enhances the multilinguality of the embeddings via multilingual adversarial training. Our experimental results based on several language pairs show that our specialized embeddings outperform the state-of-the-art multilingual sentence embedding model on the task of cross-lingual intent classification using only monolingual labeled data.