研究論文

ACL
2019
Semantic Cross-lingual Sentence Embedding 
Wataru Hirota, Yoshihiko Suhara, Behzad Golshan, Wang-Chiew Tan
WWW
2019
Sara Evensen, Aaron Feng, Alon Halevy, Jinfeng Li, Vivian Li, Yuliang Li, Huining Liu, George Mihaila, John Morales, Natalie Nuno, Ekaterina Pavlovic, Wang-Chiew Tan, Xiaolan Wang
We describe Voyageur, which is an application of experiential search to the domain of travel. Unlike traditional search engines for online services, experiential search focuses on the experiential aspects of the service under consideration. In particular, Voyageur needs to handle queries for subjective aspects of the service (e.g., quiet hotel, friendly staff) and combine these with objective attributes, such as price and location. Voyageur also highlights interesting facts and tips about the services the user is considering to provide them with further insights into their choices.
言語処理学会(NLP)
2019
松田 寛, 大村 舞(国立国語研究所), 浅原 正幸(国立国語研究所)
近年、オープンソース・ソフトウェア(以下 OSS)として Stanford Core-NLP1や spaCy2のような高機能な NLP フレームワークが利用可能となっている。これらは商用利用も可能3なライセンス形態で供与されている。特に商用アプリケーションでは i18n 対応コストが重視されることが多く、NLP フレームワークには(プログラムを書き換えることなく)リソース切り替えのみで様々な言語に対応可能であることが要請される。Stanford Core-NLP や spaCy では英語以外の多くの言語リソースが提供されているが、日本語には未対応の状況が⾧く続いており、日本国内での NLP フレームワーク普及促進を妨げる要因となるばかりでなく、データサイエンス領域における日本語のプレゼンス低下に繋がることが懸念される。本稿では Universal Dependencies (Zeman[1])(UD)に基づいて設計された spaCy を NLP フレームワークとして採用し、その日本語版リソースの実現に不可欠な学習系・解析系の機能実装と精度評価を行う。UD に基づく正解コーパスには現代日本語書き言葉均衡コーパス BCCWJ (Maekawa[2])を UD 化した UDJapanese BCCWJ (Omura[3])を用いる。日本語の平文を UD に基づいてトークン化するには形態素解析器が必要となる。spaCy は Python ライブラリとして提供されるため、本稿では形態素解析器Sudachi (Takaoka[4])の Python クローンである SudachiPy4 を使用することで言語リソースの PurePython 化を実現する。Sudachi の辞書は UniDic 短単位品詞体系 (伝[5])をベースとするため、UniDic 体系に基づいて設計された UD-Japanese BCCWJ との親和性は高い。ただし、UD-Japanese BCCWJ の構築には後述のように UniDic ⾧単位品詞の参照が必要となるため、UniDic 短単位品詞体系に含まれる可能性に基づく品詞の解決(短単位品詞の用法曖昧性解決)が必要となる。本稿では依存関係ラベルに正解品詞を埋め込むことで、短単位品詞の用法曖昧性解決と依存構造解析を同時学習する方式を提案・評価する。
言語処理学会(NLP)
2019
林部 祐太
オンラインでの宿予約では,フォームに条件を入力して検索し,検索結果から選んで決めるという流れが一般的である.フォームには,日付・エリア・人数・予算などの基本的な条件のほか,食事の有無,部屋のタイプ,喫煙の可否,大浴場の有無などの「こだわり条件」が設定できることがある.しかし,それらの「こだわり条件」は操作性やスペースの制約のため,多くの人が気にすると思われる一般的な条件の中からしか選べるようになっていない.そのため,「自動販売機のビールが安い」「キッズスペースがある」といった,よりきめ細やかな条件では宿を探せない.そこで我々は,ユーザの旅行の状況を聞き出し,その状況に合わせた宿を提案する対話システムの構築を目指している.本研究では,的確な推薦のために,「肯定的事実」と「推薦対象」という言語知識を宿のレビューテキストから抽出する手法を提案する.例えば,「自販機のビールがかなり安いので,酒飲みには嬉しい」や「子どもたちには、キッズスペースや図書館など楽しかったようです」というレビュー文から「自販機のビールがかなり安い」ことは「酒飲み」に肯定的であることや,「キッズスペースや図書館」が「子どもたち」に肯定的である,という知識を抽出する.提案手法を用いて,旅行情報サイト「じゃらん net」1に投稿された宿のレビューから肯定的事実と推薦対象を 7,701 組抽出した.そして,そのうち 2,439 組に対して肯定的事実のアスペクトのアノテーションと,肯定的事実を推薦対象にアピールポイントとして用いることが妥当であるかの評価を,クラウドソーシングを 用いて実施した.
言語処理学会(NLP)
2019
川島 寛乃, 松田 寛, 毛利 研
口コミや記事などの文書を分析する際に単語の極性情報は重要であり,日本語では小林ら [1] の日本語評価極性辞書や高村ら [2] の単語感情極性対応表,また梶ら [3] の評価表現辞書などが利用できる.しかしこれらの評価極性辞書に含まれる単語は,特定のドメインにおけるテキストに対しては適合しないことがある.一例として,本研究でデータセットとして用いる旅行情報サイトの口コミデータでは,評価極性辞書に含まれる 13,625 個の単語のうち約 6 割強の単語はデータの語彙に含まれず,ドメインに適応した単語が辞書に含まれているとは言い難い.また「静か」のように極性辞書では negative な極性単語として登録されているが,実際の口コミデータにおいては肯定的な意味合いで使われる単語も存在する.このようなドメインに特化した辞書を用いたい場合,ドメインに応じて低コストで辞書の拡張を行う必要がある.そこで,本研究では極性既知の単語のベクトルに対して類似度の高いベクトルを,評価極性辞書への新たな追加候補語として取得する方法を提案する.まず極性が既知で対義関係にある単語対の差分を差分ベクトルとして算出し,元の極性既知の単語ベクトルに差分ベクトルを加えたベクトルに対する高類似度単語を取得することで,反対極性の単語を除外した追加候補語の獲得を試みる.
EMNLP
2018
Dan Iter, Alon Y. Halevy, Wang-Chiew Tan
A common need of NLP applications is to extract structured data from text corpora in order to perform analytics or trigger an appropriate action. The ontology defining the structure is typically application dependent and in many cases it is not known a priori. We describe the FrameIt System that provides a workflow for (1) quickly discovering an ontology to model a text corpus and (2) learning an SRL model that extracts the instances of the ontology from sentences in the corpus. FrameIt exploits data that is obtained in the ontology discovery phase as weak supervision data to bootstrap the SRL model and then enables the user to refine the model with active learning. We present empirical results and qualitative analysis of the performance of FrameIt on three corpora of noisy user-generated text.
PVLDB
2018
Xiaolan Wang, Aaron Feng, Behzad Golshan, Alon Y. Halevy, George A. Mihaila, Hidekazu Oiwa, Wang-Chiew Tan
We present the KOKO system that takes declarative information extraction to a new level by incorporating advances in natural language processing techniques in its extraction language. KOKO is novel in that its extraction language simultaneously supports conditions on the surface of the text and on the structure of the dependency parse tree of sentences, thereby allowing for more refined extractions. KOKO also supports conditions that are forgiving to linguistic variation of expressing concepts and allows to aggregate evidence from the entire document in order to filter extractions. To scale up, KOKO exploits a multi-indexing scheme and heuristics for efficient extractions. We extensively evaluate KOKO over publicly available text corpora. We show that KOKO indices take up the smallest amount of space, are notably faster and more effective than a number of prior indexing schemes. Finally, we demonstrate KOKO’s scalability on a corpus of 5 million Wikipedia articles.
PVLDB
2018
Xiaolan Wang, Jiyu Komiya, Yoshihiko Suhara, Aaron Feng, Behzad Golshan, Alon Y. Halevy, Wang-Chiew Tan
KOKO is a declarative information extraction system that incorporates advances in natural language processing techniques in its extraction language. KOKO’s extraction language supports simultaneous specification of conditions over the surface syntax and on the structure of the dependency parse tree of sentences, thereby allowing for more refined extractions. Furthermore, the KOKO extraction language allows for aggregating evidence from an input document and supports conditions that are tolerant of linguistic variation of expressing concepts. In this demo, we outline the design of KOKO, a system for extracting information and understanding the results of the extraction. KOKO provides an interactive interface that allows participants to write queries, understand the input and results of the queries. In particular, the user can customize the input text, visualize the input text’s dependency parse trees, and understand the correspondences between query components, dependency tree nodes, text tokens, and the computation and associated scores that led to an extraction.
IEEE
2018
Chen Chen, Behzad Golshan, Alon Y. Halevy, Wang-Chiew Tan, AnHai Doan
We present BIGGORILLA, an open-source resource for data scientists who need data preparation and integration tools, and the vision underlying the project. We then describe four packages that we contributed to BIGGORILLA: KOKO (an information extraction tool), FLEXMATCHER (a schema matching tool), MAGELLAN and DEEPMATCHER (two entity matching tools). We hope that as more software packages are added to BIGGORILLA, it will become a one-stop resource for both researchers and industry practitioners, and will enable our community to advance the state of the art at a faster pace.
LREC
2018
Akari Asai, Sara Evensen, Behzad Golshan, Alon Y. Halevy, Vivian Li, Andrei Lopatenko, Daniela Stepanov, Yoshihiko Suhara, Wang-Chiew Tan, Yinzhan Xu
The science of happiness is an area of positive psychology concerned with understanding what behaviors make people happy in a sustainable fashion. Recently, there has been interest in developing technologies that help incorporate the findings of the science of happiness into users’ daily lives by steering them towards behaviors that increase happiness. With the goal of building technology that can understand how people express their happy moments in text, we crowd-sourced HappyDB, a corpus of 100,000 happy moments that we make publicly available. This paper describes HappyDB and its properties, and outlines several important NLP problems that can be studied with the help of the corpus. We also apply several state-of-the-art analysis techniques to analyze HappyDB. Our results demonstrate the need for deeper NLP techniques to be developed which makes HappyDB an exciting resource for follow-on research. Keywords:science of happiness, positive psychology, happyDB corpus, crowdsourcing