LLMと自然言語処理

大規模言語モデル(LLM)の革新により、自然言語処理(NLP)はタスク固有の手法から汎用的なデータ駆動型アプローチへと移行し、研究と応用に革命をもたらしました。現代の LLM は、検索エンジン、API、シンボリック推論システムなどの外部ツールと統合され、専門知識を要する複雑なタスクに対応できるようになっています。しかし、LLM の利用が拡大するにつれ、公平性、制御性、透明性、説明可能性といった課題が浮き彫りになっています。特に、人事(HR)、法律、金融、医療といった分野では、これらの要素が極めて重要です

Megagon Labs では、LLM の可能性を最大限に活用しつつ、これらの課題を克服することを目指しています。私たちの研究は、以下の 3 つの主要分野に焦点を当てています。

  1. LLM の挙動と制約の理解: 実世界のプロダクション環境における LLM の性能と、その課題を調査。
  2. LLM の能力向上: 新たなシステム、ハイブリッドなニューロンシンボリックアプローチ、ドメイン固有の技術革新を開発し、LLM のパフォーマンスを向上。
  3. 堅牢な評価手法: 複雑な実世界のタスクにおける LLM の評価手法を確立し、多様なアプリケーションにおいて信頼性と有効性を確保。

これらの手法を活用し、HR や関連分野に適した AI ソリューションの品質、一貫性、公平性、真実性を向上させ、研究と実践の両面で有意義な進展を促進します。私たちの取り組みは、基礎研究、応用プロジェクト、オープンソース貢献を含み、研究所内外での実際的な影響を生み出すことを目指しています。

ハイライト

プロジェクト

「抽出して評価(Extract then Evaluate)」という革新的な手法を提案。これにより、LLM を用いた長文要約の評価コストを大幅に削減し、人間による評価との整合性を向上させる。

自然言語生成(NLG)タスクの指示に含まれる曖昧な仕様を特定し、より良い出力品質を実現するために指示を明確化する手法を提案。

複数選択式質問応答タスクにおける LLM の感度を調査。このタスクは、LLM の推論能力や事実検索能力を評価するためによく使用される。

LLM のパフォーマンス向上において、検索がどのように機能するのかを評価し、検索が有効な場合と逆効果になる場合を明らかにするベンチマークおよび調査を実施。本研究の知見は、信頼性の高い検索拡張型言語モデル(RAG)ベースの QA システムの開発に貢献する。

関連

研究論文

LT4HALA
2026
Hiroshi Matsuda, Masayuki Asahara
omnes flores is an NLP framework based on Universal Dependencies (UD) that utilizes multilingual Large Language Models (LLMs), and its default model is trained on data from 40 UD languages comprising 40 treebanks. For the EvaLatin 2026 Dependency Parsing Tasks, we extended the training data of omnes flores by incorporating six public Latin treebanks from UD and trained a dependency parsing model using the extended training data. The dependency parser of omnes flores normally takes a list of word FORM values as input. However, since the EvaLatin 2026 test data includes an UPOS column, we investigated whether incorporating both FORM and UPOS during both training and inference could improve parsing accuracy. Our experiments show that training using both FORM and UPOS improves performance by 0.5-1.0 LAS points on Prose compared with training using only FORM, but decreases performance by 5 points on Poetry.
UDW
2026
Hiroshi Matsuda, Masayuki Asahara
In this research, we introduce LoRA probing, a lightweight approach for observing how core syntactic abilities emergeduring LLM pretraining. Leveraging OLMo-2’s public intermediate checkpoints, we trace learning curves across 24 pretraining stages on 33 Universal Dependencies languages by fine-tuning LoRA with step-by-step parsing instructions and a simple tabular output. To fit the relatively short context length of the OLMo-2, we design a compact 2-step-no-form prompt template and this matches the baseline in average accuracy while halving the context length and substantially increasing throughput, enabling efficient large-scale evaluation. Token Recall surpasses 0.9 within the first 1–2K pretraining steps, indicating that stable output formatting emerges early. Despite OLMo-2-7B’s English-centric pretraining, LAS exceeds 80 points in 29 of 33 languages; however, relations such as iobj and csubj show delayed onset and instability across many languages. LoRA probing thus provides a practical, reproducible lens on the cross-lingual dynamics of syntactic acquisition during LLM pretraining.
人工知能学会
2026
金子 正弘, 松田 寛, 鈴木 久美, 関根 聡
大規模言語モデルは差別的な社会的バイアスを含む情報を生成するリスクがあり,その評価が必要である.しかし,何を「差別的な社会的バイアス」とみなすかは社会的文脈に依存するため,普遍的な価値基準を定義することは難しく,個々の社会的文脈において合意可能な価値基準に基づいた安全性担保を行う必要がある.社会的バイアスのベンチマーク構築に関する先行研究では,社会的文脈の一つである国による価値基準の相違に対応するため,当該国のアノテーターを用いてローカライゼーションを行っているが,この手法はアノテーターの主観に強く依存しており,安全性の判断が当該国において合意可能なものであることを明確には担保していない.本研究では,各国において合意された「差別的な社会的バイアス」の最低限の基準として,法令とその判例等を根拠とする安全性担保のローカライゼーションを提案し,日本の雇用関連領域および医療提供関連領域の法令において差別と判断された事例を収集して,社会的バイアスデータセット – JLawBias を構築し,6 つの日本語対応 LLM に対して簡易な評価を実施して手法の有効性を確認した. ここに掲載した著作物の利用に関する注意 本著作物の著作権は人工知能学会に帰属します。本著作物は著作権者である人工知能学会の許可のもとに掲載するものです。ご利用に当たっては「著作権法」に従うことをお願いいたします。 Notice for the use of this material. The copyright of this material is retained by the Japanese Society for Artificial Intelligence (JSAI). This material is published here with the agreement of JSAI. Please be complied with Copyright Law of Japan if any users wish to reproduce, make derivative work, distribute or make available to the public any part or whole thereof. All Rights Reserved, Copyright (C) The Japanese Society for Artificial Intelligence.
Moin Amin-Naseri, Hannah Kim, Estevam Hruschka
The extraction of structured information from raw text is a fundamental component of many NLP applications, including document retrieval, ranking, and relevance estimation. High-quality extractions often require domain-specific accuracy, up-to-date understanding of specialized taxonomies, and the ability to incorporate emerging jargon and rare outliers. In many domains–such as medical, legal, and HR–the extraction model must also adapt to shifting terminology and benefit from explicit reasoning over structured knowledge. We propose DySECT, a Dynamic Self-Evolving Extraction and Curation Toolkit, which continually improves as it is used. The system incrementally populates a versatile, self-expanding knowledge base (KB) with triples extracted by the LLM. The KB further enriches itself through the integration of probabilistic knowledge and graph-based reasoning, gradually accumulating domain concepts and relationships. The enriched KB then feeds back into the LLM extractor via prompt tuning, sampling of relevant few-shot examples, or fine-tuning using KB-derived synthetic data. As a result, the system forms a symbiotic closed-loop cycle in which extraction continuously improves knowledge, and knowledge continuously improves extraction.
ACL
2026
Tool-augmented Language Models (TaLMs) can invoke external tools to solve problems beyond their parametric capacity. However, it remains unclear whether these tool-enabled gains reflect trustworthy reasoning. Focusing on the Code Interpreter tool, we show that even when tools are selected and executed correctly, TaLMs treat tool outputs as substitutes for reasoning, producing solutions that appear correct but lack coherent justification. We term this failure mode Tool-Induced Myopia (TIM), and study it using PYMATH, a benchmark of 1,679 competition-level mathematical problems for which Python code is helpful but not sufficient. We further develop a multi-dimensional evaluation suite to quantify reasoning degradation in TaLMs relative to their non-tool counterparts. Our findings reveal that while TaLMs achieve up to a 19.3 percentage point gain in final-answer accuracy, their reasoning behavior consistently deteriorates (e.g., non-tool LLMs win up to 41.5% more often in pairwise comparisons of reasoning process). This degradation intensifies with tool use; the more frequently a model invokes tools, the less coherent its reasoning becomes. Moreover, tool use shifts errors from arithmetic mistakes toward global reasoning failures (logic, assumption, creativity); with TIM present in ~55% of high-risk cases. Finally, we propose a preference-optimizationbased framework that realigns TaLMs to use tools as assistive evidence, improving both final-answer accuracy and reasoning depth under tool use. Codes and data are available at: https://github.com/megagonlabs/TIM.
言語処理学会 (NLP)
2026
大規模言語モデルにおけるプロンプト変動が出力に対する影響について、様々な文脈で研究されており、用語は多数存在する。本研究は、既存研究で混在してきた概念を「頑健性」と「可制御性」の二軸から再構造化する。さらに、公開データセットを前提とした従来の分析とは異なり、複雑なタスク構成や追加知識の記述を要するビジネスサービス環境に着目し、両概念の重要度を体系的に評価した。実験の結果、我々が考察したタスクにおいては、先行研究で強調されてきた頑健性よりも、プロンプト意図を確実に反映し必要情報を安定して引き出す可制御性が実運用において本質的であることが明らかとなった。本研究は、ビジネス環境に適したプロンプト設計指針の再考に寄与するとともに、将来の評価指標構築やモデル改善への示唆を提供する。
13 Min Read
November 7, 2025
「混合シグナル(Mixed Signals)」は、視覚言語モデル(VLM)の隠れたバイアスを明らかにし、ヘルスケア、RAG システム、AI の安全性に対して重大な示唆を与えています。
7 Min Read
May 5, 2025
私たちは、自然言語処理における喫緊かつ未開拓のトピックである「複数文書推論」に取り組む3つの新しい論文を紹介します。これらの論文は、大規模言語モデル(LLM)が複数の情報源にまたがる複雑性をどのように扱うかについて、厳密なベンチマーク、新しい方法論、経験的洞察を提供します。
11 Min Read
February 5, 2025
MCRankベンチマークとEXSIR手法を用いることで、構造化された推論によりLLMの性能がこれらの難解なタスクで大幅に向上することを示しました。