Research
Overview
To make use of diverse kinds of knowledge, AI needs to combine approaches suited to the nature of that knowledge: learning it during training, retrieving it from external sources when needed, and updating it as circumstances change.
We study large language models (LLMs) that flexibly handle diverse kinds of knowledge and search agents that autonomously find and integrate the information they need. Through this research, we aim to develop AI that draws on knowledge to reason, plan, and act autonomously.
Search Agents
We study search agents that autonomously alternate between search and reasoning, drawing on diverse sources such as the web, academic literature, and internal organizational documents. Our goal is to develop AI that searches for the information it needs based on the model’s knowledge and search results, integrates information while considering its reliability and recency, and provides answers grounded in evidence.
We also work on using information across languages and from sources containing figures, tables, and mathematical expressions, improving search efficiency, and developing evaluation datasets and methods to measure these capabilities. We aim to create environments where people and AI can collaborate on research and problem solving, drawing on human knowledge and judgment.
Selected Achievements
As foundational retrieval technologies, we proposed BPR, a retrieval model with reduced memory requirements, and KPR, a retrieval model that can incorporate additional entity knowledge without retraining. For the 2025 NeurIPS MMU-RAG Competition, we developed a search agent that repeatedly searches and reasons, integrating multiple sources to generate detailed answers. The agent won the static evaluation in the open-source division of the Text-to-Text track.
We also contribute to question answering evaluation and competition organization. Yamada co-organized the MIA Workshop at NAACL 2022, which hosted a competition on open-retrieval question answering across 16 languages. At the EMM-QA Workshop at ICML 2026, Yamada served as a co-organizer and also helped organize a competition on multimodal question answering using text and images.
Our results in international competitions include:
-
2025: NeurIPS MMU-RAG Competition
Our search agent won the static evaluation in the open-source division of the Text-to-Text track.
Paper: An Open and Reproducible Deep Research Agent for Long-Form Question Answering (Yamada et al., Preprint, 2025) -
2020: NeurIPS EfficientQA Competition
We placed second behind Facebook in the 6GB track and third behind Microsoft and Facebook in the unrestricted track.
Paper: NeurIPS 2020 EfficientQA Competition: Systems, Analyses and Lessons Learned (Min et al., PMLR 2021) -
2017: NIPS Human-Computer QA Competition
Our system won the competition among AI systems and defeated a team of six U.S. quiz champions in a live match at the NIPS 2017 workshop.
Paper: Studio Ousia’s Quiz Bowl Question Answering System (Yamada et al., The NIPS ’17 Competition: Building Intelligent Systems, 2018)


Related Papers
- Dynamic Injection of Entity Knowledge into Dense Retrievers (Yamada et al., Findings of EMNLP 2025)
- MIA 2022 Shared Task: Evaluating Cross-lingual Open-Retrieval Question Answering for 16 Diverse Languages (Asai et al., MIA 2022)
- Efficient Passage Retrieval with Hashing for Open-domain Question Answering (Yamada et al., ACL-IJCNLP 2021)
- Trick Me If You Can: Human-in-the-Loop Generation of Adversarial Examples for Question Answering (Wallace et al., TACL 2019)
Large Language Models (LLMs)
We aim to develop large language models (LLMs) that can flexibly handle diverse kinds of knowledge, including advanced domain expertise, knowledge specific to organizations or individuals, and knowledge that requires frequent updates. Our research includes, for example, methods for learning knowledge efficiently, updating and correcting it while preserving existing capabilities, and applying learned knowledge to a variety of tasks and situations.
We also study the use and transfer of knowledge across languages. For example, we develop methods for applying knowledge learned in one language to another and analyze the factors behind performance differences across languages.
Selected Achievements
We have proposed LUKE, a language model that uses entity information; mLUKE, a multilingual model; and LEIA, a method for facilitating cross-lingual knowledge transfer in language models. The LUKE paper has been cited more than 1,000 times, and the model is also included in Hugging Face Transformers. Our foundational work on knowledge representations supporting these efforts includes Wikipedia2Vec, which learns vector representations of words and entities from Wikipedia; NTEE, which embeds texts and entities in a shared vector space; and EASE, which learns sentence representations using information about related entities.
Related Papers
- LEIA: Facilitating Cross-lingual Knowledge Transfer in Language Models with Entity-based Data Augmentation (Yamada and Ri, Findings of ACL 2024)
- Entity Embedding Completion for Wide-Coverage Entity Disambiguation (Oba et al., Findings of EMNLP 2022)
- A Multilingual Bag-of-Entities Model for Zero-Shot Cross-Lingual Text Classification (Nishikawa et al., CoNLL 2022)
- EASE: Entity-Aware Contrastive Learning of Sentence Embedding (Nishikawa et al., NAACL 2022)
- Global Entity Disambiguation with BERT (Yamada et al., NAACL 2022)
- mLUKE: The Power of Entity Representations in Multilingual Pretrained Language Models (Ri et al., ACL 2022)
- LUKE: Deep Contextualized Entity Representations with Entity-aware Self-attention (Yamada et al., EMNLP 2020)
- Wikipedia2Vec: An Efficient Toolkit for Learning and Visualizing the Embeddings of Words and Entities from Wikipedia (Yamada et al., EMNLP 2020 System Demonstrations)
- Learning Distributed Representations of Texts and Entities from Knowledge Base (Yamada et al., TACL 2017)
- Joint Learning of the Embedding of Words and Entities for Named Entity Disambiguation (Yamada et al., CoNLL 2016)
Putting Research into Practice
We work on applications in academic research and industry to put our findings into practice. We aim to support activities such as reviewing academic literature, organizing claims and evidence across studies, and conducting technical investigations and R&D using organizational documents and advanced domain knowledge.
Alongside our papers, we release models, code, and evaluation datasets to make our research broadly accessible and reusable. Students participate in research projects throughout the full research process, from implementation, experimentation, and evaluation to publishing papers and releasing research outputs.