Megagon Labs investigates this question through projects like WiTQA, a benchmark that systematically evaluates when retrieval helps versus degrades LLM question-answering performance. The findings reveal that RAG benefits depend heavily on the type of question, the relevance of retrieved documents, and the model’s existing knowledge — making RAG a context-dependent rather than universal improvement.