Order Matters: Assessing LLM Sensitivity in Multiple-Choice Tasks

Order Matters: Assessing LLM Sensitivity in Multiple-Choice Tasks

Explore the relationship between option arrangement and performance variations in Large Language Models (LLMs) during multiple-choice tasks. Through meticulous analysis, we uncovered substantial sensitivity of LLMs to the order of answer options, with performance fluctuations of up to 75% across different benchmarks.

Towards Enterprise Compound AI Systems

Towards Enterprise Compound AI Systems

Researchers at Megagon Labs have been exploring how we can address the challenges of building compound AI systems for enterprises. In this blog post, we introduce three projects that we have undertaken: (1) developing a suitable architecture for productizing compound AI systems, (2) optimizing agentic workflows with real-world constraints, and (3) benchmarking the performance of agents within a compound AI system, specifically in an enterprise setting.

Watchog: Leveraging Contrastive Learning for Enhanced Table Understanding and Column Annotation

By enabling robust and accurate column annotation, this innovative framework holds the potential to revolutionize data-driven decision-making processes across a multitude of industries. Watchog could empower businesses to extract valuable insights from product catalogs, pricing tables, and customer data repositories and use it to optimize their pricing strategies, and deliver personalized recommendations to enhance customer satisfaction and loyalty.

Less Is More for Long Document Summary Evaluation by LLMs

In the realm of text generation and summarization, the evaluation of generated summaries, especially for long documents, has always been a challenging task. To address these challenges, we evaluate long models using an innovative approach that significantly reduces evaluation costs and aligns more closely with human evaluations.