Starmie: Table Discovery in Data Lakes – Exploring State-of-the-Art Success and Future Directions

In this work, we propose an end-to-end framework named Starmie. Dataset discovery from data lakes is a critical way to utilize open-domain data within the enterprise. To overcome the issues stemming from data quality and incomplete metadata in data lakes, it is essential to support the problem of table union search, which aims to find all tables that are unionable with the query table, given a query table and a collection of data lake tables.
SIGMOD 2023 Highlights

The ACM SIGMOD conference is the leading forum for the principles, techniques, and applications of database management systems and data management technology. There are 26 sponsors for SIGMOD this year, and Megagon Labs was a Silver sponsor. The conference consisted of the research track, the industry track, the demonstration track, 11 tutorials, and 10 workshops.
Megagon KnowledgeHub: A versatile knowledge repository to bridge the gap in HR AI applications

At Megagon Labs, we are working on symbiotic models and systems (Figure 1) that take advantage of LLMs as well as structured (knowledge bases [KBs], knowledge graphs [KGs], databases [DBs], etc.) and unstructured (texts) information in a continuous and (semi-) automated machine-learning paradigm. In this post we will describe Megagon KnowledgeHub and how our research and development benefits from it.
Innovative Machine Learning Takes Center Stage

We shine a spotlight on three cutting-edge AI projects that have been making waves in the industry: ZETT, CoCoSum, and ESE. These groundbreaking initiatives offer a glimpse into the future of AI and the transformative impact it holds across various domains.
Megagon Team Feature: Vishwas Mruthyunjaya

We’d like to introduce you to Vishwas Mruthyunjaya, Senior Data Scientist at Megagon Labs. We’ll discuss his growth at Megagon, the advice he’d give to aspiring data scientists and engineers and his interesting journey from Robotics to AI.
The Intern Experience at Megagon Labs: What You Can Expect

At Megagon Labs, we see bringing on interns as more than just hiring a short-term helping hand. As we welcome spring and summer interns, we’d like to share with you how we foster an environment of growth and career development for both mentors and interns.
Megagon Team Profile: Dan Zhang

Dan Zhang, Research Manager and Senior Research Engineer at Megagon Labs, gives us a recount of her journey from childhood to and through her career as a research engineer.
Weedle: Composable Dashboard for Data-Centric NLP in Computational Notebooks

To help NLP researchers and practitioners understand and improve their data, we introduce Weedle, an exploratory text analysis tool for data-centric NLP. Here are Weedle’s biggest strengths…
Feature Stores: Deep Learning, NLP, and Knowledge Graphs

We will introduce feature stores and examine the implications of deep learning on feature stores as well as discuss the role of feature stores as part of the emerging MLOps stack.
Sudowoodo: Contrastive Self-supervised Learning for Data Integration Applications

We introduce Sudowoodo, an end-to-end framework for a variety of data integration applications to resolve the limitations of data integration. Sudowoodo addresses the label requirement by leveraging contrastive learning to learn a data representation model from a large collection of unlabeled data items. This is realized by the contrastive objective that allows the model to learn how to distinguish pairs of similar data items from dissimilar ones that are likely to be distinct.