Enterprise context and ontology
Build an ontology out of real enterprise applications.
You get roughly half a million documents drawn from nine different sources: Slack, Gmail, Linear, Google Drive, HubSpot, Fireflies, GitHub, Jira and Confluence. They arrive with all the noise an actual company has, including misfiled documents, near duplicates and statements that flatly contradict each other.
Your job is to turn that into a clean, queryable ontology in HydraDB and then answer questions ranging from simple lookups to multi-hop reasoning, conflict resolution, and correctly recognising when the answer just is not in there.
Extraction is the easy part now that LLMs are cheap. The hard part is entity resolution and ontology alignment: deciding that “Sam”, “@soham” and “S. Ratnaparkhi” are one person, and figuring out which of two contradictory statements to trust.