Join us for an evening focused on how knowledge graphs work and how they are applied in practice. Whether you write code, organize data, or simply want to understand what graph technology offers, this meetup is a welcoming space to learn and connect. Our focus tonight is on making AI stick to the facts and using knowledge graphs to enforce accuracy. Whether by keeping an LLM honest through strict rules, or by bypassing language models entirely for pure computation.
Agenda
When: 30th of September 2026 Where: Crowd Collective Office, Keilaranta 4, Espoo
- 16:30Doors open & some snacks
- 17:00Welcome & opening
- 17:10The Classical Music Knowledge Graph: Rules, Reasoning and an Idea for Ontology-Driven MCP Tools by Katariina Kari
- 18:10Break
- 18:30AICHAX: Grounded, Not Generated — A Graph Engine Without an LLM, a Drug-Repurposing Benchmark, and an Idea for Grading Your Data Sources by Ricky Sun
- 19:30Onsite networking
- 20:30Event ends
Talks
Katariina Kari
The Classical Music Knowledge Graph: Rules, Reasoning, and an Idea for Ontology-Driven MCP Tools
Classical music listings live in the worst possible format for anyone trying to answer a simple question like "when is this conductor next performing, and what's on the programme?" — scattered across dozens of orchestra websites, each with its own layout, language, and update cadence. The Semantic Score project turns that mess into a linked-data knowledge graph: a Custom Music Ontology (CMO) covering performances, programmes, performers, composers and venues, populated by a self-correcting extraction pipeline that sends each orchestra's concert page to an LLM and asks for RDF back — then checks its work.
Part 1 — the work: rules that keep an LLM honest. Getting an LLM to emit valid, consistent RDF isn't a one-shot prompt, it's a feedback loop. The pipeline validates every extraction against SHACL shapes (a source-layer shape for the raw extraction, an inference-layer shape for completeness after reasoning), checks that every asserted URL actually resolves, and feeds failures back to the model for correction. Once data is accepted, a library of SPARQL CONSTRUCT rules — one file per inference, from schema.org typing to composer nationality to "performs-with" relationships — runs to a fixpoint over the whole assertions store, materialising everything the raw extraction implies but never states outright. The result is a hard boundary between asserted fact and derived fact, with the derived layer fully disposable and reproducible.
Part 2 — the idea: ontologies as the contract for MCP tools, not just the data. Most MCP tool servers are hand-written: someone decides "there should be a search_performances tool" and hard-codes its arguments, return shape, and the query behind it. The idea this talk puts forward is that when a domain already has an ontology and SHACL shapes describing what's true and how it's structured, that same artefact should be doing double duty — defining the tool's contract as well as the data's shape. A tool's inputs mirror the ontology's classes and properties; its outputs are constrained by the same shapes that already validate the graph; and the SPARQL behind it is generated from the ontology rather than written by hand and left to drift out of sync. A small proof-of-concept — a stdio MCP server exposing one tool over the live SPARQL endpoint, walked through step by step from human question to tool call to answer — is offered not as a finished product but as the smallest possible sketch of that idea, and a starting point for discussing what a genuinely ontology-driven MCP server would need.
Ricky Sun
AICHAX: Grounded, Not Generated — A Graph Engine Without an LLM, a Drug-Repurposing Benchmark, and an Idea for Grading Your Data Sources
Some questions can be answered with plausible text. Others have to be defended — to a reviewer, a regulator, a clinician, or anyone whose next words are “how do you know?” For the second kind, a language model is the wrong instrument: it produces the shape of an answer, fluently and confidently, with nothing underneath that can be checked. AICHAX is built on the opposite premise — that a defensible answer is computed, by traversing evidence, and arrives with its derivation attached. It runs on Ultipa GQLDB, a graph engine that speaks ISO GQL, over a billion-scale general knowledge graph and many domain-specific graphs serving individual verticals. There is no language model anywhere in the answer path: no embeddings, no retriever, no generation step.
Part 1 — the work: what it takes to compute an answer instead of retrieving one. Three properties carry the weight, and none of them is free. Reasoning runs inside the database, at query time — GQLDB derives inverse, symmetric, transitive, sub-property and property-chain relationships as virtual edges while the query executes, so nothing is materialized and a correction to the data changes the inferred answer on the very next read. For anyone who has run CONSTRUCT rules to a fixpoint, that is the same clean boundary between asserted and derived fact, without the re-computation — and GQLDB will import your OWL/RDF, export it losslessly, and federate to a remote SPARQL endpoint from inside a GQL query. Every answer carries its derivation, down to the source record it rests on. And the same question returns the same answer — which sounds trivial, and is in fact something you engineer deliberately and then test for. Drug repurposing is where we chose to prove all this, because it is a domain that will actually score you: rank candidate compounds by traversing declared biological mechanisms, then hide every known drug across 75 diseases and measure how many the engine brings back in its top ten. Across 752 held-out pairs it recovers them 8.6× more often than a random draw of the same size — an unglamorous number, stated the unglamorous way on purpose. The vertical is one studio among several on the same engine; the machinery underneath it is not biomedical.
Part 2 — the idea: your benchmark should be grading your data, not just your engine. Most of us build a retrospective benchmark to answer one question — does the thing work? — and then we leave it there. The idea this talk puts forward is that the same artifact should be doing double duty. Run it again with one source database removed, and the drop in recall is that source’s contribution, measured in the units the application is actually judged by rather than asserted in a datasheet. Across the integrated biomedical graph behind that benchmark, the result was uncomfortable. Three-quarters of the edges are never traversed at all. Thirteen integrated source databases move the ranking by exactly zero, while the single smallest contributor carries two-thirds of the ranking quality. Size is not value. This is offered not as a finished methodology but as an instrument any graph team can build in days, and as a starting point for two conversations: what we owe the sources we integrate, and where the line sits between a structural hypothesis and a causal claim — a line this engine deliberately does not cross.
Speakers

Katariina Kari
Co-founder, Knowledge Graph Academy
Katariina Kari is an expert in semantic web technologies and enterprise knowledge graphs. She has collaborated with major retail brands building enterprise-level knowledge graphs that power search, recommendations, and richer customer experiences. Recognized as one of the world’s top talents with hands-on expertise in semantic technologies, Katariina is a frequent speaker at industry events and serves as co-chair of Connected Data London, a leading conference in the field. She recently founded the Knowledge Graph Academy with other semantic web experts such as Tony Seale and Jessica Talisman. Her work bridges deep technical skill with a creative, human-centered approach to technology. Holding both a Master of Science and a Master of Music, Katariina proudly identifies as a musically-inclined art-loving, tech-savvy nerd. Before her current career in AI and data, she ran her own consultancy from 2012 to 2016, advising classical music organizations and artists on digital outreach strategies. Outside of work, she finds joy in baking sourdough rye bread, playing the cello, and swimming. Her latest craft fixation has to do with poppana!

Ricky Sun
Founder & CEO, Ultipa
Mr. Ricky Sun is a serial entrepreneur, world-class high-performance storage and computing system expert, started his career in the heart of Silicon Valley with his professor a year before his graduation from SCU. Over the past 20+ years, he went through 3 M&A, while Ultipa is his 4th venture. Ricky was formally CTO of EMC's largest global R&D center, Chief architect of Splashtop, a Pre-IPO unicorn startup that built real-time operating system that directly inspired Google's Chrome OS and real-time SaaS-grade remote desktop products. Ricky launched Ultipa with the belief and aim that real-time graph database is the ultimate form of database that empowers smart enterprise with graph-augmented AI (a.k.a, XAI). Ricky is the holder of more than 50 U.S/CN patents. Ricky graduated from SCU, majored in MSCE with distinction, and BSCS from Tsinghua University.
Save your spot
Ready to join us? Sign up below and we'll save your spot.
