I accidentally turned LLM memory into program analysis

The author, a vulnerability researcher, found that LLM agents lose track of established facts during long investigations, leading to hallucinations. To solve this, they built Lemmalog, a Datalog engine that maintains a structured knowledge base. The LLM extracts facts from messy sources, while Lemmalog derives conclusions and tracks dependencies, automatically invalidating affected results when facts change. This approach provides provenance for conclusions and distinguishes between relevant past information and current truth, improving reliability in complex security research.
The amusing part is that our parser is probabilistic, while everything after it does not necessarily have to be.
- sim04ful
I reached a similar conclusion: LLMs should only really sit at the terminals of request fulfilment.
1. User request understanding: natural language -> a more rigorous representation, in my case Datalog.
2. Result interpretation: facts and derived facts -> natural language.
Between those terminals, the work should be mechanical reasoning over some ontology or formal knowledge structure.
That connects to another principle I've been thinking about, which I call Weathering: useful reasoning should change the shape of the system. If an LLM has already had to infer a relation, mapping, rule, or abstraction, repeated use should wear that inference into the system so that the next similar request doesn't require discovering it again from scratch.
With continued use, a weathering-capable system should therefore require less and less probabilistic intelligence for recurring work. Put another way, there should be a declining marginal cost of cognition since the products of intelligence harden into structure that can subsequently be reused and evaluated mechanically.
- akkad33
Has anyone tried formal verification with AI generated code? I can't convince my company to use it but I realise it's very easy to ask Claude to add a verification step locally on my own PRs
- Animats
So he's using an LLM to generate data stored in an "is_a" representation.
That's so classic AI.
Soon, he'll discover that he needs quantifiers. Then that "for all" is too strong sometimes, and he needs "for most". That way lies Cyc.
It's not a bad idea. But it does have a history.
- akkad33
I had tried to get long term memory out of Claude by indexing my notes with keywords and putting that in a sqllite database and Claude queries using full text search. Don't know how good it is, it seems to find things alright. My goal was to keep context small and only get Claude to ask for what it needs. Datalog seems like a great idea, will definitely try it out
- keeda
Very cool. I recall an HN submission (which I can't find offhand unfortunately) that did something similar -- it used an LLM to decompose articles into a set of statements which were used to construct an entity-relationship graph of facts and events. It then queried that using conventional graph query methods, much like DataLog / Lemmalog is doing here. I remember it was particularly effective at answering timeline-based queries that LLMs (back then) sucked at.
(See also Cyc: https://en.wikipedia.org/wiki/Cyc)
I think approaches like this are going to be (or maybe already are?) the basis of effective grounding of LLM responses in authoritative data sources. It should be possible to pinpoint any error to an incorrect traversal or an incorrect "fact." This would work best for concrete, unambiguous facts, however; fuzzy, ambiguous or opinion-based information will probably remain the purview of LLMs.