Джорди Зомер опубликовал техническую статью о памяти LLM-агентов, в которой на примере исследований уязвимостей объясняет, почему агенту важно поддерживать актуальное знание, а не просто извлекать фрагменты прошлых разговоров.
В Lemmalog LLM преобразует естественный язык, код и вывод отладчика в структурированные факты. Datalog-движок детерминированно применяет правила и поддерживает производные факты.
В LongMemEval Lemmalog получил F1 0,463 и accuracy 0,575. Контекст для отвечающей модели был примерно в 38 раз меньше полного.
В категории Knowledge Update Lemmalog набрал 0,579 против 0,528 у PropMem. В multi-session reasoning он получил 0,211 против 0,582 у PropMem, поскольку нужные факты часто вообще не извлекались.
Проверка утверждений:
- 28 августа 2026 года Джорди Зомер опубликовал техническую статью «I accidentally turned LLM memory into program analysis», в которой на примере исследований уязвимостей объясняет, почему агенту важно поддерживать актуальное знание, а не просто извлекать фрагменты прошлых разговоров. (подтверждено первоисточником: доказательство; «I accidentally turned LLM memory into program analysis [ 28 Aug 2026 ] // JORDY ZOMER // 19 MIN READ Over the past few months I have been playing around quite a bit with LLM agents, particularly for vulnerability research. They are becoming surprisingly good at navigating large codebases, explaining unfamiliar subsystems and helping explore potential attack surfaces. However, once an investigation starts taking a few hours, I kept running into the same problem: the model would slowly lose track of what we had actually established.»)
- Lemmalog разделяет работу так: LLM преобразует естественный язык, код и вывод отладчика в структурированные факты, а Datalog-движок детерминированно применяет правила и поддерживает производные факты. (подтверждено первоисточником: доказательство; «This means that the LLM is still responsible for understanding natural language, source code, debugger output and all the other messy information that appears during an investigation. LLMs happen to be quite good at this. But once that information has been converted into structured facts, we no longer need the model to repeatedly determine all of its consequences. The database can do that instead.»)
- При удалении факта Lemmalog учитывает несколько независимых оснований одного вывода: вывод сохраняется, пока остаётся хотя бы одно основание, а зависимости позволяют автоматически инвалидировать затронутые заключения. (подтверждено первоисточником: доказательство; «Here c has two separate reasons for being true. If we remove a , we cannot simply remove c , because b still provides another derivation for it. However, if we remove both a and b , c should disappear as well. This turns out to be quite important during vulnerability research, because a conclusion may be supported by multiple observations.»)
- Движок хранит происхождение производных фактов, поэтому у заключения можно запросить цепочку оснований, а при опровержении наблюдения автоматически удалить зависимые выводы. (подтверждено первоисточником: доказательство; «Because Lemmalog already tracks the dependencies of derived facts, we can ask it for the provenance of a conclusion.»)
- Lemmalog поддерживает интервалы действительности фактов, что позволяет отдельно отвечать, верно ли утверждение сейчас и почему оно считалось верным раньше. (подтверждено первоисточником: доказательство; «For this reason Lemmalog can associate facts with validity intervals. Conceptually, we can represent the state as something like: viable(primitive_a) [10:14, 12:37) not_viable(primitive_a) [12:37, …) This allows us to answer both: Is primitive_a viable now? and: Why did we think primitive_a was viable earlier?»)
- Автор различает семантический поиск релевантной информации и поддержание текущей истины: векторная база решает первую задачу, а Lemmalog экспериментирует со второй; оба подхода можно совмещать. (подтверждено первоисточником: доказательство; «Retrieval is very good at the first problem. Lemmalog is mostly an experiment in solving the second one. The two can also be combined, which is what I currently do.»)
- На момент публикации движок поддерживал инкрементальные вычисления, откаты, provenance, временные факты, агрегации, согласование сущностей, гибридный поиск и demand-driven queries; для прямого доступа агентов был также MCP-сервер. (подтверждено первоисточником: доказательство; «The engine itself now supports incremental evaluation, retractions, provenance, temporal facts, aggregations, entity reconciliation, hybrid retrieval, demand-driven queries and a bunch of other things that I probably added because implementing Datalog features is more fun than I expected. There is also an MCP server which allows agents to use Lemmalog directly.»)
- Автор протестировал Lemmalog на LongMemEval и LoCoMo в стандартной схеме MemEval; при загрузке данных извлечение выполняла Claude Sonnet 4.6, а последующие этапы использовали стандартные модели чтения и оценки бенчмарков. (подтверждено первоисточником: доказательство; «So I plugged it into MemEval and tested it on both LongMemEval and LoCoMo using their standardized reader models and evaluation setup. Extraction during ingestion is Claude Sonnet 4.6 (chunked and file-cached, so it is paid once per conversation); everything after extraction uses the benchmark’s own standardized readers and judges.»)
- В трёх запусках LongMemEval Lemmalog получил F1 0,463 ± 0,010 и accuracy 0,575 ± 0,004; контекст для отвечающей модели был примерно в 38 раз меньше полного — около 2 700 против 104 000 токенов на вопрос. (подтверждено первоисточником: доказательство; «Lemmalog F1: 0.463 +/- 0.010 Accuracy: 0.575 +/- 0.004 For comparison, the published memory-system results are: PropMem 0.550 SimpleMem 0.480 Lemmalog 0.463 +/- 0.010 OpenClaw 0.244 Full Context 0.222 My own full-context GPT-4.1 run scored 0.197 F1. So Lemmalog is not beating PropMem yet, and it is still slightly behind SimpleMem, but it gets more than twice the F1 of giving GPT-4.1 the entire conversation. More amusingly, the context passed to the answering model is roughly 38 times smaller . Full context: ~104,000 tokens/question Lemmalog: ~2,700 tokens/question»)
- В категории LongMemEval Knowledge Update Lemmalog набрал 0,579 против 0,528 у PropMem и 0,202 у full context, но в multi-session reasoning заметно уступил: 0,211 против 0,582 и 0,382 соответственно; автор связал провал с тем, что нужные факты часто вообще не извлекались. (подтверждено первоисточником: доказательство; «The result I found most interesting was Knowledge Update. Lemmalog scored 0.579 , compared with 0.528 for PropMem and 0.202 for full context. Knowledge Update is basically the situation I originally cared about: we believed A | later we learn that A is no longer true | what should we believe now? So seeing Lemmalog top the published field on the category that most closely resembles maintained program state was rather satisfying. Single-session factual memory also worked surprisingly well. Lemmalog reached 0.790 on user facts and 0.672 on assistant facts, while temporal reasoning reached 0.416 , almost identical to PropMem’s 0.424 in that run. The obvious remaining problem is multi-session reasoning: PropMem 0.582 SimpleMem 0.382 Lemmalog 0.211 Diagnosing those failures was interesting: the information usually was not mis-connected, it was simply never extracted. If the extractor never emits a fact for the Airbnb booking, no amount of derivation is going to answer a question about it.»)
- На полном LoCoMo из 1 986 вопросов Lemmalog показал F1 0,533 ± 0,001 по трём запускам; среди выделенных систем памяти он оказался позади PropMem и OpenClaw. (подтверждено первоисточником: доказательство; «Lemmalog LoCoMo: 0.533 +/- 0.001 F1 The published comparison looks like this: System F1 PropMem 0.605 OpenClaw 0.557 Full Context 0.542 Lemmalog 0.533 ± 0.001 Hindsight 0.489 Graphiti 0.416 Memory-R1 0.389 SimpleMem 0.358 So Lemmalog currently sits third among the dedicated memory systems in this comparison, behind PropMem and OpenClaw.»)
- После исправления сравнения дат как внутренних идентификаторов символов и нормализации дат во временные метки результат temporal reasoning на LoCoMo вырос с 0,257 до 0,454 — почти на 20 пунктов F1. (подтверждено первоисточником: доказательство; «The initial version of Lemmalog scored: 0.257 After fixing temporal normalization and retrieval: 0.454 The bug was actually quite funny. At one point I was comparing date-like values as interned Datalog symbols. The engine’s < operator on symbols compares their internal ids. Internal ids are obviously not dates :) After normalising extracted dates into comparable integers and deriving happened_before from actual timestamps, temporal performance jumped by almost twenty F1 points.»)
Первоисточники:
оценка 80.4 · тип guide · ревизия 1 · истории st-n3di2f