SourcesZep Temporal Knowledge Graph Architecture for Agent MemoryarXiv·high signalXBlueskyLinkedInCopy linkGraphiti engine builds temporally-aware knowledge graphs outperforming MemGPT on Deep Memory Retrieval benchmarkSourceSource pagearXiv↳ Follow the threadStack layer / Update threadNCP-ArchPreview reaches OLMo-3-7B's final pretraining loss on 51.3% of the tokens by predicting multi-token conceptsarXiv / HuggingFace Daily PapersStack layer / ContrastEvoSafeHarness searches policies and code together to build a per-model safety harness, cutting attack success from 45.6% to 10.0%arXiv (2609.05903)Stack layer / Update thread100 LLM agents running a simulated town economy for 26 weeks froze the money supply: 0.3% of prices ever changed, and memory deletion made no differencearXivStack layer / ContrastSkill optimization via contextual bandits cut optimization cost 55-58% using only 50 examples per benchmarkarXiv 2609.11682Policy dependency / Follow-up threadMr.LHDR benchmark: the best deep-research agent gets 43.1% of final answers right but only 34.3% with a correct intermediate chainarXivStack layer / Follow-up threadΦ-Bench tests whether LLMs can engineer their own serving and training stack, from kernels to end-to-end optimizationarXivStack layerBenchShield raises reward-hacking detection in agent benchmarks from 23-94% to 77-100% full-chain recall, using a lifecycle model of reward eventsarXivStack layerThe Power Flexibility Index Measures How Much Throughput an LLM Training Job Loses When You Cut Its PowerarXiv 2609.11542