Memory-Augmented Evidence Graphs for Verifiable AI Reasoning in High-Stakes Domains

Authors

  • Sahaj Tushar Gandhi Independent Researcher, San Francisco, CA, USA Author

DOI:

https://doi.org/10.15662/IJEETR.2026.0803016

Keywords:

Memory-Augmented Retrieval-Augmented Generation (RAG), Evidence Graphs, Verifiable AI, Provenance Tracking, Trustworthy Artificial Intelligence, Knowledge Graph Memory, High-Stakes Decision Support

Abstract

Training Large Language Models (LLMs) to be more objective by leveraging outside knowledge to create answers has shown to dramatically enhance their capabilities as Retrieval-Augmented Generation (RAG). But in most cases today, RAG systems are predominantly stateless and query evidence individually by query and do not store prior-proven knowledge, evidence provenance, history of contradiction and temporal integrity. These limits limit the reliability and accountability of high stakes sectors like health services, cybersecurity, money, lawful intelligence and compliance. This paper outlines a Memory- Augmented Evidence Graph (MAEG) system, which compiles Short-Term Evidence Memory (STEM) to empower active taskspecific reasoning with Long-Term Evidence Memory (LTEM), to empower persistent, provenanceaware organizational knowledge. The proposed architecture incorporates hybrid retrieval, evidence in the form of graphs, claim extraction, multi-stages persistent validation of evidence, contradiction detection, temporal consistency control and full audit log, to enable plausible and explainable AI reasoning. Key contributions include a two-layer evidence memory architecture, a provenance-backed evidence graph for persistent knowledge management, a verification pipeline to ensure only verified claims are stored, and temporal reasoning to detect evidence staleness or contradictions. Experimental outcomes on five benchmark datasets comprising 80,917 samples demonstrate that MAEG outperforms Vanilla RAG, Citation-based RAG, GraphRAG, RAPTOR, LightRAG, and HippoRAG, achieving 96.7% answer accuracy, 98.1% claim support, 97.5% evidence accuracy, and 96.8% contradiction detection. These results reveal that Long-Term Evidence memory is an important addition to reliability, transparency, and auditability of AI reasoning to enterprise decision-support systems.

References

[1] B. J. Gutierrez, Y. Shu, Y. Gu, M. Yasunaga, and Y. Su, “HippoRAG: Neurobiologically inspired long-term memory for large language models,” in Proc. 38th Conf. Neural Information Processing Systems (NeurIPS), 2024.

[2] D. Edge, H. Trinh, N. Cheng, J. Bradley, A. Chao, A. Mody, S. Truitt, and J. Larson, “From Local to Global: A GraphRAG Approach to Query-Focused Summarization,” arXiv preprint arXiv:2404.16130, 2024.

[3] Z. Guo, L. Xia, Y. Yu, T. Ao, and C. Huang, “LightRAG: Simple and Fast Retrieval-Augmented Generation,” arXiv preprint arXiv:2410.05779, 2024.

[4] P. Sarthi, S. Abdullah, A. Tuli, S. Khanna, A. Goldie, and C. D. Manning, “RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval,” in Proc. Int. Conf. Learning Representations (ICLR), 2024.

[5] Y. Wang, R. Ren, J. Li, X. Zhao, J. Liu, and J. Wen, “REAR: A Relevance-Aware Retrieval-Augmented Framework for Open-Domain Question Answering,” in Proc. Conf. Empirical Methods in Natural Language Processing (EMNLP), pp. 5613–5626, 2024.

[6] A. Mallen, A. Asai, V. Zhong, R. Das, D. Khashabi, and H. Hajishirzi, “When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories,” in Proc. 61st Annu. Meeting Assoc. Comput. Linguistics (ACL), pp. 9802–9822, 2023.

[7] H. Chen, R. Pasunuru, J. Weston, and A. Celikyilmaz, “Walking Down the Memory Maze: Beyond Context Limit Through Interactive Reading,” arXiv preprint arXiv:2310.05029, 2023.

[8] J. Xie, K. Zhang, J. Chen, R. Lou, and Y. Su, “Adaptive Chameleon or Stubborn Sloth: Revealing the Behavior of Large Language Models in Knowledge Conflicts,” in Proc. Int. Conf. Learning Representations (ICLR), 2024.

[9] H. Shi, Z. Xu, H. Wang, W. Qin, W. Wang, Y. Wang, Z. Wang, S. Ebrahimi, and H. Wang, “Continual Learning of Large Language Models: A Comprehensive Survey,” arXiv preprint arXiv:2404.16789, 2024.

[10] G. Izacard, M. Caron, L. Hosseini, S. Riedel, P. Bojanowski, A. Joulin, and E. Grave, “Unsupervised Dense Information Retrieval with Contrastive Learning,” Trans. Mach. Learn. Res., 2022.

[11] J. Ni, C. Qu, J. Lu, Z. Dai, G. Hernandez Abrego, J. Ma, V. Zhao, Y. Luan, K. Hall, M.-W. Chang, and Y. Yang, “Large Dual Encoders Are Generalizable Retrievers,” in Proc. Conf. Empirical Methods in Natural Language Processing (EMNLP), pp. 9844–9855, 2022.

[12] X. H. Lu, “BM25S: Orders of Magnitude Faster Lexical Search via Eager Sparse Scoring,” arXiv preprint arXiv:2407.03618, 2024.

[13] H. Trivedi, N. Balasubramanian, T. Khot, and A. Sabharwal, “MuSiQue: Multihop Questions via Single-Hop Question Composition,” Trans. Assoc. Comput. Linguistics, vol. 10, pp. 539–554, 2022.

[14] T. Yuan, X. Ning, D. Zhou, Z. Yang, S. Li, M. Zhuang, Z. Tan, Z. Yao, D. Lin, B. Li, G. Dai, S. Yan, and Y. Wang, “LV-Eval: A Balanced Long-Context Benchmark with Five Length Levels up to 256K,” arXiv preprint arXiv:2402.05136, 2024.

[15] C. Lee, R. Roy, M. Xu, J. Raiman, M. Shoeybi, B. Catanzaro, and W. Ping, “NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models,” in Proc. Int. Conf. Learning Representations (ICLR), 2025

Downloads

Published

2026-06-09

How to Cite

Memory-Augmented Evidence Graphs for Verifiable AI Reasoning in High-Stakes Domains. (2026). International Journal of Engineering & Extended Technologies Research (IJEETR), 8(3), 5160-5173. https://doi.org/10.15662/IJEETR.2026.0803016