NUS researchers unveil MRAgent framework, reducing agentic memory token consumption 27x

What is the MRAgent LLM Memory Framework?

In June 2024, researchers at the National University of Singapore (NUS) published details of MRAgent, a new agentic memory framework designed to fundamentally change how large language model (LLM) agents handle long-term context. The framework, formally described as a Memory Reasoning Architecture for LLM Agents, was benchmarked against industry-standard evaluation suites including LongMemEval, where it achieved a 27-fold reduction in token consumption compared to LangChain’s LangMem framework. In practical terms, MRAgent required 118,000 tokens per query against LangMem’s 3.26 million, while simultaneously halving runtime and outperforming competing systems on accuracy.

The development matters because token consumption is not merely a technical metric. For any organisation deploying autonomous AI agents in production environments, tokens translate directly to API cost, response latency, and the practical upper limit of what an agent can reason over in a single session. For professional services firms, including environmental consultancies, engineering practices, legal teams, and local government technical units, these constraints have previously made long-horizon agentic workflows either prohibitively expensive or operationally unstable. MRAgent addresses that barrier at the architectural level rather than through incremental tuning.

The research arrives at a point where enterprise adoption of agentic AI is accelerating. Many organisations are moving beyond simple chatbot interfaces toward autonomous agents that manage document repositories, maintain project histories across months of interaction, and coordinate multi-step technical workflows. The scalability ceiling imposed by conventional retrieval architectures has been a well-recognised obstacle. MRAgent’s approach, if it holds up in production deployments beyond benchmark conditions, represents a meaningful shift in what becomes economically viable for sustained autonomous AI operation.

Key details of MRAgent architecture and benchmark performance

The central innovation in MRAgent is its treatment of memory as an interactive, navigable environment rather than a static database. Conventional Retrieval-Augmented Generation (RAG) systems and memory frameworks operate on a retrieve-then-reason model: when an agent needs context, it executes a vector search or keyword lookup, retrieves a large block of potentially relevant content, and loads the entire result into the prompt. This approach is computationally blunt. It frequently includes irrelevant material, inflates token counts, and degrades both speed and reasoning quality when the agent is working across extended interaction histories or large document sets.

MRAgent replaces this with a structure called a Cue-Tag-Content Graph. Memory is organised so that lightweight associative semantic tags act as indexed bridges between fine-grained user cues and the actual substantive memory content. When a query arrives, the backbone LLM, tested in the NUS research using Claude Sonnet 4.5 and Gemini 2.5 Flash, does not execute a single flat retrieval operation. Instead, it uses its own reasoning steps to traverse the memory graph iteratively. It infers search constraints from the query, follows the most contextually promising paths through the tag layer, and prunes branches that do not yield relevant connections before ever loading the heavier episodic content into the prompt. The result is that only the specifically relevant memory fragments are reconstructed and presented to the model at query time.

A further efficiency gain comes from the construction phase. Competing frameworks such as Mem0, LangMem, and A-MEM perform significant summarisation and relationship-analysis work during memory ingestion, which is computationally intensive and occurs regardless of whether those relationships will ever be queried. MRAgent defers this complex relational reasoning to query time, performing it on demand and in a query-specific manner. This design choice reduces ingestion overhead while concentrating computational effort where it produces the most value. The benchmark figures are concrete: 118,000 tokens per query for MRAgent against 3.26 million tokens for LangMem, a 27x reduction, alongside reported runtime halving and superior accuracy scores on long-horizon reasoning tasks across the LongMemEval benchmark suite.

The choice of backbone models is also noteworthy. Testing against both Claude Sonnet 4.5 and Gemini 2.5 Flash suggests the architecture is not tightly coupled to a single model provider. The graph traversal and active reconstruction mechanism relies on the reasoning capability of the backbone LLM, which means performance will scale with improvements in the underlying models over time. For enterprise deployments where model provider selection is driven by factors including data residency, commercial agreements, or regulatory requirements, this provider-agnostic design is a practical advantage.

arxiv.org
Image source: arxiv.org

Australian context: AI efficiency frameworks and professional services implications

Australia’s professional services sector, including environmental consulting, engineering, legal, and local government technical advisory, is in the early stages of integrating agentic AI into substantive workflows. Firms are beginning to deploy AI agents to assist with document review, regulatory correspondence management, project data synthesis, and long-running client engagement support. The primary constraint reported by technical leads at these firms has consistently been the cost and latency of maintaining meaningful context across extended project lifecycles. A contaminated land assessment, for example, may span 12 to 36 months, involve hundreds of laboratory reports, multiple regulatory exchanges, and evolving site conceptual models. An autonomous agent capable of maintaining coherent, accurate context across that kind of extended, document-heavy history has until now been difficult to justify on cost grounds alone. MRAgent’s token efficiency figures suggest that threshold may be shifting.

References and related sources

How iEnvi can help

iEnvi integrates technology and data-driven approaches into environmental consulting. We monitor AI and technology developments that affect how environmental professionals deliver services to clients.


This is an iEnvi Machete news summary. Prepared by iEnvi to summarise the source article for environmental professionals tracking AI, data, and technology developments that affect consulting and project delivery.

Published: 29 Jun 2026

Need advice on this topic? Speak to an iEnvi expert at info@ienvi.com.au or 1300 043 684, or contact us online.

Need advice on this issue? iEnvi provides practical, senior-led environmental consulting across contaminated land, remediation, ecology and environmental risk.

Contaminated land services Talk to iEnvi