Ramp Labs Introduces Multi-Agent Memory Sharing Solution, Token Consumption Reduced by Up to 65%

By: theblockbeats.news|2026/04/11 14:19:36
0
Share
copy

BlockBeats News, April 11th, AI infrastructure company Ramp Labs released research results on "Latent Briefing", achieving efficient memory sharing among multi-agent systems through direct compression of large-scale model KV cache, significantly reducing Token consumption without sacrificing accuracy.


In mainstream multi-agent architectures, the Orchestrator decomposes tasks and repeatedly calls Worker models. As the inference chain extends, Token usage exponentially inflates. The core idea of Latent Briefing is to leverage the attention mechanism to identify the truly critical parts in the context, directly discard redundant information at the representation layer, rather than relying on the slow-speed LLM summary or the unstable RAG retrieval.


In the LongBench v2 benchmark test, this method performed remarkably: Worker model Token consumption decreased by 65%, the median Token savings for medium-length documents (32k to 100k) reached 49%, the overall accuracy improved by approximately 3 percentage points compared to the baseline, and the additional time for each compression was only about 1.7 seconds, achieving a speedup of about 20 times compared to the original algorithm.


The experiment used Claude Sonnet 4 as the Orchestrator, and Qwen3-14B as the Worker model, covering various document scenarios such as academic papers, legal documents, novels, and government reports. The research also found that the optimal compression threshold varies depending on task difficulty and document length—difficult tasks are suitable for aggressive compression to filter out speculative reasoning noise, while long documents are more suitable for mild compression to retain scattered key information.

-- Price

--

You may also like

Morning News | The draft amendment to the People's Bank of China Law aims to clarify the legal status of digital renminbi; South Korea will transfer about 40 unregistered virtual asset service providers to law enforcement agencies

Overview of Important Market Events on June 24

The cryptocurrency industry has entered the "Show Me" era: merely relying on vision is no longer enough

The awareness level of the audience in the cryptocurrency industry—including media, institutions, and retail investors—is steadily increasing, and this trend has become a foregone conclusion.

Interpreting the Ethereum Foundation's new structure: Reaffirming self-sovereignty amid institutional trends

The Ethereum Foundation has announced a new five-layer working framework, clarifying the focus of future development and reaffirming its commitment to decentralized core values amidst the wave of institutionalization.

Former SpaceX engineer reconstructs the financial execution system using first principles

Plan Execution Lab completes angel round financing for Singapore family office, with a valuation of 50 million USD.

Tidal Investment: We still have a positive outlook on the AI industry chain, but the reasons have changed

The intense financing by tech giants has triggered a panic of "AI peak," but the soaring capital expenditures of the five major cloud vendors and the bottlenecks in physical infrastructure indicate that the AI investment cycle is far from over; the second half of this grand performance has just begu...

Standard Chartered Bank sings a 50x rhapsody again, aiming for AAVE to reach 3500 USD

The throne of DeFi lending still exists, but the foundation beneath the throne needs to undergo a reconstruction or reinforcement.

Contents

Popular coins

Latest Crypto News

Read more
iconiconiconiconiconiconicon
Customer Support:@weikecs
Business Cooperation:@weikecs
Quant Trading & MM:[email protected]
VIP Program:[email protected]