Skip to main navigation Skip to search Skip to main content

Efficient optimization of large language models: a hybrid approach combining linear attention, chunk, and recurrent

  • Cheng Zhang*
  • , Linlin Shen
  • , Yudong Li
  • *Corresponding author for this work

Research output: Journal PublicationArticlepeer-review

Abstract

This research proposes a hybrid approach that combines linear attention, chunking, and recurrent mechanisms to address the efficiency issues of Large Language Models(LLMs) within the traditional transformer framework. Our approach integrates three key innovations: We use linear attention to employ kernel function mapping to reduce time and space complexity from O(n2) to O(n); The proposed dynamic chunk-based processing, can compress 5 times KV cache with mean pooling; Through 3 different ways, our hard thresholding, adaptive gating, and hierarchical chunking, can filter token and reduce load. The result shows that it can actually improve the efficiency of LLM, and performs excellently among some evaluation tools. Experiments demonstrate that our 3.2B parameter model achieves excellent performance in multiple benchmark tests, outperforming dense models of similar scale and even matching the performance of larger models in certain tasks, which provides a theoretically grounded and empirically validated framework for efficient LLM optimization.

Original languageEnglish
Article number163
JournalComplex and Intelligent Systems
Volume12
Issue number6
DOIs
Publication statusPublished - Jun 2026
Externally publishedYes

Free Keywords

  • Chunk-based KV compression
  • Large language model
  • Linear attention
  • Recurrent

ASJC Scopus subject areas

  • Information Systems
  • Engineering (miscellaneous)
  • Computational Mathematics
  • Artificial Intelligence

Fingerprint

Dive into the research topics of 'Efficient optimization of large language models: a hybrid approach combining linear attention, chunk, and recurrent'. Together they form a unique fingerprint.

Cite this