
Solves the signal explosion problem of existing 'hyper-connections' by projecting onto a 'doubly stochastic matrix'... Combines infrastructure technologies such as 'kernel fusion' to prevent system overload... Verified in a 27B model
Chinese AI company DeepSeek has unveiled a new architecture technology called 'mHC (Manifold-Constrained Hyper-Connections)' that can simultaneously resolve the chronic issues of 'training instability' and 'memory bottlenecks' in large language models (LLMs). This technology is characterized by its ability to mathematically and precisely control the flow of information inside the AI model, maximizing computational efficiency even as the model's size increases.
According to a paper published by DeepSeek researchers on the 31st of last month, the 'hyper-connection (HC)' method, which broadens information transmission pathways to improve model performance, has recently been gaining attention in the AI academic community. However, this method had a fatal limitation in that as the layers deepen, the signal is exponentially amplified (Explosion) or lost, breaking the 'Identity Mapping' property during large-scale training.
The newly developed 'mHC' introduces a novel concept called 'Manifold Constraints' to resolve this instability. The researchers projected the Residual Connection matrix into the 'Birkhoff polytope' space, converting it into a 'Doubly Stochastic Matrix' where the sum of the rows and columns of the matrix always equals 1.
Simply put, it acts as a mathematical 'balancing device' to prevent the total volume of signals from arbitrarily increasing or decreasing while data is transferred. Through this, mHC activates the exchange of information while stably maintaining the size of the signal, successfully controlling the signal amplification phenomenon that soared up to 3000 times in the existing HC method.
In particular, this paper draws attention as it goes beyond simply presenting theories and implements optimization at the system infrastructure level. To solve the 'Memory Wall' problem caused by the surge in data input/output (I/O) when applying hyper-connections, the researchers applied ▲Kernel Fusion ▲Recomputing ▲Communication Overlap (DualPipe) technologies.
As a result, despite expanding the residual stream width by 4 times (n=4) in a 27 billion (27B) parameter model, DeepSeek suppressed the additional training time cost to a mere 6.7%. This proves that high-performance models can be trained at a commercially viable cost level.
Actual performance indicators also showed a distinct improvement. As a result of the 27B model test, mHC recorded scores that were 2.1% and 2.3% higher than existing hyper-connection models in the 'BBH' benchmark, which requires complex reasoning, and the 'DROP' benchmark, which evaluates reading comprehension ability, respectively. The training loss gap also decreased faster than the existing models, demonstrating excellent convergence speed.
DeepSeek representatives stated, "mHC is a study that has identified the impact of the topological architecture of large-scale models on training optimization," and predicted, "It will present a new direction for the development of next-generation foundation models in the future." This research achievement is expected to become a core technology for catching the two birds of efficiency and stability amidst the trend of increasing AI model sizes.
Company financial data, investment reports, and startup analysis — all in one place
Explore PitchdeckCurated news, every week — straight to your inbox
Every Friday · Unsubscribe anytime