securitylab_nJuly 12, 2026🇷🇺Translated from Russian

Doubling GPT-3 Inference Speed and Eliminating Silicon Furnaces: Vertical Memory Architectures V-Die and MOSAIC Revolutionize HBM Cooling and Bandwidth for AI Accelerators

Future AI accelerators may soon feature memory chips literally standing on their edges. Researchers from South Korea and Japan have proposed rotating DRAM dies vertically to boost both the speed and capacity of HBM without turning multilayer stacks into poorly cooled silicon furnaces.

Two new architectures, V-Die and MOSAIC, were presented in June 2026 at the IEEE Symposium on VLSI Circuits. Both teams abandoned the conventional approach of stacking dynamic memory dies horizontally on top of a base die and connecting them through TSV (through-silicon via) channels. Instead, the engineers assemble the package first and then rotate the entire structure so that individual dies stand vertically, functioning like the fins of a radiator.

Current HBM places multiple memory dies above a base die and links the layers with short, wide TSV buses that deliver several terabytes per second of bandwidth, making it the foundation of powerful AI accelerators. However, increasing the number of layers makes heat removal increasingly difficult. Lower dies become extremely hot, and heat must travel through silicon, solder, insulating materials, and additional package layers. The TSVs themselves also consume valuable die area that could otherwise be used for memory cells.

The Korean V-Die architecture removes TSVs from the memory dies altogether. Each vertical die receives its own input/output lines along the bottom edge and connects directly to the substrate. According to the researchers’ calculations, the design provides four times more connections than HBM4 and cuts data read time by 37 percent. Microfluidic channels running between the dies maintain temperatures around 45 °C, whereas dense conventional HBM packages can exceed 80 °C under heavy load.

Simulations of a 16-die system demonstrated significant performance gains. When processing GPT-3-scale workloads, the V-Die architecture achieved 540 tokens per second compared with 296 tokens per second for an equivalent-capacity HBM4 configuration. First-token latency dropped by 32 percent, or roughly 24 milliseconds. These results are currently simulation-based; the research team is now preparing physical prototypes to validate thermal and electrical characteristics.

The Japanese MOSAIC project addresses a different challenge of vertical assembly: even small thickness variations among dozens of dies can misalign contact pads. Engineers from the University of Tokyo, Tohoku University, and RIKEN proposed transmitting data without direct metal contact. Miniature coils on the dies and substrate exchange signals through electromagnetic induction, eliminating the need for perfect physical alignment during assembly.

The experimental MOSAIC interface reached speeds of up to 4 Gbit/s per channel. Researchers expect to double the memory volume of HBM4 by placing the vertical block directly above the GPU while increasing maximum temperature by only about one degree. One configuration accommodates 98 dies and 294 GB of memory; further thinning of the dies could theoretically push capacity to 882 GB.

Neither V-Die nor MOSAIC is yet ready to replace production HBM. The Korean architecture exists primarily as calculations, while the Japanese prototype must still demonstrate acceptable cost, reliability, and high manufacturing yield. Nevertheless, both developments outline a promising route around the thermal barrier that increasingly constrains memory scaling for AI accelerators. Instead of endlessly stacking dies higher, engineers are proposing to rotate the entire structure and turn the memory dies themselves into an integrated heat-dissipation system.

Related articles

HabrOther

Deploying Self-Hosted Hysteria 2 Proxy on Debian-Based Linux VPS via Terminal

A detailed guide explains how to set up a personal Hysteria 2 proxy server on a KVM VPS running Debian or Ubuntu without any web panels. The process begins with generating ed25519 SSH keys, hardening the sshd_config file, and restricting access with ufw to only TCP port 22 and UDP port 443. Hysteria 2 is downloaded from GitHub, made executable, and configured using a TOML file that enables salamander obfuscation and a self-signed TLS certificate. A custom systemd unit ensures the service restarts on failure. The client configuration includes SHA256 pinning of the server certificate to prevent MITM attacks. The guide emphasizes manual CLI operations that apply equally to other services such as Nginx and stresses checking local laws before deployment.

AntiMalwareOther

Rostec Scales PCAT Platform Nationwide as Russia's First Industrial Marketplace

Rostec has expanded its PCAT platform to every organization within the state corporation that manufactures civilian products. Operating since 2025 and upgraded in September 2026, the platform now unites more than 180 enterprises and research organizations. Its catalog contains over 1,250 finished products along with 370 technological and manufacturing competencies. Visitors can locate not only equipment and components but also partners able to design, test, or produce required solutions. The portal receives more than 23,000 weekly visits, 60 percent of them from corporations and large enterprises. Rostec is extending the network into the regions through supply-chain agreements already signed with Krasnodar Krai and the oblasts of Tver, Tula, and Ryazan. In parallel the corporation launched the Robot Management System in November 2025 for centralized control of robots, sensors, and related IT services.

AntiMalwareOther

Kate Mobile Loses VK API Access After New Request Limits Exhaust Quota in 1.5 Days

Popular third-party Android client Kate Mobile has been cut off from VK services following the introduction of strict monthly API request caps. VK implemented the new limits on September 7, offering verified partners up to 100 million requests per month while requiring payment for additional access by third-party services. Kate Mobile developers had requested pricing details in advance but received no response from VK. Calculations showed that the app's real user base would consume the entire 100-million-request allowance in roughly 36 hours, with the messages.send method alone generating twice the allowed volume. Caching optimizations cannot mitigate the issue because message sending cannot be cached. Developers view the change as an effort to eliminate alternative clients rather than a genuine monetization strategy. Users expressed disappointment, praising the app's long-term support and criticizing the official VK client for excessive features and advertising.

AntiMalwareOther

Russian AI Research Ranks High in Global Science but Struggles with Commercialization

Russia has secured third place among BRICS nations and twentieth worldwide in the number of scientific papers presented at ten leading international conferences on machine learning and artificial intelligence. According to a study by the Scientometric Center of HSE University, Russian organizations contributed 560 papers between 2020 and 2025 that received over 12,300 citations. The average international citation rate reached 3.59, surpassing India despite fewer total publications. Russian strengths are most evident in the mathematics of machine learning, optimization, and formal concept analysis, with notable results also in computer vision and speech technologies. More than 40 percent of domestic publications involve business participation, led by Yandex among companies, HSE University and Skoltech among universities, and AIRI among non-profit organizations. Significant barriers remain, including shortages of computing power, limited access to high-quality data, and weak transfer of research into commercial products, particularly in natural language processing, AI agents, and infrastructure technologies. The Ministry of Digital Development has announced plans to stimulate demand for domestic AI solutions, expand computing infrastructure, improve regulation, and accelerate the implementation of scientific developments.