Doubling GPT-3 Inference Speed and Eliminating Silicon Furnaces: Vertical Memory Architectures V-Die and MOSAIC Revolutionize HBM Cooling and Bandwidth for AI Accelerators
Future AI accelerators may soon feature memory chips literally standing on their edges. Researchers from South Korea and Japan have proposed rotating DRAM dies vertically to boost both the speed and capacity of HBM without turning multilayer stacks into poorly cooled silicon furnaces.
Two new architectures, V-Die and MOSAIC, were presented in June 2026 at the IEEE Symposium on VLSI Circuits. Both teams abandoned the conventional approach of stacking dynamic memory dies horizontally on top of a base die and connecting them through TSV (through-silicon via) channels. Instead, the engineers assemble the package first and then rotate the entire structure so that individual dies stand vertically, functioning like the fins of a radiator.
Current HBM places multiple memory dies above a base die and links the layers with short, wide TSV buses that deliver several terabytes per second of bandwidth, making it the foundation of powerful AI accelerators. However, increasing the number of layers makes heat removal increasingly difficult. Lower dies become extremely hot, and heat must travel through silicon, solder, insulating materials, and additional package layers. The TSVs themselves also consume valuable die area that could otherwise be used for memory cells.
The Korean V-Die architecture removes TSVs from the memory dies altogether. Each vertical die receives its own input/output lines along the bottom edge and connects directly to the substrate. According to the researchers’ calculations, the design provides four times more connections than HBM4 and cuts data read time by 37 percent. Microfluidic channels running between the dies maintain temperatures around 45 °C, whereas dense conventional HBM packages can exceed 80 °C under heavy load.
Simulations of a 16-die system demonstrated significant performance gains. When processing GPT-3-scale workloads, the V-Die architecture achieved 540 tokens per second compared with 296 tokens per second for an equivalent-capacity HBM4 configuration. First-token latency dropped by 32 percent, or roughly 24 milliseconds. These results are currently simulation-based; the research team is now preparing physical prototypes to validate thermal and electrical characteristics.
The Japanese MOSAIC project addresses a different challenge of vertical assembly: even small thickness variations among dozens of dies can misalign contact pads. Engineers from the University of Tokyo, Tohoku University, and RIKEN proposed transmitting data without direct metal contact. Miniature coils on the dies and substrate exchange signals through electromagnetic induction, eliminating the need for perfect physical alignment during assembly.
The experimental MOSAIC interface reached speeds of up to 4 Gbit/s per channel. Researchers expect to double the memory volume of HBM4 by placing the vertical block directly above the GPU while increasing maximum temperature by only about one degree. One configuration accommodates 98 dies and 294 GB of memory; further thinning of the dies could theoretically push capacity to 882 GB.
Neither V-Die nor MOSAIC is yet ready to replace production HBM. The Korean architecture exists primarily as calculations, while the Japanese prototype must still demonstrate acceptable cost, reliability, and high manufacturing yield. Nevertheless, both developments outline a promising route around the thermal barrier that increasingly constrains memory scaling for AI accelerators. Instead of endlessly stacking dies higher, engineers are proposing to rotate the entire structure and turn the memory dies themselves into an integrated heat-dissipation system.
Related articles
How the Lorenz SZ 42 Teleprinter Cipher Machine Worked: Nazi High Command Encryption and Its 1941 Breakthrough
The Lorenz SZ 42 was a teleprinter attachment used by the German high command for encrypting top-secret communications during World War II, operating on the Vernam cipher principle with twelve wheels generating a keystream. Unlike the portable Enigma, Lorenz integrated directly between teletypes for automatic five-bit ITA2 encryption. British interceptors at Knockholt first encountered its signals in 1940, later named Tunny. A critical operator error on 30 August 1941 allowed cryptanalysts at Bletchley Park to deduce the machine's structure. This led to the development of the Colossus computer in 1944 for automated decryption. The article details the χ, ψ, and μ wheel groups, the stuttering psi mechanism, and Python implementations of ITA2 encoding and XOR operations.
Inside the Fortress: Why Perimeter Security Tools Fall Short and How Microsegmentation Protects Networks Internally
Companies invest heavily in perimeter defenses such as firewalls and intrusion detection systems, yet these measures no longer guarantee safety as attackers increasingly operate from within networks. Traditional L2 domains leave virtual machines unisolated, enabling traffic interception, lateral movement, and malware spread similar to an apartment building with poor soundproofing. Microsegmentation powered by SDN divides VLANs into isolated microsegments down to individual VM ports, enforcing granular policies based on ports, IP addresses, and protocols. This approach implements Zero Trust by placing virtual packet filters directly at VM network interfaces on the hypervisor, independent of guest OS actions. Performance remains high because filtering runs on powerful virtualization servers, and scaling occurs naturally as additional hypervisors absorb new workloads without extra configuration. A real-world case from the oil and gas sector shows one customer creating up to 5,000 new microsegmentation rules per week via open REST API. The technology complements rather than replaces perimeter firewalls, delivering both strict internal controls and operational agility.
Good Bear 1.0 Released: Firefox-Based Browser with Isolated Russian PKI Trust Container
Good Bear 1.0 is a Russian-language browser built on Firefox 156.0 that provides an isolated container for handling Russian PKI certificates without mixing trust contexts or user data with the standard browsing session. The release includes .deb packages for Ubuntu 24.04 LTS amd64 and Windows x64 installers, using Mozilla Public License 2.0 and reproducible build processes from pinned Firefox sources. Instead of globally importing root certificates, the browser performs secondary chain validation only inside a dedicated userContextId container with strict OriginAttributes isolation for caches, storage, and connections. Password autofill and sensitive session data are disabled in the container when separation cannot be guaranteed, and POST requests trigger explicit user choice before reopening in the isolated context. The interface shows both a persistent container marker and a separate RU indicator only when Russian PKI is actively used, along with detailed security panels explaining the trust source. Updates, crash reporting, and automatic MAR mechanisms are intentionally omitted to avoid creating unverified trust chains for the distribution itself.
Survey of 254 Russian Domains Shows 89% DMARC Adoption but Highlights Gaps in Reporting and Subdomain Policies
A manual review of public DNS records across 254 prominent Russian domains from 17 sectors found strong baseline adoption of email authentication mechanisms. MX records appeared in 96.1% of domains, SPF in 93.7%, DMARC in 89.0%, and DKIM records via common selectors in 62.2%. Among domains with DMARC, 40.7% published a reject policy and 42.9% used quarantine, while 16.4% remained at none. Notably, 19% of DMARC-enabled domains lacked any rua address for aggregate reports, including 33 domains enforcing reject or quarantine. The study also identified cases of inconsistent policies between parent domains and subdomains, as well as SPF records ending in ~all paired with strict DMARC settings. Researchers emphasized that DNS data alone cannot confirm actual mail flow alignment or report consumption.