AntiMalware•August 20, 2026•🇷🇺Translated from Russian

Rubytech Tests Russian LLM Cotype 3 on Chinese GPUs Matching NVIDIA H100 Performance

Rubytech has completed testing of the Russian large language model Cotype 3 from MWS AI on Chinese graphics processing units, demonstrating performance comparable to systems based on NVIDIA H100 accelerators.

The experiments were conducted on the Skala^r software-hardware complex designed for generative artificial intelligence workloads. In the test configuration, eight Chinese GPUs were deployed. When processing contexts of 27,000 tokens, the system achieved an average first-token generation time of approximately 8 seconds and inter-token latency ranging from 111 milliseconds.

According to Rubytech, these metrics are sufficient to support enterprise-grade AI services. Engineers performed extensive adaptations of drivers, the software stack, and orchestration mechanisms. In selected inference scenarios, these optimizations delivered speedups of 2 to 2.2 times relative to the default Chinese GPU environment.

The containerized architecture is expected to simplify model deployment and accelerate the transition from laboratory testing to production environments. The primary conclusion drawn from the trials is that corporate AI infrastructure can be constructed without exclusive dependence on NVIDIA accelerators.

This finding holds particular significance for government agencies, state-participated enterprises, and organizations operating critical information infrastructure that require localized model hosting, controlled infrastructure, and the ability to scale without supply-chain uncertainties associated with NVIDIA H100 GPUs.

Rubytech also indicated that Chinese GPUs can reduce total cost of ownership in specific use cases. The company stressed that the results represent an additional viable option rather than a full departure from NVIDIA technology. Future plans include testing newer generations of Chinese accelerators and expanding the range of supported Russian models and enterprise AI services.

Related articles

Securitylab•Other

Pivoting in Legacy Hell: Navigating MIPS Servers, BusyBox, and 2014 Kernels During Internal Network Assessments

The article details a complete methodology for pivoting from an initial SSH compromise on an old Debian MIPS server to reach a hidden web admin panel inside a segmented network. It covers environment enumeration with commands like uname -a and ip route, followed by setting up a Chisel-based SOCKS5 proxy when standard SSH dynamic forwarding is disabled by server configuration. Scanning proceeds via proxychains with nmap using -sT, -Pn, and -n flags or by deploying static MIPS binaries directly on the host. Traffic is then routed through Burp Suite chained to the SOCKS proxy for password guessing against the web interface. The piece emphasizes practical constraints such as BusyBox limitations, kernel version incompatibilities with modern binaries, and the need for careful subnet identification. Readers are directed to replicate the full chain in the Forgotten Server task from the free White Hacker Profession course.

Securitylab•Other

Neuromorphic Processors Deliver Reflex-Like Responses for Robots, Drones and Edge Sensors

Neuromorphic chips are optimized for sparse, event-driven data rather than dense matrix operations, making them ideal for always-on peripheral devices that must react instantly while conserving power. The technology pairs naturally with event cameras and temporal sensors in robotics, drones, automotive systems, medical wearables, industrial monitoring and space applications. Platforms such as Intel Loihi 2, BrainChip Akida, SynSense Speck and SpiNNaker2 already demonstrate working prototypes that activate only on meaningful changes in the input stream. Researchers at TU Delft have flown autonomous drones using spiking networks on Loihi, while NASA has tested radiation-tolerant neuromorphic designs for onboard decision making. The approach complements rather than replaces GPUs and NPUs, creating hybrid systems where the neuromorphic layer handles fast reflexes and conventional accelerators manage complex models.

Habr•Other

K2 Cloud Adds Native OVN Traffic Mirroring for NTA/NDR Deployment in Public Cloud

K2 Cloud has implemented traffic mirroring for virtual machine interfaces inside overlay networks built on OVN, solving a long-standing gap that prevented NTA/NDR systems from operating in Russian public clouds. The company contributed the feature upstream, and the patch was accepted into the main OVN codebase. The solution supports source, destination, and mirroring session objects together with optional match/action filters that eliminate duplicate packets, reduce load on sensors, and allow traffic distribution across multiple analyzers. Testing with Positive Technologies PT NAD showed sustained throughput above 3 Gbit/s with approximately 400 000 packets per second and minimal packet loss. The feature works inside a single availability zone; multi-AZ deployments require one PT NAD sensor per zone. Management is available through the web console, an EC2-compatible API, and the official Terraform provider.

Securitylab•Other

Bitrix24 Introduces Cowork/Code AI Agent for Corporate Task Automation and App Building

Bitrix24 has launched Cowork/Code, an AI application that combines file management, company data access, and application development inside a controlled corporate environment. The tool features an AI agent capable of executing multi-step workflows such as locating records, comparing documents, generating tables, and saving results to shared folders. It operates in two modes: Cowork for one-time tasks like overdue task reports or client preparation, and Code for creating reusable tools such as dashboards or notification bots. A memory technology called Radiant stores context from chats, tasks, meetings, and employee data to deliver more accurate, personalized responses over time. The platform addresses common risks of vibe coding by keeping code, data access, and distribution within the Bitrix24 ecosystem hosted on Russian infrastructure. A free tier provides limited usage, with paid plans required for sustained team operation.