September 2026 AI Model Rankings: Fable 5.1 Tops Intelligence Index as Competition Tightens Across GPT-5.6 Sol, Grok 4.6 and Muse Spark 1.3
The start of September 2026 proved to be a rare moment when the roster of the strongest language models had to be almost completely rewritten. Anthropic introduced Fable 5.1 and the restricted Mythos 5.1, Meta updated Muse Spark to version 1.3, Google launched Gemini 3.8 Flash, and Alibaba refreshed Qwen3.8-Max. Summer releases such as GPT-5.6 Sol, Grok 4.6, Kimi K3, GLM-5.3 and DeepSeek V4 Pro stayed in the race and continue to compete with the newcomers.
Compiling a conventional ranking from smartest to least capable has become difficult. Contemporary models run in several reasoning-depth modes, and the gap between low, high and max settings can exceed ten points on a single test. Speed, token consumption, long-context pricing, image support, tool-use quality and the ability to execute long autonomous action sequences now vary widely.
It is therefore more useful to treat the market as several overlapping races. Fable 5.1 currently claims the highest quality for complex reasoning. GPT-5.6 Sol and Grok 4.6 deliver comparable performance at lower cost. Muse Spark 1.3 unexpectedly entered the top tier on capability-to-cost ratio. Gemini 3.8 Flash bets on speed and multimodality. Kimi, GLM, Qwen and DeepSeek have nearly erased the former boundary between Chinese and American models.
How to Compare Models Without Misleading Yourself
The Artificial Analysis Intelligence Index serves as a convenient common scale. This independent metric combines tests on knowledge, programming, scientific reasoning, tool use and long-horizon agent tasks. The methodology is more reliable than vendor tables because models are evaluated under comparable conditions.
Scores cannot be read as an intelligence quotient. A result of 66 versus 61 does not mean the first model is eight percent smarter. The index reflects average performance on a specific test suite. On Russian-language article writing, accounting-document analysis, website layout or debugging enterprise applications the ordering may differ.
Reasoning modes alter the picture even more. GPT-5.6 Sol scores around 61 in max mode, 59 in xhigh and 57 in high. Fable 5.1 reaches 66 only at maximum depth. Gemini 3.8 Flash scores 59 in high, 57 in medium and 52 in low. Comparing Gemini against Claude without specifying the mode is about as useful as comparing cars without stating engine power.
Market Leaders at the Start of September 2026
The current top tier is unusually dense. After Fable 5.1 the gap between several flagships fits within a few points. Expensive models do not always complete a given task more cheaply; some systems generate far more intermediate text, reason longer or invoke external tools more frequently.
- Claude Fable 5.1, max – 66 points, 1 M context, $10/$50 per million tokens, available
- Claude Opus 5, max – 63 points, 1 M context, $5/$25, available
- Muse Spark 1.3, max – 62 points, 1 M context, price not announced, limited access
- GPT-5.6 Sol, max – 61 points, ~1.05 M context, $4/$20, available
- Grok 4.6, high – 61 points, 500 k context, $2/$6, available
- Muse Spark 1.3, xhigh – 61 points, 1 M context, $1.25/$4.25, available
- Kimi K3, max – 60 points, ~1 M context, $3/$15, open weights
- GLM-5.3, max – 60 points, up to 1 M context, $1.40/$4.40, open weights
- Gemini 3.8 Flash, high – 59 points, ~1.05 M context, $0.75/$3.75, available
- Qwen3.8-Max – 58 points (previous snapshot), 1 M context, ~$2/$6, updated to 0902
- DeepSeek V4 Pro 0813, max – 53 points, 1 M context, $0.66–1.32/$1.98–3.96, open weights
Additional caveats apply. Independent tests of the updated Qwen3.8-Max-0902 have not yet produced a comparable final score, so the 58-point figure refers to the prior version. Muse Spark 1.3 maximum mode remains in limited testing; the public xhigh mode scores 61. For Grok 4.6 the best independent result occurs in high mode, while xhigh unexpectedly scores slightly lower.
Pricing also requires notes. The Gemini 3.8 Flash tariff is temporarily reduced until the end of 2026. GPT-5.6 Sol is sold at a temporarily lowered price. DeepSeek pricing changes between peak and off-peak hours. Grok 4.6 doubles in price for requests longer than 200 k tokens. A simple API-price column therefore hides half of the real economics.
Related articles
InfoWatch Acquires Web Control DC Team and Rebrands sPACE PAM as InfoWatch Privilege Control
InfoWatch has expanded its product portfolio by incorporating the Web Control DC development team and rebranding its flagship sPACE PAM solution. The new product, InfoWatch Privilege Control, is designed to manage and monitor privileged accounts belonging to system administrators, contractors, external specialists, and business users. These accounts provide access to servers, databases, network equipment, and critical applications, making them high-value targets for attackers. According to InfoWatch data, approximately 40% of critical information security incidents in Russia in 2025 were linked to the leakage or misuse of privileged credentials, while another 30% of confirmed cyberattacks occurred through compromised IT contractors. The solution enables time-limited privilege issuance, connection management, and detailed activity logging to prevent unauthorized actions. The original sPACE PAM product remains listed in the Russian software registry and holds FSTEC Russia certification at the fourth trust level, with compatibility for Astra Linux, Alt, and RED OS operating systems.
Secure Custom Domain Setup for Client Status Pages Using CNAME, Certbot and Go Instead of ACME On-Demand
A detailed case study describes how a monitoring service implemented white-label status pages on customer domains without relying on ACME on-demand certificate issuance. The approach uses pre-validated domains stored in a database, background DNS checks via CNAME or A-record matching, and a root helper script running on a systemd timer to handle certbot issuance and nginx configuration. Key design choices separate privileges so the Go application never touches certificates or nginx directly, while loopback endpoints are protected against proxy header spoofing. The solution explicitly addresses risks such as DoS through malicious Host headers exhausting Let's Encrypt rate limits, private key exposure, and first-visitor latency during TLS handshakes. Hysteresis in DNS status prevents temporary resolution glitches from disabling active customer pages. The entire implementation stays under a few hundred lines of code and runs reliably on a single server with nginx in front of the Go application.
Durev VPN Accused of Plagiarizing Independent Researchers' Articles for Commercial YouTube Promotion
Independent Russian cybersecurity researcher zarazaexe has publicly accused the commercial VPN service Durev VPN of systematically copying technical articles about Russian internet censorship systems, messenger analysis, and government-issued certificates. The team behind Durev VPN allegedly rewrote the original research into YouTube video scripts, replaced first-person statements with references to their own specialists, and removed all attribution to the source author or project. One video titled СРОЧНО УДАЛИ СЕРТИФИКАТ МИНЦИФРЫ reportedly gained 795,000 views in two days, while the service's Telegram bot shows 260,000 users and charges a minimum subscription of 313 rubles per month. Specific copied elements include scanning results of 46 million Russian IP addresses that yielded 63,000 entries in TSPU whitelists, architectural descriptions of default-deny and fail-closed policies, ECH handling by DPI systems, and detailed reverse-engineering findings from the MAX messenger client. The researcher documented the reuse of his earlier errors about UDP blocking inside whitelists and identical analogies comparing root certificates to apartment keys. When confronted, the Durev VPN representative initially demanded patents, invoked fair use, and later claimed the matter would be reviewed by their lawyer within a week. The researcher published the full chat logs and removed the disputed segment from one video after the complaint.
Russia's Rassvet LEO Satellite System Enters Real-World Testing Phase
The Russian low-orbit satellite constellation Rassvet has moved from laboratory development into active field trials with real consumers across multiple regions. Bureau 1440, part of IKS Holding, has begun delivering satellite internet services and installed the first user terminal on a long-distance Russian Railways train. CEO Alexey Shelobkov stated that the project advanced to service validation in 2026, ahead of schedule, and will now test performance over large territories and in motion. The initiative aims to provide high-speed connectivity to remote areas, reduce the digital divide, and create an independent Russian hybrid communications system. Shelobkov emphasized that Rassvet is not merely a response to Starlink but a strategic necessity for sovereign infrastructure free from foreign hardware and policy dependencies. Roscosmos is simultaneously increasing engagement with private companies to commercialize space activities.