securitylab_n•July 16, 2026•🇷🇺Translated from Russian

Former OpenAI CTO Mira Murati Launches Thinking Machines' Inkling: Open-Weights Multimodal AI Model with 975 Billion Parameters and Self-Training Demo

Releasing yet another powerful language model is no longer enough to stand out. Thinking Machines, the startup founded by former OpenAI technical director Mira Murati, has chosen a different approach to attract developers. The company unveiled Inkling, its first open-weights model that can be freely used and adapted for specific tasks.

Inkling is a multimodal model capable of processing text, images, and audio without separate processing modules for each modality. It is built on a mixture-of-experts architecture containing 975 billion parameters, of which only 41 billion are active at any given time. This design significantly reduces computational costs. The model supports a maximum context length of 1 million tokens and was trained on a dataset of 45 trillion tokens that included text, images, audio, and video. Alongside the main model, the company released a preliminary Inkling Small version with 12 billion active parameters optimized for lower-cost and faster deployment.

One of Inkling’s standout features is the ability to regulate reasoning depth. Developers can choose how many computational resources the model should allocate to solving a problem. For simple queries the system responds faster and consumes fewer tokens, while for complex tasks it can increase internal computation. According to the company, in certain programming scenarios Inkling uses approximately three times fewer tokens than several other open models while maintaining comparable quality.

Thinking Machines acknowledges that Inkling does not yet aim to be the strongest model on the market. Closed models from OpenAI, Anthropic, and Google continue to lead most comprehensive benchmarks, while Chinese models remain ahead in certain disciplines. Instead of competing for top rankings, the company focused on creating a versatile foundation for subsequent training and customization for enterprise use cases.

To support further development, Thinking Machines offers its own Tinker platform. Through this platform, developers can fine-tune Inkling on their own data without building complex infrastructure themselves. To demonstrate the platform’s capabilities, the company conducted an unusual experiment: Inkling was tasked with training itself. The model autonomously created a training task, executed the process via Tinker, evaluated the results, and switched to the updated version. In the demonstration, it was given the unusual objective of learning to answer questions while completely avoiding the use of one letter of the English alphabet.

The company has also made the model weights available on Hugging Face and added support for popular inference frameworks including Transformers, vLLM, SGLang, and llama.cpp. Inkling is designed not only for cloud services but also for continued training, creation of specialized assistants, and development of autonomous agents capable of executing complex sequences of actions.

The release of Inkling represents an important development for the Western open-weights AI community. After Meta reduced its activity in this area and many organizations began turning to Chinese models, developers now have another major Western-origin solution that can be freely run, modified, and adapted to their own tasks thanks to its open weights.

Related articles

Habr•Other

Why Defending a Company Costs Millions While Attacks Can Succeed for Just Hundreds of Dollars

In the latest episode of Belyaev Podcast, CISO Vyacheslav Kasimov of Tochka Bank and Boris Evdokimov of ASNA pharmacy chain discussed the persistent asymmetry in cybersecurity spending. Attackers increasingly rely on affordable cloud services, automation, and rented infrastructure, while defenders must invest heavily in monitoring, access controls, backups, and skilled teams. The experts stressed that the absence of known breaches does not equal security, as undetected incidents or delayed discovery remain common risks. They advocated shifting from a "no" culture to risk-based decision making that helps business leaders understand potential losses, mitigation costs, and residual risk. The conversation also covered responsible use of AI in SOC operations and the long-term damage caused by loss of customer trust after incidents.

AntiMalware•Other

Beeline Offers One Month Free Access to Six Services for Prepaid Customers

Beeline has launched a promotional campaign allowing home users on prepaid plans to try up to six digital services for free over 30 days. The offer, tied to the operator's second annual Cellular Independence Day, runs from October 2 to October 9 and includes services such as Virtual Assistant PRO, unlimited mobile data, internet sharing without speed reduction, custom network name display, 250 GB of cloud storage, and access to over 650,000 e-books and audiobooks. Each selected service activates its own free period starting from the moment of connection and deactivates automatically afterward. Customers already paying for four or more of the listed services will receive 300 bonus rubles for communication instead. The unlimited data option is unavailable in the Chukotka Autonomous Okrug and Norilsk. Activation is handled exclusively through the Beeline mobile app, and users with existing paid subscriptions to any service cannot activate the free trial version of the same service.

Habr•Other

Enterprise-Grade Web Protection on a Budget: How Cloud WAF Lowers Barriers for SMBs

A new overview from Reg.cloud explains how cloud-based Web Application Firewalls reduce the cost and complexity of protecting websites, APIs, and web applications for small and medium-sized Russian businesses. According to Positive Technologies data cited in the article, 75% of successful web application attacks in 2025 disrupted organizational operations, while 82% of SMBs faced cyber incidents in the past year. The piece details the differences between traditional on-premises WAF deployments and cloud offerings, emphasizing ready-made protection profiles for CMS platforms, SaaS services, and digital agencies. It outlines a three-stage operational model covering preparation, DNS-based traffic redirection, and ongoing policy tuning that can be handled by existing DevOps or development teams without dedicated security staff. The service currently offers a free tier supporting up to three applications at 50 requests per second, along with seven preconfigured security profiles and dual audit/blocking modes. The article concludes by stressing that WAF remains only one layer and must be combined with patching, access controls, and separate DDoS or anti-bot solutions.

AntiMalware•Other

Yandex B2B Tech Integrates Hybrid Full-Text and Vector Search in Single YDB Query

Yandex B2B Tech has added hybrid search to its YDB database, allowing full-text and vector approaches to run together inside one SQL query. The update helps small and medium businesses as well as large corporations locate exact document identifiers while also matching semantic meaning in descriptions, even when wording differs. Full-text search handles precise elements such as policy numbers, codes, and names, whereas vector search identifies conceptual similarity. Results from both methods are merged and ranked within the same transaction, keeping all data inside a single database instance. This removes the need to maintain a separate search engine and vector store or to reconcile information between them. The technology is aimed at chatbots, recommendation systems, and AI assistants that process technical content where both exact codes and human-readable problem descriptions matter equally. Hybrid search is now available in the on-premises YDB 26.3 release and in the cloud-based Managed Service for YDB.