Habr•August 11, 2026•🇷🇺Translated from Russian

Building Secure MLOps Platforms in Air-Gapped Environments for DevOps Engineers

The article presents a practical guide for DevOps engineers tasked with building an MLOps platform inside a fully closed, air-gapped environment. Up to 80 percent of machine-learning projects reportedly never reach production because of missing reproducible processes for data, models, and deployment.

MLOps extends classic DevOps practices to data and models. Data becomes a first-class artifact that must be versioned with DVC and stored in MinIO, while experiments are tracked in MLflow backed by PostgreSQL. A Model Registry with champion and challenger aliases allows safe model rollouts without changing application code.

The platform runs on two physical segments: a Kubernetes cluster handling serving, storage, and monitoring, and a separate GPU server used exclusively for training via Docker containers with NVIDIA runtime. This separation prevents expensive GPU resources from being locked to Kubernetes nodes.

Key components include ArgoCD with Apps-of-Apps pattern, GitLab CI/CD, FastAPI inference service, Celery and Redis for GPU task queuing, JupyterLab for experiments, OpenBao instead of Vault, External Secrets Operator, Trivy and Bandit scanning, and Prometheus plus Grafana monitoring. All Helm charts are vendored inside the Git repository to eliminate external dependencies.

Security measures specific to closed contours cover private Harbor registry, SOPS encryption with age keys delivered via ArgoCD CMP sidecar, wildcard TLS certificates distributed by ClusterExternalSecret, and a Docker Socket Proxy that restricts container operations to the minimum required privileges.

The author deliberately avoids Kubeflow because of its cloud-oriented design and heavy CRDs, and replaces Airflow with the lighter Celery queue. The resulting stack provides a reproducible, auditable foundation that can later be scaled when project volume increases.

Related articles

Habr•Other

Yandex Disk Files Become Partially Inaccessible After Sasovo Data Center Incident

A detailed user report reveals that approximately one percent of files stored on Yandex Disk are currently unavailable for download following reported incidents at the Sasovo data center. The problems manifest in three distinct states: missing thumbnails with downloadable originals, visible thumbnails with inaccessible originals returning 504 Gateway Time-out errors, and cases where both thumbnails and originals fail to load. Technical analysis using curl requests traced the failures to specific storage nodes such as s418klg.storage.yandex.net, indicating that some data shards may reside in affected infrastructure while others remain operational. Yandex support requested original files for diagnosis but closed the ticket without providing an official explanation or confirming data integrity. The author emphasizes that the issue affects files across both the Photos and Files sections and recommends maintaining offline backups due to the lack of guaranteed availability during data center failures.

AntiMalware•Other

Keurig K-Supreme Smart Coffee Maker Generates Nearly 1 TB of Outbound Traffic in Ten Days

A Keurig K-Supreme Smart coffee maker unexpectedly produced around 1008 GB of outgoing traffic over ten days, overwhelming a home UniFi access point while generating only 9.94 GB of inbound data. The device had been placed on a separate network segment, yet the traffic remained largely internal to the home LAN rather than traversing the internet connection. The anomaly was discovered by user Nomad while assisting family members with network maintenance through the UniFi dashboard. After the coffee maker was powered off, the issue could not be reproduced in subsequent testing, and no packet captures were available to determine the content or root cause of the traffic. The model requires internet connectivity for remote control, scheduling, capsule recognition, and automatic reordering of coffee supplies. No similar incidents have been reported by other users, and the manufacturer has not issued any statement regarding the event.

Habr•Other

Hash Functions Part 1: Core Properties, Security Requirements and Practical Applications

The article provides a detailed introduction to hash functions, explaining how they map arbitrary-length input to fixed-length output while satisfying three fundamental security properties. It covers preimage resistance, second preimage resistance, and collision resistance, along with the avalanche effect that makes even minor input changes produce unrecognizable output. The text explains why a 256-bit digest is required to achieve 128-bit collision resistance, referencing the birthday paradox and its implications for MD5 and SHA-1. Practical guidance includes using OpenSSL for hashing, applying hashes in commitment schemes, enforcing subresource integrity on web pages, and securely storing passwords with Argon2 and bcrypt. The post emphasizes that hash functions alone do not guarantee integrity without proper transmission of the digest and announces a follow-up on SHA-2 and SHA-3 internals.

Habr•Other

Digital Twins Enable Pre-Deployment Testing and Post-Change Control in Complex Multi-Vendor Networks

UserGate and Hadal Project experts presented a joint approach at Saint HighLoad++ that combines physical labs, emulation, and simulation into a single lifecycle for validating network changes. The method addresses recurring failures such as IPsec tunnel outages after routine software updates that pass vendor checks yet break branch connectivity. Three complexity sources—multi-vendor environments, historical configuration debt, and continuous dynamic updates—are mitigated by maintaining an always-current network model. Physical laboratories provide hardware-level accuracy for critical devices, while uInfraTwin emulation allows rapid, repeatable testing of configuration scenarios with traffic generators. Simulation tools including Batfish, Hadal, Forward Networks, and IP Fabric deliver end-to-end reachability analysis across tens of thousands of nodes without sending test traffic on production networks. The integrated digital twin continuously updates from live infrastructure, feeds selected segments into safe test environments, and verifies policy compliance after deployment.