Version: v26.09

Version Overview ​

Component Change Description ​

Table 1 openFuyao v26.09 Component Change Description

Component Name (Helm Chart Package Name)Component Change Type
(New, Enhancement, Fix, Change, Deprecated, Removed)
Change DescriptionSIG
hermes-routerEnhancementAdded cacheIndexer.blockSize configuration parameter for unified configuration of cache block size delivery by policies such as kv-cache-aware. Fixed prefix cache statistics issue: only complete Blocks are counted to avoid scoring deviation caused by semantic inconsistency with cache-indexer.sig-ai-inference
cache-indexerEnhancementAdded KV Cache awareness capability for Decode nodes, incorporating decode Pods into the L1 index discovery chain; L1 kvevent subscription is compatible with vLLM 0.28's map encoding events to avoid cache index missing.sig-ai-inference
infernexEnhancement1. infernex enhancement. Added vLLM engine PodMonitor metric collection, collecting engine metrics of prefill/decode inference instances through Prometheus. Supports traceparent request header pass-through, enabling eagle-eye distributed tracing; monitoring and tracing components can be toggled independently.
2. infernex-bridge enhancement. Aligned with KServe scheduler contract, removed scheduler configuration webhook patch, streamlined KServe deployment example and compatible with restricted PSA policy.
3. infernex-checker enhancement. Added optional network connectivity check switch; msprof benchmark testing adds HCCL communication data collection to assist slow card localization.
sig-ai-inference
opensandboxNewAdded opensandbox component, compatible with opensandbox API, providing fluxsandbox provider routing requests to the high-performance sandbox scheduling engine.sig-agent-sandbox
fluxsandboxNewAdded fluxsandbox component, a high-performance Kubernetes sandbox scheduling engine that serves as the runtime backend of OpenSandbox, providing low-latency, high-throughput sandbox lifecycle management capabilities for AI Agent workloads.
fluxSandbox does not directly expose APIs to users; instead, it is accessed through the flux runtime of the OpenSandbox Server: users create sandboxes using the OpenSandbox SDK, and the Server forwards them via gRPC to the fluxSandbox controller to complete scheduling and lifecycle management.
sig-agent-sandbox
npu-dra-pluginEnhancementSupports multiple scheduling methods by resource capacity (core, video memory), by device ID, and by device type; simultaneously supports multiple scheduling modes including full-card, hard partitioning, and soft partitioning; supports intra-node topology-affinity scheduling to improve inference performance.sig-orchestration-engine
eagle-eyeEnhancementAdded non-intrusive distributed tracing capability for vLLM-Ascend; fixed the issue that chart unconditionally creates a namespace, now only creates it when enabled.sig-ai-inference
kubevirtEnhancementAdded support for kubevirt advanced capabilities in Kunpeng environment, including CDI integration and memorydump capability, as well as custom parameter installation and deployment capability.sig-orchestration-engine
ubs-k8s-enableEnhancementFixed conflict issues in concurrent update and cleanup scenarios; upgraded and adapted to underlying UBS-related SDK versions, enhancing availability; supports observability of super pods, reporting metrics such as super pod topology, memory borrowing, memory sharing, and URMA devices; supports topology-aware scheduling of super pods.sig-orchestration-engine
superpod-exporterNewAdded superpod observability, supporting observability capabilities of super pods, reporting metrics such as super pod topology, memory borrowing, memory sharing, and URMA devices.sig-orchestration-engine
compliance-operatorEnhancementAdded security hardening and rollback tool based on Compliance Operator, which directly hardens FAIL items in scan results that have no direct impact on the cluster, helping users perform security remediation on demand after installing Kubernetes using the installation and deployment tool.sig-installation
cluster-api-provider-bkeEnhancement1. Supports immutable KubeOS for worker nodes during installation.
2. Refactored DAG upgrade framework, supporting both Yaml and Helm component upgrade access.
3. Completed a series of hardening including concurrency safety, installation state normalization, configuration documentation consolidation, log observability, performance, and exception scenarios, improving installation and deployment performance and reliability.
sig-installation
bke-manifestsEnhancementSupports version component yaml change adaptation.sig-installation
bkeadmEnhancement1. Supports immutable KubeOS for worker nodes during installation.
2. bkeadm tool hardening, configuration documentation consolidation, log observability hardening, improving installation and deployment reliability and usability.
3. Unified pre-check capability, supporting bke preflight to guide node initialization and pre-check validation for cluster creation.
sig-installation
release-imageEnhancementModified the componentversion field to support Yaml and Helm upgrades, adding version content.sig-installation
upgrade-pathEnhancementAdded version upgrade paths, supporting automatic orchestration, validation, and execution of upgrades based on the component manifest and dependency relationships of community release packages.sig-installation
weight-dispatcherEnhancementRefactored RDMA data transmission code implementation, optimizing RDMA link data transmission performance.sig-ai-inference
npu-operatorEnhancementAdded one-click installation, management, and uninstallation capabilities for the VNPU and NPU DRA plugin components, enabling the cluster to access different resource request and partitioning capabilities.sig-orchestration-engine

Interface Change Description ​

None

Version Feature Description ​

The main features of openFuyao v26.09 are shown in Table 2 and Table 3. For detailed information about the features, please refer to Installation and Deployment, Management Plane Operations, and Feature User Guide.

Table 2 Container Platform

Feature NameBrief IntroductionFeature DescriptionSIG
Installation and DeploymentInstallation and DeploymentAn installation and deployment tool compatible with the standard Cluster-API, supporting one-click installation of business clusters. The management cluster provides multi-scenario interactive business cluster lifecycle management capabilities on a unified management plane, including single/multi-node installation (including high availability), online/offline installation, cluster scaling, and Kubernetes in-place upgrades.sig-installation
Container OrchestrationContainer OrchestrationProvides openFuyao Kubernetes, compatible with K8s 1.34, offering enhanced features such as high-density deployment, startup acceleration, log enhancement, and certificate management enhancement.sig-orchestration-engine
Management PlaneManagement PlaneProvides an out-of-the-box console, supporting application management, application marketplace, extension component management, resource management, repository management, monitoring, alerting, user management, command-line interaction, and other functions.sig-container-platform

Table 3 Independent Components

Feature NameBrief IntroductionFeature DescriptionSIG
infernexAI Inference Acceleration SuiteInferNex is an end-to-end integrated deployment solution for AI inference in cloud-native scenarios. It seamlessly integrates core components such as open-source gateways, intelligent routing, high-performance inference backends, KVCache index management, scaling decision frameworks, and inference observability systems through Helm Chart, providing a complete acceleration chain from request access, dynamic routing, inference execution to resource management and monitoring. It aims to improve inference throughput and reduce inference latency, achieving a one-stop efficient AI service deployment and usage experience.sig-ai-inference
eagle-eyeAI Inference ObservabilityEagle Eye is an observability system for AI inference scenarios, implementing full-link metric collection, near-real-time transmission, and intelligent diagnosis from AI gateways, inference engines, and Mooncake to infrastructure (Ray, K8s, hardware). This system integrates Prometheus's periodic metric collection with the low-latency push mechanism of distributed message queue systems, supporting both trend analysis for scaling decisions and meeting the need for second-level data updates of modules with high timeliness requirements (such as intelligent routing). Through an independent hardware health diagnosis module, it achieves continuous monitoring and anomaly identification of underlying metrics such as NPU/GPU, temperature, power consumption, and error codes, building a closed-loop monitoring capability of "collection—transmission—diagnosis—evaluation", providing solid data support for the stability, performance optimization, and resource scheduling of AI inference systems.sig-ai-inference
hermes-routerAI Inference Intelligent RoutingHermes-router is an AI inference intelligent routing component based on the K8s GIE framework, providing multiple routing strategies such as KVCache aware and latency prediction. It is used to receive inference requests and forward them to the optimal inference service backend, helping users improve AI inference performance, cluster resource utilization, and service stability in various scenarios.sig-ai-inference
cache-indexerAI Inference KVCache Index Managementcache-indexer is a global KVCache index management component for AI inference scenarios, uniformly maintaining the two-layer KVCache view of local HBM and memory of inference instances, supporting intelligent routing to query the KVCache hit rate of candidate inference instances for more accurate request-level scheduling.sig-ai-inference
pd-orchestratorElastic Scheduling Component SetAn elastic scheduling component set for Kubernetes, containing 3 core components: Elastic Scaler provides general scaling decision capabilities, ResourceScalingGroup provides workload group scaling and resource orchestration capabilities, and Tidal provides time-rule-based tidal scaling capabilities.sig-ai-inference
weight-dispatcherModel Weight DistributionWeight-Dispatcher is a model weight preheating and distribution component for AI inference scenarios, used to distribute model weights from source nodes or external model repositories to the local cache directory of target inference nodes before inference instances start, reducing the waiting time and bandwidth pressure caused by repeatedly downloading large model weights during the instance startup phase.sig-ai-inference
kae-operatorKAE Access EnablementImplements minute-level automated management capabilities for Kunpeng KAE hardware, including KAE hardware feature discovery, and automated management and installation capabilities for components such as drivers, firmware, and hardware device plugins. KAE can be deployed and made available within five minutes.sig-orchestration-engine
logging-packageLogging ComponentAs an extension component, logging can view the log information of Pods and containers, and supports configuring log collection sources, collection tasks, and alert rules. The openFuyao logging system can effectively improve the platform's efficiency in identifying errors and enhance the monitoring capability of the entire platform.sig-container-platform
colocation-packageOnline/Offline Colocation Scheduling EnhancementSupports mixed deployment of online/offline workloads, ensuring scheduling of online workloads during peak usage periods and suppression of offline workloads, while enabling offline workloads to use oversold resources during low periods of online workloads to improve cluster resource utilization, with utilization increased by 30%~50%, no significant QoS impact, and jitter below 5%.sig-orchestration-engine
many-core-orchestratorMany-Core Colocation SchedulingFor many-core colocation clusters, provides host interference metric collection, node interference analysis, Volcano interference-aware scheduling, and optional Kata Containers VM-level isolation capabilities. The goal is to reduce the long-tail latency jitter of online workloads in I/O, LLC cache, and memory pressure colocation scenarios, and improve the secure colocation density of nodes.sig-orchestration-engine
matrixagentMemory Borrowing Node-Side Functionmatrixagent periodically collects node memory usage and reports it to matrixcontroller, while listening for escape decisions or memory borrowing, and calls VirtAgent to perform memory borrowing, balancing cluster memory usage and improving resource utilization.sig-ub-enable
matrixcontrollerMemory Borrowing Central-Side Functionmatrixcontroller listens to data reported by matrixagent, calculates memory watermarks, and automatically initiates memory borrowing requests based on thresholds, balancing cluster memory usage and improving resource utilization.sig-ub-enable
matrixshmShared Memory CSIProvides a memory sharing solution based on CR declarative isolation for Pods, supporting only full mapping and unmapping. Capacity is constrained by two factors: configured block size and NUMA lending ratio, and Pod scheduling is controlled by node label sharing domains. Its lifecycle management is strongly bound to Pod exit, preventing resource leakage through multi-layer mechanisms such as shm-agent listening, reference counting, operation failure rollback, and startup state recovery, and ensuring security and high availability through path validation, strict permission control, atomic write cache, and task queue overload protection.sig-ub-enable
superpod-exporterSuperpod ObservabilitySupports observability capabilities of super pods, reporting metrics such as super pod topology, memory borrowing, memory sharing, and URMA devices.sig-orchestration-engine
monitoring-dashboardCustom Monitoring DashboardUsed to present extensible metrics, meeting enterprises' needs for highly customizable and extensible monitoring systems.sig-container-platform
multi-cluster-serviceMulti-Cluster ManagementProvides a cluster list interface and full lifecycle management functionality, supporting cluster onboarding, scaling, configuration updates, and secure destruction. Users can achieve quick access through cluster credentials and flexibly edit cluster labels through label management. This component implements cross-cluster secure access, facilitating efficient and unified management of multiple Kubernetes clusters in hybrid environments.sig-container-platform
npu-operatorNPU Access EnablementNPU Operator uses the Operator Framework in Kubernetes to automatically manage all Ascend drivers and firmware required to configure Ascend devices, enabling the full-process operation of cluster construction, supporting functions such as cluster job scheduling, O&M monitoring, and fault recovery. By installing the corresponding components, NPU resource management, optimized scheduling of workloads, and containerized support for training and inference tasks can be achieved, enabling AI jobs to be deployed and run on NPU devices in container form.sig-orchestration-engine
numa-affinity-packageNUMA Affinity SchedulingImplements hardware NUMA topology awareness at both cluster and node levels, and performs NUMA-affinity scheduling for applications based on NUMA affinity to improve application performance.sig-container-platform
ray-packageopenFuyao/rayDerived from the "Parallel Universe" project of ray-project/ray, a unified distributed computing framework for extending AI and Python applications, dedicated to providing stable, reliable, and domestic-hardware-affinity LTS (Long-Term Support) versions for Ray users and operators.sig-distributed-framework
compliance-operatorCompliance ScanningCompliance Operator is an automated compliance scanning tool based on the Kubernetes Operator pattern, supporting security compliance scanning of Kubernetes clusters using CIS Benchmark and DISA STIG standards.sig-installation
opensandboxSandbox AccessCompatible with opensandbox API, providing fluxsandbox provider routing requests to the high-performance sandbox scheduling engine.sig-agent-sandbox
fluxsandboxHigh-Performance Sandbox Scheduling EngineAdded fluxsandbox component, a high-performance Kubernetes sandbox scheduling engine that serves as the runtime backend of OpenSandbox, providing low-latency, high-throughput sandbox lifecycle management capabilities for AI Agent workloads.
fluxSandbox does not directly expose APIs to users; instead, it is accessed through the flux runtime of the OpenSandbox Server: users create sandboxes using the OpenSandbox SDK, and the Server forwards them via gRPC to the fluxSandbox controller to complete scheduling and lifecycle management.
sig-agent-sandbox
vNPUNPU VirtualizationImplements NPU soft and hard partitioning through the Ascend docker runtime, supporting 310P, 910B, and 910C card models, improving cluster computing resource utilization.sig-orchestration-engine
npu-dra-pluginAscend NPU-Affinity DRA PluginBased on the Kubernetes DRA mechanism, supports 310P, 910B, and 910C card models; supports multiple scheduling methods by resource capacity (core, video memory), by device ID, and by device type; simultaneously supports multiple scheduling modes including full-card, hard partitioning, and soft partitioning; supports intra-node topology-affinity scheduling to improve inference performance.sig-orchestration-engine
kubevirtContainer-VM Co-Management ComponentVM-container co-management capability of kubevirt in the Kunpeng environment, supporting full VM lifecycle management, supporting macvlan, SR-IOV, SR-IOV live migration, NIC hot-plug, VM live migration, and memory export. Verified that kubevirt currently does not support Kunpeng UEFI secure boot. Supports custom parameter installation and deployment.sig-orchestration-engine