Version: v26.09

Frequently Asked Questions (FAQ) ​

This document aims to provide common installation and deployment issues along with available solutions.

1. The install-cni container of the calico component keeps restarting, with calico-node stuck in the init phase (1/3) and in a CrashLoopbackOff state ​

  • Cause

    You can run ip route on the node to view routing information. This issue is usually caused by one of the following two situations.

    1. The node is missing a default route.
    2. The node has multiple default routes, and the route assigned for cluster deployment is not the one with the highest priority (a smaller metric value indicates a higher priority).
  • Solution

    For cause 1, add a default route by running the following command.

    bash
    ip route add default via 192.168.100.1 dev eth0 proto static metric 100

    For cause 2, delete the redundant default route by running the following command.

    bash
    ip route del default via 192.168.90.1 dev eth1
  • Extended Discussion

    Calico requires that each node has a default route with the highest priority, essentially to ensure that all traffic to and from the cluster can be correctly forwarded at the node, which acts as a Layer 3 router.

    The default route is typically configured under the /etc/sysconfig/network-scripts directory, which defines and configures persistent configuration files for system network interfaces (network cards). When the system restarts or the network service restarts, the system reads these files to set the IP address, gateway, DNS, and other information.

2. calico-node restarts frequently, and the container falls into a "start - probe failure - restart" infinite loop ​

  • Cause

    This issue is usually caused by setting initialDelaySeconds too short. The Pod has not yet completed initialization, connection establishment, configuration loading, and other startup operations when kubelet starts probing the health status of calico-node.

  • Solution

    Set an appropriate delay time to tell kubelet to wait for a period of time after the container starts before performing the first probe. Run the following command to configure it.

    bash
    # Edit the K8s resource YAML, find the readinessProbe section, and add the initialDelaySeconds field
    kubectl edit ds -n kube-system calico-node

    The recommended initialDelaySeconds settings are as follows; adjust according to actual scenarios.

    • Small clusters (fewer than 10 nodes): set to 30, which is sufficient for calico to complete basic initialization.
    • Large clusters (more than 10 nodes): set to 60, as more nodes and intense resource competition require a longer startup time.

3. calico-node is in a 0/1 Running state ​

  • Detailed Symptoms

    View the detailed information as follows.

    • Node status is NotReady: This is the most obvious symptom. Through the kubectl get nodes command, you will find that the node status is not Ready, but NotReady.
    • calico-node Pod probe failure: When viewing the details of the calico-node Pod (kubectl describe pod ...) or logs, error messages similar to the following will appear repeatedly.
      • Readiness probe failed: calico/node is not ready: BIRD is not ready
      • BGP not established with X.X.X.X (X.X.X.X is usually the IP of another node)
      • Error querying BIRD: unable to connect to BIRDv4 socket
    • Cross-node Pod communication is interrupted: Because the BGP session cannot be established, routing information cannot be synchronized. The most direct consequence is that Pods on the current node cannot communicate normally with Pods on other nodes in the cluster.
  • Cause

    The IP_AUTODETECTION_METHOD field of calico-node is set incorrectly, causing node network inaccessibility.

  • Solution

    Run the following command to modify the value of the environment variable.

    bash
    # Edit the K8s resource YAML and set the IP_AUTODETECTION_METHOD environment variable
    kubectl edit ds -n kube-system calico-node

    Common settings are as follows.

    • skip-interface=nerdctl*: The default policy set by openFuyao, which skips network cards with the nerdctl prefix and selects the IP of the first valid network card.
    • can-reach=192.168.100.5: Directly specifies the network card IP of the target address.
    • interface=eth4: Uses the IP on the eth4 network card.
  • Extended Discussion

    The IP_AUTODETECTION_METHOD field determines how Calico automatically selects the correct IP address for establishing BGP neighbors and encapsulating traffic on nodes with multiple network cards.

4. When the bootstrap cluster installs a service cluster, the bkeagent of the service cluster cannot connect to the APIServer of the bootstrap cluster, causing cluster installation to fail ​

  • Cause

    When bkeagent starts, it specifies the --kube-config parameter to configure the APIServer it monitors. After bkeagent starts, it checks whether the CRD it reconciles has been deployed in the monitored APIServer, which triggers bkeagent to attempt to connect to the bootstrap cluster's APIServer, resulting in the following error log.

    # For v25.12 and earlier versions, the log is /var/log/bkeagenbt.log; for later versions, it is /var/log/openFuyao/bkeagent.log
    The CRD cannot be installed in the target cluster, xxx

    When viewing the configuration file (for v25.12 and earlier versions, the log is at /etc/bkeagent/config; for later versions, it is at /etc/openFuyao/bkeagent/config), you will find that the IP address corresponding to server is not the address of the given bootstrap node.

  • Solution

    Run the following command to reset the bootstrap node, then specify the IP address for initialization.

    bash
    # Reset the bootstrap node
    bke reset --all --mount
    # Specify the IP address for initialization
    bke init --hostIP=1.2.3.4

5. When the bootstrap cluster deploys the openFuyao management plane, coredns is in a CrashLoopbackOff state ​

  • Details

    View the detailed logs of coredns (kubectl logs -n kube-system coredns-xxx), and the following logs appear (the actual log IP addresses and Port values differ).

    [ERROR] plugin/errors: 2 . NS: read udp 100.20.0.15:59690->100.10.0.10:53: i/o timeout
    [FATAL] plugin/loop: Loop(127.0.0.1:34812 -> :53) detected for zone ".", see https://coredns.io/plugins/loop#troubleshooting. Query: "HINFO ***"
  • Cause

    The above logs indicate that coredns falls into a loop infinite cycle during domain name resolution.

  • Solution

    kubectl get cm -n kube-system coredns -o yaml shows that forward is set to /etc/resolv.conf, which causes coredns's server ip to be used as the server for domain name resolution when no other DNS server is available, ultimately resulting in an infinite loop.

    Run the following command to set the server.

    bash
    # Set forward to "forward . 8.8.8.8"; if a DNS server is available, set it to the corresponding IP address
    kubectl edit cm -n kube-system coredns
  • Extended Discussion

    The forward plugin forwards DNS requests that cannot be resolved within the cluster to the specified upstream DNS server.

6. Deploying a high-availability cluster using a virtual IP fails; after resetting nodes with bke reset, redeploying the high-availability cluster using the occupied virtual IP fails again ​

  • Cause

    The keepalived component of the high-availability cluster binds the virtual IP to a node in the high-availability cluster. After resetting the environment on each node using bke reset, the virtual IP is not unbound from the node it was bound to, causing an error when reused, ultimately preventing the cluster from being brought up.

  • Solution

    Log in to the Master node of the high-availability cluster from the first installation, and run the following command to unbind the virtual IP.

    bash
    # View the IP addresses bound to the node's network cards
    ip addr
    # If a virtual IP is found to be bound, run the following command to unbind it; replace vip with the actual virtual IP used, and eth0 with the actual bound network card
    ip addr del <vip> dev <eth0>

7. For clusters installed by openFuyao, besides deleting a cluster through the management plane, how can the backend delete a cluster ​

For the formal steps, deletion order, cleanup scope, and success criteria, see Cluster Uninstallation and Delete a Service Cluster via Backend (Command Line).

The installer-service of the bootstrap cluster or management cluster reads BKECluster data from the APIServer. In addition to deletion through the management plane, you can also execute the following in the bootstrap node or management cluster terminal:

bash
# Query cluster information
kubectl get bc -A
# Replace the placeholders with the actual namespace and name, then edit
kubectl edit bc -n <namespace> <name>

In the opened YAML editor, set the following two items (when performing a thorough cleanup of the target node):

yaml
metadata:
  annotations:
    bke.bocloud.com/ignore-target-cluster-delete: "false"
spec:
  reset: true

8. How does the backend perform scaling operations for an openFuyao cluster ​

The openFuyao management plane provides cluster lifecycle management capabilities, including cluster scaling in, scaling out, upgrade, installation, and uninstallation. Here we provide backend procedures for cluster scaling operations, which must be performed when the cluster is in a healthy state. When the cluster is in an unhealthy state, scaling operations may result in errors.

  • Scale in operation: remove nodes from an existing cluster.

    View existing BKENode resources.

    bash
    # Replace bke-cluster with the actual cluster information
    kubectl get bn -n bke-cluster

    Delete the corresponding BKENode resource.

    bash
    # Replace bke-cluster-n1 with the actual node name
    kubectl delete bn -n bke-cluster bke-cluster-n1

    Run the command to view existing BKENode resources again; if the corresponding node is no longer present, the deletion was successful.

  • Scale out operation: add new nodes to an existing cluster.

    Write the configuration file (newNode.yaml) for the new node.

    yaml
    apiVersion: bke.bocloud.com/v1beta1
    kind: BKENode
    metadata:
      name: bke-cluster-n1
      namespace: bke-cluster
      labels:
        cluster.x-k8s.io/cluster-name: bke-cluster
    spec:
      hostname: n1
      ip: <node-ip>
      password: '<encrypted>'
      port: "22"     
      role:
      - node
      username: root

    Run the following command to perform the scale out operation.

    bash
    kubectl apply -f newNode.yaml

    Run the command to view existing BKENode resources; if the corresponding node is Ready, the scale out was successful.

Note: For v25.12 and earlier versions, follow the guide below.

  • Scale in operation: remove nodes from an existing cluster.

    Edit the BKECluster resource.

    bash
    # Replace bke-cluster with the actual cluster information
    kubectl edit bc -n bke-cluster bke-cluster

    Set the nodes scheduled for deletion.

    yaml
    metadata:
      annotations:
        # Node scheduled deletion; deleting a node is a dangerous action, so this annotation is added for secondary confirmation; none by default
        # When deleting a node, in addition to removing it from spec, you also need to fill in the IP of the node to be deleted; separate multiple IPs with ','
        # Both steps are required; missing either will not trigger node deletion
        bke.bocloud.com/appointment-deleted-nodes: "172.100.200.10"

    Remove the node information from Spec.

    yaml
    spec:
      clusterConfig:
        nodes:                
        - hostname: master-1  
          ip: 172.100.200.10
          username: root      
          password: password0
          port: "22"
          role:             
          - master
          - etcd
  • Scale out operation: add new nodes to an existing cluster.

    Edit the BKECluster resource.

    bash
    # Replace bke-cluster with the actual cluster information
    kubectl edit bc -n bke-cluster bke-cluster

    Add the node information to Spec.

    yaml
    spec:
      clusterConfig:
        nodes:                
        - hostname: master-1  
          ip: 172.100.200.10
          username: root      
          password: password0
          port: "22"
          role:             
          - master
          - etcd

9. When the default policy of the node iptables FORWARD chain is DROP, it causes abnormal container network communication and cluster initialization or creation failures ​

  • Detailed Symptoms

    When iptables exists on a node and the default policy of the FORWARD chain is DROP, one or more of the following symptoms may occur.

    1. After bke init initializes the bootstrap node, the container services (image repository, YUM repository, Chart repository) on the bootstrap node start normally, but inter-Pod communication in the bootstrap cluster (k3s) is abnormal, or Pod access to external services times out.
    2. When the bootstrap cluster deploys the openFuyao management plane, some Pods (such as the cluster-api-provider-bke controller) experience network inaccessibility, with errors such as connection timeouts or refusals in the logs.
    3. The openFuyao frontend management plane access times out; accessing the backend service via curl on the bootstrap node works normally, but running curl to access it on other nodes times out or is unreachable.
    4. During service cluster node initialization or the node joining process, image pulling times out, and the node cannot join the cluster.
    5. After the cluster is created, cross-node Pod communication fails, CoreDNS resolution times out, and Calico BGP sessions cannot be established.
  • Solution

    On the node experiencing the above symptoms, manually change the default policy of the iptables FORWARD chain to ACCEPT.

    bash
    # View the current default policy of the FORWARD chain
    iptables -t filter -nvL FORWARD
    
    # Set the default policy of the FORWARD chain to ACCEPT
    iptables -t filter -P FORWARD ACCEPT

    If you need to set this before initializing the bootstrap node, you can run the above command before executing bke init.

  • Extended Discussion

    Common sources of the iptables FORWARD chain default policy being DROP include:

    • firewalld firewall was enabled during system installation, and firewalld sets the FORWARD chain policy to DROP by default. Even if the firewalld service is subsequently stopped, the iptables policies already set will not automatically revert to ACCEPT.
    • In security hardening scenarios, administrators manually set the FORWARD chain policy to DROP to restrict unauthorized traffic forwarding.
    • Some Linux distributions (such as certain minimal installations of CentOS/RHEL) set the FORWARD policy to DROP by default.

10. During installation, /etc/resolv.conf is refreshed to default content by NetworkManager, causing node network abnormalities ​

  • Cause

    During installation, the NetworkManager (NM) service manages the node's DNS configuration and refreshes /etc/resolv.conf to default content, overwriting the configured valid DNS servers. This causes abnormal domain name resolution on the node, which in turn leads to network inaccessibility, image pulling failures, component communication timeouts, and other issues.

  • Solution

    Modify the configuration of NetworkManager and set dns=none to prevent NetworkManager from managing the DNS configuration.

    bash
    # Edit the NetworkManager main configuration file
    vi /etc/NetworkManager/NetworkManager.conf

    Add or modify the following configuration in the [main] section.

    ini
    [main]
    dns=none

    After saving, restart the NetworkManager service for the configuration to take effect, and confirm that the DNS configuration in /etc/resolv.conf is correct.

    bash
    # Restart the NetworkManager service
    systemctl restart NetworkManager
    # Confirm that the resolv.conf content has not been overwritten with default configuration
    cat /etc/resolv.conf
  • Extended Discussion

    dns=none means that NetworkManager will no longer write to or update /etc/resolv.conf, and the DNS configuration is maintained by the administrator. It is recommended to configure this before installation to avoid unexpected refresh during the installation process.

11. Bootstrap node k3s container keeps restarting intermittently ​

  • Detailed Phenomenon

    On the bootstrap node, execute ps -ef | grep k3s and nerdctl ps -a, and find that the k3s server restarts intermittently. Using nerdctl logs kubernetes, the log shows failed to run iptables command to create KUBE-ROUTER-OUTPUT chain due to running [/bin/aux/iptables -t filter -S KUBE-ROUTER-OUTPUT 1 --wait]: exit status 3.

  • Cause

    Execute lsmod | grep -E 'ip_tables|iptable_filter' with no output. The cause of this phenomenon is that when k3s starts the network component, the iptables related modules of the host kernel are missing, causing the network policy controller initialization to fail, triggering a panic and exit.

  • Solution

    The corresponding modules can be loaded with the following commands:

    bash
    modprobe ip_tables
    modprobe iptables_filter
    modprobe iptables_nat

12. On a single-node cluster, one of the two CoreDNS Pods occasionally stays in the Pending state ​

  • Detailed Phenomenon

    On the created single-node cluster, run kubectl get pod -n kube-system | grep coredns, and you will find that one CoreDNS Pod is in the Running state while the other is in the Pending state.

  • Solution

    Run the following command to scale down the CoreDNS replicas to 1. For a single-node environment, one CoreDNS replica is fully sufficient to provide DNS resolution services.

    bash
    kubectl scale deployment coredns --replicas=1 -n kube-system

13. chartRepo validation fails when configuring chart-type addons, or chart component pull fails during installation ​

  • Detailed Phenomenon

    1. When creating a cluster with a type: chart addon (for example, logging-package), bke cluster create fails with an error similar to:

      text
      failed calling webhook "vbkecluster.kb.io": ... context deadline exceeded

      or the Validating Webhook times out (about 10s by default), so the BKECluster cannot be created.

    2. Cluster creation has passed, but installing a chart component fails while pulling the chart, with a log similar to:

      text
      failed to pull chart: ... Get "https://<chartRepo-ip>:443/v2/charts/...": connection reset by peer

      The request falls back to the chartRepo IP instead of the domain name.

  • Cause

    1. After a chart-type addon is configured, the Validating Webhook in bke-controller-manager performs a reachability check on chartRepo. The check may take longer than the webhook default timeout (10s), causing admission to fail with a timeout.
    2. bke-controller-manager defaults to dnsPolicy: ClusterFirst and depends on cluster DNS (CoreDNS). If CoreDNS is not ready or unavailable, or the Pod cannot resolve the external domain, ResolveReachableChartRepo falls back to chartRepo.ip. HTTPS access to a public registry by bare IP often fails with TLS/SNI/CDN issues such as connection reset by peer, so chart pull fails.
  • Solution

    Scenario 1: chartRepo validation / Webhook timeout

    Increase the timeout of the vbkecluster.kb.io webhook to 30s (adjust the webhook index as needed; use the command below to confirm the index).

    bash
    # List webhook names and their 0-based indexes
    kubectl get validatingwebhookconfiguration bke-validating-webhook-configuration \
      -o jsonpath='{range .webhooks[*]}{.name}{"\n"}{end}' | nl -v 0
    
    # Set timeoutSeconds to 30 for the matching index (example assumes index 0)
    kubectl patch validatingwebhookconfiguration bke-validating-webhook-configuration --type json -p '[
      {"op":"replace","path":"/webhooks/0/timeoutSeconds","value":30}
    ]'

    Then rerun bke cluster create.

    Scenario 2: chart component pull fails during installation

    Change the dnsPolicy of bke-controller-manager in the cluster-system namespace from the default ClusterFirst to None, so the Pod uses the configured dnsConfig.nameservers (for example, 8.8.8.8) to resolve external domains and avoid incorrectly falling back to the chartRepo IP.

    bash
    kubectl -n cluster-system patch deployment bke-controller-manager --type strategic -p '{
      "spec": {
        "template": {
          "spec": {
            "dnsPolicy": "None"
          }
        }
      }
    }'
    kubectl -n cluster-system rollout status deployment/bke-controller-manager

    Confirm the Pod has taken effect:

    bash
    kubectl -n cluster-system get pod -l control-plane=controller-manager \
      -o jsonpath='{.items[0].spec.dnsPolicy}{"\n"}'
    # Expected output: None

    Then retry chart installation in one of the following ways:

    1. Recommended: after dnsPolicy / chartRepo are reachable, delete and recreate the cluster, then reinstall.

    2. Remove the failed entry from status.addonStatus, then add the retry annotation to trigger reconcile:

      bash
      # Step 1: remove the failed chart entry (example: logging-package)
      kubectl -n <bke-ns> get bkecluster <bke-name> -o json | \
        jq '(.status.addonStatus) |= map(select(.name != "logging-package"))' | \
        kubectl -n <bke-ns> replace --subresource=status -f -
      
      # Step 2: add the retry annotation to trigger reconcile
      kubectl -n <bke-ns> annotate bkecluster <bke-name> bke.bocloud.com/retry=
    3. Manually install the chart into the target cluster with helm/OCI.

  • Extended Discussion

    • Extending the webhook timeout only mitigates admission failures caused by slow reachability checks; it does not replace real reachability of the chart repository.
    • After setting dnsPolicy: None, the Pod fully depends on the nameservers in dnsConfig. Ensure bke-controller-manager has dnsConfig.nameservers configured and that those DNS servers are reachable.
    • In offline environments or environments without public network access, upload charts to the bootstrap node local chart repository (for example, port 38080) and point chartRepo to that reachable address, instead of depending on the public cr.openfuyao.cn.