Frequently Asked Questions (FAQ)
This document aims to provide common installation and deployment issues along with available solutions.
1. The install-cni container of the calico component keeps restarting, with calico-node stuck in the init phase (1/3) and in a CrashLoopbackOff state
Cause
You can run
ip routeon the node to view routing information. This issue is usually caused by one of the following two situations.- The node is missing a default route.
- The node has multiple default routes, and the route assigned for cluster deployment is not the one with the highest priority (a smaller metric value indicates a higher priority).
Solution
For cause 1, add a default route by running the following command.
baship route add default via 192.168.100.1 dev eth0 proto static metric 100For cause 2, delete the redundant default route by running the following command.
baship route del default via 192.168.90.1 dev eth1Extended Discussion
Calico requires that each node has a default route with the highest priority, essentially to ensure that all traffic to and from the cluster can be correctly forwarded at the node, which acts as a Layer 3 router.
The default route is typically configured under the
/etc/sysconfig/network-scriptsdirectory, which defines and configures persistent configuration files for system network interfaces (network cards). When the system restarts or the network service restarts, the system reads these files to set theIPaddress, gateway,DNS, and other information.
2. calico-node restarts frequently, and the container falls into a "start - probe failure - restart" infinite loop
Cause
This issue is usually caused by setting
initialDelaySecondstoo short. The Pod has not yet completed initialization, connection establishment, configuration loading, and other startup operations whenkubeletstarts probing the health status ofcalico-node.Solution
Set an appropriate delay time to tell
kubeletto wait for a period of time after the container starts before performing the first probe. Run the following command to configure it.bash# Edit the K8s resource YAML, find the readinessProbe section, and add the initialDelaySeconds field kubectl edit ds -n kube-system calico-nodeThe recommended
initialDelaySecondssettings are as follows; adjust according to actual scenarios.- Small clusters (fewer than 10 nodes): set to 30, which is sufficient for
calicoto complete basic initialization. - Large clusters (more than 10 nodes): set to 60, as more nodes and intense resource competition require a longer startup time.
- Small clusters (fewer than 10 nodes): set to 30, which is sufficient for
3. calico-node is in a 0/1 Running state
Detailed Symptoms
View the detailed information as follows.
- Node status is
NotReady: This is the most obvious symptom. Through thekubectl get nodescommand, you will find that the node status is notReady, butNotReady. calico-nodePod probe failure: When viewing the details of thecalico-nodePod (kubectl describe pod ...) or logs, error messages similar to the following will appear repeatedly.Readiness probe failed: calico/node is not ready: BIRD is not readyBGP not established with X.X.X.X(X.X.X.X is usually the IP of another node)Error querying BIRD: unable to connect to BIRDv4 socket
- Cross-node Pod communication is interrupted: Because the BGP session cannot be established, routing information cannot be synchronized. The most direct consequence is that Pods on the current node cannot communicate normally with Pods on other nodes in the cluster.
- Node status is
Cause
The
IP_AUTODETECTION_METHODfield ofcalico-nodeis set incorrectly, causing node network inaccessibility.Solution
Run the following command to modify the value of the environment variable.
bash# Edit the K8s resource YAML and set the IP_AUTODETECTION_METHOD environment variable kubectl edit ds -n kube-system calico-nodeCommon settings are as follows.
skip-interface=nerdctl*: The default policy set byopenFuyao, which skips network cards with thenerdctlprefix and selects the IP of the first valid network card.can-reach=192.168.100.5: Directly specifies the network card IP of the target address.interface=eth4: Uses the IP on theeth4network card.
Extended Discussion
The
IP_AUTODETECTION_METHODfield determines how Calico automatically selects the correct IP address for establishingBGPneighbors and encapsulating traffic on nodes with multiple network cards.
4. When the bootstrap cluster installs a service cluster, the bkeagent of the service cluster cannot connect to the APIServer of the bootstrap cluster, causing cluster installation to fail
Cause
When
bkeagentstarts, it specifies the--kube-configparameter to configure theAPIServerit monitors. Afterbkeagentstarts, it checks whether theCRDit reconciles has been deployed in the monitoredAPIServer, which triggersbkeagentto attempt to connect to the bootstrap cluster'sAPIServer, resulting in the following error log.# For v25.12 and earlier versions, the log is /var/log/bkeagenbt.log; for later versions, it is /var/log/openFuyao/bkeagent.log The CRD cannot be installed in the target cluster, xxxWhen viewing the configuration file (for v25.12 and earlier versions, the log is at
/etc/bkeagent/config; for later versions, it is at/etc/openFuyao/bkeagent/config), you will find that theIPaddress corresponding toserveris not the address of the given bootstrap node.Solution
Run the following command to reset the bootstrap node, then specify the
IPaddress for initialization.bash# Reset the bootstrap node bke reset --all --mount # Specify the IP address for initialization bke init --hostIP=1.2.3.4
5. When the bootstrap cluster deploys the openFuyao management plane, coredns is in a CrashLoopbackOff state
Details
View the detailed logs of
coredns(kubectl logs -n kube-system coredns-xxx), and the following logs appear (the actual logIPaddresses andPortvalues differ).[ERROR] plugin/errors: 2 . NS: read udp 100.20.0.15:59690->100.10.0.10:53: i/o timeout [FATAL] plugin/loop: Loop(127.0.0.1:34812 -> :53) detected for zone ".", see https://coredns.io/plugins/loop#troubleshooting. Query: "HINFO ***"Cause
The above logs indicate that
corednsfalls into aloopinfinite cycle during domain name resolution.Solution
kubectl get cm -n kube-system coredns -o yamlshows thatforwardis set to/etc/resolv.conf, which causescoredns'sserver ipto be used as the server for domain name resolution when no otherDNSserver is available, ultimately resulting in an infinite loop.Run the following command to set the server.
bash# Set forward to "forward . 8.8.8.8"; if a DNS server is available, set it to the corresponding IP address kubectl edit cm -n kube-system corednsExtended Discussion
The
forwardplugin forwardsDNSrequests that cannot be resolved within the cluster to the specified upstreamDNSserver.
6. Deploying a high-availability cluster using a virtual IP fails; after resetting nodes with bke reset, redeploying the high-availability cluster using the occupied virtual IP fails again
Cause
The
keepalivedcomponent of the high-availability cluster binds the virtualIPto a node in the high-availability cluster. After resetting the environment on each node usingbke reset, the virtualIPis not unbound from the node it was bound to, causing an error when reused, ultimately preventing the cluster from being brought up.Solution
Log in to the Master node of the high-availability cluster from the first installation, and run the following command to unbind the virtual
IP.bash# View the IP addresses bound to the node's network cards ip addr # If a virtual IP is found to be bound, run the following command to unbind it; replace vip with the actual virtual IP used, and eth0 with the actual bound network card ip addr del <vip> dev <eth0>
7. For clusters installed by openFuyao, besides deleting a cluster through the management plane, how can the backend delete a cluster
For the formal steps, deletion order, cleanup scope, and success criteria, see Cluster Uninstallation and Delete a Service Cluster via Backend (Command Line).
The installer-service of the bootstrap cluster or management cluster reads BKECluster data from the APIServer. In addition to deletion through the management plane, you can also execute the following in the bootstrap node or management cluster terminal:
# Query cluster information
kubectl get bc -A
# Replace the placeholders with the actual namespace and name, then edit
kubectl edit bc -n <namespace> <name>In the opened YAML editor, set the following two items (when performing a thorough cleanup of the target node):
metadata:
annotations:
bke.bocloud.com/ignore-target-cluster-delete: "false"
spec:
reset: true8. How does the backend perform scaling operations for an openFuyao cluster
The openFuyao management plane provides cluster lifecycle management capabilities, including cluster scaling in, scaling out, upgrade, installation, and uninstallation. Here we provide backend procedures for cluster scaling operations, which must be performed when the cluster is in a healthy state. When the cluster is in an unhealthy state, scaling operations may result in errors.
Scale in operation: remove nodes from an existing cluster.
View existing
BKENoderesources.bash# Replace bke-cluster with the actual cluster information kubectl get bn -n bke-clusterDelete the corresponding BKENode resource.
bash# Replace bke-cluster-n1 with the actual node name kubectl delete bn -n bke-cluster bke-cluster-n1Run the command to view existing
BKENoderesources again; if the corresponding node is no longer present, the deletion was successful.Scale out operation: add new nodes to an existing cluster.
Write the configuration file (newNode.yaml) for the new node.
yamlapiVersion: bke.bocloud.com/v1beta1 kind: BKENode metadata: name: bke-cluster-n1 namespace: bke-cluster labels: cluster.x-k8s.io/cluster-name: bke-cluster spec: hostname: n1 ip: <node-ip> password: '<encrypted>' port: "22" role: - node username: rootRun the following command to perform the scale out operation.
bashkubectl apply -f newNode.yamlRun the command to view existing
BKENoderesources; if the corresponding node is Ready, the scale out was successful.
Note: For v25.12 and earlier versions, follow the guide below.
Scale in operation: remove nodes from an existing cluster.
Edit the
BKEClusterresource.bash# Replace bke-cluster with the actual cluster information kubectl edit bc -n bke-cluster bke-clusterSet the nodes scheduled for deletion.
yamlmetadata: annotations: # Node scheduled deletion; deleting a node is a dangerous action, so this annotation is added for secondary confirmation; none by default # When deleting a node, in addition to removing it from spec, you also need to fill in the IP of the node to be deleted; separate multiple IPs with ',' # Both steps are required; missing either will not trigger node deletion bke.bocloud.com/appointment-deleted-nodes: "172.100.200.10"Remove the node information from
Spec.yamlspec: clusterConfig: nodes: - hostname: master-1 ip: 172.100.200.10 username: root password: password0 port: "22" role: - master - etcdScale out operation: add new nodes to an existing cluster.
Edit the
BKEClusterresource.bash# Replace bke-cluster with the actual cluster information kubectl edit bc -n bke-cluster bke-clusterAdd the node information to
Spec.yamlspec: clusterConfig: nodes: - hostname: master-1 ip: 172.100.200.10 username: root password: password0 port: "22" role: - master - etcd
9. When the default policy of the node iptables FORWARD chain is DROP, it causes abnormal container network communication and cluster initialization or creation failures
Detailed Symptoms
When iptables exists on a node and the default policy of the FORWARD chain is DROP, one or more of the following symptoms may occur.
- After
bke initinitializes the bootstrap node, the container services (image repository, YUM repository, Chart repository) on the bootstrap node start normally, but inter-Pod communication in the bootstrap cluster (k3s) is abnormal, or Pod access to external services times out. - When the bootstrap cluster deploys the openFuyao management plane, some Pods (such as the cluster-api-provider-bke controller) experience network inaccessibility, with errors such as connection timeouts or refusals in the logs.
- The openFuyao frontend management plane access times out; accessing the backend service via
curlon the bootstrap node works normally, but runningcurlto access it on other nodes times out or is unreachable. - During service cluster node initialization or the node joining process, image pulling times out, and the node cannot join the cluster.
- After the cluster is created, cross-node Pod communication fails, CoreDNS resolution times out, and Calico BGP sessions cannot be established.
- After
Solution
On the node experiencing the above symptoms, manually change the default policy of the iptables FORWARD chain to ACCEPT.
bash# View the current default policy of the FORWARD chain iptables -t filter -nvL FORWARD # Set the default policy of the FORWARD chain to ACCEPT iptables -t filter -P FORWARD ACCEPTIf you need to set this before initializing the bootstrap node, you can run the above command before executing
bke init.Extended Discussion
Common sources of the iptables FORWARD chain default policy being DROP include:
- firewalld firewall was enabled during system installation, and firewalld sets the FORWARD chain policy to DROP by default. Even if the firewalld service is subsequently stopped, the iptables policies already set will not automatically revert to ACCEPT.
- In security hardening scenarios, administrators manually set the FORWARD chain policy to DROP to restrict unauthorized traffic forwarding.
- Some Linux distributions (such as certain minimal installations of CentOS/RHEL) set the FORWARD policy to DROP by default.
10. During installation, /etc/resolv.conf is refreshed to default content by NetworkManager, causing node network abnormalities
Cause
During installation, the
NetworkManager(NM) service manages the node'sDNSconfiguration and refreshes/etc/resolv.confto default content, overwriting the configured validDNSservers. This causes abnormal domain name resolution on the node, which in turn leads to network inaccessibility, image pulling failures, component communication timeouts, and other issues.Solution
Modify the configuration of
NetworkManagerand setdns=noneto preventNetworkManagerfrom managing theDNSconfiguration.bash# Edit the NetworkManager main configuration file vi /etc/NetworkManager/NetworkManager.confAdd or modify the following configuration in the
[main]section.ini[main] dns=noneAfter saving, restart the
NetworkManagerservice for the configuration to take effect, and confirm that theDNSconfiguration in/etc/resolv.confis correct.bash# Restart the NetworkManager service systemctl restart NetworkManager # Confirm that the resolv.conf content has not been overwritten with default configuration cat /etc/resolv.confExtended Discussion
dns=nonemeans thatNetworkManagerwill no longer write to or update/etc/resolv.conf, and theDNSconfiguration is maintained by the administrator. It is recommended to configure this before installation to avoid unexpected refresh during the installation process.
11. Bootstrap node k3s container keeps restarting intermittently
Detailed Phenomenon
On the bootstrap node, execute
ps -ef | grep k3sandnerdctl ps -a, and find that the k3s server restarts intermittently. Usingnerdctl logs kubernetes, the log shows failed to run iptables command to create KUBE-ROUTER-OUTPUT chain due to running [/bin/aux/iptables -t filter -S KUBE-ROUTER-OUTPUT 1 --wait]: exit status 3.Cause
Execute
lsmod | grep -E 'ip_tables|iptable_filter'with no output. The cause of this phenomenon is that when k3s starts the network component, theiptablesrelated modules of the host kernel are missing, causing the network policy controller initialization to fail, triggering a panic and exit.Solution
The corresponding modules can be loaded with the following commands:
bashmodprobe ip_tables modprobe iptables_filter modprobe iptables_nat
12. On a single-node cluster, one of the two CoreDNS Pods occasionally stays in the Pending state
Detailed Phenomenon
On the created single-node cluster, run
kubectl get pod -n kube-system | grep coredns, and you will find that one CoreDNS Pod is in theRunningstate while the other is in thePendingstate.Solution
Run the following command to scale down the CoreDNS replicas to 1. For a single-node environment, one CoreDNS replica is fully sufficient to provide DNS resolution services.
bashkubectl scale deployment coredns --replicas=1 -n kube-system
13. chartRepo validation fails when configuring chart-type addons, or chart component pull fails during installation
Detailed Phenomenon
When creating a cluster with a
type: chartaddon (for example,logging-package),bke cluster createfails with an error similar to:textfailed calling webhook "vbkecluster.kb.io": ... context deadline exceededor the Validating Webhook times out (about 10s by default), so the
BKEClustercannot be created.Cluster creation has passed, but installing a chart component fails while pulling the chart, with a log similar to:
textfailed to pull chart: ... Get "https://<chartRepo-ip>:443/v2/charts/...": connection reset by peerThe request falls back to the chartRepo IP instead of the domain name.
Cause
- After a chart-type addon is configured, the Validating Webhook in
bke-controller-managerperforms a reachability check onchartRepo. The check may take longer than the webhook default timeout (10s), causing admission to fail with a timeout. bke-controller-managerdefaults todnsPolicy: ClusterFirstand depends on cluster DNS (CoreDNS). If CoreDNS is not ready or unavailable, or the Pod cannot resolve the external domain,ResolveReachableChartRepofalls back tochartRepo.ip. HTTPS access to a public registry by bare IP often fails with TLS/SNI/CDN issues such asconnection reset by peer, so chart pull fails.
- After a chart-type addon is configured, the Validating Webhook in
Solution
Scenario 1: chartRepo validation / Webhook timeout
Increase the timeout of the
vbkecluster.kb.iowebhook to 30s (adjust the webhook index as needed; use the command below to confirm the index).bash# List webhook names and their 0-based indexes kubectl get validatingwebhookconfiguration bke-validating-webhook-configuration \ -o jsonpath='{range .webhooks[*]}{.name}{"\n"}{end}' | nl -v 0 # Set timeoutSeconds to 30 for the matching index (example assumes index 0) kubectl patch validatingwebhookconfiguration bke-validating-webhook-configuration --type json -p '[ {"op":"replace","path":"/webhooks/0/timeoutSeconds","value":30} ]'Then rerun
bke cluster create.Scenario 2: chart component pull fails during installation
Change the
dnsPolicyofbke-controller-managerin thecluster-systemnamespace from the defaultClusterFirsttoNone, so the Pod uses the configureddnsConfig.nameservers(for example,8.8.8.8) to resolve external domains and avoid incorrectly falling back to the chartRepo IP.bashkubectl -n cluster-system patch deployment bke-controller-manager --type strategic -p '{ "spec": { "template": { "spec": { "dnsPolicy": "None" } } } }' kubectl -n cluster-system rollout status deployment/bke-controller-managerConfirm the Pod has taken effect:
bashkubectl -n cluster-system get pod -l control-plane=controller-manager \ -o jsonpath='{.items[0].spec.dnsPolicy}{"\n"}' # Expected output: NoneThen retry chart installation in one of the following ways:
Recommended: after
dnsPolicy/ chartRepo are reachable, delete and recreate the cluster, then reinstall.Remove the failed entry from
status.addonStatus, then add the retry annotation to trigger reconcile:bash# Step 1: remove the failed chart entry (example: logging-package) kubectl -n <bke-ns> get bkecluster <bke-name> -o json | \ jq '(.status.addonStatus) |= map(select(.name != "logging-package"))' | \ kubectl -n <bke-ns> replace --subresource=status -f - # Step 2: add the retry annotation to trigger reconcile kubectl -n <bke-ns> annotate bkecluster <bke-name> bke.bocloud.com/retry=Manually install the chart into the target cluster with helm/OCI.
Extended Discussion
- Extending the webhook timeout only mitigates admission failures caused by slow reachability checks; it does not replace real reachability of the chart repository.
- After setting
dnsPolicy: None, the Pod fully depends on the nameservers indnsConfig. Ensurebke-controller-managerhasdnsConfig.nameserversconfigured and that those DNS servers are reachable. - In offline environments or environments without public network access, upload charts to the bootstrap node local chart repository (for example, port
38080) and pointchartRepoto that reachable address, instead of depending on the publiccr.openfuyao.cn.