![]()
Kubernetes Cluster API Provider GCP
Kubernetes-native declarative infrastructure for GCP.
What is the Cluster API Provider GCP?
The Cluster API brings declarative Kubernetes-style APIs to cluster creation, configuration and management. The API itself is shared across multiple cloud providers allowing for true Google Cloud hybrid deployments of Kubernetes.
Documentation
Please see our book for in-depth documentation.
Quick Start
Checkout our Cluster API Quick Start to create your first Kubernetes cluster on Google Cloud Platform using Cluster API.
Getting Involved and Contributing
Are you interested in contributing to cluster-api-provider-gcp? We, the maintainers and the community would love your suggestions, support and contributions! The maintainers of the project can be contacted anytime to learn about how to get involved.
Before starting with the contribution, please go through prerequisites of the project.
To set up the development environment, checkout the development guide.
In the interest of getting new people involved, we have issues marked as good first issue. Although
these issues have a smaller scope but are very helpful in getting acquainted with the codebase.
For more, see the issue tracker. If you’re unsure where to start, feel free to reach out to discuss.
See also: Our own contributor guide and the Kubernetes community page.
We also encourage ALL active community participants to act as if they are maintainers, even if you don’t have ‘official’ written permissions. This is a community effort and we are here to serve the Kubernetes community. If you have an active interest and you want to get involved, you have real power!
Office hours
- Join the SIG Cluster Lifecycle Google Group for access to documents and calendars.
- Participate in the conversations on Kubernetes Discuss
- Provider implementers office hours (CAPI)
- Weekly on Wednesdays @ 10:00 am PT (Pacific Time) on Zoom
- Previous meetings: [ notes | recordings ]
- Cluster API Provider GCP office hours (CAPG)
- Monthly on first Thursday @ 05:00 am PT (Pacific Time) on Zoom
- Previous meetings: [ notes|recordings ]
Other ways to communicate with the contributors
Please check in with us in the #cluster-api-gcp on Slack.
Github Issues
Bugs
If you think you have found a bug, please follow the instruction below.
- Please give a small amount of time giving due diligence to the issue tracker. Your issue might be a duplicate.
- Get the logs from the custom controllers and please paste them in the issue.
- Open a bug report.
- Remember users might be searching for the issue in the future, so please make sure to give it a meaningful title to help others.
- Feel free to reach out to the community on slack.
Tracking new feature
We also have an issue tracker to track features. If you think you have a feature idea, that could make Cluster API provider GCP become even more awesome, then follow these steps.
- Open a feature request.
- Remember users might be searching for the issue in the future, so please make sure to give it a meaningful title to help others.
- Clearly define the use case with concrete examples. Example: type
thisand cluster-api-provider-gcp doesthat. - Some of our larger features will require some design. If you would like to include a technical design in your feature, please go ahead.
- After the new feature is well understood and the design is agreed upon, we can start coding the feature. We would love for you to code it. So please open up a WIP (work in progress) PR and happy coding!
Code of conduct
Participation in the Kubernetes community is governed by the Kubernetes Code of Conduct.
Getting started with CAPG
In this section we’ll cover the basics of how to prepare your environment to use Cluster API Provider for GCP.
Before installing CAPG, your Kubernetes cluster has to be transformed into a CAPI management cluster. If you have already done this, you can jump directly to the next section: Installing CAPG. If, on the other hand, you have an existing Kubernetes cluster that is not yet configured as a CAPI management cluster, you can follow the guide from the CAPI book.
Requirements
- Linux or MacOS (Windows isn’t supported at the moment).
- A Google Cloud account.
- Packer and Ansible to build images
maketo useMakefiletargets- Install
coreutils(for timeout) on OSX
Credentials
To create and manage clusters, CAPG uses a GCP service account to authenticate with GCP’s APIs. There are two supported authentication methods.
First, create a service account with Editor permissions.
Service Account JSON Key
Generate a JSON Key for the service account and store it somewhere safe. This key will be base64-encoded and provided to CAPG at installation time (see Installing CAPG).
Workload Identity Federation (GKE management clusters)
If your CAPI management cluster runs on GKE, Workload Identity Federation is the preferred authentication method. It eliminates the need to manage JSON key files by binding the CAPG Kubernetes ServiceAccount to a GCP service account.
Enable Workload Identity on your GKE management cluster (if not already enabled):
gcloud container clusters update <MANAGEMENT_CLUSTER> \
--workload-pool=<PROJECT_ID>.svc.id.goog \
--region <REGION>
Grant the Workload Identity User role so the CAPG Kubernetes ServiceAccount can impersonate the GCP service account:
export GCP_SA_EMAIL=<gcp-service-account>@<PROJECT_ID>.iam.gserviceaccount.com
gcloud iam service-accounts add-iam-policy-binding "${GCP_SA_EMAIL}" \
--role roles/iam.workloadIdentityUser \
--member "serviceAccount:<PROJECT_ID>.svc.id.goog[capg-system/capg-manager]"
Then deploy CAPG using the config/wif overlay, substituting your GCP service account email:
export GCP_SA_EMAIL=<gcp-service-account>@<PROJECT_ID>.iam.gserviceaccount.com
kustomize build config/wif/ | envsubst | kubectl apply -f -
CAPG will authenticate via the GKE metadata server automatically — no credentials secret is needed.
Installing CAPG
There are two major provider installation paths: using clusterctl or the Cluster API Operator.
clusterctl is a command line tool that provides a simple way of interacting with CAPI and is usually the preferred alternative for those who are getting started. It automates fetching the YAML files defining provider components and installing them.
The Cluster API Operator is a Kubernetes Operator built on top of clusterctl and designed to empower cluster administrators to handle the lifecycle of Cluster API providers within a management cluster using a declarative approach. It aims to improve user experience in deploying and managing Cluster API, making it easier to handle day-to-day tasks and automate workflows with GitOps. Visit the CAPI Operator quickstart if you want to experiment with this tool.
You can opt for the tool that works best for you or explore both and decide which is best suited for your use case.
clusterctl
The Service Account you created will be used to interact with GCP and it must be base64 encoded and stored in a environment variable before installing the provider via clusterctl.
export GCP_B64ENCODED_CREDENTIALS=$( cat /path/to/gcp-credentials.json | base64 | tr -d '\n' )
Finally, let’s initialize the provider.
clusterctl init --infrastructure gcp
This process may take some time and, once the provider is running, you’ll be able to see the capg-controller-manager pod in your CAPI management cluster.
Cluster API Operator
You can refer to the Cluster API Operator book here to learn about the basics of the project and how to install the operator.
When using Cluster API Operator, secrets are used to store credentials for cloud providers and not environment variables, which means you’ll have to create a new secret containing the base64 encoded version of your GCP credentials and it will be referenced in the yaml file used to initialize the provider. As you can see, by using Cluster API Operator, we’re able to manage provider installation declaratively.
Create GCP credentials secret.
export CREDENTIALS_SECRET_NAME="gcp-credentials"
export CREDENTIALS_SECRET_NAMESPACE="default"
export GCP_B64ENCODED_CREDENTIALS=$( cat /path/to/gcp-credentials.json | base64 | tr -d '\n' )
kubectl create secret generic "${CREDENTIALS_SECRET_NAME}" --from-literal=GCP_B64ENCODED_CREDENTIALS="${GCP_B64ENCODED_CREDENTIALS}" --namespace "${CREDENTIALS_SECRET_NAMESPACE}"
Define CAPG provider declaratively in a file capg.yaml.
apiVersion: v1
kind: Namespace
metadata:
name: capg-system
---
apiVersion: operator.cluster.x-k8s.io/v1alpha2
kind: InfrastructureProvider
metadata:
name: gcp
namespace: capg-system
spec:
version: v1.8.0
configSecret:
name: gcp-credentials
After applying this file, Cluster API Operator will take care of installing CAPG using the set of credentials stored in the specified secret.
kubectl apply -f capg.yaml
Prerequisites
Before provisioning clusters via CAPG, there are a few extra tasks you need to take care of, including configuring the GCP network and building images for GCP virtual machines.
Set environment variables
export GCP_REGION="<GCP_REGION>"
export GCP_PROJECT="<GCP_PROJECT>"
# Make sure to use same kubernetes version here as building the GCE image
export KUBERNETES_VERSION=1.22.3
export GCP_CONTROL_PLANE_MACHINE_TYPE=n1-standard-2
export GCP_NODE_MACHINE_TYPE=n1-standard-2
export GCP_NETWORK_NAME=<GCP_NETWORK_NAME or default>
export CLUSTER_NAME="<CLUSTER_NAME>"
Configure Network and Cloud NAT
Google Cloud accounts come with a default network which can be found under
VPC Networks.
If you prefer to create a new Network, follow these instructions.
Cloud NAT
This infrastructure provider sets up Kubernetes clusters using a Global Load Balancer with a public ip address.
Kubernetes nodes, to communicate with the control plane, pull container images from registered (e.g. gcr.io or dockerhub) need to have NAT access or a public ip. By default, the provider creates Machines without a public IP.
To make sure your cluster can communicate with the outside world, and the load balancer, you can create a Cloud NAT in the region you’d like your Kubernetes cluster to live in by following these instructions.
NB: The following commands needs to be run if
${GCP_NETWORK_NAME}is set todefault
# Ensure if network list contains default network
gcloud compute networks list --project="${GCP_PROJECT}"
gcloud compute networks describe "${GCP_NETWORK_NAME}" --project="${GCP_PROJECT}"
# Ensure if firewall rules are enabled
$ gcloud compute firewall-rules list --project "$GCP_PROJECT"
# Create routers
gcloud compute routers create "${CLUSTER_NAME}-myrouter" --project="${GCP_PROJECT}" --region="${GCP_REGION}" --network="default"
# Create NAT
gcloud compute routers nats create "${CLUSTER_NAME}-mynat" --project="${GCP_PROJECT}" --router-region="${GCP_REGION}" --router="${CLUSTER_NAME}-myrouter"
--nat-all-subnet-ip-ranges --auto-allocate-nat-external-ips
Building images
NB: The following commands should not be run as
rootuser.
# Export the GCP project id you want to build images in.
export GCP_PROJECT_ID=<project-id>
# Export the path to the service account credentials created in the step above.
export GOOGLE_APPLICATION_CREDENTIALS=</path/to/serviceaccount-key.json>
# Clone the image builder repository if you haven't already.
git clone https://github.com/kubernetes-sigs/image-builder.git image-builder
# Change directory to images/capi within the image builder repository
cd image-builder/images/capi
# Run the Make target to generate GCE images.
make build-gce-ubuntu-2204
# Check that you can see the published images.
gcloud compute images list --project ${GCP_PROJECT_ID} --no-standard-images --filter="family:capi-ubuntu-2204-k8s"
# Export the IMAGE_ID from the above
export IMAGE_ID="projects/${GCP_PROJECT_ID}/global/images/<image-name>"
Clean-up
Delete the NAT gateway
gcloud compute routers nats delete "${CLUSTER_NAME}-mynat" --project="${GCP_PROJECT}" \
--router-region="${GCP_REGION}" --router="${CLUSTER_NAME}-myrouter" --quiet || true
Delete the router
gcloud compute routers delete "${CLUSTER_NAME}-myrouter" --project="${GCP_PROJECT}" \
--region="${GCP_REGION}" --quiet || true
Self-managed clusters
This section contains information about how you can provision self-managed Kubernetes clusters hosted in GCP’s Compute Engine.
Provisioning a self-managed Cluster
This guide uses an example from the ./templates folder of the CAPG repository. You can inspect the yaml file here.
Configure cluster parameters
While inspecting the cluster definition in ./templates/cluster-template.yaml you probably noticed that it contains a number of parameterized values that must be substituted with the specifics of your use case. This can be done via environment variables and clusterctl and effectively makes the template more flexible to adapt to different provisioning scenarios. These are the environment variables that you’ll be required to set before deploying a workload cluster:
export GCP_REGION=us-east4
export GCP_PROJECT=cluster-api-gcp-project
export CONTROL_PLANE_MACHINE_COUNT=1
export WORKER_MACHINE_COUNT=1
export KUBERNETES_VERSION=1.29.3
export GCP_CONTROL_PLANE_MACHINE_TYPE=n1-standard-2
export GCP_NODE_MACHINE_TYPE=n1-standard-2
export GCP_NETWORK_NAME=default
export IMAGE_ID=projects/cluster-api-gcp-project/global/images/your-image
Generate cluster definition
The sample cluster templates are already prepared so that you can use them with clusterctl to create a self-managed Kubernetes cluster with CAPG.
clusterctl generate cluster capi-gcp-quickstart -i gcp > capi-gcp-quickstart.yaml
In this example, capi-gcp-quickstart will be used as cluster name.
Create cluster
The resulting file represents the workload cluster definition and you simply need to apply it to your cluster to trigger cluster creation:
kubectl apply -f capi-gcp-quickstart.yaml
Kubeconfig
When creating an GCP cluster 2 kubeconfigs are generated and stored as secrets in the management cluster.
User kubeconfig
This should be used by users that want to connect to the newly created GCP cluster. The name of the secret that contains the kubeconfig will be [cluster-name]-user-kubeconfig where you need to replace [cluster-name] with the name of your cluster. The -user-kubeconfig in the name indicates that the kubeconfig is for the user use.
To get the user kubeconfig for a cluster named managed-test you can run a command similar to:
kubectl --namespace=default get secret managed-test-user-kubeconfig \
-o jsonpath={.data.value} | base64 --decode \
> managed-test.kubeconfig
Cluster API (CAPI) kubeconfig
This kubeconfig is used internally by CAPI and shouldn’t be used outside of the management server. It is used by CAPI to perform operations, such as draining a node. The name of the secret that contains the kubeconfig will be [cluster-name]-kubeconfig where you need to replace [cluster-name] with the name of your cluster. Note that there is NO -user in the name.
The kubeconfig is regenerated every sync-period as the token that is embedded in the kubeconfig is only valid for a short period of time.
CNI
By default, no CNI plugin is installed when a self-managed cluster is provisioned. As a user, you need to install your own CNI (e.g. Calico with VXLAN) for the control plane of the cluster to become ready.
This document describes how to use Flannel as your CNI solution.
Modify the Cluster resources
Before deploying the cluster, change the KubeadmControlPlane value at spec.kubeadmConfigSpec.clusterConfiguration.controllerManager.extraArgs.allocate-node-cidrs to "true"
apiVersion: controlplane.cluster.x-k8s.io/v1beta1
kind: KubeadmControlPlane
spec:
kubeadmConfigSpec:
clusterConfiguration:
controllerManager:
extraArgs:
allocate-node-cidrs: "true"
Modify Flannel Config
(NOTE): This is based off of the instruction at: deploying-flannel-manually
You need to make an adjustment to the default flannel configuration so that the CIDR inside your CAPG cluster matches the Flannel Network CIDR.
View your capi-cluster.yaml and make note of the Cluster Network CIDR Block. For example:
apiVersion: cluster.x-k8s.io/v1beta1
kind: Cluster
spec:
clusterNetwork:
pods:
cidrBlocks:
- 192.168.0.0/16
Download the file at https://raw.githubusercontent.com/coreos/flannel/master/Documentation/kube-flannel.yml and modify the kube-flannel-cfg ConfigMap. Set the value at data.net-conf.json.Network value to match your Cluster Network CIDR Block.
wget https://raw.githubusercontent.com/coreos/flannel/master/Documentation/kube-flannel.yml
Edit kube-flannel.yml and change this section so that the Network section matches your Cluster CIDR
kind: ConfigMap
apiVersion: v1
metadata:
name: kube-flannel-cfg
data:
net-conf.json: |
{
"Network": "192.168.0.0/16",
"Backend": {
"Type": "vxlan"
}
}
Apply kube-flannel.yml
kubectl apply -f kube-flannel.yml
GKE Support in the GCP Provider
- Feature status: Experimental
- Feature gate (required): GKE=true
Overview
The GCP provider supports creating GKE based cluster. Currently the following features are supported:
- Provisioning/managing a GCP GKE Cluster
- Upgrading the Kubernetes version of the GKE Cluster
- Creating a managed node pool and attaching it to the GKE cluster
The implementation introduces the following CRD kinds:
- GCPManagedCluster - presents the properties needed to provision and manage the general GCP operating infrastructure for the cluster (i.e project, networking, iam)
- GCPManagedControlPlane - specifies the GKE Cluster in GCP and used by the Cluster API GCP Managed Control plane
- GCPManagedMachinePool - defines the managed node pool for the cluster
And a new template is available in the templates folder for creating a managed workload cluster.
SEE ALSO
Provisioning a GKE cluster
This guide uses an example from the ./templates folder of the CAPG repository. You can inspect the yaml file here.
Configure cluster parameters
While inspecting the cluster definition in ./templates/cluster-template-gke.yaml you probably noticed that it contains a number of parameterized values that must be substituted with the specifics of your use case. This can be done via environment variables and clusterctl and effectively makes the template more flexible to adapt to different provisioning scenarios. These are the environment variables that you’ll be required to set before deploying a workload cluster:
export GCP_PROJECT=cluster-api-gcp-project
export GCP_REGION=us-east4
export GCP_NETWORK_NAME=default
export WORKER_MACHINE_COUNT=1
Generate cluster definition
The sample cluster templates are already prepared so that you can use them with clusterctl to create a GKE cluster with CAPG.
To create a GKE cluster with a managed node group (a.k.a managed machine pool):
clusterctl generate cluster capi-gke-quickstart --flavor gke -i gcp > capi-gke-quickstart.yaml
In this example, capi-gke-quickstart will be used as cluster name.
Create cluster
The resulting file represents the workload cluster definition and you simply need to apply it to your cluster to trigger cluster creation:
kubectl apply -f capi-gke-quickstart.yaml
Kubeconfig
When creating an GKE cluster 2 kubeconfigs are generated and stored as secrets in the management cluster.
User kubeconfig
This should be used by users that want to connect to the newly created GKE cluster. The name of the secret that contains the kubeconfig will be [cluster-name]-user-kubeconfig where you need to replace [cluster-name] with the name of your cluster. The -user-kubeconfig in the name indicates that the kubeconfig is for the user use.
To get the user kubeconfig for a cluster named managed-test you can run a command similar to:
kubectl --namespace=default get secret managed-test-user-kubeconfig \
-o jsonpath={.data.value} | base64 --decode \
> managed-test.kubeconfig
Cluster API (CAPI) kubeconfig
This kubeconfig is used internally by CAPI and shouldn’t be used outside of the management server. It is used by CAPI to perform operations, such as draining a node. The name of the secret that contains the kubeconfig will be [cluster-name]-kubeconfig where you need to replace [cluster-name] with the name of your cluster. Note that there is NO -user in the name.
The kubeconfig is regenerated every sync-period as the token that is embedded in the kubeconfig is only valid for a short period of time.
GKE Cluster Upgrades
Control Plane Upgrade
Upgrading the Kubernetes version of the control plane is supported by the provider. To perform an upgrade you need to update the controlPlaneVersion in the spec of the GCPManagedControlPlane. Once the version has changed the provider will handle the upgrade for you.
Enabling GKE Support
Enabling GKE support is done via the GKE feature flag by setting it to true. This can be done before running clusterctl init by using the EXP_CAPG_GKE environment variable:
export EXP_CAPG_GKE=true
clusterctl init --infrastructure gcp
Disabling GKE Support
Support for GKE is disabled by default when you use the GCP infrastructure provider.
Network Configuration
spec.clusterNetwork on GCPManagedControlPlane also controls two GKE cluster-networking features: the datapath provider (Dataplane V2) and the in-cluster DNS provider.
Datapath provider (Dataplane V2)
datapathProvider selects the implementation of the Kubernetes networking model GKE uses for service resolution and network policy enforcement:
apiVersion: infrastructure.cluster.x-k8s.io/v1beta1
kind: GCPManagedControlPlane
metadata:
name: my-cluster
spec:
clusterNetwork:
datapathProvider: advanced
datapathProvider accepts:
advanced— GKE Dataplane V2, an eBPF-based dataplane. This is Google’s recommended default for new clusters.legacy— the IPTables-based implementation built on kube-proxy.
Omitting datapathProvider leaves GKE’s default in place. This field is immutable once the cluster is created — GKE does not support switching datapath providers on an existing cluster. It also cannot be set when enableAutopilot is true: Autopilot clusters always use Dataplane V2.
Cluster DNS
dnsConfig chooses which DNS provider serves in-cluster DNS records:
apiVersion: infrastructure.cluster.x-k8s.io/v1beta1
kind: GCPManagedControlPlane
metadata:
name: my-cluster
spec:
clusterNetwork:
dnsConfig:
clusterDNS: cloud-dns
clusterDNSScope: cluster
clusterDNSDomain: cluster.local
clusterDNS—platform(GKE’s default),cloud-dns, orkube-dns.clusterDNSScope—cluster(records resolvable from within the cluster only) orvpc(records resolvable from anywhere in the VPC). Only meaningful withclusterDNS: cloud-dns.clusterDNSDomain— the suffix used for cluster service records.
dnsConfig is mutable — you can change its contents on an existing cluster. However, once set it cannot be removed entirely; there’s no clean way to revert to “no DNS config” once GKE has one configured.
ClusterClass
- Feature status: Experimental
- Feature gate:
ClusterTopology=true
ClusterClass is a collection of templates that define a topology (control plane and machine deployments) to be used to continuously reconcile one or more Clusters. It is built on top of the existing Cluster API resources and provides a set of tools and operations to streamline cluster lifecycle management while maintaining the same underlying API.
CAPG supports the creation of clusters via Cluster Topology for self-managed clusters only.
Provisioning a Cluster via ClusterClass
This guide uses an example from the ./templates folder of the CAPG repository. You can inspect the yaml file for the ClusterClass here and the cluster definition here.
Templates and clusters
ClusterClass makes cluster templates more flexible and versatile as it allows users to create cluster flavors that can be reused for cluster provisioning.
In this case, while inspecting the sample files, you probably noticed that there are references to two different yaml:
./templates/cluster-template-clusterclass.yamlis the class definition. It represents the template that define a topology: control plane and machine deployment but it won’t provision the cluster../templates/cluster-template-topology.yamlis the cluster definition that references the class. This workload cluster definition is considerably simpler than a regular CAPI cluster template that does not use ClusterClass, as most of the complexity of defining the control plane and machine deployment has been removed by the class.
Configure ClusterClass
While inspecting the templates you probably noticed that they contain a number of parameterized values that must be substituted with the specifics of your use case. This can be done via environment variables and clusterctl and effectively make the templates more flexible to adapt to different provisioning scenarios. These are the environment variables that you’ll be required to set before deploying a class and a workload cluster from it:
export CLUSTER_CLASS_NAME=sample-cc
export GCP_PROJECT=cluster-api-gcp-project
export GCP_REGION=us-east4
export GCP_NETWORK_NAME=default
export IMAGE_ID=projects/cluster-api-gcp-project/global/images/your-image
Generate ClusterClass definition
The sample ClusterClass template is already prepared so that you can use it with clusterctl to create a CAPI ClusterClass with CAPG.
clusterctl generate cluster capi-gcp-quickstart-clusterclass --flavor clusterclass -i gcp > capi-gcp-quickstart-clusterclass.yaml
In this example, capi-gcp-quickstart-clusterclass will be used as class name.
Create ClusterClass
The resulting file represents the class template definition and you simply need to apply it to your cluster to make it available in the API:
kubectl apply -f capi-gcp-quickstart-clusterclass.yaml
Create a cluster from a class
ClusterClass is a powerful feature of CAPI because we can now create one or multiple clusters that are based on the same class that is available in the CAPI Management Cluster. This base template can be parameterized so clusters created from it can make slight changes to the original configuration and adapt to the specifics of the use case, e.g. provisioning clusters for different development, staging and production environments.
Now that the class is available to be referenced by cluster objects, let’s configure the workload cluster and provision it.
export CLUSTER_NAME=sample-cluster
export CLUSTER_CLASS_NAME=sample-cc
export KUBERNETES_VERSION=1.29.3
export CONTROL_PLANE_MACHINE_COUNT=1
export WORKER_MACHINE_COUNT=1
export GCP_REGION=us-east4
export GCP_CONTROL_PLANE_MACHINE_TYPE=n1-standard-2
export GCP_NODE_MACHINE_TYPE=n1-standard-2
export CNI_RESOURCES=./cni-resource
You can take a look at CAPG’s CNI requirements here
You can use clusterctl to create a cluster definition.
clusterctl generate cluster capi-gcp-quickstart-topology --flavor topology -i gcp > capi-gcp-quickstart-topology.yaml
And by simply applying the resulting template, the cluster will be provisioned based on the existing ClusterClass.
kubectl apply -f capi-gcp-quickstart-topology.yaml
You can now experiment with creating more clusters based on this class while applying different configurations to each workload cluster.
Enabling ClusterClass Support
Enabling ClusterClass support is done via the ClusterTopology feature flag by setting it to true. This can be done before running clusterctl init by using the CLUSTER_TOPOLOGY environment variable:
export CLUSTER_TOPOLOGY=true
clusterctl init --infrastructure gcp
Disabling ClusterClass Support
Support for ClusterClass is disabled by default when you use the GCP infrastructure provider.
General Topics
This section contains information about relevant CAPG features and how to use them.
Additional Labels
The additionalLabels field lets you attach arbitrary GCP resource labels to
infrastructure objects managed by CAPG.
Supported resources
| CRD | Labels applied to |
|---|---|
GCPCluster | Load balancer forwarding rules, disks |
GCPMachine | Compute Engine instances and their root disks |
GCPManagedCluster | GKE cluster (ResourceLabels) |
GCPManagedMachinePool | GKE node pool (ResourceLabels) |
GCPMachinePool | Managed instance group instances |
Usage
Set additionalLabels under spec on the relevant infrastructure object.
Label keys and values must conform to
GCP label requirements.
apiVersion: infrastructure.cluster.x-k8s.io/v1beta1
kind: GCPManagedCluster
metadata:
name: my-cluster
spec:
additionalLabels:
env: production
team: platform
Precedence
When both a GCPCluster/GCPManagedCluster and a machine-level object
(GCPMachine, GCPMachinePool) define the same key, the machine-level
value takes precedence.
Semantics for GKE clusters
For GCPManagedCluster, label management is opt-in. The controller only
reconciles ResourceLabels on the GKE cluster when additionalLabels is
explicitly set. Clusters with no additionalLabels field are left untouched,
so any labels applied directly in GCP are not disturbed.
Once opted in, additionalLabels is treated as the complete desired label
set. The GKE SetLabels API replaces all user-defined ResourceLabels, so
labels applied outside of CAPI that are not present in the spec will be
removed on the next reconcile. This is consistent with CAPI’s source-of-truth
approach.
Label changes to an existing cluster are applied on the next reconcile cycle after any pending cluster updates have completed, since GKE does not permit concurrent cluster operations.
Alias IP Ranges
Configure secondary IP ranges for instances via the aliasIPRanges field in GCPMachineTemplate.
This enables CNI plugins like Cilium to use Native Routing by allocating pod and service IPs from the alias ranges.
---
apiVersion: infrastructure.cluster.x-k8s.io/v1beta1
kind: GCPMachineTemplate
metadata:
name: mygcpmachinetemplate
namespace: mynamespace
spec:
template:
spec:
image: projects/myproject/global/images/myimage
instanceType: n1-standard-2
aliasIPRanges:
- ipCidrRange: /24
subnetworkRangeName: pods
- ipCidrRange: 10.96.0.0/16
subnetworkRangeName: services
The ipCidrRange accepts:
- CIDR notation:
10.0.0.0/24 - IP address only:
10.0.0.1 - Netmask only:
/24
The subnetworkRangeName is optional and references a secondary IP range configured on the subnet.
https://cloud.google.com/vpc/docs/alias-ip
Autoscaling from Zero
Overview
CAPG supports autoscaling from zero replicas by populating Status.Capacity and Status.NodeInfo on GCPMachineTemplate. This enables cluster-autoscaler to scale NodePools from 0 replicas without requiring existing nodes.
This follows the CAPI autoscaling-from-zero proposal and matches the implementations in CAPA (AWS) and CAPZ (Azure).
Key benefits:
- Cost optimization by scaling unused node pools to zero
- Efficient resource utilization for dev/test environments
- Support for batch workloads that scale between job runs
When do I use autoscaling from zero?
Autoscaling from zero is useful when you want to:
- Reduce costs by scaling NodePools down to zero replicas when not in use
- Let cluster-autoscaler automatically create nodes when workloads need them
- Support dynamic workloads that may need specialized node pools (high-memory, high-CPU) only occasionally
How It Works
CAPG’s GCPMachineTemplate controller automatically populates status fields when a template is created or reconciled:
- The controller queries the GCP Compute API for machine type specifications
- It extracts capacity information (CPU cores, memory) from the machine type
- It determines node architecture (amd64/arm64) from the machine type’s CPU platform
- It queries the GCP Images API to detect the operating system (linux/windows) from image metadata
- This information is written to
status.capacityandstatus.nodeInfofields
The cluster-autoscaler reads these status fields to simulate node capacity for pending pods, enabling scale-from-zero decisions without requiring actual nodes to exist.
The controller respects cluster pause annotations and requires the template to have an owner reference to a Cluster resource.
Example
After creating a GCPMachineTemplate, CAPG automatically populates the status:
apiVersion: infrastructure.cluster.x-k8s.io/v1beta1
kind: GCPMachineTemplate
metadata:
name: worker-node-pool
spec:
template:
spec:
instanceType: n1-standard-2
imageFamily: ubuntu-2004-lts
imageProject: gke-node-images
status:
capacity:
cpu: "2"
memory: 7680Mi
nodeInfo:
architecture: amd64
operatingSystem: linux
No manual configuration needed — CAPG queries GCP and populates these values automatically based on the machine type and image specified in .spec.template.spec.
Status Fields
The CAPG controller populates the following fields in GCPMachineTemplate status:
| Field | Description | Example | Source |
|---|---|---|---|
status.capacity.cpu | Number of vCPUs | "2", "4", "96" | GCP Compute API (MachineTypes) |
status.capacity.memory | Memory size | "7680Mi", "16Gi" | GCP Compute API (MachineTypes) |
status.nodeInfo.architecture | CPU architecture | amd64, arm64 | GCP Compute API (CPU platform) |
status.nodeInfo.operatingSystem | OS type | linux, windows | GCP Images API (image metadata) |
Inspect the status of a GCPMachineTemplate:
kubectl get gcpmachinetemplate worker-node-pool -o jsonpath='{.status}' | jq
Example output:
{
"capacity": {
"cpu": "2",
"memory": "7680Mi"
},
"nodeInfo": {
"architecture": "amd64",
"operatingSystem": "linux"
}
}
Using with Cluster Autoscaler
Prerequisites
Cluster-autoscaler requires:
-
RBAC permissions in the management cluster:
- Read/write access to CAPI resources (
MachineDeployment,MachineSet,Machine) - Read access to infrastructure templates (
GCPMachineTemplate) for scale-from-zero - See RBAC setup for scale-from-zero
- Read/write access to CAPI resources (
-
Workload cluster kubeconfig:
- Cluster-autoscaler needs access to both management cluster (for CAPI resources) and workload cluster (for pod scheduling)
- See kubeconfig configuration for setup options
Configuration
To enable scale-from-zero:
- Deploy cluster-autoscaler in your management cluster with
--cloud-provider=clusterapi - Configure autoscaler to target your MachineDeployment (see example below)
- Set MachineDeployment replicas to 0 (or allow autoscaler to scale down)
- When pods are unschedulable, autoscaler reads
GCPMachineTemplate.Status.Capacityand scales up
Example MachineDeployment configuration:
apiVersion: cluster.x-k8s.io/v1beta1
kind: MachineDeployment
metadata:
name: worker-md-0
annotations:
cluster.x-k8s.io/cluster-api-autoscaler-node-group-min-size: "0"
cluster.x-k8s.io/cluster-api-autoscaler-node-group-max-size: "10"
spec:
clusterName: my-cluster
replicas: 0
template:
spec:
infrastructureRef:
apiVersion: infrastructure.cluster.x-k8s.io/v1beta1
kind: GCPMachineTemplate
name: worker-node-pool
The annotations tell cluster-autoscaler the min/max bounds. With replicas: 0, the pool starts at zero and autoscaler scales up when needed.
Supported Machine Types
All GCP machine types are supported — standard, high-memory, high-CPU, ARM (T2A/T2D), and custom machine types. CAPG queries the GCP Compute API for each machine type to get accurate CPU/memory values.
For ARM-based machine types (e.g., t2a-standard-2), Status.NodeInfo.Architecture is automatically set to arm64.
Supported Operating Systems
CAPG detects the operating system from the image specified in GCPMachineTemplate.Spec.Template.Spec.Image:
- Linux — Ubuntu, Debian, COS, RHEL, Fedora, Rocky Linux (default)
- Windows — Windows Server images
The OS is detected from GCP image metadata and populated in Status.NodeInfo.OperatingSystem.
Related Resources
- Cluster API Autoscaling - Cluster API autoscaling overview
- Autoscaling from Zero Proposal - CAPI proposal for scale-from-zero
- Kubernetes Cluster Autoscaler - Cluster autoscaler documentation
- Cluster Autoscaler with Cluster API - Using cluster-autoscaler with CAPI
- GCP Machine Types - Available GCP machine types and specifications
Running Conformance tests
Required environment variables
- Set the GCP region
export GCP_REGION=us-east4
- Set the GCP project to use
export GCP_PROJECT=your-project-id
- Set the path to the service account
export GOOGLE_APPLICATION_CREDENTIALS=path/to/your/service-account.json
Optional environment variables
- Set a specific name for your cluster
export CLUSTER_NAME=test1
- Set a specific name for your network
export NETWORK_NAME=test1-mynetwork
- Skip cleaning up the project resources
export SKIP_CLEANUP=1
Running the conformance tests
scripts/ci-conformance.sh
Firewall Rules
Cluster API Provider GCP (CAPG) allows you to configure GCP VPC firewall rules for your clusters through the GCPCluster and GCPManagedCluster (GKE) resources. This feature provides fine-grained control over network access to your cluster infrastructure.
Overview
Firewall rules are configured through the network.firewall field in the GCPCluster or GCPManagedCluster spec. The firewall configuration supports:
- Default rule management: Control whether the provider creates default firewall rules
- Custom firewall rules: Define additional firewall rules to meet your specific security requirements
Lifecycle Behavior
CAPG creates firewall rules during cluster provisioning and deletes them when the cluster is deleted. On every reconcile loop it also compares each rule in the spec against the rule that exists in GCP and updates the rule in place when they differ, so changes to the spec are rolled out automatically. Because the rule in GCP is replaced with the spec, any out-of-band change made directly in GCP to a rule that CAPG manages is reverted on the next reconcile.
CAPG matches rules by name, and every rule it reconciles is recorded in status.network.firewallRules. A recorded rule that the spec no longer asks for is deleted on the next reconcile, so removing a rule from the spec removes it from GCP, and renaming one deletes the rule that carried the old name. A rule whose name does not match any rule in the spec is never recorded and never touched, so you are free to manage your own rules in the same network by hand.
A rule that omits name is named after its contents, and that name is written back into the spec by the defaulting webhook. Editing the rule afterwards therefore changes its contents but not its name, and it is updated in place like any other rule. Clusters owned by a ClusterClass topology are the exception: the webhook leaves their spec untouched, because the topology controller strips anything it writes on its next apply, so the name is derived again on every reconcile. For those clusters, editing the contents of an unnamed rule changes its name, and CAPG creates the rule under the new name and deletes the one under the old. Give the rule an explicit name if you want it updated in place instead.
Warning: Because rules are matched by name, a rule that already exists in GCP is adopted as soon as a rule in the spec resolves to the same name — CAPG does not check who created it. Remember that CAPG prefixes rule names with the cluster name, so a spec rule named
webmatches an existing<cluster-name>-web. Once adopted, the rule is recorded in the status, overwritten to match the spec, and deleted when it leaves the spec or the cluster is deleted. Do not point the spec at a pre-existing rule you intend to keep managing yourself.
Because the spec is the desired state, a rule that is deleted directly in GCP is recreated on the next reconcile. To remove a rule for good, remove it from the spec.
The network a rule belongs to cannot be changed in place; it is not compared and is left untouched.
Upgrading from CAPG v1.13 and Earlier
Rules that omitted name used to be named <cluster-name>-ingress or <cluster-name>-egress, one per direction no matter how many rules the spec held. They are now named after their contents, so the first reconcile after the upgrade creates each rule under its own name and deletes the rule left behind under the old one.
Immutable Fields
GCP cannot modify the name, the network, the direction of traffic or the action on match (allowed versus denied) of an existing firewall rule. Since CAPG updates a rule in place as long as it keeps its name, changing direction or switching between allowed and denied on a named rule is rejected by the webhook: the reconciler would otherwise retry an update GCP always refuses. To make one of those changes, rename the rule. The rule under the old name is deleted and the new one is created in its place.
Rules that omit name are exempt, because their generated name is derived from their contents: changing the direction of an unnamed rule already changes its name, which replaces the rule instead of updating it.
Default Firewall Rules
By default, the provider creates two firewall rules to enable cluster functionality:
-
Health Check Rule (
allow-<cluster-name>-healthchecks):- Allows TCP ingress traffic on port 6443 from GCP health check IP ranges
- Source ranges:
35.191.0.0/16,130.211.0.0/22 - Target: control-plane nodes
-
Cluster Internal Rule (
allow-<cluster-name>-cluster):- Allows all ingress traffic between cluster nodes
- Source/Target: control-plane and worker nodes
Managing Default Rules
You can control the creation of default firewall rules using the defaultRulesManagement field:
apiVersion: infrastructure.cluster.x-k8s.io/v1beta1
kind: GCPCluster
metadata:
name: my-cluster
spec:
network:
firewall:
defaultRulesManagement: "Managed" # or "Unmanaged"
Values:
Managed(default): CAPG creates and manages default firewall rulesUnmanaged: CAPG does not create or manage default firewall rules
Important Notes:
- CAPG allows switching from
ManagedtoUnmanagedand vice versa. Switching fromUnmanagedtoManagedmakes CAPG create the default rules if they don’t already exist. Switching fromManagedtoUnmanageddeletes the default rules CAPG already created: they are recorded instatus.network.firewallRules, and every recorded rule the spec no longer asks for is pruned on the next reconcile. SetUnmanagedfrom the start if you want to own rules with those names yourself - GCPManagedCluster (GKE): The
defaultRulesManagementfield is ignored. CAPG does not create default firewall rules for GKE clusters because the default rules target CAPI-specific instance tags (e.g.<cluster>-control-plane,<cluster>-node) that do not exist on GKE nodes. Only custom firewall rules defined infirewallRulesare reconciled - When using a shared VPC (
HostProject), CAPG will not create, modify, or delete any firewall rules (default or custom)
Custom Firewall Rules
You can define up to 50 additional firewall rules using the firewallRules field:
apiVersion: infrastructure.cluster.x-k8s.io/v1beta1
kind: GCPCluster
metadata:
name: my-cluster
spec:
network:
firewall:
defaultRulesManagement: "Managed"
firewallRules:
- name: "custom-ingress-rule"
description: "Allow SSH and custom application traffic"
direction: "Ingress"
priority: 1000
allowed:
- IPProtocol: "TCP"
ports:
- "22"
- "8080"
- "8443"
sourceRanges:
- "10.0.0.0/8"
- "172.16.0.0/12"
targetTags:
- "web-servers"
FirewallRule Fields
| Field | Type | Required | Description |
|---|---|---|---|
name | string | Optional | Rule name (1-63 chars, must match [a-z]([-a-z0-9]*[a-z0-9])?). If not prefixed with cluster name, it will be prepended automatically. When omitted, the name is the cluster name followed by a suffix derived from the rule, truncated to 63 characters. |
description | string | Optional | Description of the rule (max 2000 chars). Defaults to “Created by Cluster API GCP Provider”. |
direction | string | Optional | Traffic direction: Ingress (default) or Egress. |
priority | integer | Optional | Rule priority (1-65535). Lower values = higher priority. Defaults to 1000 when omitted. |
allowed | []FirewallDescriptor | Optional | List of ALLOW rules (max 1024). Cannot be set together with denied. |
denied | []FirewallDescriptor | Optional | List of DENY rules (max 1024). Cannot be set together with allowed. |
sourceRanges | []string | Optional | Source IP ranges in CIDR format (max 1024). Supports IPv4 and IPv6. Not valid for Egress rules. |
sourceTags | []string | Optional | Source instance tags (max 30, 1-63 chars each). Only applies to traffic between instances in the same VPC. Not valid for Egress rules. |
destinationRanges | []string | Optional | Destination IP ranges in CIDR format (max 1024). Only valid for Egress rules. |
targetTags | []string | Optional | Target instance tags (max 70, 1-63 chars each). If empty, rule applies to all instances. |
FirewallDescriptor Fields
| Field | Type | Required | Description |
|---|---|---|---|
IPProtocol | string | Yes | Protocol: TCP, UDP, ICMP, ESP, AH, IPIP, or SCTP. |
ports | []string | Optional | Port numbers or ranges (e.g., ["22"], ["80","443"], ["12345-12349"]). Only applicable for TCP/UDP. Max 500 entries. |
Best Practices
- Use Descriptive Names: Choose meaningful names that describe the rule’s purpose
- Set Appropriate Priorities: Use lower values (higher priority) for more critical security rules
- Minimize Source Ranges: Restrict access to only necessary IP ranges
- Use Tags Strategically: Leverage instance tags for flexible rule targeting
- Document Rules: Always include a description explaining the rule’s purpose
- Test Before Production: Verify firewall rules in a development environment first
- Avoid Priority 65535: GCP reserves this priority for implied rules
- Consider DENY Rules: Use DENY rules for explicit blocking with higher priority
Examples
Example 1: Allow SSH from Specific IP Range
apiVersion: infrastructure.cluster.x-k8s.io/v1beta1
kind: GCPCluster
metadata:
name: my-cluster
spec:
network:
firewall:
firewallRules:
- name: "allow-ssh"
description: "Allow SSH from office network"
direction: "Ingress"
priority: 900
allowed:
- IPProtocol: "TCP"
ports:
- "22"
sourceRanges:
- "203.0.113.0/24"
targetTags:
- "ssh-enabled"
Example 2: Allow Application Traffic
apiVersion: infrastructure.cluster.x-k8s.io/v1beta1
kind: GCPCluster
metadata:
name: my-cluster
spec:
network:
firewall:
firewallRules:
- name: "web-traffic"
description: "Allow HTTP and HTTPS"
direction: "Ingress"
priority: 1000
allowed:
- IPProtocol: "TCP"
ports:
- "80"
- "443"
sourceRanges:
- "0.0.0.0/0"
targetTags:
- "web-servers"
Example 3: Deny Rule with Higher Priority
apiVersion: infrastructure.cluster.x-k8s.io/v1beta1
kind: GCPCluster
metadata:
name: my-cluster
spec:
network:
firewall:
firewallRules:
- name: "deny-telnet"
description: "Block telnet traffic"
direction: "Ingress"
priority: 500
denied:
- IPProtocol: "TCP"
ports:
- "23"
sourceRanges:
- "0.0.0.0/0"
Example 4: Egress Rule for External Services
apiVersion: infrastructure.cluster.x-k8s.io/v1beta1
kind: GCPCluster
metadata:
name: my-cluster
spec:
network:
firewall:
firewallRules:
- name: "allow-external-api"
description: "Allow egress to external API"
direction: "Egress"
priority: 1000
allowed:
- IPProtocol: "TCP"
ports:
- "443"
destinationRanges:
- "198.51.100.0/24"
targetTags:
- "api-client"
Example 5: Multiple Protocols and Ports
apiVersion: infrastructure.cluster.x-k8s.io/v1beta1
kind: GCPCluster
metadata:
name: my-cluster
spec:
network:
firewall:
firewallRules:
- name: "multi-service"
description: "Allow multiple services"
direction: "Ingress"
priority: 1000
allowed:
- IPProtocol: "TCP"
ports:
- "80"
- "443"
- "8080-8090"
- IPProtocol: "UDP"
ports:
- "53"
- IPProtocol: "ICMP"
sourceRanges:
- "10.0.0.0/8"
targetTags:
- "multi-service-node"
Example 6: GKE Managed Cluster with Firewall Rules
Note:
defaultRulesManagementis ignored forGCPManagedCluster. Only custom firewall rules infirewallRulesare reconciled.
apiVersion: infrastructure.cluster.x-k8s.io/v1beta1
kind: GCPManagedCluster
metadata:
name: my-gke-cluster
spec:
project: my-gcp-project
region: us-central1
network:
name: gke-network
firewall:
firewallRules:
- name: "allow-nodeport-range"
description: "Allow NodePort traffic to GKE nodes"
direction: "Ingress"
priority: 1000
allowed:
- IPProtocol: "TCP"
ports:
- "30000-32767"
sourceRanges:
- "10.0.0.0/8"
targetTags:
- "gke-node"
Complete Example with Default Rule Management
apiVersion: infrastructure.cluster.x-k8s.io/v1beta1
kind: GCPCluster
metadata:
name: production-cluster
spec:
project: my-gcp-project
region: us-central1
network:
name: production-network
firewall:
# Manage default health check and cluster internal rules
defaultRulesManagement: "Managed"
# Add custom rules
firewallRules:
# Allow monitoring from Prometheus
- name: "monitoring"
description: "Allow Prometheus scraping"
direction: "Ingress"
priority: 900
allowed:
- IPProtocol: "TCP"
ports:
- "9090"
- "9100"
sourceRanges:
- "10.128.0.0/16"
targetTags:
- "prometheus-target"
# Allow database access from application tier
- name: "database-access"
description: "Allow app to database communication"
direction: "Ingress"
priority: 1000
allowed:
- IPProtocol: "TCP"
ports:
- "5432"
- "3306"
sourceTags:
- "app-tier"
targetTags:
- "database-tier"
Troubleshooting
Rules Not Applied
The controller logs each firewall rule it creates, updates, skips, or fails on, along with the reason. Start there:
kubectl -n capg-system logs deploy/capg-controller-manager | grep -i firewall
Failures are also recorded as Warning events on the cluster object (kubectl describe gcpcluster my-cluster).
If the logs don’t explain it:
- Check that
defaultRulesManagementis set toManagedif you expect default rules - Verify you’re not using a shared VPC (HostProject), which disables custom firewall rules
- Ensure rule names are unique and follow GCP naming conventions. CAPG prepends the cluster name and truncates to 63 characters, so two long names can collide into one rule
- Compare
spec.network.firewall.firewallRuleswithstatus.network.firewallRules, which lists the rules CAPG manages - Check GCP Cloud Console for any conflicting VPC firewall rules
Connection Issues
If experiencing connection problems:
- Verify source/destination ranges include the correct IP addresses
- Check that priority values don’t conflict with DENY rules
- Ensure target tags match your instance tags
- Review that the correct protocol and ports are specified
- Check GCP VPC firewall logs for blocked traffic
Validation Errors
Common validation errors:
- Invalid CIDR: Ensure IP ranges use valid CIDR notation
- Name too long: Rule names are limited to 63 characters
- Too many rules: Maximum 50 custom rules per cluster
- Invalid port format: Use format like
"80","443", or"8080-8090"
Gateway API
Configure GKE’s managed Gateway API controller via the gatewayAPIChannel field on GCPManagedControlPlane’s clusterNetwork.
apiVersion: infrastructure.cluster.x-k8s.io/v1beta1
kind: GCPManagedControlPlane
metadata:
name: mygcpmanagedcontrolplane
spec:
clusterNetwork:
gatewayAPIChannel: standard
gatewayAPIChannel accepts:
standard— enables the GKE-managed Gateway API controller on the standard release channel.disabled— disables the Gateway API controller.
Omitting gatewayAPIChannel leaves GKE’s default behavior in place. The field is mutable and can be changed on an existing cluster.
gatewayAPIChannel cannot be set when enableAutopilot is true, since Autopilot clusters manage Gateway API themselves.
GPUs
Add GPUs via the guestAccelerators field in GCPMachineTemplate.
---
apiVersion: infrastructure.cluster.x-k8s.io/v1beta1
kind: GCPMachineTemplate
metadata:
name: mygcpmachinetemplate
namespace: mynamespace
spec:
template:
spec:
image: projects/myproject/global/images/myimage
instanceType: n1-standard-2
guestAccelerators:
- type: projects/myproject/zones/us-central1-c/acceleratorTypes/nvidia-tesla-t4
count: 1
https://cloud.google.com/compute/docs/gpus
NOTE: Instances with accelerators/GPUs do NOT support live migration.
Therefore, the onHostMaintenance event is always TERMINATE.
https://cloud.google.com/compute/docs/instances/setting-vm-host-options
Machine Locations
This document describes how to configure the location of a CAPG cluster’s compute resources. By default, CAPG requires the user to specify a GCP region for the cluster’s machines by setting the GCP_REGION environment variable as outlined in the CAPI quickstart guide. The provider then picks a zone to deploy the control plane and worker nodes in and generates the according portions of the cluster’s YAML manifests.
It is possible to override this default behaviour and exercise more fine-grained control over machine locations as outlined in the rest of this document.
Control Plane Machine Location
Before deploying the cluster, add a failureDomains field to the spec of your GCPCluster definition, containing a list of allowed zones:
apiVersion: infrastructure.cluster.x-k8s.io/v1alpha4
kind: GCPCluster
metadata:
name: capi-quickstart
spec:
network:
name: default
project: cyberscan2
region: europe-west3
+ failureDomains:
+ - europe-west3-b
In this example configuration, only a single zone has been added, ensuring the control plane is provisioned in europe-west3-b.
Node Pool Location
Similar to the above, you can override the auto-generated GCP zone for your MachineDeployment, by changing the value of the failureDomain field at spec.template.spec.failureDomain:
apiVersion: cluster.x-k8s.io/v1alpha4
kind: MachineDeployment
metadata:
name: capi-quickstart-md-0
spec:
clusterName: capi-quickstart
# [...]
template:
spec:
# [...]
clusterName: capi-quickstart
- failureDomain: europe-west3-a
+ failureDomain: europe-west3-b
When combined like this, the above configuration effectively instructs CAPG to deploy the CAPI equivalent of a zonal GKE cluster.
Preemptible Virtual Machines
GCP Preemptible Virtual Machines allows user to run a VM instance at a much lower price when compared to normal VM instances.
Compute Engine might stop (preempt) these instances if it requires access to those resources for other tasks. Preemptible instances will always stop after 24 hours.
When do I use Preemptible Virtual Machines?
A Preemptible VM works best for applications or systems that distribute processes across multiple instances in a cluster. While a shutdown would be disruptive for common enterprise applications, such as databases, it’s hardly noticeable in distributed systems that run across clusters of machines and are designed to tolerate failures.
How do I use Preemptible Virtual Machines?
To enable a machine to be backed by Preemptible Virtual Machine, add preemptible option to GCPMachineTemplate and set it to True.
apiVersion: infrastructure.cluster.x-k8s.io/v1beta1
kind: GCPMachineTemplate
metadata:
name: capg-md-0
spec:
template:
spec:
instanceType: n1-standard-2
rootDeviceSize: 30
rootDeviceType: pd-standard
preemptible: true
Spot VMs
Spot VMs are the latest version of preemptible VMs.
To use a Spot VM instead of a Preemptible VM, add provisioningModel to GCPMachineTemplate and set it to Spot.
apiVersion: infrastructure.cluster.x-k8s.io/v1beta1
kind: GCPMachineTemplate
metadata:
name: capg-md-0
spec:
template:
spec:
instanceType: n1-standard-2
rootDeviceSize: 30
rootDeviceType: pd-standard
provisioningModel: Spot
NOTE: specifying preemptible: true and provisioningModel: Spot is equivalent to only provisioningModel: Spot. Spot takes priority.
GKE Managed Node Pools
For GKE clusters, interruptible capacity is configured directly on GCPManagedMachinePool using two distinct boolean fields:
spot— creates nodes as Spot VMs. Spot VMs can be reclaimed at any time and are the recommended choice for new workloads.preemptible— creates nodes as Preemptible VMs. Preemptible VMs have a maximum lifetime of 24 hours and are the legacy offering.
Both fields are immutable — they cannot be changed after the node pool is created.
Spot node pool
apiVersion: infrastructure.cluster.x-k8s.io/v1beta1
kind: GCPManagedMachinePool
metadata:
name: spot-nodepool
spec:
spot: true
Preemptible node pool
apiVersion: infrastructure.cluster.x-k8s.io/v1beta1
kind: GCPManagedMachinePool
metadata:
name: preemptible-nodepool
spec:
preemptible: true
Developer Guide
Everything you need to know about contributing to CAPG.
If you are new to the project and want to help but don’t know where to start, you can refer to the Cluster API contributing guide.
Developing Cluster API Provider GCP
Setting up
Base requirements
- Install go
- Get the latest patch version for go v1.18.
- Install jq
brew install jqon macOS.sudo apt install jqon Windows + WSL2.sudo apt install jqon Ubuntu Linux.
- Install gettext package
brew install gettext && brew link --force gettexton macOS.sudo apt install gettexton Windows + WSL2.sudo apt install gettexton Ubuntu Linux.
- Install KIND
GO111MODULE="on" go get sigs.k8s.io/kind@v0.14.0.
- Install Kustomize
brew install kustomizeon macOS.- install instructions on Windows + WSL2, Linux and macOS.
- Install Python 3.x, if neither is already installed.
- Install make.
brew install makeon MacOS.sudo apt install makeon Windows + WSL2.sudo apt install makeon Linux.
- Install timeout
brew install coreutilson macOS.
When developing on Windows, it is suggested to set up the project on Windows + WSL2 and the file should be checked out on as wsl file system for better results.
Get the source
git clone https://github.com/kubernetes-sigs/cluster-api-provider-gcp
cd cluster-api-provider-gcp
Get familiar with basic concepts
This provider is modeled after the upstream Cluster API project. To get familiar with Cluster API resources, concepts and conventions (such as CAPI and CAPG), refer to the Cluster API Book.
Dev manifest files
Part of running cluster-api-provider-gcp is generating manifests to run. Generating dev manifests allows you to test dev images instead of the default releases.
Dev images
Container registry
Any public container registry can be leveraged for storing cluster-api-provider-gcp container images.
CAPG Node images
In order to deploy a workload cluster you will need to build the node images to use, for that you can reference the image-builder project, also you can read the image-builder book
Please refer to the image-builder documentation in order to get the latest requirements to build the node images.
To build the node images for GCP: https://image-builder.sigs.k8s.io/capi/providers/gcp.html
Developing
Change some code!
Modules and Dependencies
This repository uses Go Modules to track vendor dependencies.
To pin a new dependency:
- Run
go get <repository>@<version> - (Optional) Add a replace statement in
go.mod
Makefile targets and scripts are offered to work with go modules:
make verify-moduleschecks whether go modules are out of date.make modulesrunsgo mod tidyto ensure proper vendoring.hack/ensure-go.shchecks that the Go version and environment variables are properly set.
Setting up the environment
Your environment must have the GCP credentials, check Authentication Getting Started
Tilt Requirements
Install Tilt:
brew install tilt-dev/tap/tilton macOS or Linuxscoop bucket add tilt-dev https://github.com/tilt-dev/scoop-bucket&scoop install tilton Windows
After the installation is done, verify that you have installed it correctly with: tilt version
Install Helm:
brew install helmon MacOSchoco install kubernetes-helmon Windows- Install instructions for Linux
As the project lacks a lot of feature for windows, it would be suggested to follow the above steps on Windows + WSL2 rather than Windows.
Using Tilt
Both of the Tilt setups below will get you started developing CAPG in a local kind cluster. The main difference is the number of components you will build from source and the scope of the changes you’d like to make. If you only want to make changes in CAPG, then follow CAPG instructions. This will save you from having to build all of the images for CAPI, which can take a while. If the scope of your development will span both CAPG and CAPI, then follow the CAPI and CAPG instructions.
Tilt for dev in CAPG
If you want to develop in CAPG and get a local development cluster working quickly, this is the path for you.
From the root of the CAPG repository, run the following to generate a tilt-settings.json file with your GCP
service account credentials:
$ cat <<EOF > tilt-settings.json
{
"kustomize_substitutions": {
"GCP_B64ENCODED_CREDENTIALS": "$(cat PATH_FOR_GCP_CREDENTIALS_JSON | base64 -w0)"
}
}
EOF
Set the following environment variables with the appropriate values for your environment:
$ export GCP_REGION="<GCP_REGION>" \
$ export GCP_PROJECT="<GCP_PROJECT>" \
$ export CONTROL_PLANE_MACHINE_COUNT=1 \
$ export WORKER_MACHINE_COUNT=1 \
# Make sure to use same kubernetes version here as building the GCE image
$ export KUBERNETES_VERSION=1.23.3 \
$ export GCP_CONTROL_PLANE_MACHINE_TYPE=n1-standard-2 \
$ export GCP_NODE_MACHINE_TYPE=n1-standard-2 \
$ export GCP_NETWORK_NAME=<GCP_NETWORK_NAME or default> \
$ export CLUSTER_NAME="<CLUSTER_NAME>" \
To build a kind cluster and start Tilt, just run:
make tilt-up
Alternatively, you can also run:
./scripts/setup-dev-enviroment.sh
It will setup the network, if you already setup the network you can skip this step for that just run:
./scripts/setup-dev-enviroment.sh --skip-init-network
By default, the Cluster API components deployed by Tilt have experimental features turned off.
If you would like to enable these features, add extra_args as specified in The Cluster API Book.
Once your kind management cluster is up and running, you can deploy a workload cluster.
To tear down the kind cluster built by the command above, just run:
make kind-reset
And if you need to cleanup the network setup you can run:
./scripts/setup-dev-enviroment.sh --clean-network
Tilt for dev in both CAPG and CAPI
If you want to develop in both CAPI and CAPG at the same time, then this is the path for you.
To use Tilt for a simplified development workflow, follow the instructions in the cluster-api repo. The instructions will walk you through cloning the Cluster API (CAPI) repository and configuring Tilt to use kind to deploy the cluster api management components.
you may wish to checkout out the correct version of CAPI to match the version used in CAPG
Note that tilt up will be run from the cluster-api repository directory and the tilt-settings.json file will point back to the cluster-api-provider-gcp repository directory. Any changes you make to the source code in cluster-api or cluster-api-provider-gcp repositories will automatically redeployed to the kind cluster.
After you have cloned both repositories, your folder structure should look like:
|-- src/cluster-api-provider-gcp
|-- src/cluster-api (run `tilt up` here)
After configuring the environment variables, run the following to generate your tilt-settings.json file:
cat <<EOF > tilt-settings.json
{
"default_registry": "${REGISTRY}",
"provider_repos": ["../cluster-api-provider-gcp"],
"enable_providers": ["gcp", "docker", "kubeadm-bootstrap", "kubeadm-control-plane"],
"kustomize_substitutions": {
"GCP_B64ENCODED_CREDENTIALS": "$(cat PATH_FOR_GCP_CREDENTIALS_JSON | base64 -w0)"
}
}
EOF
$REGISTRYshould be in the formatdocker.io/<dockerhub-username>
The cluster-api management components that are deployed are configured at the /config folder of each repository respectively. Making changes to those files will trigger a redeploy of the management cluster components.
Debugging
If you would like to debug CAPG you can run the provider with delve, a Go debugger tool. This will then allow you to attach to delve and troubleshoot the processes.
To do this you need to use the debug configuration in tilt-settings.json. Full details of the options can be seen here.
An example tilt-settings.json:
{
"default_registry": "gcr.io/your-project-name-her",
"provider_repos": ["../cluster-api-provider-gcp"],
"enable_providers": ["gcp", "kubeadm-bootstrap", "kubeadm-control-plane"],
"debug": {
"gcp": {
"continue": true,
"port": 30000,
"profiler_port": 40000,
"metrics_port": 40001
}
},
"kustomize_substitutions": {
"GCP_B64ENCODED_CREDENTIALS": "$(cat PATH_FOR_GCP_CREDENTIALS_JSON | base64 -w0)"
}
}
Once you have run tilt (see section below) you will be able to connect to the running instance of delve.
For vscode, you can use the a launch configuration like this:
{
"version": "0.2.0",
"configurations": [
{
"name": "Core CAPI Controller GCP",
"type": "go",
"request": "attach",
"mode": "remote",
"remotePath": "",
"port": 30000,
"host": "127.0.0.1",
"showLog": true,
"trace": "log",
"logOutput": "rpc"
}
]
}
Create a new configuration and add it to the “Debug” menu to configure debugging in GoLand/IntelliJ following these instructions.
Alternatively, you may use delve straight from the CLI by executing a command like this:
delve -a tcp://localhost:30000
Deploying a workload cluster
After your kind management cluster is up and running with Tilt, ensure you have all the environment variables set as described in Tilt for dev in CAPG, and deploy a workload cluster with the following:
make create-workload-cluster
To delete the cluster:
make delete-workload-cluster
Submitting PRs and testing
Pull requests and issues are highly encouraged! If you’re interested in submitting PRs to the project, please be sure to run some initial checks prior to submission:
Do make sure to set the GOOGLE_APPLICATION_CREDENTIALS environment variable with the path to your JSON file. Check out the this doc to generate the credential.
make lint # Runs a suite of quick scripts to check code structure
make test # Runs tests on the Go code
Executing unit tests
make test executes the project’s unit tests. These tests do not stand up a
Kubernetes cluster, nor do they have external dependencies.
Nightly Builds
Nightly builds are regular automated builds of the CAPG source code that occur every night.
These builds are generated directly from the latest commit of source code on the main branch.
Nightly builds serve several purposes:
- Early Testing: They provide an opportunity for developers and testers to access the most recent changes in the codebase and identify any issues or bugs that may have been introduced.
- Feedback Loop: They facilitate a rapid feedback loop, enabling developers to receive feedback on their changes quickly, allowing them to iterate and improve the code more efficiently.
- Preview of New Features: Users and can get a preview of upcoming features or changes by testing nightly builds, although these builds may not always be stable enough for production use.
Overall, nightly builds play a crucial role in software development by promoting user testing, early bug detection, and rapid iteration.
CAPG Nightly build jobs run in Prow.
Usage
To try a nightly build, you can download the latest built nightly CAPG manifests, you can find the available ones by executing the following command:
curl -sL -H 'Accept: application/json' "https://storage.googleapis.com/storage/v1/b/k8s-staging-cluster-api-gcp/o" | jq -r '.items | map(select(.name | startswith("components/nightly_main"))) | .[] | [.timeCreated,.mediaLink] | @tsv'
The output should look something like this:
2024-05-03T08:03:09.087Z https://storage.googleapis.com/download/storage/v1/b/k8s-staging-cluster-api-gcp/o/components%2Fnightly_main_2024050x?generation=1714723389033961&alt=media
2024-05-04T08:02:52.517Z https://storage.googleapis.com/download/storage/v1/b/k8s-staging-cluster-api-gcp/o/components%2Fnightly_main_2024050y?generation=1714809772486582&alt=media
2024-05-05T08:02:45.840Z https://storage.googleapis.com/download/storage/v1/b/k8s-staging-cluster-api-gcp/o/components%2Fnightly_main_2024050z?generation=1714896165803510&alt=media
Now visit the link for the manifest you want to download. This will automatically download the manifest for you.
Once downloaded you can apply the manifest directly to your testing CAPI management cluster/namespace (e.g. with kubectl), as the downloaded CAPG manifest will already contain the correct, corresponding CAPG nightly image reference.
Creating cluster without clusterctl
This document describes how to create a management cluster and workload cluster without using clusterctl. For creating a cluster with clusterctl, checkout our Cluster API Quick Start
For creating a Management cluster
-
Build required images by using the following commands:
docker build --tag=gcr.io/k8s-staging-cluster-api-gcp/cluster-api-gcp-controller:e2e .make docker-build-all
-
Set the required environment variables. For example:
export GCP_REGION=us-east4 export GCP_PROJECT=k8s-staging-cluster-api-gcp export CONTROL_PLANE_MACHINE_COUNT=1 export WORKER_MACHINE_COUNT=1 export KUBERNETES_VERSION=1.21.6 export GCP_CONTROL_PLANE_MACHINE_TYPE=n1-standard-2 export GCP_NODE_MACHINE_TYPE=n1-standard-2 export GCP_NETWORK_NAME=default export GCP_B64ENCODED_CREDENTIALS=$( cat /path/to/gcp_credentials.json | base64 | tr -d '\n' ) export CLUSTER_NAME="capg-test" export IMAGE_ID=projects/k8s-staging-cluster-api-gcp/global/images/cluster-api-ubuntu-2204-v1-27-3-nightly
You can check for other images to set the IMAGE_ID of your choice.
- Run
make create-management-clusterfrom root directory.
Jobs
This document provides an overview of our jobs running via Prow and Github actions.
Builds and tests running on the default branch
Legend
🟢 REQUIRED - Jobs that have to run successfully to get the PR merged.
Presubmits
Prow Presubmits:
-
🟢pull-cluster-api-provider-gcp-test
./scripts/ci-test.sh -
🟢pull-cluster-api-provider-gcp-build
../scripts/ci-build.sh -
🟢pull-cluster-api-provider-gcp-make
runner.sh./scripts/ci-make.sh -
🟢pull-cluster-api-provider-gcp-e2e-test
"BOSKOS_HOST"="boskos.test-pods.svc.cluster.local" ./scripts/ci-e2e.sh -
pull-cluster-api-provider-gcp-conformance-ci-artifacts
"BOSKOS_HOST"="boskos.test-pods.svc.cluster.local" ./scripts/ci-conformance.sh --use-ci-artifacts -
pull-cluster-api-provider-gcp-conformance
"BOSKOS_HOST"="boskos.test-pods.svc.cluster.local" ./scripts/ci-conformance.sh -
pull-cluster-api-provider-gcp-capi-e2e
"BOSKOS_HOST"="boskos.test-pods.svc.cluster.local" GINKGO_FOCUS="Cluster API E2E tests" ./scripts/ci-e2e.sh -
pull-cluster-api-provider-gcp-test-release-0-4
./scripts/ci-test.sh -
pull-cluster-api-provider-gcp-build-release-0-4
./scripts/ci-build.sh -
pull-cluster-api-provider-gcp-make-release-0-4
runner.sh./scripts/ci-make.sh -
pull-cluster-api-provider-gcp-e2e-test-release-0-4
"BOSKOS_HOST"="boskos.test-pods.svc.cluster.local" ./scripts/ci-e2e.sh -
pull-cluster-api-provider-gcp-make-conformance-release-0-4
"BOSKOS_HOST"="boskos.test-pods.svc.cluster.local" ./scripts/ci-conformance.sh --use-ci-artifactsGithub Presubmits Workflows:
-
Markdown-link-check
find . -name \*.md | xargs -I{} markdown-link-check -c .markdownlinkcheck.json {} -
🟢Lint-check
make lint
Postsubmits
Github Postsubmit Workflows:
- Code-coverage-check
make test-cover
Periodics
Prow Periodics:
- periodic-cluster-api-provider-gcp-build
runner.sh./scripts/ci-build.sh - periodic-cluster-api-provider-gcp-test
runner.sh./scripts/ci-test.sh - periodic-cluster-api-provider-gcp-make-conformance-v1alpha4
"BOSKOS_HOST"="boskos.test-pods.svc.cluster.local" ./scripts/ci-conformance.sh - periodic-cluster-api-provider-gcp-make-conformance-v1alpha4-k8s-ci-artifacts
"BOSKOS_HOST"="boskos.test-pods.svc.cluster.local" ./scripts/ci-conformance.sh --use-ci-artifacts - periodic-cluster-api-provider-gcp-conformance-v1alpha4
"BOSKOS_HOST"="boskos.test-pods.svc.cluster.local" ./scripts/ci-conformance.sh - periodic-cluster-api-provider-gcp-conformance-v1alpha4-k8s-ci-artifacts
"BOSKOS_HOST"="boskos.test-pods.svc.cluster.local" ./scripts/ci-conformance.sh --use-ci-artifacts
Adding new E2E test
E2E tests verify a complete, real-world workflow ensuring that all parts of the system work together as expected. If you are introducing a new feature that interconnects with other parts of the software, you will likely be required to add a verification step for this functionality with a new E2E scenario (unless it is already covered by existing test suites).
Create a cluster template
The test suite will provision a cluster based on a pre-defined yaml template (stored in ./test/e2e/data) which is then sourced in ./test/e2e/config/gcp-ci.yaml. New cluster definitions for E2E tests have to be added and sourced before being available to use in the E2E workflow.
Add test case
When the template is available, you can reference it as a flavor in Go. For example, adding a new test for self-managed cluster provisioning would look like the following:
Context("Creating a control-plane cluster with an internal load balancer", func() {
It("Should create a cluster with 1 control-plane and 1 worker node with an internal load balancer", func() {
By("Creating a cluster with internal load balancer")
clusterctl.ApplyClusterTemplateAndWait(ctx, clusterctl.ApplyClusterTemplateAndWaitInput{
ClusterProxy: bootstrapClusterProxy,
ConfigCluster: clusterctl.ConfigClusterInput{
LogFolder: clusterctlLogFolder,
ClusterctlConfigPath: clusterctlConfigPath,
KubeconfigPath: bootstrapClusterProxy.GetKubeconfigPath(),
InfrastructureProvider: clusterctl.DefaultInfrastructureProvider,
Flavor: "ci-with-internal-lb",
Namespace: namespace.Name,
ClusterName: clusterName,
KubernetesVersion: e2eConfig.MustGetVariable(KubernetesVersion),
ControlPlaneMachineCount: ptr.To[int64](1),
WorkerMachineCount: ptr.To[int64](1),
},
WaitForClusterIntervals: e2eConfig.GetIntervals(specName, "wait-cluster"),
WaitForControlPlaneIntervals: e2eConfig.GetIntervals(specName, "wait-control-plane"),
WaitForMachineDeployments: e2eConfig.GetIntervals(specName, "wait-worker-nodes"),
}, result)
})
})
In this case, the flavor ci-with-internal-lb is a reference to the template cluster-template-ci-with-internal-lb.yaml which is available in ./test/e2e/data/infrastructure-gcp/cluster-template-ci-with-internal-lb.yaml.
Release Process
Change milestone
- Create a new GitHub milestone for the next release
- Change milestone applier so new changes can be applied to the appropriate release
- Open a PR in https://github.com/kubernetes/test-infra to change this line
- Example PR: https://github.com/kubernetes/test-infra/pull/16827
- Open a PR in https://github.com/kubernetes/test-infra to change this line
Ensure that CI is stable
- Before releasing always ensure CI is stable. Check Prow CAPG dashboard
Ensure you have a GITHUB_TOKEN and are logged into GCP
-
If you don’t have a GitHub token, create one by going to your GitHub settings in Personal access tokens. Make sure you give the token the
reposcope. If you have one, make sure it has the right scope and it is not expired. -
Configure gcloud authentication. If you have multiple gcloud configurations, make sure to activate the one relevant to CAPG first:
# Only needed if you have multiple gcloud configurations gcloud config configurations list gcloud config configurations activate <your-capg-config> gcloud auth login <your-community-email-address>
Create the branch, new version tag, staging image
-
Please fork
https://github.com/kubernetes-sigs/cluster-api-provider-gcpand clone your own repository with e.g.git clone git@github.com:YourGitHubUsername/cluster-api-provider-gcp.git. kpromo uses the fork to build images from. -
Add a git remote to the upstream project. git remote add upstream
git@github.com:kubernetes-sigs/cluster-api-provider-gcp.git -
If this is a major or minor release, create a new release branch
git checkout -b release-1.7. Otherwise if it is a patch release the branch should already exist to it, fetch the latest and sync up your local branch: e.g.git fetch upstream release-1.7 && git checkout release-1.7 && git reset upstream/release-1.7 --hard. -
If this is a major or minor release, update
metadata.yamlby adding a new section with the version, and make a commit. -
Push new/existing release branch to your repository’s fork, e.g.
git push origin HEAD:release-1.7. origin refers to the remote git reference to your fork. -
Push new/existing release branch to the upstream repository, e.g.
git push upstream HEAD:release-1.7. upstream refers to the upstream git reference. -
Make sure your repo is clean by git standards.
-
Set the environment variable VERSION which is the current release that you are making, e.g.
export VERSION=v1.7.0, orexport VERSION=v1.7.1). Note: the version MUST contain a v in front. Note: you must have a gpg signing configured with git and registered with GitHub. -
Create a tag locally
git tag -s -a $VERSION -m $VERSION-s flag is for GNU Privacy Guard (GPG) signing. -
Make sure you have push permissions to the upstream CAPG repo. Push tag you’ve just created
git push <upstream-repo-remote> $VERSION.<upstream-repo-remote>must be the remote pointing togithub.com/kubernetes-sigs/cluster-api-provider-gcp. -
Pushing this will create the tag and this will automatically trigger a ProwJob to publish images to the staging image repository.
Generate release manifests and release notes
-
make releasefrom repo, this will create the release artifacts in theout/folder. It is recommended to verify that the artifact fileinfrastructure-components.yamlpoints to the new image. -
Install the
release-notestool according to instructions -
Generate release-notes (requires exported
GITHUB_TOKENvariable, ensure the TOKEN is not expired!):Commits range from the first commit after the previous non-beta release to the newest commit of the release branch. Set branch to the release branch you are cutting this release from. For example if this is release
v1.11.z, branch is going to berelease-1.11.First, compute the required variables. Review the output and verify the values are correct before proceeding:
source ./hack/compute-release-notes-vars.shOnce the values look correct, run the release-notes tool. When this finishes it will log the path to the temporary file where the notes have been written:
release-notes --org kubernetes-sigs --repo cluster-api-provider-gcp \ --start-rev "${PREVIOUS_TAG}" \ --end-rev "${VERSION}" \ --branch "${RELEASE_BRANCH}" \ --required-author "" \ --go-template "go-template:hack/release-notes.tpl"Note: The
--required-author ""flag is needed because the tool defaults to only processing commits authored byk8s-ci-robot, which would skip most PRs since CAPG uses squash merges that preserve the original author. -
Open the output temporary file logged by the tool and manually format and categorize the release notes.
Prepare release in GitHub
Create the GitHub release in the UI:
- Go to: https://github.com/kubernetes-sigs/cluster-api-provider-gcp/releases
- Create a draft release with the output from above in GitHub and associate it with the tag that was created
- Copy paste the release notes
- Upload artifacts from the
out/folder - Leave everything unchecked and click “Save Draft”
Promote image to prod repo
Images are built by the push images job after pushing a tag.
To promote images from the staging repository to the production registry (registry.k8s.io/cluster-api-gcp):
-
Wait until images for the tag have been built and pushed to the staging repository by the push images job.
-
If you don’t have a GitHub token, create one by going to your GitHub settings in Personal access tokens. Make sure you give the token the
reposcope. -
Create a PR to promote the images to the production registry:
# Export the tag of the release to be cut, e.g.: export USER_FORK=<personal GitHub handle> # if needed (see notes below) export VERSION=v1.11.0-beta.0 export GITHUB_TOKEN=<your GH token> make promote-imagesNotes:
make promote-imagestarget tries to figure out your Github user handle in order to find the forked k8s.io repository. If you have not forked the repo, please do it before running the Makefile target.- if
make promote-imagesfails with an error likeFATAL while checking fork of kubernetes/k8s.ioyou may be able to solve it by manually setting the USER_FORK variable i.e.export USER_FORK=<personal GitHub handle>. kpromousesgit@github.com:...as remote to push the branch for the PR. If you don’t havesshset up you can configure git to usehttpsinstead viagit config --global url."https://github.com/".insteadOf git@github.com:.- This will automatically create a PR in k8s.io and assign the CAPG maintainers.
-
Merge the PR (/lgtm + /hold cancel) and verify the images are available in the production registry: - Wait for the promotion prow job to complete successfully. Then verify that the production images are accessible:
docker pull registry.k8s.io/cluster-api-gcp/cluster-api-gcp-controller:${VERSION}
Location of image: https://console.cloud.google.com/gcr/images/k8s-staging-cluster-api-gcp/GLOBAL/cluster-api-gcp-controller?rImageListsize=30
Release in GitHub
Go back to the GitHub release in the UI
- Edit the draft release previously created in GitHub Releases
- Check the “Set as the latest release” checkbox
- Check the “Create a Discussion for this release” in Announcements checkbox
- ONLY CHECK the “Set as a pre-release” checkbox IF it is a
vXX.XX.XX-beta.Xrelease - Now, hit “Publish release”
- Announce the release
Versioning
cluster-api-provider-gcp follows the semantic versioning specification.
Example versions:
- Pre-release:
v0.1.1-alpha.1 - Minor release:
v0.1.0 - Patch release:
v0.1.1 - Major release:
v1.0.0
Expected artifacts
- A release yaml file
infrastructure-components.yamlcontaining the resources needed to deploy to Kubernetes - A
cluster-templates.yamlfor each supported flavor - A
metadata.yamlwhich maps release series to cluster-api contract version - Release notes
Communication
Patch Releases
- Announce the release in Kubernetes Slack on the #cluster-api-gcp channel.
Minor/Major Releases
- Follow the communications process for pre-releases
- An announcement email is sent to
sig-cluster-lifecycle@kubernetes.iowith the subject[ANNOUNCE] cluster-api-provider-gcp <version> has been released
Bumping Go
This document describes how to bump the Go version across the project. It is primarily intended to be consumed by an AI coding agent (e.g. via /bump-go 1.26), but the steps can also be followed manually.
Convention
godirective in go.mod: useX.Y.0(the minor version with patch 0)toolchaindirective in go.mod: usegoX.Y.Zwhere Z is the latest available patch
Step 1: Research
Perform these lookups before making any changes.
1a. Latest patch version
Find the latest Go X.Y.x patch release at https://go.dev/doc/devel/release.
Determine the full version string (e.g. 1.26.4). Call this FULL_VERSION.
1b. Docker image digest
Pull the official golang image and extract its digest:
docker pull golang:FULL_VERSION
docker inspect --format='{{index .RepoDigests 0}}' golang:FULL_VERSION
Extract the sha256:... digest. Call this GOLANG_DIGEST.
1c. Upstream cluster-api references
Look up what kubernetes-sigs/cluster-api uses at the latest release tag on the main CAPI minor version this project depends on (check go.mod for sigs.k8s.io/cluster-api to find the CAPI minor, then find the latest tag for that minor, e.g. v1.13.2):
# GCB image digest + tag comment
gh api "repos/kubernetes-sigs/cluster-api/contents/cloudbuild.yaml?ref=<TAG>" --jq '.content' | base64 -d
# golangci-lint version
gh api "repos/kubernetes-sigs/cluster-api/contents/.github/workflows/pr-golangci-lint.yaml?ref=<TAG>" --jq '.content' | base64 -d | grep 'version:'
Call these GCB_DIGEST, GCB_TAG_COMMENT, and CAPI_GOLANGCI_VER.
1d. golangci-lint compatibility
The golangci-lint version must be built with the target Go version or newer. Versions built with an older Go will refuse to lint. Test candidate versions:
go run github.com/golangci/golangci-lint/v2/cmd/golangci-lint@<MINOR_VERSION> version
Look for built with goX.Y in the output. Pick a stable minor release that resolves to a linter built with the target Go minor or newer. Call this GOLANGCI_MINOR (e.g. v2.12). Both CI and make lint use this minor-version selector, so its resolved patch version may change.
1e. Delve version
Delve major version should match the Go minor version (e.g. Go 1.26 → dlv v1.26). Verify at https://github.com/go-delve/delve/releases that a matching tag exists. If not, use the latest available. Call this DLV_VERSION.
Step 2: Update files
go.mod (root) and hack/tools/go.mod
go X.Y.0
toolchain goFULL_VERSION
Makefile
GOLANG_VERSION := FULL_VERSION
GOLANG_DIRECTIVE_VERSION ?= X.Y.0
GOLANGCI_LINT_VER is read from .github/workflows/lint.yml; do not set it separately in the Makefile.
Dockerfile
Update the FROM line, replacing both the tag and the digest:
FROM golang:FULL_VERSION@sha256:GOLANG_DIGEST AS builder
Tiltfile
Two FROM golang: lines (tilt-helper and tilt) and the delve install:
FROM golang:FULL_VERSION as tilt-helper
RUN go install github.com/go-delve/delve/cmd/dlv@DLV_VERSION
...
FROM golang:FULL_VERSION as tilt
netlify.toml
GO_VERSION = "FULL_VERSION"
.github/workflows/lint.yml
go-version: "X.Y"
...
version: GOLANGCI_MINOR
cloudbuild.yaml and cloudbuild-nightly.yaml
Update the GCB image digest and tag comment:
- name: 'gcr.io/k8s-staging-test-infra/gcb-docker-gcloud@sha256:GCB_DIGEST' # GCB_TAG_COMMENT
Step 3: go mod tidy
Run in both module directories:
go mod tidy
cd hack/tools && go mod tidy
Step 4: Lint and test
make lint
make test
If make lint fails:
- If the failure is
the Go language version used to build golangci-lint is lower than the targeted Go version, the chosenGOLANGCI_MINORresolves to an incompatible linter. Go back to Step 1d. - If the failure shows new lint findings, fix them. These are pre-existing issues surfaced by the newer linter version, not caused by the Go bump itself. Common categories:
- goconst in test files: add an exclusion in
.golangci.yml(upstream cluster-api excludes goconst from_test.go) - goconst in production code: extract repeated string literals into package-level constants
- prealloc: change
var s []Tors := []T{}tos := make([]T, 0, len(source)) - perfsprint: replace string concatenation in loops with
strings.Builder - staticcheck: follow the suggested fix (e.g.
fmt.Fprintfinstead ofWriteString(fmt.Sprintf))
- goconst in test files: add an exclusion in
Step 5: Verify build
go build ./...
Step 6: Commit
Create three separate commits in this order:
- Lint fixes (if any):
fix(lint): resolve <linter-names> lint issues- Only the source files with lint fixes
- Linter bump (if version changed):
chore(bump): bump golangci-lint from <old> to <new>.github/workflows/lint.yml(version only), and.golangci.ymlif its configuration changes
- Go version bump:
chore(bump): bump Go to FULL_VERSION- All remaining files: go.mod, hack/tools/go.mod, Makefile (GOLANG_VERSION + GOLANG_DIRECTIVE_VERSION), Dockerfile, Tiltfile, netlify.toml, cloudbuild*.yaml, .github/workflows/lint.yml (go-version only)
If there are no lint fixes or no linter version change, skip that commit.
Do NOT push or create a PR unless the user asks.
Bumping Kubernetes and Cluster API
This document describes how to bump the Kubernetes and Cluster API (CAPI) versions across the project. It is primarily intended to be consumed by an AI coding agent (e.g. via /bump-k8s-capi 1.35 1.13), but the steps can also be followed manually.
The two arguments are the target Kubernetes minor version (e.g. 1.35) and the target CAPI minor version (e.g. 1.13).
Convention
- Kubernetes module version:
v0.MINOR.PATCH(e.g.v0.35.4for k8s 1.35) - CAPI version:
v1.MINOR.PATCH(e.g.v1.13.2) - Infrastructure provider dev version:
v1.CAPI_MINOR.99(e.g.v1.13.99) - CCM version major matches the Kubernetes minor (e.g. k8s 1.35 → CCM
v35.x.y)
Step 1: Research
Perform these lookups before making any changes. All values determined here are referenced in later steps.
Kubernetes-version-derived variables in test/e2e/config/gcp-ci.yaml
(KUBERNETES_VERSION, KUBERNETES_VERSION_GKE, CCM_VERSION,
KUBERNETES_VERSION_MANAGEMENT, and the upgrade-test FROM/TO/etcd/coredns/
image variables) are not looked up here — they derive automatically at
CI time from KUBERNETES_MINOR via hack/resolve-e2e-versions.sh. The one
thing to check by hand before setting KUBERNETES_MINOR is that GKE
actually supports it yet (1f below) — that’s a precondition for the bump,
not something to route around if it doesn’t.
1a. Latest CAPI patch
gh api repos/kubernetes-sigs/cluster-api/tags --paginate -q '.[].name' | grep '^v1.MINOR\.' | head -5
Pick the latest stable tag. Call it CAPI_VERSION (e.g. v1.13.2).
1b. CAPI dependency versions
Fetch CAPI’s go.mod to determine aligned dependency versions:
gh api "repos/kubernetes-sigs/cluster-api/contents/go.mod?ref=CAPI_VERSION" --jq '.content' | base64 -d
Extract:
sigs.k8s.io/controller-runtimeversion →CONTROLLER_RUNTIME_VERk8s.io/apiversion →K8S_MODULE_VER(e.g.v0.35.4)
1c. setup-envtest version
Fetch CAPI’s Makefile to align the setup-envtest CLI version:
gh api "repos/kubernetes-sigs/cluster-api/contents/Makefile?ref=CAPI_VERSION" --jq '.content' | base64 -d | grep '^SETUP_ENVTEST_VER :='
Call this SETUP_ENVTEST_VER. The controller-gen and conversion-gen versions are selected by hack/tools/go.mod, not copied from CAPI’s Makefile.
1d. GCP k8s-cloud-provider version
gh api repos/GoogleCloudPlatform/k8s-cloud-provider/tags -q '.[].name' | head -5
Pick the latest matching the target k8s minor. Call it K8S_CLOUD_PROVIDER_VER.
1e. kind version
Kind follows the version selected by the root go.mod; do not choose a separate Makefile pin. After updating dependencies in Step 3, check the selected version with go list -m sigs.k8s.io/kind.
1f. Confirm GKE supports the target minor
gcloud container get-server-config --region=us-central1 --format=json | \
jq -r '.channels[] | select(.channel=="REGULAR") | .validVersions[]' | \
grep '^K8S_MINOR\.' | sort -V | tail -5
If nothing matches, GKE hasn’t caught up to this minor yet. Don’t bump
KUBERNETES_MINOR past it — wait for GKE to catch up instead. This
project bumps to the trailing edge of Kubernetes releases, not the
bleeding edge, so this should be rare; the more likely failure mode over
time is the opposite one (GKE eventually dropping support for a minor
this project sat on too long), which the same check catches just as well.
Step 2: Review CAPI migration guide
Read the upstream CAPI migration guide for the version jump being performed:
https://cluster-api.sigs.k8s.io/developer/providers/migrations/v1.OLD_CAPI_MINOR-to-v1.NEW_CAPI_MINOR
For example, for a jump from CAPI v1.12 to v1.13: https://cluster-api.sigs.k8s.io/developer/providers/migrations/v1.12-to-v1.13
Do NOT apply everything listed there blindly. Instead:
- Removals and API changes: these are mandatory. Fix anything that applies to this project — the build will likely fail otherwise.
- Deprecation, Cluster API Contract changes, and Suggested changes for providers: review these carefully. Evaluate whether each suggestion applies to this project and whether it makes sense to adopt now. Some may be best deferred to a follow-up.
Step 3: Update go.mod (root)
Update the direct dependencies:
sigs.k8s.io/cluster-api CAPI_VERSION
sigs.k8s.io/cluster-api/test CAPI_VERSION
sigs.k8s.io/controller-runtime CONTROLLER_RUNTIME_VER
k8s.io/api K8S_MODULE_VER
k8s.io/apimachinery K8S_MODULE_VER
k8s.io/client-go K8S_MODULE_VER
k8s.io/component-base K8S_MODULE_VER
github.com/GoogleCloudPlatform/k8s-cloud-provider K8S_CLOUD_PROVIDER_VER
Then run:
go mod tidy
Indirect dependencies, including Kind, will be resolved automatically.
Step 4: Update hack/tools/go.mod
Update the sigs.k8s.io/cluster-api/hack/tools pseudo-version. To find the right version:
GOPROXY=https://proxy.golang.org go list -m -json "sigs.k8s.io/cluster-api/hack/tools@CAPI_VERSION"
Update the setup-envtest tool pin to the version found in Step 1c, then run:
cd hack/tools
go get -tool sigs.k8s.io/controller-runtime/tools/setup-envtest@SETUP_ENVTEST_VER
go mod tidy
Step 5: Verify derived tool versions
Do not manually edit KUBEBUILDER_ENVTEST_KUBERNETES_VERSION, CONTROLLER_GEN_VER, CONVERSION_GEN_VER, KIND_VER, KUBECTL_VER, or SETUP_ENVTEST_VER in the Makefile. They follow the root or tools module:
go list -m k8s.io/client-go sigs.k8s.io/kind github.com/onsi/ginkgo/v2
(cd hack/tools && go list -m sigs.k8s.io/controller-tools k8s.io/code-generator sigs.k8s.io/controller-runtime/tools/setup-envtest)
The k8s.io/client-go version v0.MINOR.PATCH determines kubectl v1.MINOR.PATCH and the envtest 1.MINOR selector. Verify the kubectl binary and a matching envtest release are available. The Ginkgo CLI follows the root go.mod; controller-gen, conversion-gen, and setup-envtest follow hack/tools/go.mod. Check that their selected versions are appropriate before regenerating files.
Do NOT update the independent pins GOLANG_VERSION, KUSTOMIZE_VER, CERT_MANAGER_VER, or CALICO_VERSION as part of the k8s/CAPI bump. GOLANGCI_LINT_VER follows .github/workflows/lint.yml and is handled by the Go bump guide.
Step 6: Update metadata.yaml
Add a new release series entry at the end of the list for the new CAPG minor version:
- major: 1
minor: NEW_CAPI_MINOR
contract: v1beta1
Step 7: Update test/e2e/data/shared/v1beta1/metadata.yaml
Add a new release series entry at the top of the list (this file is ordered newest-first):
- major: 1
minor: NEW_CAPI_MINOR
contract: v1beta1
Step 8: Update test/e2e/config/gcp-ci.yaml
Provider versions
Update CAPI core, bootstrap, and control-plane provider versions and URLs from OLD_CAPI_VERSION to CAPI_VERSION.
Update the GCP infrastructure provider dev version from v1.OLD_CAPI_MINOR.99 to v1.NEW_CAPI_MINOR.99.
Variables
KUBERNETES_MINOR: "K8S_MINOR"
That’s the only line to touch here. Every other Kubernetes-version
variable in this file (KUBERNETES_VERSION, KUBERNETES_VERSION_GKE,
CCM_VERSION, KUBERNETES_VERSION_MANAGEMENT, the upgrade-test FROM/TO/
etcd/coredns/image variables) is already "${VAR}" with no default, and
derives automatically at CI time via hack/resolve-e2e-versions.sh — don’t
add a hand-pinned fallback to any of them; that would just recreate the
staleness problem this whole scheme exists to avoid.
Optionally, run hack/resolve-e2e-versions.sh locally first (needs
gcloud/docker/git access; E2E_FLAVOR=all GCP_PROJECT=... GCP_REGION=...)
to catch a missing nightly image, CCM tag, or kindest/node image before
pushing, rather than waiting on a full CI run to find out.
Step 9: Update CCM manifest
test/e2e/data/ccm/gce-cloud-controller-manager.yaml already references
${CCM_VERSION} with no default — nothing to change here either, for the
same reason as Step 8.
Step 10: Regenerate CRDs
Changes to the derived controller-gen or conversion-gen versions may update generated files:
make generate
make manifests
Step 11: Build and verify
go build ./...
If go build fails with API changes (e.g. breaking changes in controller-runtime or CAPI), fix the Go source files to match the new API using the migration guide from Step 2.
Step 12: Fix lint issues
make lint
If lint reports deprecation warnings (e.g. SA1019 for deprecated interfaces or types), fix what can be fixed (migrate to new APIs) and add .golangci.yml exclusions for deprecations that cannot be resolved yet (e.g. upstream CAPI types still using the deprecated form).
Step 13: Run tests
make test
Step 14: Branch and commit
Create a branch named bump-k8s-MINOR-capi-MINOR (e.g. bump-k8s-135-capi-113) and commit the changes as two separate commits:
-
The version bump itself:
chore(bump): bump k8s to K8S_MINOR, CAPI to CAPI_VERSION -
Lint and build fixes (if any):
fix(lint): <describe the migration or fix>
Do NOT push or create a PR unless the user asks. When creating a PR, follow the template in .github/PULL_REQUEST_TEMPLATE.md.
Bumping the Ubuntu Image
This document describes how to bump the Ubuntu version of the VM images used by the e2e and conformance jobs. It is primarily intended to be consumed by an AI coding agent (e.g. via /bump-ubuntu-image 2404), but the steps can also be followed manually.
Convention
The Ubuntu version lives in one place: CAPG_UBUNTU_VERSION in ../../../../test/e2e/config/gcp-ci.yaml. Its value is the suffix of the image-builder target, so "2204" means make build-gce-ubuntu-2204.
../../../../hack/resolve-e2e-versions.sh, scripts/ci-e2e.sh and scripts/ci-conformance.sh all read the pin, so none of them should be edited for a bump. Everything Kubernetes-related is bumped separately, see Bumping Kubernetes and Cluster API.
Let UBUNTU_VERSION be the target version (e.g. 2404).
Step 1: Research
Perform these lookups before making any changes.
1a. image-builder supports the target
gh api repos/kubernetes-sigs/image-builder/contents/images/capi/packer/gce --jq '.[].name' | grep ubuntu
gh api repos/kubernetes-sigs/image-builder/contents/images/capi/Makefile --jq .content | base64 -d | grep '^GCE_BUILD_NAMES'
ubuntu-UBUNTU_VERSION.json must exist and gce-ubuntu-UBUNTU_VERSION must be in GCE_BUILD_NAMES. If not, stop: there is nothing to bump to yet.
1b. Nightly images exist for the target
The e2e jobs consume prebuilt images from the k8s-staging-cluster-api-gcp project, built daily by the image-builder project’s nightly job. The job builds one set of images per config file in images/capi/packer/gce/ci/nightly (one per Kubernetes minor), using build-gce-all, so every target in GCE_BUILD_NAMES should be published.
A given Kubernetes patch only gets a nightly image if someone has added config for it there. ../../../../hack/resolve-e2e-versions.sh searches for the latest patch that has one; it can’t conjure a new one into existence.
Check that the target’s images exist for the current minor (KUBERNETES_MINOR in gcp-ci.yaml) and the minors used by the upgrade tests:
gcloud compute images list --project k8s-staging-cluster-api-gcp --no-standard-images \
--filter="name~cluster-api-ubuntu-UBUNTU_VERSION-.*-nightly" --format="value(name)"
If any minor the tests need is missing, the fix is in image-builder, not here. Don’t bump until it is published, or the resolver will fail with “no nightly image published for minor …”.
1c. Current version
hack/tools/bin/yq -e '.variables.CAPG_UBUNTU_VERSION' test/e2e/config/gcp-ci.yaml
Step 2: Update files
test/e2e/config/gcp-ci.yaml
CAPG_UBUNTU_VERSION: "UBUNTU_VERSION"
That’s the only line to touch for the e2e and conformance jobs.
Docs
These name the image directly and are not derived from the pin:
../prerequisites.md: themake build-gce-ubuntu-*target and thefamily:capi-ubuntu-*-k8sfiltercluster-creation.md: the exampleIMAGE_ID
Leave ../topics/autoscaling.md alone: its imageFamily is a Google public image family in an unrelated example.
Check nothing is left behind
rg -n 'ubuntu-?[0-9]{4}' scripts hack docs test --glob '!*.sum'
Every hit for the old version should be gone, apart from the unrelated autoscaling example above.
Step 3: Verify
shellcheck scripts/ci-e2e.sh scripts/ci-conformance.sh hack/resolve-e2e-versions.sh
Only pre-existing info-level findings (SC1091, SC2329) are expected.
Then run the resolver to confirm the nightly images resolve (needs gcloud/docker/git access):
E2E_FLAVOR=all GCP_PROJECT=<project> GCP_REGION=<region> source hack/resolve-e2e-versions.sh
The summary it prints should show IMAGE_ID and the upgrade images using the new version.
Optionally, build an image locally with scripts/ci-e2e.sh --init-image --build-image-only. This needs GCP credentials and an image-builder checkout in $GOPATH/src/sigs.k8s.io/image-builder, and creates billable resources, so only do it if asked.
Step 4: Commit
Create a single commit: chore(bump): bump e2e Ubuntu image to UBUNTU_VERSION, containing gcp-ci.yaml and the docs changes.
Do NOT push or create a PR unless the user asks.
Cluster API GCP roadmap
This roadmap is a constant work in progress, subject to frequent revision. Dates are approximations. Features are listed in no particular order.
v0.4 (v1Alpha4)
| Description | Issue/Proposal/PR |
|---|
v1beta1/v1
Proposal awaits.
Lifecycle frozen
Items within this category have been identified as potential candidates for the project and can be moved up into a milestone if there is enough interest.