
Operators and CRDs: Install One, Read Its Schema, Stop Its Controller
Operators and CRDs: Install One, Read Its Schema, Stop Its Controller
Part 9 of the CKA roadmap Β· Phase 2 β Cluster Architecture Β· 25%
Operators came into the CKA with the 2025 curriculum revision. The objective reads "Understand CRDs, install and configure operators" β and nothing in that line asks you to write a controller or design a CRD. What it does ask for, taken at its word:
- install an operator someone else wrote
- find out what new resource types it added
- write a custom resource for a schema you have never seen, without a template
- work out why a custom resource you created is not doing anything
This post does all four on the two-node cluster from part 1.
The one idea to hold
An operator is a CRD plus a controller. Everything else in this post follows from which half you are looking at.
| Piece | What it is | What it does |
|---|---|---|
| CRD (CustomResourceDefinition) | A schema, registered with the API server | Teaches the API a new kind β nothing more |
| CR (custom resource) | An object of that kind, stored in etcd | Nothing. It is data: a statement of what you want |
| Controller | A Deployment running a reconcile loop | Watches CRs and changes the world to match them |
The API server validates and stores CRs by itself. Only the controller ever acts on one. Keep that split in mind and both the install step and the troubleshooting step become obvious.
Why cert-manager
Any operator would do. I picked cert-manager for three reasons:
- One Helm command to install, with no Operator Lifecycle Manager (OLM) to set up first
- Three small pods, which fit comfortably next to everything else on an
e2-medium - A controller whose work you can see. Change a
Certificateand a new Secret, a new CertificateRequest and a burst of log lines appear. That visible reaction is the whole point of the middle of this lab.
The task
Install cert-manager from the Helm repository
https://charts.jetstack.iointo the namespacecert-manager, including its CRDs. List the resource types it added. For the type that represents a TLS certificate, find whichspecfields are required.In namespace
op-lab, create a self-signedIssuernamedselfsignedand aCertificatenameddemo-certfordemo.op-lab.svc, stored in the Secretdemo-cert-tls. Then add a second DNS name and prove the certificate was reissued with it.
Written from scratch for this series, like every task in it.
The solution
Step 0 β A baseline
A namespace for the lab, and a list of the CRDs that exist before you install anything:
kubectl create ns op-lab
kubectl get crd -o name | sort > crd-before.txt
wc -l crd-before.txtThe count depends on your CNI. A Flannel cluster usually has zero CRDs. A Calico cluster already has a dozen or more β *.projectcalico.org and *.operator.tigera.io β because Calico is itself installed by an operator. If you followed the Calico switch in part 1, you installed an operator before you knew what one was.
Step 1 β Install the operator
helm repo add jetstack https://charts.jetstack.io
helm repo update
helm search repo jetstack/cert-manager --versions | head -5helm install cert-manager jetstack/cert-manager \
-n cert-manager --create-namespace \
--set crds.enabled=truecrds.enabled=true is the current flag, from chart v1.15 on. Much of what you will find online still says installCRDs=true β that is the old name, deprecated but still accepted.
Wait for all three Deployments:
kubectl wait --for=condition=Available deploy --all -n cert-manager --timeout=180s
kubectl get deploy,pods -n cert-managerThe three are worth telling apart, because each one fails differently:
| Deployment | Role |
|---|---|
cert-manager | The controller β the reconcile loop that does the work |
cert-manager-webhook | Validates CRs before the API server stores them β admission control |
cert-manager-cainjector | Injects CA data into webhook and API configurations |
Finding an operator
Outside the exam, Artifact Hub is where you look for a chart β search cert-manager, pick the one from the verified publisher, and the Install button gives you the repo URL. helm search hub cert-manager queries the same index from the terminal.
Artifact Hub is not among the sites you can open during the exam, so an exam task would hand you the repository. Practise with it to learn what a chart page tells you β repo, values, CRDs β not to memorise URLs.
Step 2 β What did it add?
Take the same snapshot again and print only the new lines:
kubectl get crd -o name | sort > crd-after.txt
comm -13 crd-before.txt crd-after.txtSix new CRDs across two API groups:
| Group | CRDs |
|---|---|
cert-manager.io | certificates, certificaterequests, issuers, clusterissuers |
acme.cert-manager.io | orders, challenges |
A CRD name tells you where it lives, but not how to use it. For that, ask for the group as resources:
kubectl api-resources --api-group=cert-manager.io
kubectl api-resources --api-group=acme.cert-manager.ioIf you memorise one command from this post, make it this one. A single line of output answers four questions a task is likely to depend on: the plural name, the short name, the API version, and whether the type is namespaced. Here it shows, for example, that Issuer is namespaced and ClusterIssuer is not.
Step 3 β Read the schema with kubectl explain
Do this before writing any YAML, because custom resources have no generator. There is no kubectl create certificate, so the $do trick from part 2 cannot help. Every line of a CR comes from reading its schema.
Top down:
kubectl explain certificate
kubectl explain certificate.specRead two things in the output: the type in angle brackets, and the -required- label.
To find where a field lives, print the whole tree without descriptions, and search it instead of scrolling:
kubectl explain certificate.spec --recursive | head -60
kubectl explain certificate.spec --recursive | grep -n -i dnsOnce you know where a field is, go back to plain explain to read what it means:
kubectl explain certificate.spec.issuerRef
kubectl explain certificate.spec.secretName
kubectl explain issuer.spec --recursive | head -40The apiVersion: line comes from the CRD itself, or from the APIVERSION column of api-resources:
kubectl get crd certificates.cert-manager.io -o jsonpath='{.spec.versions[*].name}{"\n"}'Answer these before you write anything:
- Which
specfields of aCertificateare required? - What values does
issuerRef.kindaccept? - Which field decides the name of the Secret that ends up holding the key pair?
Answers
secretNameandissuerRef.IssuerorClusterIssuer. It defaults toIssuerwhen left out.spec.secretName.
When `explain` gives you less than you need
the server doesn't have a resource typeβ the name is wrong. Use a name fromapi-resources: singular, plural and short name all work.- A field shows
<Object>and no children β the CRD marks itx-kubernetes-preserve-unknown-fields: true, which means it has no schema.explainhas nothing to show, and this is the one case where you have to read the operator's own docs. - Two CRDs define the same kind in different groups β add
--api-version=<group>/<version>to pick one.
Step 4 β Create the CR, and watch the controller take it
Open a second SSH session first and leave the controller's log running for the rest of the lab:
kubectl logs -n cert-manager deploy/cert-manager -f --since=1sdeploy/<name> saves you from copying a pod name with a hash in it.
Back in the first session, write the two resources from what explain told you:
cat > cert.yaml <<'YAMLEOF'
apiVersion: cert-manager.io/v1
kind: Issuer
metadata:
name: selfsigned
namespace: op-lab
spec:
selfSigned: {}
---
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: demo-cert
namespace: op-lab
spec:
secretName: demo-cert-tls
commonName: demo.op-lab.svc
dnsNames:
- demo.op-lab.svc
issuerRef:
name: selfsigned
kind: Issuer
YAMLEOF
kubectl apply -f cert.yamlThe second session should fill with lines mentioning op-lab/demo-cert. That is the controller picking the work up.
Now look at what it produced:
kubectl get issuer,certificate -n op-lab
kubectl get certificaterequest -n op-lab
kubectl get secret demo-cert-tls -n op-labREADY on the Certificate should be True.
You did not create demo-cert-tls. You declared that you wanted a certificate. The controller generated the key, signed it, and wrote the Secret. That is the operator pattern in one command.
Step 5 β Change the CR
Add a second DNS name:
kubectl patch certificate demo-cert -n op-lab --type=merge \
-p '{"spec":{"dnsNames":["demo.op-lab.svc","demo2.op-lab.svc"]}}'kubectl edit certificate demo-cert -n op-lab works just as well. Custom resources take edit, patch, describe and get -o yaml exactly like built-in ones.
Then check three places:
kubectl get certificaterequest -n op-lab
kubectl describe certificate demo-cert -n op-lab | tail -15kubectl get secret demo-cert-tls -n op-lab -o jsonpath='{.data.tls\.crt}' \
| base64 -d | openssl x509 -noout -ext subjectAltNameWhat to expect:
- A new CertificateRequest ending in
-2. Every reissue is a new request, numbered by revision. On cert-manager 1.18 and later,revisionHistoryLimitdefaults to1, so the-1request is cleaned up and you will likely see only-2. On older releases both stay. - A fresh round of events in
describe, showing the certificate being issued again. - Both DNS names in the certificate's subject alternative names.
- New lines in the log session, arriving the moment you patched.
Troubleshooting: the CR that does nothing
This half of the lab belongs as much to the troubleshooting domain as to this one. The situation is a common one in real clusters: the custom resource exists, kubectl get shows it, it looks correct β and nothing happens.
Stop the controller
kubectl scale deploy cert-manager -n cert-manager --replicas=0Change the Certificate again:
kubectl patch certificate demo-cert -n op-lab --type=merge \
-p '{"spec":{"dnsNames":["demo.op-lab.svc","demo2.op-lab.svc","demo3.op-lab.svc"]}}'
kubectl get certificaterequest -n op-labPredict before you press Enter. Will the patch be accepted? Will a third request appear?
What happens
The patch is accepted. The API server and the webhook are both still running, so the change is validated and stored in etcd.
But no -3 request appears, and the Secret does not change. The CR is data. With no controller running, nobody reads it.
Start it again
kubectl scale deploy cert-manager -n cert-manager --replicas=1
kubectl rollout status deploy cert-manager -n cert-manager
kubectl get certificaterequest -n op-labWithin moments of starting, the controller reissues the certificate with demo3 β and you did not re-apply anything.
It is not replaying an event it missed while it was down. It reads the CR as it is now, compares that with what actually exists, and closes the gap. That is what a reconcile loop means, and it is why an operator that restarts catches up on its own.
The order to check in
When a custom resource "does nothing", look at the controller before you look at the CR:
kubectl describe <kind> <name>β readStatusandEventskubectl get deploy,pods -n <operator-namespace>β is the controller running at all?kubectl logs deploy/<controller> -n <operator-namespace>β does it see the CR? Is it reporting errors?- If the log says
forbiddenβ the controller's ServiceAccount is missing RBAC permissions
The opposite symptom: a dead webhook
I only stopped the controller in this run, not the webhook β but the contrast is worth knowing.
With cert-manager-webhook down, kubectl apply of a cert-manager resource is rejected immediately, with an error mentioning failed calling webhook.
So the two failures look nothing alike:
| What is down | What you see |
|---|---|
| Controller | The CR is stored, but nothing acts on it |
| Webhook | The CR cannot be stored at all |
Recognise which one you are looking at and you know where to go.
A CRD with no controller at all
The same point, from the other side. Here is a CRD I wrote for this post: a made-up Snapshot type, with nothing behind it.
cat > snapshot-crd.yaml <<'YAMLEOF'
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
name: snapshots.lab.example.com
spec:
group: lab.example.com
scope: Namespaced
names:
kind: Snapshot
singular: snapshot
plural: snapshots
shortNames:
- snap
versions:
- name: v1
served: true
storage: true
schema:
openAPIV3Schema:
type: object
properties:
spec:
type: object
required: ["schedule", "pvcName"]
properties:
schedule:
type: string
pvcName:
type: string
keep:
type: integer
minimum: 1
YAMLEOF
kubectl apply -f snapshot-crd.yamlRead the new type exactly as you read cert-manager's:
kubectl api-resources --api-group=lab.example.com
kubectl explain snapshot.specexplain marks pvcName and schedule as -required- β the schema's required list, surfaced back to you. Now create one:
cat > snapshot.yaml <<'YAMLEOF'
apiVersion: lab.example.com/v1
kind: Snapshot
metadata:
name: data-nightly
namespace: op-lab
spec:
schedule: "0 2 * * *"
pvcName: data
keep: 7
YAMLEOF
kubectl apply -f snapshot.yaml
kubectl get snap -n op-lab
kubectl describe snapshot data-nightly -n op-labIt is stored, listed and described like any built-in object. And nothing will ever act on it. No Job, no backup, no event from a controller, because no controller exists. This is the stopped cert-manager case, made permanent.
"Install the operator" β but is there one?
One practice exercise I worked through handed me a manifest described as installing a database operator. It contained exactly one object: a CustomResourceDefinition. No Deployment, no controller.
Before you trust any file sold as an "operator", check what is actually in it:
grep -c '^kind:' <file>.yaml
kubectl get deploy -A | grep -i <name>A CRD on its own gives you half an operator. Custom resources of that type will be stored and never acted on.
The schema is enforced, not decorative
The API server uses the CRD's schema to check every write β even with no controller anywhere.
A value of the wrong type:
kubectl patch snap data-nightly -n op-lab --type=merge -p '{"spec":{"pvcName":123}}'Rejected, because the schema says pvcName is a string. The same happens with a value below the declared minimum:
kubectl patch snap data-nightly -n op-lab --type=merge -p '{"spec":{"keep":0}}'A field the schema does not know goes a different way. Add retention: 30 under spec in snapshot.yaml and apply it twice:
kubectl apply -f snapshot.yaml
kubectl apply -f snapshot.yaml --validate=warn
kubectl get snap data-nightly -n op-lab -o jsonpath='{.spec}{"\n"}'- By default, kubectl validates strictly, and the first apply fails with an
unknown fielderror. - With
--validate=warn, the apply goes through with a warning β andretentionis still not in the output.
The API server prunes fields that are not in the schema before storing the object. An unknown field never reaches etcd.
Verification
Each of these should hold before you call the lab done:
kubectl get crd | grep -c cert-manager.io
kubectl get certificate demo-cert -n op-lab
kubectl get secret demo-cert-tls -n op-lab -o jsonpath='{.data.tls\.crt}' \
| base64 -d | openssl x509 -noout -ext subjectAltName
kubectl get snap -n op-lab- Six cert-manager CRDs
demo-certisREADY True- The certificate lists all three DNS names β
demo3included, which proves the controller caught up after its restart data-nightlyexists, and nothing has ever acted on it
Traps
This lab ran clean for me. It added no rows to my error log. So these are not mistakes I made. They are the edges the lab was built around β each one quick to hit if you don't know it is there.
The old install flag
--set installCRDs=true appears in most older tutorials. Since v1.15 the flag is crds.enabled. The old one still works for now, but it is deprecated.
A dot in a JSON path key
Secret keys like tls.crt contain a dot, and JSON path reads dots as separators. Write {.data.tls\.crt}. Without the backslash you get empty output β and no error to tell you why.
An unquoted cron expression
Always quote cron schedules in YAML, in a CR and in a CronJob alike. "0 2 * * *" happens to parse without quotes. * * * * * and */5 * * * * do not, because a value starting with * is YAML syntax for an alias. Quote every one, and you never have to remember which is which.
helm uninstall leaves the CRDs behind
The cert-manager chart marks its CRDs with helm.sh/resource-policy: keep (the chart value crds.keep, on by default). helm uninstall removes the controller and leaves every CRD in place.
That is deliberate. The alternative is worse.
kubectl delete crd deletes more than a CRD
Deleting a CRD deletes every custom resource of that kind, in every namespace, with no confirmation. On a lab cluster that is the cleanup you want. On a real one it can wipe out every certificate in the cluster in a single command.
Speed
I have no times for this lab β I ran it clean and moved on without recording them, and my own study log calls that out. So there is no number here. Instead, the rule I hold every lab to: any part that takes over seven minutes goes on my list of slowest tasks, to be re-run against a clock before the exam.
What actually saves time here:
kubectl api-resources --api-group=<group>β plural, short name, version and scope, in one linekubectl explain <kind>.spec --recursive | grep -n <word>β find the field, don't scroll to itdeploy/<name>inkubectl logsβ no pod name to copy- Short names β
snap, and whateverapi-resourceslists for other types, are just as valid in a task as the full name
The checklist:
The docs pages
- Extend the Kubernetes API with CustomResourceDefinitions β the one to bookmark. It has a complete CRD to copy from, and sections on validation and pruning. You can open it during the exam, and adapting its example is faster than writing
apiextensions.k8s.io/v1from memory. - Custom Resources β the concepts page: what CRDs are, and when to use one
- Operator pattern β short, and worth reading once
The cert-manager documentation is not on the exam's allowed list. Inside the exam, kubectl explain is the operator documentation you have.
Say it out loud
Before you move on, answer these without notes:
- What are the parts of an operator, and which part does the work?
- A custom resource does nothing. Where do you look first?
- Which command shows a custom resource's schema without opening any docs?
- You stop the controller, change a CR, then start the controller again. What happens, and why?
- What does
kubectl delete crdtake with it?
Clean up
kubectl delete ns op-lab
helm uninstall cert-manager -n cert-manager
kubectl get crd | grep cert-manager.ioThe CRDs are still there β the keep policy at work. Remove them, and the one from this post, by hand:
kubectl get crd -o name | grep cert-manager.io | xargs kubectl delete
kubectl delete crd snapshots.lab.example.com
kubectl delete ns cert-managerPrevious: Helm and Kustomize Β· Back to the CKA roadmap Β· Related:Build a two-node kubeadm cluster on GCP, where Calico is installed by an operator