Preview environments are easy to start and surprisingly easy to forget. A branch gets a Deployment, somebody checks it through a port-forward, and the next thing on the backlog takes over. The pods are still there a week later.
That is a useful problem for learning Kubernetes operators. There is a desired state, a lifecycle, and a cleanup rule that should keep working after your laptop closes. You can also see the result without connecting six other services first.
In this exercise, we will build a small Go operator for an existing Amazon EKS cluster. One PreviewApp describes an application image, a replica count, and a lifetime. The operator creates a Deployment and an internal Service, repairs changes to the resources it owns, and removes them when the lifetime ends.
The goal is to understand the reconciliation loop well enough to debug it. This is a learning demo with a real platform use case. It manages one namespace and needs no AWS API permissions, load balancer, DNS automation, or GPU.
What we are actually building
A custom resource definition, or CRD, teaches the Kubernetes API about a new kind of object. A custom resource is an instance of that kind. Our controller watches those instances and repeatedly compares what exists with what should exist. Together, the CRD and lifecycle-aware controller implement the operator pattern.
The API we want to use is deliberately small:
apiVersion: platform.neuralops.ca/v1alpha1
kind: PreviewApp
metadata:
name: checkout
namespace: preview-demo
spec:
image: nginxinc/nginx-unprivileged:1.28-alpine
replicas: 1
ttlSeconds: 600
Think of this as a request to the platform: “Run this preview for ten minutes.” The image must serve HTTP on port 8080 as a non-root user. The sample is a static NGINX page; swap in your own application once you understand the loop.
The operator owns the Deployment and Service. Kubernetes' Deployment controller owns the ReplicaSets and pods. The Service stays inside the cluster, and we use a port-forward to inspect it. Deleting an owned resource triggers repair; deleting the PreviewApp means the preview should disappear.
There is also a time-based transition. At expiry, the operator deletes its Deployment and Service and leaves the custom resource with phase: Expired. That record makes cleanup visible. An expired preview stays expired; create a new PreviewApp to start another one.
Get the runnable project
Download the complete operator demo. It contains the Go API types, controller, generated CRD, namespace-scoped deployment manifests, tests, and a README. There are no missing controller methods to fill in before you can try it.
unzip eks-preview-operator.zip
cd preview-operator
go version
make test
make build
Use Go 1.25.3 or newer, GNU-compatible make, and kubectl compatible with your cluster. Docker with Buildx and AWS CLI v2 are needed for the later ECR deployment. The project was scaffolded with Kubebuilder v4.13.0 and pins controller-runtime v0.23.1 and Kubernetes client libraries v0.35.0 in go.mod; retain go.sum when you make changes. Start with an EKS 1.35 learning cluster or check your chosen release against the client library's compatibility guidance.
If you want the scaffolding exercise too, install that Kubebuilder release and run these commands in a separate empty directory:
kubebuilder init --domain neuralops.ca \
--repo neuralops.ca/preview-operator --plugins go/v4
kubebuilder create api --group platform --version v1alpha1 \
--kind PreviewApp --resource --controller
That produces the starting files, not the completed demo. Compare them with the download. The Kubebuilder quick start explains the generated project layout. Our small additions live mainly in api/v1alpha1/previewapp_types.go, internal/controller/previewapp_controller.go, and the namespace configuration in cmd/main.go.
Read the API before the controller
The API types enforce a few useful boundaries: one to three replicas, a lifetime between 60 seconds and 24 hours, and a nonempty image. The defaults are one replica and 600 seconds. Schema validation rejects an invalid request before the controller sees it.
The spec is what you want. The status is what the controller has observed: the phase, expiry time, Service name, and a standard Ready condition. Keep those responsibilities separate. Writing “Ready” into a resource is useful only when you have checked the actual Deployment.
After changing API fields or Kubebuilder markers, regenerate the files:
make manifests generate
make manifests rebuilds the CRD and RBAC from markers. make generate rebuilds the DeepCopy methods. Editing the generated YAML directly is a quick way to lose a change on your next build.
Walk through one reconciliation
Open internal/controller/previewapp_controller.go. Its flow is the part worth understanding:
PreviewApp. A missing resource ends this reconciliation; a deleting resource should not have its children recreated.metadata.creationTimestamp and spec.ttlSeconds.The deadline calculation is intentionally boring:
deadline := app.CreationTimestamp.Add(
time.Duration(app.Spec.TTLSeconds) * time.Second,
)
Using “now plus ten minutes” on every reconcile would give the preview a fresh lifetime whenever anything changed. Keeping a timer only in memory would lose the deadline when the controller restarted. The API server's creation timestamp solves both problems. Changing the TTL before expiry changes the deadline relative to the original creation time; it does not reset the start time.
The controller uses the resource UID in child names and selectors, so even a long preview name does not create an invalid Service name. It also checks ownership before updating or deleting an existing child. A name collision is an error to investigate, not permission to adopt somebody else's workload.
Owner references let Kubernetes clean up dependents when you delete the PreviewApp. Expiry is handled explicitly because we keep the expired record. Cleanup is asynchronous: pods can take a short time to terminate after their Deployment disappears. Kubernetes garbage collection
There are no external resources to clean up here, so this version needs no custom finalizer. Adding a DNS record or an AWS resource later introduces an external cleanup lifecycle, which is where finalizers become useful.
Point kubectl at your learning cluster
Use an existing sandbox cluster with spare capacity and permission to install a CRD. This walkthrough does not provision an EKS cluster. Keep its infrastructure in Terraform as usual.
export AWS_REGION=us-east-1
export EKS_CLUSTER=your-learning-cluster
aws eks update-kubeconfig --region "$AWS_REGION" --name "$EKS_CLUSTER"
kubectl config current-context
kubectl get nodes
Confirm that the displayed context is the cluster you intend to use. Your local controller uses that same kubeconfig. The EKS kubeconfig guide covers authentication and access prerequisites.
Start locally, with EKS as the API server
The quickest feedback loop runs the controller on your workstation. In the project directory:
kubectl create namespace preview-demo
make install
WATCH_NAMESPACE=preview-demo make run
Leave that command in the foreground. In another terminal, from the same project directory:
kubectl apply -f config/samples/platform_v1alpha1_previewapp.yaml
kubectl get previewapps -n preview-demo
kubectl wait -n preview-demo --for=condition=Ready \
previewapp/checkout --timeout=120s
export PREVIEW_SERVICE=$(kubectl get previewapp checkout -n preview-demo \
-o jsonpath='{.status.serviceName}')
kubectl port-forward -n preview-demo "service/$PREVIEW_SERVICE" 8080:80
Open http://localhost:8080. You should see the NGINX welcome page once the pod is ready. A slow image pull or a Pending pod may need investigation before the wait succeeds; a controller cannot supply missing node capacity.
This local mode authenticates as your kubeconfig user. The in-cluster mode below uses a namespace-scoped service account instead. Stop the local controller before deploying the in-cluster one, so two controllers do not compete over the same resources.
Change it, break it, watch it recover
Stop the port-forward with Ctrl-C and use these commands in the second terminal while the local controller keeps running:
kubectl patch previewapp checkout -n preview-demo --type merge \
-p '{"spec":{"replicas":2}}'
kubectl get deployments -n preview-demo
kubectl scale deployment "$PREVIEW_SERVICE" -n preview-demo --replicas=1
kubectl get deployments -n preview-demo -w
The operator should restore two replicas because the custom resource still asks for two. This is why you should update the PreviewApp, rather than edit the generated Deployment. Stop the watch with Ctrl-C, then delete the child:
kubectl delete deployment "$PREVIEW_SERVICE" -n preview-demo
kubectl get deployments -n preview-demo -w
The Deployment should come back. That recovery uses the child watch registered with Owns(). The controller is responsible for creating a missing Deployment; the Deployment controller already handles replacing an individual failed pod.
For the expiry exercise, apply the short-lived second sample:
kubectl apply -f config/samples/expiry-demo.yaml
kubectl get previewapp expiry-demo -n preview-demo -w
It expires 60 seconds after creation. The expected transition is Provisioning or Ready, followed by Expired; a slow startup can go straight from Provisioning to Expired. Stop the watch, then inspect kubectl get deployments,services -n preview-demo. The checkout preview may still exist, but the expiry demo's children should be gone.
Try stopping and restarting the local controller during that minute. If the deadline passes while the controller is down, cleanup should happen when it returns. This is eventual reconciliation, not a guarantee that deletion happens at an exact wall-clock second during an outage.
Package it for EKS
Once the local exercise makes sense, stop make run. Build and push the manager image to an existing private ECR repository named preview-operator:
export AWS_ACCOUNT_ID=$(aws sts get-caller-identity --query Account --output text)
export ECR_REGISTRY="$AWS_ACCOUNT_ID.dkr.ecr.$AWS_REGION.amazonaws.com"
export IMG="$ECR_REGISTRY/preview-operator:demo-v1"
aws ecr get-login-password --region "$AWS_REGION" | \
docker login --username AWS --password-stdin "$ECR_REGISTRY"
docker buildx create --name preview-operator-builder \
--driver docker-container --use
docker buildx inspect --bootstrap
docker buildx build --platform linux/amd64,linux/arm64 \
--tag "$IMG" --push .
make deploy-demo IMG="$IMG"
kubectl rollout status -n preview-demo deployment/preview-operator --timeout=120s
kubectl logs -n preview-demo deployment/preview-operator
The multi-architecture build covers ordinary x86 and Graviton worker pools. We create a Docker-container builder because the default driver does not support this workflow on every Docker installation. If that named builder already exists, select it with docker buildx use preview-operator-builder instead of creating it again. Your nodes need network access and permission to pull from ECR; ImagePullBackOff is an image-access problem to resolve separately. Create the repository through your normal infrastructure workflow if it does not exist. Docker multi-platform builds, ECR push prerequisites
deploy-demo installs the CRD and the custom config/demo manifests. Those manifests use a Role and RoleBinding in preview-demo, one manager replica, and a namespace-scoped cache. The controller calls only the Kubernetes API, so it needs no EKS Pod Identity association or IAM permissions for AWS APIs. The person installing the CRD still needs cluster-level installation privileges.
Use this demo target for the walkthrough. Kubebuilder's original make deploy target is still present as scaffold reference and uses its broader generated configuration. Do not confuse the two.
Delete and recreate an expired sample if you want to repeat the exercise with the in-cluster controller. Reapplying the same expired custom resource does not create a new lifetime.
When something does not happen
If no Deployment appears, inspect the custom resource's Ready condition and the controller logs. Check the namespace first, then RBAC and ownership errors. An Expired phase means the lifetime is over; confirm that cleanup succeeded using the condition reason and child-resource listing. ReconcileFailed means cleanup needs attention even though the preview is terminal. If the Deployment exists but never becomes ready, inspect its pods with kubectl describe pod and check image pulls, scheduling, probes, and application logs.
If a direct Deployment edit sticks, make sure the controller is still running and watching preview-demo. If cleanup seems late, check status.expiresAt, the controller's availability, and resource termination. A resource retained by another finalizer needs its own investigation.
What the tests prove
The download uses envtest: a local API server and etcd, with no Docker daemon needed for that suite. It exercises admission validation, default values, idempotent reconciliation, replica repair, child recreation, readiness after a spec change, the exact expiry boundary, restart behavior, and refusing to touch unowned resources. A fake clock makes expiry tests fast and repeatable.
Those tests do not run kubelet, the Deployment controller, or Kubernetes garbage collection. They cannot prove an image starts, pods terminate, or EKS networking works. Use the live commands above to validate those parts on your own cluster. This article does not claim a live EKS deployment was performed while preparing the demo.
Clean up the exercise
Delete the two previews while the controller is still running. For an in-cluster installation, then remove the demo manager and its RBAC:
kubectl delete previewapp checkout expiry-demo -n preview-demo --ignore-not-found
make undeploy-demo
kubectl get deployments,services,pods -n preview-demo
For local mode, stop the controller after deleting the previews. undeploy-demo removes only the manager, service account, Role, and RoleBinding; it leaves the namespace and CRD. If this was a dedicated exercise and no other previews depend on the CRD, remove the remaining namespace and run make uninstall. Deleting a CRD deletes its custom resources across the cluster, so inspect them first. Delete the demo ECR image separately through your normal registry workflow when you no longer need it.
Where I would take it next
The first useful extension is a pull-request workflow that creates a PreviewApp and deletes it when the PR closes. Keep the TTL as a fallback for missed events. After that, add application configuration, namespace isolation, quotas, image policy, and ingress if your teams need browser-accessible previews.
A shared production service would also need deliberate ownership rules, leader election, observable reconciliation failures, and an approach to conflicts with admission mutations or other controllers. Start with the small loop you can explain. When you can predict what happens after a child is deleted or the operator restarts, you have learned the part that makes the next operator easier.

