cd ../blog
KubernetesOctober 7, 202616 min read

Building an EKS Operator for Expiring Preview Environments

Build a small Go operator that creates preview apps, repairs drift, and removes its workloads at expiry. A hands-on EKS walkthrough with a runnable project, tests, and cleanup steps.

Roger Vasconcelos
Roger Vasconcelos
AWS DevOps Architect
Building an EKS Operator for Expiring Preview Environments

AI-generated editorial illustration by NeuralOps

Preview environments are easy to start and surprisingly easy to forget. A branch gets a Deployment, somebody checks it through a port-forward, and the next thing on the backlog takes over. The pods are still there a week later.

That is a useful problem for learning Kubernetes operators. There is a desired state, a lifecycle, and a cleanup rule that should keep working after your laptop closes. You can also see the result without connecting six other services first.

In this exercise, we will build a small Go operator for an existing Amazon EKS cluster. One PreviewApp describes an application image, a replica count, and a lifetime. The operator creates a Deployment and an internal Service, repairs changes to the resources it owns, and removes them when the lifetime ends.

The goal is to understand the reconciliation loop well enough to debug it. This is a learning demo with a real platform use case. It manages one namespace and needs no AWS API permissions, load balancer, DNS automation, or GPU.

What we are actually building

A custom resource definition, or CRD, teaches the Kubernetes API about a new kind of object. A custom resource is an instance of that kind. Our controller watches those instances and repeatedly compares what exists with what should exist. Together, the CRD and lifecycle-aware controller implement the operator pattern.

The API we want to use is deliberately small:

apiVersion: platform.neuralops.ca/v1alpha1
kind: PreviewApp
metadata:
  name: checkout
  namespace: preview-demo
spec:
  image: nginxinc/nginx-unprivileged:1.28-alpine
  replicas: 1
  ttlSeconds: 600

Think of this as a request to the platform: “Run this preview for ten minutes.” The image must serve HTTP on port 8080 as a non-root user. The sample is a static NGINX page; swap in your own application once you understand the loop.

The operator owns the Deployment and Service. Kubernetes' Deployment controller owns the ReplicaSets and pods. The Service stays inside the cluster, and we use a port-forward to inspect it. Deleting an owned resource triggers repair; deleting the PreviewApp means the preview should disappear.

There is also a time-based transition. At expiry, the operator deletes its Deployment and Service and leaves the custom resource with phase: Expired. That record makes cleanup visible. An expired preview stays expired; create a new PreviewApp to start another one.

Get the runnable project

Download the complete operator demo. It contains the Go API types, controller, generated CRD, namespace-scoped deployment manifests, tests, and a README. There are no missing controller methods to fill in before you can try it.

unzip eks-preview-operator.zip
cd preview-operator
go version
make test
make build

Use Go 1.25.3 or newer, GNU-compatible make, and kubectl compatible with your cluster. Docker with Buildx and AWS CLI v2 are needed for the later ECR deployment. The project was scaffolded with Kubebuilder v4.13.0 and pins controller-runtime v0.23.1 and Kubernetes client libraries v0.35.0 in go.mod; retain go.sum when you make changes. Start with an EKS 1.35 learning cluster or check your chosen release against the client library's compatibility guidance.

If you want the scaffolding exercise too, install that Kubebuilder release and run these commands in a separate empty directory:

kubebuilder init --domain neuralops.ca \
  --repo neuralops.ca/preview-operator --plugins go/v4
kubebuilder create api --group platform --version v1alpha1 \
  --kind PreviewApp --resource --controller

That produces the starting files, not the completed demo. Compare them with the download. The Kubebuilder quick start explains the generated project layout. Our small additions live mainly in api/v1alpha1/previewapp_types.go, internal/controller/previewapp_controller.go, and the namespace configuration in cmd/main.go.

Read the API before the controller

The API types enforce a few useful boundaries: one to three replicas, a lifetime between 60 seconds and 24 hours, and a nonempty image. The defaults are one replica and 600 seconds. Schema validation rejects an invalid request before the controller sees it.

The spec is what you want. The status is what the controller has observed: the phase, expiry time, Service name, and a standard Ready condition. Keep those responsibilities separate. Writing “Ready” into a resource is useful only when you have checked the actual Deployment.

After changing API fields or Kubebuilder markers, regenerate the files:

make manifests generate

make manifests rebuilds the CRD and RBAC from markers. make generate rebuilds the DeepCopy methods. Editing the generated YAML directly is a quick way to lose a change on your next build.

Walk through one reconciliation

Open internal/controller/previewapp_controller.go. Its flow is the part worth understanding:

  • Read the latest PreviewApp. A missing resource ends this reconciliation; a deleting resource should not have its children recreated.
  • Compute the deadline from metadata.creationTimestamp and spec.ttlSeconds.
  • If the preview is expired, delete only children controlled by that exact resource UID, then record the expired state.
  • Otherwise, create or update the desired Deployment and Service with controller owner references.
  • Check the Deployment's current generation, updated replicas, and ready replicas before publishing readiness.
  • Schedule another reconciliation no later than the deadline. Watches on the owned children handle changes sooner.
  • The deadline calculation is intentionally boring:

    deadline := app.CreationTimestamp.Add(
        time.Duration(app.Spec.TTLSeconds) * time.Second,
    )
    

    Using “now plus ten minutes” on every reconcile would give the preview a fresh lifetime whenever anything changed. Keeping a timer only in memory would lose the deadline when the controller restarted. The API server's creation timestamp solves both problems. Changing the TTL before expiry changes the deadline relative to the original creation time; it does not reset the start time.

    The controller uses the resource UID in child names and selectors, so even a long preview name does not create an invalid Service name. It also checks ownership before updating or deleting an existing child. A name collision is an error to investigate, not permission to adopt somebody else's workload.

    Owner references let Kubernetes clean up dependents when you delete the PreviewApp. Expiry is handled explicitly because we keep the expired record. Cleanup is asynchronous: pods can take a short time to terminate after their Deployment disappears. Kubernetes garbage collection

    There are no external resources to clean up here, so this version needs no custom finalizer. Adding a DNS record or an AWS resource later introduces an external cleanup lifecycle, which is where finalizers become useful.

    Point kubectl at your learning cluster

    Use an existing sandbox cluster with spare capacity and permission to install a CRD. This walkthrough does not provision an EKS cluster. Keep its infrastructure in Terraform as usual.

    export AWS_REGION=us-east-1
    export EKS_CLUSTER=your-learning-cluster
    aws eks update-kubeconfig --region "$AWS_REGION" --name "$EKS_CLUSTER"
    kubectl config current-context
    kubectl get nodes
    

    Confirm that the displayed context is the cluster you intend to use. Your local controller uses that same kubeconfig. The EKS kubeconfig guide covers authentication and access prerequisites.

    Start locally, with EKS as the API server

    The quickest feedback loop runs the controller on your workstation. In the project directory:

    kubectl create namespace preview-demo
    make install
    WATCH_NAMESPACE=preview-demo make run
    

    Leave that command in the foreground. In another terminal, from the same project directory:

    kubectl apply -f config/samples/platform_v1alpha1_previewapp.yaml
    kubectl get previewapps -n preview-demo
    kubectl wait -n preview-demo --for=condition=Ready \
      previewapp/checkout --timeout=120s
    export PREVIEW_SERVICE=$(kubectl get previewapp checkout -n preview-demo \
      -o jsonpath='{.status.serviceName}')
    kubectl port-forward -n preview-demo "service/$PREVIEW_SERVICE" 8080:80
    

    Open http://localhost:8080. You should see the NGINX welcome page once the pod is ready. A slow image pull or a Pending pod may need investigation before the wait succeeds; a controller cannot supply missing node capacity.

    This local mode authenticates as your kubeconfig user. The in-cluster mode below uses a namespace-scoped service account instead. Stop the local controller before deploying the in-cluster one, so two controllers do not compete over the same resources.

    Change it, break it, watch it recover

    Stop the port-forward with Ctrl-C and use these commands in the second terminal while the local controller keeps running:

    kubectl patch previewapp checkout -n preview-demo --type merge \
      -p '{"spec":{"replicas":2}}'
    kubectl get deployments -n preview-demo
    kubectl scale deployment "$PREVIEW_SERVICE" -n preview-demo --replicas=1
    kubectl get deployments -n preview-demo -w
    

    The operator should restore two replicas because the custom resource still asks for two. This is why you should update the PreviewApp, rather than edit the generated Deployment. Stop the watch with Ctrl-C, then delete the child:

    kubectl delete deployment "$PREVIEW_SERVICE" -n preview-demo
    kubectl get deployments -n preview-demo -w
    

    The Deployment should come back. That recovery uses the child watch registered with Owns(). The controller is responsible for creating a missing Deployment; the Deployment controller already handles replacing an individual failed pod.

    For the expiry exercise, apply the short-lived second sample:

    kubectl apply -f config/samples/expiry-demo.yaml
    kubectl get previewapp expiry-demo -n preview-demo -w
    

    It expires 60 seconds after creation. The expected transition is Provisioning or Ready, followed by Expired; a slow startup can go straight from Provisioning to Expired. Stop the watch, then inspect kubectl get deployments,services -n preview-demo. The checkout preview may still exist, but the expiry demo's children should be gone.

    Try stopping and restarting the local controller during that minute. If the deadline passes while the controller is down, cleanup should happen when it returns. This is eventual reconciliation, not a guarantee that deletion happens at an exact wall-clock second during an outage.

    Package it for EKS

    Once the local exercise makes sense, stop make run. Build and push the manager image to an existing private ECR repository named preview-operator:

    export AWS_ACCOUNT_ID=$(aws sts get-caller-identity --query Account --output text)
    export ECR_REGISTRY="$AWS_ACCOUNT_ID.dkr.ecr.$AWS_REGION.amazonaws.com"
    export IMG="$ECR_REGISTRY/preview-operator:demo-v1"
    aws ecr get-login-password --region "$AWS_REGION" | \
      docker login --username AWS --password-stdin "$ECR_REGISTRY"
    docker buildx create --name preview-operator-builder \
      --driver docker-container --use
    docker buildx inspect --bootstrap
    docker buildx build --platform linux/amd64,linux/arm64 \
      --tag "$IMG" --push .
    make deploy-demo IMG="$IMG"
    kubectl rollout status -n preview-demo deployment/preview-operator --timeout=120s
    kubectl logs -n preview-demo deployment/preview-operator
    

    The multi-architecture build covers ordinary x86 and Graviton worker pools. We create a Docker-container builder because the default driver does not support this workflow on every Docker installation. If that named builder already exists, select it with docker buildx use preview-operator-builder instead of creating it again. Your nodes need network access and permission to pull from ECR; ImagePullBackOff is an image-access problem to resolve separately. Create the repository through your normal infrastructure workflow if it does not exist. Docker multi-platform builds, ECR push prerequisites

    deploy-demo installs the CRD and the custom config/demo manifests. Those manifests use a Role and RoleBinding in preview-demo, one manager replica, and a namespace-scoped cache. The controller calls only the Kubernetes API, so it needs no EKS Pod Identity association or IAM permissions for AWS APIs. The person installing the CRD still needs cluster-level installation privileges.

    Use this demo target for the walkthrough. Kubebuilder's original make deploy target is still present as scaffold reference and uses its broader generated configuration. Do not confuse the two.

    Delete and recreate an expired sample if you want to repeat the exercise with the in-cluster controller. Reapplying the same expired custom resource does not create a new lifetime.

    When something does not happen

    If no Deployment appears, inspect the custom resource's Ready condition and the controller logs. Check the namespace first, then RBAC and ownership errors. An Expired phase means the lifetime is over; confirm that cleanup succeeded using the condition reason and child-resource listing. ReconcileFailed means cleanup needs attention even though the preview is terminal. If the Deployment exists but never becomes ready, inspect its pods with kubectl describe pod and check image pulls, scheduling, probes, and application logs.

    If a direct Deployment edit sticks, make sure the controller is still running and watching preview-demo. If cleanup seems late, check status.expiresAt, the controller's availability, and resource termination. A resource retained by another finalizer needs its own investigation.

    What the tests prove

    The download uses envtest: a local API server and etcd, with no Docker daemon needed for that suite. It exercises admission validation, default values, idempotent reconciliation, replica repair, child recreation, readiness after a spec change, the exact expiry boundary, restart behavior, and refusing to touch unowned resources. A fake clock makes expiry tests fast and repeatable.

    Those tests do not run kubelet, the Deployment controller, or Kubernetes garbage collection. They cannot prove an image starts, pods terminate, or EKS networking works. Use the live commands above to validate those parts on your own cluster. This article does not claim a live EKS deployment was performed while preparing the demo.

    Clean up the exercise

    Delete the two previews while the controller is still running. For an in-cluster installation, then remove the demo manager and its RBAC:

    kubectl delete previewapp checkout expiry-demo -n preview-demo --ignore-not-found
    make undeploy-demo
    kubectl get deployments,services,pods -n preview-demo
    

    For local mode, stop the controller after deleting the previews. undeploy-demo removes only the manager, service account, Role, and RoleBinding; it leaves the namespace and CRD. If this was a dedicated exercise and no other previews depend on the CRD, remove the remaining namespace and run make uninstall. Deleting a CRD deletes its custom resources across the cluster, so inspect them first. Delete the demo ECR image separately through your normal registry workflow when you no longer need it.

    Where I would take it next

    The first useful extension is a pull-request workflow that creates a PreviewApp and deletes it when the PR closes. Keep the TTL as a fallback for missed events. After that, add application configuration, namespace isolation, quotas, image policy, and ingress if your teams need browser-accessible previews.

    A shared production service would also need deliberate ownership rules, leader election, observable reconciliation failures, and an approach to conflicts with admission mutations or other controllers. Start with the small loop you can explain. When you can predict what happens after a child is deleted or the operator restarts, you have learned the part that makes the next operator easier.

    Amazon EKSKubernetes OperatorsGoKubebuilderPlatform Engineering

    Share this article

    Roger Vasconcelos

    Roger Vasconcelos

    AWS DevOps Architect

    Senior AWS DevOps Architect with 10 AWS certifications and 20+ years of experience, running a boutique practice augmented by a fleet of AI agents.

    Learn more about me

    Need Help with Your Cloud Infrastructure?

    Let's discuss how I can help you implement these best practices in your organization.

    Get in Touch