Manual snapshots (Part 2) are for the moment before a risky change. Day-to-day protection is a schedule. NDK splits that into three objects, and this part builds them for the journal application, watches the first run fire, and then looks at a plan that has been running on our lab for 25 days to see what retention really does.
Lab: NKP 2.18, NDK 2.3.0, all outputs from 2026-09-19.
All manifests in this part are in github.com/Fen0l/ndk-examples, the same files we applied on the lab.
Three objects, one job each
| Object | Answers | Scope |
|---|---|---|
JobScheduler |
When? | namespace, reusable by several plans |
ProtectionPlan |
What kind of protection, how many to keep, replicate where? | namespace |
AppProtectionPlan |
Which Application gets which plans? | namespace, one per Application |
The split looks heavy for a first plan. It pays off when you have twenty applications on the same "hourly, keep 24" policy: one scheduler, one plan, twenty bindings.
The 60-minute floor
Before writing the schedule, know the limit. NDK refuses anything more frequent than once an hour, on every schedule type. We tried both:
The JobScheduler "journal-every-5m" is invalid: spec.cronSchedule: Invalid value: "*/5 * * * *":
Cron schedule of lower interval than 60 minutes is not allowed
The JobScheduler "journal-30m" is invalid: spec.interval.minutes: Invalid value: 30:
Interval time must be a positive integer greater than or equal to 60
That is a validating webhook, so the object never lands in etcd. If your RPO requirement is below one hour, snapshots are the wrong tool and you are looking at NearSync or sync replication, which need a PE-level topology we do not have in the lab.
The manifests
apiVersion: scheduler.nutanix.com/v1alpha1
kind: JobScheduler
metadata:
name: journal-hourly
namespace: ndk-howto
spec:
interval:
minutes: 60
startTime: "2026-09-19T14:05:00Z"
timeZoneName: Etc/UTC
---
apiVersion: dataservices.nutanix.com/v1alpha1
kind: ProtectionPlan
metadata:
name: journal-local-hourly
namespace: ndk-howto
spec:
protectionType: async
scheduleName: journal-hourly
retentionPolicy:
retentionCount: 2
---
apiVersion: dataservices.nutanix.com/v1alpha1
kind: AppProtectionPlan
metadata:
name: journal-protection
namespace: ndk-howto
spec:
applicationName: journal
protectionPlanNames:
- journal-local-hourly
A few notes on the fields:
startTimeis optional. Without it, an interval schedule starts counting from the moment you create the object. We set it three minutes ahead so we could watch the first run.daily,weekly,monthlyandcronScheduleare the other schedule types, all with the same 60-minute floor.protectionType: asyncwith noreplicationConfigsmeans local snapshots only. Adding areplicationConfigslist with areplicationTargetNamemakes every scheduled snapshot replicate (Part 4).retentionCountis between 1 and 15. It counts successful snapshots on this cluster only.- The
AppProtectionPlanis where protection actually starts. Nothing happens until it exists.
kubectl apply -f 05-scheduled-protection.yaml
kubectl -n ndk-howto get jobscheduler,protectionplan,appprotectionplan
NAME LASTACTIVATION NEXTACTIVATION
jobscheduler.scheduler.nutanix.com/journal-hourly 2026-09-19T14:05:00Z
NAME SCHEDULE-NAME RETENTION-COUNT AVAILABLE DEGRADED PROTECTION-TYPE
protectionplan.dataservices.nutanix.com/journal-local-hourly journal-hourly 2 True False async
NAME APPLICATIONNAME PROTECTIONPLANS-APPLIED AVAILABLE DEGRADED
appprotectionplan.dataservices.nutanix.com/journal-protection journal ["journal-local-hourly"] True False
The NEXTACTIVATION column is your first check. If it is empty, the schedule spec was not understood.
Binding a plan also changes the Application. Its finalizers went from one to two:
["dataservices.nutanix.com/app","dataservices.nutanix.com/app-protection-plan"]
This is what we pointed at in Part 1: an Application with a plan cannot be deleted until the AppProtectionPlan is gone. Delete in reverse order of creation.
The first run
kubectl -n ndk-howto get applicationsnapshot -w
At 14:05:28Z, 28 seconds after the scheduled time:
NAME AGE READY-TO-USE BOUND-SNAPSHOTCONTENT SNAPSHOT-AGE CONSISTENCY-TYPE
journal-before-change 20m true asc-fd129a47-fc53-4572-91b3-2c473555a3c9 19m CrashConsistent
journal-f40c53315adf8d8c-1c72d2d 27s false asc-f5ba48b0-929c-411e-898d-9bf893bdf11c
The snapshot's creationTimestamp is 2026-09-19T14:05:00Z. Not 14:05:03, not 14:05:12. The scheduler fires on the second, and we saw the same on the 25-day-old plan below (every run at exactly 16:07:00Z). Half a minute later it was READY-TO-USE: true, CrashConsistent.
The scheduler and the binding both record the run:
kubectl -n ndk-howto get jobscheduler journal-hourly -o jsonpath='{.status}'
kubectl -n ndk-howto get appprotectionplan journal-protection -o jsonpath='{.status.protectionPlanExecutionStatus}'
{"lastActivation":"2026-09-19T14:05:00Z","lastUpdatedAt":"2026-09-19T14:05:00Z","nextActivation":"2026-09-19T15:05:00Z"}
[{"lastExecutionTime":"2026-09-19T14:05:00Z","lastScheduledExecutionTime":"2026-09-19T14:05:00Z","protectionPlanName":"journal-local-hourly"}]
How to tell a scheduled snapshot from a manual one
The generated name (journal-f40c53315adf8d8c-1c72d2d) is application name, a hash, and a counter. More useful are the labels NDK puts on it:
kubectl -n ndk-howto get applicationsnapshot journal-f40c53315adf8d8c-1c72d2d -o jsonpath='{.metadata.labels}' | jq .
{
"dataservices.nutanix.com/app-protection-plan": "journal-protection",
"dataservices.nutanix.com/application-name": "journal",
"dataservices.nutanix.com/application-namespace": "ndk-howto",
"dataservices.nutanix.com/protection-plan": "journal-local-hourly"
}
(plus the UIDs of both plans). So kubectl get applicationsnapshot -l dataservices.nutanix.com/protection-plan=journal-local-hourly lists exactly what one plan produced. And spec.expiresAfter is empty on a scheduled snapshot: retention is the plan's job, which is why the webhook in Part 2 only demands expiresAfter on manual ones.
Retention, observed over 25 days
One hourly run does not show retention. Our older ndk-demo application does. It has had a daily plan with retentionCount: 3 since 2026-08-24, and nobody touched it since.
kubectl -n ndk-demo get jobscheduler,protectionplan
NAME LASTACTIVATION NEXTACTIVATION
jobscheduler.scheduler.nutanix.com/demo-daily 2026-09-18T16:07:00Z 2026-09-19T16:07:00Z
NAME SCHEDULE-NAME RETENTION-COUNT AVAILABLE DEGRADED PROTECTION-TYPE
protectionplan.dataservices.nutanix.com/demo-local-plan demo-daily 3 True False async
25 daily runs. Here is what is left:
kubectl -n ndk-demo get applicationsnapshot -o custom-columns='NAME:.metadata.name,CREATED:.metadata.creationTimestamp,READY:.status.readyToUse,CONS:.status.consistencyType'
NAME CREATED READY CONS
ndk-demo-e51de6e006b519b4-1c6c2c7 2026-08-31T16:07:00Z false <none>
ndk-demo-e51de6e006b519b4-1c6d947 2026-09-04T16:07:00Z false <none>
ndk-demo-e51de6e006b519b4-1c71cc7 2026-09-16T16:07:00Z true CrashConsistent
ndk-demo-e51de6e006b519b4-1c72267 2026-09-17T16:07:00Z true CrashConsistent
ndk-demo-e51de6e006b519b4-1c72807 2026-09-18T16:07:00Z true CrashConsistent
Retention works: exactly three ready snapshots, the three most recent. The other 20 runs are gone, pruned as newer ones arrived.
But there are five objects, not three. Two runs, on 08-31 and 09-04, never reached READY. Their content objects say why:
kubectl get applicationsnapshotcontent asc-fc169e04-bc6d-4d4c-bf30-23f8a8e16614 -o jsonpath='{.status.conditions}' | jq -c '.[] | {type,status,reason}'
{"type":"Progressing","status":"False","reason":"VolumeSnapshotCreationFailedDueToBlockVolumes"}
{"type":"AppConfigAcquired","status":"True","reason":"AcquiredAppConfig"}
{"type":"VolumeSnapshotsCreated","status":"False","reason":"VolumeSnapshotCreationFailedDueToBlockVolumes"}
The full message carries the Prism API response: AUTHENTICATION_REQUIRED, a 401. On those two evenings, Prism Central rejected the service account for a few minutes (we saw the same thing during this write-up, it cleared on its own). NDK asked for the volume snapshot, got a 401, and marked the run failed. No retry.
Two conclusions from that, and they matter more than the manifests:
- Failed snapshots are not retried and not pruned. They sit outside the retention count, with
readyToUse: false, until you delete them. Over a year, a plan with occasional failures accumulates objects.readyToUseis a status field, so a label or field selector cannot find them; list them withjqand delete by name:
kubectl -n ndk-demo get applicationsnapshot -o json \
| jq -r '.items[] | select(.status.readyToUse != true) | .metadata.name'
ndk-demo-e51de6e006b519b4-1c6c2c7
ndk-demo-e51de6e006b519b4-1c6d947
kubectl -n ndk-demo delete applicationsnapshot ndk-demo-e51de6e006b519b4-1c6c2c7 ndk-demo-e51de6e006b519b4-1c6d947
Check the list before you pipe it into a delete. A snapshot that is still in progress also has readyToUse: false.
2. The plan reported healthy the whole time. AVAILABLE: True, DEGRADED: False, on both the ProtectionPlan and the AppProtectionPlan, before, during and after the failures. Those conditions describe the plan's configuration, not its results. If your monitoring watches plan status, it will never page.
What to watch instead: the age of the newest applicationsnapshot with readyToUse: true per application. If it is older than your schedule interval plus a margin, protection is broken. Part 5 looks at what NDK exposes to Prometheus today, and it is less than you would hope.
Changing a plan: three webhooks you will meet
We wanted to add replication to the running plan. That turned into a tour of NDK's admission rules, all observed on the lab.
A ProtectionPlan is immutable.
kubectl -n ndk-howto patch protectionplan journal-local-hourly --type=merge \
-p '{"spec":{"replicationConfigs":[{"replicationTargetName":"demo-wkl-02"}]}}'
The ProtectionPlan "journal-local-hourly" is invalid: spec: Invalid value: Spec is immutable for protectionPlan.dataservices.nutanix.com
Retention, schedule, replication: none of it can change after creation. You create a new plan.
Plans can be added to an AppProtectionPlan, not removed.
kubectl -n ndk-howto patch appprotectionplan journal-protection --type=merge \
-p '{"spec":{"protectionPlanNames":["journal-hourly-to-demo-wkl-02"]}}'
admission webhook "vappprotectionplan.kb.io" denied the request: removing protection plans from an AppProtectionPlan is not allowed: [journal-local-hourly]. Delete the AppProtectionPlan instead to remove protection
A bound ProtectionPlan will not delete.
kubectl -n ndk-howto delete protectionplan journal-local-hourly --timeout=30s
protectionplan.dataservices.nutanix.com "journal-local-hourly" deleted from ndk-howto namespace
error: timed out waiting for the condition on protectionplans/journal-local-hourly
The object sits in Terminating with the finalizer dataservices.nutanix.com/app-protection-plan-journal-protection, named after the binding that holds it. It goes away the moment the binding does.
So the working sequence to replace a plan is:
kubectl apply -f 09-replicated-plan.yaml # new ProtectionPlan
kubectl -n ndk-howto delete appprotectionplan journal-protection # old binding (old plan finalizes now)
kubectl apply -f 10-appprotectionplan-replicated.yaml # new binding
Five seconds after the binding was deleted, the old plan was gone and the Application was back to a single finalizer. Existing snapshots stayed: they belong to the Application, not to the plan. Then one more thing happened that is worth knowing.
Binding a plan can snapshot immediately
Our first binding, at 14:02 with a startTime three minutes ahead, did nothing until 14:05:00. The second binding, at 14:12, against a scheduler whose startTime was already in the past, took a snapshot right away:
NAME CREATED READY PLAN
journal-316d190781390250-1c72d2d 2026-09-19T14:12:53Z false journal-hourly-to-demo-wkl-02
journal-f40c53315adf8d8c-1c72d2d 2026-09-19T14:05:00Z true journal-local-hourly
journal-before-change 2026-09-19T13:45:13Z true <none>
Reasonable behaviour (a newly protected application should not wait an hour for its first copy), but plan for it on a large application, and do not rebind twenty applications at once on a busy afternoon.
Three hours later, the same list showed retention at work on this plan too:
NAME CREATED READY PLAN
journal-316d190781390250-1c72d69 2026-09-19T15:05:00Z true journal-hourly-to-demo-wkl-02
journal-316d190781390250-1c72da5 2026-09-19T16:05:00Z true journal-hourly-to-demo-wkl-02
journal-f40c53315adf8d8c-1c72d2d 2026-09-19T14:05:00Z true journal-local-hourly
journal-before-change 2026-09-19T13:45:13Z true <none>
retentionCount: 2, three runs (14:12, 15:05, 16:05), two left: the 14:12 snapshot was pruned when the 16:05 one became ready. The 14:05 snapshot from the deleted plan is still there. Retention only counts snapshots of the plan that made them, so a plan you delete leaves its snapshots behind until they expire or you remove them.
Removing scheduled protection
Same order: binding, plan, scheduler.
kubectl -n ndk-howto delete appprotectionplan journal-protection
kubectl -n ndk-howto delete protectionplan journal-hourly-to-demo-wkl-02
kubectl -n ndk-howto delete jobscheduler journal-hourly
Summary
Scheduled protection in NDK is three objects: a JobScheduler (never more often than hourly), a ProtectionPlan (type, retention, replication, immutable once created), and an AppProtectionPlan that binds a plan to an Application. The schedule fires on the second and retention keeps exactly the count you asked for. What it does not do is retry or clean up failed runs, and the plan's own status stays green through them, so monitor snapshot freshness, not plan conditions.
In Part 4 the same snapshot goes to a second cluster: Remote, ReplicationTarget, and a restore on the other side.


