In Part 1 we built a small journal application and an NDK Application around it. Now we do the thing you install NDK for: take a snapshot, break the application on purpose, and bring it back. The journal writes a timestamped line every 10 seconds, so we can see to the second what came back.
All timings and outputs are from our lab (NKP 2.18, NDK 2.3.0, Nutanix Volumes on AOS 7.6) on 2026-09-19.
All manifests in this part are in github.com/Fen0l/ndk-examples, the same files we applied on the lab.
Prerequisites
- The
journalApplication from Part 1, in namespacendk-howto, withACTIVE: True kubectlaccess to the clusterjq, because NDK puts the useful information in.status
The snapshot object
A manual snapshot is one small CR. It references the Application by name and says how long to keep the snapshot.
apiVersion: dataservices.nutanix.com/v1alpha1
kind: ApplicationSnapshot
metadata:
name: journal-before-change
namespace: ndk-howto
spec:
source:
applicationRef:
name: journal
expiresAfter: 72h
expiresAfter is mandatory in 2.3
In NDK 2.2 the field was optional. In 2.3 the admission webhook refuses a manual snapshot without it:
Error from server (Forbidden): error when creating "STDIN": admission webhook "vapplicationsnapshot.kb.io"
denied the request: spec.expiresAfter: Invalid value: null: should be provided for snapshots not created by AppProtectionPlans
The message is precise: manual snapshots must carry their own expiry, scheduled ones get it from the plan's retention policy (Part 3). Accepted values are <n>m or <n>h, up to 1440h (60 days). If you have automation from 2.2 that creates snapshots, this is the field to add. And as we saw in Part 1, the value is immutable once set.
Take the snapshot
Note the last journal line first. That is our reference point.
kubectl -n ndk-howto exec deploy/journal-writer -- tail -1 /data/journal.log
kubectl apply -f 03-manual-snapshot.yaml
kubectl -n ndk-howto get applicationsnapshot -w
2026-09-19T13:45:10 entry
applicationsnapshot.dataservices.nutanix.com/journal-before-change created
NAME AGE READY-TO-USE BOUND-SNAPSHOTCONTENT SNAPSHOT-AGE CONSISTENCY-TYPE
journal-before-change 5s false asc-fd129a47-fc53-4572-91b3-2c473555a3c9
journal-before-change 26s false asc-fd129a47-fc53-4572-91b3-2c473555a3c9
journal-before-change 31s true asc-fd129a47-fc53-4572-91b3-2c473555a3c9 2s CrashConsistent
31 seconds from apply to READY-TO-USE: true, for one 2 GiB volume plus three manifests. Most of that time is Prism creating the volume group snapshot.
The BOUND-SNAPSHOTCONTENT column shows the cluster-scoped object NDK created and bound. It holds the step-by-step conditions:
kubectl get applicationsnapshotcontent asc-fd129a47-fc53-4572-91b3-2c473555a3c9 \
-o jsonpath='{.status.conditions}' | jq -c '.[] | {type,status,reason}'
{"type":"Progressing","status":"False","reason":"ApplicationSnapshotReady"}
{"type":"AppConfigAcquired","status":"True","reason":"AcquiredAppConfig"}
{"type":"VolumeSnapshotsCreated","status":"True","reason":"VolumeSnapshotCreationSucceeded"}
{"type":"ApplicationSnapshotFinalized","status":"True","reason":"ApplicationSnapshotFinalized"}
Two phases: acquire the application config (the manifests), then snapshot the volumes. When a snapshot fails, VolumeSnapshotsCreated is where the Prism error message lands. We had one earlier that day when Prism Central briefly rejected the service account, and the full 401 body was in that condition. Look there before looking at pod logs.
The snapshot's own status summarises what is inside:
kubectl -n ndk-howto get applicationsnapshot journal-before-change -o yaml
status:
boundApplicationSnapshotContentName: asc-fd129a47-fc53-4572-91b3-2c473555a3c9
consistencyType: CrashConsistent
creationTime: "2026-09-19T13:45:42Z"
expirationTime: "2026-09-22T13:45:42Z"
readyToUse: true
summary:
snapshotArtifactsByNamespace:
ndk-howto:
apps/v1/Deployment:
- name: journal-writer
v1/ConfigMap:
- name: journal-config
v1/PersistentVolumeClaim:
- name: journal-data
creationTime is the moment Prism took the volume snapshot. Keep that timestamp in mind, it matters below.
Break the application
We delete the Deployment and the PVC. Both. The PVC deletion removes the Nutanix volume group behind it, so this is a real data loss, not a pod restart.
kubectl -n ndk-howto exec deploy/journal-writer -- sh -c 'wc -l /data/journal.log; tail -1 /data/journal.log'
kubectl -n ndk-howto delete deployment journal-writer
kubectl -n ndk-howto delete pvc journal-data
kubectl -n ndk-howto get deploy,pvc,pod
132 /data/journal.log
2026-09-19T13:47:10 entry
deployment.apps "journal-writer" deleted from ndk-howto namespace
persistentvolumeclaim "journal-data" deleted from ndk-howto namespace
No resources found in ndk-howto namespace.
The journal had 132 lines and the last one was written at 13:47:10. The ConfigMap is still there on purpose: we want to see how NDK handles an object that already exists.
Restore
Again one CR. It references the snapshot, nothing else.
apiVersion: dataservices.nutanix.com/v1alpha1
kind: ApplicationSnapshotRestore
metadata:
name: journal-restore-1
namespace: ndk-howto
spec:
applicationSnapshotName: journal-before-change
kubectl apply -f 04-restore.yaml
kubectl -n ndk-howto get applicationsnapshotrestore -w
NAME SNAPSHOT-NAME COMPLETED
journal-restore-1 journal-before-change false
journal-restore-1 journal-before-change false
journal-restore-1 journal-before-change true
Started 13:51:50, finished 13:52:05. Fifteen seconds. The conditions show the order of operations:
kubectl -n ndk-howto get applicationsnapshotrestore journal-restore-1 \
-o jsonpath='{.status.conditions}' | jq -r '.[] | "\(.lastTransitionTime) \(.type): \(.message)"'
2026-09-19T13:51:50Z PrechecksPassed: All prechecks passed and finalizers on dependent resources set. Skipped restoring resources since they already exist: ["/v1, Kind=ConfigMap, ndk-howto/journal-config"]
2026-09-19T13:51:50Z VolumeRestoreRequestsSubmitted: Restore requests for all eligible volumes submitted
2026-09-19T13:51:50Z ApplicationConfigRestored: All eligible application configs restored
2026-09-19T13:52:05Z VolumesRestored: All eligible volumes restored
2026-09-19T13:52:05Z ApplicationRestoreFinalised: Application restore successfully finalised
Read the first line carefully. The ConfigMap existed, NDK skipped it and kept the live one. Manifests came back in under a second, the volume took 15 seconds. The Deployment was recreated, Kubernetes scheduled a new pod, and it mounted a new PVC backed by a clone of the snapshot.
kubectl -n ndk-howto get deploy,pvc,pod
NAME READY UP-TO-DATE AVAILABLE AGE
deployment.apps/journal-writer 1/1 1 1 72s
NAME STATUS VOLUME CAPACITY ACCESS MODES STORAGECLASS
persistentvolumeclaim/journal-data Bound pvc-ea32fe25-ccb1-4f8f-842c-1398b1fda24e 2Gi RWO nutanix-volume
NAME READY STATUS RESTARTS AGE
pod/journal-writer-6fb7556dcc-6976d 1/1 Running 0 72s
Same PVC name, new volume ID. Same Deployment, new pod.
An existing PVC blocks the restore, an existing ConfigMap does not
We hit this the first time by mistake, and it is worth a section. Our first attempt deleted only the Deployment. The PVC was still there. NDK refused:
PrechecksPassed=False ResourcesAlreadyExist:
Resources to restore already exist in the kubernetes cluster: ["/v1, Kind=PersistentVolumeClaim, ndk-howto/journal-data"]
The restore object stays with COMPLETED: false and this message, and nothing on the cluster is touched. That is the behaviour you want from a restore tool: it will not silently replace a volume that still holds data.
The rule is not a heuristic, it is a flag on the controller. The 2.3 chart starts the manager with:
--skip_restoring_existing_group_kinds=ServiceAccount,ConfigMap,Secret,RoleBinding.rbac.authorization.k8s.io,Role.rbac.authorization.k8s.io,...
Objects of those kinds that already exist are skipped and the live copy wins. Anything else that already exists, the PVC in our case, fails the precheck. Note what is on the list: identity and configuration, the things you typically manage outside the application (a Secret from an external secrets operator, a ServiceAccount bound to cloud IAM). If you want a clean restore of a half-broken application, delete the PVC yourself first, or restore into another namespace or cluster (Part 4).
One more practical note: kubectl delete pvc blocks while a pod still uses the volume (the kubernetes.io/pvc-protection finalizer). Delete the workload first, wait for the pod to go, then the PVC goes immediately.
The data: what came back, to the second
This is the part a datasheet cannot give you.
kubectl -n ndk-howto exec deploy/journal-writer -- sh -c 'wc -l /data/journal.log; sed -n "119,122p" /data/journal.log'
126 /data/journal.log
2026-09-19T13:45:00 entry
2026-09-19T13:45:10 entry
2026-09-19T13:52:11 entry
2026-09-19T13:52:21 entry
Three facts in four lines:
- 120 lines were restored. The file had 120 lines at 13:45:10, and 132 when we deleted it at 13:47. The 12 lines written after the snapshot are gone, as expected. The snapshot is a point in time, not a sync.
- The writer resumed at 13:52:11, six seconds after
VolumesRestored. Line 121 onward is new data on the restored volume. - The last restored line is 13:45:10, but the snapshot was taken at 13:45:42. Thirty seconds of writes are missing from a snapshot that says
CrashConsistent.
That third point is the one to understand. Our writer appends with a shell redirect and never calls fsync. Those three entries (13:45:20, 13:45:30, 13:45:40) were in the Linux page cache, not on the block device, when Prism snapshotted the volume group. A crash-consistent snapshot captures the disk exactly as it would look after a power loss. Anything the application had not flushed is not there.
For a log writer that is harmless. For a database it is the difference between a clean start and a crash recovery on the restored volume. Real databases are built for that: their crash recovery replays the journal or WAL and comes up consistent, which is why crash-consistent snapshots are an accepted backup method for them. But an application that buffers writes in memory and never flushes will lose those writes in every crash-consistent snapshot, from any vendor. If that matters to you, the answer is application-consistent snapshots (quiesce hooks) or sync replication, both outside the scope of this part.
Cleaning up snapshots
A manual snapshot expires on its own after expiresAfter. To remove one early:
kubectl -n ndk-howto delete applicationsnapshot journal-before-change
We tested it with a throwaway snapshot: the bound ApplicationSnapshotContent was gone within a few seconds of deleting the ApplicationSnapshot. The restore object (journal-restore-1) is just a record. We deleted it right after the restore and the pod and PVC were not affected.
Summary
Two small CRs give you a point-in-time copy and a restore: ApplicationSnapshot (31 seconds on our lab) and ApplicationSnapshotRestore (15 seconds). NDK refuses to overwrite an existing volume and skips configuration objects that already exist, which is exactly the safety you want. And a crash-consistent snapshot holds what was on disk, not what was in memory, which our timestamped journal shows to the second.
In Part 3 we stop taking snapshots by hand: JobScheduler, ProtectionPlan and AppProtectionPlan, with a retention policy we have watched run for four weeks.


