In Part 1 we built a small journal application and an NDK Application around it. Now we do the thing you install NDK for: take a snapshot, break the application on purpose, and bring it back. The journal writes a timestamped line every 10 seconds, so we can see to the second what came back.

All timings and outputs are from our lab (NKP 2.18, NDK 2.3.0, Nutanix Volumes on AOS 7.6) on 2026-09-19.

All manifests in this part are in github.com/Fen0l/ndk-examples, the same files we applied on the lab.

Prerequisites

  • The journal Application from Part 1, in namespace ndk-howto, with ACTIVE: True
  • kubectl access to the cluster
  • jq, because NDK puts the useful information in .status

The snapshot object

A manual snapshot is one small CR. It references the Application by name and says how long to keep the snapshot.

yaml
apiVersion: dataservices.nutanix.com/v1alpha1
kind: ApplicationSnapshot
metadata:
  name: journal-before-change
  namespace: ndk-howto
spec:
  source:
    applicationRef:
      name: journal
  expiresAfter: 72h

expiresAfter is mandatory in 2.3

In NDK 2.2 the field was optional. In 2.3 the admission webhook refuses a manual snapshot without it:

output
Error from server (Forbidden): error when creating "STDIN": admission webhook "vapplicationsnapshot.kb.io"
denied the request: spec.expiresAfter: Invalid value: null: should be provided for snapshots not created by AppProtectionPlans

The message is precise: manual snapshots must carry their own expiry, scheduled ones get it from the plan's retention policy (Part 3). Accepted values are <n>m or <n>h, up to 1440h (60 days). If you have automation from 2.2 that creates snapshots, this is the field to add. And as we saw in Part 1, the value is immutable once set.

Take the snapshot

Note the last journal line first. That is our reference point.

bash
kubectl -n ndk-howto exec deploy/journal-writer -- tail -1 /data/journal.log
kubectl apply -f 03-manual-snapshot.yaml
kubectl -n ndk-howto get applicationsnapshot -w
output
2026-09-19T13:45:10 entry
applicationsnapshot.dataservices.nutanix.com/journal-before-change created
NAME                    AGE   READY-TO-USE   BOUND-SNAPSHOTCONTENT                      SNAPSHOT-AGE   CONSISTENCY-TYPE
journal-before-change   5s    false          asc-fd129a47-fc53-4572-91b3-2c473555a3c9
journal-before-change   26s   false          asc-fd129a47-fc53-4572-91b3-2c473555a3c9
journal-before-change   31s   true           asc-fd129a47-fc53-4572-91b3-2c473555a3c9   2s             CrashConsistent

31 seconds from apply to READY-TO-USE: true, for one 2 GiB volume plus three manifests. Most of that time is Prism creating the volume group snapshot.

The BOUND-SNAPSHOTCONTENT column shows the cluster-scoped object NDK created and bound. It holds the step-by-step conditions:

bash
kubectl get applicationsnapshotcontent asc-fd129a47-fc53-4572-91b3-2c473555a3c9 \
  -o jsonpath='{.status.conditions}' | jq -c '.[] | {type,status,reason}'
output
{"type":"Progressing","status":"False","reason":"ApplicationSnapshotReady"}
{"type":"AppConfigAcquired","status":"True","reason":"AcquiredAppConfig"}
{"type":"VolumeSnapshotsCreated","status":"True","reason":"VolumeSnapshotCreationSucceeded"}
{"type":"ApplicationSnapshotFinalized","status":"True","reason":"ApplicationSnapshotFinalized"}

Two phases: acquire the application config (the manifests), then snapshot the volumes. When a snapshot fails, VolumeSnapshotsCreated is where the Prism error message lands. We had one earlier that day when Prism Central briefly rejected the service account, and the full 401 body was in that condition. Look there before looking at pod logs.

The snapshot's own status summarises what is inside:

bash
kubectl -n ndk-howto get applicationsnapshot journal-before-change -o yaml
yaml
status:
  boundApplicationSnapshotContentName: asc-fd129a47-fc53-4572-91b3-2c473555a3c9
  consistencyType: CrashConsistent
  creationTime: "2026-09-19T13:45:42Z"
  expirationTime: "2026-09-22T13:45:42Z"
  readyToUse: true
  summary:
    snapshotArtifactsByNamespace:
      ndk-howto:
        apps/v1/Deployment:
        - name: journal-writer
        v1/ConfigMap:
        - name: journal-config
        v1/PersistentVolumeClaim:
        - name: journal-data

creationTime is the moment Prism took the volume snapshot. Keep that timestamp in mind, it matters below.

Break the application

We delete the Deployment and the PVC. Both. The PVC deletion removes the Nutanix volume group behind it, so this is a real data loss, not a pod restart.

bash
kubectl -n ndk-howto exec deploy/journal-writer -- sh -c 'wc -l /data/journal.log; tail -1 /data/journal.log'
kubectl -n ndk-howto delete deployment journal-writer
kubectl -n ndk-howto delete pvc journal-data
kubectl -n ndk-howto get deploy,pvc,pod
output
132 /data/journal.log
2026-09-19T13:47:10 entry
deployment.apps "journal-writer" deleted from ndk-howto namespace
persistentvolumeclaim "journal-data" deleted from ndk-howto namespace
No resources found in ndk-howto namespace.

The journal had 132 lines and the last one was written at 13:47:10. The ConfigMap is still there on purpose: we want to see how NDK handles an object that already exists.

Restore

Again one CR. It references the snapshot, nothing else.

yaml
apiVersion: dataservices.nutanix.com/v1alpha1
kind: ApplicationSnapshotRestore
metadata:
  name: journal-restore-1
  namespace: ndk-howto
spec:
  applicationSnapshotName: journal-before-change
bash
kubectl apply -f 04-restore.yaml
kubectl -n ndk-howto get applicationsnapshotrestore -w
output
NAME                SNAPSHOT-NAME           COMPLETED
journal-restore-1   journal-before-change   false
journal-restore-1   journal-before-change   false
journal-restore-1   journal-before-change   true

Started 13:51:50, finished 13:52:05. Fifteen seconds. The conditions show the order of operations:

bash
kubectl -n ndk-howto get applicationsnapshotrestore journal-restore-1 \
  -o jsonpath='{.status.conditions}' | jq -r '.[] | "\(.lastTransitionTime) \(.type): \(.message)"'
output
2026-09-19T13:51:50Z PrechecksPassed: All prechecks passed and finalizers on dependent resources set. Skipped restoring resources since they already exist: ["/v1, Kind=ConfigMap, ndk-howto/journal-config"]
2026-09-19T13:51:50Z VolumeRestoreRequestsSubmitted: Restore requests for all eligible volumes submitted
2026-09-19T13:51:50Z ApplicationConfigRestored: All eligible application configs restored
2026-09-19T13:52:05Z VolumesRestored: All eligible volumes restored
2026-09-19T13:52:05Z ApplicationRestoreFinalised: Application restore successfully finalised

Read the first line carefully. The ConfigMap existed, NDK skipped it and kept the live one. Manifests came back in under a second, the volume took 15 seconds. The Deployment was recreated, Kubernetes scheduled a new pod, and it mounted a new PVC backed by a clone of the snapshot.

bash
kubectl -n ndk-howto get deploy,pvc,pod
output
NAME                             READY   UP-TO-DATE   AVAILABLE   AGE
deployment.apps/journal-writer   1/1     1            1           72s

NAME                                 STATUS   VOLUME                                     CAPACITY   ACCESS MODES   STORAGECLASS
persistentvolumeclaim/journal-data   Bound    pvc-ea32fe25-ccb1-4f8f-842c-1398b1fda24e   2Gi        RWO            nutanix-volume

NAME                                  READY   STATUS    RESTARTS   AGE
pod/journal-writer-6fb7556dcc-6976d   1/1     Running   0          72s

Same PVC name, new volume ID. Same Deployment, new pod.

An existing PVC blocks the restore, an existing ConfigMap does not

We hit this the first time by mistake, and it is worth a section. Our first attempt deleted only the Deployment. The PVC was still there. NDK refused:

output
PrechecksPassed=False ResourcesAlreadyExist:
Resources to restore already exist in the kubernetes cluster: ["/v1, Kind=PersistentVolumeClaim, ndk-howto/journal-data"]

The restore object stays with COMPLETED: false and this message, and nothing on the cluster is touched. That is the behaviour you want from a restore tool: it will not silently replace a volume that still holds data.

The rule is not a heuristic, it is a flag on the controller. The 2.3 chart starts the manager with:

output
--skip_restoring_existing_group_kinds=ServiceAccount,ConfigMap,Secret,RoleBinding.rbac.authorization.k8s.io,Role.rbac.authorization.k8s.io,...

Objects of those kinds that already exist are skipped and the live copy wins. Anything else that already exists, the PVC in our case, fails the precheck. Note what is on the list: identity and configuration, the things you typically manage outside the application (a Secret from an external secrets operator, a ServiceAccount bound to cloud IAM). If you want a clean restore of a half-broken application, delete the PVC yourself first, or restore into another namespace or cluster (Part 4).

One more practical note: kubectl delete pvc blocks while a pod still uses the volume (the kubernetes.io/pvc-protection finalizer). Delete the workload first, wait for the pod to go, then the PVC goes immediately.

The data: what came back, to the second

This is the part a datasheet cannot give you.

bash
kubectl -n ndk-howto exec deploy/journal-writer -- sh -c 'wc -l /data/journal.log; sed -n "119,122p" /data/journal.log'
output
126 /data/journal.log
2026-09-19T13:45:00 entry
2026-09-19T13:45:10 entry
2026-09-19T13:52:11 entry
2026-09-19T13:52:21 entry

Three facts in four lines:

  1. 120 lines were restored. The file had 120 lines at 13:45:10, and 132 when we deleted it at 13:47. The 12 lines written after the snapshot are gone, as expected. The snapshot is a point in time, not a sync.
  2. The writer resumed at 13:52:11, six seconds after VolumesRestored. Line 121 onward is new data on the restored volume.
  3. The last restored line is 13:45:10, but the snapshot was taken at 13:45:42. Thirty seconds of writes are missing from a snapshot that says CrashConsistent.

That third point is the one to understand. Our writer appends with a shell redirect and never calls fsync. Those three entries (13:45:20, 13:45:30, 13:45:40) were in the Linux page cache, not on the block device, when Prism snapshotted the volume group. A crash-consistent snapshot captures the disk exactly as it would look after a power loss. Anything the application had not flushed is not there.

For a log writer that is harmless. For a database it is the difference between a clean start and a crash recovery on the restored volume. Real databases are built for that: their crash recovery replays the journal or WAL and comes up consistent, which is why crash-consistent snapshots are an accepted backup method for them. But an application that buffers writes in memory and never flushes will lose those writes in every crash-consistent snapshot, from any vendor. If that matters to you, the answer is application-consistent snapshots (quiesce hooks) or sync replication, both outside the scope of this part.

Cleaning up snapshots

A manual snapshot expires on its own after expiresAfter. To remove one early:

bash
kubectl -n ndk-howto delete applicationsnapshot journal-before-change

We tested it with a throwaway snapshot: the bound ApplicationSnapshotContent was gone within a few seconds of deleting the ApplicationSnapshot. The restore object (journal-restore-1) is just a record. We deleted it right after the restore and the pod and PVC were not affected.

Summary

Two small CRs give you a point-in-time copy and a restore: ApplicationSnapshot (31 seconds on our lab) and ApplicationSnapshotRestore (15 seconds). NDK refuses to overwrite an existing volume and skips configuration objects that already exist, which is exactly the safety you want. And a crash-consistent snapshot holds what was on disk, not what was in memory, which our timestamped journal shows to the second.

In Part 3 we stop taking snapshots by hand: JobScheduler, ProtectionPlan and AppProtectionPlan, with a retention policy we have watched run for four weeks.