Empire AI: Storage Architecture & User Workflow Guide

Created by Cesar Arias, Modified on Thu, 1 Oct at 5:24 PM by Cesar Arias

Empire AI: Storage Architecture & User Workflow Guide

Environment: Phase Alpha++ & Phase Beta  |  Last updated: September 2026

System overview: Empire AI pairs an all-flash VAST DataStore (~20 PB) with a parallel DDN EXAScaler Lustre filesystem (~10 PB). Following best practices from leading supercomputing facilities, both tiers are presented as static, distinct mount points across compute nodes for stability, predictable I/O, and maximum GPU throughput.

1. Storage hierarchy summary

/projects and /ddn/scratch are organized around different owners. /projects is project-level storage, namespaced by PI: <institution>/<pi_username>/<project_keyword>. /ddn/scratch is ephemeral, self-service working space namespaced by the individual researcher: <institution>/<user_name>/<project>, with no admin or PI involvement required to create project subdirectories.

Tip: Naming your scratch project folder to match your project_keyword in /projects can help research peers find your working files. This is a naming convention only; it doesn't change who can access the files.

TierPathApplianceQuotaBacked up?Purge policy
Home/mnt/home/$USERVAST (All-Flash)100 GB (hard limit)NoNone
Project/projects/<institution>/<pi_username>/<project_keyword>/VAST (All-Flash)Per project allocation (based on approved Beta request)NoNone
Scratch/ddn/scratch/<institution>/<user_name>/<project>/DDN EXAScaler (Lustre)5 TB per user (not SU-charged)No30-day unaccessed-file purge
Cluster availability: /ddn/scratch is currently available on the Beta and Grace nodes only as of September 2026. On Alpha nodes, use /projects for all job I/O until the Alpha Lustre rollout is complete.

A shared, read-only dataset tier for consortium foundation models may be added in the future.

Service Units (SUs): /projects usage counts against your project's SU allocation at 8.333 SU per TB per month (capped at 100 TB per project). /ddn/scratch is not SU-charged, but it's still shared, finite capacity, not a place for long-term free storage.
Purge policy warning: Scratch is temporary. Anything under /ddn/scratch/... that goes unaccessed for 30 days is subject to automatic removal. Never store the only copy of a result there. Always promote important files to /projects/....

2. How the two tiers work together

Once your dataset is on /projects, jobs read it directly. No staging step is needed by default, so you don't have to copy data onto scratch first (for repeat-heavy, I/O-bound jobs, see the FAQ). On the way out, only the files you actually want to keep need to move from scratch to /projects, while the rest can be removed or left to age out.

Storage tierMount pathRoleWhat happens during a job
VAST All-Flash/projects/<inst>/<pi>/<project>/Persistent storage — input and final outputCode and training data stay here long-term and are read directly over NFS, with no pre-copying required.
DDN Lustre/ddn/scratch/<inst>/<user>/<project>/Working storage — active checkpoints, self-managed per researcherThe training script writes checkpoints here every epoch or on a timer. Lustre spreads large writes across many storage targets so checkpoints finish quicker and training spends less time waiting on them.

3. VAST vs. DDN Lustre

  • VAST — All-flash, low-latency storage mounted directly over NFS. Your code, data, and curated weights are kept here long-term (no purge).
  • DDN Lustre — A parallel filesystem built for fast, large checkpoint writes, whether a single rank writes the whole checkpoint or many ranks write their pieces at once. Lustre spreads each file across many storage targets so the write finishes quickly.
  • The scratch fileset has a progressive file layout (PFL) applied by default, scaling stripe width with file size automatically—see Section 6.

4. Step-by-step: what happens during a training job

Step 1 — Read directly from VAST

Point your dataset path straight at /projects/.... VAST is mounted directly, so training begins the moment the job is scheduled. You still need to upload your dataset to /projects initially. Use Globus for large transfers — see Transfer Data With Globus.

Step 2 — Write checkpoints to Lustre (/ddn/scratch)

In sharded training, many ranks can each write their piece of a checkpoint at the same second, creating a large synchronized burst. Standard DDP training usually gathers the model state and writes the checkpoint from a single rank, which also works well here. Lustre stripes the write across many OSTs, allowing the burst to finish quickly. If a job hits its wallclock limit or a node fails, the next job can resume from the latest checkpoint in /ddn/scratch/....

Step 3 — Promote the winner back to VAST (/projects)

A run might leave 50 intermediate checkpoints behind, consuming several terabytes of scratch. Select the best checkpoint and copy that file back to /projects/.../models/, where it is retained for fine-tuning, evaluation, or publication. If you are interested in training data like loss curves, use a tool like Weights & Biases or save to a local file and also copy that back to VAST.

Step 4 — Remove or let scratch expire

Once you've promoted what you need, clean up interim checkpoints yourself. The reference script only deletes scratch after the copy back to /projects succeeds. Or leave them. The 30-day unaccessed-file purge on /ddn/scratch removes anything left behind.

5. Reference Slurm template

This sample pattern reads from VAST, checkpoints to Lustre, promotes the result and logs, and cleans up only after the copies succeed:

#!/bin/bash
#SBATCH --job-name=my_training_job
#SBATCH --nodes=2
#SBATCH --gpus-per-node=8
#SBATCH --time=08:00:00

# Stop on any error, so cleanup below only runs if every copy succeeds
set -euo pipefail

# 1. Project storage (VAST) is namespaced by your PI
INSTITUTION="<institution>"
PI="<pi_username>"
PROJECT="<project_keyword>"
PROJECT_DIR="/projects/$INSTITUTION/$PI/$PROJECT"

# 2. Scratch (Lustre) is namespaced by YOU ($USER is automatically supplied)
# Reusing $PROJECT here makes it easy for research peers to find working files.
SCRATCH_DIR="/ddn/scratch/$INSTITUTION/$USER/$PROJECT/$SLURM_JOB_ID"
CKPT_DIR="$SCRATCH_DIR/checkpoints"
LOG_DIR="$SCRATCH_DIR/logs"

# 3. Create job's scratch workspace on Lustre (no admin request needed)
mkdir -p "$CKPT_DIR" "$LOG_DIR"

# 4. Run training
# - Read data and code directly from VAST (NFS-mounted, no pre-copy)
# - Write checkpoints and logs to Lustre scratch
srun python train.py \
  --train_data "$PROJECT_DIR/data/train.parquet" \
  --pretrained_model "$PROJECT_DIR/models/pretrained" \
  --checkpoint_dir "$CKPT_DIR" \
  --log_dir "$LOG_DIR"

# 5. Copy the best weights AND logs back to permanent VAST storage
# rsync -a can resume if a copy is interrupted
rsync -a "$CKPT_DIR/best_model.pt" "$PROJECT_DIR/models/"
rsync -a "$LOG_DIR/" "$PROJECT_DIR/models/logs_$SLURM_JOB_ID/"

# 6. Clean up scratch (only reached if both copies succeeded)
rm -rf "$SCRATCH_DIR"

6. File layout on /ddn/scratch (PFL)

Unlike VAST, where file layout is fully automated, Lustre lets administrators define how files are striped across Object Storage Targets (OSTs). The scratch fileset uses a Progressive File Layout (PFL) by default. It applies at the fileset root and is inherited by every directory underneath:

Byte rangeStripe countOST pool
0 – 1 GB1tlc_pool
1 GB – 64 GB8tlc_pool
64 GB – EOF16qlc_pool

tlc_pool and qlc_pool are two distinct flash tiers. Files under 64 GB stay on the TLC tier, while data past 64 GB is striped 16-wide on the higher-capacity QLC tier. This default applies to files created after the policy took effect; existing files keep their original layout.

To verify inherited settings for any directory:


lfs getstripe -d /ddn/scratch/<institution>/$USER/

If you think you need something different

Highly parallel writes to a single shared file (many MPI ranks or HDF5 collective I/O) are the main case where the default may not be ideal, because it scales by file size rather than by number of writers. If you're doing this and suspect a bottleneck, contact your Empire AI administrator rather than changing the layout yourself.

Directory organization tip: Nest pretrained weights, sweep results, and checkpoint variants inside a project's own tree rather than creating sibling directories like <project>-pretrained or <project>-v2. Keep container images in the project too, since they're reused across runs and would be lost to the scratch purge: /projects/<institution>/<pi_username>/<project_keyword>/{data,models,src,containers}

7. FAQ

Do I need my PI to create my scratch directory?

No. Unlike /projects, /ddn/scratch is namespaced by your individual username. You can manage it yourself: mkdir -p /ddn/scratch/<institution>/$USER/<project>

Do I need to copy my training dataset to /ddn/scratch before starting?

No. VAST is mounted directly and available to every job. Reading straight from /projects/... can save pre-staging time and prevents unnecessary scratch usage. One exception: if a job re-reads the same dataset over many epochs and looks I/O-bound, staging it to /ddn/scratch once may be faster than repeated reads from /projects. This hasn't been benchmarked on Empire AI, so treat direct reads as the default and test staging if needed.

Why not write checkpoints straight to /projects on VAST?

Lustre is the recommended default for checkpoints because it's built for fast, large writes, whether from one rank or many. The two tiers haven't been formally benchmarked against each other on Empire AI, so treat this as guidance rather than a measured result.

What happens if my job is killed before the final model gets copied back?

Scratch supports restarting: your checkpoints stay in /ddn/scratch/... until the 30-day unaccessed-file purge, so a new job pointed at the same directory can resume. But scratch is not backed up or protected against storage incidents. For a run you can't afford to lose, promote important checkpoints to /projects periodically during the run.

Why use rsync or cp instead of Globus to move files between scratch and /projects?

Both filesystems are mounted on the compute node, so a local copy is the right tool inside a job. Globus runs asynchronously through separate data-transfer nodes and can't be called mid-script, and /ddn/scratch isn't a Globus endpoint. Use Globus to bring external data onto Empire AI (see Step 1).

8. Do's and don'ts

DoDon't
Read training data straight from /projects/...Pre-copy datasets to scratch "just in case" — it wastes pool capacity and staging time
Write frequent checkpoints to /ddn/scratch/.../$SLURM_JOB_ID/checkpoints/Write high-frequency checkpoints directly to VAST
Promote final/best model weights and training logs back to /projectsLeave your only copy of a result sitting on scratch past 30 days
Keep one project_keyword per logical project, nesting variants underneathCreate sibling top-level directories like project-pretrained or project-v2
Name your /ddn/scratch project subdirectory to match your /projects keywordAssume peers can find scratch files if using inconsistent or private naming
Use /mnt/home only for dotfiles, small scripts, and configsStore datasets or model weights in /mnt/home (100 GB hard limit)
Treat /ddn/scratch as temporary working space within your 5 TB quotaUse scratch as free long-term storage because it isn't SU-charged

Was this article helpful?

That’s Great!

Thank you for your feedback

Sorry! We couldn't be helpful

Thank you for your feedback

Let us know how can we improve this article!

Select at least one of the reasons

Feedback sent

We appreciate your effort and will try to fix the article