4.2 KiB
AWS SageMaker HyperPod Example
SageMaker HyperPod clusters in TypeScript, covering both orchestrators and
every workload tier — from sbatch on a Slurm login node to effectful
TypeScript jobs arbitrated by task governance.
HyperPod is a persistent, resilient fleet of ML compute. How work lands on it depends on the orchestrator you pick at creation:
| Slurm (default) | EKS (orchestrator: { Eks }) |
|
|---|---|---|
| Provision | alchemy.run.ts |
eks.run.ts |
| Low-level workloads | ssm start-session → sbatch |
Kubernetes.Manifest (or kubectl) |
| High-level workloads | — (no submission API) | Kubernetes.Job / Kubernetes.Deployment via HyperPod attributes |
| Governance | Slurm accounting | ClusterSchedulerConfig + ComputeQuota (Kueue) |
The Slurm stack (alchemy.run.ts)
src/infra.ts— the lifecycle-script bucket and the instance execution role (inline grant built withOutput.interpolate).src/lifecycle.ts— a deploy-timeAlchemy.Actionthat uploadson_create.shwith the Effect-native distilled SDK (s3.putObject). Bucket → script → cluster ordering is inferred from the data flow.- Slurm requires a
LifeCycleConfigper instance group — that's where real clusters install the scheduler, mount FSx, and wire observability.
bun run --filter aws-hyperpod-example deploy # ~5 minutes at this size
bun run --filter aws-hyperpod-example destroy
Workloads are submitted on the cluster — each node is an SSM target:
aws sagemaker list-cluster-nodes --cluster-name <clusterName output>
aws ssm start-session \
--target sagemaker-cluster:<cluster-id>_controller-<instance-id>
# then, on the node:
sbatch --nodes=1 train.sbatch
The EKS stack (eks.run.ts)
-
src/eks-infra.ts— network (private subnets + NAT), a plain EKS control plane (HyperPod supplies the nodes), the HyperPod instance role (managed policy + the EKS networking/ECR/pod identity grants), the HyperPod cluster attached viaorchestrator: { Eks: { ClusterArn } }, theamazon-sagemaker-hyperpod-taskgovernanceadd-on, a scheduler policy, and the research team's compute quota.LifeCycleConfigis required here too — the API enforces it for EKS-orchestrated instance groups. -
Low level (
eks.run.ts) — a raw batch/v1 Job applied withKubernetes.Manifest, pinned to HyperPod nodes with the well-known labels and submitted through governance with the Kueue labels:nodeSelector: hyperpod.instanceGroups.workers.nodeSelector, labels: { [AWS.SageMaker.KUEUE_QUEUE_NAME_LABEL]: researchQuota.queueName, [AWS.SageMaker.KUEUE_PRIORITY_CLASS_LABEL]: "training-priority", }, -
High level (
src/TrainJob.ts) — an effectfulKubernetes.Jobbundled from TypeScript, written in plain Kubernetes vocabulary — the HyperPod resources expose the derived values as attributes referenced through the graph: the instance-group keys carry through to the cluster's attributes as types (a typo'd name is a compile error), and the quota materializes the governed namespace and Kueue queue:yield* Kubernetes.Job("TrainJob", { cluster: eks, main: import.meta.url, namespace: researchQuota.namespace, // hyperpod-ns-research labels: { [AWS.SageMaker.KUEUE_QUEUE_NAME_LABEL]: researchQuota.queueName, [AWS.SageMaker.KUEUE_PRIORITY_CLASS_LABEL]: "training-priority", }, podTemplate: { spec: { // health-checked nodes of the `workers` group (key-typed) nodeSelector: hyperpod.instanceGroups.workers.nodeSelector, }, }, });Bindings resolve in init and land IAM on the pod-identity role, exactly like any other Kubernetes Job or Deployment on EKS.
bun alchemy deploy --config ./eks.run.ts # EKS ~10-15 min + HyperPod ~10-20 min
bun alchemy destroy --config ./eks.run.ts
Inspection
aws sagemaker describe-cluster --cluster-name <name>
aws eks update-kubeconfig --name <eksClusterName output>
kubectl get nodes -l sagemaker.amazonaws.com/node-health-status=Schedulable
kubectl get workloads -n hyperpod-ns-research # Kueue admission