t3-code-android-nightly/.repos/alchemy-effect/examples/aws-hyperpod/README.md
Julius Marminge e3c85ead63
chore(refs): sync Effect and Alchemy references to rc.115 and beta.78 (#12327)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-17 23:21:25 -07:00

4.2 KiB

AWS SageMaker HyperPod Example

SageMaker HyperPod clusters in TypeScript, covering both orchestrators and every workload tier — from sbatch on a Slurm login node to effectful TypeScript jobs arbitrated by task governance.

HyperPod is a persistent, resilient fleet of ML compute. How work lands on it depends on the orchestrator you pick at creation:

Slurm (default) EKS (orchestrator: { Eks })
Provision alchemy.run.ts eks.run.ts
Low-level workloads ssm start-session → sbatch Kubernetes.Manifest (or kubectl)
High-level workloads — (no submission API) Kubernetes.Job / Kubernetes.Deployment via HyperPod attributes
Governance Slurm accounting ClusterSchedulerConfig + ComputeQuota (Kueue)

The Slurm stack (alchemy.run.ts)

  • src/infra.ts — the lifecycle-script bucket and the instance execution role (inline grant built with Output.interpolate).
  • src/lifecycle.ts — a deploy-time Alchemy.Action that uploads on_create.sh with the Effect-native distilled SDK (s3.putObject). Bucket → script → cluster ordering is inferred from the data flow.
  • Slurm requires a LifeCycleConfig per instance group — that's where real clusters install the scheduler, mount FSx, and wire observability.
bun run --filter aws-hyperpod-example deploy    # ~5 minutes at this size
bun run --filter aws-hyperpod-example destroy

Workloads are submitted on the cluster — each node is an SSM target:

aws sagemaker list-cluster-nodes --cluster-name <clusterName output>
aws ssm start-session \
  --target sagemaker-cluster:<cluster-id>_controller-<instance-id>
# then, on the node:
sbatch --nodes=1 train.sbatch

The EKS stack (eks.run.ts)

  • src/eks-infra.ts — network (private subnets + NAT), a plain EKS control plane (HyperPod supplies the nodes), the HyperPod instance role (managed policy + the EKS networking/ECR/pod identity grants), the HyperPod cluster attached via orchestrator: { Eks: { ClusterArn } }, the amazon-sagemaker-hyperpod-taskgovernance add-on, a scheduler policy, and the research team's compute quota. LifeCycleConfig is required here too — the API enforces it for EKS-orchestrated instance groups.

  • Low level (eks.run.ts) — a raw batch/v1 Job applied with Kubernetes.Manifest, pinned to HyperPod nodes with the well-known labels and submitted through governance with the Kueue labels:

    nodeSelector: hyperpod.instanceGroups.workers.nodeSelector,
    labels: {
      [AWS.SageMaker.KUEUE_QUEUE_NAME_LABEL]: researchQuota.queueName,
      [AWS.SageMaker.KUEUE_PRIORITY_CLASS_LABEL]: "training-priority",
    },
    
  • High level (src/TrainJob.ts) — an effectful Kubernetes.Job bundled from TypeScript, written in plain Kubernetes vocabulary — the HyperPod resources expose the derived values as attributes referenced through the graph: the instance-group keys carry through to the cluster's attributes as types (a typo'd name is a compile error), and the quota materializes the governed namespace and Kueue queue:

    yield* Kubernetes.Job("TrainJob", {
      cluster: eks,
      main: import.meta.url,
      namespace: researchQuota.namespace,  // hyperpod-ns-research
      labels: {
        [AWS.SageMaker.KUEUE_QUEUE_NAME_LABEL]: researchQuota.queueName,
        [AWS.SageMaker.KUEUE_PRIORITY_CLASS_LABEL]: "training-priority",
      },
      podTemplate: {
        spec: {
          // health-checked nodes of the `workers` group (key-typed)
          nodeSelector: hyperpod.instanceGroups.workers.nodeSelector,
        },
      },
    });
    

    Bindings resolve in init and land IAM on the pod-identity role, exactly like any other Kubernetes Job or Deployment on EKS.

bun alchemy deploy --config ./eks.run.ts    # EKS ~10-15 min + HyperPod ~10-20 min
bun alchemy destroy --config ./eks.run.ts

Inspection

aws sagemaker describe-cluster --cluster-name <name>
aws eks update-kubeconfig --name <eksClusterName output>
kubectl get nodes -l sagemaker.amazonaws.com/node-health-status=Schedulable
kubectl get workloads -n hyperpod-ns-research    # Kueue admission