Instructor

Run the room

One account stack serves the whole class, organised by company: the LiteLLM gateway, participant roles, budgets, guardrails and this website. Model access uses server-side virtual keys that are never handed out: the gateway recognises each participant by IAM role.

The account stack

instructor.py up applies terraform/account once for the whole class. Each participant's AgentCore stack (home page diagram) is a separate terraform workspace on their own laptop.

Account stack The account stack that instructor.py up applies. Participants' laptops and agent runtimes call the LiteLLM gateway with a presigned STS identity through CloudFront and an HTTPS load balancer into an ECS Fargate service of two to four tasks. Each task runs a keyless proxy, which maps the caller's role to the participant's virtual key, allows participants only the model routes, refuses guardrails that are not their own, retries empty model streams with a fallback model and gates the admin UI behind Google sign-in, in front of LiteLLM, which keeps keys, team budgets and spend in RDS Postgres and calls Amazon Bedrock. A second CloudFront distribution serves this site from S3. A guard Lambda enforces per-participant limits on tool sessions, a scheduled clean-up Lambda alerts on or purges idle participant stacks, and CloudWatch alarms and an AWS budget all alert through the SNS topic bootcamp-alerts. IAM holds the participant, evaluation and instructor roles with a workload permissions boundary; Transaction Search sends every participant's spans to aws/spans; terraform state lives in a versioned, encrypted S3 bucket. Each participant's own AgentCore stack is a separate terraform workspace on their laptop. VPC · PUBLIC + PRIVATE SUBNETS GUARDRAILS AND ALERTS Participants laptops + agent runtimes presigned STS identity Instructor instructor.py · master key /ui: Google sign-in first CloudFront · gateway llm.<prefix>.<zone> · ACM 120 s origin timeout CloudFront · site <prefix>.<zone> · Route 53 S3 bucket: this website ALB HTTPS · CloudFront only RDS Postgres keys · budgets · spend ECS Fargate · litellm-gateway 2–4 tasks, CPU + memory autoscaling keyless-proxy STS role → participant → their key model routes · own guardrails only empty stream: retry, fallback model admin gate (Google) on /ui LiteLLM virtual keys · company team caps Amazon Bedrock roster models + participants' guardrails Secrets Manager master key · DB URL bootcamp/participants/ <p>/litellm (gateway URL) Guard Lambda infra_limits per participant / team CloudTrail session starts + 5 min sweep auto-stop sessions (enforce) bootcamp-cleanup Lambda cleanup.schedule (daily) idle > after_hours: alert, or purge ships the CLI's purge.py CloudWatch alarms tool + runtime sessions, invocations gateway ECS CPU · RDS CPU guard + clean-up errors AWS Budget Project tag 80% · 100% fcst SNS bootcamp-alerts → alert_email IAM bootcamp-participant-<p>: ABAC on Participant tag bootcamp-eval-<p> · bootcamp-instructor workload boundary: no direct Bedrock Transaction Search every participant's spans in aws/spans (Task 5) up --transaction-search Terraform state in S3 bootcamp-tfstate-<account>-<region> versioned · encrypted · lock file bootstrapped by instructor.py up Participant stacks awsworkshop-<p>-* (bootcamp.py up) local state on each laptop reset · purge · clean-up remove admin site models
Solid arrows are requests, the dashed arrow is the site upload, dotted lines are alerts. On a phone, scroll the diagram sideways.

Set up the class

  1. List companies and participants

    Edit roster.json. Participants are grouped by company under teams. Each company becomes a LiteLLM team bootcamp-<company> whose budget_usd caps what its participants spend together. Each participant gets default_budget_usd unless they set their own budget_usd; optionally add principal_arn to restrict who may assume their role. Participant roles are tagged Team=<company> as well as Participant=<name>.

    roster.json
    {
      "default_budget_usd": 1,
      "dns_zone": "workshop.example.com",
      "dns_prefix": "bootcamp",
      "teams": {
        "acme": {
          "budget_usd": 20,
          "participants": {
            "alice": {},
            "bob": {"budget_usd": 2}
          }
        },
        "globex": {
          "budget_usd": 15,
          "participants": {"carol": {}}
        }
      },
      "alert_email": "instructor@example.com",
      "infra_limits": {"enforce": true, "monthly_budget_usd": 100, "per_participant": {"code_interpreter_sessions": 3, "browser_sessions": 3}},
      "cleanup": {"after_hours": 48, "enforce": false, "schedule": "rate(1 day)"}
    }
  2. Bring up the account stack (once)

    From the instructor checkout (participants clone only the bootcamp-participant branch, which has no instructor tooling):

    terminal
    uv run instructor.py up

    Deploys the shared platform, creates a bootcamp-participant-<name> role per participant and syncs their server-side virtual keys and budgets. Adding a participant later also needs instructor.py up, because it creates their IAM role.

  3. Run the pre-flight (a week ahead, and the morning of day 1)

    terminal
    uv run instructor.py preflight

    Read-only, about a minute. Fails on blockers: gateway /readyz and site not answering, an invalid roster, a roster model Bedrock will not serve (one tiny call per model with the master key), a missing participant role or key, the state bucket, and Transaction Search being off (Task 5 traces need it). Warns about an unconfirmed alert email, the evaluator quota (L-876938BB) against the class size, the budget, the Project cost tag and the clean-up schedule. It also prints the Identity Center policy the participants' permission set needs, which it cannot verify from the workshop account.

  4. Sync server-side virtual keys and budgets

    terminal
    uv run instructor.py keys

    Only syncs LiteLLM virtual keys and budgets with roster.json, for example to top up someone who ran out. Keys stay on the server and are never handed out: the gateway maps each IAM role to its key.

  5. Hand out names and SSO details

    Give each participant their PARTICIPANT name, the SSO start URL and region (see below), and point them at this site. The CLI derives their role; BOOTCAMP_ROLE_ARN is only an optional override. A few days ahead, ask everyone to follow Before the workshop and send you the result of uv run bootcamp.py doctor.

  6. Watch the class during the day

    terminal
    uv run instructor.py class --watch 30
    uv run instructor.py spend

    class prints one row per participant: company, stage (from their deployed AgentCore resources), LiteLLM spend and budget, tracked Code Interpreter / Browser sessions, open guardrail alerts and last activity. spend prints spend per company team, then per participant.

  7. Reset a broken participant stack

    terminal
    uv run instructor.py reset <name>           # dry run: lists their resources
    uv run instructor.py reset <name> --force   # deletes them

    Like purge, scoped to one participant. They then run uv run bootcamp.py forget to drop their stale local terraform state and start again with up 0. Spend is not reset; raise their budget in roster.json and run keys instead.

  8. Tear down

    Ask participants to run uv run bootcamp.py down, then check for leftovers and remove the account stack:

    terminal
    uv run instructor.py purge           # dry run: lists leftover participant resources
    uv run instructor.py purge --force   # deletes them (only if the list wasn't empty)
    uv run instructor.py down

    instructor.py down refuses while any participant resources still exist and lists them; instructor.py down --force purges first, then destroys. Forgotten stacks are also caught by the scheduled clean-up (see Infra guardrails).

Participant AWS access (SSO)

Participants sign in with AWS IAM Identity Center: share the SSO start URL and region; they run aws configure sso and put the profile in AWS_PROFILE. The CLI then assumes bootcamp-participant-<name>.

Each role's trust defaults to the account root, so the participants' permission set must allow sts:AssumeRole on arn:aws:iam::<account>:role/bootcamp/bootcamp-participant-*. To pin a role to one person, set principal_arn for that participant in roster.json.

Infra guardrails

LLM spend is capped by LiteLLM: per-participant budgets inside per-company team budgets (HTTP 429 budget_exceeded). AWS spend is watched by an infra guard Lambda, CloudWatch alarms and an AWS Budget, all alerting to the SNS topic bootcamp-alerts.

Per-participant limits

The guard reacts within seconds to new Code Interpreter / Browser sessions (a dedicated CloudTrail trail, bootcamp-agentcore-sessions, logs only those two session-start data events; nothing else in the sandbox is logged) and sweeps live resources every 5 minutes. It attributes usage by caller role (bootcamp-participant-<p> / awsworkshop-<p>-*) and resource name. Alerts are de-duplicated per hour.

Limit (infra_limits.per_participant)DefaultWhen exceeded
runtimes4alert
gateways2alert
memories2alert
code_interpreters2alert
browsers2alert
code_interpreter_sessions3auto-stop newest (if enforce)
browser_sessions3auto-stop newest (if enforce)
runtime_invocations_5min300alert

Override per team (teams.<t>.infra_limits) or per participant (teams.<t>.participants.<p>.infra_limits). Class-wide totals default to per-participant × headcount (override with infra_limits.class). Auto-stop needs infra_limits.enforce: true; runtimes, gateways and memories are alert-only.

Alarms

CloudWatch alarm → bootcamp-alertsThreshold
Code Interpreter ActiveSessionCount (account)25
Browser ActiveSessionCount (account)25
Runtime ActiveSessionCount (account)200
InvokeAgentRuntime (account, 5 min)1500
LiteLLM gateway ECS CPU> 85%
LiteLLM RDS CPU> 80%
Guard Lambda errorsany
Clean-up Lambda errorsany

The 5-minute sweep also reports when the LiteLLM gateway autoscaling is at its max task count.

Cost

AWS Budget bootcamp-genai-bootcamp-tagged on cost tag Project=genai-bootcamp (infra_limits.monthly_budget_usd, default $100), notifying at 80% actual and 100% forecast. AWS-managed tool sessions can't be tagged, so set infra_limits.account_budget_usd > 0 if you also want an account-wide budget.

Scheduled clean-up

The Lambda bootcamp-cleanup runs on cleanup.schedule (daily by default). It finds participant stacks with the same discovery as instructor.py purge (it ships the CLI's own purge.py) and treats a participant as idle when none of their roles (bootcamp-participant-<p>, awsworkshop-<p>-*) was used for cleanup.after_hours (default 48, minimum 6). Idle stacks are emailed to bootcamp-alerts; with cleanup.enforce: true they are purged first, in purge's dependency order, and only then does the Lambda get delete rights. Test it any time with a dry run that never deletes:

terminal
aws lambda invoke --function-name bootcamp-cleanup --cli-binary-format raw-in-base64-out \
  --payload '{"dry_run": true}' out.json && cat out.json

Manual steps

  1. Set alert_email in roster.json and run uv run instructor.py up.
  2. Confirm the SNS subscription email. Until then nobody is emailed.
  3. Activate Project as a cost allocation tag in the payer account.
  4. Optional billing alarm: enable billing alerts, put enable_billing_alarm = true in terraform/account/terraform.tfvars, then run instructor.py up.
What instructor.py up creates and common issues
  • LiteLLM gateway on ECS Fargate (2 to 4 tasks, CPU and memory autoscaling; gateway_min_tasks / gateway_max_tasks) behind an ALB and CloudFront with a 120 s origin timeout, the only path to models. Participants call it with an OpenAI-compatible API and a presigned STS identity header (X-Amz-Sts-Identity-Url); the keyless proxy in each task maps bootcamp-participant-<name> and awsworkshop-<name>-* roles to that participant's key and forwards to LiteLLM and Bedrock. Participants may call only the model routes and /key/info (HTTP 403 otherwise), may name only guardrails tagged with their own Participant, and an empty Bedrock stream is retried once, then sent to a fallback model. The LiteLLM admin UI (/ui) asks for a Google Workspace sign-in (admin_sso) before the master key.
  • RDS database for LiteLLM keys, budgets and spend tracking. The gateway URL is stored per participant in the secret bootcamp/participants/<name>/litellm, from which the CLI fills in LITELLM_BASE_URL.
  • Per-participant IAM roles with tag-based isolation (Participant=<name>, Team=<company>, names awsworkshop-<name>-*) and a permissions boundary on every role they create that blocks direct Bedrock access. Participants can't assume their workload roles, and every role they create must carry the bootcamp-workload-boundary (an allowlist) and their Participant tag; the platform adds both automatically.
  • Observability: instructor.py up leaves the account's Transaction Search setting alone; run uv run instructor.py up --transaction-search to switch it on so participants' traces show up in GenAI Observability (Phase 2, Task 5).
  • This website, served from S3 through CloudFront (custom domain and ACM certificate when dns_zone is set).
  • Terraform state in the S3 bucket bootcamp-tfstate-<account>-<region> (versioned, encrypted, S3 lock file), created by instructor.py up before terraform init. Participant stacks keep local state on their laptops.

Common issues: the first up 0 waits about 5 minutes for Cognito DNS (domains look like ds-<10 hex>-<account>); if someone queried the domain too early, flush their DNS cache. instructor.py up warns loudly when alert_email is empty. Stage 5 checks may WARN for about 10 minutes until spans arrive. A participant whose own budget or company team budget is exhausted gets HTTP 429 budget_exceeded from LiteLLM: check instructor.py spend and raise their budget with instructor.py keys.