Instructor
Run the room
One account stack serves the whole class, organised by company: the LiteLLM gateway, participant roles, budgets, guardrails and this website. Model access uses server-side virtual keys that are never handed out: the gateway recognises each participant by IAM role.
The account stack
instructor.py up applies terraform/account once for the whole class. Each participant's AgentCore stack (home page diagram) is a separate terraform workspace on their own laptop.
Set up the class
List companies and participants
Edit
roster.json. Participants are grouped by company underteams. Each company becomes a LiteLLM teambootcamp-<company>whosebudget_usdcaps what its participants spend together. Each participant getsdefault_budget_usdunless they set their ownbudget_usd; optionally addprincipal_arnto restrict who may assume their role. Participant roles are taggedTeam=<company>as well asParticipant=<name>.roster.json{ "default_budget_usd": 1, "dns_zone": "workshop.example.com", "dns_prefix": "bootcamp", "teams": { "acme": { "budget_usd": 20, "participants": { "alice": {}, "bob": {"budget_usd": 2} } }, "globex": { "budget_usd": 15, "participants": {"carol": {}} } }, "alert_email": "instructor@example.com", "infra_limits": {"enforce": true, "monthly_budget_usd": 100, "per_participant": {"code_interpreter_sessions": 3, "browser_sessions": 3}}, "cleanup": {"after_hours": 48, "enforce": false, "schedule": "rate(1 day)"} }Bring up the account stack (once)
From the instructor checkout (participants clone only the
bootcamp-participantbranch, which has no instructor tooling):terminaluv run instructor.py upDeploys the shared platform, creates a
bootcamp-participant-<name>role per participant and syncs their server-side virtual keys and budgets. Adding a participant later also needsinstructor.py up, because it creates their IAM role.Run the pre-flight (a week ahead, and the morning of day 1)
terminaluv run instructor.py preflightRead-only, about a minute. Fails on blockers: gateway
/readyzand site not answering, an invalid roster, a roster model Bedrock will not serve (one tiny call per model with the master key), a missing participant role or key, the state bucket, and Transaction Search being off (Task 5 traces need it). Warns about an unconfirmed alert email, the evaluator quota (L-876938BB) against the class size, the budget, theProjectcost tag and the clean-up schedule. It also prints the Identity Center policy the participants' permission set needs, which it cannot verify from the workshop account.Sync server-side virtual keys and budgets
terminaluv run instructor.py keysOnly syncs LiteLLM virtual keys and budgets with
roster.json, for example to top up someone who ran out. Keys stay on the server and are never handed out: the gateway maps each IAM role to its key.Hand out names and SSO details
Give each participant their
PARTICIPANTname, the SSO start URL and region (see below), and point them at this site. The CLI derives their role;BOOTCAMP_ROLE_ARNis only an optional override. A few days ahead, ask everyone to follow Before the workshop and send you the result ofuv run bootcamp.py doctor.Watch the class during the day
terminaluv run instructor.py class --watch 30 uv run instructor.py spendclassprints one row per participant: company, stage (from their deployed AgentCore resources), LiteLLM spend and budget, tracked Code Interpreter / Browser sessions, open guardrail alerts and last activity.spendprints spend per company team, then per participant.Reset a broken participant stack
terminaluv run instructor.py reset <name> # dry run: lists their resources uv run instructor.py reset <name> --force # deletes themLike
purge, scoped to one participant. They then runuv run bootcamp.py forgetto drop their stale local terraform state and start again withup 0. Spend is not reset; raise their budget inroster.jsonand runkeysinstead.Tear down
Ask participants to run
uv run bootcamp.py down, then check for leftovers and remove the account stack:terminaluv run instructor.py purge # dry run: lists leftover participant resources uv run instructor.py purge --force # deletes them (only if the list wasn't empty) uv run instructor.py downinstructor.py downrefuses while any participant resources still exist and lists them;instructor.py down --forcepurges first, then destroys. Forgotten stacks are also caught by the scheduled clean-up (see Infra guardrails).
Participant AWS access (SSO)
Participants sign in with AWS IAM Identity Center: share the SSO start URL and region; they run aws configure sso and put the profile in AWS_PROFILE. The CLI then assumes bootcamp-participant-<name>.
Each role's trust defaults to the account root, so the participants' permission set must allow sts:AssumeRole on arn:aws:iam::<account>:role/bootcamp/bootcamp-participant-*. To pin a role to one person, set principal_arn for that participant in roster.json.
Infra guardrails
LLM spend is capped by LiteLLM: per-participant budgets inside per-company team budgets (HTTP 429 budget_exceeded). AWS spend is watched by an infra guard Lambda, CloudWatch alarms and an AWS Budget, all alerting to the SNS topic bootcamp-alerts.
Per-participant limits
The guard reacts within seconds to new Code Interpreter / Browser sessions (a dedicated CloudTrail trail, bootcamp-agentcore-sessions, logs only those two session-start data events; nothing else in the sandbox is logged) and sweeps live resources every 5 minutes. It attributes usage by caller role (bootcamp-participant-<p> / awsworkshop-<p>-*) and resource name. Alerts are de-duplicated per hour.
Limit (infra_limits.per_participant) | Default | When exceeded |
|---|---|---|
runtimes | 4 | alert |
gateways | 2 | alert |
memories | 2 | alert |
code_interpreters | 2 | alert |
browsers | 2 | alert |
code_interpreter_sessions | 3 | auto-stop newest (if enforce) |
browser_sessions | 3 | auto-stop newest (if enforce) |
runtime_invocations_5min | 300 | alert |
Override per team (teams.<t>.infra_limits) or per participant (teams.<t>.participants.<p>.infra_limits). Class-wide totals default to per-participant × headcount (override with infra_limits.class). Auto-stop needs infra_limits.enforce: true; runtimes, gateways and memories are alert-only.
Alarms
CloudWatch alarm → bootcamp-alerts | Threshold |
|---|---|
| Code Interpreter ActiveSessionCount (account) | 25 |
| Browser ActiveSessionCount (account) | 25 |
| Runtime ActiveSessionCount (account) | 200 |
| InvokeAgentRuntime (account, 5 min) | 1500 |
| LiteLLM gateway ECS CPU | > 85% |
| LiteLLM RDS CPU | > 80% |
| Guard Lambda errors | any |
| Clean-up Lambda errors | any |
The 5-minute sweep also reports when the LiteLLM gateway autoscaling is at its max task count.
Cost
AWS Budget bootcamp-genai-bootcamp-tagged on cost tag Project=genai-bootcamp (infra_limits.monthly_budget_usd, default $100), notifying at 80% actual and 100% forecast. AWS-managed tool sessions can't be tagged, so set infra_limits.account_budget_usd > 0 if you also want an account-wide budget.
Scheduled clean-up
The Lambda bootcamp-cleanup runs on cleanup.schedule (daily by default). It finds participant stacks with the same discovery as instructor.py purge (it ships the CLI's own purge.py) and treats a participant as idle when none of their roles (bootcamp-participant-<p>, awsworkshop-<p>-*) was used for cleanup.after_hours (default 48, minimum 6). Idle stacks are emailed to bootcamp-alerts; with cleanup.enforce: true they are purged first, in purge's dependency order, and only then does the Lambda get delete rights. Test it any time with a dry run that never deletes:
aws lambda invoke --function-name bootcamp-cleanup --cli-binary-format raw-in-base64-out \
--payload '{"dry_run": true}' out.json && cat out.jsonManual steps
- Set
alert_emailinroster.jsonand runuv run instructor.py up. - Confirm the SNS subscription email. Until then nobody is emailed.
- Activate
Projectas a cost allocation tag in the payer account. - Optional billing alarm: enable billing alerts, put
enable_billing_alarm = trueinterraform/account/terraform.tfvars, then runinstructor.py up.
What instructor.py up creates and common issues
- LiteLLM gateway on ECS Fargate (2 to 4 tasks, CPU and memory autoscaling;
gateway_min_tasks/gateway_max_tasks) behind an ALB and CloudFront with a 120 s origin timeout, the only path to models. Participants call it with an OpenAI-compatible API and a presigned STS identity header (X-Amz-Sts-Identity-Url); the keyless proxy in each task mapsbootcamp-participant-<name>andawsworkshop-<name>-*roles to that participant's key and forwards to LiteLLM and Bedrock. Participants may call only the model routes and/key/info(HTTP 403 otherwise), may name only guardrails tagged with their ownParticipant, and an empty Bedrock stream is retried once, then sent to a fallback model. The LiteLLM admin UI (/ui) asks for a Google Workspace sign-in (admin_sso) before the master key. - RDS database for LiteLLM keys, budgets and spend tracking. The gateway URL is stored per participant in the secret
bootcamp/participants/<name>/litellm, from which the CLI fills inLITELLM_BASE_URL. - Per-participant IAM roles with tag-based isolation (
Participant=<name>,Team=<company>, namesawsworkshop-<name>-*) and a permissions boundary on every role they create that blocks direct Bedrock access. Participants can't assume their workload roles, and every role they create must carry thebootcamp-workload-boundary(an allowlist) and theirParticipanttag; the platform adds both automatically. - Observability:
instructor.py upleaves the account's Transaction Search setting alone; runuv run instructor.py up --transaction-searchto switch it on so participants' traces show up in GenAI Observability (Phase 2, Task 5). - This website, served from S3 through CloudFront (custom domain and ACM certificate when
dns_zoneis set). - Terraform state in the S3 bucket
bootcamp-tfstate-<account>-<region>(versioned, encrypted, S3 lock file), created byinstructor.py upbeforeterraform init. Participant stacks keep local state on their laptops.
Common issues: the first up 0 waits about 5 minutes for Cognito DNS (domains look like ds-<10 hex>-<account>); if someone queried the domain too early, flush their DNS cache. instructor.py up warns loudly when alert_email is empty. Stage 5 checks may WARN for about 10 minutes until spans arrive. A participant whose own budget or company team budget is exhausted gets HTTP 429 budget_exceeded from LiteLLM: check instructor.py spend and raise their budget with instructor.py keys.