Task 4 · 8 tasks
Your agent on AgentCore Runtime
Deploy the Phase 1 orchestrator, wired to Gateway, Memory and LiteLLM. This is where you code.
The “Agent Deployment and Tool Access” challenge
Everything is in place: identity, a tool server, a Gateway and memory. Now Alice's executive assistant itself goes live. It's the Phase 1 orchestrator with three swaps: tools come from the Gateway, sessions come from AgentCore Memory, and the human approval prompt becomes a read-only guard.
The weather specialist no longer calls the National Weather Service itself: it finds Weather___get_us_forecast with the Gateway's semantic search and calls the Lambda target you met in Task 2 (same US-only rule, same few lines of forecast).
What the platform provisions for you
uv run bootcamp.py up 4- Your agent from
phase2/app/agent/, deployed to AgentCore Runtime (HTTP protocol) behind a JWT authorizer that accepts only user tokens from your pool (the user client from Task 0), and forwards only theAuthorizationheader to your code - Environment variables so the code has no hardcoded ids:
GATEWAY_URL,MEMORY_ID,MEMORY_STRATEGY_ID,CREDENTIAL_PROVIDER_NAME,MODEL_IDandLITELLM_BASE_URL - An execution role
awsworkshop-<name>-agent-runtime, which the LiteLLM gateway maps to your budget. No API key is deployed anywhere
Adding this stage usually takes ~5.5 min; the CLI prints progress (and full terraform output with --verbose).
What you do as a developer
Read the entrypoint
phase2/app/agent/agent.py@app.entrypoint @requires_access_token(scopes=[OAUTH_SCOPE], provider_name=CREDENTIAL_PROVIDER_NAME, auth_flow="M2M") async def invoke(payload: dict, context: RequestContext, access_token: str) -> dict | AsyncIterator[dict]: """`{"prompt": ...}` -> one JSON reply; with `"stream": true`, the events of `run_agent` as server-sent events. An optional `"variant"` picks a prompt variant (prompt_variants.py) for A/B comparison; default: the control. """ actor_id, session_id = actor_from(context), context.session_id or "default_session" variant = variant_name(payload.get("variant")) events = run_agent(str(payload.get("prompt", "")), session_id, actor_id, access_token, variant) return events if payload.get("stream") else await last_event(events) async def run_agent( prompt: str, session_id: str, actor_id: str, access_token: str, variant: str = "" ) -> AsyncIterator[dict]: """Stream `{"delta": text}` and `{"tool": name}` events as the orchestrator works, then the reply object.""" specialist_results: list[str] = [] result = None with gateway_client(access_token) as gateway: orchestrator = build_orchestrator(gateway, session_id, actor_id, specialist_results) variant = apply_variant(orchestrator, variant) async for event in orchestrator.stream_async(prompt): for item in stream_items(event): yield item result = event.get("result", result) answer = final_answer(str(result or ""), specialist_results) yield {"response": answer, "actor_id": actor_id, "session_id": session_id, "variant": variant} def stream_items(event: dict) -> list[dict]: """What a streaming client sees of one Strands event: text deltas, and the tools the model decided to call.""" if "data" in event: return [{"delta": event["data"]}] message = event.get("message") or {} if message.get("role") != "assistant": return [] return [{"tool": block["toolUse"]["name"]} for block in message.get("content", []) if "toolUse" in block] async def last_event(events: AsyncIterator[dict]) -> dict: """Drain a stream and keep only its final event (the reply object).""" reply: dict = {} async for event in events: reply = event return reply def final_answer(text: str, specialist_results: list[str]) -> str: """The orchestrator's answer; if the model returned nothing, the last specialist answer (never an empty reply).""" if text.strip(): return text return next((answer for answer in reversed(specialist_results) if answer.strip()), EMPTY_ANSWER)actor_from(Task 3) names the user from the verified token.@requires_access_tokenfetches a Gateway token through AgentCore Identity: the agent's own M2M identity, separate from the user's.run_agentiterates Strands'stream_asyncand turns its events into{"delta": ...}(text) and{"tool": ...}(a tool the model decided to call), then yields the reply object. Without"stream": truein the payload,last_eventkeeps only that reply, soinvoke,testandevalstill get one JSON object.build_specialistscollects each specialist's answer, so if the model ends with empty text the reply falls back to the last specialist answer, or to “Sorry, the model returned an empty answer (possibly an upstream timeout). Please ask again.” Model calls use the same keyless gateway client as Phase 1:phase2/app/agent/agent.pyLITELLM_BASE_URL = os.environ["LITELLM_BASE_URL"] def make_model() -> OpenAIModel: """All model calls go through the LiteLLM gateway, keyless (the runtime role identifies you), guarded if enabled.""" if guardrail.enabled(): return guardrail.make_model(MODEL_ID, LITELLM_BASE_URL) return litellm_gateway.make_model(MODEL_ID, LITELLM_BASE_URL)Talk to your deployed agent
Each
invokeusually takes 20 s to 2 min, up to about 5 min on a cold start or a long tool chain (then it gives up), and prints its session id.terminaluv run bootcamp.py invoke "How many employees are in Engineering?" --actor alice-chen uv run bootcamp.py invoke "And in Sales?" --actor alice-chen --session last # same warm session: faster, keeps the conversationWhy so slow? Every new session cold-starts a fresh microVM for your agent, and every Gateway tool call starts another one for the MCP server (Task 5 measures it).
--session lastreuses your previous invoke's session, so the agent's microVM is already warm and it remembers the conversation; leave it off to start fresh.--jsonprints the raw reply.Stream the answer
A plain
invokeis silent until the whole answer exists. With--streamthe CLI sends"stream": true; the entrypoint then returns an async generator, andBedrockAgentCoreAppanswers withContent-Type: text/event-stream, onedata: {...}line per event, as the model produces it.terminaluv run bootcamp.py invoke "Summarise our departments by headcount." --actor alice-chen --streamwhat you see (text arrives word by word)read onlyAsking your agent as alice-chen (usually 20s-2m, up to ~5 min on a cold start)... [calling data_agent] | Department | Headcount | |---|---:| | Engineering | 481 | | Sales | 360 | ... (tools: data_agent; session: bootcamp-66772bb19e334da88e9450304bdbfaf5, actor: alice-chen; same session: --session last; spans: `traces`)bootcamp_cli/agentcore.pydef stream_agent(outputs: dict[str, str], token: str, prompt: str, session_id: str) -> Iterator[dict]: """Ask with `"stream": true` and yield each server-sent event as it arrives (the last one is the reply object). An agent that does not stream yet answers with plain JSON, which is yielded as a single event. Raises: RuntimeError: The runtime answered with a non-2xx status. """ payload, headers = {"prompt": prompt, "stream": True}, invoke_headers(token, session_id) url = runtime_url(outputs) with requests.post(url, json=payload, headers=headers, timeout=INVOKE_TIMEOUT, stream=True) as response: if not response.ok: raise RuntimeError(f"agent returned HTTP {response.status_code}: {response.text[:300]}") if "text/event-stream" not in response.headers.get("Content-Type", ""): yield response.json() return yield from sse_events(response.iter_lines(decode_unicode=True)) def sse_events(lines: Iterator[str]) -> Iterator[dict]: """The JSON objects of `data: ...` lines in a server-sent event stream (non-object data wrapped as "delta").""" for line in lines: if not line or not line.startswith("data:"): continue event = json.loads(line.removeprefix("data:").strip()) yield event if isinstance(event, dict) else {"delta": str(event)}The total time is about the same; what changes is the perceived latency: the first words appear as soon as the orchestrator starts its final answer, and tool calls show up while they run.
--stream --jsonprints every raw event.Prove memory works across sessions and users
Each
invokewithout--sessionstarts a new session.terminaluv run bootcamp.py invoke "Remember that I prefer reports as Python code." --actor alice-chen # give extraction about a minute uv run bootcamp.py invoke "How do I like my reports?" --actor alice-chen uv run bootcamp.py invoke "How do I like my reports?" --actor jordan-leeAlice gets her preference back; Jordan doesn't.
Try to impersonate Alice
Before this fix, any caller could set the actor header to
alice-chen. Now--actor jordan-leesigns in as Jordan, and nothing Jordan sends besides his token changes who he is. The stage 4 check proves it: no token, the M2M token and a hand-made JWT with Alice'susernameare all rejected at the runtime, and Jordan's real token with the header forged to another user still gets Jordan's identity and none of that user's memories.uv run bootcamp.py test --only 4 (identity lines)read only[PASS] stage 4 identity from the verified token only: HTTP {'none': 401, 'm2m': 401, 'forged-jwt': 403} | jordan-lee token + header alice-chen -> actor_id=jordan-lee [PASS] stage 4 memory recall across sessions, no cross-user leak: bootcamp-check-0cb3dc07 recalled: Your favorite color is saffron. | jordan-lee forging bootcamp-check-0cb3dc07: leaked=FalseQuestion: Two locks protect this. Which?
Answer
The runtime no longer forwards the header at all (
request_header_allowlistinterraform/participant/task4_agent_runtime.tflists onlyAuthorization), andactor_fromignores headers anyway. The authorizer'sallowed_clientsadmits only the user client, so a machine token cannot reach the agent without a user.Personalise the assistant
Challenge
Change
orchestrator_promptso the assistant greets the user by name, answers in short bullet points, and politely declines requests unrelated to DataStream.Hint 1
The prompt is built per request and already receives
actor_id. Keep the routing rules and the “Always finish with a final answer” line; add style and scope rules.Hint 2
Be explicit about the refusal (“If a request has nothing to do with DataStream, say so in one sentence”), but say what is in scope. Redeploy with
uv run bootcamp.py deployand test with an off-topic question.Solution
phase2/app/agent/agent.py (starting point)def orchestrator_prompt(actor_id: str, sandbox_tools: list) -> str: """System prompt for the orchestrator, plus one hint per enabled optional sandbox tool.""" return ( f"You are CEO Alice's executive assistant at DataStream Corp, talking to {actor_id}.\n" "Route database questions to data_agent and weather questions to weather_agent; " "answer simple company questions directly. Use what you remember about the user.\n" "<user_context> holds that memory: facts, preferences and summaries of earlier sessions; always follow the " "user's stated preferences (language, format, units) in your final answer.\n" "weather_agent covers US locations only; when it says it can only look up US weather, pass that on " "unchanged.\n" "Always finish with a final answer for the user that combines the specialists' results.\n" "If a specialist reports that it failed, tell the user which part could not be answered; never invent " "numbers or use placeholders such as [count].\n" f"{prompt_hints(sandbox_tools)}" f"\n{guardrail.prompt_hint()}" ).strip()one possible solutiondef orchestrator_prompt(actor_id: str, sandbox_tools: list) -> str: """System prompt for the orchestrator, plus one hint per enabled optional sandbox tool.""" return ( f"You are CEO Alice's executive assistant at DataStream Corp, talking to {actor_id}. " "Greet them by name. Answer in at most five short bullet points.\n" "Route database questions to data_agent and weather questions to weather_agent; " "answer simple company questions directly. Use what you remember about the user.\n" "<user_context> holds that memory: facts, preferences and summaries of earlier sessions; always follow the " "user's stated preferences (language, format, units) in your final answer.\n" "weather_agent covers US locations only; when it says it can only look up US weather, pass that on " "unchanged.\n" "Always finish with a final answer for the user that combines the specialists' results.\n" "If a specialist reports that it failed, tell the user which part could not be answered; never invent " "numbers or use placeholders such as [count].\n" "Requests about DataStream's people, data, analysis or travel weather are in scope; " "if a request has nothing to do with DataStream, decline politely in one sentence.\n" f"{prompt_hints(sandbox_tools)}" f"\n{guardrail.prompt_hint()}" ).strip()Show which tools answered
Challenge
“Where did that number come from?” Make the reply list the tools the orchestrator called for this request (
tools_used, e.g.["data_agent"]), so a UI can show “answered from the database” and you can spot answers made up without tools.Hint 1
The last event of
stream_asynccarries the StrandsAgentResult;run_agentkeeps it inresult. Itsmetricshold per-tool metrics for the run.Hint 2
result.metrics.tool_metricsis a dict keyed by tool name. Return a sorted list.Solution
phase2/app/agent/agent.py (end of run_agent)# last lines of run_agent(): answer = final_answer(str(result or ""), specialist_results) tools_used = sorted(result.metrics.tool_metrics) if result else [] yield {"response": answer, "actor_id": actor_id, "session_id": session_id, "tools_used": tools_used}Asked as the same actor twice, the agent may answer from long-term memory and call no tool at all (
tools_used=[]); that's why the check signs in as a brand-new user.invoke --streamalready shows tool calls live; this field puts them in the JSON reply for clients that don't stream. This combines with the Task 3 challenge:{"response": ..., "actor_id": ..., "memories": ..., "tools_used": ...}.Experiments
- Set
MODEL_ID=claude-sonnet-5-5,deploy, and compare answers and cost (uv run bootcamp.py llm). - Ask for something destructive:
uv run bootcamp.py invoke "Delete employee 5" --actor alice-chen. TheReadOnlyGuardHookcancels the call and the agent explains why. Even without the guard the database would refuse the write (Task 1). Keep this in mind for Task 7. - Time it: run the same question with and without
--stream. When do the first words appear, and when does the whole answer? Which parts of the wait are tool calls? - Reopen the hole on purpose: make
actor_fromreturn the actor header, add the header back torequest_header_allowlist,up 4, and runtest --only 4. The identity check fails and the recall check reportsleaked=True. Restore both,up 4again.
- Set
Deploy and re-test
deployrebuilds your code and re-applies the current stage with theMODEL_IDfrom.env. Then talk to it and re-run the stage checks:terminaluv run bootcamp.py deploy uv run bootcamp.py invoke "How many people work in Sales?" --actor alice-chen uv run bootcamp.py invoke "What is the capital of France?" --actor alice-chen uv run bootcamp.py test --only 4expected [CHALLENGE] lineread only[CHALLENGE] SKIP stage 4 tools_used in the reply: the reply has no tools_used field yet # after your change: [CHALLENGE] PASS stage 4 tools_used in the reply: asked a database question: tools_used=['data_agent']
Check your work
uv run bootcamp.py test --only 4
uv run bootcamp.py invoke "How many employees are in Engineering?" --actor alice-chenPasses when the agent answers, only a verified user token gets in (and decides the actor), a fact told in one session is recalled in another but not by a user forging that actor, and the model spend landed on your own LiteLLM budget. Each check signs in as a brand-new throwaway user, deleted when the run ends. Each agent question is retried once in a new session if the answer comes back empty; if both are empty the line says model returned empty text (upstream timeout?) - re-run or try MODEL_ID=claude-sonnet-5-5.
Under the hood
The agent runs as an HTTP-protocol runtime: BedrockAgentCoreApp serves /invocations and the decorated function receives the JSON payload plus a RequestContext (session id, allowlisted headers). Inbound, the runtime's customJWTAuthorizer fetches your pool's keys from the OIDC discovery URL and checks each bearer token's signature, issuer, expiry and client_id before your code runs. A handler that returns a generator is served as server-sent events (text/event-stream); WebSocket bidirectional streaming is the option when the client must also talk mid-answer (voice, interruptions). The runtime sticks a session to one micro-VM, so consecutive calls in a session stay warm. Outbound, the agent's workload identity asks AgentCore Identity for an M2M token (credential provider), uses it to open an MCP session with the Gateway, and sends model calls to LiteLLM with a presigned STS identity from the runtime's execution role. The gateway maps awsworkshop-<name>-agent-runtime to your budgeted key, forwards to Bedrock and debits your budget.