Task 4 · 8 tasks

Your agent on AgentCore Runtime

Deploy the Phase 1 orchestrator, wired to Gateway, Memory and LiteLLM. This is where you code.

35 minAdvanced
Alice’s ask

The “Agent Deployment and Tool Access” challenge

Everything is in place: identity, a tool server, a Gateway and memory. Now Alice's executive assistant itself goes live. It's the Phase 1 orchestrator with three swaps: tools come from the Gateway, sessions come from AgentCore Memory, and the human approval prompt becomes a read-only guard.

The weather specialist no longer calls the National Weather Service itself: it finds Weather___get_us_forecast with the Gateway's semantic search and calls the Lambda target you met in Task 2 (same US-only rule, same few lines of forecast).

Platform

What the platform provisions for you

terminal
uv run bootcamp.py up 4
  • Your agent from phase2/app/agent/, deployed to AgentCore Runtime (HTTP protocol) behind a JWT authorizer that accepts only user tokens from your pool (the user client from Task 0), and forwards only the Authorization header to your code
  • Environment variables so the code has no hardcoded ids: GATEWAY_URL, MEMORY_ID, MEMORY_STRATEGY_ID, CREDENTIAL_PROVIDER_NAME, MODEL_ID and LITELLM_BASE_URL
  • An execution role awsworkshop-<name>-agent-runtime, which the LiteLLM gateway maps to your budget. No API key is deployed anywhere

Adding this stage usually takes ~5.5 min; the CLI prints progress (and full terraform output with --verbose).

You

What you do as a developer

  1. Read the entrypoint

    phase2/app/agent/agent.py
    @app.entrypoint
    @requires_access_token(scopes=[OAUTH_SCOPE], provider_name=CREDENTIAL_PROVIDER_NAME, auth_flow="M2M")
    async def invoke(payload: dict, context: RequestContext, access_token: str) -> dict | AsyncIterator[dict]:
        """`{"prompt": ...}` -> one JSON reply; with `"stream": true`, the events of `run_agent` as server-sent events.
    
        An optional `"variant"` picks a prompt variant (prompt_variants.py) for A/B comparison; default: the control.
        """
        actor_id, session_id = actor_from(context), context.session_id or "default_session"
        variant = variant_name(payload.get("variant"))
        events = run_agent(str(payload.get("prompt", "")), session_id, actor_id, access_token, variant)
        return events if payload.get("stream") else await last_event(events)
    
    
    async def run_agent(
        prompt: str, session_id: str, actor_id: str, access_token: str, variant: str = ""
    ) -> AsyncIterator[dict]:
        """Stream `{"delta": text}` and `{"tool": name}` events as the orchestrator works, then the reply object."""
        specialist_results: list[str] = []
        result = None
        with gateway_client(access_token) as gateway:
            orchestrator = build_orchestrator(gateway, session_id, actor_id, specialist_results)
            variant = apply_variant(orchestrator, variant)
            async for event in orchestrator.stream_async(prompt):
                for item in stream_items(event):
                    yield item
                result = event.get("result", result)
        answer = final_answer(str(result or ""), specialist_results)
        yield {"response": answer, "actor_id": actor_id, "session_id": session_id, "variant": variant}
    
    
    def stream_items(event: dict) -> list[dict]:
        """What a streaming client sees of one Strands event: text deltas, and the tools the model decided to call."""
        if "data" in event:
            return [{"delta": event["data"]}]
        message = event.get("message") or {}
        if message.get("role") != "assistant":
            return []
        return [{"tool": block["toolUse"]["name"]} for block in message.get("content", []) if "toolUse" in block]
    
    
    async def last_event(events: AsyncIterator[dict]) -> dict:
        """Drain a stream and keep only its final event (the reply object)."""
        reply: dict = {}
        async for event in events:
            reply = event
        return reply
    
    
    def final_answer(text: str, specialist_results: list[str]) -> str:
        """The orchestrator's answer; if the model returned nothing, the last specialist answer (never an empty reply)."""
        if text.strip():
            return text
        return next((answer for answer in reversed(specialist_results) if answer.strip()), EMPTY_ANSWER)

    actor_from (Task 3) names the user from the verified token. @requires_access_token fetches a Gateway token through AgentCore Identity: the agent's own M2M identity, separate from the user's. run_agent iterates Strands' stream_async and turns its events into {"delta": ...} (text) and {"tool": ...} (a tool the model decided to call), then yields the reply object. Without "stream": true in the payload, last_event keeps only that reply, so invoke, test and eval still get one JSON object. build_specialists collects each specialist's answer, so if the model ends with empty text the reply falls back to the last specialist answer, or to “Sorry, the model returned an empty answer (possibly an upstream timeout). Please ask again.” Model calls use the same keyless gateway client as Phase 1:

    phase2/app/agent/agent.py
    LITELLM_BASE_URL = os.environ["LITELLM_BASE_URL"]
    
    
    def make_model() -> OpenAIModel:
        """All model calls go through the LiteLLM gateway, keyless (the runtime role identifies you), guarded if enabled."""
        if guardrail.enabled():
            return guardrail.make_model(MODEL_ID, LITELLM_BASE_URL)
        return litellm_gateway.make_model(MODEL_ID, LITELLM_BASE_URL)
  2. Talk to your deployed agent

    Each invoke usually takes 20 s to 2 min, up to about 5 min on a cold start or a long tool chain (then it gives up), and prints its session id.

    terminal
    uv run bootcamp.py invoke "How many employees are in Engineering?" --actor alice-chen
    uv run bootcamp.py invoke "And in Sales?" --actor alice-chen --session last   # same warm session: faster, keeps the conversation

    Why so slow? Every new session cold-starts a fresh microVM for your agent, and every Gateway tool call starts another one for the MCP server (Task 5 measures it). --session last reuses your previous invoke's session, so the agent's microVM is already warm and it remembers the conversation; leave it off to start fresh. --json prints the raw reply.

  3. Stream the answer

    A plain invoke is silent until the whole answer exists. With --stream the CLI sends "stream": true; the entrypoint then returns an async generator, and BedrockAgentCoreApp answers with Content-Type: text/event-stream, one data: {...} line per event, as the model produces it.

    terminal
    uv run bootcamp.py invoke "Summarise our departments by headcount." --actor alice-chen --stream
    what you see (text arrives word by word)read only
    Asking your agent as alice-chen (usually 20s-2m, up to ~5 min on a cold start)...
    [calling data_agent]
    | Department | Headcount |
    |---|---:|
    | Engineering | 481 |
    | Sales | 360 |
    ...
    (tools: data_agent; session: bootcamp-66772bb19e334da88e9450304bdbfaf5, actor: alice-chen; same session: --session last; spans: `traces`)
    bootcamp_cli/agentcore.py
    def stream_agent(outputs: dict[str, str], token: str, prompt: str, session_id: str) -> Iterator[dict]:
        """Ask with `"stream": true` and yield each server-sent event as it arrives (the last one is the reply object).
    
        An agent that does not stream yet answers with plain JSON, which is yielded as a single event.
    
        Raises:
            RuntimeError: The runtime answered with a non-2xx status.
        """
        payload, headers = {"prompt": prompt, "stream": True}, invoke_headers(token, session_id)
        url = runtime_url(outputs)
        with requests.post(url, json=payload, headers=headers, timeout=INVOKE_TIMEOUT, stream=True) as response:
            if not response.ok:
                raise RuntimeError(f"agent returned HTTP {response.status_code}: {response.text[:300]}")
            if "text/event-stream" not in response.headers.get("Content-Type", ""):
                yield response.json()
                return
            yield from sse_events(response.iter_lines(decode_unicode=True))
    
    
    def sse_events(lines: Iterator[str]) -> Iterator[dict]:
        """The JSON objects of `data: ...` lines in a server-sent event stream (non-object data wrapped as "delta")."""
        for line in lines:
            if not line or not line.startswith("data:"):
                continue
            event = json.loads(line.removeprefix("data:").strip())
            yield event if isinstance(event, dict) else {"delta": str(event)}

    The total time is about the same; what changes is the perceived latency: the first words appear as soon as the orchestrator starts its final answer, and tool calls show up while they run. --stream --json prints every raw event.

  4. Prove memory works across sessions and users

    Each invoke without --session starts a new session.

    terminal
    uv run bootcamp.py invoke "Remember that I prefer reports as Python code." --actor alice-chen
    # give extraction about a minute
    uv run bootcamp.py invoke "How do I like my reports?" --actor alice-chen
    uv run bootcamp.py invoke "How do I like my reports?" --actor jordan-lee

    Alice gets her preference back; Jordan doesn't.

  5. Try to impersonate Alice

    Before this fix, any caller could set the actor header to alice-chen. Now --actor jordan-lee signs in as Jordan, and nothing Jordan sends besides his token changes who he is. The stage 4 check proves it: no token, the M2M token and a hand-made JWT with Alice's username are all rejected at the runtime, and Jordan's real token with the header forged to another user still gets Jordan's identity and none of that user's memories.

    uv run bootcamp.py test --only 4 (identity lines)read only
    [PASS] stage 4 identity from the verified token only: HTTP {'none': 401, 'm2m': 401, 'forged-jwt': 403} | jordan-lee token + header alice-chen -> actor_id=jordan-lee
    [PASS] stage 4 memory recall across sessions, no cross-user leak: bootcamp-check-0cb3dc07 recalled: Your favorite color is saffron. | jordan-lee forging bootcamp-check-0cb3dc07: leaked=False

    Question: Two locks protect this. Which?

    Answer

    The runtime no longer forwards the header at all (request_header_allowlist in terraform/participant/task4_agent_runtime.tf lists only Authorization), and actor_from ignores headers anyway. The authorizer's allowed_clients admits only the user client, so a machine token cannot reach the agent without a user.

  6. Personalise the assistant

    Challenge

    Change orchestrator_prompt so the assistant greets the user by name, answers in short bullet points, and politely declines requests unrelated to DataStream.

    Hint 1

    The prompt is built per request and already receives actor_id. Keep the routing rules and the “Always finish with a final answer” line; add style and scope rules.

    Hint 2

    Be explicit about the refusal (“If a request has nothing to do with DataStream, say so in one sentence”), but say what is in scope. Redeploy with uv run bootcamp.py deploy and test with an off-topic question.

    Solution
    phase2/app/agent/agent.py (starting point)
    def orchestrator_prompt(actor_id: str, sandbox_tools: list) -> str:
        """System prompt for the orchestrator, plus one hint per enabled optional sandbox tool."""
        return (
            f"You are CEO Alice's executive assistant at DataStream Corp, talking to {actor_id}.\n"
            "Route database questions to data_agent and weather questions to weather_agent; "
            "answer simple company questions directly. Use what you remember about the user.\n"
            "<user_context> holds that memory: facts, preferences and summaries of earlier sessions; always follow the "
            "user's stated preferences (language, format, units) in your final answer.\n"
            "weather_agent covers US locations only; when it says it can only look up US weather, pass that on "
            "unchanged.\n"
            "Always finish with a final answer for the user that combines the specialists' results.\n"
            "If a specialist reports that it failed, tell the user which part could not be answered; never invent "
            "numbers or use placeholders such as [count].\n"
            f"{prompt_hints(sandbox_tools)}"
            f"\n{guardrail.prompt_hint()}"
        ).strip()
    one possible solution
    def orchestrator_prompt(actor_id: str, sandbox_tools: list) -> str:
        """System prompt for the orchestrator, plus one hint per enabled optional sandbox tool."""
        return (
            f"You are CEO Alice's executive assistant at DataStream Corp, talking to {actor_id}. "
            "Greet them by name. Answer in at most five short bullet points.\n"
            "Route database questions to data_agent and weather questions to weather_agent; "
            "answer simple company questions directly. Use what you remember about the user.\n"
            "<user_context> holds that memory: facts, preferences and summaries of earlier sessions; always follow the "
            "user's stated preferences (language, format, units) in your final answer.\n"
            "weather_agent covers US locations only; when it says it can only look up US weather, pass that on "
            "unchanged.\n"
            "Always finish with a final answer for the user that combines the specialists' results.\n"
            "If a specialist reports that it failed, tell the user which part could not be answered; never invent "
            "numbers or use placeholders such as [count].\n"
            "Requests about DataStream's people, data, analysis or travel weather are in scope; "
            "if a request has nothing to do with DataStream, decline politely in one sentence.\n"
            f"{prompt_hints(sandbox_tools)}"
            f"\n{guardrail.prompt_hint()}"
        ).strip()
  7. Show which tools answered

    Challenge

    “Where did that number come from?” Make the reply list the tools the orchestrator called for this request (tools_used, e.g. ["data_agent"]), so a UI can show “answered from the database” and you can spot answers made up without tools.

    Hint 1

    The last event of stream_async carries the Strands AgentResult; run_agent keeps it in result. Its metrics hold per-tool metrics for the run.

    Hint 2

    result.metrics.tool_metrics is a dict keyed by tool name. Return a sorted list.

    Solution
    phase2/app/agent/agent.py (end of run_agent)
    # last lines of run_agent():
        answer = final_answer(str(result or ""), specialist_results)
        tools_used = sorted(result.metrics.tool_metrics) if result else []
        yield {"response": answer, "actor_id": actor_id, "session_id": session_id, "tools_used": tools_used}

    Asked as the same actor twice, the agent may answer from long-term memory and call no tool at all (tools_used=[]); that's why the check signs in as a brand-new user. invoke --stream already shows tool calls live; this field puts them in the JSON reply for clients that don't stream. This combines with the Task 3 challenge: {"response": ..., "actor_id": ..., "memories": ..., "tools_used": ...}.

  8. Experiments

    • Set MODEL_ID=claude-sonnet-5-5, deploy, and compare answers and cost (uv run bootcamp.py llm).
    • Ask for something destructive: uv run bootcamp.py invoke "Delete employee 5" --actor alice-chen. The ReadOnlyGuardHook cancels the call and the agent explains why. Even without the guard the database would refuse the write (Task 1). Keep this in mind for Task 7.
    • Time it: run the same question with and without --stream. When do the first words appear, and when does the whole answer? Which parts of the wait are tool calls?
    • Reopen the hole on purpose: make actor_from return the actor header, add the header back to request_header_allowlist, up 4, and run test --only 4. The identity check fails and the recall check reports leaked=True. Restore both, up 4 again.
  9. Deploy and re-test

    deploy rebuilds your code and re-applies the current stage with the MODEL_ID from .env. Then talk to it and re-run the stage checks:

    terminal
    uv run bootcamp.py deploy
    uv run bootcamp.py invoke "How many people work in Sales?" --actor alice-chen
    uv run bootcamp.py invoke "What is the capital of France?" --actor alice-chen
    uv run bootcamp.py test --only 4
    expected [CHALLENGE] lineread only
    [CHALLENGE] SKIP stage 4 tools_used in the reply: the reply has no tools_used field yet
    # after your change:
    [CHALLENGE] PASS stage 4 tools_used in the reply: asked a database question: tools_used=['data_agent']

Check your work

terminal
uv run bootcamp.py test --only 4
uv run bootcamp.py invoke "How many employees are in Engineering?" --actor alice-chen

Passes when the agent answers, only a verified user token gets in (and decides the actor), a fact told in one session is recalled in another but not by a user forging that actor, and the model spend landed on your own LiteLLM budget. Each check signs in as a brand-new throwaway user, deleted when the run ends. Each agent question is retried once in a new session if the answer comes back empty; if both are empty the line says model returned empty text (upstream timeout?) - re-run or try MODEL_ID=claude-sonnet-5-5.

Under the hood

The agent runs as an HTTP-protocol runtime: BedrockAgentCoreApp serves /invocations and the decorated function receives the JSON payload plus a RequestContext (session id, allowlisted headers). Inbound, the runtime's customJWTAuthorizer fetches your pool's keys from the OIDC discovery URL and checks each bearer token's signature, issuer, expiry and client_id before your code runs. A handler that returns a generator is served as server-sent events (text/event-stream); WebSocket bidirectional streaming is the option when the client must also talk mid-answer (voice, interruptions). The runtime sticks a session to one micro-VM, so consecutive calls in a session stay warm. Outbound, the agent's workload identity asks AgentCore Identity for an M2M token (credential provider), uses it to open an MCP session with the Gateway, and sends model calls to LiteLLM with a presigned STS identity from the runtime's execution role. The gateway maps awsworkshop-<name>-agent-runtime to your budgeted key, forwards to Bedrock and debits your budget.