Task 5 · 8 tasks

Observability

See every model call, tool call and millisecond your agent spends, without adding code.

20 minEasy
Alice’s ask

The “Black Box” panic

“This agent serves 1,200 people and I have no idea what happens inside it!” Before something breaks, Alice wants traces: which tools ran, how long each step took, where tokens went.

Platform

What the platform provisions for you

terminal
uv run bootcamp.py up 5
  • Nothing new to create: tracing has been on since Task 4 (the agent runs under opentelemetry-instrument) and its log group already exists
  • Stage 5 switches on the check that your spans arrive in Transaction Search

Adding this stage usually takes ~0.5 min; the CLI prints progress (and full terraform output with --verbose).

Transaction Search is an account-wide switch; your instructor turned it on once for everyone.

You

What you do as a developer

  1. Generate some traffic

    terminal
    uv run bootcamp.py invoke "Which department has the most employees?" --actor alice-chen
    uv run bootcamp.py invoke "Headcount in Sales and the weather in Chicago?" --actor alice-chen

    Spans arrive 2-10 minutes after an invoke.

  2. Look at a trace

    Two ways, same data. From the terminal (by default your latest invoke on this machine, or pass the session id invoke printed):

    terminal
    uv run bootcamp.py traces
    uv run bootcamp.py traces SESSION_ID
    uv run bootcamp.py traces --args   # also each tool call's input, e.g. the SQL sent to query_db

    It prints a span tree with durations and token counts. --args adds each tool call's input, read from your runtime's log group (it lags a few minutes too). In the AWS console: CloudWatch → GenAI Observability → Bedrock AgentCore, pick your agent, then a session and its trace.

  3. Find where the time goes

    Challenge

    For the second question, find the slowest span and decide what dominates latency: model calls, the Gateway tool calls (database and weather Lambda), or the weather API behind the Lambda.

    Hint 1

    Expand the data_agent and weather_agent tool spans: each contains the specialist's own model calls. Compare those with the Gateway query_db spans inside.

    Hint 2

    Look at the duration of each individual Gateway tool call. Is it what you'd expect for a SQLite query on a small file? Think back to the Task 1 experiment about microVMs.

    Solution

    Model calls add up (the orchestrator's routing call plus each specialist's calls, run one after another), but the surprise is the tool calls: each Gateway query_db call takes about 6-7 seconds. The query itself takes milliseconds. AgentCore Runtime starts a new microVM per MCP session (a cold start of about 5.5 s: boot, download the database, start Python), and the Gateway opens a new MCP session for every call. Requests that reuse a session take about 0.5 s.

  4. Experiments

    Switch MODEL_ID in .env, deploy, send the same prompts and compare latency and token counts in the traces.

    What should I expect?

    claude-sonnet-5-5 tends to route with fewer, better tool calls but each call is slower and pricier; gpt-6-luna is faster per call but may need extra round-trips. Fewer tool calls also means fewer cold starts. Measure, don't guess.

Check your work

terminal
uv run bootcamp.py test --only 5

Spans can take about 10 minutes to appear. A WARN here just means “not yet”; re-run later.

Under the hood

AgentCore Runtime ships the AWS Distro for OpenTelemetry. Strands emits OpenTelemetry spans for agent cycles, model calls and tool calls, and the runtime exports them to X-Ray. With Transaction Search enabled, spans land in the account-wide aws/spans log group, where the GenAI Observability dashboards, bootcamp.py traces and the evaluations in Task 6 read them. Spans carry metadata (names, timings, token counts), not your prompts or answers. Application logs go to /aws/bedrock-agentcore/runtimes/<runtime-id>-DEFAULT; your role can only read your own awsworkshop_<you>_* log groups, so another participant's logs return AccessDenied by design.