Task 5 · 8 tasks
Observability
See every model call, tool call and millisecond your agent spends, without adding code.
The “Black Box” panic
“This agent serves 1,200 people and I have no idea what happens inside it!” Before something breaks, Alice wants traces: which tools ran, how long each step took, where tokens went.
What the platform provisions for you
uv run bootcamp.py up 5- Nothing new to create: tracing has been on since Task 4 (the agent runs under
opentelemetry-instrument) and its log group already exists - Stage 5 switches on the check that your spans arrive in Transaction Search
Adding this stage usually takes ~0.5 min; the CLI prints progress (and full terraform output with --verbose).
Transaction Search is an account-wide switch; your instructor turned it on once for everyone.
What you do as a developer
Generate some traffic
terminaluv run bootcamp.py invoke "Which department has the most employees?" --actor alice-chen uv run bootcamp.py invoke "Headcount in Sales and the weather in Chicago?" --actor alice-chenSpans arrive 2-10 minutes after an invoke.
Look at a trace
Two ways, same data. From the terminal (by default your latest
invokeon this machine, or pass the session idinvokeprinted):terminaluv run bootcamp.py traces uv run bootcamp.py traces SESSION_ID uv run bootcamp.py traces --args # also each tool call's input, e.g. the SQL sent to query_dbIt prints a span tree with durations and token counts.
--argsadds each tool call's input, read from your runtime's log group (it lags a few minutes too). In the AWS console: CloudWatch → GenAI Observability → Bedrock AgentCore, pick your agent, then a session and its trace.Find where the time goes
Challenge
For the second question, find the slowest span and decide what dominates latency: model calls, the Gateway tool calls (database and weather Lambda), or the weather API behind the Lambda.
Hint 1
Expand the
data_agentandweather_agenttool spans: each contains the specialist's own model calls. Compare those with the Gatewayquery_dbspans inside.Hint 2
Look at the duration of each individual Gateway tool call. Is it what you'd expect for a SQLite query on a small file? Think back to the Task 1 experiment about microVMs.
Solution
Model calls add up (the orchestrator's routing call plus each specialist's calls, run one after another), but the surprise is the tool calls: each Gateway
query_dbcall takes about 6-7 seconds. The query itself takes milliseconds. AgentCore Runtime starts a new microVM per MCP session (a cold start of about 5.5 s: boot, download the database, start Python), and the Gateway opens a new MCP session for every call. Requests that reuse a session take about 0.5 s.Experiments
Switch
MODEL_IDin.env,deploy, send the same prompts and compare latency and token counts in the traces.What should I expect?
claude-sonnet-5-5tends to route with fewer, better tool calls but each call is slower and pricier;gpt-6-lunais faster per call but may need extra round-trips. Fewer tool calls also means fewer cold starts. Measure, don't guess.
Check your work
uv run bootcamp.py test --only 5Spans can take about 10 minutes to appear. A WARN here just means “not yet”; re-run later.
Under the hood
AgentCore Runtime ships the AWS Distro for OpenTelemetry. Strands emits OpenTelemetry spans for agent cycles, model calls and tool calls, and the runtime exports them to X-Ray. With Transaction Search enabled, spans land in the account-wide aws/spans log group, where the GenAI Observability dashboards, bootcamp.py traces and the evaluations in Task 6 read them. Spans carry metadata (names, timings, token counts), not your prompts or answers. Application logs go to /aws/bedrock-agentcore/runtimes/<runtime-id>-DEFAULT; your role can only read your own awsworkshop_<you>_* log groups, so another participant's logs return AccessDenied by design.