Task 2 · 8 tasks
AgentCore Gateway
One governed front door for every tool: two target types, and semantic search to find the right tool.
The “Tool Discovery” challenge
Today there's one tool server. Tomorrow there'll be ten, each with its own auth. Alice's agent should talk to one endpoint that knows every tool and handles credentials for it.
Part A puts your database tools behind the Gateway. Part B moves the weather tool there too, as a different kind of target (a Lambda function), and shows how an agent finds the right tool among many without loading them all.
What the platform provisions for you
uv run bootcamp.py up 2- An AgentCore Gateway (MCP) with inbound JWT auth and semantic tool search enabled
- An OAuth credential provider, so the Gateway can get its own token to call your MCP runtime
- A target named
DataStreamDatabasepointing at your MCP runtime (an MCP server target) - A Lambda function
awsworkshop-<name>-weather(Python 3.13, arm64) built fromphase2/app/weather_lambda/andshared/nws_weather.py, and a second target namedWeatherin front of it (a Lambda target; the Gateway's IAM role invokes it)
Adding this stage usually takes ~1 min; the CLI prints progress (and full terraform output with --verbose).
From now on, every deploy that changes your MCP server also re-synchronises the Gateway's tool list (about 10-20 s), so new tools show up without any extra step. A changed Lambda or tools.json is re-applied by the same deploy.
What you do as a developer
Find your tools through the Gateway
Through the Gateway every tool is named
<target>___<tool>(target name, three underscores, tool name):DataStreamDatabase___query_dbfrom your MCP runtime andWeather___get_us_forecastfrom the Lambda. If you addedlist_tablesin Task 1 it appears asDataStreamDatabase___list_tables. The Gateway also lists one tool of its own,x_amz_bedrock_agentcore_search(more on that below).terminaluv run bootcamp.py test --only 2Read the code: a guard that survives the rename
The agent's read-only guard (you deploy it in Task 4) must recognise the database tool whether it is called directly or through the Gateway:
phase2/app/agent/agent.pyclass ReadOnlyGuardHook(HookProvider): """Blocks destructive SQL. Replaces Phase 1's interactive approval: a runtime has no human to ask. shared/sql_guard.py decides: only SELECT / EXPLAIN / read-only WITH queries pass. """ def register_hooks(self, registry: HookRegistry, **kwargs) -> None: """Run `guard` before every tool call.""" registry.add_callback(BeforeToolCallEvent, self.guard) def guard(self, event: BeforeToolCallEvent) -> None: """Cancel any `query_db` call (direct or via the gateway target) that may write.""" if not event.tool_use.get("name", "").endswith("query_db"): return query = str(event.tool_use.get("input", {}).get("query", "")) if sql_guard.is_write(query): event.cancel_tool = "Destructive SQL requires human approval and is blocked in production."Question: Why
endswith("query_db")and not== "query_db"?Answer
Behind the Gateway the tool is called
DataStreamDatabase___query_db. An equality check would never match, and the guard would silently let every query through. A suffix check works in both places.Describe a table through the Gateway
Challenge
The Gateway is the single front door to your tools: add a tool on the MCP server and the agent discovers it through the Gateway without any agent change. Add
describe_table(table_name), which returns the columns (name, type) of one table, deploy, and confirm the Gateway exposes it asDataStreamDatabase___describe_table.Hint 1
SQLite's
pragma_table_info('<table>')table-valued function returns one row per column (name,type, ...).Hint 2
Never paste
table_nameinto the SQL string. Pass it as a parameter:execute("... (?)", (table_name,)).deployrefreshes the Gateway's tool list for you; the Gateway caches tools per target.Solution
phase2/app/mcp_server/mcp_server.py (add below query_db)@mcp.tool() def describe_table(table_name: str) -> str: """List the columns (name, type) of one table. Call it before writing SQL against that table.""" with lock: return execute("SELECT name, type FROM pragma_table_info(?)", (table_name,))Part B: Gateway target types
A Gateway target is anything the Gateway can turn into MCP tools. Your agent never sees the difference: every target's tools arrive in the same
tools/list, called the same way.Target type Wraps Tool schema from Outbound auth Here MCP server (AgentCore Runtime or any HTTPS MCP) An MCP server you run Listed by the server; snapshot at create/sync OAuth2 (credential provider) or none DataStreamDatabase(Task 1 runtime)Lambda A function you write; the event is the tool's arguments Inline JSON schema (or a file in S3) Gateway IAM role ( lambda:InvokeFunction)Weather(this task)OpenAPI An existing REST API, one tool per operation OpenAPI 3 spec (inline or S3) API key or OAuth2 credential provider no: see “Why Lambda?” Smithy An AWS-style API described by a Smithy model Smithy model (inline or S3) Gateway IAM role (SigV4) no API Gateway REST API A stage of an existing API Gateway API Filtered/overridden operations IAM, API key or none no Connectors and integrations Managed tools (e.g. web search) and templates for SaaS APIs Provided Per connector no The weather tool used to run inside the agent (
tools=[nws_weather.get_us_forecast]). Now it is a Lambda target: the sameshared/nws_weather.pycode, but deployed once, governed by the Gateway (auth, Cedar policy in Task 7, traces) and reusable by any agent that can reach the Gateway.Why Lambda and not an OpenAPI target for api.weather.gov?
An OpenAPI target would expose the raw NWS operations. A forecast takes two calls (
/points/{lat},{lon}, then the forecast URL it returns), the NWS requires an identifyingUser-Agent, and each raw GeoJSON answer costs the model 8K-75K input tokens. OpenAPI targets also authenticate outbound with an API-key or OAuth credential provider, and the NWS has neither. A small Lambda does both calls and returns about 100 tokens of text. Use OpenAPI targets for APIs whose operations already are the tools you want.Read the code: a Lambda target
The Gateway invokes the function with the tool's arguments as the event, and passes the tool it was called as in the Lambda context:
phase2/app/weather_lambda/handler.pyTOOL_NAME_KEY = "bedrockAgentCoreToolName" def get_us_forecast(arguments: dict[str, Any]) -> str: """The next few NWS forecast periods for a US latitude/longitude (compact text, US-only reply elsewhere).""" return nws_weather.get_us_forecast(float(arguments["latitude"]), float(arguments["longitude"])) TOOLS: dict[str, Callable[[dict[str, Any]], str]] = {"get_us_forecast": get_us_forecast} """Tool name (as in tools.json) -> implementation.""" def tool_name(context: Any) -> str: """The bare tool name the Gateway called (`Weather___get_us_forecast` -> `get_us_forecast`), or "".""" custom = getattr(getattr(context, "client_context", None), "custom", None) or {} return str(custom.get(TOOL_NAME_KEY, "")).split(TARGET_SEPARATOR)[-1] def mcp_text(text: str, is_error: bool = False) -> dict[str, Any]: """An MCP tool result. The Gateway passes `content` through as is; a bare string would reach the model JSON-quoted, with every newline escaped.""" return {"content": [{"type": "text", "text": text}], "isError": is_error} def handler(event: dict[str, Any], context: Any) -> dict[str, Any]: """Lambda entry point: run the requested tool; problems come back as an `Error: ...` text the model can read. Args: event: The tool's arguments, as the model sent them through the Gateway. context: The Lambda context; carries the Gateway's tool name. Returns: The tool's answer as an MCP tool result. """ name = tool_name(context) tool = TOOLS.get(name) if tool is None: return mcp_text(f"Error: unknown tool {name!r}; this target serves {sorted(TOOLS)}.", is_error=True) try: return mcp_text(tool(event)) except (KeyError, TypeError, ValueError) as error: return mcp_text(f"Error: bad arguments for {name} ({type(error).__name__}: {error}).", is_error=True)The tool schemas the Gateway lists (and indexes for search) come from one file, which Terraform turns into the target's inline schema and, in Task 7, into the Cedar permit for these tools:
phase2/app/weather_lambda/tools.json[ { "name": "get_us_forecast", "description": "Get the National Weather Service forecast (next few periods: temperature, wind, conditions) for a location in the United States. US only.", "inputSchema": { "type": "object", "properties": { "latitude": { "type": "number", "description": "Latitude of the US location in decimal degrees, e.g. 47.6062 for Seattle." }, "longitude": { "type": "number", "description": "Longitude of the US location in decimal degrees, e.g. -122.3321 for Seattle." } }, "required": ["latitude", "longitude"] } } ]Question: Why does
handlerreturn{"content": [...]}instead of the forecast string?Answer
The Gateway JSON-encodes whatever the function returns. A bare string reaches the model as
"NWS forecast for Seattle, WA:\nToday: ...", quoted and with every newline escaped. An MCP tool result (contentwith a text block) is passed through as is, so the model reads the same clean lines as in Phase 1.Semantic tool search: find, don't load
Every tool schema an agent is given is sent to the model on every call. With two targets that's cheap; a Gateway in front of a company's APIs can hold hundreds of tools, and loading them all costs tokens on every turn and makes the model pick the wrong tool more often. With
search_type = "SEMANTIC"the Gateway addsx_amz_bedrock_agentcore_search: give it a description of what you need and it returns the matching tools, best match first. Your specialists use it instead oflist_tools_sync():phase2/app/agent/agent.pySEARCH_TOOL = "x_amz_bedrock_agentcore_search" """The Gateway's built-in semantic tool search (`search_type = "SEMANTIC"`, terraform/participant/task2_gateway.tf).""" DATA_TOOLS = ("Run read-only SQL on the DataStream company database: employees, departments, projects, tables", 3) WEATHER_TOOLS = ("Weather forecast for a location in the United States", 1) """(what the specialist asks the Gateway's tool search for, how many of the best-ranked tools it keeps). Search ranks every tool, best match first, and has no cut-off: the specialist decides how many it needs (query_db plus the list_tables/describe_table challenge tools; one weather tool).""" def search_tools(gateway: MCPClient, need: tuple[str, int]) -> list[MCPAgentTool]: """The best-ranked Gateway tools for `need` (query, how many), found by the Gateway's semantic search. Only these tools' schemas go into the specialist's context, not every tool behind the Gateway: each schema costs input tokens on every model call, and a long tool list makes the model pick the wrong tool more often. """ query, top_k = need result = gateway.call_tool_sync(f"search-{uuid.uuid4().hex[:12]}", SEARCH_TOOL, {"query": query}) if result.get("status") != "success": raise RuntimeError(f"Gateway tool search failed: {result.get('content')}") found = json.loads("".join(part.get("text", "") for part in result.get("content", []))).get("tools", []) ranked = [spec for spec in found if spec.get("name") != SEARCH_TOOL][:top_k] return [MCPAgentTool(McpTool.model_validate(spec), gateway) for spec in ranked]Measure it:
toolslists your Gateway's tools per target, runs one search and sends the same question to your model three times (no tools, every tool, the search result), printing the prompt tokens LiteLLM billed for each.terminaluv run bootcamp.py tools uv run bootcamp.py tools --search "employee headcount by department"example outputread only$ uv run bootcamp.py tools Gateway tools (3, plus the built-in x_amz_bedrock_agentcore_search): DataStreamDatabase: query_db Weather: get_us_alerts, get_us_forecast Semantic search "weather forecast for a US city" -> Weather___get_us_forecast Prompt tokens for "What will the weather be like in Seattle tomorrow?" on gpt-6-luna: no tools 16 every tool (3) 212 (+196, ~65 per tool) search, top 1 123 (+107) At 200 tools, loading every tool would add ~13,067 tokens to every model call; with search the agent pays only for the few tools that match. $ uv run bootcamp.py tools --search "employee headcount by department" Semantic search "employee headcount by department" -> Weather___get_us_alerts ...Question: Why does each specialist keep only the top few results?
Answer
Search ranks every tool and has no relevance cut-off, so “weather forecast” still returns
query_db, just last. The specialist decides how many it needs: the data specialist keeps three (room forlist_tablesanddescribe_table), the weather specialist one.Add a Lambda tool: weather alerts
Challenge
Alice travels a lot: add a second tool to the
Weathertarget,get_us_alerts(area), that returns the active NWS alerts (warnings, watches, advisories) for a US state such asCA. Deploy, and confirm the Gateway exposesWeather___get_us_alerts. No Terraform and no IAM change: the target, its schema and (in Task 7) its Cedar permit all followtools.json.Hint 1
Two files: a function in
TOOLSinphase2/app/weather_lambda/handler.py, and its schema intools.json(samename). The NWS endpoint is/alerts/active?area=CA; itsfeatures[].properties.headlineis one line per alert.Hint 2
Reuse
nws_weather.make_client()(it sets the User-Agent the NWS requires) andnws_weather.fetch_json. Return a few short lines, not the GeoJSON: every character is model input. In Task 4, let the weather specialist keep two search results so it gets the new tool too.Solution
phase2/app/weather_lambda/handler.py (add below get_us_forecast, replace TOOLS)def get_us_alerts(arguments: dict[str, Any]) -> str: """Active NWS alerts (warnings, watches, advisories) for a US state, e.g. "CA".""" area = str(arguments["area"]).upper() with nws_weather.make_client() as client: alerts = nws_weather.fetch_json(client, f"/alerts/active?area={area}")["features"] if not alerts: return f"No active NWS alerts for {area}." lines = [alert["properties"].get("headline") or alert["properties"]["event"] for alert in alerts[:5]] return "\n".join([f"{len(alerts)} active NWS alerts for {area}:", *lines]) TOOLS: dict[str, Callable[[dict[str, Any]], str]] = { "get_us_forecast": get_us_forecast, "get_us_alerts": get_us_alerts, }phase2/app/weather_lambda/tools.json (second entry of the list){ "name": "get_us_alerts", "description": "Active National Weather Service alerts (warnings, watches, advisories) for a US state.", "inputSchema": { "type": "object", "properties": { "area": {"type": "string", "description": "Two-letter US state code, e.g. CA."} }, "required": ["area"] } }after up 4: phase2/app/agent/agent.pyWEATHER_TOOLS = ("Weather forecast for a location in the United States", 2) """(what the specialist asks the Gateway's tool search for, how many of the best-ranked tools it keeps). Search ranks every tool, best match first, and has no cut-off: the specialist decides how many it needs (query_db plus the list_tables/describe_table challenge tools; one weather tool)."""expected [CHALLENGE] lineread only[CHALLENGE] SKIP stage 2 get_us_alerts Lambda tool: no Weather___get_us_alerts on your Gateway yet # after your change: [CHALLENGE] PASS stage 2 get_us_alerts Lambda tool: Weather___get_us_alerts(CA) -> 28 active NWS alerts for CA: Red Flag Warning issued October 8 at 9:11AM PDT until October 9 at 7:00PM PDT...Experiments
- Run
tools --searchwith a few needs (“employee headcount by department”, “who works on which project?”, “is it going to rain in Denver?”). How often is the right tool first? Sharpen a vague docstring ortools.jsondescription, deploy, and search again: descriptions are what search indexes. - Multiply the per-tool cost
toolsprinted by 200 tools and by the number of model calls in oneinvoke(Task 5 shows them). What does loading every tool cost per question? - Inbound auth (agent to Gateway) and outbound auth (Gateway to target) are separate. Which outbound mode does each of your two targets use, and why is that split useful for third-party APIs?
What should I expect?
Search matches the need against each tool's description, not its data.
query_db's docstring says “SQL query on the DataStream Corp database” but never “employees” or “departments”, so in the dry run “employee headcount by department” rankedWeather___get_us_alertsfirst. Name the tables in the docstring,deploy, and it comes first. Vague needs (“help me”) rank by accident. The per-tool cost is roughly the length of its description and schema; at a few hundred tools that's tens of thousands of input tokens on every model call, before the question is even read. The database target authenticates with OAuth (a Cognito client-credentials token from the credential provider); the Lambda target with the Gateway's own IAM role. The agent only ever holds a token for the Gateway.- Run
Deploy and re-test
terminaluv run bootcamp.py deploy uv run bootcamp.py test --only 2expected [CHALLENGE] lineread only[CHALLENGE] SKIP stage 2 describe_table via the Gateway: no DataStreamDatabase___describe_table on your Gateway yet # after your change: [CHALLENGE] PASS stage 2 describe_table via the Gateway: DataStreamDatabase___describe_table(employees) -> 14 rows, updated_at column found=True
Check your work
uv run bootcamp.py test --only 2Passes when the Gateway lists both targets' tools and a query through DataStreamDatabase___query_db works, Weather___get_us_forecast answers with an NWS forecast through the Lambda, and semantic search ranks the weather tool first for a weather need. Then the [CHALLENGE] lines for describe_table and get_us_alerts.
Under the hood
AgentCore Gateway is a managed MCP endpoint that aggregates targets (MCP servers, Lambda functions, OpenAPI or Smithy APIs, API Gateway stages, connectors) into one tool catalogue. It validates inbound JWTs, then authenticates each outbound call the way its target needs: an OAuth2 credential provider from AgentCore Identity fetches a client-credentials token for your MCP runtime, and the Gateway's IAM role invokes your Lambda. Semantic search adds a built-in search tool backed by embeddings of the tool descriptions, which is another reason docstrings matter.
An MCP target snapshots the server's tool list when it is created or synchronised. The platform re-synchronises it after every MCP redeploy (python -m bootcamp_cli.agentcore sync-gateway), which is why a new tool appears after deploy. A Lambda target has no server to ask: its tool list is the inline schema (tools.json), updated by the same deploy.