Optional add-on

Optional: Code Interpreter

Give the agent a sandboxed Python environment so it computes answers instead of guessing at arithmetic.

25 minAdvanced
Alice’s ask

The “Numbers, Not Prose” request

Alice wants more than counts: “What's the average and median team size per department?” SQL gets the rows; models are unreliable at arithmetic. AgentCore Code Interpreter runs model-written Python in an isolated sandbox, so the agent can compute without touching your runtime.

Platform

What the platform provisions for you

.env
ENABLE_CODE_INTERPRETER=true
terminal
uv run bootcamp.py deploy
  • Session permissions for your agent role on the AWS-managed aws.codeinterpreter.v1
  • The run_python tool, handed to your orchestrator automatically

Flags default to false and need stage 4 or higher. uv run bootcamp.py status shows CODE_INTERPRETER_ENABLED and BROWSER_ENABLED.

You

What you do as a developer

  1. Read the tool

    phase2/app/agent/sandbox_tools.py
    @tool
    def run_python(code: str) -> str:
        """Execute Python code in a secure AgentCore Code Interpreter sandbox and return its output.
    
        Use it for calculations, statistics and data transformations. print() every value you need to see.
    
        Args:
            code: Python source code to run.
        """
        with code_session(REGION, identifier=CODE_INTERPRETER_ID) as client:
            return stream_text(client.execute_code(code))[:OUTPUT_LIMIT]

    One call, one fresh sandbox session: code_session from bedrock_agentcore.tools starts it and stops it when the block exits.

    terminal
    uv run bootcamp.py invoke "Use the code interpreter to compute the mean and median of 3, 5, 10 and 22." --actor alice-chen
  2. Add a data-analysis specialist

    Challenge

    The orchestrator can call run_python, but it has no data. Add an analysis_agent specialist that fetches rows through the Gateway and then computes with run_python, and register it with the orchestrator.

    Hint 1

    It's the agents-as-tools pattern again. Look at data_agent in build_specialists: a @tool that builds an Agent. AgentCore Code Interpreter docs.

    Hint 2

    Import run_python from sandbox_tools and give the specialist tools=[*gateway_tools, run_python]. Tell it: query first, then compute, print every result, report numbers rather than code. Return it from build_specialists.

    Solution
    phase2/app/agent/agent.py (inside build_specialists)
    from sandbox_tools import run_python
    
    
    @tool
    def analysis_agent(question: str) -> str:
        """Analyse DataStream data with Python: statistics, distributions, trends. Use for anything beyond one SQL answer.
    
        Args:
            question: The analysis to perform.
        """
        agent = Agent(
            model=make_model(),
            system_prompt=(
                "You are a data analyst. First fetch the rows you need with the database tool, "
                "then compute with run_python (print every result). Report numbers, not code."
            ),
            tools=[*gateway_tools, run_python],
            hooks=[ReadOnlyGuardHook()],
            callback_handler=None,
        )
        return consult(agent, question, "analysis")

    Return [data_agent, weather_agent, analysis_agent] and add a routing line to orchestrator_prompt.

  3. Deploy and ask for analysis

    terminal
    uv run bootcamp.py deploy
    uv run bootcamp.py invoke "What is the average and median team size per department?" --actor alice-chen
  4. Experiments

    • Look at the trace (uv run bootcamp.py traces): how many run_python calls did one question take, and how long did each sandbox session last?
    • Each call gets a fresh session. What does that mean for code that builds on a previous call's variables?
    • What could go wrong if model-written code ran inside your runtime instead?

Check your work

With the flag on, test --only 4 adds a code interpreter tool check. It sends exactly this prompt:

terminal · the check's prompt, to try by hand
uv run bootcamp.py invoke 'DataStream analytics needs an exact checksum for the quarterly report: use the run_python tool to run `print(sum(range(12345)) % 65521)` and reply with the exact number it prints.' --actor alice-chen

It passes when the answer contains 57938 (thousands separators such as 57,938 are fine): a number the model can't produce reliably without running the code.

terminal
uv run bootcamp.py test --only 4
expected output (abridged)read only
[PASS] stage 4 ...
[PASS] stage 4 code interpreter tool: The script prints 57938.

The check asks as a fresh actor in a fresh session and retries once on an empty answer. The prompt is framed as DataStream work so a Task 4 “decline off-topic requests” prompt keeps it; if your prompt still refuses it, loosen the scope rule.

Under the hood

Both tools use the AWS-managed built-ins aws.codeinterpreter.v1 and aws.browser.v1. Every call starts a fresh, isolated sandbox session and stops it afterwards, so nothing leaks between participants or requests. When a flag is on, your agent role gets only the session permissions it needs, in the inline policy awsworkshop-<you>-sandbox-tools; the platform sets CODE_INTERPRETER_ENABLED / BROWSER_ENABLED on the runtime, and optional_tools() plus prompt_hints() hand the tools to the orchestrator with a one-line usage hint each.

phase2/app/agent/sandbox_tools.py
def optional_tools() -> list:
    """The sandbox tools this runtime is configured for (empty when both capabilities are off)."""
    tools = []
    if capability_enabled("CODE_INTERPRETER_ENABLED"):
        tools.append(run_python)
    if capability_enabled("BROWSER_ENABLED"):
        tools.append(browse_web)
    return tools